Compositions and methods for detecting predisposition to cardiovascular disease - Patents.com

By providing a kit for detecting specific CpG binucleotides and SNP genotypes, the problem of limited effectiveness of DNA methylation signatures in the prior art in detecting cardiovascular diseases is solved, and more accurate cardiovascular disease identification and prediction are achieved.

JP7672192B2Active Publication Date: 2025-05-07THE UNIVERSITY OF IOWA RESEARCH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2018564383
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2017-02-06
Filing Date
2017-06-08
Publication Date
2025-05-07
Estimated Expiration
2037-06-08

AI Technical Summary

Technical Problem

The prior art has limited effects when applying DNA methylation signatures to detect cardiovascular disease (CVD), possibly due to masking of the interaction effect between genes and methylation.

Method used

A kit containing specific CpG binucleotide and single nucleotide polymorphism (SNP) genotypes is provided to detect the methylation status of CpG binucleotides and the genotype of SNP by aligning bisulfite-transformed DNA sequences.

Benefits of technology

By accurately detecting the methylation status of CpG binucleotides and the genotype of SNP, the recognition and prediction ability of cardiovascular diseases is improved, and the influence of gene-methylation interaction effects is overcome.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007672192000035
    Figure 0007672192000035
  • Figure 0007672192000036
    Figure 0007672192000036
  • Figure 0007672192000037
    Figure 0007672192000037
Patent Text Reader

Abstract

Methods and compositions are provided for detecting a predisposition to cardiovascular disease in an individual.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority under 35 USC § 119(e) of U.S. Application No. 62 / 347,479, filed June 8, 2016, and U.S. Application No. 62 / 455,468, filed February 6, 2017, both of which are incorporated herein in their entireties.

[0002] STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH This invention was made with government support under Nos. R01DA037648 and R44DA041014 awarded by the National Institutes of Health. The United States Government has certain rights in this invention. [Background technology]

[0003] 2. Background of the Invention Cardiovascular disease (CVD), consisting of coronary heart disease (CHD), congestive heart failure (CHF), and stroke, is the leading cause of death in the United States. Although effective treatments exist to reduce the morbidity and mortality of CVD, their clinical implementation is hindered by inefficient screening techniques. Recently, others and the present inventors have shown that DNA methylation signatures can predict the presence of various disorders related to CVD, such as smoking. Unfortunately, when these epigenetic techniques are applied to CVD itself, the power of these methods is reduced, which limits their clinical usefulness. One possible reason for these failures may be the obscuring of epigenetic signatures of CVD by gene x methylation interaction effects.

[0004] A reliable laboratory test would have practical value in clinical practice, for example in assisting physicians in prescribing appropriate treatments for patients. Thus, there is a need for methods to identify subjects who have or are at risk of developing CVD. Summary of the Invention

[0005] In certain embodiments, the present disclosure provides a kit for determining the methylation status of at least one CpG dinucleotide and the genotype of at least one single nucleotide polymorphism (SNP), comprising at least one first nucleic acid primer at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes the CpG dinucleotide of a gene described in Figure 15, or the CpG dinucleotide of a first CpG site described in Figure 16, or the CpG dinucleotide of a second CpG site that is collinear (e.g., R>0.3) with the first CpG site described in Figure 16, wherein the first nucleic acid primer detects the unmethylated CpG dinucleotide; and at least one second nucleic acid primer at least 8 nucleotides in length that is complementary to a DNA sequence or a bisulfite converted DNA sequence of a first SNP described in Figure 21 or a second SNP that is in linkage disequilibrium with the first SNP described in Figure 21. In some embodiments, the linkage disequilibrium has a value of R>0.3.

[0006] In certain embodiments, the present disclosure provides a kit for determining the methylation status of at least one CpG dinucleotide and the genotype of at least one single nucleotide polymorphism (SNP), comprising at least one first nucleic acid primer of at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a gene set forth in Figure 17, a first CpG dinucleotide set forth in Figure 18, or a second CpG dinucleotide that is collinear (e.g., R>0.3) with the first CpG site set forth in Figure 18, where the first nucleic acid primer detects the unmethylated CpG dinucleotide; and at least one second nucleic acid primer of at least 8 nucleotides in length that is complementary to a DNA sequence or a bisulfite converted DNA sequence of a first SNP set forth in Figure 22 or a second SNP that is in linkage disequilibrium with the first SNP set forth in Figure 22. In some embodiments, the linkage disequilibrium has a value of R>0.3.

[0007] In certain embodiments, the present disclosure provides a kit for determining the methylation status of at least one CpG dinucleotide and the genotype of at least one single nucleotide polymorphism (SNP), comprising at least one first nucleic acid primer of at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes the CpG dinucleotide of a gene listed in Figure 19, or the CpG dinucleotide of a first CpG site in Figure 20, or a second CpG dinucleotide that is collinear (R>0.3) with the first CpG site listed in Figure 20, wherein the first nucleic acid primer detects the unmethylated CpG dinucleotide; and at least one second nucleic acid primer of at least 8 nucleotides in length that is complementary to a DNA sequence or a bisulfite converted DNA sequence of a first SNP listed in Figure 23 or a second SNP that is in linkage disequilibrium with the first SNP listed in Figure 23. In some embodiments, the linkage disequilibrium has a value of R>0.3.

[0008] In certain aspects, the disclosure provides a kit for determining the methylation status of at least one CpG dinucleotide and the presence of at least one single nucleotide polymorphism (SNP), the kit comprising at least one first nucleic acid primer at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide at position 92203667 of chromosome 1 in the transforming growth factor beta receptor III (TGFBR3) gene, the first nucleic acid primer detecting the unmethylated CpG dinucleotide; and at least one second nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs347027.

[0009] In certain embodiments, the present disclosure provides a kit for determining the methylation status of at least one CpG dinucleotide and the presence of at least one single nucleotide polymorphism (SNP), comprising: at least one first nucleic acid primer at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide at position 38364951 within an intergenic region of chromosome 15, the first nucleic acid primer detecting the unmethylated CpG dinucleotide; and at least one second nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs4937276.

[0010] In certain aspects, the disclosure provides a kit for determining the methylation status of at least one CpG dinucleotide and the presence of at least one single nucleotide polymorphism (SNP), comprising at least one first nucleic acid primer at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide at position 84206068 of chromosome 4 in the coenzyme Q2 4-hydroxybenzoate polyprenyltransferase (COQ2) gene, the first nucleic acid primer detecting the unmethylated CpG dinucleotide; and at least one second nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs17355663.

[0011] In certain aspects, the disclosure provides a kit for determining the methylation status of at least one CpG dinucleotide and the presence of at least one single nucleotide polymorphism (SNP), the kit comprising at least one first nucleic acid primer at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide at position 26146070 of chromosome 16 in the heparan sulfate 3-O-sulfotransferase 4 (HS3ST4) gene, the first nucleic acid primer detecting the unmethylated CpG dinucleotide; and at least one second nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs235807.

[0012] In certain aspects, the disclosure provides a kit for determining the methylation status of at least one CpG dinucleotide and the presence of at least one single nucleotide polymorphism (SNP), the kit comprising: at least one first nucleic acid primer at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide at position 91171013 in an intergenic region of chromosome 1, the first nucleic acid primer detecting the unmethylated CpG dinucleotide; and at least one second nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs11579814.

[0013] In certain aspects, the disclosure provides a kit for determining the methylation status of at least one CpG dinucleotide and the presence of at least one single nucleotide polymorphism (SNP), comprising at least one first nucleic acid primer at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide at position 39491936 of chromosome 1 in the NADH dehydrogenase (ubiquinone) Fe-S protein 5 (NDUFS5) gene, wherein the first nucleic acid primer detects the unmethylated CpG dinucleotide; and at least one second nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs2275187.

[0014] In certain aspects, the disclosure provides a kit for determining the methylation status of at least one CpG dinucleotide and the presence of at least one single nucleotide polymorphism (SNP), comprising: at least one first nucleic acid primer at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide at position 186426136 mapped to chromosome 1 within the phosducin gene, the first nucleic acid primer detecting the unmethylated CpG dinucleotide; and at least one second nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs4336803.

[0015] In certain aspects, the disclosure provides a kit for determining the methylation status of at least one CpG dinucleotide and the presence of at least one single nucleotide polymorphism (SNP), comprising at least one first nucleic acid primer at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide at position 205475130 of chromosome 1 in the cyclin dependent kinase 18 (CDK18) gene, wherein the first nucleic acid primer detects the unmethylated CpG dinucleotide; and at least one second nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs4951158.

[0016] In certain aspects, the disclosure provides a kit for determining the methylation status of at least one CpG dinucleotide and the presence of at least one single nucleotide polymorphism (SNP), comprising at least one first nucleic acid primer at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide at position 130614013 on chromosome 3 in the ATPase, Ca++ Transporting, Type 2C, Member 1 (ATP2C1) gene, wherein the first nucleic acid primer detects the unmethylated CpG dinucleotide, and at least one second nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs925613.

[0017] In certain aspects, the disclosure provides a kit for determining the methylation status of at least one CpG dinucleotide, comprising at least one first nucleic acid primer of at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide at position 92203667 of chromosome 1 in the transforming growth factor beta receptor III (TGFBR3) gene, the at least one first nucleic acid primer comprising one or more nucleotide analogs or one or more synthetic or non-natural nucleotides, and that detects either the unmethylated or methylated CpG dinucleotide.

[0018] In certain aspects, the disclosure provides a kit for determining the methylation status of at least one CpG dinucleotide, the kit comprising at least one first nucleic acid primer at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide at position 92203667 of chromosome 1 in the transforming growth factor beta receptor III (TGFBR3) gene, the at least one first nucleic acid primer detecting either the unmethylated CpG dinucleotide or the methylated CpG dinucleotide; and a detectable label selected from the group consisting of an enzymatic label, a fluorescent label, and a chromogenic label.

[0019] In certain aspects, the disclosure provides a kit for determining the methylation status of at least one CpG dinucleotide, the kit comprising at least one first nucleic acid primer at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide at position 92203667 of chromosome 1 in the transforming growth factor beta receptor III (TGFBR3) gene, the at least one first nucleic acid primer detecting either the unmethylated CpG dinucleotide or the methylated CpG dinucleotide; and a solid substrate to which the at least one first nucleic acid primer is attached.

[0020] In certain embodiments, the disclosure provides a method for detecting whether a subject is predisposed to or has coronary heart disease, comprising the steps of: (a) providing a biological sample from the subject; (b) contacting DNA from the biological sample with bisulfite under alkaline conditions; (c) contacting the bisulfite-treated DNA with at least one first oligonucleotide probe of at least 8 nucleotides in length that is complementary to a sequence comprising a CpG dinucleotide at position 92203667 of chromosome 1 within transforming growth factor beta receptor III (TGFBR3), comprising at least one oligonucleotide probe having a length of at least 8 nucleotides. (d) determining the genotype at single nucleotide polymorphism rs347027; and (e) detecting either an unmethylated CpG dinucleotide or a methylated CpG dinucleotide, wherein methylation of the CpG dinucleotide at position 92203667 of chromosome 1 is associated with coronary heart disease when rs347027 is genotyped.

[0021] In certain embodiments, the disclosure provides a method for measuring the presence of a biomarker in a biological sample from a patient, comprising the steps of: (a) contacting DNA from the biological sample with bisulfite under alkaline conditions; and (b) contacting the bisulfite-treated DNA with at least one first oligonucleotide probe of at least 8 nucleotides in length that is complementary to a sequence comprising a CpG dinucleotide at position 92203667 of chromosome 1 located within the transforming growth factor beta receptor III (TGFBR3) gene, wherein the at least one first oligonucleotide probe detects either the CpG dinucleotide that is unmethylated or the CpG dinucleotide that is methylated, for use in predicting whether a patient has coronary heart disease or is likely to develop coronary heart disease.

[0022] In certain embodiments, the disclosure provides a method of predicting the presence of a biomarker associated with cardiovascular disease (CVD) in a biological sample from a patient, comprising the steps of: (a) preparing a first aliquot from the biological sample and contacting DNA from the first biological sample with bisulfite under alkaline conditions; and (b) preparing a second aliquot from the biological sample; (c) (i) contacting the first aliquot with a first oligonucleotide probe at least 8 nucleotides in length that is complementary to a sequence comprising a CpG dinucleotide at position 92203667 of chromosome 1 within the transforming growth factor beta receptor III (TGFBR3) gene, and contacting the second aliquot with a nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs347027, (ii) contacting the first aliquot with a first oligonucleotide probe at least 8 nucleotides in length that is complementary to a sequence comprising a CpG dinucleotide at position 38364951 within an intergenic region of chromosome 15, and contacting the second aliquot with a nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs347027. (iii) contacting the first aliquot with a first oligonucleotide probe at least 8 nucleotides in length that is complementary to a sequence comprising a CpG dinucleotide at position 84206068 on chromosome 4 in the coenzyme Q2 4-hydroxybenzoate polyprenyltransferase (COQ2) gene, and contacting the second aliquot with a nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs17355663; (iv) contacting the first aliquot with a first oligonucleotide probe at least 8 nucleotides in length that is complementary to a sequence comprising a CpG dinucleotide at position 26146070 on chromosome 16 in the heparan sulfate 3-O-sulfotransferase 4 (HS3ST4) gene, and contacting the second aliquot with a nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs235807; (v)(vi) contacting the first aliquot with a first oligonucleotide probe at least 8 nucleotides in length that is complementary to a sequence that includes a CpG dinucleotide at position 39491936 of chromosome 1 located within the NADH dehydrogenase (ubiquinone) Fe-S protein 5 (NDUFS5) gene, and contacting the second aliquot with a nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs2275187; and (vii) (viii) contacting the first aliquot with a first oligonucleotide probe at least 8 nucleotides in length that is complementary to a sequence comprising a CpG dinucleotide at position 205475130 of chromosome 1 located in the cyclin-dependent kinase 18 (CDK18) gene and contacting the second aliquot with a nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs4951158; and / or (ix) contacting the first aliquot with a nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs4951158, which ...and contacting the second aliquot with a first oligonucleotide probe of at least 8 nucleotides in length that is complementary to a sequence that includes a CpG dinucleotide at position 130614013 on chromosome 3 within the TGFBR3 gene, cg20636912, cg16947 ... The present invention relates to a method in which methylation of 05916059, cg04567738, cg16603713, cg05709437, cg12081870, and / or cg18070470, and a G at position 91618766 on chromosome 1, or polymorphisms at rs4937276, rs17355663, rs235807, rs11579814, rs2275187, rs4336803, rs4951158, and / or rs925613 are associated with CVD.

[0023] In certain embodiments, the biological sample is a saliva sample.

[0024] In certain aspects, the disclosure provides a method for detecting one or more copies of the G allele at rs347027 and the methylation status of cg13078798 on a nucleic acid sample from a subject at risk for cardiovascular disease (CVD), the method comprising: (a) performing a genotyping assay on the nucleic acid sample of the human subject to detect the presence of one or more copies of the G allele of the rs347027 polymorphism; and (b) performing a methylation assessment of cg13078798 on the nucleic acid sample of the human to detect the methylation status to determine whether cg13078798 is unmethylated.

[0025] In certain embodiments, the present disclosure provides a method of predicting the presence of a biomarker associated with cardiovascular disease (CVD) in a biological sample from a patient, comprising predicting the presence of a biomarker associated with cardiovascular disease (CVD) in a biological sample from a patient, the method comprising predicting the presence of a biomarker associated with cardiovascular disease (CVD) in a biological sample from a patient, the method comprising predicting the presence of a biomarker associated with cardiovascular disease (CVD) in a biological sample from a patient, the method comprising predicting the presence of a biomarker associated with cardiovascular disease (CVD) in a biological sample from a patient, the method comprising predicting the presence of a biomarker associated with cardiovascular disease (CVD) in a biological sample from a patient, the method comprising predicting the presence of a biomarker associated with cardiovascular disease (CVD) in a biological sample from a patient, the method comprising predicting the presence of a biomarker associated with cardiovascular disease (CVD) in a biological sample from a patient, the method comprising predicting the presence of a biomarker associated with cardiovascular disease (CVD) in a biological sample from a patient, the method comprising predicting the presence of a biomarker associated with cardiovascular disease (CVD) in a biological sample from a patient, the method comprising predicting the presence of a biomarker associated with cardiovascular disease (CVD) in a biological sample from a patient, the method comprising predicting the presence of a biomarker associated with cardiovascular disease (CVD) in a biological sample from a patient the combination of CpG cg12081870 and SNP rs4951158; and / or the combination of CpG cg18070470 and SNP rs925613).

[0026] In certain embodiments, the CVD is coronary heart disease (CHD), congestive heart failure (CHF), and / or stroke.

[0027] In certain embodiments, the disclosure provides a method of determining the presence of a biomarker associated with CHD in a patient sample, comprising the steps of: (a) isolating a nucleic acid sample from the patient sample; (b) performing a genotyping assay on a first aliquot of the nucleic acid sample to detect the presence of at least one SNP to obtain genotype data, wherein the SNP is a SNP selected from a first SNP set forth in FIG. 21 and / or a second SNP that is in linkage disequilibrium (e.g., R>0.3) with the first SNP set forth in FIG. 21 ; and / or (c) bisulfite converting the nucleic acid in the second aliquot of the nucleic acid to determine the methylation status of at least one gene set forth in FIG. 15 or a second SNP set forth in FIG. performing a methylation assessment on a second aliquot of the nucleic acid sample to detect the methylation status of the first CpG site and / or the methylation status of a second CpG site that is collinear (e.g., R>0.3) with the first CpG as described in Figure 16, and obtaining methylation data regarding whether a particular CpG residue is unmethylated; and (d) inputting the genotype from step (b) and / or the methylation data from step (c) into an algorithm that reveals the contribution of at least one main effect of SNP and / or at least one main effect of CpG and / or at least one interaction effect (e.g., SNP x SNP, CpG x CpG, SNP x CpG). In some embodiments, the algorithm is Random Forest™ or another algorithm that can account for linear and nonlinear effects.

[0028] In certain embodiments, the disclosure provides a method of determining the presence of a biomarker associated with stroke in a patient sample, comprising the steps of: (a) isolating a nucleic acid sample from the patient sample; (b) performing a genotyping assay on a first aliquot of the nucleic acid sample to detect the presence of at least one SNP to obtain genotype data, wherein the SNP is a SNP selected from a first SNP set forth in FIG. 22 and / or a second SNP that is in linkage disequilibrium (e.g., R>0.3) with the first SNP set forth in FIG. 22; and / or (c) bisulfite converting the nucleic acid in a second aliquot of the nucleic acid to determine the methylation status of at least one gene set forth in FIG. 17 or a second SNP set forth in FIG. and / or a second CpG site that is collinear (e.g., R>0.3) with the first CpG as set forth in FIG. 18, to obtain methylation data regarding whether a particular CpG residue is unmethylated; and (d) inputting the genotype from step (b) and / or the methylation data from step (c) into an algorithm that accounts for the contribution of at least one SNP main effect and / or at least one CpG main effect and / or at least one interaction effect (e.g., SNP x SNP, CpG x CpG, SNP x CpG). In some embodiments, the algorithm is Random Forest™ or another algorithm that can account for linear and nonlinear effects.

[0029] In certain embodiments, the disclosure provides a method of determining the presence of a biomarker associated with CHF in a patient sample, comprising the steps of: (a) isolating a nucleic acid sample from the patient sample; (b) performing a genotyping assay on a first aliquot of the nucleic acid sample to detect the presence of at least one SNP to obtain genotype data, wherein the SNP is a SNP selected from a first SNP in FIG. 23 and / or a second SNP that is in linkage disequilibrium (e.g., R>0.3) with the first SNP set forth in FIG. 23; and / or (c) bisulfite converting the nucleic acid in the second aliquot of the nucleic acid to determine the methylation status of at least one gene set forth in FIG. 19 or a second SNP set forth in FIG. performing a methylation assessment on a second aliquot of the nucleic acid sample to detect the methylation status of a CpG site and / or a second CpG site that is collinear (e.g., R>0.3) with the first CpG as described in Figure 20, to obtain methylation data regarding whether a particular CpG residue is unmethylated; and (d) inputting the genotype from step (b) and / or the methylation data from step (c) into an algorithm that reveals the contribution of at least one SNP main effect and / or at least one CpG main effect and / or at least one interaction effect (e.g., SNP x SNP, CpG x CpG, SNP x CpG). In some embodiments, the algorithm is Random Forest™ or another algorithm that can account for linear and nonlinear effects.

[0030] In certain embodiments, the results include a gene-environment interaction effect (SNP×CpG) between a second CpG site that is collinear (e.g., R>0.3) with the first CpG site listed in FIG. 16 and a first SNP listed in FIG. 21 or a second SNP that is in linkage disequilibrium (e.g., R>0.3) with the first SNP listed in FIG. 21. In certain embodiments, the results include at least one environment-environment interaction effect (CpG×CpG) between at least two CpG sites listed in FIG. 16 and / or at least two genes listed in FIG. 15. In certain embodiments, the results include at least one environment-environment interaction effect (CpG×CpG) between at least two CpG sites that are collinear with the first CpG site listed in FIG. 16. In certain embodiments, the results include a gene-environment interaction effect (SNP×CpG) between a CpG site that is collinear (e.g., R>0.3) with a first CpG site listed in FIG. 18 and a first SNP listed in FIG. 22 or a second SNP that is in linkage disequilibrium (e.g., R>0.3) with the first SNP listed in FIG. 22. In certain embodiments, the results include at least one environment-environment interaction effect (CpG×CpG) between at least two CpG sites listed in FIG. 18 and / or a gene listed in FIG. 17. In certain embodiments, the results include at least one environment-environment interaction effect (CpG×CpG) between at least two CpG sites that are collinear with a first CpG site listed in FIG. 18. In certain embodiments, the results include a gene-environment interaction effect (SNP×CpG) between a second CpG site that is collinear (e.g., R>0.3) with the first CpG site listed in FIG. 20 and a first SNP listed in FIG. 23 or a second SNP that is in linkage disequilibrium (e.g., R>0.3) with the first SNP listed in FIG. 23. In certain embodiments, the results include at least one environment-environment interaction effect (CpG×CpG) between at least two CpG sites listed in FIG. 20 and / or a gene listed in FIG. 19. In certain embodiments, the results include at least one environment-environment interaction effect (CpG×CpG) between at least two CpG sites that are collinear with the first CpG site listed in FIG. 20.

[0031] In certain embodiments of the present disclosure, the blood cells are lymphocytes, such as monocytes, basophils, eosinophils, and / or neutrophils. In certain embodiments, the type of lymphocyte is B-lymphocyte. In certain embodiments, the B-lymphocyte is immortalized. In certain embodiments, the type of blood cell is a mixture of peripheral white blood cells. In certain embodiments, the peripheral blood cells are transformed into cell lines.

[0032] In certain embodiments, the analysis process comprises comparing the obtained profile with a reference profile.In certain embodiments, the reference profile comprises data obtained from one or more healthy control subjects, or comprises data obtained from one or more subjects diagnosed with substance use disorder.In certain embodiments, the method further comprises obtaining a statistical measure of the similarity of the obtained profile with the reference profile.In certain embodiments, the blood cell or blood cell derivative is a peripheral blood cell.In certain embodiments, the profile is obtained by sequencing methylated DNA, for example by digital sequencing.

[0033] In certain embodiments, the present disclosure may also take the form of a PCR (polymerase chain reaction) assay. In some cases, this takes the form of a real-time PCR assay (RTPCR) or digital PCR assay. In certain embodiments of these PCR assays, the kit may contain two primers that specifically amplify a region of a target gene and a gene-specific probe that selectively recognizes the amplified region. Collectively, the primers and gene-specific probe are referred to as a primer-probe set. By measuring the amount of gene-specific probe hybridized to the amplified segment at a given time point in the PCR reaction or throughout the PCR reaction, one of skill in the art can infer the amount of nucleic acid originally present at the start of the reaction. In some cases, the amount of hybridized probe is measured through fluorescence spectrophotometry. The number of primer-probe sets can be any integer number between 1 and 10,000 probes, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, ..., 9997, 9998, 9999, 10,000, etc. In one kit, all the probes can be physically located in a single reaction well or in multiple reaction wells. The probes can be in dry or liquid form. They can be used in a single reaction or in a series of reactions. In certain embodiments, the probes are oligonucleotide probes. In certain embodiments, the probes are nucleic acid derivative probes.

[0034] Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the subject method and composition belongs.Methods and materials similar or equivalent to those described herein can be used to carry out or test the subject method and composition, and suitable methods and materials are described below.In addition, the materials, methods, and examples are only illustrative and are not intended to be limiting.All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety. [The present invention 1001] A kit for determining the methylation status of at least one CpG dinucleotide and the genotype of at least one single nucleotide polymorphism (SNP), comprising: at least one first nucleic acid primer at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide of a gene listed in FIG. 15, a CpG dinucleotide of a CpG site listed in FIG. 16, or a CpG dinucleotide that is collinear (R>0.3) with a CpG site listed in FIG. 16, wherein the first nucleic acid primer detects the unmethylated CpG dinucleotide; and At least one second nucleic acid primer at least 8 nucleotides in length that is complementary to the DNA sequence or bisulfite converted DNA sequence of the first SNP set forth in FIG. 21 or a second SNP that is in linkage disequilibrium with the first SNP set forth in FIG. 21, wherein the linkage disequilibrium has a value of R>0.3. The kit comprising: [The present invention 1002] A kit for determining the methylation status of at least one CpG dinucleotide and the genotype of at least one single nucleotide polymorphism (SNP), comprising: at least one first nucleic acid primer at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a gene as set forth in FIG. 17, a CpG dinucleotide at a CpG site as set forth in FIG. 18, or a CpG dinucleotide that is collinear (R>0.3) with a CpG site as set forth in FIG. 18, wherein the first nucleic acid primer detects the unmethylated CpG dinucleotide; and At least one second nucleic acid primer at least 8 nucleotides in length that is complementary to the DNA sequence or bisulfite converted DNA sequence of the first SNP set forth in FIG. 22 or a second SNP that is in linkage disequilibrium with the first SNP set forth in FIG. 22, wherein the linkage disequilibrium has a value of R>0.3. The kit comprising: [The present invention 1003] A kit for determining the methylation status of at least one CpG dinucleotide and the genotype of at least one single nucleotide polymorphism (SNP), comprising: at least one first nucleic acid primer at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide of a gene listed in FIG. 19, a CpG dinucleotide of a CpG site listed in FIG. 20, or a CpG dinucleotide that is collinear (R>0.3) with a CpG site listed in FIG. 20, wherein the first nucleic acid primer detects the unmethylated CpG dinucleotide; and At least one second nucleic acid primer at least 8 nucleotides in length that is complementary to the DNA sequence or bisulfite converted DNA sequence of the first SNP set forth in FIG. 23 or a second SNP that is in linkage disequilibrium with the first SNP set forth in FIG. 23, wherein the linkage disequilibrium has a value of R>0.3. The kit comprising: [The present invention 1004] 1. A kit for determining the methylation status of at least one CpG dinucleotide and the presence of at least one single nucleotide polymorphism (SNP), comprising: at least one first nucleic acid primer at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide at position 92203667 of chromosome 1 located in the transforming growth factor beta receptor III (TGFBR3) gene, wherein the at least one first nucleic acid primer detects the unmethylated CpG dinucleotide; and At least one second nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs347027. The kit comprising: [The present invention 1005] The kit of the present invention 1004, wherein rs347027 includes a G allele. [The present invention 1006] The kit of the present invention 1004 further comprises at least one third nucleic acid primer of at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide at position 92203667 of chromosome 1 in the TGFBR gene, wherein said at least one second nucleic acid primer detects the CpG dinucleotide that is methylated. [The present invention 1007] 1. A kit for determining the methylation status of at least one CpG dinucleotide and the presence of at least one single nucleotide polymorphism (SNP), comprising: at least one first nucleic acid primer at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide at position 38,364,951 within an intergenic region of chromosome 15, wherein the at least one first nucleic acid primer detects the unmethylated CpG dinucleotide; and At least one second nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs4937276. The kit comprising: [The present invention 1008] The kit of the present invention 1007 further comprises at least one third nucleic acid primer of at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide at position 38364951 within an intergenic region of chromosome 15, wherein the at least one second nucleic acid primer detects the CpG dinucleotide that is methylated. [The present invention 1009] 1. A kit for determining the methylation status of at least one CpG dinucleotide and the presence of at least one single nucleotide polymorphism (SNP), comprising: at least one first nucleic acid primer at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide at position 84206068 of chromosome 4 in the coenzyme Q2 4-hydroxybenzoate polyprenyltransferase (COQ2) gene, the at least one first nucleic acid primer detecting the unmethylated CpG dinucleotide; and At least one second nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs17355663. The kit comprising: [The present invention 1010] The kit of the present invention 1009 further comprises at least one third nucleic acid primer of at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence containing a CpG dinucleotide at position 84206068 of chromosome 4 in the coenzyme Q2 4-hydroxybenzoate polyprenyltransferase (COQ2) gene, wherein the at least one second nucleic acid primer detects the CpG dinucleotide that is methylated. [The present invention 1011] 1. A kit for determining the methylation status of at least one CpG dinucleotide and the presence of at least one single nucleotide polymorphism (SNP), comprising: at least one first nucleic acid primer at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide at position 26146070 of chromosome 16 located in the heparan sulfate 3-O-sulfotransferase 4 (HS3ST4) gene, wherein the at least one first nucleic acid primer detects the unmethylated CpG dinucleotide; and At least one second nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs235807. The kit comprising: [The present invention 1012] The kit of the present invention further comprises at least one third nucleic acid primer of at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide at position 26146070 of chromosome 16 in the heparan sulfate 3-O-sulfotransferase 4 (HS3ST4) gene, wherein the at least one second nucleic acid primer detects the CpG dinucleotide that is methylated. [The present invention 1013] 1. A kit for determining the methylation status of at least one CpG dinucleotide and the presence of at least one single nucleotide polymorphism (SNP), comprising: at least one first nucleic acid primer at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide at position 91171013 in an intergenic region of chromosome 1, wherein the first nucleic acid primer detects the unmethylated CpG dinucleotide; and At least one second nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs11579814. The kit comprising: [The present invention 1014] The kit of the present invention further comprises at least one third nucleic acid primer of at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence comprising a CpG dinucleotide at position 91171013 in an intergenic region of chromosome 1, wherein said at least one second nucleic acid primer detects the CpG dinucleotide in a methylated state. [The present invention 1015] 1. A kit for determining the methylation status of at least one CpG dinucleotide and the presence of at least one single nucleotide polymorphism (SNP), comprising: at least one first nucleic acid primer at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide at position 39491936 of chromosome 1 located in the NADH dehydrogenase (ubiquinone) Fe-S protein 5 (NDUFS5) gene, wherein the at least one first nucleic acid primer detects the unmethylated CpG dinucleotide; and At least one second nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs2275187. The kit comprising: [The present invention 1016] The kit of the present invention 1015 further comprises at least one third nucleic acid primer of at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide at position 39491936 of chromosome 1 in the NADH dehydrogenase (ubiquinone) Fe-S protein 5 (NDUFS5) gene, wherein the at least one second nucleic acid primer detects the CpG dinucleotide that is methylated. [The present invention 1017] 1. A kit for determining the methylation status of at least one CpG dinucleotide and the presence of at least one single nucleotide polymorphism (SNP), comprising: at least one first nucleic acid primer at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide at position 186426136 mapped to chromosome 1 within the phosducin gene, wherein the at least one first nucleic acid primer detects the unmethylated CpG dinucleotide; and At least one second nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs4336803. The kit comprising: [The present invention 1018] The kit of the present invention 1017 further comprises at least one third nucleic acid primer of at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide at position 186426136 mapped to chromosome 1 within the phosducin gene, wherein the at least one second nucleic acid primer detects the CpG dinucleotide that is methylated. [The present invention 1019] 1. A kit for determining the methylation status of at least one CpG dinucleotide and the presence of at least one single nucleotide polymorphism (SNP), comprising: at least one first nucleic acid primer at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide at position 205475130 of chromosome 1 located in the cyclin-dependent kinase 18 (CDK18) gene, wherein the first nucleic acid primer detects the unmethylated CpG dinucleotide; and At least one second nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs4951158. The kit comprising: [The present invention 1020] The kit of the present invention further comprises at least one third nucleic acid primer of at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide at position 205475130 of chromosome 1 in the cyclin-dependent kinase 18 (CDK18) gene, wherein the at least one second nucleic acid primer detects the CpG dinucleotide that is methylated. [The present invention 1021] 1. A kit for determining the methylation status of at least one CpG dinucleotide and the presence of at least one single nucleotide polymorphism (SNP), comprising: at least one first nucleic acid primer at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide at position 130614013 on chromosome 3 within the ATPase, Ca++ Transporting, Type 2C, Member 1 (ATP2C1) gene, wherein the first nucleic acid primer detects the unmethylated CpG dinucleotide; and At least one second nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs925613. The kit comprising: [The present invention 1022] The kit of the present invention further comprises at least one third nucleic acid primer of at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide at position 130614013 of chromosome 3 in the ATPase, Ca++ Transporting, Type 2C, Member 1 (ATP2C1) gene, wherein the at least one second nucleic acid primer detects the CpG dinucleotide that is methylated. [The present invention 1023] The kit of any of claims 1001 to 1022, wherein the at least one first primer is at least 10 nucleotides in length, and the at least one second primer is at least 10 nucleotides in length. [The present invention 1024] The kit of any of claims 1001 to 1022, wherein the at least one first primer is at least 12 nucleotides in length, and the at least one second primer is at least 12 nucleotides in length. [The present invention 1025] The kit of any one of claims 1001 to 1024, wherein the at least one first nucleic acid primer comprises one or more nucleotide analogues. [The present invention 1026] The kit of any of claims 1001 to 1024, wherein the at least one first nucleic acid primer comprises one or more synthetic or non-natural nucleotides. [The present invention 1027] The kit of any of claims 1001 to 1026, further comprising a solid substrate to which said at least one first nucleic acid primer is attached. [The present invention 1028] The kit of claim 1027, wherein the substrate is a polymer, glass, semiconductor, paper, metal, gel, or hydrogel. [The present invention 1029] The kit of claim 1027, wherein the solid substrate is a microarray or a microfluidic card. [The present invention 1030] The kit of any of claims 1001 to 1029, further comprising a detectable label. [The present invention 1031] A kit for determining the methylation status of at least one CpG dinucleotide, comprising: At least one first nucleic acid primer at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide at position 92203667 of chromosome 1 located in the transforming growth factor beta receptor III (TGFBR3) gene, the at least one first nucleic acid primer including one or more nucleotide analogs or one or more synthetic or non-natural nucleotides, and that detects either the unmethylated CpG dinucleotide or the methylated CpG dinucleotide. The kit comprising: [The present invention 1032] The kit of the present invention 1031, further comprising at least one second nucleic acid primer of at least 8 nucleotides in length complementary to a bisulfite converted nucleic acid sequence comprising a CpG dinucleotide at position 92203667 of chromosome 1 located in the TGFBR gene, wherein the at least one second nucleic acid primer detects the unmethylated CpG dinucleotide or the methylated CpG dinucleotide, whereas the opposite CpG dinucleotide is detected by the at least one first nucleic acid primer. [The present invention 1033] The kit of claim 1031 or 1032, wherein said at least one first nucleic acid primer detects said CpG dinucleotide that is unmethylated. [The present invention 1034] The kit of claim 1031 or 1032, wherein said at least one first nucleic acid primer detects said CpG dinucleotide which is methylated. [The present invention 1035] The kit of the present invention, wherein the at least one first nucleic acid primer detects the CpG dinucleotide that is unmethylated and the at least one second nucleic acid primer detects the CpG dinucleotide that is methylated. [The present invention 1036] The kit of the present invention, wherein the at least one first nucleic acid primer detects the CpG dinucleotide that is methylated and the at least one second nucleic acid primer detects the CpG dinucleotide that is unmethylated. [The present invention 1037] The kit of claim 1032, further comprising at least a third nucleic acid primer at least 8 nucleotides in length that is complementary to a nucleic acid sequence upstream of the CpG dinucleotide at position 92203667 of chromosome 1 located in the TGFBR gene. [The present invention 1038] The kit of claim 1037, further comprising at least a fourth nucleic acid primer at least 8 nucleotides in length that is complementary to a nucleic acid sequence downstream of the CpG dinucleotide at position 92203667 of chromosome 1 located in the TGFBR gene. [The present invention 1039] The kit of claim 1037, wherein the at least third nucleic acid primer is complementary to a bisulfite converted nucleic acid sequence. [The present invention 1040] The kit of claim 1038, wherein the at least fourth nucleic acid primer is complementary to a bisulfite converted nucleic acid sequence. [The present invention 1041] The kit of claim 1032, wherein the at least one second nucleic acid primer comprises one or more nucleotide analogues. [The present invention 1042] The kit of claim 1033, wherein the at least one second nucleic acid primer comprises one or more synthetic or non-naturally occurring nucleotides. [The present invention 1043] The kit of claim 1032, further comprising a solid substrate to which said at least one first nucleic acid primer is attached. [The present invention 1044] The kit of claim 1043, wherein the substrate is a polymer, glass, semiconductor, paper, metal, gel, or hydrogel. [The present invention 1045] The kit of claim 1043, wherein the solid substrate is a microarray or a microfluidic card. [The present invention 1046] The kit of any one of claims 1031 to 1045, further comprising a detectable label. [The present invention 1047] A kit for determining the methylation status of at least one CpG dinucleotide, comprising: at least one first nucleic acid primer at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide at position 92203667 of chromosome 1 located in the transforming growth factor beta receptor III (TGFBR3) gene, wherein the first nucleic acid primer detects either the unmethylated CpG dinucleotide or the methylated CpG dinucleotide; and A detectable label selected from the group consisting of an enzyme label, a fluorescent label, and a chromogenic label. The kit comprising: [The present invention 1048] The kit of claim 1047, further comprising at least one second nucleic acid primer of at least 8 nucleotides in length complementary to a bisulfite converted nucleic acid sequence comprising a CpG dinucleotide at position 92203667 of chromosome 1 in the TGFBR gene, wherein the at least one second nucleic acid primer detects the unmethylated CpG dinucleotide or the methylated CpG dinucleotide, and the opposite CpG dinucleotide is detected by the at least one first nucleic acid primer. [The present invention 1049] The kit of any one of claims 1047 to 1048, wherein said at least one first nucleic acid primer detects said CpG dinucleotide that is unmethylated. [The present invention 1050] The kit of any one of claims 1047 to 1048, wherein said at least one first nucleic acid primer detects said CpG dinucleotide which is methylated. [The present invention 1051] The kit of the present invention, wherein the at least one first nucleic acid primer detects the CpG dinucleotide that is unmethylated and the at least one second nucleic acid primer detects the CpG dinucleotide that is methylated. [The present invention 1052] The kit of the present invention, wherein the at least one first nucleic acid primer detects the CpG dinucleotide that is methylated and the at least one second nucleic acid primer detects the CpG dinucleotide that is unmethylated. [The present invention 1053] The kit of claim 1048, further comprising at least a third nucleic acid primer at least 8 nucleotides in length that is complementary to a nucleic acid sequence upstream of the CpG dinucleotide at position 92203667 of chromosome 1 located in the TGFBR gene. [The present invention 1054] The kit of claim 1053, further comprising at least a fourth nucleic acid primer at least 8 nucleotides in length that is complementary to a nucleic acid sequence downstream of the CpG dinucleotide at position 92203667 of chromosome 1 located in the TGFBR gene. [The present invention 1055] The kit of claim 1053, wherein the at least a third nucleic acid primer is complementary to a bisulfite converted nucleic acid sequence. [The present invention 1056] The kit of claim 1054, wherein the at least fourth nucleic acid primer is complementary to a bisulfite converted nucleic acid sequence. [The present invention 1057] The kit of any one of claims 1047 to 1056, wherein the at least one first nucleic acid primer comprises one or more nucleotide analogues. [The present invention 1058] The kit of any of claims 1047 to 1056, wherein the at least one first nucleic acid primer comprises one or more synthetic or non-natural nucleotides. [The present invention 1059] The kit of any of claims 1047 to 1058, further comprising a solid substrate to which said at least one first nucleic acid primer is attached. [The present invention 1060] The kit of claim 1059, wherein the substrate is a polymer, glass, semiconductor, paper, metal, gel, or hydrogel. [The present invention 1061] The kit of claim 1059, wherein the solid substrate is a microarray or a microfluidic card. [The present invention 1062] A kit for determining the methylation status of at least one CpG dinucleotide, comprising: at least one first nucleic acid primer at least 8 nucleotides in length that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide at position 92203667 of chromosome 1 located in the transforming growth factor beta receptor III (TGFBR3) gene, wherein the first nucleic acid primer detects either the unmethylated CpG dinucleotide or the methylated CpG dinucleotide; and a solid substrate to which the at least one first nucleic acid primer is attached. The kit comprising: [The present invention 1063] The kit of claim 1062, further comprising at least one second nucleic acid primer of at least 8 nucleotides in length complementary to a bisulfite converted nucleic acid sequence comprising a CpG dinucleotide at position 92203667 of chromosome 1 in the TGFBR gene, wherein the at least one second nucleic acid primer detects the unmethylated CpG dinucleotide or the methylated CpG dinucleotide, and the opposite CpG dinucleotide is detected by the at least one first nucleic acid primer. [The present invention 1064] The kit of claim 1062 or 1063, wherein said at least one first nucleic acid primer detects said CpG dinucleotide that is unmethylated. [The present invention 1065] 1064. The kit of claim 1062 or 1063, wherein said at least one first nucleic acid primer detects said CpG dinucleotide that is methylated. [The present invention 1066] The kit of the present invention, wherein the at least one first nucleic acid primer detects the CpG dinucleotide that is unmethylated and the at least one second nucleic acid primer detects the CpG dinucleotide that is methylated. [The present invention 1067] The kit of the present invention, wherein the at least one first nucleic acid primer detects the CpG dinucleotide that is methylated and the at least one second nucleic acid primer detects the CpG dinucleotide that is unmethylated. [The present invention 1068] The kit of claim 1063, further comprising at least a third nucleic acid primer at least 8 nucleotides in length that is complementary to a nucleic acid sequence upstream of the CpG dinucleotide at position 92203667 of chromosome 1 located within the TGFBR gene. [The present invention 1069] The kit of claim 1068, further comprising at least a fourth nucleic acid primer at least 8 nucleotides in length that is complementary to a nucleic acid sequence downstream of the CpG dinucleotide at position 92203667 on chromosome 1 located in the TGFBR gene. [The present invention 1070] The kit of claim 1068, wherein the at least a third nucleic acid primer is complementary to a bisulfite converted nucleic acid sequence. [The present invention 1071] The kit of claim 1069, wherein the at least fourth nucleic acid primer is complementary to a bisulfite converted nucleic acid sequence. [The present invention 1072] The kit of any one of claims 1062 to 1071, wherein the at least one first nucleic acid primer comprises one or more nucleotide analogues. [The present invention 1073] The kit of any of claims 1062 to 1071, wherein the at least one first nucleic acid primer comprises one or more synthetic or non-natural nucleotides. [The present invention 1074] The kit of any one of claims 1062 to 1073, wherein the substrate is a polymer, glass, semiconductor, paper, metal, gel, or hydrogel. [The present invention 1075] The kit of any one of claims 1062 to 1073, wherein the solid substrate is a microarray or a microfluidic card. [The present invention 1076] The kit of any one of claims 1062 to 1075, further comprising a detectable label. [The present invention 1077] 1. A method for predicting the presence of a biomarker associated with cardiovascular disease (CVD) in a biological sample from a patient, comprising: (a) providing a first aliquot from the biological sample and contacting DNA from the first biological sample with bisulfite under alkaline conditions; (b) providing a second aliquot from the biological sample; and (c)(i) contacting the first aliquot with a first oligonucleotide probe at least 8 nucleotides in length that is complementary to a sequence that includes a CpG dinucleotide at position 92203667 on chromosome 1 located in the transforming growth factor beta receptor III (TGFBR3) gene, and contacting the second aliquot with a nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs347027; (ii) contacting the first aliquot with a first oligonucleotide probe at least 8 nucleotides in length that is complementary to a sequence that includes a CpG dinucleotide at position 38364951 within an intergenic region of chromosome 15, and contacting the second aliquot with a nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs4937276; (iii) contacting the first aliquot with a first oligonucleotide probe at least 8 nucleotides in length that is complementary to a sequence that includes a CpG dinucleotide at position 84206068 on chromosome 4 located in the coenzyme Q2 4-hydroxybenzoate polyprenyltransferase (COQ2) gene, and contacting the second aliquot with a nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs17355663; (iv) contacting the first aliquot with a first oligonucleotide probe at least 8 nucleotides in length that is complementary to a sequence that includes a CpG dinucleotide at position 26146070 on chromosome 16, which is located in the heparan sulfate 3-O-sulfotransferase 4 (HS3ST4) gene, and contacting the second aliquot with a nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs235807; (v) contacting the first aliquot with a first oligonucleotide probe at least 8 nucleotides in length that is complementary to a sequence that includes a CpG dinucleotide at position 91171013 in the intergenic region of chromosome 1, and contacting the second aliquot with a nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs11579814; (vi) contacting the first aliquot with a first oligonucleotide probe at least 8 nucleotides in length that is complementary to a sequence that includes a CpG dinucleotide at position 39491936 on chromosome 1 within the NADH dehydrogenase (ubiquinone) Fe-S protein 5 (NDUFS5) gene, and contacting the second aliquot with a nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs2275187; (vii) contacting the first aliquot with a first oligonucleotide probe at least 8 nucleotides in length that is complementary to a sequence that includes a CpG dinucleotide at position 186426136 mapped to chromosome 1 within the phosducin gene, and contacting the second aliquot with a nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs4336803; (viii) contacting the first aliquot with a first oligonucleotide probe at least 8 nucleotides in length that is complementary to a sequence that includes a CpG dinucleotide at position 205475130 of chromosome 1 located in the cyclin-dependent kinase 18 (CDK18) gene, and contacting the second aliquot with a nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs4951158; and / or (ix) contacting the first aliquot with a first oligonucleotide probe at least 8 nucleotides in length that is complementary to a sequence that includes a CpG dinucleotide at position 130614013 on chromosome 3 located in the ATPase, Ca++ Transporting, Type 2C, Member 1 (ATP2C1) gene, and contacting the second aliquot with a nucleic acid primer at least 8 nucleotides in length that is complementary to rs925613; The method, wherein methylation of the CpG dinucleotide at position 92203667 on chromosome 1, cg20636912, cg16947947, cg05916059, cg04567738, cg16603713, cg05709437, cg12081870, and / or cg18070470 in the TGFBR3 gene and a G at position 1618766 on chromosome 1, or polymorphisms at rs4937276, rs17355663, rs235807, rs11579814, rs2275187, rs4336803, rs4951158, and / or rs925613 are associated with CVD. [The present invention 1078] The method of claim 1077, wherein the biological sample is a saliva sample. [The present invention 1079] 1. A method for predicting the presence of a biomarker associated with cardiovascular disease (CVD) in a biological sample from a patient, the method comprising detecting one or more pairs of SNPs and CpGs set forth in Table 3. [The present invention 1080] 1. A method for detecting one or more copies of the G allele at rs347027 and the methylation status of the CpG at position 92203667 of chromosome 1 in a nucleic acid sample from a subject at risk for cardiovascular disease (CVD), comprising: (a) performing a genotyping assay on a nucleic acid sample of the human subject to detect the presence of one or more copies of the G allele of the rs347027 polymorphism; and (b) performing a methylation assessment on the nucleic acid sample of said human to determine whether the CpG at position 92203667 of chromosome 1 is unmethylated. The method comprising: [The present invention 1081] The method according to any one of claims 1077 to 1080, wherein the CVD is coronary heart disease (CHD). [The present invention 1082] The method according to any one of claims 1077 to 1080, wherein the CVD is congestive heart failure (CHF). [The present invention 1083] The method according to any one of claims 1077 to 1080, wherein the CVD is stroke. [The present invention 1084] 1. A method for determining the presence of a biomarker associated with CHD in a patient sample, comprising: (a) isolating a nucleic acid sample from said patient sample; (b) performing a genotyping assay on a first aliquot of the nucleic acid sample to detect the presence of at least one SNP to obtain genotype data, wherein the at least one SNP is a first SNP set forth in FIG. 21 and / or a second SNP that is in linkage disequilibrium (R>0.3) with the first SNP set forth in FIG. 21; and / or (c) bisulfite converting the nucleic acid in a second aliquot of the nucleic acid and performing a methylation assessment on the second aliquot of the nucleic acid sample to detect the methylation status of at least one gene set forth in FIG. 15 and / or the methylation status of a first CpG site set forth in FIG. 16 and / or the methylation status of a second CpG site that is collinear with the first CpG set forth in FIG. 16, to obtain methylation data regarding whether a particular CpG residue is unmethylated; and (d) inputting the genotype data from step (b) and / or the methylation data from step (c) into at least one algorithm that determines the contribution of the main effect of at least one SNP and / or the main effect of at least one CpG and / or at least one interaction effect. The method comprising: [The present invention 1085] 1. A method for determining the presence of a biomarker associated with stroke in a patient sample, comprising: (a) isolating a nucleic acid sample from said patient sample; (b) performing a genotyping assay on a first aliquot of the nucleic acid sample to detect the presence of at least one SNP to obtain genotype data, wherein the at least one SNP is a first SNP set forth in FIG. 22 and / or a second SNP that is in linkage disequilibrium with the first SNP set forth in FIG. 22; and / or (c) bisulfite converting the nucleic acid in a second aliquot of the nucleic acid and performing a methylation assessment on the second aliquot of the nucleic acid sample to detect the methylation status of at least one gene set forth in FIG. 17 and / or the methylation status of a first CpG site set forth in FIG. 18 and / or the methylation status of a second CpG site that is collinear with the first CpG set forth in FIG. 18, to obtain methylation data regarding whether a particular CpG residue is unmethylated; and (d) inputting the genotype data from step (b) and / or the methylation data from step (c) into an algorithm that determines the contribution of the main effect of at least one SNP and / or the main effect of at least one CpG and / or at least one interaction effect. The method comprising: [The present invention 1086] 1. A method for determining the presence of a biomarker associated with CHF in a patient sample, comprising: (a) isolating a nucleic acid sample from said patient sample; (b) performing a genotyping assay on a first aliquot of the nucleic acid sample to detect the presence of at least one SNP to obtain genotype data, wherein the SNP is a first SNP set forth in FIG. 23 and / or a second SNP that is in linkage disequilibrium (R>0.3) with the first SNP set forth in FIG. 23; and / or (c) bisulfite converting the nucleic acid in a second aliquot of the nucleic acid and performing a methylation assessment on the second aliquot of the nucleic acid sample to detect the methylation status of at least one gene set forth in FIG. 19 and / or the methylation status of a first CpG site set forth in FIG. 20 and / or the methylation status of a second CpG site that is collinear with the first CpG set forth in FIG. 20, to obtain methylation data regarding whether a particular CpG residue is unmethylated; and (d) inputting the genotype data from step (b) and / or the methylation data from step (c) into an algorithm that determines the contribution of the main effect of at least one SNP and / or the main effect of at least one CpG and / or at least one interaction effect. The method comprising: [The present invention 1087] Any of the methods of claims 1084 to 1086, wherein the at least one interaction effect is selected from the group consisting of gene-environment interaction (SNP x CpG) effects, gene-gene interaction (SNP x SNP) effects, and environment-environment interaction (CpG x CpG) effects. [The present invention 1088] The method of the present invention 1084, wherein the results include a gene-environment interaction effect (SNP x CpG) between a second CpG site that is collinear with the first CpG site listed in Figure 16 and a SNP listed in Figure 21 or a second SNP that is in linkage disequilibrium with the first SNP listed in Figure 21. [The present invention 1089] The method of the present invention 1084, wherein the results include at least one environment-environment interaction effect (CpG x CpG) between at least two genes listed in Figure 15 and / or at least two CpG sites listed in Figure 16. [The present invention 1090] The method of the present invention 1084, wherein the results include at least one environment-environment interaction effect (CpG x CpG) between at least two CpG sites that are collinear with the first CpG site described in Figure 16. [The present invention 1091] The method of the present invention 1085, wherein the results include a gene-environment interaction effect (SNP x CpG) between a second CpG site that is collinear (R>0.3) with the first CpG site listed in Figure 18 and a SNP listed in Figure 22 or a second SNP that is in linkage disequilibrium with the first SNP listed in Figure 22. [The present invention 1092] The method of the present invention 1085, wherein the results include at least one environment-environment interaction effect (CpG x CpG) between at least two genes listed in Figure 17 and / or at least two CpG sites listed in Figure 18. [The present invention 1093] The method of the present invention 1085, wherein the results include at least one environment-environment interaction effect (CpG x CpG) between at least two CpG sites that are collinear with the first CpG site described in Figure 18. [The present invention 1094] The method of the present invention 1086, wherein the results include a gene-environment interaction effect (SNP x CpG) between a second CpG site that is collinear with the first CpG site listed in Figure 20 and a first SNP listed in Figure 23 or a second SNP that is in linkage disequilibrium with the first SNP listed in Figure 23. [The present invention 1095] The method of the present invention 1086, wherein the results include at least one environment-environment interaction effect (CpG x CpG) between at least two genes listed in Figure 19 and / or at least two CpG sites listed in Figure 20. [The present invention 1096] The method of the present invention 1086, wherein the results include at least one environment-environment interaction effect (CpG x CpG) between at least two CpG sites that are collinear with the first CpG site described in Figure 20. [Brief description of the drawings]

[0035] [Figure 1]Area under the receiver operating characteristic curve for cg05575921 (A), age+sex+batch+cg05575921 (B), self-reported smoking status (C), and age+sex+batch+self-reported smoking status (D). [Diagram 2] Area under the receiver operating characteristic curve for CHD prediction models (not optimized). [Diagram 3] Protein-protein interactome of CHD. Network of the top 1000 genes with at least one DNA methylation probe significantly associated with syndromic CHD. [Figure 4] Venn diagram of DNA methylation probes significantly associated with symptomatic CHD and its traditional modifiable risk factors. [Diagram 5] Venn diagram of genes with at least one DNA methylation probe significantly associated with symptomatic CHD and its traditional modifiable risk factors. [Figure 6] ROC curve of the integrated gene-epigenetic model with the highest average 10-fold cross-validation AUC value. [Figure 7] ROC curve of the conventional risk factor model with the highest average 10-fold cross-validation AUC value. [Figure 8] Partial dependence plots of DNA methylation sites and SNPs. [Figure 9] Two-dimensional histograms of sensitivity and specificity for 10,000 permutations of DNA methylation sites and SNPs. [Figure 10] ROC curves for main effects of the CHF classification model. [Figure 11] ROC curves for interaction effects in the CHF classification model. [Figure 12] ROC curves for main effects of stroke classification models. [Figure 13] ROC curves for interaction effects in stroke classification models. [Figure 14] 1 is a flow chart of a particular embodiment of the method of the present invention. [Figure 15-1] List of genes whose methylation is associated with CHD. [Figure 15-2]List of genes whose methylation is associated with CHD. [Figure 15-3] List of genes whose methylation is associated with CHD. [Figure 15-4] List of genes whose methylation is associated with CHD. [Figure 15-5] List of genes whose methylation is associated with CHD. [Figure 15-6] List of genes whose methylation is associated with CHD. [Figure 15-7] List of genes whose methylation is associated with CHD. [Figure 15-8] List of genes whose methylation is associated with CHD. [Figure 15-9] List of genes whose methylation is associated with CHD. [Figure 15-10] List of genes whose methylation is associated with CHD. [Figure 15-11] List of genes whose methylation is associated with CHD. [Figure 15-12] List of genes whose methylation is associated with CHD. [Figure 15-13] List of genes whose methylation is associated with CHD. [Figure 15-14] List of genes whose methylation is associated with CHD. [Figure 15-15] List of genes whose methylation is associated with CHD. [Figure 15-16] List of genes whose methylation is associated with CHD. [Figure 15-17] List of genes whose methylation is associated with CHD. [Figure 15-18] List of genes whose methylation is associated with CHD. [Figure 15-19] List of genes whose methylation is associated with CHD. [Figure 15-20] List of genes whose methylation is associated with CHD. [Figure 15-21] List of genes whose methylation is associated with CHD. [Figure 16-1]List of CpGs whose methylation is associated with CHD. [Figure 16-2] List of CpGs whose methylation is associated with CHD. [Figure 16-3] List of CpGs whose methylation is associated with CHD. [Figure 16-4] List of CpGs whose methylation is associated with CHD. [Figure 16-5] List of CpGs whose methylation is associated with CHD. [Figure 16-6] List of CpGs whose methylation is associated with CHD. [Figure 16-7] List of CpGs whose methylation is associated with CHD. [Figure 16-8] List of CpGs whose methylation is associated with CHD. [Figure 16-9] List of CpGs whose methylation is associated with CHD. [Figure 16-10] List of CpGs whose methylation is associated with CHD. [Figure 16-11] List of CpGs whose methylation is associated with CHD. [Figure 16-12] List of CpGs whose methylation is associated with CHD. [Figure 16-13] List of CpGs whose methylation is associated with CHD. [Figure 16-14] List of CpGs whose methylation is associated with CHD. [Figure 16-15] List of CpGs whose methylation is associated with CHD. [Figure 16-16] List of CpGs whose methylation is associated with CHD. [Figure 16-17] List of CpGs whose methylation is associated with CHD. [Figure 16-18] List of CpGs whose methylation is associated with CHD. [Figure 16-19] List of CpGs whose methylation is associated with CHD. [Figure 16-20] List of CpGs whose methylation is associated with CHD. [Figure 16-21]List of CpGs whose methylation is associated with CHD. [Figure 16-22] List of CpGs whose methylation is associated with CHD. [Figure 16-23] List of CpGs whose methylation is associated with CHD. [Figure 16-24] List of CpGs whose methylation is associated with CHD. [Figure 16-25] List of CpGs whose methylation is associated with CHD. [Figure 17-1] List of genes whose methylation is associated with stroke. [Figure 17-2] List of genes whose methylation is associated with stroke. [Figure 17-3] List of genes whose methylation is associated with stroke. [Figure 17-4] List of genes whose methylation is associated with stroke. [Figure 17-5] List of genes whose methylation is associated with stroke. [Figure 17-6] List of genes whose methylation is associated with stroke. [Figure 17-7] List of genes whose methylation is associated with stroke. [Figure 17-8] List of genes whose methylation is associated with stroke. [Figure 17-9] List of genes whose methylation is associated with stroke. [Figure 17-10] List of genes whose methylation is associated with stroke. [Figure 17-11] List of genes whose methylation is associated with stroke. [Figure 18-1] List of CpGs whose methylation is associated with stroke. [Figure 18-2] List of CpGs whose methylation is associated with stroke. [Figure 18-3] List of CpGs whose methylation is associated with stroke. [Figure 18-4] List of CpGs whose methylation is associated with stroke. [Figure 18-5]List of CpGs whose methylation is associated with stroke. [Figure 18-6] List of CpGs whose methylation is associated with stroke. [Figure 18-7] List of CpGs whose methylation is associated with stroke. [Figure 18-8] List of CpGs whose methylation is associated with stroke. [Figure 18-9] List of CpGs whose methylation is associated with stroke. [Figure 18-10] List of CpGs whose methylation is associated with stroke. [Figure 18-11] List of CpGs whose methylation is associated with stroke. [Figure 18-12] List of CpGs whose methylation is associated with stroke. [Figure 18-13] List of CpGs whose methylation is associated with stroke. [Figure 19-1] List of genes whose methylation is associated with CHF. [Figure 19-2] List of genes whose methylation is associated with CHF. [Figure 19-3] List of genes whose methylation is associated with CHF. [Figure 19-4] List of genes whose methylation is associated with CHF. [Figure 19-5] List of genes whose methylation is associated with CHF. [Figure 20-1] List of CpGs whose methylation is associated with CHF. [Figure 20-2] List of CpGs whose methylation is associated with CHF. [Figure 20-3] List of CpGs whose methylation is associated with CHF. [Figure 20-4] List of CpGs whose methylation is associated with CHF. [Figure 20-5] List of CpGs whose methylation is associated with CHF. [Figure 20-6] List of CpGs whose methylation is associated with CHF. [Figure 20-7]List of CpGs whose methylation is associated with CHF. [Figure 20-8] List of CpGs whose methylation is associated with CHF. [Figure 20-9] List of CpGs whose methylation is associated with CHF. [Figure 20-10] List of CpGs whose methylation is associated with CHF. [Figure 20-11] List of CpGs whose methylation is associated with CHF. [Figure 20-12] List of CpGs whose methylation is associated with CHF. [Figure 20-13] List of CpGs whose methylation is associated with CHF. [Figure 21-1] List of SNPs associated with CHD. [Figure 21-2] List of SNPs associated with CHD. [Figure 21-3] List of SNPs associated with CHD. [Figure 21-4] List of SNPs associated with CHD. [Figure 21-5] List of SNPs associated with CHD. [Figure 21-6] List of SNPs associated with CHD. [Figure 21-7] List of SNPs associated with CHD. [Figure 21-8] List of SNPs associated with CHD. [Figure 21-9] List of SNPs associated with CHD. [Figure 21-10] List of SNPs associated with CHD. [Figure 22-1] List of SNPs associated with stroke. [Figure 22-2] List of SNPs associated with stroke. [Figure 22-3] List of SNPs associated with stroke. [Figure 22-4] List of SNPs associated with stroke. [Figure 22-5] List of SNPs associated with stroke. [Figure 22-6]List of SNPs associated with stroke. [Figure 22-7] List of SNPs associated with stroke. [Figure 23-1] List of SNPs associated with CHF. [Figure 23-2] List of SNPs associated with CHF. [Figure 23-3] List of SNPs associated with CHF. [Figure 23-4] List of SNPs associated with CHF. [Figure 23-5] List of SNPs associated with CHF. [Figure 23-6] List of SNPs associated with CHF. [Figure 23-7] List of SNPs associated with CHF. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0036] Detailed Description of the Invention The present disclosure provides a method and kit for determining whether a subject has a predisposition to cardiovascular disease (CVD) or has the possibility of having or developing the disease.As shown herein, the methylation status of one or more CpG dinucleotides is associated with CVD, either alone or in combination with genotype and / or the interaction between genotype and methylation status (e.g., CH3xSNP).As used herein, the term "predisposition" is defined as the tendency or susceptibility of a subject to develop a certain condition.For example, a subject is more likely to develop a certain condition than a control subject.

[0037] DNA methylation DNA does not exist as a naked molecule in cells. For example, DNA associates with proteins called histones to form a complex substance known as chromatin. Chemical modifications of DNA or histones change the structure of chromatin without changing the nucleotide sequence of the DNA. Such modifications are described as "epigenetic" modifications of DNA. Changes in the structure of chromatin can have a profound effect on gene expression. When chromatin is condensed, factors involved in gene expression may not be able to access the DNA, and genes will be switched off. Conversely, when chromatin is "open," genes can be switched on. Some important forms of epigenetic modifications are DNA methylation and histone deacetylation. DNA methylation is a chemical modification of the DNA molecule itself, carried out by enzymes called DNA methyltransferases. Methylation can directly switch off gene expression by preventing transcription factors from binding to the promoter. A more common effect is the attraction of methyl-binding domain (MBD) proteins. These interact with further enzymes called histone deacetylases (HDACs), which function to chemically modify histones and alter chromatin structure. Chromatin containing acetylated histones is gapped and accessible to transcription factors, and genes are potentially active. Histone deacetylation leads to condensation of chromatin, making it inaccessible to transcription factors and resulting in gene silencing.

[0038] CpG islands are short stretches of DNA in which the frequency of CpG sequences is higher than in other regions. The "p" in the term CpG indicates that a cysteine ​​("C") and a guanine ("G") are connected by a phosphodiester bond. CpG islands are often located around the promoters of housekeeping genes and many regulatory genes. At these locations, the CG sequences are not methylated. In contrast, CG sequences in inactive genes are usually methylated, thereby repressing their expression.

[0039] As used herein, the term "methylation status" refers to the determination of whether a particular target DNA (e.g., a CpG dinucleotide) is methylated. As used herein, the term "CpG dinucleotide repeat motif" refers to two or more consecutive CpG dinucleotides located in a DNA sequence.

[0040] Approximately 56% of human genes and 47% of mouse genes are associated with CpG islands. Often, CpG islands overlap with promoters and extend downstream to transcription units for approximately 1000 base pairs. Identification of potential CpG islands during sequence analysis helps define the furthest 5' end of genes, which is notoriously difficult to do with cDNA-based approaches. Methylation of CpG islands can be determined by those skilled in the art using any method suitable for determining such methylation. For example, those skilled in the art can use a method based on bisulfite reaction to determine such methylation.

[0041] The present disclosure provides methods for assessing nucleic acid methylation of TGFBR3 in a patient to predict the clinical course and ultimate outcome of the patient predisposed to or suspected of having CHD.

[0042] In particular, in certain embodiments of the present disclosure, the method may be carried out as follows: A sample, such as a blood sample, is taken from a patient. In certain embodiments, a single cell type, such as lymphocytes, basophils, or monocytes, isolated from the blood may be isolated for further testing. DNA is taken from the sample and examined to determine whether the TGFBR3 region is methylated. For example, the DNA of interest may be treated with bisulfite to deaminate unmethylated cytosine residues to uracil. Uracil base pairs with adenosine, so that thymidine is incorporated into the subsequent DNA strand in place of unmethylated cytosine residues during partial sequence PCR amplification. The target sequence is then amplified by PCR and probed with a TGFBR3-specific probe. Only methylated DNA from the patient binds to the probe. A particular profile is associated with a particular disease state.

[0043] Methods for determining patient nucleic acid profile are well known to those skilled in the art and include any of the well-known detection methods.Various PCR methods are described, for example, in PCR Primer: A Laboratory Manual, Dieffenbach 7 Dveksler, Eds., Cold Spring Harbor Laboratory Press, 1995.Other analytical methods include, but are not limited to, nucleic acid quantification, restriction enzyme digestion, DNA sequencing, hybridization techniques, such as Southern blotting, amplification methods, such as ligase chain reaction (LCR), nucleic acid sequence-based amplification (NASBA), self-sustained sequence replication (SSR or 3SR), strand displacement amplification (SDA), and transcription-mediated amplification (TMA), quantitative PCR (qPCR), or other DNA analysis, as well as RT-PCR, in vitro translation, Northern blotting, and other RNA analysis.In another embodiment, hybridization to microarray is used.

[0044] Single Nucleotide Polymorphism (SNP) Genotyping Traditional methods for screening for genetic diseases rely on the identification of either an abnormal gene product (e.g., sickle cell anemia) or an abnormal phenotype (e.g., mental retardation). With the development of simple and inexpensive genetic screening methodologies, it is now possible to identify polymorphisms that indicate a propensity to develop a disease, even when the disease is of polygenic origin.

[0045] Single nucleotide polymorphism (SNP) genotyping measures the genetic variation of SNPs among members of a species. SNPs are single base pair mutations at a specific locus, usually consisting of two alleles (in this case the rare allele frequency is >1%), and are very common. Because SNPs are conserved during evolution, they have been proposed as markers for use in quantitative trait locus (QTL) analysis and association studies, instead of microsatellites. Many different SNP genotyping methods are known, including hybridization-based methods (such as dynamic allele-specific hybridization, molecular beacons, and SNP microarrays), enzyme-based methods (including restriction fragment length polymorphism, PCR-based methods, flap endonuclease, primer extension, 5'-nuclease, and oligonucleotide ligation assays), other post-amplification methods based on the physical properties of DNA (such as single-strand conformation polymorphism, temperature gradient gel electrophoresis, denaturing high performance liquid chromatography, high resolution melting of the entire amplicon, use of DNA mismatch binding protein, SNPlex and surveyor nuclease assays), and sequencing (such as "next-generation" sequencing).See, for example, U.S. Patent No. 7,972,779.

[0046] Multiple alleles with distinct organ functions (e.g., high and low levels of expression in the heart, or high, medium, and low levels of expression in the heart) can result from one or more polymorphisms in the region of the gene that encodes the polypeptide, or can be in a regulatory control sequence that affects the expression of the polypeptide, such as a promoter or polyadenylation sequence. Alternatively, related alleles can result from one or more polymorphisms in a locus distal to the gene that has a direct effect on the identified property, where the product of the distal locus has an indirect effect on the property. Related alleles can affect the polypeptide at the transcription or translation level, and can affect the transcription rate, translation rate, degradation rate, or activity of the polypeptide. The difference between alleles in brain function genes can be characterized in samples from one subject or from multiple subjects by methods for assaying any of the above known to those skilled in the art. Such methods can include, but are not limited to, measuring the amount of encoded polypeptide and measuring the potential of the polynucleotide sequence to be expressed. The assay method can detect proteins or nucleic acids directly or indirectly. The suitability of an upstream promoter region for directing transcription of a coding region of a polynucleotide encoding a polypeptide can be assessed, or the suitability of a coding region encoding a functional polypeptide can be assessed. It is specifically contemplated that the assay method includes screening for the presence of a particular sequence or structure of a nucleic acid or polypeptide, for example, using any of a variety of known microarray technologies.

[0047] One of skill in the art will appreciate that an allele need not have been previously shown to have any association or connection with a disorder phenotype, instead, an allele and a pathogenic environmental risk factor may interact to predict predisposition to a disorder phenotype, even if neither has any direct relationship to the disorder phenotype.

[0048] Genetic screening (also called genotyping or molecular screening) can be broadly defined as a test to determine whether a patient has a mutation (or allele or polymorphism) that either causes a disease state or is "linked" to a mutation that causes a disease state. Linkage refers to the phenomenon whereby DNA sequences that are near each other in the genome have a tendency to be inherited together. Two sequences may be linked due to some selective advantage of co-inheritance. More typically, however, two polymorphic sequences are co-inherited because meiotic recombination events occur relatively infrequently within the region between the two polymorphisms. Co-inherited polymorphic alleles are said to be in "linkage disequilibrium" with each other in a given population because they tend to either occur together in any particular member of the population or not at all. Indeed, when multiple polymorphisms in a given chromosomal region are found to be in linkage disequilibrium with each other, they define a metastable genetic "haplotype." In contrast, recombination events that occur between two polymorphic loci leave them separated on separate homologous chromosomes. When meiotic recombination occurs frequently enough between two physically linked polymorphisms, the two polymorphisms appear to segregate independently and are said to be in linkage equilibrium.

[0049] It will be understood that linkage equilibrium / disequilibrium can be quantified (e.g., using Pearson correlation (R) or allele co-inheritance (D')). For example, a low level of linkage can be reflected by a correlation (e.g., R value) of about 0.1 or less, a medium level of linkage can be reflected by an R value of about 0.3, and a high level of linkage can be reflected by an R value of 0.5 or more. It will also be understood that when referring to methylation (i.e., CpGs), collinearity (having an R value) is used as a measure of the strength of the linear association between two CpGs (e.g., a low level of collinearity can be reflected by an R value of about 0.1 or less; a medium level of collinearity can be reflected by an R value of about 0.3; and a high level of collinearity can be reflected by an R value of about 0.5 or more).

[0050] The frequency of meiotic recombination between two markers is generally proportional to the physical distance between them on a chromosome, but the occurrence of "hot spots" and areas of suppressed chromosomal recombination can lead to discrepancies between the physical and recombination distances between two markers. Thus, in a particular chromosomal region, multiple polymorphic loci spanning a wide chromosomal domain may be in linkage disequilibrium with each other, thereby defining a wide-range gene haplotype. Furthermore, if a disease-causing mutation is found within or linked to this haplotype, one or more polymorphic alleles of the haplotype can be used as a diagnostic or prognostic indicator of the likelihood of developing the disease. This linkage between other benign polymorphisms and disease-causing polymorphisms occurs when the disease mutation has arisen so recently that not enough time has passed for equilibrium to be achieved through recombination events. Therefore, the identification of a haplotype that spans or is associated with a disease-causing mutational change serves as a means of predicting the likelihood of an individual inheriting the disease-causing mutation. Such prognostic or diagnostic procedures can be utilized without the need to identify and isolate the actual disease-causing lesion. This is important because precise determination of the molecular defect involved in a disease process can be difficult and challenging, especially in the case of multifactorial diseases.

[0051] A statistical correlation between a disorder and a polymorphism does not necessarily indicate that the polymorphism directly causes the disorder. Rather, a correlated polymorphism may be a harmless allelic variant that is evolutionarily recent and is associated with (i.e., in linkage disequilibrium with) a disorder-causing mutation such that not enough time has passed for equilibrium to be achieved with intervening chromosomal segments through recombination events. Thus, for the purposes of diagnostic and prognostic assays for a particular disease, detection of a polymorphic allele associated with that disease can be utilized without considering whether the polymorphism is directly involved in the etiology of the disease. Furthermore, if a given harmless polymorphic locus is in linkage disequilibrium with an apparent disease-causing polymorphic locus, it is likely that other polymorphic loci in linkage disequilibrium with the harmless polymorphic locus are also in linkage disequilibrium with the disease-causing polymorphic locus. Thus, these other polymorphic loci are also used in the prognosis or diagnosis of the likelihood that the disease-causing polymorphic locus has been inherited. Once an association has been derived between a particular disease or condition and the corresponding haplotype, the extended haplotype (referring to a typical pattern of co-inheritance of alleles of a set of linked polymorphic markers) can be targeted for diagnostic purposes. Thus, by characterizing one or more disease-associated polymorphic alleles (or even one or more disease-associated haplotypes), the likelihood of an individual developing a particular disease or condition can be determined, without necessarily determining or characterizing the causative genetic variation.

[0052] Many methods are available for detecting specific alleles of polymorphic loci. A particular method for detecting a specific polymorphic allele will depend in part on the molecular nature of the polymorphism. For example, various allelic forms of a polymorphic locus may differ by a single base pair of DNA. Such single nucleotide polymorphisms (or SNPs) are the major contributors to genetic variation, comprising approximately 80% of all known polymorphisms, and their density in the genome is estimated to be an average of 1 per 1,000 base pairs. SNPs are frequently biallelic or occur in only two different forms (although up to four different forms of SNPs are theoretically possible, corresponding to the four different nucleotide bases present in DNA). Nevertheless, SNPs are more mutationally stable than other polymorphisms, which makes them suitable for association studies that use linkage disequilibrium between markers and unknown variants to map disease-causing mutations. In addition, because SNPs typically have only two alleles, they can be genotyped by simple plus / minus assays rather than length measurements, which makes them more amenable to automation.

[0053] In one embodiment, allele profiling can be performed using nucleic acid microarrays, which can be commercially available alone or in combination with one or more kit components.The genetic testing field is rapidly evolving, and therefore, those skilled in the art will recognize that a wide range of profiling tests exist and will be developed to determine the allele profile of an individual according to the present disclosure.

[0054] Nucleic Acids and Polypeptides The term "nucleic acid" refers to deoxyribonucleotides or ribonucleotides and polymers thereof in either single-stranded or double-stranded form, made of monomers (nucleotides) containing sugars, phosphates, and bases (either purines or pyrimidines). Unless specifically limited, the term encompasses nucleic acids containing known analogs of natural nucleotides that have similar binding properties as the reference nucleic acid and are metabolized similarly to naturally occurring nucleotides. Unless otherwise indicated, a particular nucleic acid sequence also encompasses not only the sequence specified, but also conservatively modified variants thereof (e.g., degenerate codon substitutions) and complementary sequences. Specifically, degenerate codon substitutions can be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed bases and / or deoxyinosine residues. The terms "nucleic acid", "nucleic acid molecule" or "polynucleotide" are used interchangeably and may also be used interchangeably with gene, cDNA, DNA and / or RNA encoded by a gene.

[0055] The term "nucleotide sequence" refers to a polymer of DNA or RNA that can be single- or double-stranded, optionally containing synthetic, non-natural or altered nucleotide bases that can be incorporated into the DNA or RNA polymer. A DNA molecule or polynucleotide is a polymer of deoxyribonucleotides (A, G, C, and T), and an RNA molecule or polynucleotide is a polymer of ribonucleotides (A, G, C, and U).

[0056] "Gene", for the purposes of this disclosure, includes DNA regions that code for a gene product, as well as all DNA regions that regulate the production of the gene product, whether or not such regulatory sequences are adjacent to the coding and / or transcriptional sequences. The term "gene" is used broadly to refer to any segment of nucleic acid associated with a biological function. Genes include coding sequences and / or regulatory sequences required for their expression. Thus, genes include, but need not be limited to, promoter sequences, terminators, translational regulatory sequences, such as ribosome binding sites and internal ribosome entry sites, enhancers, silencers, insulators, boundary elements, origins of replication, matrix attachment sites, and locus control regions. For example, "gene" refers to an mRNA, functional RNA, or nucleic acid fragment that expresses a specific protein, including regulatory sequences. "Functional RNA" refers to sense RNA, antisense RNA, ribozyme RNA, siRNA, or other RNA that cannot be translated but still has an effect on at least one cellular process. "Gene" also includes non-expressed DNA segments, for example, forming recognition sequences for other proteins. A "gene" can be obtained from a variety of sources, including cloning from a source of interest or synthesis from known or predicted sequence information, and may contain a sequence designed to have desired parameters.

[0057] "Gene expression" refers to the conversion of information contained in a gene into a gene product. It refers to the transcription and / or translation of an endogenous gene, a heterologous gene or nucleic acid segment, or a transgene in a cell. In addition, expression refers to the transcription and stable accumulation of sense (mRNA) or functional RNA. Expression may also refer to the production of a protein. The term "altered levels of expression" refers to a level of expression in a transgenic cell or organism that differs from that in a normal or non-transformed cell or organism.

[0058] A gene product can be the direct transcription product of a gene (e.g., mRNA, tRNA, rRNA, antisense RNA, ribozyme, structural RNA, or any other type of RNA) or a protein produced by translation of an mRNA. Gene products also include RNA modified by processes such as capping, polyadenylation, methylation, and editing, as well as proteins modified by, for example, methylation, acetylation, phosphorylation, ubiquitination, ADP-ribosylation, myristylation, and glycosylation. The term "RNA transcript" refers to the product resulting from RNA polymerase-catalyzed transcription of a DNA sequence. When the RNA transcript is a perfect complementary copy of the DNA sequence, it is referred to as the primary transcript, or it can be an RNA sequence derived from post-transcriptional processing of the primary transcript and is referred to as the mature RNA. "Messenger RNA" (mRNA) refers to an RNA that is free of introns and can be translated into a protein by a cell. "cDNA" refers to a single- or double-stranded DNA that is complementary to and derived from an mRNA. "Functional RNA" refers to sense RNA, antisense RNA, ribozyme RNA, siRNA, or other RNA that cannot be translated yet still has an effect on at least one cellular process.

[0059] A "coding sequence," or a sequence that "encodes" a selected polypeptide, is a nucleic acid molecule that is transcribed (in the case of DNA) and translated (in the case of mRNA) into a polypeptide in vivo when placed under the control of appropriate regulatory sequences. The boundaries of the coding sequence are determined by a start codon at the 5' (amino) terminus and a translation stop codon at the 3' (carboxy) terminus. Coding sequences can include, but are not limited to, cDNA from viral, prokaryotic or eukaryotic mRNA, genomic DNA sequences from viral (e.g., DNA viruses and retroviruses) or prokaryotic DNA, and synthetic DNA sequences, among others. A transcription termination sequence may be located 3' to the coding sequence.

[0060] Certain embodiments of the present disclosure encompass isolated or substantially purified nucleic acid compositions. In the context of the present disclosure, an "isolated" or "purified" DNA or RNA molecule is a DNA or RNA molecule that exists apart from its native environment and is therefore not a product of nature. An isolated DNA or RNA molecule may exist in purified form or in a non-native environment, such as in a transgenic host cell. For example, an "isolated" or "purified" nucleic acid molecule is substantially free of other cellular materials or culture medium when produced by recombinant techniques, or substantially free of chemical precursors or other chemicals when chemically synthesized. In one embodiment, an "isolated" nucleic acid does not include sequences that naturally flank the nucleic acid in the genomic DNA of the organism from which the nucleic acid is derived (i.e., sequences located at the 5' and 3' ends of the nucleic acid).

[0061] By "fragment" is intended a polypeptide consisting of only a portion of the intact full-length polypeptide sequence and structure. Fragments can include C-terminal deletions, N-terminal deletions, and / or internal deletions of the native polypeptide. A fragment of a protein generally contains at least about 5-10 contiguous amino acid residues of the full-length molecule, preferably at least about 15-25 contiguous amino acid residues of the full-length molecule, and most preferably at least about 20-50 or more contiguous amino acid residues of the full-length molecule, or any integer between 5 amino acids and the full-length sequence.

[0062] Certain embodiments of the present disclosure encompass isolated or substantially purified nucleic acid compositions. In the context of the present disclosure, an "isolated" or "purified" DNA or RNA molecule is a DNA or RNA molecule that exists apart from its native environment and is therefore not a product of nature. An isolated DNA or RNA molecule may exist in purified form or in a non-native environment, such as in a transgenic host cell. For example, an "isolated" or "purified" nucleic acid molecule is substantially free of other cellular materials or culture medium when produced by recombinant techniques, or substantially free of chemical precursors or other chemicals when chemically synthesized. In one embodiment, an "isolated" nucleic acid does not include sequences that naturally flank the nucleic acid in the genomic DNA of the organism from which the nucleic acid is derived (i.e., sequences located at the 5' and 3' ends of the nucleic acid).

[0063] "Naturally occurring" is used to describe a composition that can be found in nature, rather than being artificially produced. For example, a nucleotide sequence present in an organism that can be isolated from a natural source and has not been artificially modified in a laboratory is naturally occurring.

[0064] "Regulatory sequence" and "suitable regulatory sequence" refer, respectively, to a nucleotide sequence located upstream of a coding sequence (5' non-coding sequence), within a coding sequence, or downstream of a coding sequence (3' non-coding sequence), and which influence the transcription, RNA processing or stability, or translation of the associated coding sequence. Regulatory sequences include enhancers, promoters, translation leader sequences, introns, and polyadenylation signal sequences. These include both natural and synthetic sequences, as well as sequences which may be a combination of synthetic and natural sequences.

[0065] "5' non-coding sequence" refers to a nucleotide sequence located 5' (upstream) of a coding sequence. It is present in the fully processed mRNA upstream of the start codon and may affect the processing of the primary transcript into mRNA, mRNA stability or translation efficiency. "3' non-coding sequence" refers to a nucleotide sequence located 3' (downstream) of a coding sequence and may include polyadenylation signal sequences and other sequences encoding regulatory signals that may affect mRNA processing or gene expression. Polyadenylation signals are usually characterized as affecting the addition of polyadenylic acid tracts to the 3' end of the mRNA precursor. The term "translation leader sequence" refers to the portion of the DNA sequence of a gene between the promoter and the coding sequence to be transcribed into RNA and is present in the fully processed mRNA upstream (5') of the translation start codon. The translation leader sequence may affect the processing of the primary transcript into mRNA, mRNA stability or translation efficiency.

[0066] "Promoter" refers to a nucleotide sequence, usually upstream (5') of a coding sequence, that directs and / or controls the expression of a coding sequence by providing recognition for RNA polymerase and other factors required for proper transcription. "Promoter" includes a minimal promoter, which is a short DNA sequence consisting of a TATA box and other sequences that help specify the site of transcription initiation, to which regulatory elements are added for control of expression. "Promoter" also refers to a nucleotide sequence that contains a minimal promoter plus regulatory elements that can control the expression of a coding sequence or functional RNA. This type of promoter sequence consists of proximal and more distal upstream elements, the latter elements often referred to as enhancers. Thus, an "enhancer" is a DNA sequence that can stimulate promoter activity and can be a native element of the promoter or a heterologous element inserted to enhance the level or tissue specificity of the promoter. It can operate in both orientations (normal or inverted) and can even function when moved either upstream or downstream from the promoter. Both enhancers and other upstream promoter elements bind sequence-specific DNA-binding proteins that mediate their effects. Promoters may be derived in their entirety from native genes, be composed of different elements derived from different promoters found in nature, or even be composed of synthetic DNA segments. Promoters may also contain DNA sequences involved in the binding of protein factors that control the effectiveness of transcription initiation in response to physiological or developmental conditions. "Constitutive expression" refers to expression using a constitutive promoter. "Conditional expression" and "regulated expression" refer to expression controlled by a regulated promoter.

[0067] "Operably linked" refers to the interaction of nucleic acid sequences on a single nucleic acid fragment such that the function of one of the sequences is affected by the other. For example, a regulatory DNA sequence is said to be "operably linked to" or "interacts with" a DNA sequence that encodes an RNA or polypeptide if the two sequences are positioned such that the regulatory DNA sequence affects expression of a coding DNA sequence (i.e., the coding sequence or functional RNA is under the transcriptional control of a promoter). A coding sequence can be operably linked to a regulatory sequence in a sense or antisense orientation.

[0068] "Expression" refers to the transcription and / or translation of an endogenous gene, a heterologous gene or nucleic acid segment, or a transgene in a cell. In addition, expression refers to the transcription and stable accumulation of sense (mRNA) or functional RNA. Expression may also refer to the production of a protein. The term "altered levels of expression" refers to a level of expression in a cell or organism that is different from that of a normal cell or organism.

[0069] For sequence comparison, typically, one sequence acts as a reference sequence with which test sequence is compared.When using sequence comparison algorithm, test sequence and reference sequence are input into computer, partial sequence coordinates are designated as necessary, and sequence algorithm program parameters are designated.Then, sequence comparison algorithm calculates the percent sequence identity of test sequence to reference sequence based on designated program parameters.

[0070] The following terms are used to describe the sequence relationship between two or more nucleic acids or polynucleotides: (a) "reference sequence", (b) "comparison window", (c) "sequence identity", (d) "percentage of sequence identity", and (e) "substantial identity". As used herein, a "reference sequence" is a defined sequence used as a standard for sequence comparison. A reference sequence may be a subset or the entirety of a specified sequence (e.g., as a segment of a full-length cDNA or gene sequence, or a complete cDNA or gene sequence). As used herein, a "comparison window" refers to a contiguous and specified segment of a polynucleotide sequence, where the polynucleotide sequence in the comparison window may include additions or deletions (i.e., gaps) compared to the reference sequence (which does not include additions or deletions) for optimal alignment of the two sequences. Generally, the comparison window is at least 20 contiguous nucleotides in length, and may be 30, 40, 50, 100, or longer in some cases. Those skilled in the art will appreciate that to avoid high similarity to a reference sequence due to the inclusion of gaps in a polynucleotide sequence, a gap penalty is typically introduced to subtract from the number of matches.

[0071] Methods for alignment of sequences for comparison are well known in the art. Thus, a mathematical algorithm can be used to determine the percent identity between any two sequences. Non-limiting examples of such mathematical algorithms include the Myers and Miller algorithm (Myers and Miller, CABIOS, 4, 11 (1988)); the Smith et al. local homology algorithm (Smith et al., Adv. Appl. Math., 2, 482 (1981)); the Needleman and Wunsch homology alignment algorithm (Needleman and Wunsch, JMB, 48, 443 (1970)); the Pearson and Lipman similarity search method (Pearson and Lipman, Proc. Natl. Acad. Sci. USA, 85, 2444 (1988)); the Karlin and Altschul algorithm (Karlin and Altschul, Proc. Natl. Acad. Sci. USA, 87, 2264 (1990)), modified Karlin and Altschul (Karlin and Altschul, Proc. Natl. Acad. Sci. USA 90, 5873 (1993).

[0072] Computer implementations of these mathematical algorithms can be used to compare sequences to determine sequence identity. Such implementations include, but are not limited to: CLUSTAL (available from Intelligenetics, Mountain View, Calif.) in the PC / Gene program; ALIGN program (Version 2.0), and GAP, BESTFIT, BLAST, FASTA, and TFASTA in the Wisconsin Genetics Software Package, Version 8 (available from Genetics Computer Group (GCG), 575 Science Drive, Madison, Wis., USA). Alignment using these programs can be performed using default parameters. The CLUSTAL program has been described in detail by Higgins et al. (Higgins et al., CABIOS, 5, 151 (1989)); Corpet et al. (Corpet et al., Nucl. Acids Res., 16, 10881 (1988)); Huang et al. (Huang et al., CABIOS, 8, 155 (1992)); and Pearson et al. (Pearson et al., Meth. Mol. Biol., 24, 307 (1994)). The ALIGN program is based on the algorithm of Myers and Miller, supra. The BLAST program of Altschul et al. (Altschul et al., JMB, 215, 403 (1990)) is based on the algorithm of Karlin and Altschul, supra.

[0073] Software for performing BLAST analysis is publicly available through the National Center for Biotechnology Information. The algorithm involves first identifying high-scoring sequence pairs (HSPs) by identifying short words of length "W" in the query sequence that match or meet some positive threshold score T when aligned with words of the same length in a database sequence. "T" is referred to as the neighbor word score threshold. These initial neighbor word hits act as seeds to initiate searches to find longer HSPs containing them. The word hits are then extended in both directions along each sequence for as far as the cumulative alignment score can be increased. Cumulative scores are calculated using the parameters "M" (reward score for a pair of matching residues; always >0) and "N" (penalty score for mismatching residues; always <0) for nucleotide sequences. For amino acid sequences, a scoring matrix is ​​used to calculate the cumulative scores. Extension of the word hits in each direction is stopped when the cumulative alignment score deviates from its maximum achieved value by an amount "X", when the cumulative score falls below zero due to the accumulation of one or more negatively scored residue alignments, or when the end of either sequence is reached.

[0074] In addition to calculating sequence identity percentage, BLAST algorithm also performs statistical analysis of the similarity between two sequences.One measure of similarity provided by BLAST algorithm is the minimum sum probability (P(N)), which provides an indication of the probability that a match between two nucleotide or amino acid sequences occurs by chance.For example, when comparing test nucleic acid sequence with reference nucleic acid sequence, if the minimum sum probability is less than about 0.1, less than about 0.01, or even less than about 0.001, the test nucleic acid sequence is considered to be similar to the reference sequence.

[0075] To obtain gapped alignments for comparison purposes, Gapped BLAST (in BLAST 2.0) can be used. Alternatively, PSI-BLAST (in BLAST 2.0) can be used to perform an iterative search that detects distant relationships between molecules. When using BLAST, Gapped BLAST, or PSI-BLAST, the default parameters of the respective programs (e.g., BLASTN for nucleotide sequences, BLASTX for proteins) can be used. The BLASTN program (for nucleotide sequences) uses as defaults a word length (W) of 11, an expectation (E) of 10, a cutoff of 100, M=5, N=-4, and a comparison of both strands. For amino acid sequences, the BLASTP program uses as defaults a word length (W) of 3, an expectation (E) of 10, and the BLOSUM62 scoring matrix. Alignments can also be performed manually by inspection.

[0076] For purposes of this disclosure, comparison of nucleotide sequences to determine percent sequence identity to the promoter sequences disclosed herein may be performed using the BlastN program (version 1.4.7 or later) with its default parameters, or any equivalent program. By "equivalent program" is intended any sequence comparison program that generates an alignment that has identical nucleotide or amino acid residue matches and identical percent sequence identity for any two sequences in question when compared to the corresponding alignment generated by the program.

[0077] As used herein, "sequence identity" or "identity" in the context of two nucleic acid or polypeptide sequences refers to a certain percentage of residues in the two sequences that are the same when aligned for maximum correspondence over a specified comparison window, as measured by a sequence comparison algorithm or visual inspection. When percentage sequence identity is used in the context of proteins, it is recognized that non-identical residue positions often differ by conservative amino acid substitutions (amino acid residues are replaced with other amino acid residues that have similar chemical properties (e.g., charge or hydrophobicity), and therefore do not change the functional properties of the molecule). When sequences differ by conservative substitutions, the percentage sequence identity may be adjusted upwards to correct for the conservative nature of the substitution. Sequences that differ by such conservative substitutions are said to have "sequence similarity" or "similarity". Means for making this adjustment are well known to those skilled in the art. Typically, this involves scoring conservative substitutions as partial rather than complete mismatches, thereby increasing the percentage of sequence identity. Thus, for example, where identical amino acids are given a score of 1 and non-conservative substitutions are given a score of 0, conservative substitutions are given a score between 0 and 1. Scoring of conservative substitutions is calculated, for example, as performed in the program PC / GENE (Intelligenetics, Mountain View, Calif.).

[0078] As used herein, "percentage of sequence identity" refers to a value determined by comparing two optimally aligned sequences over a comparison window, where the portion of the polynucleotide sequence in the comparison window may contain additions or deletions (i.e., gaps) compared to the reference sequence (which does not contain additions or deletions) due to optimal alignment of the two sequences. The percentage is calculated by determining the number of positions in both sequences where the same nucleic acid base or amino acid residue exists, determining the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window, and multiplying the result by 100 to determine the percentage of sequence identity.

[0079] The term "substantial identity" of a polynucleotide sequence means that the polynucleotide comprises a sequence having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, or 94%, or even at least 95%, 96%, 97%, 98%, or 99% sequence identity with respect to a reference sequence using one of the alignment programs described with standard parameters. Those skilled in the art will recognize that these values ​​can be appropriately adjusted to determine the corresponding identity of proteins encoded by two nucleotide sequences by taking into account codon degeneracy, amino acid similarity, reading frame arrangement, and the like. Substantial identity of amino acid sequences for these purposes usually means at least 70%, 80%, 90%, or even at least 95% sequence identity.

[0080] The term "substantial identity" in the context of peptides indicates that the peptide comprises a sequence having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, or 94%, or even 95%, 96%, 97%, 98% or 99% sequence identity to the reference sequence over a specified comparison window. In certain embodiments, optimal alignment is performed using the Needleman and Wunsch homology alignment algorithm (Needleman and Wunsch, JMB, 48, 443 (1970)). An indication that two peptide sequences are substantially identical is that one peptide is immunologically reactive with an antibody raised against the second peptide. Thus, for example, one peptide is substantially identical to a second peptide where the two peptides differ only by conservative substitutions. Thus, the present disclosure also provides nucleic acid molecules and peptides that are substantially identical to the nucleic acid molecules and peptides presented herein.

[0081] Another indication that nucleotide sequences are substantially identical is if two molecules hybridize to each other under stringent conditions. Nucleic acid hybridization is discussed in more detail below.

[0082] Oligonucleotide Probes As used herein, "primer", "probe" and "oligonucleotide" are used interchangeably. The term "nucleic acid probe" or "nucleic acid specific probe" refers to a nucleic acid sequence having at least about 80%, such as at least about 90%, such as at least about 95% contiguous sequence identity or homology to a nucleic acid sequence encoding a targeting sequence of interest. A probe (or oligonucleotide or primer) of the present disclosure is at least about 8 nucleotides in length (e.g., at least about 8-50 nucleotides in length, such as at least about 10-40, such as at least about 15-35 nucleotides in length). An oligonucleotide probe or primer of the present disclosure may include at least about 8 nucleotides 3' of the oligonucleotide having at least about 80%, such as at least about 85%, such as at least about 90% contiguous identity to a targeting sequence of interest.

[0083] Primer pairs are useful for determining the nucleotide sequence of a particular SNP using PCR. A pair of single-stranded DNA primers can be annealed to sequences within or surrounding the SNP to prime the amplification DNA synthesis of the SNP itself.

[0084] The first step of the process involves contacting a physiological sample obtained from a patient, the sample containing nucleic acid, with an oligonucleotide probe to form hybridized DNA. Oligonucleotide probes useful in the methods of the disclosure can be any probe of about 4 or 6 bases up to about 80 or 100 bases or more. In one embodiment of the disclosure, the probe is about 10 to about 20 bases.

[0085] The primer itself can be synthesized using techniques well known in the art. Generally, primers can be produced using oligonucleotide synthesizing machines, which are commercially available.

[0086] The primer or probe of the present disclosure can be labeled using techniques known to those skilled in the art. For example, the label used in the disclosed assay can be a primary label (where the label contains an element that is directly detected) or a secondary label (where the label that is detected is bound to the primary label, e.g., as is common in immunological labeling). A general overview of labels (also called "tags"), tagging or labeling procedures, and detection of labels can be found in Polak and Van Noorden (1997) Introduction to Immunocytochemistry, second edition, Springer Verlag, NY, and Haugland (1996) Handbook of Fluorescent Probes and Research Chemicals, a combined handbook and catalogue Published by Molecular Probes, Inc., Eugene, Oreg. Primary and secondary labels can include non-detectable elements as well as detectable elements. Primary and secondary labels useful in the present disclosure include spectral labels, e.g., fluorescent dyes (e.g., fluorescein and derivatives, such as fluorescein isothiocyanate (FITC) and Oregon Green™, rhodamine and derivatives (e.g., Texas red, tetramethylrhodamine isothiocyanate (TRITC), etc.), digoxigenin, biotin, phycoerythrin, AMCA, CyDyes™, etc.), radiolabels (e.g., 3 H, 125 I, 35 S, 14 C. 32 P, 33P), enzymes (e.g., horseradish peroxidase, alkaline phosphatase), spectral chromogenic labels, such as colloidal gold or colored glass or plastic (e.g., polystyrene, polypropylene, latex) beads. The labels may be directly or indirectly coupled to components of the detection assay (e.g., labeled nucleic acid) according to methods well known in the art. As indicated above, a wide variety of labels may be used, with the choice of label depending on the required sensitivity, ease of conjugation with the compound, stability requirements, available instrumentation, and disposal regulations.

[0087] Generally, the detector that monitors probe-substrate nucleic acid hybridization is adapted to the specific label used.Typical detectors include spectrophotometers, phototubes and photodiodes, microscopes, scintillation counters, cameras, films, etc., and combinations thereof.Examples of suitable detectors are widely available from various commercial sources known to those skilled in the art.Usually, the optical image of the substrate containing the bound labeled nucleic acid is digitized for subsequent computer analysis.

[0088] Preferred labels are: (1) chemiluminescence (using horseradish peroxidase and / or alkaline phosphatase with a substrate that produces photons as a breakdown product) (kits are available, e.g., from Molecular Probes, Amersham, Boehringer-Mannheim, and Life Technologies / Gibco BRL); (2) color generation (using both horseradish peroxidase and / or alkaline phosphatase with a substrate that produces a colored precipitate) (kits are available, e.g., from Life Technologies / Gibco BRL). BRL, and Boehringer-Mannheim; (3) hemifluorescence (e.g., using alkaline phosphatase and its substrate AttoPhos (Amersham) or other substrates that produce a fluorescent product); (4) fluorescence (e.g., using Cy-5 (Amersham), fluorescein, and other fluorescent labels); (5) radioactivity (using kinase enzymes or other end-labeling approaches, nick translation, random priming, or PCR to incorporate radioactive molecules into the labeled nucleic acid). Other methods for labeling and detection will be readily apparent to those of skill in the art.

[0089] Fluorescent labels can be used and have the advantage that few handling precautions are required and are adaptable to high throughput imaging techniques (optical analysis including digitization of images for analysis in integrated systems including computers). Preferred labels are typically characterized by one or more of the following: high sensitivity, high stability, low background, low environmental sensitivity, and high specificity in labeling. Fluorescent moieties incorporated into the labels of the present disclosure are commonly known, including Texas red, dixogenin, biotin, 1- and 2-aminonaphthalenes, p,p'-diaminostilbene, pyrene, quaternary phenanthridine salts, 9-aminoacridine, p,p'-diaminobenzophenone imine, anthracene, oxacarbocyanine, merocyanine, 3-aminoequilenin, perylene, bis-benzoxazole, bis-p-oxazolylbenzene, 1,2-benzophenazine, retinol, bis-3-aminopyridinium salts, hellebrigenin, tetracycline, sterophenol, benzimidazolylphenylamine, 2-oxo-3-chromene, indole, xanthene, 7-hydroxycoumarin, phenoxazine, salicylic acid, strophanthidin, porphyrin, triarylmethane, flavins and many others.Many fluorescent labels are commercially available from SIGMA Chemical Company (Saint Louis, MO), Molecular Probes, R&D systems (Minneapolis, MN), Pharmacia LKB Biotechnology (Piscataway, NJ), CLONTECH Laboratories, Inc. (Palo Alto, CA), Chem Genes Corp., Aldrich Chemical Company (Milwaukee, WI), Glen Research, Inc., GIBCO BRL Life Technologies, Inc. (Gaithersberg, MD), Fluka ChemicaBiochemika Analytika (Fluka Chemie AG, Buchs, Switzerland), and Applied Biosystems™ (Foster City, CA), as well as many other commercial sources known to those of skill in the art.

[0090] Means for detecting and quantifying the label are well known to those skilled in the art.Thus, for example, if the label is a radioactive label, detection means include a scintillation counter or photographic film, as in autoradiography.If the label is optically detectable, typical detectors include microscopes, cameras, phototubes and photodiodes, and many other detection systems that are widely available.

[0091] Oligonucleotide probes having any of a wide variety of base sequences may be prepared according to techniques well known in the art. Suitable bases for preparing oligonucleotide probes include naturally occurring nucleotide bases, such as adenine, cytosine, guanine, uracil, and thymine; and non-naturally occurring or "synthetic" nucleotide bases, such as 7-deaza-guanine, 8-oxo-guanine, 6-mercaptoguanine, 4-acetylcytidine, 5-(carboxyhydroxyethyl)uridine, 2'-O-methylcytidine, 5-carboxymethylamino-methyl-2-thiolysine (thiol). oridine), 5-carboxymethylaminomethyluridine, dihydrouridine, 2'-O-methylpseudouridine, β,D-galactosylqueosine, 2'-O-methylguanosine, inosine, N6-isopentenyladenosine, 1-methyladenosine, 1-methylpseudouridine, 1-methylguanosine, 1-methylinosine, 2,2-dimethylguanosine, 2-methyladenosine, 2-methylguanosine, 3-methyl Cytidine, 5-methylcytidine, N6-methyladenosine, 7-methylguanosine, 5-methylaminomethyluridine, 5-methoxyaminomethyl-2-thiouridine, β,D-mannosylqueosine, 5-methoxycarbonylmethyluridine, 5-methoxyuridine, 2-methylthio-N6-isopentenyladenosine, N-((9-β-D-ribofuranosyl-2-methylthiopurine-6-yl)carbamoyl)threonine, N-(( 9-β-D-ribofuranosylpurin-6-yl)N-methyl-carbamoyl)threonine, uridine-5-oxyacetic acid methyl ester, uridine-5-oxyacetic acid, wybutoxosine, pseudouridine, queosine, 2-thiocytidine, 5-methyl-2-thiouridine, 2-thiouridine, 2-thiouridine, 5-methyluridine, N-((9-β-D-ribofuranosylpurin-6-yl)carbamoyl)threonine, 2′-O-methyl-5-methyluridine, 2′-O-methyluridine, wybutoxine, and 3-(3-amino-3-carboxypropyl)uridine.Any oligonucleotide backbone may be employed, including DNA, RNA (although RNA is less preferred than DNA), modified sugars (e.g., carbocycles, and sugars containing 2' substitutions such as fluoro and methoxy). The oligonucleotide may be an oligonucleotide in which at least one or all of the internucleotide bridging phosphate residues are modified phosphonates, such as methyl phosphonates, methyl phosphorothioates, phosphoroinorpholidates, phosphoropiperazidates, and phosplioramidates (e.g., every other internucleotide bridging phosphate residue may be modified as described). The oligonucleotide may be a "peptide nucleic acid" as described in Nielsen et al., Science, 254:1497-1500 (1991).

[0092] As used herein, a "single base pair extension probe" is a nucleic acid that selectively recognizes a single nucleotide polymorphism (i.e., either the A or G of an A / G polymorphism). Generally, these probes take the form of a DNA primer (e.g., as in a PCR primer) that has been modified such that incorporation of the primer releases a fluorophore. An example of this is the Taqman® probe, which uses the 5' exonuclease activity of the Taq polymerase enzyme to measure the amount of a target sequence in a sample. TaqMan® probes consist of 18-22 bp oligonucleotide probes that are labeled with a reporter fluorophore at the 5' end and a quencher fluorophore at the 3' end. Incorporation of the probe molecule into the PCR strand (which occurs because the probe set is contained in a mixture of PCR primers) releases the reporter fluorophore from the influence of the quencher. The primer must be capable of recognizing the target binding site. Some primer extension probes can be "activated" directly by DNA polymerase without a complete PCR extension cycle.

[0093] The only requirement is that the oligonucleotide probe should have a sequence that is capable of binding, at least in part, to a portion of the DNA sample whose sequence is known. The nucleic acid probes provided by this disclosure are useful for a number of purposes.

[0094] Methods for detecting nucleic acids A. Amplification According to the method of the present disclosure, the amplification of DNA present in physiological sample can be carried out by any means known in the art.Examples of suitable amplification techniques include, but are not limited to, polymerase chain reaction (for RNA amplification, including reverse transcriptase polymerase chain reaction), ligase chain reaction, strand displacement amplification, transcription-based amplification, self-sustained sequence replication (or "3SR"), Qbeta replicase system, nucleic acid sequence-based amplification (or "NASBA"), repair chain reaction (or "RCR") and boomerang DNA amplification (or "BDA").

[0095] The bases incorporated into the amplification products may be natural or modified bases (modified before or after amplification), and the bases may be selected to optimize the subsequent electrochemical detection step.

[0096] Polymerase chain reaction (PCR) may be carried out according to known techniques.See, for example, U.S. Patent Nos. 4,683,195; 4,683,202; 4,800,159; and 4,965,188.Generally, PCR involves first treating a nucleic acid sample (e.g., in the presence of a thermostable DNA polymerase) with one oligonucleotide primer for each strand of the specific sequence to be detected under hybridization conditions such that the extension products of each primer that are complementary to each strand of nucleic acid are synthesized, and the primers are sufficiently complementary to hybridize to each strand of the specific sequence so that the extension products synthesized from each primer, when separated from their complements, can serve as templates for the synthesis of the extension products of the other primer; then treating the sample under denaturing conditions to separate the primer extension products from their templates if the sequence to be detected is present.These steps are repeated periodically until a desired degree of amplification is obtained. The detection of the amplified sequence may be carried out by adding an oligonucleotide probe (such as the oligonucleotide probe of the present disclosure) that can hybridize to the reaction product to the reaction product, the probe having a detectable label; then detecting the label according to known techniques.Various labels that can be incorporated into or functionally linked to nucleic acid, such as radioactive labels, enzyme labels, and fluorescent labels, are well known in the art.When the nucleic acid to be amplified is RNA, amplification may be carried out by initial conversion to DNA by reverse transcriptase according to known techniques.

[0097] Strand displacement amplification (SDA) may be performed according to known techniques. For example, SDA may be performed with a single amplification primer or a pair of amplification primers, with exponential amplification being achieved with the latter. In general, SDA amplification primers include, from 5' to 3', flanking sequences (whose DNA sequence is not critical), a restriction site for the restriction enzyme employed in the reaction, and an oligonucleotide sequence (e.g., an oligonucleotide probe of the present disclosure) that hybridizes to the target sequence to be amplified and / or detected. The flanking sequences, which serve to facilitate binding of the restriction enzyme to its recognition site and provide a DNA polymerase priming site after nicking the restriction site, are in one embodiment about 15-20 nucleotides in length. The restriction site is functional in the SDA reaction. The oligonucleotide probe portion is in one embodiment of the present disclosure about 13-15 nucleotides in length.

[0098] Ligase chain reaction (LCR) can also be carried out according to known techniques. In general, the reaction is carried out with two pairs of oligonucleotide probes: one pair binds to one strand of the sequence to be detected; the other pair binds to the other strand of the sequence to be detected. Each pair together completely overlaps with its corresponding strand. The reaction is carried out by first denaturing (e.g., separating) the strand of the sequence to be detected, then reacting it with the two pairs of oligonucleotide probes in the presence of a thermostable ligase such that each pair of oligonucleotide probes is ligated together, then separating the reaction products, and then repeating the process cyclically until the sequence is amplified to the desired extent. Detection can then be carried out in the same manner as described above for PCR.

[0099] According to the method of the present disclosure, the specific SNP of this locus is detected.The techniques that are useful in the method of the present disclosure include, but are not limited to, direct DNA sequencing, PFGE analysis, allele-specific oligonucleotide (ASO), dot blot analysis and denaturing gradient gel electrophoresis, and are well known to those skilled in the art.

[0100] There are several methods that can be used to detect DNA sequence variations. Direct DNA sequencing, either manual or automated fluorescent sequencing, can detect sequence variations. Another approach is the single-strand conformation polymorphism assay (SSCA). This method does not detect all sequence changes, but can be optimized to detect the largest DNA sequence variations, especially when the DNA fragment size is larger than 200 bp. Although reduced detection sensitivity is a drawback, the increased throughput possible with SSCA makes it an attractive and viable alternative to direct sequencing for mutation detection on a research basis. Fragments with altered mobility on the SSCA gel are then sequenced to determine the exact nature of the DNA sequence variation. Other approaches based on the detection of mismatches between two complementary DNA strands include clamped denaturing gel electrophoresis (CDGE), heteroduplex analysis (HA) and chemical mismatch cleavage (CMC). Once the mutation is known, a large number of other samples can be easily screened for that same mutation using allele-specific detection approaches such as allele-specific oligonucleotide (ASO) hybridization. Such techniques can utilize gold nanoparticle-labeled probes to produce visible color results.

[0101] Detection of SNPs can be achieved by sequencing the desired target region using techniques well known in the art. Alternatively, gene sequences can be directly amplified from genomic DNA preparations from patient tissues using known techniques.The DNA sequence of the amplified sequence can then be determined.

[0102] There are six well-known methods for indirect but more complete tests to confirm the presence of mutant alleles: 1) single-strand conformation analysis (SSCA); 2) denaturing gradient gel electrophoresis (DGGE); 3) RNase protection assay; 4) allele-specific oligonucleotide (ASO); 5) use of a protein that recognizes nucleotide mismatches, such as the E. coli mutS protein; and 6) allele-specific PCR. For allele-specific PCR, primers are used that hybridize to a specific DNM1 mutation at their 3' ends. If that specific mutation is not present, no amplification product is observed. Amplification refractory mutation system (ARMS) can also be used. Gene insertions and deletions can also be detected by cloning, sequencing and amplification. In addition, restriction fragment length polymorphism (RFLP) probes for the gene or surrounding marker genes can be used to score allele changes or insertions in the polymorphic fragment. Other techniques for detecting insertions and deletions can be used as known in the art.

[0103] The first three methods (SSCA, DGGE and RNase protection assay) result in the appearance of new electrophoretic bands. SSCA detects differentially migrating bands due to differences in single-stranded intramolecular base pairing caused by sequence changes. RNase protection involves cleavage of mutant polynucleotides into two or more smaller fragments. DGGE uses denaturing gradient gels to detect differences in migration rates of mutant sequences compared to wild-type sequences. In allele-specific oligonucleotide assays, oligonucleotides are designed to detect specific sequences, and the assay is performed by detecting the presence or absence of a hybridization signal. In the mutS assay, the protein binds only to sequences that contain nucleotide mismatches in the heteroduplex between the mutant and wild-type sequences.

[0104] A mismatch according to the present disclosure is a hybridized nucleic acid duplex where the two strands are not 100% complementary. The lack of complete homology can be due to deletions, insertions, inversions or substitutions. Mismatch detection can be used to detect point mutations in a gene or its mRNA product. These techniques are less sensitive than sequencing, but they are easier to perform on a large number of samples. An example of a mismatch cleavage technique is the RNase protection method. A riboprobe and either mRNA or DNA isolated from tumor tissue are annealed (hybridized) together and then digested with RNase A enzyme, which can detect some mismatches in the duplex RNA structure. When a mismatch is detected by RNase A, RNase A cleaves at the site of the mismatch. Thus, when the annealed RNA preparation is separated on an electrophoretic gel matrix, an RNA product smaller than the full-length duplex RNA for the riboprobe and mRNA or DNA will be seen if the mismatch was detected and cleaved by RNase A. The riboprobe does not have to be the full length DNM1 mRNA or gene, but can be a segment of either. If the riboprobe contains only a segment of the DNM1 mRNA or gene, it is desirable to use a large number of these probes to screen the entire mRNA sequence for mismatches.

[0105] Similarly, DNA probes can be used to detect mismatches through enzymatic or chemical cleavage.Alternatively, mismatches can be detected by the shift in electrophoretic mobility of mismatched duplexes relative to matched duplexes.Using either riboprobes or DNA probes, the mRNA or DNA of cells that may contain mutations can be amplified using PCR before hybridization.

[0106] B. Hybridization The phrase "specifically hybridizes to" refers to a molecule that binds, duplexes, or hybridizes under stringent conditions only to a particular nucleotide sequence when present in a complex mixture (e.g., whole cell) DNA or RNA. "Substantially binds" refers to complementary hybridization between the probe and target nucleic acids, and encompasses minor mismatches that can be accommodated by reducing the stringency of the hybridization medium to achieve the desired detection of the target nucleic acid sequence.

[0107] Generally, stringent conditions are determined by the thermal melting point (T m ) is selected to be about 5°C lower than the desired temperature. However, stringent conditions encompass temperatures ranging from about 1°C to about 20°C, depending on the desired degree of stringency as limited elsewhere herein. Nucleic acids that do not hybridize to each other under stringent conditions are still substantially identical if the polypeptides they encode are substantially identical. This can occur, for example, when copies of a nucleic acid are generated using the maximum codon degeneracy permitted by the genetic code. One indication that two nucleic acid sequences are substantially identical is when a polypeptide encoded by a first nucleic acid is immunologically cross-reactive with a polypeptide encoded by a second nucleic acid.

[0108] "Stringent conditions" are those that employ (1) low ionic strength and high temperature for washing, e.g., 0.015M NaCl / 0.0015M sodium citrate (SSC); 0.1% sodium lauryl sulfate (SDS) at 50°C, or (2) a denaturing agent such as formamide, e.g., 50% formamide with 0.1% bovine serum albumin / 0.1% Ficoll / 0.1% polyvinylpyrrolidone / 50 mM sodium phosphate buffer (pH 6.5) containing 750 mM NaCl, 75 mM sodium citrate at 42°C during hybridization. Another example is the use of 50% formamide, 5×SSC (0.75 M NaCl, 0.075 M sodium citrate), 50 mM sodium phosphate (pH 6.8), 0.1% sodium pyrophosphate, 5×Denhardt's solution, sonicated salmon sperm DNA (50 μg / ml), 0.1% SDS, and 10% dextran sulfate at 42° C., with a wash in 0.2×SSC and 0.1% SDS at 42° C. Other examples of stringent conditions are well known in the art.

[0109] "Stringent hybridization conditions" and "stringent hybridization wash conditions" in the context of nucleic acid hybridization experiments such as Southern and Northern hybridization are sequence-dependent and are different under different environmental parameters. Longer sequences hybridize specifically at higher temperatures. The thermal melting point (Tm) is the temperature (under defined ionic strength and pH) at which 50% of the target sequence hybridizes to a perfectly matched probe. Specificity is typically correlated with post-hybridization washes, with the important factors being the ionic strength and temperature of the final wash solution. For DNA-DNA hybrids, T m can be approximated by the equation of Meinkoth and Wahl (1984); T m81.5°C + 16.6 (log M) + 0.41 (% GC) - 0.61 (% form) - 500 / L; where M is the molar concentration of monovalent cations, % GC is the percentage of guanosine and cytosine nucleotides in DNA, % form is the percentage of formamide in the hybridization solution, and L is the hybrid length in base pairs. m is reduced by approximately 1°C for every 1% mismatch; therefore, T m , hybridization, and / or washing conditions can be adjusted to hybridize to sequences of desired identity. For example, if one is looking for sequences with >90% identity, T m Generally, stringent conditions are those that meet the T β for a particular sequence and its complement at a defined ionic strength and pH. m However, stringent conditions are chosen to be approximately 5°C lower than T m Moderately stringent conditions can utilize hybridization and / or washing temperatures that are 1, 2, 3, or 4° C. lower than T m Low stringency conditions can be used for hybridization and / or washing at temperatures 6, 7, 8, 9, or 10° C. lower than the T m Hybridization and / or washing at temperatures 11, 12, 13, 14, 15, or 20° C. lower than the desired temperature can be utilized. Using the equations, hybridization and washing compositions, and desired temperature, one of skill in the art will appreciate that variations in the stringency of hybridization and / or washing solutions are inherently described. If the desired degree of mismatching results in temperatures below 45° C. (aqueous solution) or 32° C. (formamide solution), the SSC concentration is increased so that higher temperatures can be used. In general, highly stringent hybridization and washing conditions are those that provide a T for a particular sequence at a defined ionic strength and pH. m is selected to be approximately 5°C lower than

[0110] An example of a highly stringent wash condition is 0.15M NaCl at 72°C for about 15 minutes. An example of a stringent wash condition is a 0.2xSSC wash at 65°C for 15 minutes. Often, a high stringency wash is preceded by a low stringency wash to remove background probe signal. For example, an example of a medium stringency wash for duplexes of more than 100 nucleotides is 1xSSC at 45°C for 15 minutes. For short nucleotide sequences (e.g., about 10-50 nucleotides), stringent conditions typically include salt concentrations of less than about 1.5M, less than about 0.01-1.0M Na ion concentrations (or other salts) at pH 7.0-8.3, and temperatures of at least about 30°C and at least about 60°C for long probes (e.g., >50 nucleotides). Stringent conditions may also be achieved by the addition of destabilizing agents such as formamide. Generally, a signal-to-noise ratio of 2x (or higher) than that observed for an unrelated probe in a particular hybridization assay indicates detection of specific hybridization.Nucleic acids that do not hybridize to each other under stringent conditions are still substantially identical if the proteins they code for are substantially identical.This occurs, for example, when copies of nucleic acids are generated using the maximum codon degeneracy permitted by the genetic code.

[0111] Highly stringent conditions are defined as conditions where the T mAn example of stringent conditions for hybridization of complementary nucleic acids having more than 100 complementary residues on a filter in a Southern or Northern blot is 50% formamide, e.g., hybridization in 50% formamide, 1 M NaCl, 1% SDS at 37° C., and washing in 0.1×SSC at 60-65° C. Exemplary low stringency conditions include hybridization at 37° C. using a buffer solution of 30-35% formamide, 1 M NaCl, 1% SDS (sodium dodecyl sulfate), and washing in 1×-2×SSC (20×SSC=3.0 M NaCl / 0.3 M trisodium citrate) at 50-55° C. Exemplary moderate stringency conditions include hybridization in 40-45% formamide, 1.0 M NaCl, 1% SDS at 37°C, and washing in 0.5x-1x SSC at 55-60°C.

[0112] "Northern analysis" or "Northern blotting" is a method used to identify RNA sequences that hybridize to a known probe, such as an oligonucleotide, a DNA fragment, a cDNA or fragment thereof, or an RNA fragment. The probe is attached by biotinylation or using an enzyme. 32 The RNA to be analyzed can be typically separated by electrophoresis on an agarose or polyacrylamide gel, transferred to a nitrocellulose, nylon, or other suitable membrane, and hybridized with a probe using standard techniques well known in the art.

[0113] The nucleic acid sample may be contacted with the oligonucleotide probe by any suitable method known to those skilled in the art.For example, the DNA sample may be solubilized in a solution, and the oligonucleotide probe may be contacted with the DNA sample by solubilizing the oligonucleotide probe in a solution together with the DNA sample under conditions that allow hybridization.Suitable conditions are well known to those skilled in the art.Alternatively, the DNA sample may be solubilized in a solution together with the oligonucleotide probe immobilized on a solid support, and the DNA sample may be contacted with the oligonucleotide probe by immersing the solid support with the immobilized oligonucleotide probe in a solution containing the DNA sample.

[0114] The term "substrate" refers to any solid support to which a probe can be attached. Substrate materials may be covalently or otherwise modified with coating agents or functional groups that facilitate attachment of probes. Suitable substrate materials include polymers, glass, semiconductors, paper, metals, gels, and hydrogels, among others. Substrates may have any physical shape or size, e.g., plates, strips, or microparticles. The term "spot" refers to a discrete location on a substrate to which a probe of known sequence is attached. A spot may be an area on a planar substrate, which may be, for example, a microparticle that is distinguishable from other microparticles. The term "bound" means attached to a solid substrate. A spot is "bound" to a solid substrate if it is attached to a specific location on the substrate for the purpose of a screening assay.

[0115] In certain embodiments of the present disclosure, the substrate is a polymer, glass, semiconductor, paper, metal, gel, or hydrogel. In certain embodiments of the present disclosure, the kit can further include a solid substrate and at least one control probe, wherein the at least one control probe is bound to a distinct spot on the substrate.

[0116] In certain embodiments of the present disclosure, the solid substrate is a microarray. "Array" or "microarray" are used synonymously herein and refer to a plurality of probes attached to one or more identifiable spots on a substrate. A microarray may comprise a single substrate or multiple substrates, such as multiple beads or microspheres. A "copy" of a microarray contains the same type and arrangement of probes.

[0117] Methods for detecting coronary heart disease - Patents.com The present disclosure provides a method for determining whether a subject is likely to have CVD by determining the methylation status of CpG dinucleotide repeat or CpG dinucleotide repeat motif regions using bisulfite-treated DNA, where the methylation status of the CpG dinucleotide is associated with CVD. In certain embodiments, the method determines the methylation status of a plurality of (e.g., any integer between 1 and 10,000, e.g., at least 100) CpG dinucleotide repeat motif regions.

[0118] Various techniques and reagents are useful in the method of the present disclosure.In one embodiment of the present disclosure, blood sample or blood-derived sample, such as plasma, circulating, peripheral lymphocytes, etc., is assayed for the presence of one or more SNPs and / or the methylation status of one or more CpG dinucleotides.Biological sample can also be saliva.Typically, biological sample containing nucleic acid is prepared and tested.

[0119] As used herein, the term "healthy" means that a subject does not exhibit a particular pathological condition and is not probabilistically predisposed to a particular pathological condition.

[0120] In certain embodiments, the present disclosure provides a method for detecting whether a subject has a predisposition to or has coronary heart disease.Such a method typically includes the following steps: prepare a biological sample from a subject; contact DNA from the biological sample with bisulfite under alkaline conditions; contact the bisulfite-treated DNA with at least one first oligonucleotide probe of at least 8 nucleotides in length that is complementary to a sequence that comprises CpG dinucleotide, and the at least one first oligonucleotide probe detects either unmethylated CpG dinucleotide or methylated CpG dinucleotide; and detects either unmethylated CpG dinucleotide or methylated CpG dinucleotide, and the methylation of CpG dinucleotide is associated with coronary heart disease.Such a method can further include determining the genotype of single nucleotide polymorphism (SNP) (e.g., rs347027).

[0121] In certain embodiments, the method further comprises contacting the bisulfite-treated DNA with at least one second oligonucleotide probe of at least 8 nucleotides in length that is complementary to a sequence comprising a CpG dinucleotide, wherein the at least one second oligonucleotide probe detects either the unmethylated CpG dinucleotide or the methylated CpG dinucleotide, neither of which is detected by the at least one first oligonucleotide probe.

[0122] In certain embodiments, the method further comprises determining the ratio of methylated CpG dinucleotides to unmethylated CpG dinucleotides.In certain embodiments, the method can comprise an amplification step after the contacting step.In certain embodiments, the method can comprise a sequencing step after the contacting step.

[0123] In certain embodiments, a method is provided for measuring the presence of biomarkers in biological samples from patients.Such method can include the steps of contacting DNA from biological samples with bisulfite under alkaline conditions; and contacting the bisulfite-treated DNA with at least one first oligonucleotide probe, which is at least 8 nucleotides long and is complementary to the sequence that contains CpG dinucleotide, and at least one first oligonucleotide probe detects either the CpG dinucleotide that is unmethylated or the CpG dinucleotide that is methylated.Such method can be used to predict whether a patient has coronary heart disease or is likely to develop coronary heart disease.

[0124] In certain embodiments, a method for predicting the presence of biomarkers associated with coronary heart disease (CHD) in a biological sample from a patient is provided.Such a method typically includes the steps of preparing a first aliquot from the biological sample and contacting the DNA from the first aliquot with bisulfite under alkaline conditions. Such methods also typically include the steps of preparing a second aliquot from the biological sample; and contacting the bisulfite-treated first aliquot and the second aliquot, (i) with a first oligonucleotide probe of at least 8 nucleotides in length that is complementary to a sequence comprising a CpG dinucleotide at position 92203667 of chromosome 1 located in the transforming growth factor beta receptor III (TGFBR3) gene, and contacting the second aliquot with a nucleic acid primer of at least 8 nucleotides in length that is complementary to SNP rs347027; (ii) with a first oligonucleotide probe of at least 8 nucleotides in length that is complementary to a sequence comprising a CpG dinucleotide at position 38364951 located in an intergenic region of chromosome 15, and contacting the second aliquot with a nucleic acid primer of at least 8 nucleotides in length that is complementary to SNP rs4937276; (iii) with a first aliquot containing coenzyme Q2. (iv) contacting the first aliquot with a first oligonucleotide probe at least 8 nucleotides in length that is complementary to a sequence comprising a CpG dinucleotide at position 84206068 on chromosome 4 in the 4-hydroxybenzoate poly-prenyltransferase (COQ2) gene, and contacting the second aliquot with a nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs17355663; (iv) contacting the first aliquot with a first oligonucleotide probe at least 8 nucleotides in length that is complementary to a sequence comprising a CpG dinucleotide at position 26146070 on chromosome 16 in the heparan sulfate 3-O-sulfotransferase 4 (HS3ST4) gene, and contacting the second aliquot with a nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs235807;(v) contacting the first aliquot with a first oligonucleotide probe at least 8 nucleotides in length that is complementary to a sequence that includes a CpG dinucleotide at position 91171013 of an intergenic region of chromosome 1, and contacting the second aliquot with a nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs11579814; (vi) contacting the first aliquot with a first oligonucleotide probe at least 8 nucleotides in length that is complementary to a sequence that includes a CpG dinucleotide at position 39491936 of chromosome 1 within the NADH dehydrogenase (ubiquinone) Fe-S protein 5 (NDUFS5) gene, and contacting the second aliquot with a nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs2275187; (vii) (viii) contacting the first aliquot with a first oligonucleotide probe at least 8 nucleotides in length that is complementary to a sequence comprising a CpG dinucleotide at position 205475130 of chromosome 1 located in the cyclin-dependent kinase 18 (CDK18) gene and contacting the second aliquot with a nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs4951158; and / or (ix) contacting the first aliquot with a nucleic acid primer at least 8 nucleotides in length that is complementary to SNP rs4951158, which is complementary to SNP rs4951158, which is complementary to SNP rs4951158, which is complementary to SNP rs4951158; contacting the second aliquot with a first oligonucleotide probe at least 8 nucleotides in length that is complementary to a sequence that includes a CpG dinucleotide at position 130614013 on chromosome 3 located in the ATP2C1 gene, and contacting the second aliquot with a nucleic acid primer at least 8 nucleotides in length that is complementary to rs925613;

[0125] In certain aspects, the disclosure provides a method for detecting one or more copies of the G allele at rs347027 and the methylation status of cg13078798 on a nucleic acid sample from a subject at risk for coronary heart disease (CHD), the method comprising: (a) performing a genotyping assay on the nucleic acid sample of the human subject to detect the presence of one or more copies of the G allele of the rs347027 polymorphism; and (b) performing a methylation assessment of cg13078798 on the nucleic acid sample of the human to detect the methylation status to determine whether cg13078798 is unmethylated.

[0126] In such methods, a CpG dinucleotide at position 92203667 on chromosome 1, or a G at position 1618766 on chromosome 1, or a SNP polymorphism at rs4937276, rs17355663, rs235807, rs11579814, rs2275187, rs4336803, rs4951158, and / or rs925613, in conjunction with methylation at any of positions cg20636912, cg16947947, cg05916059, cg04567738, cg16603713, cg05709437, cg12081870, and / or cg18070470, located within the TGFBR3 gene, is associated with CHD.

[0127] Kits for detecting coronary heart disease In a further aspect of the present disclosure, there are provided articles of manufacture and kits that contain, for example, probes, oligonucleotides or antibodies that can be used for the above applications. The articles of manufacture include a container with a label. Suitable containers include, for example, bottles, vials, and test tubes. The containers can be formed from a variety of materials, such as glass or plastic. The containers contain compositions that include one or more agents that are effective for carrying out the methods described herein. The labels on the containers indicate that the compositions can be used for a particular application. The kits of the present disclosure will typically include the above containers and one or more other containers that contain materials that are desirable from a commercial and user standpoint, including buffers, diluents, filters, and package inserts with instructions for use.

[0128] In certain embodiments, the present disclosure provides a kit for determining the methylation status of at least one CpG dinucleotide and the presence of at least one single nucleotide polymorphism (SNP). In certain embodiments, the kit described herein may contain a number of primers, which may be any integer between 1 and 10,000, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, . . ., 9997, 9998, 9999, 10,000. As used herein, the term "nucleic acid primer" or "nucleic acid probe" or "oligonucleotide" encompasses both DNA primers and RNA primers. In certain embodiments, the primers or probes may be physically located on a single solid substrate or on multiple substrates.

[0129] The kit described herein can include at least one first nucleic acid primer (e.g., at least 8 nucleotides in length) that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide (e.g., at position 92203667 of chromosome 1 in the transforming growth factor beta receptor III (TGFBR3) gene) and at least one second nucleic acid primer (e.g., at least 8 nucleotides in length) that is complementary to a SNP (e.g., SNP rs347027). In some embodiments, the at least one first nucleic acid primer detects the unmethylated CpG dinucleotide. In some embodiments, the at least one second nucleic acid primer has a sequence that detects the G nucleotide in SNP rs347027.

[0130] In some embodiments, the kit can further include at least one third nucleic acid primer (e.g., at least 8 nucleotides in length) that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide (e.g., at position 92203667 of chromosome 1 in the TGFBR gene), and that detects the CpG dinucleotide that is methylated.

[0131] The kits described herein can include at least one first nucleic acid primer (e.g., at least 8 nucleotides in length) that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide (e.g., at position 38364951 within an intergenic region of chromosome 15), where the first nucleic acid primer detects the unmethylated CpG dinucleotide, and at least one second nucleic acid primer (e.g., at least 8 nucleotides in length) that is complementary to a SNP (e.g., rs4937276).

[0132] In some embodiments, the kit can further comprise at least one third nucleic acid primer (e.g., at least 8 nucleotides in length) that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide (e.g., at position 38364951 within an intergenic region of chromosome 15), where the at least one second nucleic acid primer detects the CpG dinucleotide that is methylated.

[0133] The kits described herein can include at least one first nucleic acid primer (e.g., at least 8 nucleotides in length) that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide (e.g., at position 84206068 on chromosome 4 in the coenzyme Q2 4-hydroxybenzoate polyprenyltransferase (COQ2) gene), where the first nucleic acid primer detects the unmethylated CpG dinucleotide, and at least one second nucleic acid primer (e.g., at least 8 nucleotides in length) that is complementary to a SNP (e.g., SNP rs17355663).

[0134] In some embodiments, the kit can further comprise at least one third nucleic acid primer (e.g., at least 8 nucleotides in length) that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide (e.g., at position 84206068 on chromosome 4 within the coenzyme Q2 4-hydroxybenzoate polyprenyltransferase (COQ2) gene), wherein the at least one second nucleic acid primer detects the CpG dinucleotide that is methylated.

[0135] The kits described herein can include at least one first nucleic acid primer (e.g., at least 8 nucleotides in length) that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide (e.g., at position 26146070 on chromosome 16 in the heparan sulfate 3-O-sulfotransferase 4 (HS3ST4) gene), where the first nucleic acid primer detects the unmethylated CpG dinucleotide, and at least one second nucleic acid primer (e.g., at least 8 nucleotides in length) that is complementary to a SNP (e.g., SNP rs235807).

[0136] In some embodiments, the kit can further comprise at least one third nucleic acid primer (e.g., at least 8 nucleotides in length) that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide (e.g., at position 26146070 on chromosome 16 within the heparan sulfate 3-O-sulfotransferase 4 (HS3ST4) gene), where the at least one second nucleic acid primer detects the CpG dinucleotide that is methylated.

[0137] The kits described herein can include at least one first nucleic acid primer (e.g., at least 8 nucleotides in length) that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide (e.g., at position 91171013 in an intergenic region of chromosome 1), where the first nucleic acid primer detects the unmethylated CpG dinucleotide, and at least one second nucleic acid primer (e.g., at least 8 nucleotides in length) that is complementary to a SNP (e.g., SNP rs11579814).

[0138] In some embodiments, the kit can further comprise at least one third nucleic acid primer (e.g., at least 8 nucleotides in length) that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide (e.g., at position 91171013 of the intergenic region of chromosome 1), wherein the at least one second nucleic acid primer detects the CpG dinucleotide that is methylated.

[0139] The kits described herein can include at least one first nucleic acid primer (e.g., at least 8 nucleotides in length) that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide (e.g., at position 39491936 on chromosome 1 in the NADH dehydrogenase (ubiquinone) Fe-S protein 5 (NDUFS5) gene), where the first nucleic acid primer detects the unmethylated CpG dinucleotide, and at least one second nucleic acid primer (e.g., at least 8 nucleotides in length) that is complementary to a SNP (e.g., SNP rs2275187).

[0140] In some embodiments, the kit can further comprise at least one third nucleic acid primer (e.g., at least 8 nucleotides in length) that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide (e.g., at position 39491936 of chromosome 1 within the NADH dehydrogenase (ubiquinone) Fe-S protein 5 (NDUFS5) gene), where the at least one second nucleic acid primer detects the CpG dinucleotide that is methylated.

[0141] The kits described herein can include at least one first nucleic acid primer (e.g., at least 8 nucleotides in length) that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide (e.g., at position 186426136 mapped to chromosome 1 within the phosducin gene), where the first nucleic acid primer detects the unmethylated CpG dinucleotide, and at least one second nucleic acid primer (e.g., at least 8 nucleotides in length) that is complementary to a SNP (e.g., SNP rs4336803).

[0142] In some embodiments, the kit can further comprise at least one third nucleic acid primer (e.g., at least 8 nucleotides in length) that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide (e.g., at position 186426136 mapped to chromosome 1 within the phosducin gene), wherein the at least one second nucleic acid primer detects the CpG dinucleotide that is methylated.

[0143] The kits described herein can include at least one first nucleic acid primer (e.g., at least 8 nucleotides in length) that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide (e.g., at position 205475130 of chromosome 1 in the cyclin dependent kinase 18 (CDK18) gene), where the first nucleic acid primer detects the unmethylated CpG dinucleotide, and at least one second nucleic acid primer (e.g., at least 8 nucleotides in length) that is complementary to a SNP (e.g., SNP rs4951158).

[0144] In some embodiments, the kit can further comprise at least one third nucleic acid primer (e.g., at least 8 nucleotides in length) that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide (e.g., at position 205475130 of chromosome 1 within the cyclin-dependent kinase 18 (CDK18) gene), where the at least one second nucleic acid primer detects the CpG dinucleotide that is methylated.

[0145] The kits described herein can include at least one first nucleic acid primer (e.g., at least 8 nucleotides in length) that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide (e.g., at position 130614013 on chromosome 3 in the ATPase, Ca++ Transporting, Type 2C, Member 1 (ATP2C1) gene), where the first nucleic acid primer detects the unmethylated CpG dinucleotide, and at least one second nucleic acid primer (e.g., at least 8 nucleotides in length) that is complementary to a SNP (e.g., SNP rs925613).

[0146] In some embodiments, the kit can further comprise at least one third nucleic acid primer (e.g., at least 8 nucleotides in length) that is complementary to a bisulfite converted nucleic acid sequence that includes a CpG dinucleotide (e.g., at position 130614013 on chromosome 3 within the ATPase, Ca++ Transporting, Type 2C, Member 1 (ATP2C1) gene), where the at least one second nucleic acid primer detects the CpG dinucleotide that is methylated.

[0147] It will be appreciated that any of the nucleic acid primers, probes, or oligonucleotides described herein may contain one or more nucleotide analogs and / or one or more synthetic or non-naturally occurring nucleotides.

[0148] It will also be appreciated that any of the kits described herein can include a solid substrate.In some embodiments, one or more of the nucleic acid primers can be attached to a solid support.Examples of solid supports include, but are not limited to, polymers, glass, semiconductors, paper, metals, gels, or hydrogels.Additional examples of solid supports include, but are not limited to, microarrays or microfluidic cards.

[0149] It will also be appreciated that any of the kits described herein can include one or more detectable labels.In some embodiments, one or more of the nucleic acid primers can be labeled with one or more detectable labels.Representative detectable labels include, but are not limited to, enzyme label, fluorescent label, and color label.

[0150] An algorithm for predicting postoperative cardiac events Any number of algorithms that can obtain linear effects (e.g., linear regression) or both linear and nonlinear effects (e.g., random forests, gradient boosting, neural networks (e.g., deep neural networks, extreme learning machines (ELM)), support vector machines, hidden Markov models) can be used in the methods described herein. See, e.g., McKinney et al., 2011, Appl. Bioinform., 5(2):77-88; Gunther et al., 2012, BMC Genet., 13:37; and Ogutu et al., 2011, BMC Proceedings, 5(Suppl 3):S11. Any type of machine learning or deep learning neural network algorithm (synchronous or nonsynchronous) that can obtain linear and / or nonlinear contributions of traits for prediction can be used. See, e.g., FIG. 14. In some cases, a combination of algorithms (e.g., a combination or ensemble of multiple algorithms that obtain linear and / or nonlinear contributions of traits) is used.

[0151] As just one example, Random Forest™ is a popular machine learning algorithm created by Breiman & Cutler to create “classification trees” (see, for example, “stat.berkeley.edu / ~breiman / RandomForests / cc_home.htm” on the World Wide Web). Using standard machine learning and predictive modeling techniques, a diagnostic classification algorithm was constructed following the guidelines detailed by Breiman & Cutler as implemented in the R and Python programming languages ​​(although it can be implemented in many other programming languages). The diagnostic classification algorithm was created using data from at least two traits (T) and diagnoses of interest from the population. To determine an output factor (e.g., diagnosis) for a new individual, values ​​for at least two traits (T) can be easily determined and that information can be input into an algorithm (e.g., the diagnostic classification algorithm described herein or another algorithm discussed above) that can obtain the linear and nonlinear contributions of the traits.

[0152] As described herein, the input elements are at least one genotype (e.g., SNP) and the methylation status of at least one CpG dinucleotide, and the outcome can correspond to a positive or negative probability (e.g., prediction or diagnosis) of CHD, CHF, stroke, or other disease. The trait (T) used to determine the outcome can correspond to the methylation status of at least one CpG dinucleotide or at least one genotype (e.g., of an SNP), but the trait (T) can also correspond to at least one interaction (e.g., between methylation status and genotype (CpG x SNP), between the methylation status of two different sites (CpG x CpG), or between two different genotypes (SNP x SNP)). It will be appreciated that any such interaction can be visualized using a partial dependence plot.

[0153] It will be apparent that the present disclosure provides one of skill in the art with the ability to construct matrices in which the methylation status of one or more CpG dinucleotides and one or more genotypes (e.g., SNPs; e.g., at one or more alleles) can be assessed, typically using a computer, as described herein, to identify interactions and allow prediction of postoperative cardiac events. Although such analyses are complex, undue experimentation is not required, as all the necessary information is either readily available to one of skill in the art or can be obtained by experimentation as described herein.

[0154] The present invention is further described in the following examples, which are provided for illustrative purposes and are not intended to limit the invention in any way. Standard techniques well known in the art or those specifically described below are utilized. All patents and references cited herein are incorporated by reference in their entirety. EXAMPLES

[0155] Example 1: The effect of methylation and G x methylation in predicting cardiovascular disease Methylation-based biomarkers are gaining increasing clinical appeal for use in diagnosis and guiding treatment. In an attempt to identify CpG loci that predict cardiovascular disease by their methylation status, several researchers have combined genome-wide approaches with clinical diagnosis. In particular, Brenner and coworkers identified F2RL3 residue cg03636183 as a biomarker for cardiovascular disease (Breitling et al., "Smoking, F2RL3 methylation, and prognosis in stable coronary heart disease," Eur. Heart J., 2012, 33:2841-8). Unfortunately, these analyses showed complete confounding by incomplete information on smoking status and did not consider the possibility of genetic variance causing confounding. Indeed, the coronary heart disease signal at cg03636183 disappears when using a biomarker approach that fully accounts for smoking intensity. Furthermore, using genome-wide methylation and genetic analysis in combination with biomarker-guided smoking assessment, we recently analyzed data from a large cohort of subjects providing heart disease information. We demonstrate that methylation status in the context of genetics, embodied by methylation-genotype interactions (meQTL), independent of smoking intensity status, indeed contributes better to the prediction of coronary heart disease, and that the use of algorithms that combine regional genetic variation and methylation significantly improves the prediction of coronary heart disease.

[0156] Example 2 Incorporating gene × methylation interactions increases predictive power for the presence of coronary heart disease summary Coronary heart disease (CHD) is the leading cause of death in the United States. Effective treatments exist to prevent CHD morbidity and mortality, but their clinical implementation is hindered by inefficient screening techniques. Recently, others and the present inventors have shown that DNA methylation signatures can infer the presence of various disorders related to CHD, such as smoking. Unfortunately, when these epigenetic techniques are applied to CHD itself, the power of these methods is reduced, thus limiting their clinical usefulness. One possible reason for these failures may be the obscuration of epigenetic signatures of CHD by gene x methylation interaction (meQTL) effects. To investigate this possibility, the present inventors investigated whether the incorporation of meQTLs could also be employed to improve the predictive value of prior methylation-based assessments by analyzing genetic and epigenetic data from the Framingham Heart Disease Study using a stepwise approach. In our first attempt using analysis of the area under the receiver operating characteristic (ROC) curve (AUC) focused on F2RL3, we found that the addition of a cis-meQTL and a trans-meQTL at CpG residue cg13751927, close to the locus previously described by Brenner and coworkers, significantly improved the ability of a model that only included smoking status to predict CHD in the training dataset. Subsequent genome-wide meQTL analysis identified a total of 3,265 cis-meQTLs at FDR 0.05 and 467,314 significant trans-meQTLs at FDR 0.1. Our preliminary analysis suggests that the inclusion of six additional cis-meQTLs will further improve the AUC of the existing model using only the meQTLs and smoking at F2RL3. This non-optimized model is able to predict CHD with 81.9% accuracy. We conclude that incorporating meQTL information into the prediction algorithm can significantly improve the power of the algorithm to predict CHD, and further attempts to improve the model's ability to predict CHD are possible through additional optimized machine learning models.

[0157] Introduction Coronary heart disease (CHD) is the leading cause of death in the United States, with direct costs to the U.S. economy estimated at $108 billion in 2012. 1 Over the past 50 years, several medicines and devices have been developed to treat CHD. Unfortunately, tens of thousands of Americans continue to die each year because their presence goes unnoticed until a fatal cardiac event. More effective screening methods for CHD could conceivably result in the prevention of some of these deaths. 1 However, at present, the cumbersome nature of certain techniques, such as fasting fat panels, and / or the limited predictive power of other techniques, such as electrocardiograms and C-reactive protein levels, limit the effectiveness of current approaches in identifying CHD. 1-3 .

[0158] Some investigators have proposed that genetic approaches may offer another potential avenue for preventing CHD-related morbidity and mortality. 4 Using whole-exome and genome sequencing techniques, several variants that predispose to CHD have been identified. The relative risks conferred by many of these variants are often substantial, and their presence can be useful for directing prevention and treatment strategies. 5 However, variants with large effect sizes tend to be rare and their presence may not be pathognomonic for the disease at hand. 4 Therefore, at present, genetic approaches are not commonly used in general medical practice to assess the presence or absence of ongoing CHD.

[0159] Alternatively, other researchers have proposed that epigenetic techniques may be useful in the assessment of CHD. 6-8 Since the development of replicate peripheral leukocyte DNA methylation signatures for the presence of type 2 diabetes, smoking, and alcohol use, 9-12, this suggestion has strong face validity. Notably, Brenner and colleagues used this approach to propose that DNA methylation of cg03636183, a CpG residue found in the coagulation factor II (thrombin) receptor-like 3 (F2RL3), predicts cardiovascular disease risk. 6,13 Although this is a highly biologically plausible result, their subsequent studies demonstrated that the CHD-associated signal at cg03636183 completely cosegregated with smoking status as indicated by DNA methylation at cg05575921. 14 cg05575921 is a CpG residue found in the aryl hydrocarbon receptor repressor (AHRR), and its strong predictive power for smoking status has been demonstrated in multiple studies. 15 .

[0160] However, the initially intriguing results of cg03636183, unable to uniquely identify additional risk beyond that conferred by smoking alone, do not mean that methylation approaches to assess the presence of CHD are doomed to failure. Instead, they suggest that a successful approach will need to be more granular, and that a reconsideration of our conceptualization of the relationship between methylation status and CHD would be useful. For example, the results from Brenner's group strongly suggest that current methylation algorithms for predicting CHD should include an indicator of smoking status. Given the fact that smoking is the largest preventable risk factor for CHD, 16 , which is largely logical, but in addition, they may need to consider that the long-term effects of exposure to environmental risk factors, such as smoking, or other cardiac risk factors, such as hyperlipidemia, may be obscured by gene-environment interactions.

[0161] The role of gene-environment interaction (G×E) effects in moderating vulnerability to disease is perhaps better recognized in the behavioral sciences. The basic premise of G×E effects is that environmental influences during a developmentally sensitive period, in conjunction with genetics, alter the biological properties of a system so that in the future, even in the absence of that environmental factor, there will be increased vulnerability to disease. 17 Importantly, direct effects of environmental variables are generally undetectable due to confounding by genetic variables. Instead, direct effects can only be detected when considered in conjunction with genetic variation. Although the strength of some G×E results remains controversial, many researchers continue to emphasize the importance of these G×E effects in the pathogenesis of various behavioral disorders, such as depression, post-traumatic stress disorder, and antisocial behavior. 18-20 .

[0162] The physical basis of these G×E effects is thought to be variable. For example, at the anatomical level, G×E effects on behavioral deficits may emerge due to changes in synaptic structure. 21 At the molecular level, however, the physical manifestation of the G×E effect is less certain. However, some researchers have suggested that changes in DNA methylation could be one potential mechanism by which the physical effects of the G×E effect are transmitted. 22 .

[0163] Interestingly, it has been known for many years that behavior-related changes in the environment can alter DNA methylation, and that the extent of these changes is influenced by genetic variation. In our initial candidate gene studies, we showed that smoking altered DNA methylation in the promoter region of monoamine oxidase A (MAOA), a major regulator of monoaminergic neurotransmission, and that genotype at a well-characterized promoter-associated variable nucleotide repeat (VNTR) altered percent methylation in both the presence and absence of smoking. 23,24Volkow and colleagues subsequently demonstrated that methylation changes at these loci were functional. 25 .

[0164] In current terminology, the effects of VNTRs on smoking or basal DNA methylation are today referred to as genotype-methylation interactions or methylation quantitative trait locus (meQTL) effects. When we performed our first genome-wide study, these MAOA meQTL effects had consequences on our ability to detect their association with smoking. Regardless of the magnitude of changes that smoking induces in DNA methylation in response to smoking, probes surrounding the MAOA VNTR are not included among the higher ranking probes, even in studies of DNA from subjects of only one gender. Other observations from the first study are instructive as well. First, the regional methylation response to smoking was not homogenous. Factorial analysis of the methylation status of 88 CpG residues in the promoter-associated island showed that increased methylation in some regions of the island could also be associated with demethylation in other regions. 26 Finally, the effect of smoking on DNA methylation was not static; in the future, the signature tended to decay. 23 Thus, it was apparent from these early studies that genetic variation in the MAOA promoter could modify in complex ways the effects of environmental factors on regional DNA methylation signatures.

[0165] Subsequent studies suggest that many of these subtle complexities in response to smoking are evident at the genome-wide level. For example, it is clear that genetic variation influences the magnitude of the methylation response at the genome-wide level, and that these meQTL effects may impair the ability to replicate results at a given locus in a pool of subjects of different ancestry when attempting to replicate signatures from various ancestries. 27,28 Second, and equally important, reversion of methylation signatures can be complex. 28,29Guida and colleagues specifically investigated the epigenomic response to smoking cessation in DNA from 745 pooled subjects, found two classes of CpG sites whose methylation signatures reverted over time and those that did not, and concluded that at a genome-wide level, "the dynamics of methylation changes after smoking cessation are driven by smoking-induced changes of differential and site-specific magnitude that are independent of smoking intensity and duration." 29 Taken together, a substantial body of evidence suggests that genome-wide signatures for smoking are only partially reversible, and that the majority of irreversible variation may be completely hidden within meQTL effects.

[0166] Since smoking is a major risk factor for CHD, this also suggests that some of the smoking-induced risk present in the epigenome that modulates the risk of CHD may be somewhat irreversible and hidden in the meQTL response. Additionally, since smoking is only one of several factors that can modify the risk of CHD, and these other factors may also have complex epigenetic signatures, it is possible that interrogation of DNA methylation of peripheral WBCs may reveal relatively stable meQTL that modulate the risk of CHD. In this brief report, we used a regression analysis approach and epigenetic and genetic resources from 324 subjects who participated in the Framingham Heart Study to test whether the addition of meQTL effects can contribute to the algorithm to predict CHD.

[0167] method Framingham Heart Study. Data used in this study were obtained from participants in the Framingham Heart Study (FHS). 30The FHS is a long-term study aimed at understanding the risk of cardiovascular disease (CVD) and consists of several cohorts, including the original cohort, the offspring cohort, the omnicohort, the third generation cohort, the new offspring spouse cohort, and the second generation omnicohort. Specifically, the offspring cohort, which began in 1971 and consists of the children of the original cohort and their spouses, was used in this study. This cohort consists of 2,483 men and 2,641 women (5,124 in total). 31 The specific assays described in this communication were approved by the University of Iowa Institutional Review Board.

[0168] Genome-wide DNA methylation. Of the 5,124 individuals in the Offspring cohort, only 2,567 individuals (excluding duplicates) with DNA methylation data were considered. These individuals were included in the DNA methylation study because they participated in the Framingham Offspring Round 8, consented to the genetic study, had buffy coat samples, and had sufficient quantity and quality of DNA for methylation profiling. Round 8 took place between 2005 and 2008. Genomic DNA extracted from the leukocytes of these individuals was converted to bisulfite and then profiled for genome-wide DNA methylation using the Illumina HumanMethylation450 BeadChip (San Diego, CA) at either the University of Minnesota or Johns Hopkins University. The intensity data (IDAT) files of the samples, together with their slide and array information, were used to perform DASEN normalization using the MethyLumi, WateRmelon, and IlluminaHumanMethylation450k.db R packages. 32DASEN normalization performs probe filtering, background correction, and adjustment for probe type. Samples were removed if they contained >1% of CpG sites with a detection p-value >0.05. CpG sites were removed if they had a bead count of <3 and / or >1% of samples had a detection p-value >0.05. After DASEN normalization, 2,560 samples and 484,241 sites (484,125 CpG sites) remained. CpG sites were grouped by chromosome. Methylation beta values ​​were converted to M values ​​using the beta2m R function in the Lumi package, followed by z-scores using an R script 33 .

[0169] Genome-wide genotypes. Of the 2,560 individuals remaining after DNA methylation quality control, 2,406 (1,100 males and 1,306 females) had genome-wide genotype data from the Affymetrix GeneChip HumanMapping 500K Array Set (Santa Clara, CA). This array is capable of profiling 500,568 SNPs in the genome. Quality control was performed at both the sample and SNP probe levels in PLINK. The first quality control step involved identifying individuals with discordant gender information. None were identified. Next, individuals with heterozygosity rates greater or less than the mean ± 2 SD and proportions of missing SNPs > 0.03 were excluded. Related individuals were also excluded if their ancestry identity value was > 0.185 (intermediate between second-degree and third-degree consanguinity). After these sample-level quality control steps, 1,599 individuals remained (722 men and 877 women). At the probe level, minor allele frequencies >1% and Hardy-Weinberg equilibrium p-values ​​>10 -5 and SNPs with a SNP missingness rate <5% were retained. A total of 403,192 SNPs remained after these quality control steps. 34 , genotypes were coded as 0, 1, or 2.

[0170] Phenotypes. Phenotypes considered in the methylation quantitative trait loci (meQTL) analysis included age, sex, batch, tobacco smoke exposure, and coronary heart disease (CHD) status. Of the 1,599 individuals who passed all quality control stages, 324 were recorded as having CHD at study 8. These individuals were the training set. CHD was recorded as either a confirmed or newly diagnosed diagnosis, and individuals were diagnosed with CHD if the Framingham Endpoint Review Committee (a panel of three investigators) agreed that one of the following was present: myocardial infarction, coronary insufficiency, angina, sudden death due to CHD, or non-sudden death due to CHD. For the analysis, CHD was coded as 1 if the individual had either a confirmed and / or newly diagnosed CHD, and 0 otherwise. The age used was the individual's age at the time of study 8. Batch was the methylation plate count and cigarette smoke exposure was the methylation level of the aryl hydrocarbon receptor repressor (AHRR) smoking biomarker cg05575921. The demographics of the 324 individuals in the training set are summarized in Table 1.

[0171] Table 1. Demographics of the 324 individuals in the training set. TIFF0007672192000001.tif50128

[0172] The remaining 1275 individuals constituted the study data set. The CHD status of these individuals was coded as 0 if CHD was not present, otherwise coded as 1. The demographics of these individuals are summarized in Table 2.

[0173] Table 2. Demographics of the 1275 individuals in the test dataset. TIFF0007672192000002.tif52143

[0174] Methylation quantitative trait loci. meQTL analysis in the training set was performed using the MatrixeQTL package in R. 35 To determine the significant effects of SNPs on methylation for a given CHD status (meQTL), the following models were fitted: Meth i ~ age + sex + batch + cg05575921 + SNP j +CHD +SNP j *CHD

[0175] For prediction, significant SNPs j *Cis- and trans-meQTLs with CHD terms were retained. The analysis aimed to uncover specific SNPs that significantly predict specific methylation sites in a given CHD status after adjusting for age, sex, batch, tobacco smoke exposure, and the main effects of SNPs and CHD, so the interaction terms were of particular interest. This was achieved using the model LINEAR_CROSS model type in the MatrixeQTL package. The cis distance was chosen to be 500,000 on both sides of the site and was performed at the chromosome level. The meQTL analysis was performed at the genome-wide level and specifically for the coagulation factor II receptor-like 3 (F2RL3) gene. This was done to determine if there were other meQTLs beyond those identified for F2RL3 that better predict CHD.

[0176] Receiver Operating Characteristic Curve. I wrote an R script to perform a logistic regression of the model shown below, followed by the Receiver Operating Characteristic (ROC) curve using the pROC package in R. 36 We calculated the area under the curve (AUC) for cis-meQTLs that were nominally significant at the 0.05 level and for trans-meQTLs with an FDR of 0.1. In the model below, each meQTL is represented by a SNP * meth term. CHD ~ age + sex + batch + cg05575921 + SNP j + meth i + SNP j *meth i CHD ~ age + sex + batch + cg05575921 + SNP j + meth i + SNP j *methi

[0177] Model training. The model was trained on a training dataset of 324 individuals. The variables for the model were selected based on their individual area under the ROC curve (AUC) generated from the model. A 10-fold cross-validation was performed to determine the logistic regression threshold for CHD classification. A classification threshold of 0.5 was selected based on the average accuracy.

[0178] Testing the model. After determining the training model parameters and classification threshold, the trained model was applied to an independent test dataset. The demographics of the individuals in the test dataset are described above. Testing of the model was performed in R.

[0179] result cg05575921 for smoking status. As mentioned previously, smoking is a major risk factor for CHD. While most past studies have used self-reported smoking measures, the reliability and information value of these measures are suboptimal. Therefore, to minimize the influence of unreliable self-reporting and take advantage of the ability of continuous metrics to better capture tobacco consumption, we used the well-validated smoking biomarker cg05575921. 14,15,37 While cg05575921 was a strong predictor of self-reported smoking in 324 individuals (p-value = 8.71e-9, R 2 = 0.62), the strength of cg05575921 as a predictor of CHD surpasses self-reported smoking status (p-value = 1.64e-5, R 2 =0.085 vs. p-value=0.00218, R 2 = 0.042), demonstrating that incorporation of cg05575921 to represent tobacco smoke exposure instead of self-reported smoking status further strengthens downstream models for predicting CHD.

[0180] Methylation quantitative trait loci. Genome-wide DNA methylation analysis as described in Methods demonstrated the importance of accounting for the confounding effect of interactions between methylation and genotype for predicting CHD. After adjusting for age, sex, batch, and cg05575921, CHD was not significantly associated with any of the methylated CpG sites at the FDR significance level of 0.05. From the meQTL analysis, there were 5,458,462, 3,265, 2,025, and 1,227 significant cis-meQTLs at the nominal significance level of 0.05, FDR significance level of 0.05, FDR significance level of 0.01, and FDR significance level of 0.001, respectively. Similarly, there were 467,314 significant trans-meQTLs at the FDR significance level of 0.1. The importance of some of these meQTLs is demonstrated using the area under the receiver operating characteristic curve.

[0181] Receiver operating characteristic curves (ROC) of core variables. ROC curves represent the trade-off between the sensitivity and selectivity of the model. Before introducing genetic and epigenetic variables, we determined the area under the ROC curve (AUC) for the core variables age, sex, batch, and cg05575921 used in the meQTL model. The AUCs of age, sex, batch, and cg05575921 were 0.52, 0.51, 0.50, and 0.64, respectively. Collectively, these core variables yielded an AUC of 0.65, which is almost equal to the AUC of cg05575921 alone. When self-reported smoking was used instead of cg05575921, its individual and collective AUCs were 0.55 and 0.56, respectively. The ROC curves of these analyses are shown in Figure 1. Therefore, only one core variable, cg05575921, was included in the subsequent models.

[0182] ROC for predicting CHD in the training data. Using cg05575921 and the nine SNP-methylation interaction terms for predicting CHD, an AUC of the ROC curve of 0.964 was obtained (see Figure 2). The nine interaction terms and their respective AUCs with and without adding cg05575921 to the model are summarized in Table 3.

[0183] Table 3. List of the nine meQTLs used to generate the initial predictive model. TIFF0007672192000003.tif51164

[0184] Prediction model. A preliminary logistic prediction model was used to predict CHD in the training data. After 10-fold cross-validation, the classification threshold was set to 0.5. Of the 324 individuals, 299 were included in the prediction due to the absence of missing data. Of these 299 individuals, there were 73 and 226 with and without CHD, respectively. This means that if everyone was assigned to the majority class (i.e., absence of CHD), the accuracy of the prediction would be 75.6%. The average accuracy of this preliminary model after 10-fold cross-validation was 91%, which is much higher than the baseline.

[0185] Model validation. The trained model was used to predict CHD status in a unique test dataset of 1275 individuals. The model was capable of predicting CHD with 80% accuracy. The model needs to be further optimized.

[0186] Consideration These results demonstrate that the presence of CHD can be inferred through the use of methylation-genotype interactions derived from meQTL. However, before the results can be considered, it is important to mention some limitations to this study. First, the Framingham cohort is exclusively Caucasian, with the majority of subjects in their mid-to-late 60s and 70s. Thus, the results may not be applicable to other ethnicities or different age ranges. Second, besides cg05575921, the validity of the M (or B values) for other probes has not been confirmed by independent techniques such as pyrosequencing. Third, the Illumina arrays used in the study are no longer available. Due to changes in probe design or availability in new generation arrays, the ability to replicate and expand may be affected.

[0187] The present results highlight the value of resources such as the Framingham Heart Study in furthering our understanding of heart disease. Indeed, it is plausible to say that without this resource, this type of work would be difficult, if not impossible, to carry out. Moreover, given the present results with this unique dataset, a great deal of additional work is required before screening tests such as those described in this communication can be adopted clinically. Most obviously, the present results must be replicated and refined in other datasets and then retested in study populations representative of their intended future clinical applications. The latter point is particularly important, since even well-designed cohort studies that are epidemiologically sound in nature suffer from retention bias, which enriches for less severe illness in the remaining pool. This is particularly true for substance use-related illnesses, since probands with high levels of substance use are more likely to be lost to long-term follow-up. 38 Additionally, because SNP frequencies may vary among ethnicities, the effect size of a given meQTL may also vary. Therefore, extensive testing and development in ethnically-informative, diverse cohorts is necessary.

[0188] The improvement of AUC may hit a ceiling. Ironically, this has little to do with the quality or quantity of epigenetic and genetic data. Instead, the constraint may be uncertainty in clinical characterization. Sadly, even under the best conditions, clinically relevant CHD may remain undetected. This is true even for the FHS cohort. As a result, the "criteria" in this study are themselves somewhat imprecise with respect to the actual clinical situation. This imprecision increases the error even for biomarkers precisely targeted to the relevant biological properties, so our ability to improve AUC may depend on us being able to obtain more accurate clinical evaluations. 39 .

[0189] Another limitation of using this approach is the constantly evolving epidemiology of CHD. While genetic contribution to CHD is relatively constant, diet and other environmental exposures continue to fluctuate across generations. Perhaps the best illustration of this limitation is by considering the contribution of smoking to the predictive power of the test in previous generations. Because tobacco was introduced to Europe from the New World in the early 1500s, we can state with confidence that the contribution of smoking to CHD in medieval Europe was limited, and therefore the impact of cg05575921 on predictive power was zero. In contrast, in the 1960s, more than 40% of US adults smoked, so 40 , the contribution of smoking behavior to the prediction of CHD as captured by cg05575921 may have been significantly greater for subjects of that era. However, smoking is not the only environmental factor that varies between generations and cohorts. Over the past 20 years, there has been a notable change in our understanding and public attitudes toward the amount of saturated and trans fatty acids in a healthy diet. Because these environmental factors also have a strong influence on the likelihood of CHD, we expect that the weighting of meQTLs loading on these dietary factors may vary depending on age and ethnicity.

[0190] The improved predictive power of the smoking methylation biomarker cg05575921 compared with self-reported smoking is not unexpected: our initial studies have shown it to be a strong indicator of current smoking status with an AUC of 0.99 in studies using well-screened cases and controls. 37 The unreliability of self-reported smoking, especially in high-risk cohorts, is a well-established phenomenon. 41-44 Additionally, unlike CG05575921, categorical self-report does not capture smoking intensity. 37 Finally, a large number of subjects who could have participated in the study may have been former smokers but who were non-smokers at the Wave 8 interview but still had residual demethylation of the AHRR. In each of these cases, the use of a continuous metric may capture additional vulnerability to CHD not captured by the binary smoking variable.

[0191] Alcoholism is also a risk factor for CHD 1 , we were somewhat surprised that our previously established and validated biomarker approach to assess alcohol intake did not have a greater predictive impact. 10,45 In our initial model, the addition of the methylation status of cg2313759 only improved the AUC by 0.015. One reason for this failure to show the effect of alcohol use on the risk of CHD could be that this marker is not as well validated as our smoking biomarker, but there are other reasons as well. First and foremost, in contrast to the methylation of cg05575921, which shows a persistently increased risk of reduced life expectancy at all exposure levels, the methylation of cg2313759 shows an inverted U-shaped distribution with respect to biological aging. It is unknown whether the risk of CHD also follows a U-shaped distribution with respect to alcohol intake. However, it does suggest that any successful algorithm incorporating the main effect of alcohol-related methylation cannot use a simple linear approach.

[0192] Our success in discovering meQTL predictive of CHD in the absence of genome-wide significant main effects may have important implications for exploring marker sets for other common complex disorders in adulthood. Among the top 10 leading causes of death in the United States, reliable methylation signatures have only been developed for type 2 diabetes and chronic obstructive pulmonary disease (COPD) using main effects. 12,46 Since the ability to find good biomarkers of disease is highly dependent on the reliability of clinical diagnosis, the success in these two paradigms may be secondary to the excellent diagnostic reliability of the methods used to diagnose these two disorders, namely hemoglobin A1C and spirometry. Additionally, it is important to mention that the diagnostic signature for T2DM maps primarily to pathways affected by excessive glucose levels, whereas the signature associated with COPD largely overlaps with the signature of smoking, which contributes to 95% of all cases of COPD. 12,46 Nevertheless, because many of the risk factors for other leading causes of death, such as stroke, overlap with risk factors for CHD (e.g., smoking), we are optimistic that this approach can be used to generate similar profiles.

[0193] Unfortunately, the majority of common adult-onset complex disorders do not have good existing biomarkers or large effect size etiological factors. In these cases, an approach incorporating meQTLs may be beneficial, but the real question is why. Although speculative, our experience with local and genome-wide data indicates that chronic exposure to cellular stressors leads to epigenome reorganization that may only be partially reversible. Regardless of how long it lasts, that disorganization of the genome is inevitably associated with disease, and so it can be used as a disease biomarker. Understanding the reversion time of each of these meQTLs may provide additional insight. For example, pharmacological interventions may have effects on a distinct subset of these meQTLs. By understanding the relationship between reversion at these loci and treatment outcomes, it may be possible to optimize existing medications or better tailor new combination regimens.

[0194] The fact that no main effect of methylation is observed for CHD does not necessarily point to the absence of an epigenetic signature in WBCs. On the contrary, it is evidence of the complexity of the overall genetic architecture. For example, the methylation status at thousands of CpG loci was correlated with smoking status (for review see 14,15 ), the signal at cg05575921 is one of the few whose signal is not obscured by ethnic-specific genetic differences in one population or another. 27 This report shows that the epigenomic response to smoking contains a plethora of meQTLs, but the need to measure at least two values ​​for each meQTL suggests that translating these results into improvements in diagnosis, treatment, or prevention may be more difficult.

[0195] In summary, we report that an algorithm incorporating information from meQTL can predict the presence of CHD in FCS. We suggest that further research is indicated to replicate and extend the generalizability of the approach to other ethnic cohorts. We further suggest that a similar approach could lead to the generation of methylation profiles for other common complex disorders, such as stroke.

[0196] References for Example 2 TIFF0007672192000004.tif235149TIFF0007672192000005.tif235148TIFF0007672192000006.tif248147

[0197] Example 3 Smoking-associated methylation quantitative trait loci preferentially map to neurodevelopmental pathways Smoking is the leading preventable cause of morbidity and mortality in the United States. Smoking exerts its effects indirectly by increasing susceptibility to common co-morbidities such as coronary heart disease and coronary obstructive pulmonary disease. While the association between these disorders and smoking has been widely studied, our understanding of the molecular mechanisms by which smoking increases vulnerability to co-morbidities can also be further improved. This is especially true for disorders that preferentially affect the central nervous system (CNS). Smoking is a known risk factor for the development of attention deficit hyperactivity disorder and panic disorder. We designed our study to understand the effect of smoking on DNA methylation in the presence and absence of genetic conditions in the Framingham Heart Study (FHS). Specifically, we used data from 1599 individuals from the FHS Offspring Cohort. These individuals were of European descent and were in their early to mid-60s. The self-reported smoking rate among these individuals was 7.6%. Genome-wide DNA methylation was profiled using the Illumina HumanMethylation 450k BeadChip, and genome-wide SNP data was assessed using the Affymetrix GeneChip HumanMapping 500k Array Set. To understand the effect of smoking on DNA methylation in the absence of genetic variation, we regressed smoking against DNA methylation, adjusting for age, sex, and batch. After correcting for multiple comparisons, methylation status was significant at the 0.05 level at 525 sites. Consistent with previous studies, the top probe was cg05575921 in the AHRR gene (p-value of 7.65 × 10 -155). To determine the effect of smoking on DNA methylation in the presence of genetic variation, cis and trans-methylation quantitative trait loci (meQTL) analyses were then performed to determine the significant effect of SNPs on DNA methylation for a given smoking status when adjusted for age, sex, and batch. A total of 126,369,511 cis analyses and 195,068,554,297 trans analyses were performed. Of these, 5294 (0.00419%) and 422,623 (0.00022%) significant cis and trans-meQTLs were generated after correction for multiple comparisons at a significance level of 0.05. To better visualize and compare the connectivity and gene ontology (GO) enrichment between the results of both analyses, we generated protein-protein interaction (PPI) networks. The DNA methylation analysis was mapped to inflammatory pathways, whereas the cis and trans-meQTL analysis was mapped to neurodevelopmental pathways. These neurodevelopmental pathways may also provide additional insight into the association between smoking and psychiatric disorders. Furthermore, this study demonstrates that combined genetic and epigenetic analyses may be crucial to more fully understanding the reciprocal influences of environmental variables, such as smoking, and pathophysiological outcomes.

[0198] Example 4 Integrated genetic and epigenetic prediction of coronary heart disease in the Framingham Heart Study summary Background: Coronary heart disease (CHD) is the leading cause of mortality and morbidity in the United States. Unfortunately, the first sign of CHD for some patients is fatal myocardial infarction. Although sensitive methods to detect current CHD or risk of future cardiac events could prevent some of these deaths, current biomarkers of subclinical CHD are insensitive and nonspecific. Recently, other researchers and the present inventors have shown that array-based DNA methylation assessments accurately predict tobacco consumption and smoking-related risk of CHD. However, attempts to extract additional risk for CHD information from these genome-wide assessments have not yet been successful.

[0199] Methods and Results: Given the notion that CHD risk factors are a conglomeration of genetic and environmental factors, we used machine learning techniques to integrate genetic, epigenetic and phenotypic data (n=2214) from the Framingham Heart Study to build and test a random forest classification model for CHD risk. Our final classifier was trained on n=1545 individuals and utilized four DNA methylation sites, two SNPs, age and sex, which was capable of predicting CHD status with an accuracy of 78% in the test set (n=669), and a sensitivity and specificity of 0.75 and 0.80, respectively. In contrast, a model using only CHD risk factors as predictors had an accuracy and sensitivity of only 65% ​​and 0.41, respectively. The specificity was 0.89. Regression analysis of individual clinical risk factors highlights the strong role of pathways moderated by smoking in CHD pathogenesis.

[0200] Conclusions: This study demonstrates the feasibility of an integrated approach to predict symptomatic CHD status and suggests that further work may also lead to the introduction of sensitive, easily adoptable methods for the detection of asymptomatic CHD.

[0201] Introduction Coronary heart disease (CHD) is the leading cause of death in the United States 1 Effective methods exist to prevent this mortality and associated morbidity, but when employed, they are often ineffective. In fact, sudden cardiac death is the initial manifestation in 15% of patients with CHD. 2、3 .

[0202] In an attempt to detect and treat CHD more effectively, several screening methods for both symptomatic CHD (angina, myocardial infarction) and asymptomatic CHD have been developed. For asymptomatic patients, the intensity of screening for CHD depends on the level of clinical suspicion. Although clinicians are vigilant for potential cardiac disease at any age, greater attention is given to individuals with classic risk factors for CHD, including family history of CHD, smoking, elevated systolic blood pressure, diabetes, or anything resembling angina-like chest pain, as defined in the Framingham Heart Study (FHS). 4、5 Depending on the level of suspicion of CHD, initial testing typically involves a medical exam and a fasting fat panel including low-density lipoprotein (LDL), high-density lipoprotein (HDL), and triglyceride levels. 5 The next level of response is usually an electrocardiogram (ECG), followed by more expensive and invasive measures including stress tests and angiograms. 6 .

[0203] Unfortunately, most clinical routine tests, i.e., 12-lead ECG and serum lipid screening, have very low sensitivity for CHD. For example, in a study of 479 patients admitted for acute chest pain with confirmed MI by creatine kinase-MB isoenzyme (CK-MB) and troponin T (TnT), 12-lead ECGs were positive in only 33% and 28% of patients at admission and after hospitalization, respectively. 7 Similarly, serum lipid (cholesterol and triglyceride) screening has been employed for many years. Most importantly, elevated serum cholesterol levels performed at recruitment in the Framingham Heart Study (FHS) with a cutoff of 260 mg / dl failed to identify two-thirds of all men who developed CHD over the following four years. Thus, over the past decade, there has been an ever-increasing need for biomarkers for the prediction and diagnosis of CHD.

[0204] Stimulated by the lack of sensitivity and specificity of standard techniques such as ECG and lipid profiles, numerous investigators have attempted to identify biomarkers for asymptomatic CHD and its closely-related cluster of diseases, cardiovascular disease (CVD). Diverse approaches have been used, including imaging, mechanical and bioelectrical techniques. 8-10 ,The majority of researchers have focused on blood-based methods due to 1) the proof-of-principle provided by previous studies with triglycerides and cholesterol, 2) the clear involvement of blood components such as platelets and leukocytes in the pathogenesis of CHD and CVD, and 3) the ease of incorporating blood-based approaches into current medical diagnostics.

[0205] The majority of these blood-based approaches have focused on circulating lipids and proteins, such as hemoglobin A1C (HbA1c), fibrinogen, vitamin D, C-reactive protein (CRP), apolipoprotein B (ApoB), apolipoprotein AI (ApoAI), and cholesterol (including high- and low-density, HDL and LDL) (for review, see 11、12 (See ). When appropriate cutoffs are used in the study setting, each of these markers is moderately informative regarding future disease occurrence (odds ratios or relative ratios between 1.5 and 2.5). 11 Additionally, for people with pre-existing disease, cardiac troponin (cTn) levels and high-sensitivity (HsCRP) ratios can provide future risk information. 11 However, each of these markers has challenges in their clinical implementation, such as lack of ease of measurement, ethnic variability, or limited predictive ranges, which have precluded them from routine implementation in CHD screening.

[0206] Searching for alternatives to create more effective screening methods, other investigators have identified risk-associated variants using genetic approaches, including more recently genome-wide association studies (GWAS) and exome / genome sequencing studies (for review, see O'Donnell and Nabel, 2011). 13 So far, these studies have isolated a total genetic risk of about 10% for CHD. 14、15 Notably, many of these SNPs map to lipid and inflammatory pathways, both of which are known to be important from previous CHD studies. 15 Although these studies can predict who is potentially vulnerable to CHD, they do not actually indicate whether an individual will have CHD, and meta-analyses indicate, at best, that the contribution of purely genetic approaches to predicting CHD is minimal. 16 As such, genetic approaches have not been incorporated into routine clinical practice.

[0207] Epigenetic approaches may provide a new means for assessing risk for CHD. It is already well established that epigenetic approaches can quantitatively assess tobacco consumption, which may be the largest preventable cause of CHD. 17、18 Of note, Hermann Brenner and colleagues showed that DNA methylation at cg03636183 predicted not only smoking status but also risk of MI. 19、20 Unfortunately, their group also showed that the risk of MI implied by cg03636183 was fully encompassed by smoking status as indicated by methylation of cg05575921, the best-documented epigenetic smoking biomarker in all ethnic groups, suggesting that CHD risk and smoking are not independent. 17、21 .

[0208] Methylation status markers such as cg03636183 and the GPR15 marker cg19859270 22、23The observation that one of the reasons why does not predict smoking status well in the whole population is the presence of genetic confounding of methylation changes by local genetic variants. 22 is critical to this study. It was originally described as a relatively static interaction (G × Meth) 24 Our understanding of these effects has been refined over the past few years, indicating that a subset of these interactions may be related to the degree of tobacco smoke exposure. 22、25、26 Essentially, these and other results demonstrate that at the single locus level, the methylation response to smoking can be better conceptualized as a product of both tobacco smoke exposure and genetic variation. These interaction effects appear to be widespread. Using a genome-wide approach, we recently demonstrated these smoking-associated genetic effects on DNA methylation on a genome-wide basis, showing that nearly one-quarter of all genes have genetically associated methylation changes in response to smoking (Dogan et al., submitted).

[0209] In contrast to the more easily conceptualized response of just one aspect of the methylome to a single environmental factor (smoking), the entire biological response of peripheral white blood cells (WBCs) to the wide variety of factors that contribute to CHD is more complex and potentially difficult to reproducibly capture. For example, at the RNA level, microRNAs prepared from blood samples have been shown to be highly responsive to CHD. 27、28 and mRNA 29 Although significant signatures for , have been described, their obvious utility as clinical tools has yet to come to fruition. Nevertheless, their partial success to date indicates that nucleic acids prepared from peripheral WBCs may carry a larger biological signature that may also be retrieved by a more systematic approach.

[0210] To that end, we detail the results of an integrated approach that incorporates commonly used machine learning algorithms together with both genome-wide epigenetic and genetic data from the Framingham Heart Disease Study.

[0211] method Framingham Heart Study. The Framingham Heart Study (FHS) is described in detail elsewhere. 30、31 The clinical, genetic, and epigenetic data included in this study are from the Offspring cohort. Specifically, this study included 2,741 of the 5,124 individuals from the Offspring cohort who 1) survived until the 8th surveillance cycle, which took place between 2005 and 2008, 2) consented to genetic studies, and 3) had genome-wide DNA methylation data in peripheral blood. FHS data were obtained through dbGAP (https: / / dbgap.ncbi.nlm.nih.gov). The University of Iowa Institutional Review Board approved all analyses described.

[0212] Genome-wide DNA methylation. After removing repeats, DNA methylation data were available for 2,567 individuals. The data were collected using the Illumina Infinium HumanMethylation450 BeadChips at either the University of Minnesota or Johns Hopkins University. 32 (San Diego, CA) array was used to profile genome-wide DNA methylation in the offspring cohort. The 485,577 probes in the array span 99% of the RefSeq genes, with an average of 17 CpG sites per gene, both within and outside CpG islands. 32 .

[0213] MethyLumi, WateRmelon and IlluminaHumanMethylation450k.db R packages 33Using this, filtering of probes, background correction, and adjustment for probe type were performed on the files of methylation intensity data (IDAT). Quality control was carried out for both sample and probe levels. For samples, samples having >1% of CpG sites with a detection p-value >0.05 were removed, while for CpG sites with a bead count <3 and / or >1% of samples with a detection p-value >0.05 were removed. After quality control, 2,560 unique samples and 484,125 CpG sites remained. Of these CpG sites, 472,822 were mapped to autosomes. Due to the finiteness of methylation beta values (0 ≦ β ≦ 1), a logistic transformation from beta values to M values (-inf < M value < inf) was performed using the beta2m R package, followed by conversion to z-scores using an R script 34 .

[0214] Genome-wide genotype. Genome-wide SNP data were profiled using Affymetrix GeneChip HumanMapping 500K (Santa Clara, CA) arrays. Among the 2,560 individuals remaining after quality control for DNA methylation, 2,406 individuals (1,100 males and 1,306 females) had genotype data. Again, quality control was performed at both sample and probe levels. PLINK 35 was used to examine samples for discordant sex information, heterozygosity rates greater than or less than two standard deviations from the mean, and missing SNP rates >0.03. As a result, a total of 111 samples were removed. Population stratification was also performed, and no individuals were excluded. To ensure that downstream analysis was not affected by related individuals, samples were also excluded if the value of identity by descent was >0.1875. This value is intermediate between second-degree and third-degree relatives. As a result of this criterion, a total of 696 individuals were removed, and 1,599 subjects remained for further analysis (722 males and 877 females). Minor allele frequency >1%, Hardy-Weinberg equilibrium p-value >10 -5Probes were retained if the missingness rate was <5%. After quality control, 403,192 SNPs remained (472,822 mapped to autosomes). SNPs were coded as 0, 1, or 2 per minor allele frequency.

[0215] Phenotype. For each individual, the following data were extracted from the FHS dataset: age, sex, systolic blood pressure (SBP), high-density lipoprotein (HDL) cholesterol level, total cholesterol level, hemoglobin A1C (HbA1c) level, self-reported smoking status, CHD status, and date of documented CHD.

[0216] Data analysis. To identify genome-wide DNA methylation changes associated with CHD risk factors and traditional modifiable CHD risk factors, Equation 1: Meth i ~ Age + Gender + Batch + X(1) Linear regression analysis was performed on R as shown in the formula: where X represents CHD risk factors or traditional modifiable CHD risk factors: SBP (smoking), HDL (total cholesterol) and diabetes. Batch represents the laboratory batch of DNA methylation.

[0217] Associations between DNA methylation and CHD or each risk factor were determined while adjusting for age, sex, and batch effects. Bonferroni correction for multiple comparisons with genome-wide α = 0.05 was performed for each regression analysis. 36 For each X, a total of 472,822 independent tests were performed, and therefore only Xs with a nominal p-value of 1e-07 (0.05 / 472822) were considered to be significantly associated at the genome-wide level.

[0218] Network analysis: STRING version 10 for symptomatic CHD 37Networks were generated using STRING to identify gene ontology (GO) pathways. The STRING database contains information on known and predicted physical (direct) and functional (indirect) associations between proteins. Networks included genes with at least one significant main effect DNA methylation locus after genome-wide Bonferroni correction for multiple comparisons. Networks were further reduced to include only nodes (proteins) with an edge (interaction) with a highest confidence interaction score of 0.9 or greater. PPI diagrams include nodes with at least one edge. STRING Version 10 was also used to determine GO enrichment pathways for the networks.

[0219] Training and testing datasets. The goal of this study was to develop an integrated genetic-epigenetic classifier to predict syndromic CHD. To achieve this, training and testing datasets were prepared. After DNA methylation and SNP quality control as described above, 1599 subjects remained. However, based on CHD status and due date of the 8th surveillance cycle, the number of individuals was reduced from 1599 to 1545 (694 men and 851 women), and these individuals constituted the training set.

[0220] To assess the generalizability of the trained model, data from 696 individuals who were removed for consanguinity (family identity >0.1875) were used. The CHD status and 8th survey cycle date of individuals in the test dataset as well as those in the training dataset were compared to ensure that only datasets with CHD status dates less than or equal to the 8th survey cycle date were retained. In doing so, the number of individuals in the test set was reduced from 696 to 669 (314 men and 355 women).

[0221] Variable reduction. The total number of genetic (SNP) and epigenetic (DNA methylation) probes remaining after quality control measures was 403,192 and 472,822, respectively. Due to the large number of variables (876,014 in total, excluding phenotype), we reduced the search space and minimized redundancy in predictors as follows:

[0222] Linkage disequilibrium-based SNP pruning is performed using PLINK 35 A window size of 50 SNPs, a window shift of 5 SNPs, and a pairwise SNP-SNP LD threshold of 0.5 was used. This reduced the number of SNPs from 403,192 to 161,474. To further reduce the number of SNPs, chi-squared p-values ​​were calculated between the remaining 161,474 SNPs and CHD status. SNPs with a chi-squared p-value <0.1 were retained for model training, resulting in 17,532 SNPs (~4%).

[0223] To reduce the number of DNA methylation loci, we first calculated the correlation between 472,822 CpG sites and CHD status. CpG sites were retained if the point biserial correlation was at least 0.1. A total of 138,815 CpG sites remained. We then calculated the Pearson correlation between those 138,815 sites. If the Pearson correlation between two loci was at least 0.8, the locus with the smaller point biserial correlation was discarded. Finally, 107,799 DNA methylation loci (approximately 23%) remained for training the model.

[0224] Class imbalance. Of the 1545 individuals in the training dataset, only 173 were diagnosed as having symptomatic CHD. Thus, the ratio of individuals with and without symptomatic CHD is approximately 1:8 (173:1372). This means that if we were to utilize data from all 1545 individuals simultaneously, the baseline prediction accuracy (for the majority class) where all individuals are classified as not having CHD would be approximately 89% (1372 / 1545). This indicates a dominant class imbalance in this dataset, which is quite common in medical datasets. It also suggests that accuracy is not an ideal performance metric. To address class imbalance, we undersampled individuals without CHD. 38 The 1372 individuals without CHD were randomly assigned to eight datasets: four with 171 individuals and four with 172 individuals (1372 individuals in total). In total, the eight datasets consisted of an equal number of individuals with CHD, 173. This now balanced the classes to a 1:1 ratio (i.e., 50% baseline accuracy) in each of the eight datasets.

[0225] Similarly, only 71 of the 669 test set individuals were diagnosed with CHD, indicating class imbalance, therefore 71 individuals without CHD were randomly selected to ensure a 1:1 case to control ratio.

[0226] Training and testing of the models. We used a stratified 10-fold cross-validation approach to train and test the models in Python on all eight datasets consisting of genetic, epigenetic, and phenotypic data. 40 Random Forest (RF) using scikit-learn independently in 39A classification model was constructed. SNPs with smaller chi-squared p-values ​​and methylation sites with greater correlation to CHD were systematically fed into the model. Feature importance, accuracy and AUC of the RF classifier were used to select variables important for prediction. Grid search was employed for 10-fold cross-validation hyperparameter tuning of the model. Performance metrics of the model were determined. The final model was saved for testing on the test dataset.

[0227] To compare the performance of our integrated genetic-epigenetic model with a model using traditional CHD risk factors as predictors, we took a similar approach to build a model on the training data and then tested it on the test dataset.

[0228] An alternative approach was implemented using the RandomForest™ package on R. Stratified sampling of the minority classes was performed using the "strata" and "sampsize" arguments. This is a simpler implementation of the undersampling approach described above. The number of trees (ntree) parameter of this alternative RF classifier was adjusted. This classifier was trained and tested using the same n=1545 training set and n=142 test set.

[0229] result The clinical characteristics of the 1545 subjects used in the primary analysis in this study are shown in Table 4. There were more women (n = 851) than men (n = 694), and all were of Northern European ancestry. A total of 115 men (approximately 17%) and 58 women (approximately 7%) were diagnosed with symptomatic CHD. Subjects with symptomatic CHD were, on average, older, tending to be in their early 70s, in contrast to subjects without symptomatic CHD, who tended to be in their mid-60s.

[0230] Table 4. Demographics and CHD risk factors of 1,545 individuals. TIFF0007672192000007.tif184128SBP: Systolic blood pressure HbA1c: Hemoglobin A1c

[0231] Mean HDL and total cholesterol levels were higher in women and subjects without symptomatic CHD. The overall mean for total cholesterol was <200 mg / dL, but only women without symptomatic CHD had HDL cholesterol levels >60 mg / dL. More importantly, the mean HDL to mean total cholesterol ratios were 1:3.4 and 1:3.5 for men with and without symptomatic CHD, respectively, and 1:2.9 and 1:3.1 for women with and without symptomatic CHD, respectively. The target ratio of total cholesterol to HDL cholesterol for prevention of cardiovascular disease was <4.5 for men and <4.0 for women. 41 .

[0232] Subjects diagnosed with symptomatic CHD had higher HbA1c levels on average (6%) than subjects not diagnosed with symptomatic CHD (5.7%). However, whereas women with CHD had higher SBP than women without, the opposite was true for men. All SBP means were greater than 120 mmHg.

[0233] Another well-known risk factor for CHD is smoking. Based on self-reported current smoking status, a greater proportion of smokers with symptomatic CHD than those without symptomatic CHD were found among men but not women. However, the methylation status of the smoking biomarker (cg05575921) indicates that both men and women with symptomatic CHD are actually more likely to smoke than men and women without symptomatic CHD.

[0234] Regression Analysis. As a first step of the analysis, CHD status of the 1545 subjects was regressed against age, sex, cg05575921, SBP, HDL cholesterol, total cholesterol, and HbA1c percent. A summary of the regression output for each risk factor is shown in Table 5. The analysis suggests that all traditional risk factors except SBP and HDL cholesterol are significantly associated with CHD status at the 0.05 significance level. More importantly, the trends in slope suggest that the incidence of symptomatic CHD is higher in 1) males, 2) older individuals, 3) individuals with lower total cholesterol, 4) individuals who are demethylated at cg05575921 (i.e., more smoking), and 5) individuals with higher HbA1c levels.

[0235] Table 5. Regression parameters of CHD risk factors for symptomatic CHD TIFF0007672192000008.tif37139SBP: Systolic blood pressure HbA1c: Hemoglobin A1c

[0236] As a next step, we performed a regression analysis of the relationship between syndromic CHD and genome-wide DNA methylation. After Bonferroni correction, 11,497 methylation sites (2.4%) maintained a significant association with syndromic CHD. These methylation sites were mapped to 6,319 genes. The top 30 sites are shown in Table 6. All significant sites are provided in Figure 16.

[0237] Table 6. Top 30 significant CpG sites associated with symptomatic CHD TIFF0007672192000009.tif170146*All nominal p values ​​were adjusted for multiple comparisons using the Bonferroni method.

[0238] Due to the large number of genes, network analysis and functional enrichment analysis of the network were performed using data from the top 1000 genes. The network consists of 952 proteins represented by nodes and 1,144 interactions represented by edges. The expected number of edges with PPI enrichment p-value 0, suggesting that interactions between proteins and the network may have biological relevance, was 634. The average node degree and clustering coefficient were 2.4 and 0.85, respectively. The network is shown in Figure 3. The top 10 pathways in the network are shown in Table 7.

[0239] Table 7. Top 10 significant PPI network pathways associated with symptomatic CHD TIFF0007672192000010.tif65157PPI: Protein-protein interaction

[0240] Regression analysis revealed that 44,108 (9.3%) methylation sites were significantly associated with cg05575921, SBP, HDL cholesterol, total cholesterol, and HbA1c, respectively. The top results for cg05575921, HDL, total cholesterol, and HbA1c analyses are shown in Tables 8 to 11.

[0241] Table 8. Top 30 significant CpG sites associated with cg05575921 after Bonferroni correction TIFF0007672192000011.tif165156* All nominal p values ​​were adjusted for multiple comparisons using the Bonferroni method.

[0242] Table 9. A total of 32 significant CpG sites associated with HDL cholesterol after Bonferroni correction. TIFF0007672192000012.tif179166*All nominal p values ​​were adjusted for multiple comparisons using the Bonferroni method.

[0243] Table 10. Top 30 significant CpG sites associated with total cholesterol after Bonferroni correction TIFF0007672192000013.tif169143*All nominal p values ​​were adjusted for multiple comparisons using the Bonferroni method.

[0244] Table 11. A total of six significant CpG sites associated with HbA1c after Bonferroni correction TIFF0007672192000014.tif42141*All nominal p values ​​were adjusted for multiple comparisons using the Bonferroni method.

[0245] To understand the mapping of DNA methylation sites of significant symptomatic CHD with that of its risk factors, Figures 4 and 5 were drawn. The Venn diagram in Figure 4 shows the overlap in methylation probes between symptomatic CHD and its risk factors, while Figure 5 shows the overlapping genes that map to at least one of the probes. As shown in Figure 5, the top three intersections in DNA methylation-related genes are between symptomatic CHD and smoking (5229), between smoking and total cholesterol (15), and between symptomatic CHD, smoking and total cholesterol (13). One gene, DHCR24, was significantly associated with symptomatic CHD and all risk factors.

[0246] Integrated genetic-epigenetic random forest analysis. Eight RF models were constructed on eight datasets consisting of genetic, epigenetic, age and gender data from 1545 subjects in the training dataset. Standard scikit-learn RF parameters were used to determine significant SNPs and DNA methylation loci. Based on the average accuracy and AUC of the eight classifiers and the Gini index of each variable, four CpG sites (cg26910465, cg11355601, cg16410464 and cg12091641), two SNPs (rs6418712 and rs10275666), age and gender were retained for prediction. All eight models were refitted to the training dataset with adjusted parameters (maximum feature count, minimum number of samples for each branch, information gain criterion, maximum tree depth, number of trees). The performance metrics of these stratified 10-fold cross-validation models are shown in Table 12. As shown in this table, the accuracy ranged between 70-80% among these eight models, which is an increase of between 20-30% from the baseline of 50% accuracy. More importantly, the sensitivity of the models ranged from 70-82%, while the specificity ranged from 70-79%. The ROC AUC of the eight models ranged from 0.77 to 0.87. The AUC of the 10-fold ROC of the best performing model (Model 7) is shown in Figure 6. All eight models were saved for testing on the test dataset.

[0247] Table 12. Performance metrics of 10-fold cross-validation of eight integrated genetic-epigenetic models TIFF0007672192000015.tif49128

[0248] The demographics and CHD risk factors of the individuals in the test dataset are summarized in Table 13. Of the 54 women and 88 men, 22 women (about 41%) and 49 men (about 56%) were diagnosed with symptomatic CHD. Individuals with symptomatic CHD were, on average, older, with men tending to be in their late 60s and women in their early 70s. Men and women without symptomatic CHD tended, on average, to be in their late 50s and mid 60s, respectively. Unlike men, the mean ages of women with and without symptomatic CHD were comparable between the training and test datasets.

[0249] Table 13. Demographics and CHD risk factors of 142 individuals in the study dataset TIFF0007672192000016.tif125128SBP: Systolic blood pressure HbA1c: Hemoglobin A1c

[0250] The total cholesterol means were all <200 mg / dL, and the women only had mean HDL cholesterol levels >60 mg / dL. The ratios of mean HDL cholesterol to mean total cholesterol were 1:3.1 and 1:3.7 for men with and without symptomatic CHD, respectively, and 1:2.9 and 1:3.1 for women with and without symptomatic CHD, respectively. Again, the ratios were more similar in both data sets for women than for men. However, the ratios were all lower than the target ratios of total cholesterol to HDL cholesterol for the prevention of cardiovascular disease, i.e., <4.5 for men and <4.0 for women. 41 .

[0251] In the testing dataset, women tended to have higher HbA1c percentiles than men. Additionally, women with symptomatic CHD had a mean HbA1c >6%. Women had higher SBP than men. All SBP averages were >120mmHg. Similar to the training dataset, there were more smokers without symptomatic CHD than smokers with symptomatic CHD based on self-reported current smoking status. However, when considering the smoking biomarker cg05575921, men tended to be more demethylated than women.

[0252] An ensemble of eight models was used to classify CHD in the test dataset. An individual was classified as having CHD if at least four of the eight models voted in favor of CHD. Of the 142 individuals in the test dataset (71 with and 71 without symptomatic CHD), the CHD status of 110 individuals was correctly predicted, resulting in an accuracy of 77.5%. The confusion matrix for the predictions is shown in Table 14. The sensitivity and specificity of the ensemble test set were 0.75 and 0.80, respectively.

[0253] Table 14. Confusion matrix of integrated genetic-epigenetic ensemble for test dataset TIFF0007672192000017.tif23128

[0254] Traditional CHD Risk Factor Models. To compare the performance of our integrated genetic-epigenetic model with that of traditional CHD risk factors in predicting CHD status, another eight RF models were constructed using age, sex, SBP, HbA1c, total cholesterol, self-reported smoking, and HDL cholesterol as predictors. Again, eight RF models were constructed on the training dataset using the adjusted parameters and tested on the test dataset. The performance metrics of the eight models are summarized in Table 15. The accuracy of these models on their respective training datasets ranged from 70 to 76%, while the sensitivity and specificity ranged from 67 to 74% and 72 to 79%, respectively. The ROC AUC ranged from 0.72 to 0.79. Although the accuracy and specificity were quite similar to the integrated genetic-epigenetic model, the traditional risk factor model underperformed in terms of sensitivity and ROC AUC. The AUC of the 10-fold ROC of the best performing model (model 7) among the eight models is shown in Figure 7. When the eight models were tested on the test data set, the accuracy of the test was 64.8%, which is about 13% less than our integrated genetic-epigenetic set. However, the more important metric is the sensitivity, since it indicates the degree to which individuals with CHD are correctly classified. The sensitivity on the test data set was only 41%, which is 24% less than our integrated genetic-epigenetic set. However, the specificity of the traditional risk factor set was 0.89. The confusion matrix is ​​shown in Table 16.

[0255] Table 15. Performance metrics for 10-fold cross-validation of eight traditional risk factor models TIFF0007672192000018.tif49128

[0256] Table 16. Confusion matrix for the traditional risk factor set for the test dataset TIFF0007672192000019.tif23128

[0257] Alternative Random Forest Models. To determine whether our ensemble approach of eight models outperforms the single RF model as described in the Methods section, a single RF model was constructed in R with stratified sampling based on minority classes. This model also included the same four CpGs, two SNPs, age, and gender. The classifiers were tuned and the classifier with the highest sensitivity was selected (ntree=500). The training accuracy, AUC, sensitivity, and specificity of this model were 82%, 0.83, 0.68, and 0.83, respectively. Although the accuracy, AUC, and specificity of this model were similar to our ensemble model, the ensemble model was clearly more sensitive. When tested on the test set, the single RF model performed with an accuracy, sensitivity, and specificity of 76%, 0.66, and 0.86, respectively, demonstrating the increased sensitivity but not specificity provided by the ensemble approach. The comparison between this alternative approach and the ensemble approach is based on sensitivity rather than specificity because, considering the application of the classifier to predict CHD, it is more important to maximize true positives than true negatives. In other words, the negative impact of having false negatives is much larger than false positives. However, one reason why the sensitivity of the ensemble (ntree=170,000) cannot be directly compared to that of this single-RF classifier (ntree=500) is that the effective number of trees in the ensemble is much larger than this classifier. Nevertheless, a comparison can be made between one classifier in the ensemble with 20,000 trees and the alternative-RF classifier with the same number of trees. The average accuracy, AUC, sensitivity and specificity of the classifier from the ensemble with 20,000 trees were 80%, 0.87, 0.82 and 0.77, respectively. Similarly, the accuracy, AUC, sensitivity and specificity of the alternative RF classifier with 20,000 trees were 82%, 0.83, 0.67 and 0.83. As in the previous comparison, the ensemble model performed better in terms of sensitivity than specificity.

[0258] Although age and sex were included because they are two unmodifiable risk factors for CHD, we refitted the single-RF model without age and sex to demonstrate that performance is not driven solely by these two factors. In the absence of age and sex in the model, the training accuracy, AUC, sensitivity and specificity were 81%, 0.80, 0.65 and 0.83, respectively. On the test data set, the performance of this model was 78%, 0.68 and 0.89 for accuracy, sensitivity and specificity, respectively. Thus, age and sex are not solely responsible for the performance of the integrated genetic-epigenetic model. Using the traditional risk factors from the training data set, the performance of this alternative RF model was 77%, 0.77, 0.60 and 0.79 for accuracy, AUC, sensitivity and specificity, respectively. On the test data set, its performance was 69%, 0.61 and 0.77 for accuracy, sensitivity and specificity, respectively.

[0259] As depicted by the partial dependence plot in FIG. 8, this genetic-epigenetic model was also used to show that the use of the RF model provides an additional advantage in capturing possible G×M and M×M interactions. Finally, DNA methylation sites and genotypes were permuted to compare the performance of a model consisting of four randomly selected CpG sites and two randomly selected SNPs with our integrated model and conventional risk factor models using the training data set. A two-dimensional histogram of the sensitivity and specificity of 10,000 permutations is shown in FIG. 9. The maximum sensitivity and maximum specificity among these permutations were 0.62 and 0.87, respectively. The training sensitivity and specificity of the single conventional risk factor model, 0.60 and 0.79, respectively, are well within the range of the permutation sensitivity and specificity. The training sensitivity and specificity of the single integrated genetic-epigenetic model, 0.68 and 0.83, respectively, suggest that the sensitivity, but not the specificity, deviates from the permuted values.

[0260] Consideration A better understanding of the relationship between epigenetic changes and the pathogenesis of cardiovascular disease is essential to develop improved diagnostics and therapeutics. To our knowledge, we are the first group to investigate the relationship between DNA methylation, as quantified using the Illumina 450k array, and CHD. Thus, comparisons with our results are limited. Nevertheless, our analysis demonstrates that the epigenetic signature for CHD overlaps substantially with the signature of cumulative smoking. This is consistent with the strong and well-established relationship between smoking and CHD risk, with cigarette smoking responsible for approximately 30% of annual CHD-related deaths in the United States. 42、43 This is not a claim made lightly. Smoking cessation may be one of the most beneficial, yet underutilized, common interventions in clinical medicine, and it has also been shown to substantially reduce the risk of death in people with CHD. 44、45 .

[0261] Interestingly, but not surprisingly, DNA methylation analysis of all other risk factors described in our study indicates a broad effect of smoking in remodeling the epigenome. Previous studies of the relationship between atherosclerosis and lipid levels, diabetes, or hypertension have demonstrated the impact of smoking on these clinical measures. 46-48 Our analysis not only identified changes in DNA methylation associated with HDL cholesterol, total cholesterol, and HbA1c, but also delineated specific loci whose epigenetic signatures are modified by smoking and associated with increased risk of CHD. Pending confirmation of the results by others, the increased precision afforded by extending the results to a diverse set of ethnic subjects may aid in identifying specific therapeutic interventions for CHD at the individual level.

[0262] An additional application of the methylation signature of CHD and its risk factors is as an alternative approach to assessing the risk of CHD. This idea is particularly attractive given the challenges and limitations of using traditional risk factors to predict the risk of CHD. For example, most studies have used self-reported smoking status, but others and the present inventors have shown that self-reported smoking status is unreliable in more clinical / high-risk populations. 49-52 These previous results are of particular importance given the discrepancies between self-reporting and cg05575921 methylation in the offspring cohort used in this study. Another test routinely performed to assess CHD risk is the fasting serum lipid panel, which evaluates total cholesterol, HDL cholesterol, LDL cholesterol, and triglyceride levels. While studies have shown that the ratio between total cholesterol and HDL cholesterol is particularly predictive of CHD risk, 53、54 Other studies have also shown that information from additional markers, such as C-reactive protein, is needed to improve predictive power. 55 Because these DNA methylation measures are more cumulative and less affected by daily increases and decreases in diet, they may also be able to more precisely confine the relative contribution of each of these metabolic / transcriptional pathways to CHD pathogenesis.

[0263] Over the years, the identification of traditional CHD risk factors has led to the development of several multivariate risk models. The Framingham Heart Study was pioneering in this effort, developing the Framingham Risk Score for CHD. 56The algorithm was developed using traditional risk factors (age, sex, total or LDL cholesterol, SBP, diabetes, and current smoking) and the FHS cohort of individuals of European ancestry. Thus, as expected, the model performed well for white men and women, but generalization to all other ethnic groups was difficult. Specifically, in a study validating the algorithm in an ethnically diverse cohort, the predictive model also held true for black men and women, but overestimated risk for Japanese American and Latino men and Native American women. 57 Thus, there is a need for algorithms that can be used by all members of modern society.

[0264] One plausible reason for the lack of generalizability is the possible confounding effects of genetic variation. The concept of the potential for genetic confounding of epigenetic signals is widely accepted. 58 Therefore, the goal of our study was to integrate genetic and epigenetic data to develop a classifier for predicting CHD as an alternative to existing algorithms currently available. This approach of mining predictive signals from large and complex genetic and epigenetic datasets is enabled by advances in high-performance computational systems. Computational techniques such as machine learning have been successfully employed in the fields of genomics and epigenomics. 59、60 Logistic regression is a commonly used method for developing binary classification models in medical applications and has been used to analyze microarray data. 61However, it lacks the ability to capture underlying complex nonlinear relationships. Therefore, algorithms capable of detecting complex relationships such as interactions between genetic variants and DNA methylation have additional advantages. In our study, the use of random forest ensembles allowed for highly accurate, sensitive and specific classification of individuals with CHD. However, since some genetic risk variants may be sorted with ethnic background and not mapped to pathways associated with traditional risk factors, it seems necessary to build, test and extend these random forest approaches using subjects from all ethnic groups to develop the most generalized prediction tool. 62 .

[0265] While similar integrated genetic and epigenetic studies are not available for comparison, our integrated model clearly outperforms classifiers using the Framingham score risk factors. Conventional risk factor models have demonstrated the limited predictive value of these risk factors, as shown by several studies. 63-65 Moreover, in a study of more than 2000 older black and white adults, the Framingham risk score was only able to distinguish subjects who experienced a CHD event from those who did not after 8 years of follow-up, with a C-index of 0.577 and 0.583 for women and men, respectively. 66 Traditional risk factors may perform poorly due to time variability of factors such as serum cholesterol levels and the use of a single blood pressure measurement instead of an average recorded throughout the day. 67、68 .

[0266] As demonstrated in this paper, there are several approaches to construct a classifier. Comparing the two methods presented in this paper, the ensemble model performed better than the single RF model in terms of sensitivity and vice versa for specificity. The reason why we prefer a more sensitive model is simple. For classification of diseases such as CHD, false positives require further testing, but false negative results may be more harmful to the patient. Nevertheless, a test with high sensitivity and specificity is ideal. To achieve that, a larger sample consisting of a wide variety of ethnic groups encompassing both genders is required. Also, whereas we used the RF algorithm, there are many other algorithms, such as support vector machines, that can be used as the underlying algorithm for the classifier. Nevertheless, our RF model clearly shows the nonlinearity between methylation sites and SNPs as shown in the partial dependence plot. Furthermore, we want to make it clear that the combination of methylation sites and SNPs in our ensemble is only one of many possible combinations with high predictive potential. Based on the permutation results, we demonstrate that the variable reduction step performed to enrich for highly predictive methylation and SNP probes provides an edge in terms of sensitivity. However, as the pool of diverse samples increases, highly predictive yet generalizable classifiers are required.

[0267] Our analysis did not take into account possible medication effects. This is noteworthy, because the current armamentarium of cholesterol-lowering drugs can have a dramatic effect on the level of certain risk factors, such as serum cholesterol, which is associated with the risk of CHD. Indeed, the presence of these medications may be the reason why the subjects in the training set with CHO actually have lower serum cholesterol levels than those in the training set without CHD. Unfortunately, it is very difficult to incorporate these types of data into the current analysis approach for several reasons. Additionally, even if the subjects' self-reports of prescriptions are accurate, important information needed to explain the effect, such as medication compliance and length of treatment history, is not available. However, in the future, if we are to fully understand the effect of medical interventions on epigenetic signatures, it will be crucial to have data such as "pill count" and serum drug level information.

[0268] Additionally, our study has several other limitations. First, our study only includes individuals of European ancestry. However, the incorporation of genetic variants into our model allows for generalizability across ethnic groups. Nevertheless, additional studies are required to demonstrate this. Second, while our approach predicts symptomatic CHD, the goal is to use this study as a proof of concept to build a multivariate model capable of predicting the risk of a first CHD event and subsequently the risk of recurrent CHD events. Further exploration in prospectively biosampled cohorts will be necessary to achieve that goal. Furthermore, it is important to mention that there are advantages to this integrated genetic-epigenetic approach. The use of traditional risk factors to calculate risk requires cumbersome testing procedures, collection of significant amounts of blood, and multiple laboratory tests. Perhaps the need for these often cumbersome tests and procedures would be greatly reduced by using a single genetic-epigenetic assay procedure that uses one microgram or less of DNA. More importantly, pathways associated with specific epigenetic loci with high predictive value may also be extremely useful in directing therapeutic intervention, managing risk factors and monitoring efficacy of treatments and lifestyle modifications.

[0269] References for Example 4 TIFF0007672192000020.tif70159TIFF0007672192000021.tif245160TIFF00076721920 00022.tif242151TIFF0007672192000023.tif242152TIFF0007672192000024.tif242152

[0270] Example 5 The effect of methylation and G x methylation in predicting cardiovascular disease: stroke and congestive heart failure Methylation-based biomarkers are gaining increasing clinical appeal for use in guiding diagnostics and therapeutics. Currently, Cologuard, an assay that quantifies DNA methylation in human DNA found in stool samples, has been approved by the FDA for the detection of colon cancer (Lao and Grady 2011). In addition, Smoke Signature™ (Philibert, Hollenbeck et al. 2016), a DNA methylation assay that uses blood-derived DNA to detect tobacco consumption, is available in the research market and is in preparation for FDA submission. In an attempt to identify CpG loci whose methylation status predicts cardiovascular disease, several researchers have used genome-wide approaches in combination with clinical diagnostics. In particular, Brenner and coworkers (Breitling, Salzmann et al. 2012) identified residue cg03636183 of F2RL3 as a biomarker for cardiovascular disease. Unfortunately, these analyses have been shown to be completely confounded by incomplete smoking status information and did not take into account the possibility of confounding genetic variance. Indeed, the coronary heart disease signal of cg03636183 disappears when a biomarker approach that fully accounts for smoking intensity is used (Zhang, Schoettker et al. 2015). Furthermore, using genome-wide methylation and genetic analysis in combination with biomarker-guided smoking assessment, we recently analyzed data from a large cohort of subjects providing heart disease information. We showed that methylation status in the context of genetics, embodied by methylation-genotype interactions, indeed contributes better to the prediction of coronary heart disease, independent of smoking intensity status, and that the use of an algorithm that combines regional genetic variation and methylation significantly improves the prediction of coronary heart disease (CVD, Dogan et al., submitted).

[0271] However, CVD is only one of the three major forms of cardiovascular disease (CVD). Stroke and congestive heart failure (CHF) are also prominent forms of CVD. In these examples, we extend our previous research on CVD to show how the combination of genetic variations embodied by SNPs and epigenetic markers embodied by Illumina methylation probes predicts stroke or CHF.

[0272] summary Congestive heart failure (CHF) and stroke are two of the three common types of cardiovascular disease (CVD). CHF and stroke affect a large number of Americans. Although preventive measures such as avoiding smoking can be taken to reduce the risk of stroke and CHF, options for early detection of risk for these diseases are limited. However, in recent years, the field of epigenetics has provided an alternative approach to understand complex diseases. Specifically, DNA methylation signatures may present an opportunity to develop robust clinical tests for pre-occurring CHF and stroke. The ability to utilize DNA methylation alone and generalize it to a wide variety of populations may also be limited by the presence of confounding genetic effects. Thus, we integrated genetic and epigenetic data from the Framingham Heart Disease Study to uncover SNPs and DNA methylation sites that collectively increase the predictive ability of CHF and stroke. Our preliminary analysis suggests that the incorporation of three DNA methylation sites and three SNPs has the ability to classify CHF status with receiver operating characteristic (ROC) curve area under curve (AUC) of 0.78 and 0.81 in the main effect and interaction effect models, respectively. In assessing the parameters of these models, we show that both DNA methylation and SNPs are highly predictive of CHF status when run simultaneously. Similarly, the stroke ROC curve AUC of 0.85 and 0.86 for the main effect and interaction effect models, respectively, demonstrates the importance of integrating genetic and epigenetic effects. Although these models are not optimized and were developed with a relatively small sample size of CHF and stroke, we believe that a more optimized version of this algorithm that accounts for genetic and epigenetic effects, developed using a larger cohort, can also significantly improve our ability to predict the risk of CHF and stroke before their occurrence. The inventors are also confident that the presence of genetic information in the algorithm enables its generalizability to different ethnic groups.

[0273] Introduction Cardiovascular disease (CVD) includes three distinct diagnostic entities: coronary heart disease (CVD), stroke, and congestive heart failure (CHF). CVD is itself the leading cause of death in the United States, while stroke is the fourth leading cause of death (Centers for Disease Control and Prevention). Over the past 50 years, several medications and devices have been developed to treat CVD. Unfortunately, hundreds of thousands of Americans continue to die each year due to not realizing the presence of CVD until a fatal thromboembolic or cardiac event. More effective screening methods for CVD could conceivably result in the prevention of some of these deaths (Mozaffarian, Benjamin et al. 2016). However, currently the cumbersome nature of certain techniques, such as fasting fat panels, and / or the limited predictive power of others, such as electrocardiograms and C-reactive protein levels, limit the effectiveness of current approaches in identifying CVD (Buckley, Fu et al. 2009, Auer, Bauer et al. 2012, Mozaffarian, Benjamin et al. 2016).

[0274] Some researchers have proposed that genetic approaches could provide another potential means of preventing CVD-related morbidity and mortality (Paynter, Ridker et al. 2016). Using whole-exome and genome sequencing techniques, several variants that predispose to CVD have been identified. The relative risks conferred by many of these variants are often substantial, and their presence can be useful to guide prevention and treatment strategies (Mega, Stitziel et al.). However, with isolated exceptions, variants with large effect sizes tend to be rare and often population-specific, and their presence is not characteristic of the disease at hand (Traylor, Farrall et al. 2012, Paynter, Ridker et al. 2016). Thus, at present, genetic approaches are not commonly used in general medical practice to assess the presence or absence of current CVD.

[0275] Alternatively, other researchers have proposed that epigenetic techniques may be useful for the assessment of CVD (Sharma, Kumar et al. 2008, Gluckman, Hanson et al. 2009, Breitling, Salzmann et al. 2012). This suggestion has strong face validity since replicate peripheral leukocyte DNA methylation signatures for the presence of type 2 diabetes, smoking and alcohol use have been developed (Monick, Beach et al. 2012, Toperoff, Aran et al. 2012, Zeilinger, Kuehnel et al. 2013, Philibert, Penaluna et al. 2014). Notably, Brenner and coworkers used this approach to propose that DNA methylation at cg03636183, a CpG residue found in the coagulation factor II (thrombin) receptor-like 3 (F2RL3), predicts cardiovascular risk (Breitling, Salzmann et al. 2012, Zhang, Yang et al. 2014). Although this is a highly biologically plausible result, their subsequent work demonstrated that the CVD-related signal at cg03636183 completely co-segregates with smoking status as indicated by DNA methylation at cg05575921 (Zhang, Schoettker et al. 2015), a CpG residue found in the aryl hydrocarbon receptor repressor (AHRR) whose strong predictive power for smoking status has been demonstrated in multiple studies (Andersen, Dogan et al. 2015).

[0276] However, the initial intriguing results of cg03636183, unable to independently identify additional risk beyond that conferred by smoking alone, do not mean that methylation approaches to assess the presence of CVD or other forms of CVD are doomed to fail. Instead, they suggest that a successful approach needs to be more fine-grained, and that it would be useful to rethink our conceptualization of the relationship between methylation status and CVD. For example, the results from Brenner's group strongly suggest that methylation algorithms for predicting ongoing CVD should include an indicator of smoking status. Given the fact that smoking is the largest preventable risk factor for CVD (Center for Disease Control 2005), this is highly logical. However, in addition, they may need to consider that the long-term effects of exposure to environmental risk factors such as smoking or other cardiac risk factors such as hyperlipidemia may be obscured by gene-environment interactions.

[0277] The role that gene-environment interaction (G×E) effects play in moderating vulnerability to disease is perhaps better recognized in the case of behavioral sciences. The basic premise of G×E effects is that environmental influences during developmentally sensitive periods alter the biological traits of lineages in the context of genes, such that high vulnerability to disease exists even in the absence of environmental factors in the future (Yang and Khoury 1997). Importantly, due to confounding by genetic variables, direct effects of environmental variables are generally undetectable. Instead, direct effects can only be detected when considered in the context of genetic variation. Although the strength of some G×E results remains controversial, many researchers continue to emphasize the importance of these G×E effects in the pathogenesis of diverse behavioral disorders such as depression, post-traumatic stress disorder, and antisocial behavior (Caspi, McClay et al. 2002, Caspi, Sugden et al. 2003, Kolassa, Ertl et al. 2010).

[0278] The physical basis of these G×E effects is thought to be variable. For example, at the anatomical level, the G×E effects on behavioral impairments can be manifested by changes in synaptic structures (McEwen 2007). However, at the molecular level, the physical manifestation of the G×E effects is more uncertain. However, some researchers have suggested that changes in DNA methylation could be one potential mechanism by which the physical effects of the G×E effects are conveyed (Klengel, Pape et al. 2014).

[0279] Interestingly, it has been known for many years that behavior-related changes in the environment can alter DNA methylation, and the extent of these changes is influenced by genetic variation. In our initial candidate gene studies, we showed that smoking altered DNA methylation in the promoter region of monoamine oxidase A (MAOA), a key regulator of monoaminergic neurotransmission, and that genotypes at well-characterized promoter-associated variable nucleotide repeats (VNTRs) altered methylation percentages in both the presence and absence of smoking (Philibert, Gunter et al. 2008, Philibert, Beach et al. 2010). Subsequently, Volkow and coworkers showed that the methylation changes at these loci were functional (Shumay, Logan et al. 2012).

[0280] In current terminology, the effects of smoking or VNTR on basal DNA methylation are today referred to as genotype-methylation interaction effects. When we performed our first genome-wide study, these MAOA interaction effects had consequences on our ability to detect their association with smoking. Regardless of the magnitude of changes that smoking induces in DNA methylation in response to smoking, probes surrounding the MAOA VNTR are not included in the higher ranking probes, even in studies of DNA from subjects of only one gender. Other observations from the first study are instructive as well. First, the regional methylation response to smoking was not homogenous. Factorial analysis of the methylation status of 88 CpG residues in the promoter-associated island showed that increased methylation in some regions of the island could be associated with demethylation in other regions (Beach, Brody et al. 2010). Finally, the effect of smoking on DNA methylation was not static. There was a tendency for the signature to collapse in the future (Philibert, Beach et al. 2010). Thus, it was apparent from these early studies that genetic variation in the MAOA promoter could alter the effects of environmental factors on regional DNA methylation signatures in complex ways.

[0281] Subsequent studies suggest that many of these distinct complexities in response to smoking are evident at the genome-wide level. For example, it is evident that genetic variation influences the magnitude of the methylation response at the genome-wide level, and that these interaction effects can impair the ability to replicate results at a given locus in a pool of subjects of different ancestries when attempting to replicate signatures from signatures of various ancestries (Tsaprouni, Yang et al. 2014, Dogan, Xiang et al. 2015). Second, and equally important, reversion of methylation signatures can be complex (Tsaprouni, Yang et al. 2014, Guida, Sandanger et al. 2015). Guida and colleagues specifically investigated the epigenomic response to smoking cessation in DNA from 745 pooled subjects, found two classes of CpG sites whose methylation signatures reverted over time and those that did not, and concluded that at the genome-wide level, "the dynamics of methylation changes after smoking cessation are driven by smoking-induced changes of differential and site-specific magnitude that are independent of smoking intensity and duration" (Guida, Sandanger et al. 2015). Collectively, a substantial body of evidence suggests that genome-wide signatures to smoking are only partially reversible, and that the majority of irreversible changes may be completely hidden within interaction effects.

[0282] Since smoking is a major risk factor for CVD in general, and stroke and CVD in particular, this also suggests that some of the smoking-induced risk present in the epigenome that modulates risk for CVD may be somewhat irreversible and hidden by interaction effects. Additionally, since smoking is only one of several factors that can alter CVD risk, and these other factors may also have complex epigenetic signatures, it is possible that interrogation of DNA methylation of peripheral WBCs may reveal relatively stable interaction effects that modulate CVD risk.

[0283] So, in summary, using either genetic or epigenetic information for predicting various forms of CVD does not perform well, but it is possible that combinations of these measures, especially those that produce interaction effects, may perform better.

[0284] In this brief report, we used a regression analysis approach and epigenetic and genetic resources from 324 subjects participating in the Framingham Heart Study to test whether a combination or their interaction effects of environmental information (methylation) and genetic information (SNPs) could better contribute to algorithms to predict CVD.

[0285] method Framingham Heart Study. The data used in this study were obtained from participants in the Framingham Heart Study (FHS) (Dawber, Kannel et al. 1963). The FHS is a longitudinal study with the goal of understanding cardiovascular disease (CVD) risk and consists of several cohorts, including the original cohort, the offspring cohort, the omnicohort, the third generation cohort, the new offspring spouse cohort, and the second generation omnicohort. Specifically, the offspring cohort, which began in 1971 and consists of the children of the original cohort and their spouses, was used in this study. This cohort consists of 2,483 men and 2,641 women (5,124 in total) (Mahmood, Levy et al. 2014). The specific analyses described in this brief were approved by the University of Iowa Institutional Review Board.

[0286] Genome-wide DNA methylation. Of the 5,124 individuals in the Offspring cohort, only 2,567 individuals (excluding duplicates) with DNA methylation data were considered. These individuals were included in the DNA methylation study because they participated in the Framingham Offspring Round 8, consented to the genetic study, had buffy coat samples, and had sufficient quantity and quality of DNA for methylation profiling. Round 8 was conducted between 2005 and 2008. Genomic DNA extracted from the white blood cells of these individuals was bisulfite converted and then genome-wide DNA methylation was profiled using the Illumina HumanMethylation450 BeadChip (San Diego, CA) at either the University of Minnesota or Johns Hopkins University. The intensity data (IDAT) files of the samples, together with their slide and array information, were used to perform DASEN normalization using the MethyLumi, WateRmelon and IlluminaHumanMethylation 450k.db R packages (Pidsley, Y Wong et al. 2013). DASEN normalization performs probe filtering, background correction and probe type adjustment. Samples were removed if they contained >1% CpG sites with a detection p-value >0.05. CpG sites were removed if they had a bead count of <3 and / or >1% of the samples had a detection p-value >0.05. After DASEN normalization, 2,560 samples and 484,241 sites (484,125 CpG sites) remained. CpG sites were grouped by chromosome. Of these CpG sites, 472,822 were mapped to autosomes. Methylation β values ​​were converted to M values ​​using the beta2m R function in the Lumi package and subsequently converted to z scores using an R script (Du, Kibbe et al. 2008).

[0287] Genome-wide genotypes. Of the 2,560 individuals remaining after DNA methylation quality control, 2,406 (1,100 males and 1,306 females) had genome-wide genotype data from the Affymetrix GeneChip HumanMapping 500K Array Set (Santa Clara, CA). This array is capable of profiling 500,568 SNPs in the genome. Quality control was performed at both the sample and SNP probe levels in PLINK. The first quality control step involved identifying individuals with discordant gender information. None were identified. Next, individuals with heterozygosity rates greater or less than the mean ± 2 SD and proportions of missing SNPs > 0.03 were excluded. Related individuals were also excluded if their ancestry identity value was > 0.185 (intermediate between second-degree and third-degree consanguinity). After these sample-level quality control steps, 1,599 individuals remained (722 men and 877 women). At the probe level, minor allele frequencies >1% and Hardy-Weinberg equilibrium p-values ​​>10 -5 and SNPs with a SNP missingness rate <5% were retained. A total of 403,192 SNPs remained after these quality control steps. Using the recode option in PLINK (Purcell, Neale et al. 2007), genotypes were coded as 0, 1, or 2 per minor allele frequency.

[0288] Phenotype. For individuals in the study, their stroke and congestive heart failure (CHF) status was extracted. Because biomaterial for DNA methylation was collected during the 8th survey cycle of the Offspring cohort, only individuals with an incident date of stroke or CHF before this 8th survey were included. Based on this criterion, a total of 1,540 and 1,562 individuals remained for CHF and stroke analyses, respectively.

[0289] Of 1,540 subjects available for CHF analysis, 40 were classified as having CHF. Major criteria for CHF from the Framingham Study include stroke-related nocturnal dyspnea or orthopnea, jugular venous distention, rales, progressive increase in heart size on x-ray, acute pulmonary edema on chest x-ray, ventricular S(3) gallop rhythm, increased venous pressure >16 cm H2O, hepatojugular reflux, pulmonary edema, visceral congestion, cardiomegaly documented on necropsy, or weight loss of 10 lb / 5 days in response to Rx for CHF. Minor criteria include bilateral ankle edema, nocturnal cough, dyspnea on normal exertion, hepatomegaly, pleural effusion on x-ray, one-third decrease from maximum recorded vital capacity, tachycardia (120 beats per minute or greater), or pulmonary vascular congestion on chest x-ray. To be classified as having CHF, an individual must have at least two major criteria or one major criterion and two minor criteria present simultaneously. The demographics of these 1,540 individuals are summarized in Table 17.

[0290] Table 17. Demographics of the 1,540 individuals in the CHF dataset TIFF0007672192000025.tif35128

[0291] Of the 1,562 subjects available for stroke analysis, 38 were classified as having stroke, including hemorrhagic stroke (subarachnoid or intracerebral hemorrhage), ischemic stroke (cerebral embolism or atherothrombotic cerebral infarction), transient ischemic stroke, or stroke death. The demographics of these 1,562 subjects are summarized in Table 18.

[0292] Table 18. Demographics of the 1,562 individuals in the stroke dataset TIFF0007672192000026.tif35128

[0293] Variable reduction. The total number of genetic (SNP) and epigenetic (DNA methylation) probes remaining after quality control measures was 403,192 and 472,822, respectively. Due to the large number of variables (876,014 in total, excluding possible interactions between SNPs and DNA methylation sites) and to avoid collinearity, variable reduction was performed.

[0294] Linkage disequilibrium-based SNP pruning was performed using PLINK (Purcell, Neale et al. 2007) with a window size of 50 SNPs, a window shift of 5 SNPs, and a pairwise SNP-SNP LD threshold of 0.5. This reduced the number of SNPs from 403,192 to 161,474. To further reduce the number of SNPs, chi-squared p-values ​​were calculated between the remaining 161,474 SNPs and CHF and stroke status. Those with chi-squared p-values ​​<0.1 were retained for classification analysis, resulting in 15,132 SNPs for CHF and 14,819 SNPs for stroke.

[0295] To reduce the number of DNA methylation loci, we first calculated the point biserial correlation between 472,822 CpG sites and CHF and stroke status. CpG sites were retained if the point biserial correlation was at least 0.1. A total of 19,112 and 22,837 CpG sites remained for CHF and stroke, respectively. Subsequently, the Pearson correlation between sites was calculated independently for each disease. If the Pearson correlation between two loci was at least 0.8, loci with smaller point biserial correlation were discarded. Finally, 10,707 and 9,406 DNA methylation loci remained for CHF and stroke, respectively, for classification analysis.

[0296] Receiver Operating Characteristic Curves. Receiver Operating Characteristic (ROC) curves provide a graphical display of binary classification performance with various discrimination thresholds. Thus, to assess the ability of DNA methylation and SNPs in classifying CHF and Stroke, an R script was written to perform logistic regression of the model shown below, and then calculate the area under the curve (AUC) of the ROC curve using the pROC package in R (Beck and Shultz 1986). This was done systematically using DNA methylation sites ordered in descending order of point biserialization for disease and SNPs ordered in ascending order of chi-squared p-value for disease. In the model listed below, the SNP * meth term represents gene-environment interaction. CHF ~ SNP j + meth i + SNP j *meth i Stroke ~ SNP j + meth i + SNP j *meth i

[0297] result ROC for CHF classification. A model incorporating only main effects was fitted for CHF using the top three DNA methylation sites (cg09099697, cg19679281, cg25840850) and SNPs (rs10833199, rs11728055, rs16901105). The AUC of the ROC was 0.78 and is shown in Figure 10. Model parameters are summarized in Table 19.

[0298] Table 19. Main effects CHF model parameters TIFF0007672192000027.tif36157

[0299] To further demonstrate the importance of incorporating both DNA methylation and SNPs in better predicting CHF, the interaction terms shown in the Methods section were included in the CHF model. The AUC of the ROC for this model increased from the previous model to 0.81 and is shown in Figure 11. The model parameters are summarized in Table 20.

[0300] Table 20. Parameters of the interaction effect CHF model TIFF0007672192000028.tif82157

[0301] These two models for CHF clearly demonstrate the importance of accounting for both genetic and epigenetic effects. As shown in Table 19, even though only three variables (two CpGs and one SNP) are marginally significant at the 0.05 level for CHF, incorporating gene-environment interactions in the form of SNP-meth interactions strengthens the predictions. This is shown in Table 20, where two interaction terms are significant at the 0.05 level, and one other interaction is marginally significant.

[0302] ROC for stroke classification. A main effects model was fitted for stroke using the top five DNA methylation sites (cg27209395, cg27551078, cg03130180, cg10319399, cg25861340) and the top four SNPs (rs11007270, rs17073262, rs7190657, rs2411130). The AUC of the ROC was 0.85 and is shown in Figure 12. The model parameters are summarized in Table 21.

[0303] Table 21. Main effects stroke model parameters TIFF0007672192000029.tif51157

[0304] To again demonstrate the importance of DNA methylation sites and SNPs simultaneously, a model of interaction effects was fitted. The ROC AUC for this model was 0.86 and is shown in Figure 13. The model parameters are summarized in Table 22.

[0305] Table 22. Parameters of the stroke model with interaction effects TIFF0007672192000030.tif152157

[0306] Again, these two stroke models demonstrate the importance of genetics and environment in stroke. Both DNA methylation sites and SNPs are highly significant for classifying stroke. Furthermore, the classification performance may increase in additional studies with diverse ethnic backgrounds and larger sample sizes.

[0307] Consideration The results demonstrate that the presence of stroke or CHF can be inferred through the use of algorithms that utilize a combination of SNPs, methylation values, and / or their interaction terms. However, before the results can be considered, it is important to mention some limitations to this study. First, the Framingham cohort is exclusively Caucasian, with the majority of subjects in their mid-to-late 60s and 70s. Thus, the results may not be applicable to other ethnicities or different age ranges. Second, besides cg05575921, the validity of the M (or B values) for other probes has not been confirmed by proprietary techniques such as pyrosequencing. Third, the Illumina arrays used in the study are no longer available. The ability to replicate and expand may be affected due to changes in probe design or availability in newer generation arrays.

[0308] The present results highlight the value of resources such as the Framingham Heart Study in advancing our understanding of heart disease. Indeed, it is plausible that without this resource, this type of work would be difficult, if not impossible, to carry out. Moreover, given the present results with this unique dataset, a great deal of additional work is required before screening tests such as those described in this brief can be adopted clinically. Most obviously, the present results must be replicated and refined in other datasets and then retested in study populations representative of their intended future clinical applications. The latter point is particularly important, since even well-designed cohort studies that are epidemiologically sound in nature suffer from retention bias, which enriches for less severe illness in the remaining pool. This is especially true for illnesses related to substance use, since probands with high levels of substance use are more likely to be lost to long-term follow-up (Wolke, Waylen et al. 2009). Additionally, the frequency of SNPs may vary between ethnicities, so the effect size of a given interaction may also vary. Therefore, broader testing and development in ethnically-informative diverse cohorts is necessary.

[0309] The improvement of AUC may hit a ceiling. Ironically, this has little to do with the quality or quantity of epigenetic and genetic data. Instead, the constraint may be uncertainty in clinical characterization. Sadly, even under the best conditions, clinically relevant forms of CVD may remain undetected. This is true even for the FHS cohort. As a result, the "criteria" in this study itself are somewhat imprecise with respect to the actual clinical situation. This imprecision increases the error even for biomarkers that are precisely targeted to the relevant biological properties, so our ability to improve AUC may depend on our being able to obtain more accurate clinical evaluations (Philibert, Gunter et al. 2014).

[0310] Another limitation of using this approach is the constantly evolving epidemiology of CVD. While genetic contribution to CVD is relatively constant, diet and other environmental exposures continue to vary across generations. Perhaps the best illustration of this limitation can be by considering the contribution of smoking to the predictive power of this test in previous generations. Because tobacco was introduced from the New World to Europe in the early 1500s, we can confidently state that the contribution of smoking to CVD in medieval Europe was limited, and therefore the impact of cg05575921 on predictive power was zero. In contrast, since more than 40% of US adults smoked in the 1960s (Garrett, Dube et al. 2011), the contribution of smoking behavior to the prediction of CVD as captured by cg05575921 may have been significantly greater for subjects in that era. However, smoking is not the only environmental factor that varies between generations and cohorts. Over the past 20 years, there has been a significant change in our understanding and public attitudes toward the amount of saturated and trans fatty acids in a healthy diet. Because these environmental factors also have a strong influence on the likelihood of CVD, we expect that the weighting of the interaction effects loading on these dietary factors may vary depending on age and ethnicity.

[0311] The improved predictive power of the smoking methylation biomarker cg05575921 compared to self-reported smoking is not unexpected. Our initial studies have shown that it is a strong indicator of current smoking status with an AUC of 0.99 in studies using well-screened cases and controls (Philibert, Hollenbeck et al. 2015). It is a well-established phenomenon that self-reports of smoking, especially in high-risk cohorts, are unreliable (Caraballo, Giovino et al. 2001, Webb, Boyd et al. 2003, Caraballo, Giovino et al. 2004, Shipton, Tappin et al. 2009). Moreover, unlike cg05575921, categorical self-reports do not capture smoking intensity (Philibert, Hollenbeck et al. 2015). Finally, a large number of subjects who could have participated in the study may have previously smoked but still had residual demethylation of the AHRR despite not smoking at the Wave 8 interview. In each of these cases, the use of a continuous metric may capture additional vulnerability to CVD not captured by the binary smoking variable.

[0312] Since alcoholism is also a risk for CVD (Mozaffarian, Benjamin et al. 2016), we were somewhat surprised that our previously established and validated biomarker approach to assess alcohol intake did not have a larger predictive impact (Philibert, Penaluna et al. 2014, Brueckmann, Di Santo et al. 2016). In our initial model, the addition of the methylation status of cg2313759 only improved the AUC by 0.015. One reason for this failure to show the effect of alcohol use on CVD risk could be that this marker is not as well validated as our smoking biomarker, although there are other reasons as well. First and foremost, in contrast to the methylation of cg05575921, which shows a persistently increased risk of reduced life expectancy at all exposure levels, the methylation of cg2313759 shows an inverted U-shaped distribution with respect to biological aging. Whether CVD risk also follows a U-shaped distribution with respect to alcohol intake remains unknown, but it does suggest that any successful algorithm incorporating the main effects of alcohol-related methylation cannot use a simple linear approach.

[0313] Our success in finding an algorithm to predict CVD in the absence of genome-wide significant main effects may have important implications for the search for marker sets for other common complex disorders in adulthood. Among the top 10 leading causes of death in the United States, reliable methylation signatures have only been developed for type 2 diabetes and chronic obstructive pulmonary disease (COPD) using main effects (Qiu, Baccarelli et al. 2012, Toperoff, Aran et al. 2012). Since the ability to find good biomarkers for disease is highly dependent on the reliability of clinical diagnosis, the success in these two instances may be secondary to the excellent diagnostic reliability of the methods used to diagnose these two disorders, namely hemoglobin A1C and spirometry. Additionally, it is important to mention that while the diagnostic signature for T2DM maps primarily to pathways affected by excess glucose levels, the signature associated with COPD largely overlaps with that of smoking, which contributes to 95% of all cases of COPD (Qiu, Baccarelli et al. 2012, Toperoff, Aran et al. 2012). Nevertheless, since many of the risk factors for other major causes of death, such as stroke, overlap with risk factors for CVD (e.g., smoking), we are optimistic that this approach can be used to generate similar profiles.

[0314] Unfortunately, the majority of common adult-onset complex disorders do not have good existing biomarkers or large effect size etiological factors. In these cases, an approach that incorporates interaction effects may be beneficial, but the real question is why. Although speculative, our experience with local and genome-wide data shows that chronic exposure to cellular stressors leads to epigenome reorganization that may only be partially reversible. Regardless of how long it lasts, that disorganization of the genome is inevitably associated with disease, so it can be used as a disease biomarker. Understanding the reversion time of each of these effects may provide additional insight. For example, pharmacological interventions may have effects on distinct subsets. By understanding the relationship between reversion at these loci and treatment outcomes, it may be possible to optimize existing medications or better tailor new combination regimens.

[0315] In summary, we report that an algorithm incorporating information from interaction effects can predict the presence of stroke and CHF in FCS. We suggest that further studies are indicated to replicate and extend the generalizability of the approach to other ethnic cohorts. We further suggest that a similar approach could result in the generation of methylation profiles for other common complex disorders such as stroke.

[0316] References for Example 5 TIFF0007672192000031.tif245160TIFF0007672192000032.tif242153TIFF0007672192000033.tif238160TIFF0007672192000034.tif121159

[0317] While the foregoing specification and examples fully disclose and enable the invention, they are not intended to limit the scope of the invention, which is defined by the claims appended hereto.

[0318] All publications, patents, and patent applications are incorporated herein by reference. In the foregoing specification, the invention has been described with respect to certain specific embodiments thereof, and although numerous details have been set forth for purposes of illustration, it will be apparent to those skilled in the art that the invention is susceptible to additional embodiments, and that some of the details described herein may be varied in no small way without departing from the underlying principles of the invention.

[0319] The use of the terms "a" and "an" and "the" and similar referents in the context of describing the present invention should be interpreted to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. The terms "comprising," "having," "including," and "containing" should be interpreted as open-ended terms (i.e., meaning "including without limitation"), unless otherwise noted. The recitation of ranges of values ​​herein is merely intended to serve as a shorthand method of individually referring to each separate value falling within the range, unless otherwise indicated herein, and each separate value is incorporated herein as if it were individually recited herein. All methods described herein can be performed in any suitable order, unless otherwise indicated herein or clearly contradicted by context. The use of any and all examples or exemplary language provided herein (e.g., "such as") is intended merely to better illustrate the present invention, and does not impose limitations on the scope of the present invention, unless otherwise claimed. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the invention.

[0320] Aspects of the invention are described herein, including the best mode known to the inventors for carrying out the invention. Variations of those aspects may become apparent to those skilled in the art upon reading the foregoing description. The inventors anticipate that artisans will employ such variations as necessary, and the inventors intend for the invention to be practiced otherwise than as specifically described herein. Accordingly, this invention includes all modifications and equivalents of the subject matter recited in the claims appended hereto as permitted by applicable law. Moreover, any combination of the above-described elements in all possible variations thereof is encompassed by the invention unless otherwise indicated herein or clearly contradicted by context.

Claims

1. 1. A kit for determining the presence of a biomarker associated with coronary heart disease (CHD), comprising: the biomarkers include the methylation status of at least one CpG dinucleotide and the genotype of at least one single nucleotide polymorphism (SNP); complementary to a bisulfite converted nucleic acid sequence containing the CpG dinucleotide cg26910465, cg11355601, cg16410464 or cg12091641; (a) is at least 15 nucleotides in length; or (b) is at least 10-14 nucleotides in length and contains one or more synthetic nucleotide bases; at least one first nucleic acid primer that detects an unmethylated CpG dinucleotide; and is complementary to the DNA sequence of SNP rs6418712 or rs10275666; (a) is at least 15 nucleotides in length; or (b) is at least 10-14 nucleotides in length and contains one or more synthetic nucleotide bases; At least one second nucleic acid primer The kit comprising:

2. 1. A method for determining the presence of a biomarker associated with CHD in a patient sample, comprising: (a) isolating a nucleic acid sample from the patient sample; (b) performing a genotyping assay on a first aliquot of the nucleic acid sample to detect the presence of at least one SNP to obtain genotype data, wherein the at least one SNP is rs6418712 or rs10275666; and (c) bisulfite converting the nucleic acid in a second aliquot of the nucleic acid and performing a methylation assessment on the second aliquot of the nucleic acid sample to detect the methylation status of CpG sites cg26910465, cg11355601, cg16410464 or cg12091641 to obtain methylation data regarding whether a particular CpG residue is unmethylated; (d) inputting the genotype data from step (b) and the methylation data from step (c) into at least one algorithm that accounts for the contribution of at least one SNP main effect and at least one CpG main effect and gene-environment interaction (SNP x CpG) effects to obtain results; and (e) determining the presence of a biomarker associated with CHD in the patient sample based on the results from step (d). The method comprising:

3. 3. The method of claim 2, wherein the results include gene-environment interaction effects (SNP x CpG) between CpG sites cg26910465, cg11355601, cg16410464 or cg12091641 and SNP rs6418712 or rs10275666.

Citation Information

Patent Citations

  • Genetic polymorphisms associated with coronary heart disease, their detection methods and uses

    JP2009523405A