Methods and compositions for predicting and / or monitoring cardiovascular disease and for intervention thereof.
By integrating CpG methylation and SNP genotyping with machine learning, the method provides sensitive and specific cardiovascular disease prediction and management, enabling early detection and personalized interventions through non-invasive sampling, addressing the limitations of current detection methods.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- CARDIO DIAGNOSTICS INC
- Filing Date
- 2024-03-29
- Publication Date
- 2026-05-11
AI Technical Summary
Current risk estimators and detection methods for cardiovascular disease lack sensitivity, specificity, and accessibility, often requiring invasive procedures with significant side effects and being costly, making them unsuitable for widespread use.
A method and composition that utilize the methylation status of specific CpG loci and genotype of SNPs, combined with machine learning algorithms, to predict and manage cardiovascular disease through non-invasive means, such as blood or saliva samples, providing personalized intervention strategies.
Enables early detection, personalized intervention, and comprehensive assessment of cardiovascular disease with improved sensitivity and specificity, reducing costs and invasiveness, while offering remote accessibility and survival estimates.
Smart Images

Figure 2026514405000001_ABST
Abstract
Description
[Technical Field]
[0001] This disclosure relates, in general terms, to methods and compositions related to predicting cardiovascular disease (CVD) in individuals. [Background technology]
[0002] background Cardiovascular disease (CVD), particularly coronary heart disease (CHD), is the most common type of heart disease and was the cause of more than 360,000 deaths in the United States in 2017. To reduce these deaths, numerous risk estimators and detection methods have been developed to better identify people who have or are at risk of having CVD, including CHD. These tools, including the Framingham Risk Score (FRS) and, more recently, the Pooled Cohort Equation (PCE) of ASCVD, capture fluctuations in key physiological parameters known to be associated with the risk of CVD, including CHD, such as serum lipid levels. Similarly, methods such as stress echocardiography and computed tomography angiography (CCTA) are used for detection.
[0003] Despite these considerable efforts, current risk estimators and detection tests often lack sensitivity and specificity, and are frequently inaccessible due to cost and the need to schedule in-person consultations that can take several weeks. Furthermore, some current risk estimators and detection methods, such as catheterization, can have serious side effects, including stroke and heart attack. Consequently, there is a need for alternative stratification, detection, and management approaches for CVD that offer minimal risk, scalability, and a viable outlook. [Overview of the project]
[0004] overview Methods and compositions are provided for predicting the presence and / or severity (e.g., level of occlusion) of cardiovascular disease (CVD), and for managing, monitoring, and / or treating CVD. For example, methods and compositions for predicting coronary heart disease (CHD) are described herein. General principles apply to incidence windows (e.g., 1 month, 6 months, 2 years, or 10 years), as well as to the incidence, prevalence, or severity of other types of CVD, non-limitingly including CHD, stroke, arrhythmia, cardiac arrest, and congestive heart failure. The same general principles also apply to survival from CVD or CVD events, as well as to the management of CVD or CVD events, including, but not limited to, identifying, customizing, and optimizing lifestyle (e.g., exercise, diet) and / or therapeutic interventions (e.g., specific drugs or combinations thereof) and / or medical interventions (e.g., stent placement, angioplasty). The same general principles also apply to monitoring CVD or CVD events, the severity of CVD or CVD events, and / or responses to lifestyle, therapeutic interventions, and / or medical interventions. Specifically, methods and compositions comprising determining the methylation status of at least one CpG locus and / or at least one single nucleotide polymorphism (SNP) are described.
[0005] A kit is provided for determining the methylation status of at least one CpG dinucleotide and / or the genotype of at least one single nucleotide polymorphism (SNP) in one context. Such a kit typically involves determining the first CpG dinucleotide of a GC locus selected from the group consisting of cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584, or from cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584 A bisulfite-converted nucleic acid sequence comprising a bisulfite-converted nucleic acid sequence containing a 2 CpG dinucleotide in linkage disequilibrium with a 1 CpG dinucleotide of a GC locus selected from the group, wherein the linkage disequilibrium has a value of R > 0.3, and the at least one 1 nucleic acid primer detects methylated or unmethylated CpG dinucleotides. A first nucleic acid primer, and / or a first SNP selected from the group consisting of rs2869675, rs4376434, rs12129789, rs7585056, rs710987, rs4639796, rs1333048, rs12714414, rs942317, and rs1441433, or rs2869675, rs4376434, rs12129789, rs7585 The present invention includes at least one second nucleic acid primer, at least 8 nucleotides long, that is complementary to the DNA sequence of a second SNP that is in linkage disequilibrium with a first SNP selected from the group consisting of 056, rs710987, rs4639796, rs1333048, rs12714414, rs942317, and rs1441433, wherein the linkage disequilibrium has a value of R > 0.3.
[0006] In some embodiments, at least one first nucleic acid primer detects unmethylated CpG dinucleotides. In some embodiments, at least one first nucleic acid primer detects methylated CpG dinucleotides.
[0007] In some embodiments, the kit described herein further comprises at least one third nucleic acid primer having a length of at least 8 nucleotides and being complementary to the nucleic acid sequence upstream of the CpG dinucleotide. In some embodiments, the kit further comprises at least one third nucleic acid primer having a length of at least 8 nucleotides and being complementary to the nucleic acid sequence downstream of the CpG dinucleotide.
[0008] In some embodiments, at least one first nucleic acid primer comprises one or more nucleotide analogs. In some embodiments, at least one first nucleic acid primer comprises one or more synthetic or unnatural nucleotides.
[0009] In some embodiments, the kits described herein further comprise a solid substrate conjugated to at least one first nucleic acid primer. In some embodiments, the substrate is a polymer, glass, semiconductor, paper, metal, gel, or hydrogel. In some embodiments, the solid substrate is a microarray or microfluidic card.
[0010] In some embodiments, the kits described herein further include detectable labels.
[0011] In another context, a method is provided for determining the presence of biomarkers related to predicting, treating, managing, and / or monitoring CVD in patient-derived biological samples. Such a method typically comprises the steps of (a) providing a first portion of a biological sample and a second portion of a biological sample, wherein at least the nucleic acid derived from the first portion is bisulfite-converted; and (b) providing the first portion of the biological sample with a first CpG dinucleotide of a GC locus selected from the group consisting of cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584. A step of contacting a first oligonucleotide primer of at least 8 nucleotides in length with a sequence complementary to a second CpG dinucleotide in linkage disequilibrium with a first CpG dinucleotide of a GC locus selected from the group consisting of cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584, wherein the linkage disequilibrium has a value of R>0.3 and the first nucleic acid primer - A step of detecting methylated or unmethylated CpG dinucleotides; and (c) a second portion of a biological sample, a first SNP selected from the group consisting of rs2869675, rs4376434, rs12129789, rs7585056, rs710987, rs4639796, rs1333048, rs12714414, rs942317, and rs1441433, or rs2869675, rs4376434, The step of contacting a DNA sequence or bisulfite-converted DNA sequence of a second SNP that is in linkage disequilibrium with a first SNP selected from the group consisting of rs12129789, rs7585056, rs710987, rs4639796, rs1333048, rs12714414, rs942317, and rs1441433 with a nucleic acid primer of at least 8 nucleotides length that is complementary to the DNA sequence or bisulfite-converted DNA sequence, wherein the linkage disequilibrium has a value of R > 0.3.Generally, the percentage of methylation of CpG dinucleotides at GC loci selected from the group consisting of cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584, as well as the nucleotide identity of a first SNP or a second SNP in linkage disequilibrium with the first SNP, selected from the group consisting of rs2869675, rs4376434, rs12129789, rs7585056, rs710987, rs4639796, rs1333048, rs12714414, rs942317, and rs1441433, are biomarkers associated with detecting CVD or estimating survival from CVD.
[0012] In some embodiments, the biological sample is blood or saliva.
[0013] In some embodiments, at least one first nucleic acid primer detects unmethylated CpG dinucleotides. In some embodiments, at least one first nucleic acid primer detects methylated CpG dinucleotides.
[0014] In some embodiments, at least one first nucleic acid primer comprises one or more nucleotide analogs. In some embodiments, at least one first nucleic acid primer comprises one or more synthetic or unnatural nucleotides.
[0015] In some embodiments, the incidence window for detection, severity, management, and / or monitoring is 3, 5, or 10 years.
[0016] In a further context, a method for determining the presence of a biomarker in a biological sample derived from a subject, wherein the biomarker is related to detecting CVD, determining the severity of CVD, estimating survival from CVD, identifying, customizing, and / or optimizing interventions for CVD, managing CVD, and / or monitoring CVD.Such methods typically include (a) the step of isolating a nucleic acid sample from a patient sample; and (b) the step of performing a genotyping assay on a first portion of the nucleic acid sample to detect the presence of at least one SNP in order to obtain genotyping data, wherein the at least one SNP is from rs2869675, rs4376434, rs12129789, rs7585056, rs710987, rs4639796, rs1333048, rs12714414, rs942317, and rs1441433 or Annex C (c) A first SNP selected from and / or a second SNP in linkage disequilibrium (R>0.3) with a first SNP selected from rs2869675, rs4376434, rs12129789, rs7585056, rs710987, rs4639796, rs1333048, rs12714414, rs942317, and rs1441433 or from Annex C; and / or (c) a second SNP of nucleic acid to detect the methylation status of at least one CpG site in order to obtain methylation data A step of converting nucleic acids in a portion of a nucleic acid sample to bisulfite and performing a methylation evaluation on a second portion of the nucleic acid sample, wherein the at least one CpG site is selected from cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584 or from Annex A, and / or cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and c The steps include: (d) a CpG site that is collinear (R>0.3) with a CpG site selected from g17901584 or Annex A; and (d) a step of putting genotype data from step (b) and / or methylation data from step (c) into an algorithm that explains at least one SNP main effect and / or at least one CpG main effect and / or at least one interaction effect, wherein the algorithm is a machine learning algorithm capable of explaining linear and nonlinear effects.
[0017] In some embodiments, at least one interaction effect is selected from the group consisting of gene-environment interaction (SNP×CpG) effects, gene-gene interaction (SNP×SNP) effects, and environment-environment interaction (CpG×CpG) effects. In some embodiments, at least one interaction effect occurs between a CpG site selected from cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584 or from Annex A, or between a CpG site collinear (R>0.3) with a CpG site selected from cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584 or from Annex A, and rs2869675, rs4376434, rs12129789, rs75 This is a gene-environment interaction effect (SNP × CpG) between SNPs that are within moderate linkage disequilibrium (R>0.3) from 85056, rs710987, rs4639796, rs1333048, rs12714414, rs942317, and rs1441433 or selected from Annex C, or from rs2869675, rs4376434, rs12129789, rs7585056, rs710987, rs4639796, rs1333048, rs12714414, rs942317, and rs1441433 or selected from Annex C. In some embodiments, at least one interaction effect is an environment-environment interaction effect (CpG × CpG) between at least two CpG sites selected from cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584 or from Annex A.
[0018] In some embodiments, one or both of at least two CpG sites are collinear (R>0.3) with one or both of at least two CpG sites selected from cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584 or from Annex A. In some embodiments, at least one interaction effect is a gene-gene interaction effect (SNP × SNP) between at least two SNPs selected from rs2869675, rs4376434, rs12129789, rs7585056, rs710987, rs4639796, rs1333048, rs12714414, rs942317, and rs1441433 or from Annex C. In some embodiments, one or both of at least two SNPs are collinear (R>0.3) with one or both of at least two SNPs selected from rs2869675, rs4376434, rs12129789, rs7585056, rs710987, rs4639796, rs1333048, rs12714414, rs942317, and rs1441433 or from Annex C.
[0019] In some aspects, the biological sample is a saliva sample.
[0020] In another context, a system is provided for determining the methylation state of at least one CpG dinucleotide and the genotype of at least one single nucleotide polymorphism (SNP).Such a system typically includes a nucleic acid isolation module configured to isolate a nucleic acid sample from a target sample; and a genotyping assay module configured to perform a genotyping assay on a first portion of the nucleic acid sample to detect the presence of at least one SNP to obtain genotype data, wherein the at least one SNP is rs2869675, rs4376434, rs12129789, rs7585056, rs710987, rs4639796, rs1333048, rs127144 Genotyping assay module; obtain methylation data; first SNP selected from rs942317 and rs1441433 or from Annex C, and / or second SNP in linkage disequilibrium (R>0.3) with the first SNP selected from rs2869675, rs4376434, rs12129789, rs7585056, rs710987, rs4639796, rs1333048, rs12714414, rs942317 and rs1441433 or from Annex C; A methylation assay module configured to perform bisulfite conversion of a nucleic acid in a second portion of a nucleic acid and to perform methylation evaluation on the second portion of a nucleic acid sample in order to detect the methylation status of at least one CpG site, wherein the at least one CpG site is selected from cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584 or from Annex A, and / or cg0 A methylation assay module comprising a CpG site collinear (R>0.3) with a CpG selected from 4988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584 or from Annex A; and a recognition system configured to describe at least one SNP main effect and / or at least one CpG main effect and / or at least one interaction effect based on genotype data and / or methylation data.
[0021] In some embodiments, such a system further includes an output module configured to provide an output based on an identification by an identification system, the identification explaining at least one SNP main effect and / or at least one CpG main effect and / or at least one interaction effect based on genotype data and / or methylation data.
[0022] In some embodiments, the algorithm is a machine learning algorithm capable of explaining linear effects and / or non-linear effects.
[0023] In some embodiments, dimensionality reduction (e.g., principal component analysis, partial least squares regression, etc.) can be used.
[0024] In yet another aspect, a non-temporary computer-readable medium is provided for storing instructions that can be executed by a processing device for performing an action. Such an action typically involves describing at least one SNP main effect and / or at least one CpG main effect and / or at least one interaction effect based on genotype data and / or methylation data, wherein (i) the genotype data is based on a genotyping assay on a first portion of a nucleic acid sample isolated from a sample of interest to detect the presence of at least one SNP to obtain the genotype data, and the at least one The SNP is a first SNP selected from rs2869675, rs4376434, rs12129789, rs7585056, rs710987, rs4639796, rs1333048, rs12714414, rs942317, and rs1441433 or from Annex C, and / or rs2869675, rs4376434, rs12129789, rs7585056, rs710987, rs4639796, rs1333048, (ii) The methylation data is a second SNP in linkage disequilibrium (R>0.3) with a first SNP selected from rs12714414, rs942317, and rs1441433 or from Annex C; and (ii) the methylation data is based on a methylation assay for bisulfite-converted nucleic acid in a second portion of a nucleic acid sample to detect the methylation status of at least one CpG site to obtain methylation data, wherein the at least one CpG site is cg04988 The CpG sites are at least one selected from 978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584 or from Annex A, and / or CpG sites collinear (R>0.3) with a CpG selected from cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584 or from Annex A.
[0025] In some embodiments, the operation further includes providing an output based on the explanation. Representative outputs include, without limitation, saving an explanation-based report to another non-transitory computer-readable medium, modifying an explanation-based display, activating an explanation-based audible alert, activating an explanation-based tactile or vibration alert, activating printing of an explanation-based report, or activating delivery of an explanation-based treatment, including one or more of these.
[0026] In one aspect, a kit for determining the methylation status of at least one CpG dinucleotide is provided. Such a kit typically comprises at least one first nucleic acid primer of at least 8 nucleotides in length that is complementary to a bisulfite-converted nucleic acid sequence comprising a first CpG dinucleotide of a GC locus selected from the group consisting of cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584, or a second CpG dinucleotide in linkage disequilibrium with the first CpG dinucleotide of a GC locus selected from the group consisting of cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584, wherein the linkage disequilibrium has a value of R>0.3 and the at least one first nucleic acid primer detects a methylated CpG dinucleotide or an unmethylated CpG dinucleotide.
[0027] <000,0100>In another context, a method is provided for determining the presence of a biomarker in a biological sample derived from a subject, wherein the biomarker is related to detecting CVD, determining the severity of CVD, estimating survival from CVD, identifying, customizing, and / or optimizing interventions for CVD, managing CVD, and / or monitoring CVD. Such methods typically include (a) a step of providing a biological sample derived from a subject at risk of or having a CVD or CVD event, wherein at least a portion of the nucleic acid derived from the biological sample is bisulfite-converted; and (b) a step of converting the bisulfite-converted nucleic acid to a first CpG dinucleotide of a GC locus selected from the group consisting of cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584. The procedure includes contacting a first oligonucleotide primer of at least 8 nucleotides in length with a sequence complementary to a second CpG dinucleotide in linkage disequilibrium with a first CpG dinucleotide of a GC locus selected from the group consisting of cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584, wherein the percentage of methylation of the CpG dinucleotide of the GC locus selected from the group consisting of cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584 is relevant to estimating the survival of the subject.
[0028] In another context, a method is provided for determining the presence of a biomarker in a biological sample derived from a subject, wherein the biomarker is related to detecting CVD, determining the severity of CVD, estimating survival from CVD, identifying, customizing, and / or optimizing interventions for CVD, managing CVD, and / or monitoring CVD. Such a method typically comprises the steps of (a) isolating a nucleic acid sample from a subject sample; and (b) bisulfite-converting at least a portion of the nucleic acid to determine the methylation status of at least one CpG site to obtain methylation data, and performing a methylation assessment on the bisulfite-converted nucleic acid, wherein the at least one CpG site is selected from cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584 or from Annex A. A step of (c) a CpG site, and / or a CpG site collinear (R>0.3) with a CpG site selected from cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584 or from Annex A; and a step of (c) putting the methylation data from step (b) into an algorithm that describes at least one CpG main effect, wherein the algorithm is a machine learning algorithm capable of describing linear and nonlinear effects.
[0029] In another context, a system is provided for determining the methylation status of at least one CpG dinucleotide. Such a system typically includes a nucleic acid isolation module configured to isolate a nucleic acid sample from a sample of interest; and a methylation assay module configured to bisulfite-convert at least a portion of the nucleic acid and perform a methylation evaluation on the bisulfite-converted nucleic acid in order to determine the methylation status of at least one CpG site to obtain methylation data, wherein the at least one CpG site is cg04988978, cg21161138, cg12655112, cg03725309, cg125 A methylation assay module comprising: 86707, and at least one CpG site selected from cg17901584 or Annex A, and / or a CpG site collinear (R>0.3) with cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and a CpG site selected from cg17901584 or Annex A; and a discrimination system configured to describe at least one main CpG effect based on methylation data.
[0030] In yet another aspect, a non-temporary computer-readable medium is provided for storing instructions executable by a processing device for performing an action. Such a computer-readable medium typically includes describing at least one CpG main effect based on methylation data, the methylation data being based on a methylation assay on bisulfite-converted nucleic acids in at least a portion of a nucleic acid sample for detecting the methylation state of at least one CpG site to obtain the methylation data, the at least one CpG site being cg04988978, cg21161138 , cg12655112, cg03725309, cg12586707, and cg17901584 or at least one CpG site selected from Annex A, and / or a CpG site collinear (R>0.3) with cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584 or at least one CpG site selected from Annex A.
[0031] The integrated genetic-epigenetic models described herein offer several advantages and benefits, for example: • Earlier detection: PrecisionCHD can detect molecular changes that may or may not precede clinical symptoms or the onset of disease. This means that patients can be identified as having coronary heart disease before or after the onset of symptoms, enabling earlier intervention and better outcomes. • Actionable Clinical Intelligence (trademark): Trials may or may not be linked to the Actionable Clinical Intelligence platform for providers (see, for example, U.S. Provisional Patent Application No. 63 / 488,463, incorporated herein by reference), which maps each patient's molecular markers to key elements of coronary heart disease and enables tailored recommendations for lifestyle improvements, medical interventions, estimation of the effectiveness of lifestyle improvements or medical interventions, monitoring, and secondary trials. • Individualized intervention selection: Trials can be used to select one or more interventions, such as lifestyle improvements, therapeutic interventions, and medical interventions, for each patient or group of patients at one or more time points. • Optimization of individualized interventions: Trials can be used to optimize one or more interventions, such as lifestyle improvements, therapeutic interventions, and medical interventions, for each patient or group of patients at one or more time points. • Assessment of individualized interventions: Trials can be used to evaluate the effectiveness of interventions such as lifestyle improvements, therapeutic interventions, and medical interventions for each patient or group of patients by continuously monitoring CVD, CVD events, or the severity of CVD. • More comprehensive assessment: The test simultaneously assesses and integrates robust genetic and epigenetic biomarkers to provide a more comprehensive assessment of CVD status. • Discovery of novel pathways: The approach can be used to discover novel, previously unknown biological pathways for risk assessment, detection, intervention (e.g., lifestyle, therapeutic, medical), management, and monitoring of CVD. The pathways and biomarkers can also be used for the discovery, development, and validation of novel biopharmaceuticals for the assessment and management of cardiovascular disease. The approach described herein can be used to identify novel biomarkers (e.g., methylation, SNPs, proteins, etc.) for new drug development, or the capabilities of targeted treatments such as gene editing. The approach described herein can also be used to discover biomarkers for selecting the most effective drug (e.g., statins vs. beta-blockers), the most effective drug type (e.g., hydrophilic statins vs. lipophilic statins), lifestyle changes, medical interventions, or combinations thereof for a particular individual. In addition, the approach described herein can be used to optimize the use of therapeutic agents (e.g., medications, regimens, drug combinations) or to identify the lifestyle changes that have the most effect. • Non-invasive: The test requires only the simple collection of biomaterials (e.g., blood or saliva samples), making it a non-invasive and convenient alternative to more invasive diagnostic tests such as angiography. • Accessibility: Tests can be administered remotely via lancet-based sample collection kits that can be sent to the patient's home at the time of test instruction, thereby increasing and popularizing access to CVD diagnostic testing. Alternatively, biomaterials can be collected at the provider's location via vaccineiner-based sample collection. Biomaterials in the form of saliva samples can be collected remotely or at the provider's location. • Cost-effective and timely: PrecisionCHD provides clinicians with timely and cost-effective coronary heart disease trials. PrecisionCHD costs a fraction of the cost of other heart disease trials. • Survival estimates: PrecisionCHD or PrecisionCHD-Epi can provide survival estimates for individuals that have already been diagnosed with CVD or are considered to be at risk of developing CVD.
[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those generally understood by those skilled in the art in the field to which the subject methods and compositions belong. Similar or equivalent methods and materials may be used in carrying out or testing the subject methods and compositions, but preferred methods and materials are listed below. In addition, the materials, methods, and examples are illustrative only and are not intended to be limited to predicting CHD occurrence. All publications, patent applications, patents, and other references mentioned herein are incorporated in their entirety by reference. [Brief explanation of the drawing]
[0033] [Figure 1] This is a block diagram of an exemplary cardiovascular disease classification and monitoring system. [Figure 2] This is a flowchart illustrating an exemplary process for classifying and monitoring cardiovascular diseases. [Figure 3] This is a block diagram of an exemplary computing device. [Figure 4] This schematic diagram illustrates a potential therapeutic approach for patients with positive signals for CHD and / or related comorbidities, using the methods described herein. [Figure 5A] This plot shows the survival (yes or no) of marker cg04988978 during the follow-up period after methylation. [Figure 5B] This plot shows survival (yes or no) during the follow-up period for methylation of the marker cg21161138. [Figure 5C]This plot shows survival (yes or no) during the follow-up period for methylation of the marker cg12655112. [Figure 6A] This plot shows the relationship between the number of days to death and methylation at the marker cg04988978. [Figure 6B] This plot shows the relationship between the number of days to death and methylation at the marker cg21161138. [Figure 6C] This plot shows the relationship between the number of days to death and methylation at the marker cg12655112. [Modes for carrying out the invention]
[0034] Detailed explanation Recent predictive strategies leverage rapid advances in assessing genome-wide genetic or transcriptional variations. While each of these approaches has achieved some success, their clinical impact remains limited. In particular, those relying solely on genetic information have a clear upper limit to their predictive power, are potentially susceptible to ethnic stratification, and cannot be used to monitor changes in disease status because genotypes are static.
[0035] Recent advances in genome-wide epigenetic profiling techniques are increasing the potential of assessing DNA methylation in peripheral blood DNA as a mechanism for more accurate prediction of cardiovascular disease or mortality. However, predictive models that only explain epigenetic signatures cannot account for confounding genetic variations that affect the majority of environmentally responsive methylomes. This can result in models that lack robustness in terms of generalizability, particularly across different ethnic populations.
[0036] As a result, the inventors have developed a highly sensitive, clinically implementable, integrated genetic-epigenetic tool that can identify individuals who are at risk of or have cardiovascular disease (e.g., have a heart attack or sudden cardiac death). As shown herein, the methylation status of one or more specific CpG dinucleotides can be combined with the genotype of one or more specific loci (e.g., CH3×SNP) to predict cardiovascular disease (CVD), including coronary heart disease (CHD). The inventors have also developed a highly sensitive tool that can estimate the viability of individuals who are at risk of developing CVD, at risk of having a CVD event, or have already been identified as having CVD.
[0037] As described herein, biomarkers can be used in the diagnosis and prognosis of cardiovascular diseases and events. The terms “marker” and “biomarker” are interchangeable. As used herein, a biomarker generally refers to a measurable or detectable biological component (e.g., the presence or amount of a protein, genetic (e.g., polymorphism), epigenetic (e.g., methylation), and / or histological component). As described in more detail below, the biomarkers used herein are typically related to cardiovascular diseases.
[0038] As described herein, the terms “patient,” “subject,” and “individual” may be used interchangeably.
[0039] DNA methylation DNA does not exist as a naked molecule in cells. For example, DNA associates with proteins called histones to form a complex known as chromatin. Chemical modifications of DNA or histones alter the structure of chromatin without changing the nucleotide sequence of the DNA. Such modifications are described as "epigenetic" modifications of DNA. Changes in chromatin structure can have a significant impact on gene expression. If chromatin is condensed, factors involved in gene expression may not be able to access the DNA, and the gene will be switched off. Conversely, if chromatin is "open," the gene can be switched on. Some important forms of epigenetic modification are DNA methylation and histone deacetylation.
[0040] DNA methylation is a chemical modification of the DNA molecule itself, carried out by enzymes called DNA methyltransferases. Methylation can directly switch off gene expression by preventing transcription factors from binding to promoters. A more common effect is the attraction of methyl-binding domain (MBD) proteins. These associate with further enzymes called histone deacetylases (HDACs), which function to chemically modify histones and alter chromatin structure. Chromatin containing acetylated histones is open and accessible to transcription factors, and genes are potentially active. Histone deacetylation causes chromatin condensation, which makes transcription factors inaccessible and leads to gene silencing.
[0041] A CpG island is a short stretch of DNA where the frequency of CpG sequences is higher than in other regions. The "p" in the term CpG indicates that cysteine ("C") and guanine ("G") are linked by a phosphodiester bond. CpG islands are often located around the promoters of housekeeping genes and many regulated genes. In these locations, CG sequences in active genes are often unmethylated. In contrast, CG sequences in inactive genes are usually methylated to repress their expression.
[0042] As used herein, the term "methylation status" means the determination of whether a particular target DNA, such as a CpG dinucleotide, is methylated or unmethylated. As used herein, the term "CpG dinucleotide repeat motif" means a series of two or more CpG dinucleotides located within a DNA sequence.
[0043] Approximately 56% of human genes and 47% of mouse genes are associated with CpG islands. Often, CpG islands overlap with promoters and extend approximately 1000 base pairs downstream within the transcription unit. Identifying potential CpG islands during sequence analysis is notoriously difficult with cDNA-based approaches, as they help define the furthest 5' end of a gene. Methylation of CpG islands can be determined by those skilled in the art using any method suitable for determining such methylation. For example, those skilled in the art can use bisulfite reaction-based methods to determine such methylation.
[0044] This disclosure provides a method for determining nucleic acid methylation of one or more gene loci in a subject in order to identify a subject having CVD.
[0045] Linkage refers to the phenomenon where DNA sequences located close to each other in a genome tend to be inherited together. Two sequences may be linked due to several selective advantages of co-inheritance. However, more typically, two sequences co-inherit because meiotic recombination events occur relatively rarely within the region between the two sequences. Co-inherited sequences are said to be in "linkage disequilibrium" with each other because, in a given population, they tend to either exist together or not at all in any particular member of the population. In fact, when multiple sequences in a given chromosomal region are found to be in linkage disequilibrium with each other, they define a metastable "haplotype." In contrast, recombination events occurring between two gene loci cause them to separate onto distinct homologous chromosomes. If meiotic recombination occurs frequently enough between two physically linked sequences, the two sequences appear to be independently separated and are said to be in linkage equilibrium.
[0046] It is understood that linkage disequilibrium can be quantified (for example, using Pearson correlation (R) or allele co-inheritance (D')). For example, low levels of linkage can be reflected in correlations (e.g., R values) of about 0.1 or less, moderate levels of linkage in an R value of about 0.3, and high levels of linkage in an R value of 0.5 or more. When referring to methylation (i.e., CpG sites), it is also understood that collinearity (having an R value) can be used to determine the linear strength of the association between two CpGs (for example, low levels of collinearity can be reflected in an R value of about 0.1 or less; moderate levels of collinearity in an R value of about 0.3; and high levels of collinearity in an R value of about 0.5 or more).
[0047] In particular, in certain embodiments of this disclosure, the method may be carried out as follows: A sample, such as a blood sample, is collected from the subject. In certain embodiments, a single cell type, such as lymphocytes, basophils, or monocytes isolated from blood, may be isolated for further testing. DNA is collected from the sample and examined to determine the methylation of one or more loci. For example, the DNA of interest may be treated with a bisulfite to deaminate unmethylated cytosine residues to uracil. Since uracil base-pairs with adenosine, thymidine is incorporated into the subsequent DNA strand in place of unmethylated cytosine residues during partial sequence PCR amplification. The target sequence is then amplified by PCR and probed with a locus-specific probe. Depending on the specific sequence of the probe used, only methylated or unmethylated DNA will bind to the probe.
[0048] Methods for determining the target nucleic acid profile are well known to those skilled in the art and include any of the well known detection methods. Various PCR methods are described, for example, in PCR Primer: A Laboratory Manual, Dieffenbach 7 Dveksler, Eds., Cold Spring Harbor Laboratory Press, 1995. Other methods include, but are not limited to, nucleic acid quantification, restriction enzyme digestion, DNA sequencing, hybridization techniques such as Southern blotting, amplification methods such as ligase chain reaction (LCR), nucleic acid sequence-based amplification (NASBA), auto-sustained sequence replication (SSR or 3SR), strand substitution amplification (SDA), and transcription-mediated amplification (TMA), quantitative PCR (qPCR), digital PCR (dPCR) (e.g., digital droplet PCR (ddPCR)), or other DNA analyses, as well as RT-PCR, in vitro translation, Northern blotting, and other RNA analyses. In another embodiment, hybridization to microarrays is used.
[0049] Single nucleotide polymorphism (SNP) Traditional methods for screening hereditary diseases rely on identifying either an abnormal gene product (e.g., sickle cell anemia) or an abnormal phenotype (e.g., intellectual disability). With the development of simpler and less expensive gene screening methodologies, it is now possible to identify polymorphisms that indicate a predisposition to developing a disease, even when the disease has a polygenic origin.
[0050] Single nucleotide polymorphism (SNP) genotyping measures the genetic variation of SNPs among members of a species. A SNP is a single base pair variation at a specific locus, typically consisting of two alleles (where the rare allele frequency is >1%). SNPs are very common. Because SNPs are conserved during evolution, they have been proposed as markers for use in quantitative trait locus (QTL) analysis and related studies, as an alternative to microsatellites. Many different SNP genotyping methods are known, including hybridization-based methods (such as dynamic allele-specific hybridization, molecular beacons, and SNP microarrays), enzyme-based methods (including restriction fragment length polymorphism, PCR-based methods, flap endonucleases, primer extension, 5'-nucleases, and oligonucleotide ligation assays), other post-amplification methods based on the physical properties of DNA (such as single-strand higher-order structure polymorphism, temperature gradient gel electrophoresis, denaturing high-performance liquid chromatography, high-resolution melting of whole amplicons, use of DNA mismatch-binding proteins, SNPlex, and surveyor nuclease assays), and sequencing (such as "next-generation" sequencing). See, for example, U.S. Patent No. 7,972,779.
[0051] Multiple alleles at a gene locus can arise from one or more polymorphisms in the region of the polypeptide-encoding gene or in regulatory sequences that affect polypeptide expression, such as promoters or polyadenylation sequences. Alternatively, alleles can arise from one or more polymorphisms at gene loci distal to the polypeptide-encoding gene or in regulatory sequences. Polymorphisms can affect polypeptides at the transcriptional or translational level (e.g., the transcription rate, translation rate, degradation rate, and / or activity of the polypeptide). Allele differences can be characterized in samples derived from a single subject or multiple subjects using methods known to those skilled in the art. Such methods may include, but are not limited to, measuring the potential for expression of a polynucleotide sequence and / or measuring the amount of the encoded polypeptide. Methods are available that can directly or indirectly detect proteins or nucleic acids, and assay methods are specifically intended to include, for example, screening for the presence of specific sequences or structures of nucleic acids or polypeptides using one of various known microarray techniques.
[0052] It is well recognized by those skilled in the art that an allele does not need to have been previously shown to have any association or relationship with the disorder phenotype. Instead, alleles and pathogenic environmental risk factors may interact to predict predisposition to the disorder phenotype, even if neither the allele nor the risk factor has any direct relationship to the disorder phenotype.
[0053] Genetic screening (also called genotyping or molecular screening) can be broadly defined as a test to determine whether a subject has a mutation (or allele or polymorphism) that either causes a disease condition or is "linked" to a mutation that causes a disease condition. Linkage refers to the phenomenon where DNA sequences that are close to each other in the genome tend to be inherited together. Two sequences may be linked due to several selective advantages of co-inheritance. However, more typically, two polymorphic sequences are co-inherited because meiotic recombination events are relatively rare to occur within the region between the two polymorphisms. Co-inherited polymorphic alleles are said to be "linkage disequilibrium" with each other because, in a given population, they tend to either be present together or not present at all in any particular member of the population. In fact, when multiple polymorphisms in a given chromosomal region are found to be linkage disequilibrium with each other, they define a metastable gene "haplotype". In contrast, a recombination event occurring between two polymorphic loci causes them to be separated onto distinct homologous chromosomes. When meiotic recombination occurs frequently enough between two physically linked polymorphisms, the two polymorphisms appear to be independently separated and are said to be in linkage equilibrium.
[0054] It is understood that linkage disequilibrium can be quantified (for example, using Pearson correlation (R) or allele co-inheritance (D')). For example, low levels of linkage can be reflected in correlations (e.g., R values) of about 0.1 or less, moderate levels of linkage in an R value of about 0.3, and high levels of linkage in an R value of 0.5 or more.
[0055] The frequency of meiotic recombination between two markers is generally proportional to the physical distance between them on the chromosome; however, the appearance of "hot spots" and regions of suppressed chromosomal recombination can result in a discrepancy between the physical distance and the recombination distance between the two markers. Therefore, within a particular chromosomal region, multiple polymorphic loci spanning a wide range of chromosomal domains may be in linkage disequilibrium with one another, thereby defining a broad-spectrum gene haplotype. Furthermore, if a disease-causing mutation is found within or linked to this haplotype, one or more polymorphic alleles of the haplotype can be used as a diagnostic or prognostic indicator of the likelihood of developing the disease. This association between benign polymorphisms and disease-causing polymorphisms occurs when the disease-causing mutation is recent and not enough time has elapsed for equilibrium to be achieved through recombination events. Therefore, identifying a haplotype that extends to or is linked to a disease-causing mutation serves as a means of predicting the likelihood that an individual is carrying that disease-causing mutation. Such prognostic or diagnostic procedures can be utilized without requiring the identification and isolation of the actual disease-causing lesion. This is significant because accurately identifying the molecular defects involved in the disease process can be difficult and cumbersome, especially in the case of multifactorial diseases.
[0056] Statistical correlations between disorders and polymorphisms do not necessarily indicate that the polymorphism directly causes the disorder. Rather, correlated polymorphisms may be benign allele variants that occurred in a more recent evolutionary past and are linked to (i.e., in linkage disequilibrium with) the disorder-causing mutation, meaning that sufficient time has not elapsed for equilibrium to be achieved in the intervening chromosomal segment through a recombination event. Therefore, for the purpose of diagnostic and prognostic assays for specific diseases, the detection of polymorphic alleles associated with that disease can be utilized without considering whether the polymorphism is directly involved in the pathogenesis of the disease. Furthermore, if a given benign polymorphic locus is in linkage disequilibrium with a clearly disease-causing polymorphic locus, then other polymorphic loci that are in linkage disequilibrium with the benign polymorphic locus are also likely to be in linkage disequilibrium with the disease-causing polymorphic locus. Therefore, these other polymorphic loci can also be used to prognose or diagnose the likelihood that the disease-causing polymorphic locus is inherited. A wide range of haplotypes (which describe typical patterns of co-inheritance of a set of linked polymorphic marker alleles) can be targeted for diagnostic purposes once an association is established between a particular disease or condition and the corresponding haplotype. Therefore, the likelihood of an individual developing a disease of a particular condition can be determined by characterizing one or more disease-associated polymorphic alleles (or one or more disease-associated haplotypes), without necessarily determining or characterizing the causative genetic variation.
[0057] Many methods are available to detect specific alleles of polymorphic loci. Certain methods for detecting specific polymorphic alleles depend, in part, on the molecular properties of the polymorphism. For example, the various allele forms of a polymorphic locus can differ by a single base pair of DNA. Such single nucleotide polymorphisms (i.e., SNPs) are the major contributors to genetic variation, accounting for approximately 80% of all known polymorphisms, and their density in the genome is estimated to be one on average per 1,000 base pairs. SNPs most frequently exist as biallelic or in only two distinct forms (however, theoretically, up to four distinct forms of SNPs are possible, corresponding to the four different nucleotide bases present in DNA). Nevertheless, SNPs are more mutationally stable than other polymorphisms, making them suitable for association studies where linkage disequilibrium between a marker and an unknown variant is used to map disease-causing mutations. In addition, because SNPs typically have only two alleles, they can be genotyped by a simple plus / minus assay rather than by length measurement, making them more easily followed by automation.
[0058] In one embodiment, allele profiling can be performed using nucleic acid microarrays. The field of genetic testing is evolving rapidly, and therefore, those skilled in the art will recognize that a wide range of profiling tests exist and are being developed to determine the allele profile of an individual in accordance with this disclosure.
[0059] Nucleic acids and polypeptides As described herein, the methods provided in this disclosure depend on the characteristics contained in the nucleic acids of an individual, subject, or patient. The term “nucleic acid” refers to deoxyribonucleotides or ribonucleotides and polymers thereof, either in single-stranded or double-stranded form, made up of monomers (nucleotides) containing a sugar, phosphate, and a base that is either a purine or a pyrimidine. Unless specifically limited, the term includes nucleic acids containing known analogs of native nucleotides that have binding properties similar to a reference nucleic acid and are metabolized in a manner similar to naturally occurring nucleotides. Unless otherwise indicated, a particular nucleic acid sequence also includes its conservatively modified variants (e.g., degenerate codon substitutions) and complementary sequences, as well as sequences that are clearly indicated. Specifically, degenerate codon substitutions can be achieved by producing sequences in which the third position of one or more selected (or all) codons is substituted with a mixed base and / or a deoxyinosine residue. The terms “nucleic acid,” “nucleic acid molecule,” or “polynucleotide” are interchangeable and may also be interchangeable with genes, cDNA, DNA, and / or RNA encoded by genes.
[0060] The term "nucleotide sequence" refers to a polymer of DNA or RNA that may be single-stranded or double-stranded, optionally containing synthetic, non-natural, or modified nucleotide bases that can be incorporated into the DNA polymer or RNA polymer. A DNA molecule or polynucleotide is a polymer of deoxyribonucleotides (A, G, C, and T), and an RNA molecule or polynucleotide is a polymer of ribonucleotides (A, G, C, and U).
[0061] For the purposes of this disclosure, “gene” includes DNA regions that encode a gene product, as well as all DNA regions that regulate the production of the gene product, whether such regulatory sequences are adjacent to the coding sequence and / or the sequence being transcribed. The term “gene” is used broadly to refer to any segment of nucleic acid associated with a biological function. A gene includes coding sequences and / or regulatory sequences required for their expression. Thus, a gene includes, but is not limited to, promoter sequences, terminators, translation regulatory sequences, such as ribosome binding sites and internal ribosome entry sites, enhancers, silencers, insulators, boundary elements, origins of replication, matrix attachment sites, and locus regulatory regions. For example, “gene” means mRNA, functional RNA, or nucleic acid fragment that expresses a particular protein, including regulatory sequences. “Functional RNA” means sense RNA, antisense RNA, ribozyme RNA, siRNA, or other RNA that cannot be translated but still has an effect on at least one cellular process. “Genes” also include, for example, unexpressed DNA segments that form recognition sequences for other proteins. A “gene” may include sequences that can be obtained from various sources, including cloning from a source of interest, or synthesizing from known or predicted sequence information, and that are designed to have desired parameters.
[0062] "Gene expression" refers to the conversion of information contained in a gene into a gene product. This includes the transcription and / or translation of endogenous genes, heterologous genes or nucleic acid segments, or transgenes in a cell. In addition, expression refers to the transcription and stable accumulation of sense (mRNA) or functional RNA. Expression can also refer to the production of proteins. The term "altered level of expression" refers to the level of expression in a transgenic cell or organism that differs from that of a normal or non-transformed cell or organism.
[0063] A gene product can be a transcript of a gene (e.g., mRNA, tRNA, rRNA, antisense RNA, ribozyme, structural RNA, or any other type of RNA), or a protein produced by the translation of mRNA. Gene products also include RNA modified by processes such as cap formation, polyadenylation, methylation, and editing, as well as proteins modified by processes such as methylation, acetylation, phosphorylation, ubiquitination, ADP-ribose, myristylation, and glycosylation. The term “RNA transcript” refers to the product resulting from the transcription of a DNA sequence catalyzed by RNA polymerase. If an RNA transcript is a complementary copy of a DNA sequence, it is called a primary transcript; the RNA sequence derived from the post-transcriptional processing of a primary transcript is called mature RNA. “Messenger RNA” (mRNA) refers to RNA that lacks introns and can be translated into proteins by cells. “cDNA” refers to single-stranded or double-stranded DNA that is complementary to and derived from mRNA. "Functional RNA" refers to sense RNA, antisense RNA, ribozyme RNA, siRNA, or other RNA that cannot be translated but still has an effect on at least one cellular process.
[0064] A "coding sequence," or sequence that "encodes" a polypeptide, is a nucleic acid molecule that, when controlled by an appropriate regulatory sequence, is transcribed (in the case of DNA) and / or translated (in the case of mRNA) in vivo into a polypeptide. The boundaries of a coding sequence are determined by a start codon at the 5' (amino) terminus and a translation termination codon at the 3' (carboxy) terminus. Coding sequences may include, but are not limited to, cDNA derived from viral, prokaryotic, or eukaryotic mRNA, genomic DNA sequences derived from viral (e.g., DNA viruses and retroviruses) or prokaryotic DNA, and synthetic DNA sequences. A transcription termination sequence may be located at 3' of the coding sequence.
[0065] "Regulatory sequences" and "preferred regulatory sequences" refer to nucleotide sequences located upstream of a coding sequence (5' non-coding sequence), within a coding sequence, or downstream of a coding sequence (3' non-coding sequence), and which affect the transcription, RNA processing, stability, or translation of the associated coding sequence. Regulatory sequences include enhancers, promoters, translational leader sequences, introns, and polyadenylation signal sequences. They include sequences that may be native and synthetic sequences, as well as combinations of synthetic and native sequences.
[0066] Certain aspects of this disclosure encompass isolated or substantially purified nucleic acid compositions. In the context of this disclosure, an “isolated” or “purified” DNA or RNA molecule is a DNA or RNA molecule that exists away from its native environment and is therefore not a natural product. An isolated DNA or RNA molecule may exist in a purified form or in a non-native environment, such as a transgenic host cell. For example, if an “isolated” or “purified” nucleic acid molecule is produced by recombinant technology, it may substantially contain no other cell material or culture medium, or if it is chemically synthesized, it may substantially contain no chemical precursors or other chemicals. In one embodiment, an “isolated” nucleic acid does not contain sequences that are naturally adjacent to the nucleic acid in the genomic DNA of the organism from which the nucleic acid originates (i.e., sequences located at the 5' and 3' ends of the nucleic acid).
[0067] A "fragment" refers to a polypeptide consisting of only a portion of the intact full-length polypeptide sequence and structure. Fragments may include C-terminal deletions, N-terminal deletions, and / or internal deletions of the native polypeptide. Protein fragments will generally contain at least about 5 to 100 consecutive amino acid residues of the full-length molecule (e.g., at least about 15 to 25 consecutive amino acid residues of the full-length molecule, at least about 20 to 50 or more consecutive amino acid residues of the full-length molecule, or any integer between 5 amino acids and the full-length sequence).
[0068] The term "naturally occurring" is used to describe compositions that can be found in nature, as opposed to those that are artificially produced. For example, a nucleotide sequence present in a living organism that can be isolated from a natural source and has not been intentionally modified by a human in a laboratory is considered naturally occurring.
[0069] A "5' non-coding sequence" refers to a nucleotide sequence located 5' (upstream) of a coding sequence. 5' non-coding sequences are present in fully processed mRNA upstream of the start codon and can affect the processing of the primary transcript into mRNA, mRNA stability, or translation efficiency. A "3' non-coding sequence" refers to a nucleotide sequence located 3' (downstream) of a coding sequence and may include polyadenylation signal sequences and other sequences encoding regulatory signals that can affect mRNA processing or gene expression.
[0070] A “promoter” refers to a nucleotide sequence, usually upstream (5') of a coding sequence, that directs and / or controls the expression of a coding sequence by providing recognition to RNA polymerase and other factors required for proper transcription. A “promoter” can include a minimal promoter, which is a short DNA sequence consisting of a TATA box and other sequences that work to identify the site of transcription initiation, to which regulatory elements are added for control of expression. A “promoter” can also refer to a nucleotide sequence that includes a minimal promoter + one or more regulatory elements (e.g., enhancers) that can control the expression of a coding sequence or functional RNA. Promoters may be derived entirely from native sequences, or they may consist of various elements derived from various naturally occurring promoters, or they may further consist of synthetic DNA sequences. Promoters may also contain DNA sequences involved in the binding of protein factors that control the effectiveness of transcription initiation in response to physiological or developmental conditions. “Constitutive expression” refers to expression using a constitutive promoter. “Conditional” and “regulated expression” refer to expression controlled by a regulatory promoter.
[0071] An "enhancer" is a DNA sequence that can stimulate promoter activity. Enhancers can be innate elements of the promoter, or heterogeneous elements inserted to enhance the promoter's level or tissue specificity. Enhancers can often function in both orientations and can even function when moved either upstream or downstream of the promoter. Both enhancers and other regulatory elements within the promoter bind to sequence-specific DNA-binding proteins that mediate their effects.
[0072] "Functionally linked" refers to the association of nucleic acid sequences on a single nucleic acid fragment such that the function of one sequence is influenced by another sequence. For example, if two sequences are positioned such that a regulatory DNA sequence influences the expression of a coding DNA sequence (i.e., the coding sequence or functional RNA is under the transcriptional control of a promoter), then the regulatory DNA sequence is said to be "functionally linked to" or "associated with" the DNA sequence that codes for the RNA or polypeptide. The coding sequence can be functionally linked to the regulatory sequence in either the sense or antisense direction.
[0073] "Expression" refers to the transcription and / or translation of endogenous genes, heterologous genes or nucleic acid segments, or transgenes in a cell. In addition, expression refers to the transcription and stable accumulation of sense (mRNA) or functional RNA. Expression can also refer to protein production. The term "altered level of expression" refers to a level of expression in a cell or organism that is different from that of a normal cell or organism.
[0074] For sequence comparison, typically one sequence acts as the reference sequence compared to the test sequence. When using a sequence comparison algorithm, the test sequence and reference sequence are entered into a computer, and the sequence algorithm program parameters are specified. The sequence comparison algorithm then calculates the sequence identity percentage for the test sequence relative to the reference sequence based on the specified program parameters.
[0075] To describe the sequence relationship between two or more nucleic acids or polynucleotides, the following terms are used: (a) “reference sequence,” (b) “comparison window,” (c) “sequence identity,” (d) “percentage of sequence identity,” and (e) “for sequence comparisons, remain as they are; the reference sequence may be a subset or substantial identity.” As used herein, “reference sequence” means a specified sequence as a whole; for example, a segment of a full-length cDNA or gene sequence, or a defined sequence used as a complete cDNA or gene sequence. As used herein, “comparison window” refers to a contiguous and specified segment of a polynucleotide sequence, the polynucleotide sequence in the comparison window may contain additions or deletions (i.e., gaps) compared to the reference sequence (which does not contain additions or deletions) for optimal alignment of the two sequences. Generally, the comparison window is at least 20 consecutive nucleotides long and may optionally be 30, 40, 50, 100, or longer. Those skilled in the art will understand that, in order to avoid high similarity to a reference sequence due to the inclusion of gaps in a polynucleotide sequence, a gap penalty is typically introduced and subtracted from the number of matches.
[0076] Methods for aligning sequences for comparison are well known in the art. Therefore, the determination of the identity percentage between any two sequences can be performed using mathematical algorithms. Non-restrictive examples of such mathematical algorithms include the Myers and Miller algorithm (Myers and Miller, CABIOS, 4, 11 (1988)); Smith et al.'s local homology algorithm (Smith et al., Adv. Appl. Math., 2, 482 (1981)); Needleman and Wunsch's homology alignment algorithm (Needleman and Wunsch, JMB, 48, 443 (1970)); Pearson and Lipman's similarity search method (Pearson and Lipman, Proc. Natl. Acad. Sci. USA, 85, 2444 (1988)); Karlin and Altschul's algorithm (Karlin and Altschul, Proc. Natl. Acad. Sci. USA, 87, 2264 (1990)); and the modified version of Karlin and Altschul's algorithm (Karlin and Altschul, Proc. Natl. Acad. Sci. USA 90, 5873 (1993).
[0077] Computer implementations of these mathematical algorithms can be used to compare arrays and determine array identity. Such implementations include, but are not limited to, the following: CLUSTAL in the PC / Gene program (available from Intelligenetics, Mountain View, Calif.); the ALIGN program (version 2.0); and GAP, BESTFIT, BLAST, FASTA, and TFASTA in the Wisconsin Genetics Software Package, version 8 (available from Genetics Computer Group (GCG), 575 Science Drive, Madison, Wis., USA). Alignment using these programs can be performed using default parameters. The CLUSTAL program has been described in detail by Higgins et al. (Higgins et al., CABIOS, 5, 151 (1989)); Corpet et al. (Corpet et al., Nucl. Acids Res., 16, 10881 (1988)); Huang et al. (Huang et al., CABIOS, 8, 155 (1992)); and Pearson et al. (Pearson et al., Meth. Mol. Biol., 24, 307 (1994)). The ALIGN program is based on the Myers and Miller algorithms described above. The BLAST program by Altschul et al. (Altschul et al., J. Mol. Biol., 215, 403 (1990)) is based on the Karlin and Altschul algorithms described above.
[0078] Software for performing BLAST analysis is publicly available through the National Center for Biotechnology Information. This algorithm initially identifies high-scoring sequence pairs (HSPs) by identifying short words of length "W" in the query sequence that match or satisfy a certain positive threshold score T when aligned with words of the same length in the database sequence. "T" is referred to as the adjacent word score threshold. These initial adjacent word hits act as seeds to initiate a search for longer HSPs containing them. Word hits are then extended bidirectionally along each sequence as long as the cumulative alignment score increases. For nucleotide sequences, the cumulative score is calculated using parameters "M" (reward score for a pair of matching residues; always >0) and "N" (penalty score for mismatched residues; always <0), while for amino acid sequences, a scoring matrix is used to calculate the cumulative score. The extension of word hits in each direction is stopped when the cumulative alignment score deviates by an amount "X" from its maximum achieved value, when the cumulative score becomes zero or less due to the accumulation of one or more negative scoring residue alignments, or when the end of any sequence is reached.
[0079] In addition to calculating the sequence identity percentage, the BLAST algorithm also performs a statistical analysis of the similarity between two sequences. One measure of similarity provided by the BLAST algorithm is the minimum sum probability (P(N)), which indicates the probability that a match between two nucleotide or amino acid sequences occurs by chance. For example, if the minimum sum probability in the comparison between the test nucleic acid sequence and the reference nucleic acid sequence is less than approximately 0.1, less than approximately 0.01, or even less than approximately 0.001, the test nucleic acid sequence is considered similar to the reference sequence.
[0080] To obtain gapped alignments for comparative purposes, Gapped BLAST (in BLAST 2.0) can be used. Alternatively, PSI-BLAST (in BLAST 2.0) can be used to perform iterative searches to detect intermolecular distance relationships. When using BLAST, Gapped BLAST, or PSI-BLAST, the default parameters of each program (e.g., BLASTN for nucleotide sequences, BLASTX for proteins) can be used. The BLASTN program (for nucleotide sequences) uses a default word length (W) of 11, an expected value (E) of 10, a cutoff of 100, M=5, N=-4, and a comparison of both strands. For amino acid sequences, the BLASTP program uses a default word length (W) of 3, an expected value (E) of 10, and the BLOSUM62 scoring matrix. Alignment may also be performed manually by testing.
[0081] For the purposes of this disclosure, the comparison of nucleotide sequences for determining the sequence identity percentage to the promoter sequences disclosed herein may be performed using the BlastN program (version 1.4.7 or later) with its default parameters or any equivalent program. “Equivalent program” means any sequence comparison program that, for any two sequences in question, produces alignments that, when compared to corresponding alignments generated by the program, have identical nucleotide or amino acid residue matches and identical sequence identity percentages.
[0082] Where used herein, “sequence identity” or “identity” in the context of two nucleic acid or polypeptide sequences refers to the specified percentage of residues in the two sequences that are identical when aligned to maximize correspondence across a specified comparison window, as measured by a sequence comparison algorithm or visual inspection. When the percentage of sequence identity is used in relation to proteins, it is recognized that the non-identical residue positions are often differed by conserved amino acid substitutions, in which an amino acid residue is replaced by another amino acid residue having similar chemical properties (e.g., charge or hydrophobicity), and therefore do not alter the functional properties of the molecule. If sequences differ by a conserved substitution, the percentage of sequence identity may be adjusted upward to compensate for the conservative nature of the substitution. Sequences differing by such a conservative substitution are said to have “sequence similarity” or “similarity.” Means for making this adjustment are well known to those skilled in the art. Typically, this involves scoring the conservative substitution as a partial mismatch rather than a complete mismatch, thereby increasing the percentage of sequence identity. Therefore, for example, if the same amino acid is given a score of 1 and a non-conservative substitution is given a score of 0, then a conservative substitution will be given a score between 0 and 1. The scoring of conservative substitutions is calculated, for example, as implemented in the program PC / GENE (Intelligenetics, Mountain View, Calif.).
[0083] As used herein, “sequence identity percentage” means a value determined by comparing two optimally aligned sequences across a comparison window, where portions of the polynucleotide sequences within the comparison window may contain additions or deletions (i.e., gaps) compared to a reference sequence (which does not contain additions or deletions) for optimal alignment of the two sequences. The percentage is calculated by determining the number of positions in both sequences where identical nucleic acid bases or amino acid residues exist, obtaining the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window, and multiplying the result by 100 to obtain the sequence identity percentage.
[0084] The term “substantial identity” of a polynucleotide sequence means that, using one of the alignment programs described herein with standard parameters, the polynucleotide contains sequences that have at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, or 94%, or at least 95%, 96%, 97%, 98%, 99%, or 100% sequence identity compared to a reference sequence. Those skilled in the art will recognize that these values can be appropriately adjusted to determine the corresponding identity of proteins encoded by two nucleotide sequences by considering codon degeneracy, amino acid similarity, leading frame arrangement, etc. Substantial amino acid sequence identity for these purposes typically means at least 70% (e.g., 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%), at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%), at least 90% (e.g., 91%, 92%, 93%, or 94%), or even more than 95% (e.g., 96%, 97%, 98%, 99%, or 100%) sequence identity.
[0085] In the context of peptides, the term "substantial identity" indicates that a peptide contains a sequence that has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, or 94%, or even 95%, 96%, 97%, 98%, or 99% sequence identity with respect to a reference sequence across a specified comparison window. In certain embodiments, optimal alignment is performed using the Needleman and Wunsch homology alignment algorithm (Needleman and Wunsch, J. Mol. Biol., 48, 443 (1970)). An indication that two peptide sequences are substantially identical is that one peptide is immunologically reactive with an antibody produced against the second peptide. Therefore, for example, if the two peptides differ only by conservative substitutions, one peptide is substantially identical to the second peptide. Accordingly, this disclosure also provides nucleic acid molecules and peptides that are substantially identical to the nucleic acid molecules and peptides presented herein.
[0086] Another indicator of substantially identical nucleotide sequences is whether the two molecules hybridize with each other under stringent conditions. Nucleic acid hybridization will be discussed in more detail below.
[0087] Oligonucleotide primers and probes As described herein, the methods provided in this disclosure rely on oligonucleotides, optionally called primers or probes, to identify or detect features contained in nucleic acids obtained from an individual, subject, or patient. The terms “nucleic acid probe” or “nucleic acid-specific probe” refer to a nucleic acid sequence having at least about 80%, e.g., at least about 90%, e.g., at least about 95% continuous sequence identity or homology to a nucleic acid sequence encoding a target sequence of interest. The probes (or oligonucleotides or primers) of this disclosure are at least about 8 nucleotides long (e.g., at least about 8 to 50 nucleotides long, e.g., at least about 10 to 40, e.g., at least about 15 to 35 nucleotides long). The oligonucleotide probes or primers of this disclosure may comprise at least about 8 nucleotides of the 3' of an oligonucleotide having at least about 80%, e.g., at least about 85%, e.g., at least about 90%, e.g., at least about 95% continuous identity to a target sequence of interest.
[0088] Primer pairs are useful for determining the nucleotide sequence of a specific SNP using PCR. A pair of single-stranded DNA primers can be annealed to a sequence within or around the SNP to initiate amplified DNA synthesis of the SNP itself. The first step of the process includes contacting a biological sample obtained from a subject containing nucleic acid with at least one primer to form hybridized DNA. Oligonucleotide primers useful in the method of this disclosure can be any primer consisting of about 8 bases to a maximum of about 80 or 100 bases or more. In one embodiment of this disclosure, the primer is about 10 to about 20 bases.
[0089] The primers themselves can be synthesized using techniques well known in the art. Generally, primers can be prepared using commercially available oligonucleotide synthesizers.
[0090] The primers or probes of this disclosure can be labeled using techniques known to those skilled in the art. For example, the labels used in the assays of the disclosure may be primary labels (in which case the label includes an element to be directly detected) or secondary labels (in which case the detected label is bound to the primary label, as is common in immunological labeling, for example). An overview of labels (also called "tags"), tagging or labeling procedures, and the detection of labels can be found in Polak and Van Noorden (1997) Introduction to Immunocytochemistry, second edition, Springer Verlag, NY and Haugland (1996) Handbook of Fluorescent Probes and Research Chemicals, a combined handbook and catalogue Published by Molecular Probes, Inc., Eugene, Oreg. Primary and secondary labels may include non-detectable and detectable elements. Useful primary and secondary labels in this disclosure include spectral labels, such as fluorescent dyes (e.g., fluorescein and derivatives, e.g., fluorescein isothiocyanate (FITC) and Oregon Green®, rhodamine and derivatives (e.g., Texas Red, tetramethylrhodamine isothiocyanate (TRITC), etc.), digoxigenin, biotin, phycoerythrin, AMCA, CyDyes®, etc.), and radioactive labels (e.g., 3 H, 125 I, 35 S, 14 C, 32 P, 33The label may include P), enzymes (e.g., horseradish peroxidase, alkaline phosphatase), spectrally chromogenic labels, such as gold colloid or colored glass or plastic (e.g., polystyrene, polypropylene, latex) beads. The label may be directly or indirectly coupled to the components of the detection assay (e.g., labeled nucleic acid) according to methods well known in the art. As shown above, a wide variety of labels may be used, and the choice of label depends on the required sensitivity, ease of conjugation with the compound, stability requirements, available equipment, and disposal regulations.
[0091] Generally, detectors used to monitor probe-substrate nucleic acid hybridization are adapted to the specific label used. Typical detectors include spectrophotometers, photocells and photodiodes, microscopes, scintillation counters, cameras, films, and combinations thereof. Examples of suitable detectors are widely available from various commercial sources known to those skilled in the art. Generally, optical images of the substrate containing the bound labeled nucleic acid are digitized for subsequent computer analysis.
[0092] Labeling can be done in the following ways: (1) Chemiluminescence (using horseradish peroxidase and / or alkaline phosphatase with a substrate that produces photons as degradation products) (kits are available from, for example, Molecular Probes, Amersham, Boehringer-Mannheim, and Life Technologies / Gibco BRL); (2) Color generation (using both horseradish peroxidase and / or alkaline phosphatase with a substrate that produces a colored precipitate) (kits are available from Life Technologies / Gibco Labeling methods include (3) hemifluorescence (e.g., using alkaline phosphatase and the substrate AttoPhos(Amersham) or other substrates that produce a fluorescent product), (4) fluorescence (e.g., using Cy-5(Amersham), fluorescein, and other fluorescent labels), and (5) radioactivity (using kinase enzymes or other terminal labeling approaches, nick translation, random priming, or PCR to incorporate a radioactive molecule into the labeled nucleic acid). Other methods for labeling and detection will be readily apparent to those skilled in the art.
[0093] Fluorescent labeling can be used, which has the advantages of requiring less care in handling and being suitable for high-throughput visualization techniques (optical analysis, including digitization of images for analysis in an integrated system including a computer). Preferred labels are typically characterized by one or more of the following: high sensitivity, high stability, low background, low environmental sensitivity, and high specificity in the labeling. The fluorescent moieties that can be incorporated into the label are generally known, including Texas red, dixogenin, biotin, 1- and 2-aminonaphthalene, p,p'-diaminostilbene, pyrene, quaternary phenantholidine salts, 9-aminoacridin, p,p'-diaminobenzophenoneimine, anthracene, oxacarbocyanine, merocyanine, 3-aminoecurenin, perylene, bis-benzoxazole, bis-p-oxazolylbenzene, 1,2-benzophenazine, retinol, bis-3-aminopyridinium salts, hellebrigenin, tetracycline, sterophenol, benzimidazolylphenylamine, 2-oxo-3-chromene, indole, xanthene, 7-hydroxycoumarin, phenoxazine, calicylate, strophanthidin, porphyrin, triarylmethane, flavin, and many others.Numerous fluorescent labels are commercially available from SIGMA Chemical Company (Saint Louis, MO), Molecular Probes, R&D systems (Minneapolis, MN), Pharmacia LKB Biotechnology (Piscataway, NJ), CLONTECH Laboratories, Inc. (Palo Alto, CA), Chem Genes Corp., Aldrich Chemical Company (Milwaukee, WI), Glen Research, Inc., GIBCO BRL Life Technologies, Inc. (Gaithersberg, MD), Fluka ChemicaBiochemika Analytika (Fluka Chemie AG, Buchs, Switzerland), and Applied Biosystems® (Foster City, CA), as well as many other commercial sources known to those skilled in the art.
[0094] Means for detecting and quantifying labels are well known to those skilled in the art. For example, if the label is a radioactive label, means for detection include scintillation counters or photographic film, such as in autoradiography; if the label is optically detectable, typical detectors include microscopes, cameras, photocells, photodiodes, and many other detection systems that are widely available.
[0095] Oligonucleotide primers or probes having any of a wide variety of base sequences can be prepared according to techniques well known in the art. Suitable bases for preparing oligonucleotide primers or probes are naturally occurring nucleotide bases, such as adenine, cytosine, guanine, uracil, and thymine; as well as non-naturally occurring or "synthetic" nucleotide bases, such as 7-deaza-guanine, 8-oxo-guanine, 6-mercaptoguanine, 4-acetylcytidine, 5-(carboxyhydroxyethyl)uridine, 2'-O-methylcytidine, and 5-carboxymethylaminomethyl-2-thioidine (thio ridine), 5-carboxymethylaminomethyluridine, dihydrouridine, 2'-O-methylpseuduridine, β,D-galactosylquyeosin, 2'-O-methylguanosine, inosine, N6-isopentenyladenosine, 1-methyladenosine, 1-methylpseuduridine, 1-methylguanosine, 1-methylinosine, 2,2-dimethylguanosine, 2-methyladenosine, 2-methylguanosine, 3-methylcytidine, 5-methylcytidine, N6-methyladenosine, 7-methyl Tylguanosine, 5-methylaminomethyluridine, 5-methoxyaminomethyl-2-thiouridine, β,D-mannosylqueosin, 5-methoxycarbonylmethyluridine, 5-methoxyuridine, 2-methylthio-N6-isopentenyladenosine, N-((9-β-D-ribofuranosyl-2-methylthiopurine-6-yl)carbamoyl)threonine, N-((9-β-D-ribofuranosylpurine-6-yl)N-methylcarbamoyl)threonine, uridine-5-oxyacetate methyl The following can be selected from esters, uridine-5-oxyacetic acid, wybutoxosine, pseudouridine, cueosin, 2-thiocytidine, 5-methyl-2-thiouridine, 2-thiouridine, 2-thiouridine, 5-methyluridine, N-((9-β-D-ribofuranosylpurine-6-yl)carbamoyl)threonine, 2'-O-methyl-5-methyluridine, 2'-O-methyluridine, wybutoxosin, and 3-(3-amino-3-carboxypropyl)uridine.Any oligonucleotide skeleton may be employed, including DNA, RNA (however, RNA is less preferred than DNA), modified sugars such as carbocyclic compounds, and sugars containing 2' substitutions such as fluoro and methoxy. The oligonucleotide may be an oligonucleotide in which at least one or all of the internucleotide crosslinking phosphate residues are modified phosphates, e.g., methyl phosphonate, methyl phosphonotlioate, phosphoroinorpholidate, phosphoropiperazidate, and phosplioramidate (for example, every other internucleotide crosslinking phosphate residue may be modified as described). The oligonucleotide may also be a “peptide nucleic acid” as described in Nielsen et al., Science, 254:1497-1500 (1991).
[0096] As used herein, a “single base pair extension probe” is a nucleic acid that selectively recognizes a single nucleotide polymorphism (i.e., either A or G in an A / G polymorphism). Generally, these probes take the form of DNA primers (e.g., as in PCR primers) that are modified to release a phosphor upon primer integration. An example of this is the Taqman® probe, which uses the 5' exonuclease activity of the enzyme Taq polymerase to measure the amount of a target sequence in a sample. The TaqMan® probe consists of an 18-22 bp oligonucleotide probe, labeled with a reporter phosphor at the 5' end and a quencher phosphor at the 3' end. Integration of the probe molecule into the PCR strand (which occurs because the probe set is contained in a mixture of PCR primers) releases the reporter phosphor from the influence of the quencher. The primer must be able to recognize the target binding site. Some primer extension probes can be “activated” directly by DNA polymerase without a complete PCR extension cycle.
[0097] The only requirement is that the oligonucleotide probe should have a sequence in which at least a portion can bind to a known portion of the DNA sample sequence. The nucleic acid probes provided by this disclosure are useful for a number of purposes.
[0098] Methods for detecting nucleic acids A. Amplification Amplification of DNA present in a biological sample according to the methods of this disclosure can be carried out by any means known in the art. Examples of preferred amplification techniques include, but are not limited to, polymerase chain reactions (including reverse transcriptase polymerase chain reactions for RNA amplification), ligase chain reactions, strand displacement amplification, transcription-based amplification, auto-persistent sequence replication (i.e., "3SR"), Qbeta replicase systems, nucleic acid sequence-based amplification (i.e., "NASBA"), repair chain reactions (i.e., "RCR"), and boomerang DNA amplification (i.e., "BDA").
[0099] The base incorporated into the amplification product can be a native base or a modified base (modified before or after amplification), and the base can be selected to optimize subsequent detection steps (e.g., electrochemical detection steps).
[0100] Polymerase chain reaction (PCR) can be carried out according to known techniques. See, for example, U.S. Patents 4,683,195; 4,683,202; 4,800,159; and 4,965,188. Generally, PCR involves first treating a nucleic acid sample (e.g., in the presence of a heat-stable DNA polymerase) with one oligonucleotide primer for each strand of a particular sequence to be detected, under hybridizing conditions such that an extension product of each primer complementary to each nucleic acid strand is synthesized, wherein the primers are sufficiently complementary to each strand of the particular sequence to hybridize, and so that the extension product synthesized from each primer can act as a template for the synthesis of the extension product of the other primer when separated from its complement; and then treating the sample under denaturing conditions to separate the primer extension products from their templates, if one or more sequences to be detected are present. These steps are repeated periodically until the desired degree of amplification is obtained. Detection of the amplified sequence may be carried out by adding an oligonucleotide probe (e.g., an oligonucleotide primer or probe of this disclosure) that can hybridize to the reaction product, wherein the probe has a detectable label, and then by detecting the label according to known techniques. Various labels that can be incorporated into or functionally linked to nucleic acids, such as radiolabeling, enzymatic labeling, and fluorescent labeling, are well known in the art. If the nucleic acid to be amplified is RNA, amplification may be carried out by initial conversion to DNA by reverse transcriptase according to known techniques.
[0101] Strand substitution amplification (SDA) can be carried out according to known techniques. For example, SDA can be carried out with a single amplification primer or a pair of amplification primers, with exponential amplification being achieved with the latter. Generally, SDA amplification primers include, in the 5' to 3' direction, a flanking sequence (its DNA sequence is not important), a restriction site for the restriction enzyme employed in the reaction, and an oligonucleotide sequence (e.g., an oligonucleotide primer or probe as described herein) that hybridizes to the target sequence to be amplified and / or detected. The flanking sequence, which works to facilitate binding of the restriction enzyme to the recognition site and provides a DNA polymerase priming site after the restriction site has been nicked, can be about 15–20 nucleotides long. The restriction site is functional in the SDA reaction. For example, the oligonucleotide primer or probe portion can be about 13–15 nucleotides long.
[0102] Ligase chain reaction (LCR) can also be carried out according to known techniques. Generally, the reaction is carried out with two pairs of oligonucleotide probes: one pair binds to one strand of the sequence to be detected; the other pair binds to the other strand of the sequence to be detected. Each pair together completely overlaps with the corresponding strand. The reaction is carried out by first denaturing (e.g., separating) the strand of the sequence to be detected, then reacting the strand with the two pairs of oligonucleotide probes in the presence of a heat-stable ligase so that each pair of oligonucleotide probes ligates to each other, then separating the reaction product, and then repeating the process periodically until the sequence is amplified to the desired degree. Detection can then be carried out in a manner similar to that described above with respect to PCR.
[0103] Specific SNPs at specific gene loci can be detected according to the methods described herein. Useful techniques in the methods described herein include, but are not limited to, direct DNA sequencing, PFGE analysis, allele-specific oligonucleotides (ASOs), dot blot analysis, and denaturing gradient gel electrophoresis, and are well known to those skilled in the art.
[0104] Several methods can be used to detect DNA sequence variations. Sequence variations can be detected by direct DNA sequencing, either manual or automated fluorescence sequencing. Another approach is single-stranded higher-order polymorphism assay (SSCA). This method does not detect all sequence changes, especially when the DNA fragment size is larger than 200 bp, but it can be optimized to detect most DNA sequence variations. Although a drawback is reduced detection sensitivity, the increased processing capacity possible with SSCA makes it an attractive and viable alternative to direct sequencing for variant detection on a study basis. Fragments with shifted mobility on an SSCA gel can then be sequenced to determine the precise nature of the DNA sequence variation. Other approaches based on detecting mismatches between two complementary DNA strands include clamped denatured gel electrophoresis (CDGE), heterodouble-strand analysis (HA), and chemical mismatch cleavage (CMC). Once a sequence change is identified, allele-specific detection approaches, such as allele-specific oligonucleotide (ASO) hybridization, can be used to rapidly screen numerous other samples for the same sequence change (e.g., mutation, polymorphism). Such techniques can utilize probes labeled with gold nanoparticles to produce visual color results.
[0105] SNP detection can be performed by sequencing a desired target region using techniques well known in the art. Alternatively, the sequence can be directly amplified from a genomic DNA preparation derived from the target tissue using known techniques. The DNA sequence of the amplified sequence can then be determined.
[0106] Several well-known methods exist for more complete, yet still indirect, confirmation of the presence of mutant alleles: 1) single-strand higher-order structure analysis (SSCA); 2) denaturing gradient gel electrophoresis (DGGE); 3) RNase protection assays; 4) allele-specific oligonucleotides (ASOs); 5) the use of proteins that recognize nucleotide mismatches, such as the E. coli mutS protein; and / or 6) allele-specific PCR. For allele-specific PCR, primers that hybridize to a specific allele at their 3' end are used. If no specific mutation is present, no amplification product is observed. Amplification-refractory mutagenesis systems (ARMS) can also be used. Gene insertions and deletions can also be detected by cloning, sequencing, and amplification. In addition, allele alterations or insertions in polymorphic fragments can be scored using restriction enzyme fragment length polymorphism (RFLP) probes against the gene or surrounding marker genes. Other techniques for detecting insertions and deletions, as known in the art, can be used.
[0107] In the first three methods (SSCA, DGGE, and RNase protection assays), a new electrophoretic band appears. SSCA detects differentially moving bands because the sequence change causes differences in single-strand intramolecular base pairing. RNase protection involves cleavage of the mutant polynucleotide into two or more smaller fragments. DGGE uses a denaturing gradient gel to detect differences in the migration rate of the mutant sequence compared to the wild-type sequence. In allele-specific oligonucleotide assays, oligonucleotides are designed to detect specific sequences, and the assay is performed by detecting the presence or absence of a hybridization signal. In mutS assays, the protein binds only to sequences containing nucleotide mismatches in the heteroduplex between the mutant and wild-type sequences.
[0108] A mismatch as defined in this disclosure is a hybridized nucleic acid double helix in which the two strands are not 100% complementary. The lack of complete homology may result from deletions, insertions, inversions, or substitutions. Mismatch detection can be used to detect point mutations in genes or their mRNA products. These techniques are less sensitive than sequencing but are easier to perform on a large number of samples. An example of a mismatch cleavage technique is the RNase protection method. Either a riboprobe and mRNA or DNA isolated from tumor tissue are annealed (hybridized) to each other, and then digested with the enzyme RNase A to detect some mismatches in the double-stranded RNA structure. If a mismatch is detected by RNase A, RNase A cleaves at the mismatch site. Therefore, when the annealed RNA preparations are separated on an electrophoretic gel matrix, if mismatches are detected and cleaved by RNase A, smaller RNA products than the full-length double-stranded RNA will be observed for both the riboprobe and mRNA or DNA. Riboprobes do not need to be the full length of mRNA or a gene, but can be a segment of either. If riboprobes contain only mRNA or a gene segment, it is desirable to use a large number of these probes to screen the entire mRNA sequence for mismatches.
[0109] Similarly, DNA probes can be used to detect mismatches through enzymatic or chemical cleavage. Alternatively, mismatches can be detected by the shift in the electrophoretic mobility of mismatched double helices relative to matched double helices. Either riboprobes or DNA probes can be used to amplify mRNA or DNA from potentially mutant cells using PCR before hybridization.
[0110] B. Sequencing Sanger sequencing is used in a variety of applications, from targeted sequencing to the identification of variants identified using orthogonal sequencing, due to its sensitivity and relative simplicity in both workflow and technique. Sanger sequencing utilizes a strand termination method to provide nucleotide base identity and order in a given DNA strand. This method utilizes chemical analogues of four nucleotide bases (i.e., ddNTPs) that lack the hydroxyl groups necessary for the elongation of the polynucleotide chains that form the DNA molecule. By mixing radiolabeled and later fluorescently labeled ddNTPs with template DNA, the ddNTPs are randomly incorporated, and when the strands are terminated, strands of each possible length are produced. In contrast to Sanger sequencing, Maxam and Gilbert used a chemical cleavage technique. The main advantages of the Maxam-Gilbert technique compared to the Sanger method are that sequencing can be performed from the original DNA fragment instead of enzymatic copying, PCR is not required, and this method is less susceptible to secondary structure errors or enzymatic errors. The products generated in the sequencing reaction can be separated by ion electrophoresis on an acrylamide gel or by capillary electrophoresis.
[0111] Another sequencing technique called pyrosequencing has been developed, which uses a two-enzyme process to convert pyrophosphate to ATP using adenosine triphosphate (ATP) sulfurylase, and then uses this as a substrate for luciferase, thus producing light proportional to the amount of pyrophosphate. Additional approaches for sequencing nucleic acids include emulsion polymerase chain reaction (PCR), reversible terminators, and sequencing by oligonucleotide ligation and detection. Capillary electrophoresis (CE) instruments can also be used for sequencing.
[0112] High-throughput sequencing technologies, also known as next-generation sequencing (NGS), are also being developed. NGS is a large-scale parallel processing method that sequences millions of fragments simultaneously per run. This high-throughput process allows for sequencing hundreds to thousands of genes at once. NGS also offers greater discoveries, enabling the detection of novel or rare variants through deep sequencing. The scope of NGS analysis can range from a small number of genes to the entire genome. Whole-genome sequencing (WGS) and whole-exome sequencing (WES) provide sequences of DNA bases across the genome and exome, respectively. Whole-transcriptome sequencing provides sequence information on coding and multiple non-coding types of RNA to assess variations and gene expression levels across the entire transcriptome. Targeted sequencing covers a relatively small set of genes or a target region of interest (e.g., to determine the presence or absence of SNPs). Real-time sequencing and single-molecule sequencing (SMS) have also been developed, enabling the accurate sequencing of long chains of nucleic acids without intermediates and without prior transcription or amplification.
[0113] Nanopore-based sequencing technology detects the unique electrical signals generated when different molecules pass through nanopores using semiconductor-based electronic detection systems. This technology contributes to high-throughput and cost-effective sequencing solutions. At the heart of the technology are biological nanopores, which are protein pores embedded in membranes, while the logic lies in the electronic circuits and unique chemistry of semiconductor integrated circuits. Chip-embedded electronic sensor technology enables automated membrane assembly and nanopore insertion, while simultaneously allowing for active control of individual sensors on the circuit. See, for example, Oxford Nanopore and Pacific Biosciences. Nanopore and electronic sensor sequencing technologies can be used to directly determine DNA methylation using native non-bisulfite DNA.
[0114] C. Hybridization The phrase "specifically hybridizes" refers to a molecule binding, double-stranding, or hybridizing to only a specific nucleotide sequence when that sequence is present in a complex mixture (e.g., whole cell) of DNA or RNA, under stringent conditions. "Substantially binding" refers to complementary hybridization between a primer or probe nucleic acid and a target nucleic acid, encompassing a small number of mismatches that can be adapted by reducing the stringency of the hybridization medium to achieve the desired detection of the target nucleic acid sequence.
[0115] Generally, stringent conditions are those that meet the thermal melting point (T) for a specific arrangement at a given ionic strength and pH. m The temperature is selected to be approximately 5°C lower than the specified temperature. However, stringent conditions encompass temperatures in the range of approximately 1°C to approximately 20°C, depending on the desired degree of stringency, as otherwise limited herein. Nucleic acids that do not hybridize with each other under stringent conditions are still substantially identical if the polypeptides they encode are substantially identical. This can occur, for example, when copies of nucleic acids are produced using the maximum codon degeneracy permitted by the gene code. One indication that two nucleic acid sequences are substantially identical is that the polypeptide encoded by the first nucleic acid is immunologically cross-reactive with the polypeptide encoded by the second nucleic acid.
[0116] "Stringent conditions" refer to (1) low ionic strength and high temperature for washing, for example, using 0.015 M NaCl / 0.0015 M sodium citrate (SSC); 0.1% sodium lauryl sulfate (SDS) at 50°C, or (2) using a denaturing agent such as formamide during hybridization, for example, 50% formamide, together with 0.1% bovine serum albumin / 0.1% Ficoll / 0.1% polyvinylpyrrolidone / 50 mM sodium phosphate buffer (pH 6.5) containing 750 mM NaCl and 75 mM sodium citrate at 42°C. Another example is the use of 50% formamide, 5×SSC (0.75 M NaCl, 0.075 M sodium citrate), 50 mM sodium phosphate (pH 6.8), 0.1% sodium pyrophosphate, 5× Denhardt's solution, sonicated salmon sperm DNA (50 μg / ml), 0.1% SDS, and 10% dextran sulfate at 42°C, accompanied by washing with 0.2×SSC and 0.1% SDS. Examples of other stringent conditions are well known in the art.
[0117] In the context of nucleic acid hybridization experiments such as Southern and Northern hybridization, "stringent hybridization conditions" and "stringent hybridization washing conditions" are sequence-dependent and differ under different environmental parameters. Longer sequences hybridize specifically at higher temperatures. The thermal melting point (Tm) is the temperature (under specified ionic strength and pH) at which 50% of the target sequence hybridizes to a perfectly matched primer or probe sequence. Specificity is typically a function of post-hybridization washing, with key factors being the ionic strength and temperature of the final washing solution. For DNA-DNA hybrids, Tm can be approximated by the following Meinkoth and Wahl equation (1984, Anal. Biochem., 138(2):267-84); T m81.5℃ + 16.6(log M) + 0.41(%GC) - 0.61(%form) - 500 / L; where M is the molar concentration of monovalent cations, %GC is the percentage of guanosine and cytosine nucleotides in the DNA, %form is the percentage of formamide in the hybridization solution, and L is the length of the hybrid at the base pair. Tm decreases by about 1℃ for each 1% mismatch; therefore, Tm, hybridization, and / or washing conditions can be adjusted to hybridize to sequences of desired identity. For example, if seeking sequences with >90% identity, Tm can be reduced by 10℃. Generally, stringent conditions are selected so that Tm is about 5℃ lower than for a particular sequence and its complement at a given ionic strength and pH. However, strictly stringent conditions can utilize hybridization and / or washing at temperatures 1, 2, 3, or 4°C lower than Tm; moderately stringent conditions can utilize hybridization and / or washing at temperatures 6, 7, 8, 9, or 10°C lower than Tm; and low stringency conditions can utilize hybridization and / or washing at temperatures 11, 12, 13, 14, 15, or 20°C lower than Tm. Those skilled in the art will understand that variations in the stringency of the hybridization and / or washing solutions are inherently described using the formulas, hybridization and washing compositions, and desired temperatures. If the desired degree of mismatching results in temperatures below 45°C (aqueous solution) or 32°C (formamide solution), the SSC concentration can be increased to allow the use of higher temperatures. Generally, highly stringent hybridization and washing conditions are selected to be about 5°C lower than Tm for a particular arrangement at given ionic strength and pH.
[0118] Examples of highly stringent washing conditions are 0.15 M NaCl at about 72 °C for about 15 minutes. An example of stringent washing conditions is washing with 0.2×SSC at 65 °C for 15 minutes. Often, a low stringency wash to remove background signal is performed prior to a high stringency wash. For example, an example of medium stringency washing for double-strands greater than 100 nucleotides is 1×SSC at 45 °C for 15 minutes. For short nucleotide sequences (e.g., about 10 - 50 nucleotides), stringent conditions typically include a salt concentration of less than about 1.5 M, less than about 0.01 - 1.0 M, Na ion concentration (or other salts) at pH 7.0 - 8.3, and the temperature is typically at least about 30 °C and at least about 60 °C for long oligonucleotides (e.g., >50 nucleotides). Stringent conditions can also be achieved by the addition of destabilizing agents such as formamide. Generally, a signal-to-noise ratio of 2× (or higher) of what is observed for non-related probes in a particular hybridization assay indicates detection of specific hybridization. Nucleic acids that do not hybridize to each other under stringent conditions are still substantially identical if the proteins they encode are substantially identical. This can occur, for example, when copies of a nucleic acid are made using the maximum codon degeneracy allowed by the genetic code.
[0119] Very stringent conditions are the T for a particular probe mThis can be equivalent to: An example of stringent conditions for hybridization of complementary nucleic acids having more than 100 complementary residues on a filter in Southern or Northern blotting is 50% formamide, e.g., hybridization at 37°C in 50% formamide, 1 M NaCl, 1% SDS, and washing at 60-65°C in 0.1×SSC. Exemplary low-stringency conditions include hybridization at 37°C in a buffer solution of 30-35% formamide, 1 M NaCl, 1% SDS (sodium dodecyl sulfate), and washing at 50-55°C in 1×~2×SSC (20×SSC = 3.0 M NaCl / 0.3 M trisodium citrate). Exemplary moderate stringency conditions include hybridization at 37°C with 40–45% formamide, 1.0 M NaCl, and 1% SDS, and washing at 55–60°C with 0.5 × ~ 1 × SSC.
[0120] Northern blotting, or "Northern analysis," is a method used to identify RNA sequences that hybridize to known probes, such as oligonucleotides, DNA fragments, cDNA or its fragments, or RNA fragments. The probes are, 32 The RNA to be analyzed can be labeled with radioactive isotopes such as 3P by biotinylation or enzymatically. The RNA to be analyzed is usually separated by electrophoresis on an agarose gel or polyacrylamide gel using standard techniques well known in the art, and then transferred to a nitrocellulose, nylon, or other suitable membrane for hybridization with the probe.
[0121] Nucleic acid samples may be brought into contact with oligonucleotide probes in any preferred manner known to those skilled in the art. For example, a DNA sample may be solubilized in a solution and then brought into contact with oligonucleotides by solubilizing the oligonucleotides in a solution containing the DNA sample under conditions that allow hybridization. Preferred conditions are well known to those skilled in the art. Alternatively, a DNA sample may be solubilized in a solution containing oligonucleotide probes immobilized on a solid support, and the DNA sample may be brought into contact with oligonucleotides by immersing the solid support, on which the oligonucleotides are immobilized, in a solution containing the DNA sample.
[0122] The term "substrate" refers to any solid support to which oligonucleotides can be attached. The substrate material may be covalently or otherwise modified with coatings or functional groups to facilitate the binding of oligonucleotides. Preferred substrate materials include, among others, polymers, glass, semiconductors, paper, metals, gels, and hydrogels. The substrate may have any physical shape or size, e.g., a plate, a strip, or microparticles. The term "spot" refers to a distinct location on the substrate to which an oligonucleotide of a known sequence is attached. A spot may be an area on a planar substrate, or it may be a microparticle distinguishable from other microparticles, for example. The term "bound" means attached to a solid substrate. A spot is "bound" to a solid substrate if it is attached to a specific location on the substrate for the purpose of a screening assay.
[0123] In certain embodiments of the present disclosure, the substrate is a polymer, glass, semiconductor, paper, metal, gel, or hydrogel. In certain embodiments of the present disclosure, the kit may further comprise a solid substrate and at least one control oligonucleotide, the at least one control oligonucleotide bound to the substrate at a separate spot.
[0124] In certain embodiments of this disclosure, the solid substrate is a microarray. The terms “array” or “microarray” are used herein as synonyms to refer to a plurality of primers or probes attached to one or more distinguishable spots on a substrate. A microarray may comprise a single substrate or a plurality of substrates, e.g., a plurality of beads or microspheres. A “copy” of a microarray contains primers or probes of the same type and arrangement.
[0125] Methods for detecting or predicting cardiovascular disease, or for estimating survival. Better risk assessment for cardiovascular disease, or earlier detection of cardiovascular disease, is the first step towards more effective prevention. Individuals identified as having CHD, a higher risk for CHD events (e.g., 69% PPV for CHD), or CHD can be promptly followed up for further examinations such as coronary calcium or angiography, and more aggressive interventions. They can be regularly examined or re-examined for management and monitoring to determine severity, identify, customize, and optimize interventions (e.g., lifestyle, medical, and treatments). Conversely, individuals at lower risk (e.g., 99% NPV for CHD) or those without CHD can be regularly re-examined for management and monitoring to determine severity, identify, customize, and optimize interventions (e.g., lifestyle, medical, and treatments), due to the dynamic nature of DNA methylation, to ensure continued prevention. In addition, individuals with CVD (e.g., CHD) can be assessed, and their viability can be estimated based on their genetic and / or epigenetic profiles. It will be recognized that in some circumstances (for example, for monitoring or follow-up purposes; to assess the viability of individuals already determined to be at risk of or have CVD), the step involving the detection of one or more SNPs may not be necessary, because that information may already be available (for example, from a previous determination) or may not be necessary to predict CVD or viability or severity from CVD, or to identify, customize, and optimize interventions, or to manage individuals.
[0126] Compared to the integrated genetic-epigenetic model, conventional risk factor-based calculators and / or other detection tests, such as loading tests, have generally been considerably less sensitive, less generalizable, and have also shown a gender gap in performance. In contrast, the integrated genetic-epigenetic model described herein has the ability to capture and better understand the complex nature of CVD through three angles: genetics (static genetic risk), DNA methylation (dynamic acquired risk), and genetic confounding of methylation signatures (i.e., G+M+G×M).
[0127] This disclosure provides a method for determining whether a subject may have CVD by determining the methylation status of a CpG dinucleotide repeat or CpG dinucleotide repeat motif region, wherein the methylation status of the CpG dinucleotide is associated with CVD. However, the same principle applies to assessing the prevalence and / or incidence of a number of different types of CVD, including, but not limited to, coronary heart disease (CHD) (e.g., occlusive CHD), stroke, arrhythmia, cardiac arrest, congestive heart failure, atherosclerotic cardiovascular disease (ASCVD), and its associated cardiovascular events (CVEs), including, for example, occlusive coronary artery disease (CAD), ischemia without occlusion (INOCA), myocardial infarction (MI), stroke (e.g., TIA, hemorrhagic), and cardiovascular death. This disclosure also provides a method for determining the severity of a subject who has or is at risk of having CVD, and / or estimating their survival. In a particular embodiment, the method determines the methylation status and / or SNPs of a plurality of CpG dinucleotides (e.g., any integer between 1 and 10,000, e.g., at least 100).
[0128] Where used herein, “biological sample” encompasses essentially any type of sample obtained from a subject that can be used in the manner described herein. A biological sample may be any bodily fluid, tissue, or any other sample from which clinically relevant biomarker levels can be determined. “Biological sample” may also encompass cells in culture, cell supernatants, cell lysates, blood, serum, plasma, urine, cerebrospinal fluid, biological fluids, and tissue samples. Various techniques and reagents are used in the methods of this disclosure. In one aspect of this disclosure, a blood sample, or a sample derived from blood, such as plasma, circulating, peripheral, or lymphocytes, is assayed for the presence of one or more SNPs and / or the methylation status of one or more CpG dinucleotides. A biological sample may also be saliva. Typically, a biological sample containing nucleic acids is provided and tested. Biological samples may be obtained from a subject using well-known techniques such as finger prick, venipuncture, lumbar puncture, liquid samples such as saliva or urine, or tissue biopsy.
[0129] As used herein, the term “healthy” means that the subject does not exhibit any particular condition and is not statistically likely to be susceptible to any particular condition.
[0130] Prevalence is defined by the American Psychological Association (APA) as "the total number or percentage of cases (e.g., disease or disorder) present in a population" (APA Dictionary of Psychology, (American Psychological Association, Washington, DC, 2007)). In some cases, point prevalence is used to describe the prevalence of cases at discrete points in time, while period prevalence is used to describe the number of cases present over a period of time (e.g., one month, one year). Prevalence is typically expressed as a percentage per population (e.g., number of cases per 100,000 people) rather than as an absolute number or percentage.
[0131] Similarly, incidence is defined by the APA as "the proportion of new cases of a given event or condition (e.g., disorder, disease, symptom, or injury) in a particular population over a given period of time" (APA Dictionary of Psychology, (American Psychological Association, Washington, DC, 2007)). As used herein, the term “incidence” is defined as the tendency or susceptibility of an subject to manifesting a condition, in this case CVD (e.g., CHD). In some cases, the period may be one year or less; in some cases, the period may be longer than one year (e.g., two years, five years, ten years).
[0132] Diagnosis is defined by the APA as "the process of identifying and determining the nature of a disease or disorder by its signs and symptoms, through the use of assessment techniques (e.g., tests and examinations) and other available evidence" (APA Dictionary of Psychology, (American Psychological Association, Washington, DC, 2007)). A diagnosis can refer to the present period, or to past or future periods.
[0133] Similarly, prognosis is defined by the APA as "a prediction of the course, duration, severity, and outcome of a condition, disease, or disability" (APA Dictionary of Psychology, (American Psychological Association, Washington, DC, 2007)). Prognosis can be assessed over periods such as one month, six months, one year, five years, ten years, or longer.
[0134] Risk assessment is defined as "a study of a subject that seeks to determine the probability that a person will develop a particular disease, or, if the disease is already present, the probability that the person will suffer from disease progression or death as a result" (Youngson, 2005, Collins Dictionary of Medicine). In some cases, risk assessment is based on a state or event, not on a disease. In some cases, risk assessment is made over a period of time (e.g., several months, several years).
[0135] As used herein, survival or viability typically refers to the number of days remaining in the life of an individual, measured from the time a biological sample (e.g., saliva sample, blood sample, etc.) is obtained and used to make this determination.
[0136] Those skilled in the art will understand that in some cases, for example, the "detection" of a biomarker (e.g., the presence or absence of a biomarker) can result in a "diagnosis." Similarly, those skilled in the art will understand that in some cases, "detection" and "diagnosis" are used interchangeably.
[0137] Biomarkers that can be used in methods for detecting CVD in a subject or for estimating survival (e.g., predictive or prognostic) in a subject that has CVD or is at risk of having CVD are described herein. Such methods typically include the steps of: providing a biological sample derived from the subject; contacting DNA derived from the biological sample with a bisulfite under alkaline conditions; contacting the bisulfite-treated DNA with at least one first oligonucleotide primer of at least 8 nucleotides length that is complementary to a sequence containing a CpG dinucleotide (e.g., a GC locus referred to as cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584, or another biomarker derived from Annex A); and determining the methylation status of the CpG dinucleotide. It will be understood that at least one first oligonucleotide probe can detect either unmethylated or methylated CpG dinucleotides. Such a method may further include the step of determining the genotype of a single nucleotide polymorphism (SNP) (e.g., rs2869675, rs4376434, rs12129789, rs7585056, rs710987, rs4639796, rs1333048, rs12714414, rs942317, and rs1441433, or another biomarker from Annex C), or a second SNP in linkage disequilibrium with the first SNP. As described herein, the methylation of one or more specific CpG dinucleotides and the presence of one or more specific SNPs can be used to predict CVD in a subject. Furthermore, as described herein, the methylation of one or more specific CpG dinucleotides and / or the presence of one or more specific SNPs can be used to estimate the viability of CVD.
[0138] In some embodiments, the method further comprises contacting bisulfite-treated DNA with at least one second oligonucleotide probe of at least 8 nucleotides in length, complementary to a sequence containing a CpG dinucleotide, wherein the at least one second oligonucleotide probe detects either unmethylated or methylated CpG dinucleotides, regardless of which is detected by the at least one first oligonucleotide probe.
[0139] In some embodiments, the ratio of methylated CpG dinucleotides to unmethylated CpG dinucleotides in a biological sample can be determined as part of a method described herein. By determining the ratio of methylated CpG dinucleotides to unmethylated CpG dinucleotides, it may be possible to estimate or determine risks or outcomes.
[0140] It will be recognized that any number of techniques can be used to determine the methylation status of one or more CpG dinucleotides and the presence (or absence) of SNPs, such as amplification and / or sequencing. Amplification and sequencing are well-known techniques in the art and are routinely used to determine both the methylation status and the presence / absence of SNPs of a particular sequence.
[0141] A method is provided for determining the presence of CHD-related biomarkers in a biological sample from a target. A similar approach can be used for any other form of CVD. Such a method typically comprises the steps of providing a first portion of a biological sample and contacting DNA from the first portion with a bisulfite under alkaline conditions. The bisulfite-treated first portion can be contacted with a first oligonucleotide probe that is at least 8 nucleotides long and complementary to a sequence containing a CpG dinucleotide (e.g., a CG locus referred to as cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584, or another biomarker from Annex A). Furthermore, if necessary or desired, a second portion of the biological sample may be contacted with a nucleic acid probe of at least 8 nucleotides length that is complementary to the SNP (e.g., rs2869675, rs4376434, rs12129789, rs7585056, rs710987, rs4639796, rs1333048, rs12714414, rs942317, and rs1441433, or another biomarker from Annex C).
[0142] As described herein, the percentage of methylation of a CpG dinucleotide (or a CpG dinucleotide in linkage disequilibrium with one or more such CpG dinucleotides) at one or more of the GC loci designated as cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584, as well as rs2869675, rs4376434, rs12129789, rs7585056, r Nucleotide identity at one or more SNPs designated as s710987, rs4639796, rs1333048, rs12714414, rs942317, and rs1441433 (or at SNPs in linkage disequilibrium with one or more such SNPs) is a CVD-related biomarker and can be used to predict the likelihood of an individual developing CVD and / or to determine the prognosis of an individual with respect to disease severity or outcome (e.g., survival).
[0143] In addition to the SNPs and / or CpG biomarkers identified herein, one or more clinical indicators may be used to assist in any or both of the following: diagnosis and prognosis, selection, customization or optimization, management, or monitoring of interventions (e.g., lifestyle, therapeutic, or medical). Such clinical indicators may include, without limitation,: demographics (e.g., age, sex, race); vital signs (e.g., heart rate (beats / min), systolic BP (mm Hg), diastolic BP (mm Hg)); medical history (e.g., smoking, atrial fibrillation / flutter, hypertension, coronary heart disease, myocardial infarction, heart failure, peripheral artery disease, COPD, diabetes mellitus (type 1 or type 2), CVA / TIA, chronic kidney disease, hemodialysis, angioplasty (peripheral or coronary), stent (peripheral or coronary), CABG, percutaneous coronary intervention); medications (ACE inhibitors / ARBs, beta-blockers, aldosterone antagonists, loop diuretics, nitrates, CCBs, statins, aspirin, warfarin, clopidogrel); and computed tomography angiography (e.g., atomic stenosis). stenosis), FFR-CT, plaque type, total plaque; echocardiogram results (e.g., LVEF (%), RSVP (mm Hg)); stress test results (e.g., ischemia on scan, ischemia on ECG); angiography results (e.g., coronary artery stenosis of 70% or more in two or more vessels, coronary artery stenosis of 70% or more in three or more vessels); and / or laboratory measurements (e.g., sodium, blood urea nitrogen (mg / dL), creatinine (mg / dL), eGFR (median, CKDEPI), total cholesterol (mg / dL), LDL cholesterol (mg / dL), ribitol, hemoglobin, hematocrit, triglycerides, alkaline phosphatase, HbA1c, HDL-C, non-HDL-C, ApoB, LDL-P 1 , HDL-P 1sdLDL-C, VLDL-C, Lp(a), hs-CRP, LpPLA2 activity, homocysteine, type B natriuretic peptide, glycated hemoglobin (%), glucose (mg / dL), HGB (mg / dL), C-reactive protein (mg / L), NT-proBNP, KIM-1, osteopontin, TIMP-1, kidney damage molecule-1, N-terminal pro-type B natriuretic peptide, osteopontin, tissue metalloproteinase inhibitor-1, uridine, carotene-3, ribitol, 1-stearoyl-2-adrenoyl-GPC, N-acetyl-isoptreanin, lysophosphatidylcholine, vanyl lactate, 3-ureidopropionate, serum paraoxonase, bone morphogenetic protein 1, carboxypeptidate B2 B2) Albumin, histone H2B type 1-K, versican core protein, insulin-like growth factor-binding protein 2, matrix-remolding associated protein 5, etc.
[0144] Kit for detecting coronary artery disease (CVD) Further embodiments of this disclosure provide products and kits containing probes, oligonucleotides, and / or antibodies. Such products can be used in the methods described herein. The products may include one or more containers, for example, with labels. Preferred containers include, for example, bottles, vials, and test tubes. Containers may be formed from a variety of materials, such as glass or plastic. Containers can hold compositions containing one or more active substances that are effective for carrying out the methods described herein. Labels on the containers indicate that the compositions may be used for a particular application. Kits of this disclosure typically include the above-mentioned containers, as well as one or more other containers containing materials desirable from a commercial and user perspective, including buffers, diluents, filters, and accompanying documentation with instructions for use.
[0145] In certain embodiments, the disclosure provides a kit for determining the methylation status of at least one CpG dinucleotide and, optionally, the presence of at least one single nucleotide polymorphism (SNP). In certain embodiments, a kit as described herein may contain a number of primers, any integer between 1 and 10,000, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, ... 9997, 9998, 9999, 10,000. As used herein, the terms “nucleic acid primer,” “nucleic acid probe,” or “oligonucleotide” encompass both DNA and RNA sequences. In certain embodiments, a primer or probe may be physically located on a single solid substrate or on multiple substrates.
[0146] Kits such as those described herein may include at least one first nucleic acid primer (e.g., at least 8 nucleotides long) complementary to a bisulfite-converted nucleic acid sequence containing a CpG dinucleotide (detected at GC loci designated, e.g., cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584), as well as, in some examples, at least one second nucleic acid primer (e.g., at least 8 nucleotides long) complementary to a SNP (e.g., rs2869675, rs4376434, rs12129789, rs7585056, rs710987, rs4639796, rs1333048, rs12714414, rs942317, and rs1441433). At least one first nucleic acid primer can detect methylated or unmethylated CpG dinucleotides.
[0147] It will be recognized that any of the nucleic acid primers, probes, or oligonucleotides described herein may contain one or more nucleotide analogs and / or one or more synthetic or non-natural nucleotides.
[0148] It will also be recognized that the kits described herein may include solid substrates. In some embodiments, one or more nucleic acid primers can be conjugated to a solid support. Examples of solid supports include, but are not limited to, polymers, glass, semiconductors, paper, metals, gels, or hydrogels. Additional examples of solid supports include, but are not limited to, microarrays or microfluidic cards.
[0149] It will also be recognized that any of the kits described herein may include one or more detectable labels. In some embodiments, one or more of the nucleic acid primers may be labeled with one or more detectable labels. Typical detectable labels include, but are not limited to, enzyme labels, fluorescent labels, and chromogenic labels.
[0150] Algorithms for predicting cardiovascular disease (CVD) or estimating survival rates from CVD Any number of algorithms can be used, including, without restriction, statistical algorithms (e.g., linear regression, logistic regression, proportional hazards models, etc.), machine learning algorithms (e.g., random forests, gradient boosting, support vector machines, neural networks (e.g., deep neural networks, extreme learning machines (ELMs)), Bayesian classifiers, hidden Markov models, etc.), deep learning algorithms (e.g., convolutional neural networks, recurrent neural networks, autoencoders, large language models, etc.), time series algorithms (e.g., ARIMA, etc.), Bayesian model algorithms (e.g., Bayesian networks, etc.), and / or financial algorithms (e.g., decision trees, discrete event simulations, financial impacts, etc.). See, for example, McKinney et al., 2011, Appl. Bioinform., 5(2):77-88; Gunther et al., 2012, BMC Genet., 13:37; and Ogutu et al., 2011, BMC Proceedings, 5(Suppl 3):S11. Any type of machine learning algorithm or deep learning neural network algorithm (synchronized or unsynchronized) capable of capturing the linear and / or nonlinear contributions of traits for prediction can be used. In some examples, combinations of algorithms are used (e.g., a combination or set of multiple algorithms that capture the linear and / or nonlinear contributions of traits).
[0151] Furthermore, the algorithm can implement one or more of the following: regression algorithms, instance-based methods (e.g., k-nearest neighbors, learned vector quantization, self-organizing maps, etc.), regularization methods, decision tree learning methods (e.g., classification and regression trees, chi-squared approach, random forest approach, multivariate adaptive approach, gradient boosting machine approach, etc.), Bayesian methods (e.g., naive Bayes, Bayesian belief networks, etc.), kernel methods (e.g., support vector machines, linear discriminant analysis, etc.), clustering methods (e.g., k-means clustering), association rule learning algorithms (e.g., Apriori algorithm), artificial neural network models (e.g., backpropagation, Hopfield network, learned vector quantization, etc.), deep learning algorithms (e.g., Boltzmann machines, convolutional networks, stacked autoencoders, etc.), dimensionality reduction methods (e.g., principal component analysis, partial least squares regression, etc.), ensemble methods (e.g., boosting, bootstrap aggregation, gradient boosting machine approach, etc.), and any other suitable algorithms.
[0152] As a simple example, Random Forest™ is a popular machine learning algorithm created by Breiman & Cutler to generate “classification trees” (see, for example, “stat.berkeley.edu / ~breiman / RandomForests / cc_home.htm” on the World Wide Web). Using standard machine learning and predictive modeling techniques, we wrote the diagnostic classifier algorithm to be implemented in the R and Python programming languages (but can be implemented in many other programming languages), following the guidelines detailed by Breiman & Cutler. The diagnostic classifier algorithm was generated using data from at least two traits (T) and the diagnosis of interest from that population. To determine the output (e.g., diagnosis) for a new individual, we simply determine values for at least two traits (T) and input that information into an algorithm (e.g., the diagnostic classifier algorithm described herein or another algorithm discussed above) that can capture the linear and nonlinear contributions of the traits.
[0153] Before fitting the model, the input data can be conditioned or otherwise preprocessed so that the conditioned data elements (e.g., genome reads related to the locus of interest, functional data, sensor data, other lifestyle data, etc.) are suitable for further processing. Conditioning as described herein may include filtering the data (e.g., sensor data outputs with confidence values below a threshold). Before training the model, preprocessing steps such as dimensionality reduction (e.g., principal component analysis, linear discriminant analysis, autoencoder, uniform manifold approximation and projection, partial least squares regression, etc.) can be performed on the model input.
[0154] The uncertainty of model predictions (e.g., risk scores, diagnoses, etc.) can be calculated. Methods for estimation may include bootstrap, Bayesian, Monte Carlo dropout, ensemble, and sensitivity analysis. The uncertainties of the methods and / or models described herein can be aggregated to provide the overall system uncertainty. The quantification of uncertainty can encompass all aspects of marker measurement. For example, in epigenetic measurement, the quantification of uncertainty may include, but is not limited to, sampling error, reagent quality, instrument usage, and human error.
[0155] The quantification of uncertainty as described herein can be used for quality control. For example, a known reference sample and / or measurement can be compared to a new measurement. The difference between the known and new measurement, along with its uncertainty, can be used to assess the error and / or uncertainty, and how they compare to a defined acceptable threshold. This process helps identify the source of error and / or uncertainty and ensures the reliability of the measurement by minimizing such error and / or uncertainty. For example, it can be used to determine whether a sample requires remeasurement or recollection to meet a defined acceptable threshold for the measurement and / or marker.
[0156] The returned classification regression and / or other outputs of the model may include confidence-related parameters in such classifications. In particular, confidence-related parameters may have scores (e.g., percentiles, other scores) indicating the confidence level of the returned output. Confidence can be estimated by aggregating the uncertainties of measurement and / or modeling. Transforming the output data can also be used to enhance the scalability of the previously described models. For example, SHapley Additive exPlanations (SHAP), Local Interpretable Model-agnostics Explanations, Integrated Gradients, Partial Dependence Plots, Global Surrage Models, etc., can be used to enhance the scalability of the output for the user. One or more of these approaches can be used simultaneously. In a specific example, SHAP can be used to identify the most significant contributing markers to a condition or indication.
[0157] Additionally, or alternatively, dynamic aspects of features derived from the sample (e.g., changes in markers over time, changes in frequency between instances of each feature, other temporal aspects, other frequency-related aspects, etc.) can be used to predict or otherwise forecast the state of health in order to generate personalized intervention plans.
[0158] Samples can be collected once (e.g., at a single point in time) or at multiple points in time (e.g., at random times, at regular times, in relation to triggering events, at other frequencies, etc.).
[0159] As described herein, the inputs may be at least one genotype (e.g., SNP) and / or the methylation status of at least one CpG dinucleotide and / or other data, and the outcome may represent a positive or negative probability for CVD, but severity may also be rated. The trait (T) used to determine the outcome may represent at least one methylation status of CpG dinucleotide or at least one genotype (e.g., of an SNP), but the trait (T) may also correspond to at least one interaction (e.g., between methylation status and genotype (CpG × SNP), between methylation statuses of two different sites (CpG × CpG), or between two different genotypes (SNP × SNP)). It will be recognized that any such interaction can be visualized using a partial dependency plot.
[0160] The inputs can also be data or information (e.g., a dataset) that non-limitingly include clinical diagnoses; demographics (e.g., sex, race); lifestyle; imaging (e.g., cardiac CT scan, cardiac MRI, coronary angiography); features derived from imaging (e.g., FFR, stenosis rate from CT scan); results from electrocardiograms or echocardiograms; results from stress tests; blood tests (e.g., metabolic assays, genetics, epigenetics, proteins, etc.); blood pressure; results from carotid ultrasound; and combinations thereof.
[0161] The dataset may, without limitation, include data derived from one or more of the following: weight (e.g., receiving patient weight values generated from a digital scale), body fat percentage, muscle mass, body water, height or other length measurements (e.g., from a ruler or tape measure), other body mass index (BMI) related parameters, blood chemical and biochemical information, inflammatory markers, fasting blood glucose, high-density lipids, low-density lipids, blood interleukins, c-reactive proteins, blood cell count, electrophysiological signals (e.g., electroencephalogram signals, electromyogram signals, electrocutaneous reaction signals, electrocardiogram signals, etc.), heart rate, body temperature, cardiovascular parameters, continuous glucose monitoring (blood glucose response), respiratory parameters (e.g., respiratory rate, depth / shallowness of breath, etc.), blood oxygenation signals, exercise parameters, and any other suitable physiologically relevant parameters of the patient. Additionally or alternatively, the dataset may include data derived from one or more of the following: electronic health records, health insurance claims, questionnaires, surveys, wearables, public sources (e.g., repositories), and any other direct or indirect data related to the individual. Additionally, or alternatively, datasets may contain raw, inputted, transformed, longitudinal, transverse, or temporal data.
[0162] Figure 1 is a block diagram of an exemplary CVD classification system 100. In some embodiments, the system 100 can perform CVD monitoring and / or prediction. For example, the system 100 can be used to perform one or more of the exemplary processes described herein.
[0163] In the illustrated example, subject 101 provides a subject sample 102. In some embodiments, the subject sample 102 may be a blood sample, saliva sample, mucus sample, urine or stool sample, or any other suitable biological sample derived from subject 101. In some embodiments, a healthcare professional 103 (e.g., a physician, nurse, laboratory technician, or caregiver) may assist subject 101 by obtaining the subject sample 102. In some embodiments, subject 101 may obtain the subject sample 102 from themselves (e.g., by using a portable blood collection device or a home collection kit).
[0164] The nucleic acid isolation module 110 isolates the nucleic acid sample 112 from the target sample 102. In some embodiments, the nucleic acid isolation module 110 may be a manual, semi-automatic, or automated process that performs one or more of the following: cell lysis, removal of contaminating proteins, inactivation of DNAase and / or RNAase, and recovery of DNA and / or RNA. For example, the nucleic acid isolation module 110 may be part of an automated process or analytical device configured to isolate the nucleic acid sample 112 from the target sample 102. In another example, the nucleic acid isolation module 110 may be part of one or more of the exemplary kits described herein, used by a human user, such as a healthcare professional 103.
[0165] The genotyping assay module 120 receives a portion 114a of the nucleic acid sample 112. The genotyping assay module 120 is configured to perform a genotyping assay on the portion 114a of the nucleic acid sample 112 to detect the presence of at least one SNP in order to determine, identify, or otherwise obtain a collection of genotype data 122, the at least one SNP being rs2869675, rs4376434, rs12129789, rs7585056, rs710987, rs4639796, rs1333048, rs127 A first SNP selected from 14414, rs942317, and rs1441433 or from Annex C, and / or a second SNP in linkage disequilibrium (R>0.3) with the first SNP selected from rs2869675, rs4376434, rs12129789, rs7585056, rs710987, rs4639796, rs1333048, rs12714414, rs942317, and rs1441433 or from Annex C. In some embodiments, the genotyping assay module 120 can be a manual, semi-automated, or automated process. For example, the genotyping assay module 120 can be part of an automated process or analysis device configured to perform a genotyping assay on a portion 114a. In another example, the genotyping assay module 120 may be one or more parts of the exemplary kits described in this document, used by a human user such as a healthcare professional 103 or a laboratory technician.
[0166] The methylation assay module 130 receives a portion 114b of the nucleic acid sample 112. The methylation assay module 130 is configured to bisulfite-convert the nucleic acid in the portion 114b of the nucleic acid sample 112 and to perform a methylation evaluation on the portion 114b of the nucleic acid sample 112 in order to determine, identify, or otherwise obtain the methylation data collection 132, wherein the at least one CpG site is cg04988978, cg21161138, cg CpG sites are at least one selected from 12655112, cg03725309, cg12586707, and cg17901584 or from Annex A, and / or CpG sites collinear (R>0.3) with CpG sites selected from cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584 or from Annex A.
[0167] The identification system 140 is configured to receive a collection of genotype data 122 and a collection of methylation data 132 and to identify one or more predetermined traits or characteristics of the subject 101 based on the diagnostic classifier algorithm module 142. The diagnostic classifier algorithm module 142 is configured to account for at least one SNP main effect and / or at least one CpG main effect and / or at least one interaction effect. In some embodiments, the diagnostic classifier algorithm module 142 can perform one or more of the algorithms described herein that may indicate the presence of a disease (e.g., a diagnostic indicator), a tendency to develop a disease (e.g., prediction), the severity of a disease, the selection, individualization, or optimization of one or more interventions (e.g., lifestyle interventions, medical interventions, therapeutic interventions), or the effectiveness of managing or monitoring a disease, severity, or risk. For example, the identification system may be configured to identify genetic and / or environmental characteristics that determine the presence or likelihood of a subject developing a disease (e.g., cardiovascular disease), even if the disease is of polygenic origin. In some implementations, the diagnostic classifier algorithm module 142 can be a machine learning algorithm capable of explaining linear and nonlinear effects.
[0168] The identification system 840 provides an output 150 based on diagnostic and / or prognostic indicators provided by the diagnostic classifier algorithm module 142. In some embodiments, the identification system 140 may include an output module configured to provide the output 150. In some implementations, the output 150 may be an identification of one or more diseases that subject 101 may already have. For example, the output 150 may indicate that a trait indicating the presence of cardiovascular disease was found in subject 101. In some implementations, the output 150 may be an indication of the likelihood that subject 101 may develop a disease within a predetermined time frame (for example, subject 101 may have a 43% chance of developing cardiovascular disease within 3 years, or subject 101 may have a 77% chance of having a heart attack within 2 years). In some implementations, the output 150 may include recommendations for treatment and / or prevention based on the diagnostic and / or prognostic indicators provided by the diagnostic classifier algorithm module 142. For example, in response to the identification or prediction of diabetes or a cardiac condition in subject 101, output 150 may include, in consultation with healthcare professional 103, recommendations to identify possible dietary or lifestyle changes by subject 101 to address or avoid the condition, recommendations to identify potential treatments and / or therapies for subject 101 to consider in consultation with healthcare professional 103, or a combination thereof, and / or any other appropriate information based on the output of the algorithm of the diagnostic classifier algorithm module 142. In some examples, output 150 may be the likelihood of survival for subjects determined to be at risk of or having CVD.
[0169] In the illustrated example, output 150 is provided in various formats. The information provided by output 150 can be formatted into a message 160 provided to subject 101 and / or healthcare professional 103. In some implementations, message 160 can be formatted as a report (e.g., a word processing file, a portable document format file) that is at least temporarily stored on a non-temporary storage medium (e.g., a hard drive, flash memory) and can be retrieved there by subject 101 and / or healthcare professional 103 for review. In some implementations, message 160 can be formatted as an electronic message (e.g., an email, a text message, an instant message) sent to subject 101 and / or healthcare professional 103 for review. In some implementations, message 160 can be a printed report. For example, output 150 can be provided to a printing system configured to generate a hard copy report based on output 150. A subsequent automated or manual processing system may package the report as a letter or other parcel that can be sent for physical delivery to subject 101 and / or medical personnel 103 (for example, system 100 may produce printed papers of the results and send them by mail).
[0170] The treatment device 170 can be configured to receive diagnostic and / or prognostic indicators provided by output 150 and to provide interventions (e.g., lifestyle interventions, medical interventions, therapeutic interventions) based on the diagnostic and / or prognostic indicators. For example, output 150 may indicate that subject 101 has a high probability of suffering cardiac arrest within the next two years, and the treatment device 170 may be a drug (e.g., a tablet or capsule) or an implantable drug delivery system that responds by identifying or receiving configuration settings for appropriate dosages of statins, acetylsalicylic acid (aspirin), anti-inflammatory drugs, anticoagulants, or combinations thereof, and / or any other appropriate therapeutic and / or prophylactic substances. In some embodiments, the treatment device 170 can also be configured to include one or more of the following: nucleic acid isolation module 110, genotyping assay module 120, methylation assay module 130, or identification system 140.
[0171] The storage system 180 is configured to store the output 150. For example, the information contained in the output 150 can be stored temporarily, over a predetermined period of time, or substantially permanently, in a database, a file, or as any other suitable data collection. In some embodiments, the storage system 180 can store the output 150 on a non-temporary storage medium (e.g., a hard drive, flash memory). For example, the output 150 may include some or all of the outputs 150 in a personal health record that the subject 101 can store or carry, such as a genotype data collection 122, a methylation data collection 132, and / or the output 150 in a personal health record. In some embodiments, the storage system 180 can store the output 150 as a physical medium, for example, the storage system 180 may include a printer that can generate a paper report based on the output 150 and / or store the report as a hard copy that can be physically filed for later retrieval.
[0172] The input / output device 182 is a physical device configured to display or otherwise present an output that is perceptible to a human being (e.g., subject 101, healthcare worker 103). For example, the input / output device 182 may be an electronic display device in a medical examination room. The system 100 may process the subject sample 102 and then modify the configuration of pixels on the screen to modify the information displayed by the input / output device 182 based on the output 150 (for example, the screen may be updated to display the identified diagnosis and / or prognosis for subject 101 to a healthcare worker). In another example, the input / output device 182 may be configured to provide audible (e.g., voice output) and / or tactile (e.g., Braille, tactile, vibration) outputs that modify or otherwise convert the output 150 into a physical and / or tangible output (e.g., to convey diagnostic and / or prognostic indicators in a format perceptible to a visually impaired user). In another example, the input / output device 182 may be configured to change, convert, or modify the physical structure or physical properties of a medium based on the output 150.
[0173] The user device 184 (e.g., a computer, smartphone, tablet computer, computerized terminal) is configured to display, emit, or otherwise present one or more outputs that are perceptible to a human user, such as subject 101 and / or healthcare worker 103. For example, the user device 184 may receive output 150 (e.g., as data, as a message 160) and, based on output 150, provide an alert to the user and / or provide an output (e.g., display a report, read the report aloud). In some embodiments, the user device 184 may include one or more of the storage device 180 or the input / output device 182. In some embodiments, the user device 182 may be part of the treatment device 170. In some embodiments, the user device 184 may be configured to include one or more of the nucleic acid isolation module 110, the genotyping assay module 120, the methylation assay module 130, or the identification system 140.
[0174] In some implementations, some or all of system 100 may be reused to provide additional information. For example, system 100 may be used to collect an initial set of health information about subject 101 and / or to identify information that can assist healthcare professionals 103 in initial diagnosis / prognosis. Later, patient 101 may be retested using system 100 to determine, for example, the effectiveness of prescribed medical treatment and / or lifestyle strategies over time. Since the genotype data collection 122 does not change over time for individual people, system 100 may refrain from performing the function of the genotyping assay module 120 again. In such an example, the methylation assay module 130 may be used to generate an updated version of the methylation data collection 132, and the updated methylation data collection 132, along with the previously generated genotype data collection 122, can be provided to the identification system 140 for processing. In some implementations, the target sample 102 can be periodically collected and processed based on existing genotype data collections 122 and updated methylation data collections 132 to produce an updated output 150 that can be used to provide continuous monitoring of one or more conditions identified for the target 101.
[0175] Figure 2 is a flowchart of an exemplary process 200 for cardiovascular disease classification. In some implementations, process 200 may be some or all of the exemplary processes described above. In some implementations, process 200 may be a process performed by some or all of the exemplary system 100 in Figure 1.
[0176] In step 210, the nucleic acid sample is isolated from the target sample. For example, the exemplary nucleic acid isolation module 110 may be configured to isolate and / or substantially purify a nucleic acid composition from the exemplary target sample 102 to produce the exemplary nucleic acid sample 112.
[0177] In step 220, a genotyping assay was performed on a first portion of the nucleic acid sample to detect the presence of at least one SNP in order to obtain genotype data. The at least one SNP was identified as rs2869675, rs4376434, rs12129789, rs7585056, rs710987, rs4639796, rs1333048, rs12714414, rs942317, and rs144143 A first SNP selected from 3 or Annex C, and / or a second SNP in linkage disequilibrium (R>0.3) with a first SNP selected from rs2869675, rs4376434, rs12129789, rs7585056, rs710987, rs4639796, rs1333048, rs12714414, rs942317, and rs1441433 or Annex C. For example, the exemplary genotyping assay module 120 can be used to analyze a portion 114a of an exemplary nucleic acid sample 112 to produce an exemplary genotyping data collection 122.
[0178] In 230, in order to obtain methylation data, a second portion of the nucleic acid sample is bisulfite-converted and methylation is evaluated on the second portion of the nucleic acid sample, wherein the at least one CpG site is selected from cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584 or from Annex A, and / or a CpG site that is collinear (R>0.3) with cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584 or from Annex A. For example, using the exemplary methylation assay module 130, a portion 114b of the nucleic acid sample 112 can be processed to produce an exemplary methylation data collection 132.
[0179] In step 240, genotype data from step 220 and / or methylation data from step 230 are input into the algorithm. For example, exemplary genotype data collection 122 and exemplary methylation data collection 132 are input into exemplary identification system 140 and processed using exemplary diagnostic classifier algorithm module 142.
[0180] At 250, at least one SNP main effect and / or at least one CpG main effect and / or at least one interaction effect is described. For example, the exemplary diagnostic classifier algorithm module 142 can be configured to describe at least one SNP main effect and / or at least one CpG main effect and / or at least one interaction effect. In some implementations, the diagnostic classifier algorithm module 142 can be a machine learning algorithm that can describe linear and nonlinear effects.
[0181] At 260, an output is provided. For example, an exemplary identification system 140 may provide an exemplary output 150.
[0182] In 270, another nucleic acid sample is isolated from another sample derived from the subject. For example, the exemplary nucleic acid isolation module 110 may be configured to isolate and / or substantially purify a nucleic acid composition from another sample to produce another exemplary nucleic acid sample. Since the genotype data collection 122 derived from the subject does not change over time, methylation data 132 can be obtained using the newly prepared nucleic acid sample, which is used together with the existing genotype data collection 122 to provide updated output (for example, to perform testing on subject 101 at some point later). In some implementations, this omitted process may be performed periodically or semi-periodically to provide continuous monitoring of one or more medical conditions identified for subject 101.
[0183] Figure 3 is a block diagram of exemplary computing devices 300, 350 that may be used to implement the systems and methods described herein, either as a client or as one or more servers. Computing device 300 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. Computing device 300 may also represent all or some of various forms of computerized devices, such as embedded digital controllers, media bridges, modems, network routers, network access points, network repeaters, and network interface devices, including mesh network communication interfaces. Computing device 350 is intended to represent various forms of mobile devices, such as personal digital assistants, mobile phones, smartphones, and other similar computing devices. The components, their connections and relationships, and their functions shown herein are intended to be illustrative only and are not intended to limit the implementation of the compositions and methods described herein.
[0184] The computing device 300 includes a processor 302, memory 304, storage device 306, a high-speed interface 308 connected to memory 304 and a high-speed expansion port 310, and a low-speed bus 314 and a low-speed interface 312 connected to storage device 306. Each of the components 302, 304, 306, 308, 310, and 312 can be interconnected using various buses and mounted on a common motherboard or in other manner as appropriate. The processor 302 processes instructions for execution within the computing device 300, including instructions stored in memory 304 or on storage device 306, to display graphic information for a GUI on an external input / output device, such as a display 316 coupled to the high-speed interface 308. In other implementations, multiple processors and / or multiple buses, along with multiple memories and memory types, may be used as appropriate. Also, multiple computing devices 300 may be connected, each device providing a portion of the required operation (e.g., as a server bank, a group of blade servers, or a multiprocessor system).
[0185] Memory 304 stores information within the computing device 300. In one implementation, memory 304 is a computer-readable medium. In one implementation, memory 304 is one or more volatile memory units. In another implementation, memory 304 is one or more non-volatile memory units.
[0186] The storage device 306 can provide large-capacity storage to the computing device 300. In one implementation, the storage device 306 is a computer-readable medium. In various different implementations, the storage device 306 may be an array of devices including a floppy disk device, a hard disk device, an optical disk device, or a tape device, flash memory or other similar solid-state memory device, or a device in a storage area network or other configuration. In one implementation, the computer program product is tangibly embodied on the information carrier. The computer program product contains instructions that, when executed, perform one or more methods, such as the methods described above. The information carrier is a computer-readable medium or a machine-readable medium, such as memory 304, the storage device 306, or memory on the processor 302.
[0187] The high-speed controller 308 manages the bandwidth-intensive operation of the computing device 300, while the low-speed controller 312 manages the lower bandwidth-intensive operation. Such assignments of tasks are merely illustrative. In one implementation, the high-speed controller 308 is coupled to memory 304, a display 316 (e.g., through a graphics processor or accelerator), and a high-speed expansion port 310 that can accept various expansion cards (not shown). In another implementation, the low-speed controller 312 is coupled to the storage device 306 and the low-speed expansion port 317 through a low-speed bus 314. A low-speed expansion port, which may include various communication ports (e.g., Universal Serial Bus (USB), Bluetooth, Bluetooth Low Energy (BLE), Ethernet, Wireless Ethernet (WiFi), High-Definition Multimedia Interface (HDMI), ZIGBEE, visible or infrared transceiver, Infrared Data Association (IrDA), optical fiber, laser, sound waves, ultrasound), may be coupled to one or more input / output devices, such as a keyboard, pointing device, scanner, networking device, such as a gateway, modem, switch, or router, for example, through a network adapter 313.
[0188] Peripheral devices can communicate with the high-speed controller 308 through one or more peripheral interfaces of the low-speed controller 312, including but not limited to USB stacks, Ethernet stacks, WiFi radios, Bluetooth Low Energy (BLE) radios, ZIGBEE radios, HDMI stacks, and Bluetooth radios, as appropriate for the sensor configuration. For example, a sensor that outputs readings over a USB cable can communicate through the USB stack.
[0189] The network adapter 313 can communicate with the network 315. A computer network typically has one or more gateways, modems, routers, media interfaces, media bridges, repeaters, switches, hubs, Domain Name Servers (DNS), and Dynamic Host Configuration Protocol (DHCP) servers that enable communication between devices on the network and devices on other networks (e.g., the Internet). One such gateway may be a network gateway that routes network communication traffic between devices on the network and devices off the network. One common type of network communication traffic routed through a network gateway is a Domain Name Server (DNS) request, which is a request to DNS to resolve a Uniform Resource Locator (URL) or Uniform Resource Indicated (URI) to an associated Internet Protocol (IP) address.
[0190] Network 315 may include one or more networks. The network may provide communication under various modes or protocols, including, among others, Global System for Mobile Communications (GSM) voice calls, Short Message Service (SMS), Enhanced Messaging Service (EMS), or Multimedia Messaging Service (MMS) messaging, Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Personal Digital Cellular (PDC), Wideband Code Division Multiple Access (WCDMA), CDMA2000, General Purpose Packet Radio System (GPRS), or one or more television or cable networks. For example, communication may be conducted through radio frequency transceivers. In addition, short-range communication may be conducted using, for example, Bluetooth, BLE, ZIGBEE, WiFi, IrDA, or other such transceivers.
[0191] In some embodiments, network 315 may have a hub-and-spoke network configuration. A hub-and-spoke network configuration can enable an expandable network that can accommodate components that are added, removed, fail, and replaced. This can, for example, allow for more, fewer, or different devices on network 315. For example, if a device fails or is deprecated by a newer version of the device, network 315 can be configured to allow the network adapter 313 to be updated for the replacement device.
[0192] In some embodiments, network 315 may have a mesh network configuration (e.g., ZIGBEE). A mesh configuration may be in contrast to a traditional star / tree network configuration in which networked devices are directly connected to only a small subset of other network devices (e.g., bridges / switches), and the connections between these devices are hierarchical. A mesh network configuration can enable infrastructure nodes (e.g., bridges, switches, and other infrastructure devices) to be connected to other nodes directly and non-hierarchically. The connections can dynamically self-organize and self-configure to route data. By not relying on a central coordinator, multiple nodes can participate in relaying information. In the event of a failure of one or more nodes or communication links between nodes, the mesh network can self-configure to dynamically redistribute the workload, providing fault tolerance and network robustness.
[0193] The computing device 300 can be implemented in numerous different forms, as shown in the figure. For example, it can be implemented multiple times as a standard server 320 or in a group of such servers. It can also be implemented as part of a rack server system 324. It can also be implemented as part of a network device such as a modem, gateway, router, access point, repeater, mesh node, switch, hub, or security device (e.g., camera server). In addition, it can be implemented in a personal computer such as a laptop computer 322. Alternatively, components derived from the computing device 300 can be combined with other components in a mobile device (not shown), such as device 350. In some embodiments, device 350 can be a mobile phone (e.g., a smartphone), a handheld computer, a tablet computer, a network appliance, a camera, an Extended General-Purpose Packet Radio Services (EGPRS) mobile phone, a media player, a navigation device, an email device, a game console, an interactive or so-called "smart" television, a media streaming device, or any two or more of these data processing devices or other data processing devices. In some implementations, device 350 may be included as part of an automated vehicle (e.g., a car, an emergency vehicle (e.g., a fire truck, an ambulance), or a bus). Each such device may contain one or more computing devices 300, 350, and the entire system may consist of multiple computing devices 300, 350 communicating with each other via a low-speed bus or a wired or wireless network.
[0194] The computing device 350 also includes, among other components, a processor 352, memory 364, input / output devices such as a display 354, a communication interface 366, and a transceiver 368. Device 350 may also be provided with storage devices such as a microdrive or other devices to provide additional storage. Each of components 350, 352, 364, 354, 366, and 368 is interconnected using various buses, and some of the components may be mounted on a common motherboard or in other ways as appropriate.
[0195] The processor 352 can process instructions for execution within the computing device 350, including instructions stored in memory 364. The processor may also include separate analog and digital processors. The processor may provide, for example, coordination with other components of the device 350, such as control of the user interface, applications run by the device 350, and wireless communication by the device 350.
[0196] The processor 352 may communicate with the user through a control interface 358 and a display interface 356 coupled to the display 354. The display 354 may be, for example, a TFT LCD display or an OLED display, or other suitable display technology. The display interface 356 may include appropriate circuitry for driving the display 354 to present graphic information and other information to the user. The control interface 358 may receive commands from the user and translate them for submission to the processor 352. In addition, an external interface 362 may be provided in communication with the processor 352 to enable near-area communication between the device 350 and other devices. The external interface 362 may provide, for example, wired communication (e.g., via a docking procedure) or wireless communication (e.g., via Bluetooth or other such technology).
[0197] Memory 364 stores information within the computing device 350. In one implementation, memory 364 is a computer-readable medium. In one implementation, memory 364 is one or more volatile memory units. In another implementation, memory 364 is one or more non-volatile memory units. An extension memory 374 is also provided and may be connected to the device 350, for example, through an extension interface 372 which may include a SIMM card interface. Such an extension memory 374 may provide extra storage space for the device 350 or may also store applications or other information for the device 350. Specifically, the extension memory 374 may include instructions for performing or supplementing the above processes and may also include secure information. Therefore, for example, the extension memory 374 may be provided as a security module for the device 350 and may be programmed with instructions that enable secure use of the device 350. In addition, secure applications may be provided via a SIMM card, along with additional information, for example, by placing identification information on the SIMM card in a hack-proof manner.
[0198] The memory may include, for example, flash memory and / or MRAM memory, as discussed below. In one implementation, a computer program product is tangibly embodied on an information carrier. The computer program product contains instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable or machine-readable medium, such as memory 364, extended memory 374, or memory on processor 352.
[0199] Device 350 can communicate wirelessly through a communication interface 366, which may include digital signal processing circuitry if necessary. The communication interface 366 can provide communication under various modes or protocols, including, among others, GSM voice calls, Voice Over LTE (VoLTE) calls, SMS, EMS, or MMS messaging, CDMA, TDMA, PDC, WCDMA, CDMA2000, GPRS, WiMAX, LTE, and 5G. Such communication may be conducted, for example, through a radio frequency transceiver 368. In addition, short-range communication may be conducted using, for example, Bluetooth, WiFi, or other such transceivers (not shown) configured to provide the uplink and / or downlink portions of data communication. Furthermore, a GPS receiver module 370 can provide additional radio data to device 350, which may be used as appropriate by the application running on device 350.
[0200] Device 350 may also communicate audibly using an audio codec 360 that can receive spoken information from the user and convert it into usable digital information. The audio codec 360 may also generate audible sound for the user, for example, through the speaker in the handset of device 350. Such sound may include sounds from a voice phone, recorded sounds (e.g., voice messages, music files, etc.), and sounds generated by an application running on device 350.
[0201] The computing device 350 can be implemented in a number of different forms, as shown in the figure. For example, it can be implemented as a mobile phone 380. It can also be implemented as part of a smartphone 382, a personal digital assistant, or other similar mobile device.
[0202] Various implementations of the systems and technologies described herein can be realized in digital electronic circuits, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be dedicated or general-purpose, coupled to receive and send data and instructions from a storage system, at least one input device, and at least one output device.
[0203] These computer programs (also known as programs, software, software applications, or code) include machine instructions for programmable processors and can be implemented in high-level procedural programming languages and / or object-oriented programming languages, as well as in assembly / machine languages. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, apparatus, and / or device (e.g., magnetic disks, optical disks, memory, programmable logic devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including machine-readable mediums that receive machine instructions as machine-readable signals. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0204] To provide user interaction, the systems and technologies described herein can be implemented on a computer having a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor), as well as a keyboard and pointing device (e.g., a mouse or trackball) by which the user can provide input to the computer. Other types of devices can also be used to provide user interaction; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and input from the user can be received in any form, including acoustic, voice, or tactile input.
[0205] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), middleware components (e.g., application servers), or frontend components (e.g., client computers having a graphical user interface or web browser through which a user can interact with an implementation of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., communication networks). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), and the Internet.
[0206] Some communication networks can be configured to hold power and information on the same physical medium. This allows a single cable to provide both data connectivity and power to devices. Examples of such shared media include power over network configurations, where power is provided over a medium primarily or previously used for communication. One specific manifestation of power over network is power over Ethernet (PoE), which carries power along with data over twisted-pair Ethernet cables. Other examples of such shared media include network over power configurations, where communication takes place over a medium primarily or previously used for providing power. One specific manifestation of network over power is power line communication (PLC) (also known as power line carrier, power line digital subscriber line (PDSL), mainline communication, power line telecommunications, or power line networking (PLN), or Ethernet over power (EOP)), where data is held over a conductor that is also used for AC power transmission.
[0207] A computing system can include clients and servers. Clients and servers are generally geographically separated from each other and typically interact through a communication network. The client-server relationship arises from computer programs running on each computer that have a client-server relationship with one another.
[0208] A computing system can include routers, gateways, modems, switches, hubs, bridges, and repeaters. A router is a networking device that forwards data packets between computer networks and performs traffic direction functions. A network switch is a networking device that connects devices on a network to each other by exchanging packets, receiving, processing, and forwarding data to destination devices. A gateway is a networking device that allows data to flow from one separate network to another. Some gateways can differ from routers or switches in that they can communicate using more than one protocol and can operate in one or more of the seven layers of the Open Systems Interconnection Model (OSI). A media bridge is a networking device that converts data between transmission media so that data can be sent from computer to computer. A modem is a type of media bridge typically used to connect a local area network to a wide area network, such as a telecommunications network. A network repeater is a networking device that receives a signal and retransmits it to extend the transmission and allow the signal to cover longer distances or overcome communication failures.
[0209] It will be apparent that this disclosure provides those skilled in the art with the ability to construct a matrix in which the methylation status of one or more CpG dinucleotides and / or one or more genotypes (e.g., SNPs; e.g., of one or more alleles) can be evaluated, typically using a computer, as described herein, to identify interactions and enable the prediction of the presence or incidence of CVDs. Although such analyses are complex, no excessive experimentation is required, as all necessary information is either readily available to those skilled in the art or can be obtained by experiments as described herein.
[0210] Methods for treating, managing, and / or monitoring cardiovascular disease This disclosure provides methods for determining the likelihood of a subject having CVD, methods for monitoring a subject for CVD (e.g., disease progression), methods for determining the severity of CVD (e.g., degree of occlusion), and / or methods for estimating the survival of a subject at risk of or having CVD. As used herein, CVD includes, but is not limited to, CHD, stroke, arrhythmia, cardiac arrest, and congestive heart failure. The methods and compositions described herein provide a better ability to assess a subject's risk for cardiovascular disease or to monitor the presence of cardiovascular disease, which is the first step toward more effective prevention. In addition, the methods and compositions described herein provide the ability to estimate the survival of a subject determined to be at risk of or having CVD, which may enable therapies and / or lifestyle changes that may lengthen or prolong the subject's survival.
[0211] When determining a positive prognosis for cardiac outcomes (e.g., cardiovascular death, myocardial infarction (MI), stroke, death from all causes, or a combination thereof), physicians can use the resulting prognostic information to their advantage to identify the need for intervention in a subject, or to customize or optimize interventions, such as stress tests with ECG response or myocardial perfusion imaging, computed tomography angiography, diagnostic cardiac catheterization, percutaneous coronary angioplasty (e.g., balloon angioplasty with or without stent placement), coronary artery bypass grafting (CABG), enrollment in clinical trials, and administration or monitoring of the effects of agents selected from, but not limited to, nitrates, beta-blockers, ACE inhibitors, antiplatelet agents, and lipid-lowering agents. In addition, physicians can use the resulting prognostic information to their advantage to make recommendations for lifestyle changes, including, but not limited to, dietary improvements, exercise regimens, smoking cessation and / or abstinence from alcohol, and combinations thereof. Using the information provided by the methods described herein, a physician can manage interventions and, for example, monitor individuals to observe improvements.
[0212] Individuals identified as being at higher risk (e.g., 69% PPV for CVD) or having the disease can be promptly followed up for further testing or more aggressive intervention. Conversely, individuals at lower risk can be regularly retested and monitored to ensure continuous prevention due to the dynamic nature of DNA methylation.
[0213] Interventions for cardiovascular disease may depend on the type of cardiovascular disease and the symptoms the individual is experiencing. Interventions for cardiovascular disease can be preventive, therapeutic, or palliative. Treatment for cardiovascular disease may include, for example, lifestyle changes (e.g., diet (e.g., low-fat diet), weight loss, exercise, reduction or cessation of smoking and / or alcohol consumption), treatments (e.g., beta-blockers, statins, calcium channel blockers, ACE inhibitors, vasodilators, alteplases, small molecule modulators, pre / pro / syn / post-biotics), medical interventions (e.g., angioplasty, bypass surgery, implantable devices, endarterectomy), gene therapy, gene editing, base editing, epigenetic therapy, epigenetic silencing, and / or epigenetic editing.
[0214] In accordance with this disclosure, conventional molecular biology, microbiology, biochemistry, and recombinant DNA techniques within the skills of the art may be employed. Such techniques are well described in the literature. The present invention is further described in the following examples, which do not limit the scope of the subject matter methods and compositions described in the claims. [Examples]
[0215] Example 1 - Materials and Method This study features data and / or biomaterials from three sources. The first set of anonymized genome-wide genetic data, genome-wide DNA methylation data, and clinical data are from the Framingham Heart Study (FHS) Offspring Cohort; the second set of anonymized clinical data and DNA are from the Intermountain Healthcare (IM) Biorepository; and the third set is from the Iowa Cohort, as described in more detail below. The procedures and protocols used for the analysis of the FHS data and Iowa Cohort were approved by the University of Iowa Institutional Review Board (IRB# 201503802 and IRB# 201910834), and the procedures and protocols used for the analysis of the IM material were approved by the Intermountain Healthcare Institutional Review Board (IRB# 1024811).
[0216] Example 2 - Framingham Heart Study (FHS) Offspring Cohort Details regarding the collection and preparation of clinical and biological data for the FHS cohort have been previously described (dbGAP study accession: phs000007). Briefly, demographic, risk factors, and clinical information, including coronary heart disease (CHD) status, were derived from the Offspring cohort. If an individual was diagnosed with CHD, CHD was considered present. Conversely, if an individual was not diagnosed with CHD, CHD was considered absent. Sources of clinical data in determining CHD events included subject reports, medical record reviews, and death certificates. The designation and date of CHD onset used in this study were determined by a panel of three researchers on the Framingham Endpoint Review Committee, but could be similarly applied to other CVDs.
[0217] Genome-wide DNA methylation data, profiled using the Illumina Infinium HumanMethylation450 BeadChip array (San Diego, CA, USA), was available from 2,567 subjects who underwent phlebotomy. Standard sample and probe-level quality control was performed as described in previous studies, resulting in the retention of DNA methylation data from 2,560 samples and 403,192 loci (see, for example, Dogan et al., 2018, Genes, 9:641; Pidsley et al., 2013, BMC Genomics, 14:1-10; Triche, 2014, FDb.InfiniumMethylation.hg19: Annotation package for Illumina Infinium DNA methylation probes. Vol. R package version 2.2.0; Davis et al., 2018, Handle Illumina methylation data., Vol. R package version 2.22.0; and Dogan et al., 2018, PLoS One, 13:e0190549). Genome-wide genotyping data obtained using the Affymetrix GeneChip HumanMapping 500K array (Santa Clara, CA, USA) were available for 2,406 of the remaining samples. After performing standard sample and probe-level quality control procedures on the array data in PLINK as previously described, the total number of remaining samples and SNPs was 2,295 and 472,822, respectively (Dogan et al., 2018, Genes, 9:641; Dogan et al., 2018, PLoS One, 13:e0190549; and Purcell et al., 2007, Am. J. Hum. Genet., 81:559-75).Based on the number of individuals diagnosed with CHD (cases) and individuals who were not diagnosed with or did not experience a CHD event within four years of examination (controls), the total number of subjects was 2,111. The demographics and conventional risk factors of these individuals are summarized in Table 1.
[0218] (Table 1) Summary of demographics and conventional CHD risk factors for individuals in the Framingham Heart Study Offspring Cohort TIFF2026514405000002.tif113165HDL: High-density lipoprotein, HbA1c: Hemoglobin A1c, SBP: Systolic blood pressure, DBP: Diastolic blood pressure.
[0219] Example 3 - Intermountain Healthcare Cohort The first independent validation cohort consisted of 252 subjects from the Intermountain Healthcare (IM) Heart Institute INSPIRE registry who underwent coronary angiography. CHD cases were defined as adults over 18 years of age who had no prior history of CHD or myocardial infarction (MI) prior to index coronary angiography but had a clinical diagnosis of CHD (>70% stenosis) on angiography. Control subjects were defined as adults over 18 years of age who had no prior history of CHD or myocardial infarction (MI) prior to index coronary angiography, no clinical diagnosis of CHD (<50% stenosis) on index coronary angiography, and no clinical diagnosis of CHD (>70% stenosis), MI, revascularization, or death due to CHD within 4 years of index coronary angiography.
[0220] The demographics of these individuals are summarized in Table 2.
[0221] (Table 2) Demographic summary for the Intermountain Healthcare validation set TIFF2026514405000003.tif30128
[0222] Genome-wide DNA methylation and gene evaluation for each of these 253 subjects was performed by the University of Minnesota Genome Center using the Illumina Infinium MethylationEpic Beadchip array and the Illumina Infinium Multi-Ethnic Global BeadChip array (San Diego, CA, USA), respectively. These data were then subjected to the same quality control procedures described above for FHS samples.
[0223] Example 4 - Iowa Cohort The second independent validation cohort consisted of 167 participants. Demographics are shown in Table 3. The presence or absence of a clinical diagnosis of CHD was determined by medical records.
[0224] (Table 3) Demographic summary for the Iowa validation set TIFF2026514405000004.tif30128
[0225] Example 5 - Integrated Genetic-Epigenetic Coronary Cardiac Disease Risk Prediction Model One of the objectives of this study was to rewrite array-based methylation loci into clinically implementable digital PCR (dPCR) assays, which have certain limitations in accuracy. Therefore, before exclusively data mining using data from the FHS training set, the methylation variables were reduced to include loci based on delta-beta (Δβ) (absolute difference between cases and controls). The beta values of all methylation loci were converted to M values and then scaled to have a zero mean and unit variance.
[0226] All data mining, feature selection, model development, and model tuning were performed exclusively on the FHS training set. Our data mining approach has been outlined in previous publications (Dogan et al., 2018, Genes, 9:641; Dogan et al., 2018, PLoS One, 13:e0190549). All analyses were performed in Python. Briefly, we implemented an undersampling-based approach to account for high class imbalance and combined it with a set of machine learning algorithms incorporating cross-validation to reveal nonlinear methylation-SNP interactions and high predictive biosignatures in the FHS training set (Han et al., 2011, Data Mining: Concepts and Techniques, Elsevier). As a result, a marker set consisting of six DNA methylation loci and ten SNPs was selected that exhibited the best combined performance in terms of receiver operating characteristic area (AUC), sensitivity, and specificity. This aggregate model, consisting of 16 biomarkers, was subjected to hyperparameter tuning and finalized for testing.
[0227] Example 6 - Survival Analysis and Prognostic Score Using data from FHS, Kaplan-Meier survival curves and Cox proportional hazards were fitted to express CHD as a function of risk group (high vs. low) as predicted by an integrated genetic-epigenetic model. The y-axis represents the probability of not having CHD. 95% confidence intervals (CIs) were calculated for each distribution, and the distributions of the high-risk and low-risk groups were compared using a log-rank test.
[0228] Example 7 - Results The clinical and demographic characteristics of the FHS cohort, IM cohort, and Iowa cohort are outlined in Tables 1, 2, and 3, respectively. All subjects from the FHS cohort had European ancestry, while those from the IM and Iowa cohorts showed non-European ancestry. The most significant difference was in terms of sex. Compared to the FHS and IM cohorts, CHD controls (those not diagnosed with CHD) were younger in the Iowa cohort.
[0229] Example 8 - Integrated Genetic-Epigenetic Coronary Cardiac Disease Risk Prediction Model A CHD predictive model was constructed to identify individuals with CHD using integrated genome-wide SNP and methylation data derived from a training set. All subjects had genetic (SNP) and epigenetic (DNA methylation) molecular data. All data mining, variable selection, and model development work was performed on the FHS training set. Data from the FHS trial set, as well as independent external validation sets from IM and Iowa, were used to validate the performance of the final model developed using the FHS training set. Machine learning (a subset of artificial intelligence) procedures were used with data from the FHS training set to develop a model for CHD detection. The final model was constructed using data from 10 SNPs and 6 DNA methylation loci for a total of 16 biomarkers. The performance of the model was then tuned and independently examined on the FHS trial set, as well as independent external validation sets from IM and Iowa, to better understand the generalizability of this biomarker panel at the time of final determination.
[0230] This final aggregate model consisted of a total of 16 biomarkers, 6 of which were DNA methylation biomarkers and the remaining 10 were SNPs. The 6 methylation loci were cg04988978 (5' promoter region of MPO), cg21161138 (AHRR gene entity), cg12655112 (EHD4 gene entity), cg03725309 (SARS1 gene entity), cg12586707 (CXCL1 3' intergenic region), and cg17901584 (DHCR24 gene entity). The 10 SNPs were rs2869675 (PREX1 gene entity) and rs4376434 (near LINC00972). These are the intergenetic regions of ( ), rs12129789 (the gene itself of KCND3), rs7585056 (the intergenetic region near TMEM18), rs710987 (the gene itself of LINC010019), rs4639796 (the gene itself of ZBTB41), rs1333048 (the 3' intergenetic region of CDKN2B), rs12714414 (the intergenetic region near TMEM18), rs942317 (the gene itself of KTN1-AS1), and rs1441433 (the gene itself of PPP3CA).
[0231] Table 4 shows the overall and sex-specific performance of this model for detecting CHD across all three cohorts. As expected, PrecisionCHD performed best on the FHS training dataset used to develop the model. More importantly, PrecisionCHD demonstrated robust generalizability, with sensitivity of over 75% across all cohorts, and the best validation sensitivity and specificity were 88% and 77% in the external Iowa validation cohort. Across the three sets not used to train the model (i.e., FHS trials, IM, and Iowa), the model performed overall with mean AUC, sensitivity, and specificity of 81%, 80%, and 75%, respectively. An 80% sensitivity (true positive rate) means that out of 100 individuals with CHD, 80 are accurately identified by PrecisionCHD. Similarly, a 75% specificity (true negative rate) means that out of 100 individuals without CHD, 75 are accurately identified. Similarly, the mean sensitivity and specificity for men were 81% and 73%, respectively. For women, the mean sensitivity and specificity were 76% and 75%, respectively. Overall, the model performed similarly for both men and women, and across the cohort, demonstrating minimal or no gender bias and robust generalizability.
[0232] (Table 4) Performance of PrecisionCHD® in the Framingham Heart Study cohort, Intermountain Healthcare cohort, and Iowa cohort. TIFF2026514405000005.tif65128AUC: Area under the receiver operating characteristic curve FHS: Framingham Heart Study Cohort IM: Intermountain Healthcare Cohort Iowa: Iowa Cohort Sensitivity: true positive rate; specificity: true negative rate
[0233] Example 9 - Additional Data Annex A lists CpGs whose methylation is associated with CVD. Annex B lists genes whose methylation is associated with CVD. Annex C lists SNPs associated with CVD. The values provided in Annexes A, B, and C are the mean 10x cross-validation scores, AUC ROC (area under the receiver operating characteristic curve), sensitivity, and specificity calculated by logistic regression. Sensitivity is the true positive rate, and specificity is the true negative rate.
[0234] Example 10 - PrecisionCHD PrecisionCHD is a quantitative test that supports the early detection of coronary heart disease (CHD). See Table 5. This non-invasive test assesses 10 single nucleotide polymorphisms (SNPs) and 6 DNA methylation markers in genomic DNA isolated from human peripheral whole blood. Coronary heart disease is caused by hereditary (genetic) factors as well as acquired, potentially modifiable lifestyle and environmental (epigenetic) factors. The PrecisionCHD early detection test measures certain complex genetic and epigenetic relationships associated with CHD and predicts CHD status using machine learning models.
[0235] The PrecisionCHD trial is initially intended for adults aged 35–80 years who will participate to be assessed for coronary heart disease. The results of this trial are intended to be interpreted by healthcare providers in conjunction with a comprehensive medical assessment. This trial is not intended for the independent diagnosis of coronary heart disease, nor is it intended to replace the diagnosis and treatment of coronary heart disease by healthcare professionals.
[0236] (Table 5) Summary of PrecisionCHD TIFF2026514405000006.tif78160
[0237] As described herein, PrecisionCHD assesses a total of 16 biomarkers, including 10 SNP genotypes and 6 DNA methylation biomarkers. The PrecisionCHD test uses a standard Taqman assay for genotype profiling and a proprietary methylation-sensitive digital PCR assay for DNA methylation marker profiling. The biomarkers captured by PrecisionCHD are mapped to several complex pathways related to the biology and pathogenesis of CHD, such as serine metabolism, cholesterol biosynthesis, smoking, and inflammation. Cardio Diagnostics' Actionable Clinical Intelligence® (ACI®) platform (see, for example, U.S. Patent Provisional Application 63 / 488,463, incorporated herein by reference) helps map biomarker information to modifiable risk factors for CHD, guiding personalized interventions.
[0238] Example 11 - PrecisionCHD Workflow Having your equipment tested at PrecisionCHD is easy, fast, and convenient. The process includes the following: 1. Eligibility Criteria: a. 35-80 years old b. Participate in order to be assessed for coronary heart disease. i. Exclusion: Patients who have previously received a bone marrow transplant are ineligible. 2. Sample collection: a. Option 1: A home-use lancet-based sample collection kit, mailed directly to the patient upon receiving a trial order from a clinician. b. Option 2: Blood collection using provider settings c. Option 3: Blood collection at non-donor locations such as mobile clinics, community centers, and bloodletting centers. 3. Samples are processed in a high-complexity CLIA laboratory to profile genotype and methylation biomarkers. 4. Analyze biomarkers, generate a clinical report, and share it with the prescribing clinician. 5. The clinician also receives login access to an Actionable Clinical Intelligence platform (see, for example, U.S. Patent Provisional Application No. 63 / 488,463, incorporated herein by reference), which includes supplemental mapping of molecular information to improveable risk factors related to the patient's condition. 6. Outline the results of the clinician-patient discussion and the prevention / management plan. Figure 4 outlines several potential action items that can be implemented based on a positive CHD signal using the PrecisionCHD test.
[0239] Example 12 - Individualized intervention The methods described herein can assist in the selection and / or customization of appropriate interventions. The effectiveness of interventions such as lifestyle or medication can be assessed by conducting trials before and after the intervention to quantify any changes in condition or risk, if any, and to inform other future interventions. This process can be repeated to identify the optimal intervention for each individual patient. Thus, the methods described herein can assist in the management of interventions for individual patients and, subsequently, in the subsequent monitoring of such patients to assess the effectiveness of such interventions.
[0240] Example 13 - PrecisionCHD-Epi for evaluating mortality risk in people with coronary heart disease PrecisionCHD® is a powerful, integrated genetic-epigenetic diagnostic tool for detecting the presence of coronary heart disease (CHD) in clinical populations or for post-issue health initiatives in life insurance. Its non-genetics version, PrecisionCHD-Epi, is a highly sensitive screening tool used by life insurers to screen individuals for the presence of CHD during the underwriting phase. However, in cases where the presence of CHD is already established for an individual, we conducted experiments to determine whether PrecisionCHD-Epi can further stratify the severity of CHD in relation to mortality risk.
[0241] This is a critical issue for life insurers because stratifying mortality risk in individuals already diagnosed with CHD is one of the highest-risk endeavors for health insurance underwriters. Calling for these assessments to be based on solid medical evidence, Dr. Anthony Milano, in an influential paper published 20 years ago, reviewed the existing literature and recommended using CHD class and left ventricular function as key indicators of mortality risk (Milano, 2000, J. Insur. Med. New York, 32(3):167-85). Since then, other markers of CHD mortality, such as N-terminal pro-brain natriuretic protein (NT-ProBNP), have been added to underwriting equipment to aid in prediction (Bibbins-Domingo et al., 2007, JAMA, 297(2):169-76). Nevertheless, predicting mortality in individuals with CHD remains suboptimal. In 2009, Sijbrands and colleagues used the Nationale-Nederlanden underwriting procedure to assess mortality risk in 62,334 Dutch male insured applicants, including 3,963 individuals with current cardiovascular disease (CVD) (Sijbrands et al., 2009, PloS One, 4(5):e5457). Despite the full availability of medical information and the clear relationship between CVD and mortality risk, Sijbrands concluded that "deaths could not be individually identified by the medical assessments employed."
[0242] PrecisionCHD-Epi can dynamically change this and potentially help underwriters stratify the risk of death from CHD. PrecisionCHD-Epi uses the same six methylation-sensitive digital PCR assays that power PrecisionCHD® to sensitively screen for the molecular signature of CHD. This approach is based on the understanding that not all forms of CHD are alike, and that the treatment and prognosis for patients diagnosed with CHD are all different.
[0243] To demonstrate this point and to show the power of this technique to predict the severity of CHD, the inventors analyzed clinical and epigenetic data from 244 subjects in the Framingham Heart Study (FHS) who were diagnosed with CHD by the FHS Endpoints Committee in eight waves of the FHS Offspring Cohort Study (Cupples et al., 1988, "The Framingham Heart Study, Section 35. An Epidemiological Investigation of Cardiovascular Disease Survival following Cardiovascular Events: 30 Year Follow-up," Lung and Blood Institute). Table 6 describes the characteristics of the subjects. Two-thirds of those diagnosed with CHD were male, highlighting the lack of awareness of CHD in women by current methods. The rate of self-reported smoking was higher in men but lower overall. Values of other parameters, such as blood pressure and lipid levels, were clinically normal.
[0244] (Table 6) Clinical and demographic characteristics TIFF2026514405000007.tif56128
[0245] Table 7 shows the non-digitally converted methylation values for each of the six CpG sites targeted by the MSdPCR set. Males had significantly lower methylation values at cg12655112 and cg12586707, and demethylation at each site was associated with CHD status.
[0246] (Table 7) MSdPCR DNA methylation levels of the PrecisionCHD-Epi site by sex TIFF2026514405000008.tif35128*p<0.05, **p<0.01
[0247] Next, the inventors investigated the relationship between methylation markers and survival. A total of 71 FHS subjects with CHD (48 men and 23 women) were observed to have died during the follow-up period. Using their survival time and standard proportional hazards modeling, the relationship between methylation levels at six sites and survival was analyzed using stepwise regression. A simple proportional hazards model constructed using all six MSdPCR markers predicted survival (p<0.006). The addition of age, sex, or any of the traditional predictors shown in Table 1 (lipids, HbA1c, and BP) did not improve the prediction. Stepwise removal of non-significant markers yielded three marker models: cg04988978, cg21161138, and cg12655112, with logWorth values (-log of p-value) of 2.00, 1.74, and 3.07, respectively, and survival was predicted by the chi-squared p-value with p<0.002.
[0248] Visualizing the effects of markers within a multivariate interaction model of survival can be challenging. An alternative, but less powerful, way to understand the relationship between methylation and survival is to simply plot the relationship between methylation and whether or not subjects died during the follow-up period. Figures 5A, 5B, and 5C show this relationship to mortality for each of the three markers. Overall, individuals in the lower two deciles of cg04988978 methylation were four times more likely to die during the follow-up period than those in the upper two deciles (28 out of 50 vs. 7 out of 50). Similar, but less powerful, effects were observed for cg12655112 and cg21161138 methylation. Conversely, among those who died, no significant relationship was found between cg04988978 and days of survival using bivariate analysis, as shown in Figure 6A. However, for both cg12655112 and cg21161138, a strong and significant relationship was found between survival days and methylation status using bivariate analysis, as shown in Figures 6B and 6C. In summary, these findings suggest a strong but complex relationship between methylation markers and survival in the PrecisionCHD-Epi trial.
[0249] In summary, using either PrecisionCHD or PrecisionCHD-Epi, we can now provide survival estimates for individuals at risk of developing CHD or those diagnosed with CHD.
[0250] Although the subject methods and compositions are described herein in relation to numerous different aspects, it should be understood that the foregoing descriptions of the various aspects are intended to illustrate the subject methods and compositions and not to limit their scope. Other aspects, advantages, and modifications are within the scope of the appended claims.
[0251] Methods and compositions that can be used for, in conjunction with, in preparation thereof, or are products thereof are disclosed herein. These and other materials are disclosed herein, and it is understood that combinations, subsets, interactions, groups, etc., of these methods and compositions are disclosed herein. That is, specific references to each of the various individual and collective combinations and permutations of these compositions and methods may not be explicitly disclosed, but each is specifically contemplated and described herein. For example, if a composition or method of a particular subject is disclosed and discussed, and a number of compositions or methods are discussed, each and every combination and permutation of the compositions and methods is specifically contemplated unless specifically indicated to the contrary. Similarly, any subset or combination of these is also specifically contemplated and disclosed.
[0252] TIFF2026514405000009.tif244168TIFF2026514405000010.tif244168TIFF2026514405000011.tif250168TIFF2026514405000012.tif244168TIFF202 6514405000013.tif244168TIFF2026514405000014.tif244168TIFF2026514 405000015.tif244168TIFF2026514405000016.tif244168TIFF20265144050 00017.tif244168TIFF2026514405000018.tif244168TIFF2026514405000019.tif244168TIFF2026514405000020.tif244168TIFF2026514405000021.t if244168TIFF2026514405000022.tif244168TIFF2026514405000023.tif244168TIFF2026514405000024.tif244148TIFF2026514405000025.tif250140
[0253] TIFF2026514405000026.tif245154TIFF2026514405000027.tif245155TIFF2026514405000028.tif245160TIFF2026514405000029.tif245158TIFF2026514405000030.tif245163TIFF2026514405000031.tif220156
[0254] TIFF2026514405000032.tif245164TIFF2026514405000033.tif245162TIFF2026514405000034.tif245162TIFF2026514405000035.tif245162TIFF2026514405000036.tif245162TIFF2026514405000037.tif245162TIFF2026514405000038.tif245162TIFF2026514405000039.tif245162TIFF2026514405000040.tif245162TIFF2026514405000041.tif245162TIFF2026514405000042.tif245162TIFF2026514405000043.tif245162TIFF2026514405000044.tif245162TIFF2026514405000045.tif245162TIFF2026514405000046.tif245162TIFF2026514405000047.tif180162
Claims
1. A kit for determining the methylation status of at least one CpG dinucleotide and the genotype of at least one single nucleotide polymorphism (SNP), including the following: The first CpG dinucleotide of a GC locus selected from the group consisting of cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584, or A second CpG dinucleotide in linkage disequilibrium with a first CpG dinucleotide of a GC locus selected from the group consisting of cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584. At least one first nucleic acid primer having a length of at least 8 nucleotides, complementary to a bisulfite-converted nucleic acid sequence containing, wherein the linkage disequilibrium has a value of R > 0.3, and the at least one first nucleic acid primer detects methylated CpG dinucleotides or unmethylated CpG dinucleotides. Furthermore A first SNP selected from the group consisting of rs2869675, rs4376434, rs12129789, rs7585056, rs710987, rs4639796, rs1333048, rs12714414, rs942317, and rs1441433, or A first SNP selected from the group consisting of rs2869675, rs4376434, rs12129789, rs7585056, rs710987, rs4639796, rs1333048, rs12714414, rs942317, and rs1441433, and a second SNP in linkage disequilibrium. A second nucleic acid primer having a length of at least 8 nucleotides, complementary to the DNA sequence or the bisulfite-converted DNA sequence, wherein the linkage disequilibrium has a value of R > 0.
3.
2. The kit according to claim 1, wherein the at least one first nucleic acid primer detects unmethylated CpG dinucleotides.
3. The kit according to claim 1, wherein the at least one first nucleic acid primer detects a methylated CpG dinucleotide.
4. The kit according to any one of claims 1 to 3, wherein the at least one first nucleic acid primer comprises one or more nucleotide analogs.
5. The kit according to any one of claims 1 to 4, wherein the at least one first nucleic acid primer comprises one or more synthetic or non-natural nucleotides.
6. The kit according to any one of claims 1 to 5, further comprising a solid substrate bound to at least one first nucleic acid primer.
7. The kit according to claim 6, wherein the substrate is a polymer, glass, semiconductor, paper, metal, gel, or hydrogel.
8. The kit according to claim 6, wherein the solid substrate is a microarray or a microfluidic card.
9. A kit according to any one of claims 1 to 8, further comprising a detectable label.
10. A third nucleic acid primer having a length of at least 8 nucleotides, which is complementary to the nucleic acid sequence upstream of the CpG dinucleotide. A kit according to any one of claims 1 to 9, further comprising:
11. A third nucleic acid primer having a length of at least 8 nucleotides, which is complementary to the nucleic acid sequence downstream of the CpG dinucleotide. A kit according to any one of claims 1 to 9, further comprising:
12. A method for determining the presence of a biomarker in a biological sample derived from a subject, wherein the biomarker is related to detecting CVD, determining the severity of CVD, estimating survival from CVD, identifying, customizing, and / or optimizing interventions for CVD, managing CVD, and / or monitoring CVD, and the method is (a) A step of providing a first portion of the biological sample and a second portion of the biological sample, wherein at least nucleic acids derived from the first portion are bisulfite-converted; (b) The first portion of the biological sample The first CpG dinucleotide of a GC locus selected from the group consisting of cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584, or A second CpG dinucleotide in linkage disequilibrium with a first CpG dinucleotide of a GC locus selected from the group consisting of cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584. A step of contacting a sequence containing with a first oligonucleotide primer at least 8 nucleotides long that is complementary to the sequence, wherein the chain disequilibrium has a value of R > 0.3, and the first nucleic acid primer detects methylated CpG dinucleotides or unmethylated CpG dinucleotides; and (c) The second portion of the biological sample A first SNP selected from the group consisting of rs2869675, rs4376434, rs12129789, rs7585056, rs710987, rs4639796, rs1333048, rs12714414, rs942317, and rs1441433, or A first SNP selected from the group consisting of rs2869675, rs4376434, rs12129789, rs7585056, rs710987, rs4639796, rs1333048, rs12714414, rs942317, and rs1441433, and a second SNP in linkage disequilibrium. A step of contacting a DNA sequence or a bisulfite-converted DNA sequence with a nucleic acid primer of at least 8 nucleotides length that is complementary to the DNA sequence, wherein the linkage disequilibrium has a value of R > 0.
3. Includes, The percentage of methylation of a CpG dinucleotide at a GC locus selected from the group consisting of cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584, and the nucleotide identity of a first SNP selected from the group consisting of rs2869675, rs4376434, rs12129789, rs7585056, rs710987, rs4639796, rs1333048, rs12714414, rs942317, and rs1441433, or a second SNP in linkage disequilibrium with the first SNP, are biomarkers associated with detecting CVD or estimating survival from CVD. The aforementioned method.
13. The method according to claim 12, wherein the biological sample is selected from the group consisting of blood and saliva.
14. The method according to claim 12, wherein the at least one first nucleic acid primer detects a non-methylated CpG dinucleotide.
15. The method according to claim 12, wherein the at least one first nucleic acid primer detects a methylated CpG dinucleotide.
16. The method according to claim 12, wherein the at least one first nucleic acid primer comprises one or more nucleotide analogs.
17. The method according to claim 12, wherein the at least one first nucleic acid primer comprises one or more synthetic nucleotides or non-natural nucleotides.
18. The method according to claim 12, wherein the window for the incidence rate is 3 years.
19. A method for determining the presence of a biomarker in a biological sample derived from a subject, wherein the biomarker is related to detecting CVD, determining the severity of CVD, estimating survival from CVD, identifying, customizing, and / or optimizing interventions for CVD, managing CVD, and / or monitoring CVD, and the method is (a) A step of obtaining nucleic acid samples from the target sample; (b) A step of performing a genotyping assay on a first portion of a nucleic acid sample in order to detect the presence of at least one SNP in order to obtain genotype data, The at least one SNP is A first SNP selected from rs2869675, rs4376434, rs12129789, rs7585056, rs710987, rs4639796, rs1333048, rs12714414, rs942317, and rs1441433 or from Annex C, and / or A second SNP in linkage disequilibrium (R > 0.3) with a first SNP selected from rs2869675, rs4376434, rs12129789, rs7585056, rs710987, rs4639796, rs1333048, rs12714414, rs942317, and rs1441433 or from Annex C, Process; and / or (c) A step of bisulfite conversion of the nucleic acid in a second portion of the nucleic acid and methylation evaluation of the second portion of the nucleic acid sample in order to detect the methylation status of at least one CpG site in order to obtain methylation data, The at least one CpG site is CpG sites selected from cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584 or from Annex A, and / or CpG sites that are collinear (R > 0.3) with CpGs selected from cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584 or from Annex A, Process; and (d) A step of putting genotype data from step (b) and / or methylation data from step (c) into an algorithm that explains at least one SNP main effect and / or at least one CpG main effect and / or at least one interaction effect, wherein the algorithm is a machine learning algorithm capable of explaining linear and nonlinear effects. The method, including the method described above.
20. The method according to claim 19, wherein the at least one interaction effect is selected from the group consisting of gene-environment interaction (SNP × CpG) effects, gene-gene interaction (SNP × SNP) effects, and environment-environment interaction (CpG × CpG) effects.
21. The aforementioned at least one interaction effect is A CpG site selected from cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584 or from Annex A, CpG sites collinear (R > 0.3) with CpG sites selected from cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584 or from Annex A. and, SNPs selected from rs2869675, rs4376434, rs12129789, rs7585056, rs710987, rs4639796, rs1333048, rs12714414, rs942317, and rs1441433 or from Annex C, SNPs within moderate linkage disequilibrium (R > 0.3) from rs2869675, rs4376434, rs12129789, rs7585056, rs710987, rs4639796, rs1333048, rs12714414, rs942317, and rs1441433 or selected from Annex C. The method according to claim 19, wherein the gene-environment interaction effect (SNP × CpG) between the two is a gene-environment interaction effect (SNP × CpG).
22. The method according to claim 19, wherein the at least one interaction effect is an environment-environment interaction effect (CpG × CpG) between at least two CpG sites selected from cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584 or from Annex A.
23. The method according to claim 22, wherein one or both of the at least two CpG sites are collinear (R > 0.3) with one or both of the at least two CpG sites selected from cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584 or from Annex A.
24. The method according to claim 19, wherein the at least one interaction effect is a gene-gene interaction effect (SNP × SNP) between at least two SNPs selected from rs2869675, rs4376434, rs12129789, rs7585056, rs710987, rs4639796, rs1333048, rs12714414, rs942317, and rs1441433 or from Annex C.
25. The method according to claim 24, wherein one or both of the at least two SNPs are collinear (R > 0.3) with one or both of the at least two SNPs selected from rs2869675, rs4376434, rs12129789, rs7585056, rs710987, rs4639796, rs1333048, rs12714414, rs942317, and rs1441433 or from Annex C.
26. The method according to any one of claims 19 to 25, wherein the biological sample is a saliva sample.
27. A nucleic acid isolation module configured to isolate nucleic acid samples from target samples; A genotyping assay module configured to perform a genotyping assay on a first portion of a nucleic acid sample in order to detect the presence of at least one single nucleotide polymorphism (SNP) in order to obtain genotype data, The at least one SNP is A first SNP selected from rs2869675, rs4376434, rs12129789, rs7585056, rs710987, rs4639796, rs1333048, rs12714414, rs942317, and rs1441433 or from Annex C, and / or A second SNP in linkage disequilibrium (R > 0.3) with a first SNP selected from rs2869675, rs4376434, rs12129789, rs7585056, rs710987, rs4639796, rs1333048, rs12714414, rs942317, and rs1441433 or from Annex C, Genotyping assay module; A methylation assay module configured to perform bisulfite conversion on a second portion of a nucleic acid and to perform methylation evaluation on the second portion of a nucleic acid sample in order to detect the methylation status of at least one CpG site in order to obtain methylation data, The at least one CpG site is At least one CpG site selected from cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584 or from Annex A, and / or CpG sites that are collinear (R > 0.3) with CpG sites selected from cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584 or from Annex A, Methylation assay module; and An identification system configured to explain at least one SNP main effect and / or at least one CpG main effect and / or at least one interaction effect based on the genotype data and / or the methylation data. A system for determining the methylation status of at least one CpG dinucleotide and the genotype of at least one SNP, including the above.
28. The system according to claim 27, wherein the algorithm is a machine learning algorithm capable of explaining linear and nonlinear effects.
29. Output module configured to provide output based on identification by the aforementioned identification system. It further includes, The identification explains at least one SNP main effect and / or at least one CpG main effect and / or at least one interaction effect based on genotype data from step (b) and / or methylation data from step (c). The system according to claim 27 or 28.
30. Based on genotype data and / or methylation data, explain at least one SNP main effect and / or at least one CpG main effect and / or at least one interaction effect. A non-temporary computer-readable medium for storing instructions executable by a processing device for performing operations including, The genotype data is based on a genotyping assay on a first portion of a nucleic acid sample isolated from a target sample to detect the presence of at least one SNP in order to obtain the genotype data, wherein the at least one SNP is A first SNP selected from rs2869675, rs4376434, rs12129789, rs7585056, rs710987, rs4639796, rs1333048, rs12714414, rs942317, and rs1441433 or from Annex C, and / or A second SNP that is in linkage disequilibrium (R > 0.3) with a first SNP selected from rs2869675, rs4376434, rs12129789, rs7585056, rs710987, rs4639796, rs1333048, rs12714414, rs942317, and rs1441433 or from Annex C; and The methylation data is based on a methylation assay for bisulfite-converted nucleic acids in a second portion of the nucleic acid sample to detect the methylation status of at least one CpG site in order to obtain methylation data, wherein the at least one CpG site is At least one CpG site selected from cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584 or from Annex A, and / or CpG sites that are collinear (R > 0.3) with CpG sites selected from cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584 or from Annex A, The aforementioned non-temporary computer-readable medium.
31. The non-temporary computer-readable medium according to claim 30, further comprising providing an output based on the above description.
32. The non-temporary computer-readable medium according to claim 31, wherein the output includes one or more of the following: saving a report based on the description to another non-temporary computer-readable medium; modifying a display based on the description; activating an audible alarm based on the description; activating a tactile or vibration alarm based on the description; activating printing of a report based on the description; or activating delivery of treatment based on the description.
33. A kit for determining the methylation status of at least one CpG dinucleotide, including the following: The first CpG dinucleotide of a GC locus selected from the group consisting of cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584, or A second CpG dinucleotide in linkage disequilibrium with a first CpG dinucleotide of a GC locus selected from the group consisting of cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584. A first nucleic acid primer having at least 8 nucleotides in length, which is complementary to a bisulfite-converted nucleic acid sequence containing The linkage disequilibrium has a value of R > 0.3, and the at least one first nucleic acid primer detects a methylated CpG dinucleotide or an unmethylated CpG dinucleotide.
34. A method for determining the presence of a biomarker in a biological sample derived from a subject, wherein the biomarker is related to detecting CVD, determining the severity of CVD, estimating survival from CVD, identifying, customizing, and / or optimizing interventions for CVD, managing CVD, and / or monitoring CVD, and the method is (a) A step of providing a biological sample derived from an object that is at risk of CVD or has CVD, wherein at least a portion of the nucleic acid derived from the biological sample is bisulfite-converted; and (b) The bisulfite-converted nucleic acid The first CpG dinucleotide of a GC locus selected from the group consisting of cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584, or A second CpG dinucleotide in linkage disequilibrium with a first CpG dinucleotide of a GC locus selected from the group consisting of cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584. A step of contacting a sequence containing with a first oligonucleotide primer at least 8 nucleotides long that is complementary to the sequence, wherein the chain disequilibrium has a value of R > 0.3, and the first nucleic acid primer detects methylated CpG dinucleotides or unmethylated CpG dinucleotides. Includes, The percentage of methylation of CpG dinucleotides at GC loci selected from the group consisting of cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584 is relevant to estimating the survival of the subject. The aforementioned method.
35. A method for determining the presence of a biomarker in a biological sample derived from a subject, wherein the biomarker is related to detecting CVD, determining the severity of CVD, estimating survival from CVD, identifying, customizing, and / or optimizing interventions for CVD, managing CVD, and / or monitoring CVD, and the method is (a) A step of isolating nucleic acid samples from the target sample; (b) A step of bisulfite conversion of at least a portion of the nucleic acid and methylation evaluation of the bisulfite-converted nucleic acid in order to determine the methylation status of at least one CpG site in order to obtain methylation data, The at least one CpG site is At least one CpG site selected from cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584 or from Annex A, and / or CpG sites that are collinear (R > 0.3) with CpG sites selected from cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584 or from Annex A, Process; and (c) A step of putting the methylation data from step (b) into an algorithm that explains at least one CpG main effect, wherein the algorithm is a machine learning algorithm that can explain linear and nonlinear effects. The method, including the method described above.
36. A nucleic acid isolation module configured to isolate nucleic acid samples from target samples; A methylation assay module configured to obtain methylation data by bisulfite-converting at least a portion of the nucleic acid and performing a methylation evaluation on the bisulfite-converted nucleic acid, wherein the methylation status of at least one CpG site is determined, The at least one CpG site is At least one CpG site selected from cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584 or from Annex A, and / or CpG sites that are collinear (R > 0.3) with CpG sites selected from cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584 or from Annex A, Methylation assay module; and An identification system configured to explain at least one CpG main effect based on the methylation data. A system for determining the methylation status of at least one CpG dinucleotide, including [specific component].
37. To explain at least one CpG main effect based on methylation data. A non-temporary computer-readable medium for storing instructions executable by a processing device for performing operations including, The methylation data is based on a methylation assay for bisulfite-converted nucleic acids in at least a portion of a nucleic acid sample to detect the methylation status of at least one CpG site in order to obtain methylation data. The at least one CpG site is At least one CpG site selected from cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584 or from Annex A, and / or CpG sites that are collinear (R > 0.3) with CpG sites selected from cg04988978, cg21161138, cg12655112, cg03725309, cg12586707, and cg17901584 or from Annex A, The aforementioned non-temporary computer-readable medium.