Method and system for assessing risk of ischemic heart disease
By integrating GWAS results from multiple ethnic groups, the method and system accurately predict ischemic heart disease risk, overcoming limitations of ethnic-specific genetic analyses and enhancing predictive accuracy.
Patent Information
- Application Number
- JP2021556179
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-11-15
- Filing Date
- 2020-11-13
- Publication Date
- 2025-08-13
- Estimated Expiration
- 2040-11-13
AI Technical Summary
Existing methods for predicting the risk of ischemic heart disease are limited by ethnic-specific genetic analyses, making it difficult to accurately assess risk across diverse populations, particularly when applying results from European populations to non-Western populations like East Asians.
Conduct a genome-wide association study (GWAS) on a large cohort of Japanese individuals and integrate the results with European populations through meta-analysis to identify disease susceptibility loci, creating a genomic risk score (GRS) that combines genetic data from multiple ethnic groups to enhance predictive accuracy.
The cross-ethnic GRS provides superior predictive performance compared to ethnic-specific GRS, enabling accurate risk assessment of ischemic heart disease across different populations.
Smart Images

Figure 0007722640000006 
Figure 0007722640000007 
Figure 0007722640000008
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method and system for assessing the risk of ischemic heart disease using gene sequence analysis. [Background technology]
[0002] Ischemic heart disease is a general term for heart diseases that cause ischemic conditions in the heart due to stenosis or blockage of the coronary arteries, and includes myocardial infarction, angina pectoris, etc. Because ischemic heart disease is one of the leading causes of death worldwide, there is a need to predict risk early and establish effective prevention and treatment methods.
[0003] Ischemic heart disease is a highly heritable disease, and many studies have been conducted to date, with numerous reports on the relationship between genetic mutations and ischemic heart disease. For example, Patent Document 1 (JP 2011-167124 A) discloses a method for testing myocardial infarction based on a single nucleotide polymorphism on human chromosome 5p15.3. Recently, it has become clear that the onset of the disease can be predicted with high accuracy by using a "genetic risk score (GRS)" calculated from genetic information.
[0004] Understanding disease pathogenesis, predicting risk of disease onset, and developing therapeutic drugs based on genetic information are expected to play an important role in advancing the medical science and treatment of ischemic heart disease. However, genetic analyses of ischemic heart disease to date have focused on specific ethnic groups, making it difficult to predict risk universally across races. In other words, it is known that the distribution of genetic mutations varies widely among ethnic groups. For example, it was unclear whether research results from European populations could be applied to non-Western populations, such as East Asian populations, including the Japanese. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Japanese Patent Application Laid-Open No. 2011-167124 Summary of the Invention [Problem to be solved by the invention]
[0006] An object of the present invention is to provide a method and system that can determine the risk of developing ischemic heart disease with high accuracy. [Means for solving the problem]
[0007] The present inventors have conducted extensive research to solve the above problems. First, the inventors' research group compared the genome sequences of approximately 25,892 patients with ischemic heart disease registered in Biobank Japan with those of 142,336 controls, totaling approximately 170,000 people, and conducted a genome-wide association study (GWAS) to comprehensively detect genetic mutations characteristic of patients with ischemic heart disease. As a result, 48 disease susceptibility loci and 73 genetic variants associated with ischemic heart disease were identified across the genome.
[0008] Next, the results of the GWAS on Japanese subjects (approximately 170,000 people) were integrated with those of European populations (approximately 180,000 people from the CARDioGRAMplusC4D study and approximately 300,000 people from the UK Biobank) through meta-analysis, conducting the world's largest cross-ethnic GWAS on ischemic heart disease, totaling more than 600,000 people. As a result, 175 disease susceptibility loci associated with ischemic heart disease were identified. Of these, 35 regions were newly identified, including the HMGCR gene, which is the target of statins, the most important treatment for ischemic heart disease.
[0009] Next, we used the results of this genetic variant-disease association analysis to create a genomic risk score (GRS), and examined its predictive performance in Japanese genomic cohorts (JPHC, J-MICC, and OACIS). It has long been known that GRSs calculated from GWAS results of specific ethnic groups are inapplicable to other ethnic groups and have poor predictive performance. In this study, we also confirmed that the GRS calculated from GWAS results of European populations had poorer predictive performance in Japanese individuals than the GRS calculated from GWAS results of Japanese individuals. Therefore, we created a GRS using cross-ethnic GWAS results, and surprisingly, the performance of the cross-ethnic GRS was found to be superior to that of GRSs based on Japanese and European population data. Based on the above findings, the present invention has been completed.
[0010] One aspect of the present invention is a method for determining a subject's risk of developing ischemic heart disease, comprising: preparing a data list including effect alleles and effect sizes of genetic variants associated with ischemic heart disease; obtaining genetic information of a subject; A step of calculating a risk score for developing ischemic heart disease from the genetic information of the subject based on information regarding the effect allele and effect size of the ischemic heart disease-related genetic variant in the data list; and determining the risk of developing ischemic heart disease based on the risk score; The method is characterized in that the data list is a data list that integrates the genome analysis results of European and American populations and populations other than European and American (non-Western populations) through meta-analysis. Another aspect of the present invention is a system for determining a subject's risk of developing ischemic heart disease, comprising: means for storing data including a data list including effect alleles and effect sizes of genetic variants associated with ischemic heart disease; means for receiving the subject's genetic information; A means for calculating a risk score for developing ischemic heart disease from the genetic information of a subject based on data on the effect allele and effect size of the ischemic heart disease-related genetic variant in the data list; and a means for presenting a risk of developing ischemic heart disease based on the risk score; The system is characterized in that the data list is a data list that integrates the genome analysis results of European and American populations and populations other than European and American populations through meta-analysis. Another aspect of the present invention is a method for examining ischemic heart disease, comprising the steps of: rs62190384, rs56171536, rs2380472, rs35093463, rs73386640, rs2607903, rs112735431, rs995144 7, rs11552449, rs1481345, rs7420881, rs72822411, rs7604403, rs12469628, rs2111485, rs7566501 , rs12714757, rs11133381, rs13105983, rs7680806, rs13354746, rs9351209, rs2744427, rs4723406 , rs6593297, rs9693598, rs2623168, rs13255004, rs35093463, rs11000448, rs603424, rs73392700, The present invention relates to a method for analyzing one or more single nucleotide polymorphisms selected from rs7115190, rs11230728, rs1307145, rs2607903, rs9556903, rs35224956, rs28522673, rs1879454, rs2076438, rs216158, rs12444314, rs4794213, rs2952286, rs948386, rs12625329, and rs6006426. [Effects of the Invention]
[0011] According to the method and system of the present invention, the risk of developing ischemic heart disease can be accurately determined by calculating a risk score using a data list that integrates the genome analysis results of multiple ethnic groups through meta-analysis. [Brief explanation of the drawings]
[0012] [Figure 1] (a) Overview of Japanese cohorts and GWAS, and (b) overview of cross-ethnic meta-analyses. [Figure 2] Manhattan plot results for the Japanese GWAS (25,892 CAD cases and 142,336 control cases). Two-sided P values were calculated using a logistic regression model. [Figure 3] Manhattan plot results for the Japanese GWAS (25,892 CAD cases and 142,336 control cases). Two-sided P values were calculated using a logistic regression model. [Figure 4] Manhattan plot results of the cross-ethnic meta-analysis (121,234 CAD cases and 527,824 controls) are shown. [Figure 5] Manhattan plot results of the cross-ethnic meta-analysis (121,234 CAD cases and 527,824 controls) are shown. [Figure 6] Schematics for a) K-fold cross-validation, b) performance testing of CADPRS in an independent cohort, and c) downstream analysis of CAD-PRS are shown. [Figure 7] Figure 1 shows the performance of the PRS in the study cohort (1,827 and 9,172 patients), showing a) the distribution of the PRS in case and control samples, and b) the prevalence of CAD based on CAD-PRS deciles. [Figure 8] Pairwise comparison of PRS performance in the study cohorts (1,827 patients, 9,172 patients) is shown. [Figure 9] The performance of the PRS derived from a cross-ethnic meta-analysis is shown. Each point represents the median Nagelkerke pseudo R2 for the CAD case-control status in a Japanese cohort (1,827 cases and 9,172 controls). Error bars represent 95% CI. Each point represents the race included in the discovery population. The size of each point represents the number of cases in the discovery population. [Figure 10]a) Correlations between CAD-PRS and continuous clinical measures are shown as Spearman correlation functions and 95% CI (confidence intervals). b) Associations between CAD-PRS and separate clinical measures are shown as Spearman correlation functions and 95% CI (confidence intervals). Shown as estimated β (logOR) for a 1 standard deviation increase in CAD-PRS and its 95% CI (confidence intervals). Significant correlations or associations are indicated as P<0.05 / 34. [Figure 11] The effect of CAD-PRS on long-term cardiovascular mortality is shown. a) Adjusted curves for mortality from ICD-10 (I00-I99) diseases estimated by Cox proportional hazards model are shown. b) The effect of CAD-PRS on mortality from ICD10 (I00-I99) and its subtypes is shown as the estimated HR and its 95% CI for a 1 standard deviation increase in CAD-PRS estimated by Cox proportional hazards model. DETAILED DESCRIPTION OF THE INVENTION
[0013] <Method of determining risk of developing ischemic heart disease of the present invention> The method for determining a risk of developing ischemic heart disease of the present invention comprises: A step of preparing a data list including effect alleles and effect sizes of genetic variants associated with ischemic heart disease (data list preparation step); A step of obtaining genetic information of the subject (genetic information obtaining step), A step of calculating a risk score for developing ischemic heart disease from the genetic information of the subject based on the data of ischemic heart disease-related genetic mutations in the data list (risk score calculation step); and a step of determining the risk of developing ischemic heart disease based on the risk score (determination step); The method includes:
[0014] In the present invention, ischemic heart disease means a disease in which an ischemic state occurs in the heart due to abnormal blood flow in the coronary arteries, and specific examples include myocardial infarction and angina pectoris. Each step will be explained below.
[0015] <Data list preparation process> The data list is a list of effective alleles (risk alleles) and their effect size (degree of association with disease) for a large number of variants (genetic mutations) associated with ischemic heart disease. Here, the type of genetic mutation is not particularly limited as long as it is a mutation that differs from the wild-type sequence, such as a single nucleotide polymorphism (SNP), a base deletion, or a base insertion. The total number of types of ischemic heart disease-associated variants included in the data list is preferably 1,000 or more, more preferably 5,000 or more, and even more preferably 10,000 or more.
[0016] The ischemic heart disease-associated variant is preferably a variant with a minor allele frequency of 1% or more. Furthermore, the variant associated with ischemic heart disease is preferably a variant associated with ischemic heart disease at a probability of P<0.05.
[0017] In the method of the present invention, the data list is characterized by being a data list that integrates genome analysis results of Western populations and genome analysis results of populations other than Western populations (non-Western populations) through meta-analysis. Here, Westerners include, for example, Caucasians, such as Europeans. Non-Western populations include Asian populations, such as Japanese populations. The genome analysis results are preferably large-scale, for example, results obtained from more than 10,000 subjects. Furthermore, the results are preferably the results of genome-wide association studies (GWAS). The genome analysis results of Western populations and non-Western populations may be genome analysis results that have already been conducted and are publicly available, or may be the results of newly conducted genome analysis. Examples of publicly available genome analysis results include, but are not limited to, Biobank Japan for Japanese people, and CARDioGRAMplusC4D (Coronary ARtery DIsease Genome wide Replication and Meta-analysis (CARDIoGRAM) plus The Coronary Artery Disease (C4D) Genetics) and UK Biobank for European people.
[0018] The method of meta-analysis is not particularly limited, and general meta-analysis methods can be used. A fixed-effects model that does not consider heterogeneity between studies may be used, but a random-effects model that considers heterogeneity between studies is preferably used. Meta-analysis can be performed using analysis software such as METASOFT v.2 (http: / / genetics.cs.ucla.edu / meta).
[0019] It is preferable that the data list used in the method of the present invention includes one or more, five or more, ten or more, twenty or more, or thirty or more of the variants described below. rs62190384, rs56171536, rs2380472, rs35093463, rs73386640, rs2607903, rs112735431, rs995144 7, rs11552449, rs1481345, rs7420881, rs72822411, rs7604403, rs12469628, rs2111485, rs7566501 , rs12714757, rs11133381, rs13105983, rs7680806, rs13354746, rs9351209, rs2744427, rs4723406 , rs6593297, rs9693598, rs2623168, rs13255004, rs35093463, rs11000448, rs603424, rs73392700, rs7115190, rs11230728, rs1307145, rs2607903, rs9556903, rs35224956, rs28522673, rs187945 4, rs2076438, rs216158, rs12444314, rs4794213, rs2952286, rs948386, rs12625329, rs6006426 These variants are discussed below.
[0020] <Genetic information acquisition process> In this step, genetic information of the subject, i.e., genome sequence information, is obtained. The genome sequence information may be any information that provides sequence information corresponding to each variant to be analyzed included in the data list, but whole genome sequence information is preferred.
[0021] Genetic information of a subject can be obtained by conventional gene sequence analysis. For example, genetic information of a subject can be obtained using a next-generation sequencer. Note that, in the method of the present invention, gene sequence analysis is not essential; it is sufficient to obtain data that has already been analyzed.
[0022] In the present invention, the subject of analysis is a human, for example, a human suspected of being at risk of ischemic heart disease. The subject's race is not particularly limited, and examples include Asians, including Japanese, and Caucasians, including Europeans. The sample used for genetic information analysis is not particularly limited as long as it contains chromosomal DNA derived from the subject. Examples include body fluids such as blood, urine, and cerebrospinal fluid, cells from the cervix or oral mucosa, and body hair. While these samples can be used directly, it is preferable to isolate chromosomal DNA from these samples by standard methods and analyze it.
[0023] <Risk score calculation process> In this step, a risk score for developing ischemic heart disease is calculated from the genetic information of the subject based on the risk allele and effect size data of ischemic heart disease-associated variants in the data list. Here, the risk score refers to a score (value) indicating the disease risk calculated from the genetic information, and may refer to, for example, a score calculated for each individual as the weighted sum of multiple genetic mutations associated with the disease. For example, the genetic information of the subject is compared with the variant information in the data list, and each variant, for example, the effect size (disease association) is multiplied by the number of risk alleles in the subject to calculate a value, and the sum of these values is calculated to determine the risk score for the subject.
[0024] The risk score (PRS) is, for example, expressed by the following formula: In this specification, PRS and GRS have the same meaning. PRS=β1x1+β2x2+...β k x k +β n x n Here, β k is a SNP k is the log odds ratio (OR) of the risk allele for coronary artery disease for x k is SNP k is the number of risk alleles (0, 1, or 2) for each SNP, and n is the total number of SNPs analyzed. Alternatively, it can be calculated logarithmically as follows: log(PRS)=x1*logβ1+x2*logβ2+...x k *logβ k +x n *logβ n
[0025] Preferably, the risk score is calculated using a pruning and thresholding method. That is, in the data list, variants highly associated with disease, for example, variants associated with ischemic heart disease with a probability of P<0.05, preferably P<0.01, are selected, and these variants are further classified based on linkage disequilibrium, and index variants with low P values representative of each group are selected. For each index variant, a value is calculated from the effect size of the risk allele and the number of risk alleles retained, and a risk score can be calculated from the value. Here, the index variant can be, for example, a linkage disequilibrium coefficient r 2 <0.5, preferably r 2 Variants can be grouped based on the condition of <0.8 and selected from the group.
[0026] <Judgment process> In this step, the risk of developing ischemic heart disease is determined based on the calculated risk score. That is, this step is a determination step that is based on an objective index, the risk score, and does not depend on the judgment of a doctor. For example, if the risk score exceeds a certain reference value, it can be determined that the risk of ischemic heart disease is high. As the reference value, for example, a cutoff value can be determined by calculating risk scores in advance for a group of healthy individuals and a group of ischemic heart disease patients. In addition, the disease risk can be ranked, for example, on a scale of 5 or 10, according to the magnitude of the risk score. Furthermore, a ranking can be performed in association with the severity and prognosis of ischemic heart disease.
[0027] The method of the present invention can be implemented using a computer. That is, one embodiment of the method of the present invention is a computer-implemented method for determining the risk of developing ischemic heart disease, which can be operated in a computer system including a processor and a memory, and comprises: receiving the subject's genetic information; Processing the received genetic information to determine a risk score for developing ischemic heart disease; To determine and present the risk of developing ischemic heart disease based on a risk score; The present invention relates to a method, including: Here, the processing includes comparing the received genetic information with information regarding the effect allele and effect size of each variant in a data list of ischemic heart disease-related genetic mutations obtained by integrating the genomic analysis results of Western populations and non-Western populations through meta-analysis.
[0028] The method of the present invention may be systematized as a method executed by a computer. That is, one aspect of the present invention provides a computer system for determining a subject's risk of developing ischemic heart disease.
[0029] <Ischemic heart disease risk assessment system> The system for determining a risk of developing ischemic heart disease of the present invention comprises: a data storage means including a data list including effect alleles and effect sizes of genetic variants associated with ischemic heart disease; means for receiving the subject's genetic information; A means for calculating a risk score for developing ischemic heart disease from the genetic information of the subject based on information on the effect allele and effect size of the ischemic heart disease-related genetic variant in the data list; and a means for presenting a risk of developing ischemic heart disease based on the risk score; The data list is characterized in that it is a data list that integrates the genome analysis results of European and American populations and non-European and American populations through meta-analysis.
[0030] In one embodiment, the data storage means containing the data list including the effect alleles and effect sizes of genetic variants associated with ischemic heart disease is stored in a memory area within the computer. In another embodiment, the data storage means may be stored outside the computer, such as an external server, and may be accessed by the computer when in use.
[0031] The subject's genetic information may be received by being entered directly into the computer, or may be received from a user interface coupled to the computer, or the subject's genetic information may be received from a remote device via a wireless communication network.
[0032] The risk score calculation means compares the received genetic information with the data list in the data storage means, and calculates a risk score based on the comparison result. For example, the risk score is calculated by substituting the matching results for the risk allele into a pre-stored formula for calculating the risk score. The presentation means presents a risk assessment result based on the risk score and with reference to pre-stored criteria, where presenting the assessment result includes outputting information to a user interface coupled to the computer system and transmitting information to a remote device via a wireless communication network. These means can be programmed and installed in a computer as software or an application, allowing the computer system to perform functions such as processing the received genetic information (data matching, calculation of risk scores, etc.) and presenting the results.
[0033] <Method for assessing risk of ischemic heart disease using novel SNPs associated with ischemic heart disease> Another embodiment of the method for assessing the risk of ischemic heart disease of the present invention comprises the steps of analyzing one or more SNPs selected from the following, and assessing the risk of ischemic heart disease based on the analysis results: rs62190384, rs56171536, rs2380472, rs35093463, rs73386640, rs2607903, rs112735431, rs995144 7, rs11552449, rs1481345, rs7420881, rs72822411, rs7604403, rs12469628, rs2111485, rs7566501 , rs12714757, rs11133381, rs13105983, rs7680806, rs13354746, rs9351209, rs2744427, rs4723406 , rs6593297, rs9693598, rs2623168, rs13255004, rs35093463, rs11000448, rs603424, rs73392700, rs7115190, rs11230728, rs1307145, rs2607903, rs9556903, rs35224956, rs28522673, rs187945 4, rs2076438, rs216158, rs12444314, rs4794213, rs2952286, rs948386, rs12625329, rs6006426
[0034] The rs number indicates the registration number in the dbSNP database of the National Center for Biotechnology Information (https: / / www.ncbi.nlm.nih.gov / snp / ).
[0035] rs62190384 is a SNP located near the PID1 gene on chromosome 2, and refers to a cytosine (C) / thymine (T) SNP at base 230007146 on chromosome 2 of the reference sequence hg19(GRCh37). If this base is C, the risk of developing ischemic heart disease is high. Furthermore, when genotype is taken into account in the analysis, the risk of developing ischemic heart disease increases in the order CC>TC>TT.
[0036] rs56171536 is an SNP located near the CD109 gene on chromosome 6, and refers to an adenine (A) / cytosine (C) SNP at base 74415868 of chromosome 6 of the reference sequence hg19(GRCh37). If this base is C, the risk of developing ischemic heart disease is high. Furthermore, when genotype is taken into account in the analysis, the risk of developing ischemic heart disease increases in the order CC>CA>AA.
[0037] rs2380472 is a SNP located near the C8orf34 gene on chromosome 8. It is a guanine (G) / adenine (A) SNP at base 69431711 of chromosome 8 in the reference sequence hg19(GRCh37). When this base is G, the risk of developing ischemic heart disease is high. Furthermore, when genotype is taken into account in the analysis, the order of risk of developing ischemic heart disease is GG>GA>AA.
[0038] rs35093463 is a SNP located near the ABCA1 gene on chromosome 9. It refers to a cytosine (C) / adenine (A) SNP at base 107586238 on chromosome 9 of the reference sequence hg19 (GRCh37). If this base is A, the risk of developing ischemic heart disease is high. Furthermore, when genotype is taken into account in the analysis, the risk of developing ischemic heart disease increases in the order AA > CA > CC.
[0039] rs73386640 is a SNP located near the BET1L gene on chromosome 11. It is an adenine (A) / guanine (G) SNP at base 203235 of chromosome 11 in the reference sequence hg19(GRCh37). When this base is G, the risk of developing ischemic heart disease increases. Furthermore, when genotype is taken into account in the analysis, the order of risk of developing ischemic heart disease is GG>GA>AA.
[0040] rs2607903 is a SNP located near the YBX3 gene on chromosome 12, and refers to a cytosine (C) / adenine (A) SNP at base 10876573 on chromosome 12 of the reference sequence hg19(GRCh37). If this base is C, the risk of developing ischemic heart disease is high. Furthermore, when genotype is taken into account in the analysis, the risk of developing ischemic heart disease increases in the order CC>CA>AA.
[0041] rs112735431 is a SNP located near the RNF213 gene on chromosome 17. It is a guanine (G) / adenine (A) SNP at base 78358945 on chromosome 17 of the reference sequence hg19 (GRCh37). If this base is A, the risk of developing ischemic heart disease is high. Furthermore, when genotype is taken into account in the analysis, the order of risk of developing ischemic heart disease is AA > GA > GG.
[0042] rs9951447 is a SNP located near the CTAGE1,LOC101927571 gene on chromosome 18. It is a thymine (T) / cytosine (C) SNP at base 20009691 on chromosome 18 of the reference sequence hg19(GRCh37). If this base is C, the risk of developing ischemic heart disease is high. Furthermore, when analyzing taking genotype into account, the order of increasing risk of developing ischemic heart disease is CC>TC>TT.
[0043] rs11552449 is a SNP located near the DCLRE1B gene on chromosome 1, and refers to a cytosine (C) thymine / (T) SNP at base 114448389 of chromosome 1 in the reference sequence hg19(GRCh37). If this base is C, the risk of developing ischemic heart disease is high. Furthermore, when genotype is taken into account in the analysis, the risk of developing ischemic heart disease increases in the order TT>TC>CC.
[0044] rs1481345 is a SNP located near the MIR548F3 gene on chromosome 1, and refers to a thymine (T) / cytosine (C) SNP at base 218827855 on chromosome 1 of the reference sequence hg19(GRCh37). If this base is C, the risk of developing ischemic heart disease is high. Furthermore, when genotype is taken into account in the analysis, the risk of developing ischemic heart disease increases in the order TT>TC>CC.
[0045] rs7420881 is a SNP located near the TMEM17 and EHBP1 genes on chromosome 2. It is a guanine (G) / cytosine (C) SNP at base 62878928 of chromosome 2 in the reference sequence hg19 (GRCh37), and when this base is C, the risk of developing ischemic heart disease is high. Furthermore, when genotype is taken into account in the analysis, the risk of developing ischemic heart disease increases in the order CC>GC>GG.
[0046] rs72822411 is a SNP located near the ACTR2 and SPRED2 genes on chromosome 2. It is a guanine (G) / adenine (A) SNP at base 65499468 of chromosome 2 in the reference sequence hg19 (GRCh37). If this base is A, the risk of developing ischemic heart disease is high. Furthermore, when genotype is taken into account in the analysis, the order of risk of developing ischemic heart disease is AA > GA > GG.
[0047] rs7604403 is a SNP located near the MERTK gene on chromosome 2. It is a guanine (G) / adenine (A) SNP at base 112656652 on chromosome 2 of the reference sequence hg19(GRCh37). If this base is A, the risk of developing ischemic heart disease is high. Furthermore, when genotype is taken into account in the analysis, the order of risk of developing ischemic heart disease is AA>GA>GG.
[0048] rs12469628 is a SNP located near the ARHGAP15 gene on chromosome 2, and refers to a thymine (T) / cytosine (C) SNP at base 144158418 on chromosome 2 of the reference sequence hg19 (GRCh37). If this base is C, the risk of developing ischemic heart disease is high. Furthermore, when genotype is taken into account in the analysis, the risk of developing ischemic heart disease increases in the order TT>TC>CC.
[0049] rs2111485 is a SNP located near the FAP, IFIH1 gene on chromosome 2. It is an adenine (A) / guanine (G) SNP at base 163110536 of chromosome 2 in the reference sequence hg19(GRCh37), and when this base is G, the risk of developing ischemic heart disease is high. Furthermore, when genotype is taken into account in the analysis, the order of risk of developing ischemic heart disease is GG>AG>AA.
[0050] rs7566501 is a SNP located near the PID1 gene on chromosome 2. It is a cytosine (C) / thymine (T) SNP at base 230009317 on chromosome 2 of the reference sequence hg19(GRCh37). If this base is T, the risk of developing ischemic heart disease is high. Furthermore, when genotype is taken into account in the analysis, the risk of developing ischemic heart disease increases in the order CC>TC>TT.
[0051] rs12714757 is a SNP located near the MITF gene on chromosome 3, and refers to a thymine (T) / cytosine (C) SNP at base 69820782 on chromosome 3 of the reference sequence hg19(GRCh37). If this base is C, the risk of developing ischemic heart disease is high. Furthermore, when genotype is taken into account in the analysis, the order of risk of developing ischemic heart disease is TT>TC>CC.
[0052] rs11133381 is a SNP located near the CLOCK gene on chromosome 4, and refers to a thymine (T) / cytosine (C) SNP at base 56316979 of chromosome 4 in the reference sequence hg19(GRCh37). If this base is C, the risk of developing ischemic heart disease is high. Furthermore, when genotype is taken into account in the analysis, the order of risk of developing ischemic heart disease is TT>TC>CC.
[0053] rs13105983 is a SNP located near the ADAMTS3 gene on chromosome 4. It is a guanine (G) / adenine (A) SNP at base 73420634 of chromosome 4 in the reference sequence hg19(GRCh37). If this base is A, the risk of developing ischemic heart disease is high. Furthermore, when genotype is taken into account in the analysis, the order of risk of developing ischemic heart disease is GG>AG>AA.
[0054] rs7680806 is a SNP located near the SORBS2 gene on chromosome 4, and refers to a thymine (T) / cytosine (C) SNP at base 186692853 on chromosome 4 of the reference sequence hg19(GRCh37). If this base is C, the risk of developing ischemic heart disease is high. Furthermore, when genotype is taken into account in the analysis, the risk of developing ischemic heart disease increases in the order CC>TC>TT.
[0055] rs13354746 is a SNP located near the ANKRD31 and HMGCR genes on chromosome 5. It is a cytosine (C) / thymine (T) SNP at base 74619132 of chromosome 5 in the reference sequence hg19 (GRCh37). If this base is T, the risk of developing ischemic heart disease is high. Furthermore, when genotype is taken into account in the analysis, the risk of developing ischemic heart disease increases in the order TT>TC>CC.
[0056] rs9351209 is a SNP located near the ANKRD6 gene on chromosome 6. It is an adenine (A) / guanine (G) SNP at base 90314917 of chromosome 6 in the reference sequence hg19(GRCh37), and when this base is G, the risk of developing ischemic heart disease is high. Furthermore, when the genotype is taken into account in the analysis, the order of risk of developing ischemic heart disease is AA>GA>GG.
[0057] rs2744427 is a SNP located near the TAB2 gene on chromosome 6. It is a guanine (G) / thymine (T) SNP at base 149714790 of chromosome 6 of the reference sequence hg19 (GRCh37). If this base is T, the risk of developing ischemic heart disease is high. Furthermore, when genotype is taken into account in the analysis, the order of risk of developing ischemic heart disease is GG>GT>TT.
[0058] rs4723406 is a SNP located near the TBX20 gene on chromosome 7. It is a cytosine (C) / adenine (A) SNP at base 35286471 of chromosome 7 of the reference sequence hg19(GRCh37). If this base is A, the risk of developing ischemic heart disease is high. Furthermore, when genotype is taken into account in the analysis, the risk of developing ischemic heart disease increases in the order CC>CA>AA.
[0059] rs6593297 is an SNP located near the CCT6A gene on chromosome 7, and refers to an adenine (A) / thymine (T) SNP at base 56122058 of chromosome 7 of the reference sequence hg19(GRCh37). If this base is T, the risk of developing ischemic heart disease is high. Furthermore, when the analysis takes genotype into account, the order of risk of developing ischemic heart disease is AA>TA>TT.
[0060] rs9693598 is a SNP located near the DOCK5 gene on chromosome 8. It is a guanine (G) / adenine (A) SNP at base 25064984 of chromosome 8 in the reference sequence hg19(GRCh37). When this base is A, the risk of developing ischemic heart disease is high. Furthermore, when genotype is taken into account in the analysis, the order of risk of developing ischemic heart disease is GG>GA>AA.
[0061] rs2623168 is a SNP located near the CDH17 and GEM genes on chromosome 8. It is an adenine (A) / guanine (G) SNP at base 95260225 of chromosome 8 in the reference sequence hg19 (GRCh37). When this base is G, the risk of developing ischemic heart disease increases. Furthermore, when genotype is taken into account in the analysis, the order of risk of developing ischemic heart disease is GG > GA > AA.
[0062] rs13255004 is a SNP located near the NCALD gene on chromosome 8, and refers to a cytosine (C) / thymine (T) SNP at base 102832405 on chromosome 8 of the reference sequence hg19 (GRCh37). If this base is T, the risk of developing ischemic heart disease is high. Furthermore, when genotype is taken into account in the analysis, the risk of developing ischemic heart disease increases in the order TT>TC>CC.
[0063] rs35093463 is a SNP located near the ABCA1 gene on chromosome 9. It is a cytosine (C) / adenine (A) SNP at base 107586238 on chromosome 9 of the reference sequence hg19 (GRCh37). If this base is A, the risk of developing ischemic heart disease is high. Furthermore, when genotype is taken into account in the analysis, the risk of developing ischemic heart disease increases in the order AA > AC > CC.
[0064] rs11000448 is a SNP located near the OIT3 gene on chromosome 10. It is a thymine (T) / guanine (G) SNP at base 74682633 on chromosome 10 of the reference sequence hg19(GRCh37). When this base is G, the risk of developing ischemic heart disease increases. Furthermore, when genotype is taken into account in the analysis, the order of risk of developing ischemic heart disease is TT>TG>GG.
[0065] rs603424 is a SNP located near the PKD2L1 gene on chromosome 10. It is a guanine (G) / adenine (A) SNP at base 102075479 of chromosome 10 of the reference sequence hg19(GRCh37). When this base is A, the risk of developing ischemic heart disease is high. Furthermore, when the genotype is taken into account in the analysis, the order of risk of developing ischemic heart disease is AA>AG>GG.
[0066] rs73392700 is a SNP located near the SIRT3 gene on chromosome 11, and refers to a guanine (G) / cytosine (C) SNP at base 224845 on chromosome 11 of the reference sequence hg19 (GRCh37). When this base is C, the risk of developing ischemic heart disease is high. Furthermore, when the analysis takes genotype into account, the risk of developing ischemic heart disease increases in the order CC>GC>GG.
[0067] rs7115190 is a SNP located near the WT1 gene on chromosome 11, and refers to a cytosine (C) / thymine (T) SNP at base 32441377 of chromosome 11 in the reference sequence hg19(GRCh37). When this base is T, the risk of developing ischemic heart disease is high. Furthermore, when genotype is taken into account in the analysis, the risk of developing ischemic heart disease increases in the order TT>TC>CC.
[0068] rs11230728 is a SNP located near the LRRC10B gene on chromosome 11. It is an adenine (A) / guanine (G) SNP at base 61277698 of chromosome 11 in the reference sequence hg19(GRCh37), and when this base is G, the risk of developing ischemic heart disease is high. Furthermore, when genotype is taken into account in the analysis, the order of increasing risk of developing ischemic heart disease is AA>AG>GG.
[0069] rs1307145 is a SNP located near the VPS11 gene on chromosome 11. It is a cytosine (C) / guanine (G) SNP at base 118950217 of chromosome 11 in the reference sequence hg19(GRCh37). When this base is G, the risk of developing ischemic heart disease increases. Furthermore, when genotype is taken into account in the analysis, the order of risk of developing ischemic heart disease is GG>GC>CC.
[0070] rs2607903 is a SNP located near the YBX3 gene on chromosome 12. It is a cytosine (C) / adenine (A) SNP at base 10876573 on chromosome 12 of the reference sequence hg19(GRCh37). If this base is A, the risk of developing ischemic heart disease is high. Furthermore, when genotype is taken into account in the analysis, the risk of developing ischemic heart disease increases in the order CC>AC>AA.
[0071] rs9556903 is a SNP located near the FARP1 gene on chromosome 13, and refers to a guanine (G) / adenine (A) SNP at base 98859335 of chromosome 13 of the reference sequence hg19(GRCh37). When this base is A, the risk of developing ischemic heart disease is high. Furthermore, when the analysis takes genotype into account, the order of risk of developing ischemic heart disease is GG>GA>AA.
[0072] rs35224956 is an SNP located near the MARK3 gene on chromosome 14. It is an adenine (A) / guanine (G) SNP at base 103900481 of chromosome 14 of the reference sequence hg19(GRCh37), and when this base is G, the risk of developing ischemic heart disease is high. Furthermore, when the genotype is taken into account in the analysis, the order of risk of developing ischemic heart disease is GG>AG>AA.
[0073] rs28522673 is a SNP located near the LOXL1 gene on chromosome 15. It is a guanine (G) / cytosine (C) SNP at base 74223716 of chromosome 15 in the reference sequence hg19(GRCh37). If this base is C, the risk of developing ischemic heart disease is high. Furthermore, when genotype is taken into account in the analysis, the risk of developing ischemic heart disease increases in the order CC>GC>GG.
[0074] rs1879454 is a SNP located near the TLNRD1 and CFAP161 genes on chromosome 15. It is a cytosine (C) / adenine (A) SNP at base 81377717 of chromosome 15 of the reference sequence hg19(GRCh37). If this base is A, the risk of developing ischemic heart disease is high. Furthermore, when genotype is taken into account in the analysis, the risk of developing ischemic heart disease increases in the order CC>AC>AA.
[0075] rs2076438 is a SNP located near the IFT140 and TMEM204 genes on chromosome 16. It is a thymine (T) / cytosine (C) SNP at base 1584618 of chromosome 16 in the reference sequence hg19(GRCh37), and when this base is C, the risk of developing ischemic heart disease is high. Furthermore, when genotype is taken into account in the analysis, the risk of developing ischemic heart disease increases in the order CC>TC>TT.
[0076] rs216158 is a SNP located near the MYH11 gene on chromosome 16. It is a cytosine (C) / guanine (G) SNP at base 15917838 of chromosome 16 in the reference sequence hg19(GRCh37), and when this base is G, the risk of developing ischemic heart disease is high. Furthermore, when the genotype is taken into account in the analysis, the risk of developing ischemic heart disease increases in the order CC>GC>GG.
[0077] rs12444314 is a SNP located near the FOXL1, LINC02189 gene on chromosome 16. It is an adenine (A) / guanine (G) SNP at base 86699163 of chromosome 16 of the reference sequence hg19(GRCh37), and when this base is G, the risk of developing ischemic heart disease is high. Furthermore, when genotype is taken into account in the analysis, the order of risk of developing ischemic heart disease is GG>AG>AA.
[0078] rs4794213 is a SNP located near the MBTD1 gene on chromosome 17. It is a cytosine (C) / thymine (T) SNP at base 49308707 of chromosome 17 in the reference sequence hg19(GRCh37). If this base is T, the risk of developing ischemic heart disease is high. Furthermore, when genotype is taken into account in the analysis, the risk of developing ischemic heart disease increases in the order CC>TC>TT.
[0079] rs2952286 is a SNP located near the PRKAR1A gene on chromosome 17. It is a thymine (T) / guanine (G) SNP at base 66469400 of chromosome 17 of the reference sequence hg19(GRCh37), and when this base is G, the risk of developing ischemic heart disease is high. Furthermore, when the genotype is taken into account in the analysis, the order of risk of developing ischemic heart disease is GG>TG>TT.
[0080] rs948386 is a SNP located near the CTAGE1 gene on chromosome 18, and refers to a guanine (G) / cytosine (C) SNP at base 19998810 of chromosome 18 of the reference sequence hg19(GRCh37). When this base is C, the risk of developing ischemic heart disease is high. Furthermore, when the genotype is taken into account in the analysis, the risk of developing ischemic heart disease increases in the order CC>GC>GG.
[0081] rs12625329 is a SNP located near the RGS19 gene on chromosome 20. It is a guanine (G) / adenine (A) SNP at base 62709274 of chromosome 20 of the reference sequence hg19 (GRCh37). When this base is A, the risk of developing ischemic heart disease increases. Furthermore, when genotype is taken into account in the analysis, the order of risk of developing ischemic heart disease is GG>GA>AA.
[0082] rs6006426 is a SNP located near the OSM and CASTOR1 genes on chromosome 22. It is a guanine (G) / adenine (A) SNP at base 30669883 of chromosome 22 of the reference sequence hg19 (GRCh37), and when this base is A, the risk of developing ischemic heart disease is high. Furthermore, when the analysis takes genotype into account, the order of risk of developing ischemic heart disease is AA > AG > GG.
[0083] Furthermore, the SNPs analyzed in the present invention are not limited to those mentioned above, and SNPs in linkage disequilibrium with the above SNPs may also be analyzed. 2 >0.5, preferably r 2 >0.8, more preferably r 2 SNPs that satisfy a relationship of >0.9. 2 is the linkage disequilibrium coefficient. SNPs in linkage disequilibrium can be identified using, for example, the HapMap database (http: / / www.hapmap.org / index.html.ja).
[0084] The above SNPs may be analyzed individually, but it is preferable to analyze a combination of multiple SNPs, for example, 5 or more, 10 or more, 20 or more, or 30 or more. The above SNPs may be analyzed in combination with known ischemic heart disease-related SNPs. By analyzing a combination of multiple SNPs, the accuracy of determining the risk of ischemic heart disease can be improved. Multiple SNPs may be analyzed to determine whether each is a risk allele, and the total number of risk alleles may be used as an index of ischemic heart disease risk (for example, a high risk of ischemic heart disease may be determined if there are one or more, two or more, or three or more risk alleles). Alternatively, a risk score may be calculated from the analysis results of multiple SNPs, and the obtained risk score may be used as an index of ischemic heart disease risk. For example, a high risk of ischemic heart disease can be determined if the risk score exceeds a certain reference value.
[0085] By examining the base types of one or more SNPs as described above and correlating the results with ischemic heart disease using an index such as a risk score based on the results, the risk of ischemic heart disease can be determined. In other words, the method of the present invention can provide data for diagnosing ischemic heart disease.
[0086] The results determined by the method of the present invention are provided to a physician or the like as needed. The physician or the like who receives the results can then diagnose ischemic heart disease after performing necessary tests such as electrocardiogram, blood tests, and ultrasound tests. If the physician or the like diagnoses that there is a high risk of developing ischemic heart disease, appropriate preventive measures such as administering medication can be taken, and if the physician or the like diagnoses that the disease has already developed, treatment such as administering medication or performing surgery can be performed.
[0087] SNP analysis can be performed by a conventional genetic polymorphism analysis method, including, but not limited to, sequence analysis, PCR, hybridization, and the Invader method.
[0088] The present invention also provides test reagents, such as primers and probes, for testing ischemic heart disease. Examples of such probes include probes that contain the SNP site and can determine the type of base at the SNP site based on the presence or absence of hybridization. Specific examples include probes that are 15 bases or longer and have a base sequence containing the polymorphic site of each SNP or a complementary sequence thereto, and probes that are 15 bases or longer and have a sequence containing a base in linkage disequilibrium with the base or a complementary sequence thereto. The length of the probe is preferably 15 to 35 bases, more preferably 20 to 35 bases.
[0089] Examples of primers include primers that can be used in PCR to amplify the SNP site, or primers that can be used for sequence analysis (sequencing) of the SNP site. Specifically, examples include primers that can amplify or sequence a region containing a base at each SNP site, or primers that can amplify or sequence a region containing a base that is in linkage disequilibrium with the base. The length of such primers is preferably 15 to 50 bases, more preferably 15 to 35 bases, and even more preferably 20 to 35 bases. Examples of primers for sequencing SNP sites include primers having a sequence complementary to the 5' region of the base, preferably 30 to 100 bases upstream, or the 3' region of the base, preferably 30 to 100 bases downstream. Examples of primers used to determine polymorphisms based on the presence or absence of PCR amplification include primers having a sequence containing the base and including the base on the 3' side, and primers having a complementary sequence to a sequence containing the base and including the complementary base of the base on the 3' side. [Example]
[0090] The present invention will be specifically described below with reference to examples, but the present invention is not limited to the following embodiments.
[0091] Study sample BioBank Japan (BBJ) (https: / / biobankjp.org) is a hospital-based national biobank project in Japan, containing data from approximately 200,000 patients enrolled between 2003 and 2007. Participants were recruited from 12 medical research institutes across Japan (Osaka Prefectural Adult Disease Center Hospital, Cancer Research Institute Ariake Hospital, Juntendo University, Tokyo Metropolitan Institute of Gerontology, Nippon Medical School, Nihon University School of Medicine, Iwate Medical University, Tokushukai Hospital, Shiga University of Medical Science, Fukujuji Hospital, National Hospital Organization Osaka Medical Center, and Iizuka Hospital).
[0092] The Nagahama Study (http: / / zeroji-cohort.com) is a community-based cohort study conducted in Shiga Prefecture, Japan. Participants aged 30–74 years were recruited from the general population of Nagahama City between 2008 and 2010. The Japan Public Health Center-based Prospective Study (https: / / epi.ncc.go.jp) is a community-based prospective study that has been conducted at 11 public health centers nationwide since 1990. Residents aged 40 to 69 years are enrolled in the JPHC study. In the Japan Multicenter Interdisciplinary Cohort (J-MICC) study ( http: / / www.jmicc.com ), community-dwelling individuals aged 35–69 years were recruited from 2005–2013 in 13 study areas across the country. The Osaka Acute Coronary Syndrome Study Group (OACIS) is a hospital-based registry that enrolled AMI patients between 1998 and 2014 at Osaka University and 24 collaborating hospitals in the Osaka-Hyogo area. Informed consent was obtained from all participants in each study, and the study was approved by the appropriate ethics committee at each institution.
[0093] WGS (Whole Genome Sequencing) Case samples from 1,782 patients with coronary artery disease (CAD) from the BBJ cohort and control samples from 1,007 patients from the BBJ cohort and 2,141 patients from the Nagahama cohort were used for WGS. WGS was performed using the HiSeqX platform aiming for 15× depth, using 2× 150-bp paired-end reads. The sequenced data were processed using Picard v.2.5.0 and aligned to the hs37d5 reference genome in the BBJ cohort and hg19 in the Nagahama cohort using the Burrows-Wheeler algorithm with BWA software v.0.7.5a. Genotypes for samples were called individually at each center using HaplotypeCaller implemented in GATK v.3.6 by the Genome Analysis Toolkit, which is the best method for germline SNPs and indels. GenotypeGVCF (Genomic Variant Call Format) genotype data for each sample was pooled and jointly called.
[0094] The exclusion filters for genotypes were defined as follows: (1) depth <5 and (2) assigned genotype quality <20. These genotypes were set as missing, and variants with call rates <90% were excluded before variant quality score recalibration. After variant quality score recalibration filtering, sample quality control was performed by removing excess heterozygosity (n = 2), excess missing genotypes (n = 12), excess singletons (n = 2), and closely related samples inferred based on state identity (PIHAT > 0.2, n = 34). Principal component analysis was used to restrict the samples to a cluster from mainland Japan (n = 474). After exclusion of the above samples, 1,781 CAD case samples and 2,636 control samples remained.
[0095] quality control We then performed quality control on variants by filtering those with (1) more than 5% missing data, (2) a Hardy-Weinberg equilibrium P value < 1 × 10 ‐6 (3) Data processing center (BBJ and Nagahama cohorts, Fisher's exact test P < 1 × 10 ‐6 (4) variants within low-complexity regions; and (5) variants overlapping with insertions or deletions were excluded. Following these procedures, we performed case-control association studies using genotypes from the WGS data to ensure that the quality of the variants was well controlled. Although certain variants showed significant association with CAD disproportionately to the sample size (outside established loci, P < 5 × 10 -8 , 48 variants), none of the variants showed a significant association with CAD in imputed genome-wide association studies (GWAS) (P>0.05).
[0096] Building a Reference Panel Singletons were removed from the quality-controlled WGS data, and then haplotype phasing was performed using Eagle (v2.4.1) (https: / / data.broadinstitute.org / alkesgroup / Eagle). Phased variant call format files were converted to M3VCF format using minimac3 (v2.0.1) (https: / / genome.sph.umich.edu / wiki / Minimac3) for the reference panel BBJ. CAD was constructed.
[0097] Built reference panel BBJ CAD To evaluate the performance of the 1KG project, genotypes were obtained from 1KG Phase 3 v.5 of the 1000 Genomes Project (1KG) (http: / / www.1000genomes.org) and analyzed under the same pipeline for the reference panel 1KG, separately for all individuals (n = 2,504) and East Asians (n = 504).ALL and 1KG (East Asian) EAS was constructed.
[0098] The BBJ panel was constructed to measure masked genotypes on chromosome 1 in 179,320 Japanese individuals. CAD , 1KG ALL and 1KG EAS As a result, compared to the 1KG panel, the BBJ CAD Imputation quality was significantly better in the panel, and this difference was based on improved imputation quality, especially when considering low-frequency (1% ≤ minor allele frequency (MAF) < 5%) to rare (MAF < 1%) variants.
[0099] To compare the effects of population specificity without the effect of sample size, we used BBJ CAD , 1KG ALL and 1KG EAS An additional reference panel of the same panel size (n=500) was constructed for 1KG ALL and BBJ CAD For the , we created four different panels without overlapping and tested their performance. As a result, we found that the imputation performance was consistent with population specificity (BBJ CAD >1KG EAS >1KG ALL , P<0.001).
[0100] Next, genome-wide genotype imputation was performed including all panel samples. Imputation quality and frequency (R 2 After filtering by ≥ 0.3 and MAF ≥ 0.02%), nearly twice as many testable variants were found in the BBJ. CAD Remaining on the panel (1KG ALL :10,470,888;1KG EAS :8,900,420;BBJ CAD :19,707,525). A substantial increase in the number of variants in restricted variant classes or curated pathogenic variants has been observed. For example, BBJCAD The panel imputed five times more stop-gain variants and four times more "pathogenic" variants than the 1KG reference panel, and identified many variants associated with familial hypercholesterolemia, one of the most important risk factors for the development of CAD.
[0101] Haplotype phasing and imputation of case-control samples The genome-wide association study (GWAS) samples included 25,892 CAD patients and 142,336 controls from the BBJ. GWAS samples were genotyped using the HumanOmniExpressExome v.1.0 / v.1.2 platform (Illumina) or HumanOmniExpress v.1.0 in combination with HumanExomeBeadChip v.1.0 / v.1.1 (Illumina). For variant quality control, (1) call rate <99% and (2) Hardy-Weinberg equilibrium P value <1.0 x 10 ‐6 (3) variants with heterozygosity <5 were excluded. After excluding these variants, prephasing was performed using Eagle. Phased haplotypes were imputed to the reference panel using minimac3. To assess imputation quality, a phased genotype dataset consisting of only the OmniExpress13 array was generated and analyzed. EAS , 1KG ALL , and BBJ CAD The panel was imputed. After imputation, the imputed amounts were compared with the genotypes determined directly by exome array. Correlations were assessed using Pearson's correlation coefficient. All downstream analyses were performed using R. 2 Variants with a score of <0.3 were excluded. Genotyped or imputed variants were annotated using ANNOVAR (constructed July 7, 2017) (http: / / annovar.openbioinformatics.org) or the ClinVar database (https: / / www.ncbi.nlm.nih.gov / clinvar) downloaded on February 4, 2019.
[0102] Phenotype CAD was defined as a composite of stable angina, unstable angina, and myocardial infarction. The definition of this disease relies on physician diagnosis based on applicable guidelines and common medical practice, clinical symptoms, and diagnostic tests. Quantitative trait data were obtained from medical records. Quantitative traits were normalized and adjusted as described below. Prior to normalization, samples from patients younger than 18 years of age were excluded using phenotype-specific criteria shown in Table 1.
[0103] [Table 1]
[0104] Next, we adjusted for the effects of medication as follows: For individuals taking cholesterol-lowering medications, we adjusted for TC and LDL-C values by dividing them by 0.8 and 0.7, respectively, as previously reported (Khera, AV et al. J. Am. Coll. Cardiol. 67, 2578-2589 (2016)., Benn, M. et al. Eur. Heart J. 37, 1384-1394 (2016).). For individuals taking antihypertensive medications, we added 15 mmHg to SBP and 10 mmHg to diastolic blood pressure readings. Next, linear models were constructed for these phenotypes, including sex, age, age, the top 10 principal components, and disease status. For white blood cell count and C-reactive protein (CRP), smoking status was also included in the models. Using these models, phenotypic residuals were calculated for each individual. These residuals were normalized using inverse-rank normalization and used as continuous variables. Samples of individuals under 18 years of age (n = 830), samples with missing clinical information (n = 384), samples with excess heterozygosity (n = 114), samples with excess missing genotypes (n = 30), non-Japanese outlier samples identified by principal component analysis (n = 501), and closely related samples inferred based on identity status (PIHAT > 0.2, n = 9330) were excluded from case-control or quantitative phenotype association analyses.
[0105] GWAS Case-control association analyses were performed using logistic regression analysis implemented in PLINK 2.00 (https: / / www.cog-genomics.org / plink / 1.9). Gender, age, and the top 10 principal components were included in the model as covariates. BBJ CAD Using densely imputed genotypes on the reference panel, GWAS was performed on a case-control dataset from BBJ, including 25,892 CAD cases and 142,336 controls (see Figure 1a), testing 19,707,525 variants with a MAF ≥ 0.02%. The genome-wide significance threshold was P < 5 × 10 for variants with MAF ≥ 1%. ‐8 and P < 3.93 × 10 for variants with MAF < 1% (number of variants with MAF < 1% = 12,710,563). ‐9The ratio was set to (0.05 / 12,710,563). To define loci, we created a set of genetic ranges adding 500 kb on either side for all variants with genome-wide significance and merged the overlapping regions. Within the major histocompatibility complex region (chromosome 6: 25,000,000-35,000,000 bp), we added 1 megabase to the signal on either side. Previously reported loci were further created using the same method based on the curated top variants. The loci were defined as follows: (1) they did not overlap with previously reported loci, and (2) their lead variants were LD variants (r) at the previously reported loci. 2 >0.1), the significance locus was considered novel.
[0106] lead variant r 2 To estimate the number of cases, 1KG was used for the Japanese GWAS. EAS For cross-ethnic meta-analysis, 1kg EAS and 1KG EUR For quantitative traits, linear regression analysis was performed using PLINK 2.00 on the normalized phenotypes described above. As a result, 48 loci reached genome-wide significance, 8 of which were previously unreported, as shown in Table 2, Figure 2, and Figure 3. These lead variants often had larger effect sizes in the Japanese study than in the European study. Furthermore, variants that showed comparable effect sizes between Japanese and European studies often had different allele frequencies, such as rs35093463, a lead variant at a new locus on 9q31. This variant was more frequently observed in Japanese studies (MAF 36%) than in European studies (MAF 7%). This locus harbors ABCA1, a gene important for high-density lipoprotein cholesterol (HDL-C) homeostasis, and its disruption leads to severe deficiencies in serum HDL-C. rs35093463 is associated with rs2066715 (1KG EUR In r 2=0.97, and 1KG EAS In r 2 =0), both of which are non-synonymous substitutions in ABCA1, p.Val825Ile. Next, we investigated the clinical significance of the T allele of rs2066715, a risk allele for CAD. In the Japanese population, this allele was significantly associated with increased HDL-C and total cholesterol (TC) levels (β CAD =0.060, P CAD =4.5×10 -9 ;β HDL-C =0.016, P HDL-C =4.3×10 -4 ;β TC =0.018, P TC =1.1×10 -5 The association between HDL-C levels and rs2066715 CAD risk has also been previously reported in European studies (van der Harst, P. & Verweij, N. Circ. Res. 122, 433-443 (2018)., Lu, X. et al. Nat. Genet. 49, 1722-1730 (2017).) (β CAD =0.048, P CAD =1.8×10 -5 ;βHDL-C=0.037, P HDL-C =3.0×10 -18 ;β TC =0.030, P TC =3.9×10 -12 )(http: / / csg.sph.umich.edu / willer / public). In addition, the Asian-specific lead variant rs112735431 causes a nonsynonymous substitution (p.Arg4810Lys) in RNF213, which revealed a strong signal (MAF=1.0%) on chromosome 17q25, with an OR of 1.0% for CAD. CAD ) odds ratio (OR) = 1.61, 95% confidence interval (CI) = 1.46-1.78, P = 2.3 × 10 -21 ) was detected. rs112735431 is also an established susceptibility variant for moyamoya disease, a rare cerebrovascular disease involving abnormal blood vessel formation or blockage. [Table 2]
[0107] Heritability The phenotypic variance explained by variants was estimated using the previously proposed susceptibility-threshold model (So, H.-C., Gui, AHS, Cherny, SS & Sham, PC Genet. Epidemiol. 35, 310-317 (2011)). Disease heritability was estimated to be 40%, as previously reported. To estimate heritability explained by lead or independent variants determined in the Japanese GWAS, we estimated the prevalence of CAD in the Japanese population to be 5.24% (https: / / vizhub.healthdata.org / gbd-compare / ). The heritability of lead variants determined by cross-ethnic meta-analysis was estimated using beta estimates and variant allele frequencies (AAFs) as reported in a previous European study (Nikpay, M. et al. Nat. Genet. 47, 1121-1130 (2015)). The disease prevalence was estimated to be 8.01%, according to the prevalence in the United States.
[0108] LD score regression analysis The Japanese LD score was calculated from WGS data (n = 4,417) including SNVs with a minor allele count of ≥ 5. Using this Japanese LD score, LD score regression analysis was performed for Japanese GWASs using LDSC software v.1.0.1 (https: / / github.com / bulik / ldsc / ) to estimate the bias of restricting variants to HapMap3 SNPs (http: / / hapmap.ncbi.nlm.nih.gov). For European GWASs, 1KG EUR The LD score calculated from the above was used. The observed liability scale SNP heritabilities were estimated to be 0.077 (sem = 0.007) and 0.127 (sem = 0.011, prevalence = 5.24%) in the Japanese GWAS and 0.075 (sem = 0.004) and 0.171 (sem = 0.010, prevalence = 6.77%) in a previous European GWAS (van der Harst, Circ. Res. 122, 433-443 (2018).).
[0109] Stepwise conditional analysis To identify statistically independent signals at loci, a stepwise conditional analysis of sample values was performed for each genome-wide significant locus defined above. First, the dosage of the lead variant was added as a covariate, and a logistic regression analysis was performed for all variants within the locus. Locus-wide significant associations (P < 1 × 10) were observed. ‐5 If a significant variant was found, the dose of the most significant variant was then introduced as a covariate. This procedure was repeated until no variants showed locus-wide significance.
[0110] Gene-based analysis The SAIGE-GENE method implemented in SAIGE software v.0.36.3 (ref. 64) was applied. Nonsynonymous SNVs, stop-gain / stop-loss variants, in-frame / frameshift indels, and MAF < 5% and R 2 Splice site variants with a p < 0.3 were extracted. Gene-based testing was performed in the same case-control samples with the same covariates as in the single-variant GWAS. 16,582 coding genes were tested with at least one testable variant. Significance was P = 3 × 10 -6 The FDR was set to (0.05 / 16,582), and the FDR was calculated using the Benjamini-Hochberg method.
[0111] Meta-analysis We performed a cross-ethnic meta-analysis combining the Japanese GWAS data with previously published data from two large-scale CAD GWASs, CARDioGRAMplusC4D (C4D; http: / / www.cardiogramplusc4d.org) and UKBB, which primarily include individuals of European descent (Nikpay, M. et al. Nat. Genet. 47, 1121-1130 (2015)., van der Harst, P. & Verweij, N. Circ. Res. 122, 433-443 (2018).) (Figure 1b). Summary statistics for the European CAD-GWAS were obtained from the CARDioGRAM plusC4D Consortium website (http: / / www.cardiogramplusc4d.org / data-downloads / )7 and the online supplemental data (https: / / data.mendeley.com / datasets / 2zdd47c94h / 1) of a previous report (van der Harst, P. & Verweij, N. Circ. Res. 122, 433-443 (2018).). Combining all three datasets resulted in a total of 121,234 CAD cases (BBJ: 25,892, C4D: 60,801, UKBB: 34,541) and 527,824 controls (BBJ: 142,336, C4D: 123,504, UKBB: 261,984).
[0112] Beta and allele frequencies were aligned to the variant allele of hg19, and only SNPs with a MAF of ≥ 1% in the summary statistics for Japanese individuals were merged. The bias of these three GWASs was assessed by LD score regression analysis, and each study was confirmed to be well calibrated (LD score regression intercept for BBJ = 1.035 (sem = 0.009), C4D = 0.880 (sem = 0.007 for genomic control), and UKBB = 1.016 (sem = 0.008)). Considering the racial heterogeneity in each study, the MANTRA algorithm was applied to the analysis. According to previous simulation results (Wang, X. et al. Hum. Mol. Genet. 22, 2303-2311 (2013)), log 10 Bayes factor (BF) > 6 and P Fixed effect <1×10 -5 A variant was considered to be significantly associated with CAD if 10 BF > 6 were excluded. To derive PRS, β and P values were calculated using METASOFT v.2 (http: / / genetics.cs.ucla.edu / meta) by fixed-effect and random-effect meta-analysis.
[0113] Tissue enrichment analysis and gene set enrichment analysis (GSEA) Summary statistics from the cross-ethnic analyses (BBJ, C4D, and UKBB) and the European meta-analyses (C4D and UKBB) were analyzed using DEPICT v.1, release 194 (https: / / data.broadinstitute.org / mpg / depict). The threshold was log , as in a previous study (Malik, R. et al. Nat. Genet. 50, 524-537 (2018). 10 BF > 5 was set. To reduce pathway redundancy, affinity propagation clustering was applied, implemented in the apcluster package v.1.4.8 in R. Using this algorithm, distinct clusters were obtained. Each cluster contained an "exemplar" pathway and member pathways. Exemplar tissues and gene sets were extracted and their significance was tested. 48 exemplar tissues were tested, and their significance was calculated using a P = 1 × 10 for tissue enrichment analysis. -3 (=0.05 / 48). An additional 1,157 sample gene sets were tested, and their significance was calculated using GSEA with P = 4.3 × 10 -5 (=0.05 / 1,157).
[0114] A total of 4,804,024 SNVs were examined in the meta-analysis, and 175 loci reached the genome-wide significance threshold (log 10 BF>6 and fixed effect P value<1×10 -5 ) (Figures 4 and 5). Forty of these loci have not been previously reported (Table 3), including five loci detected in the current Japanese GWAS. Overall, 43 previously unreported loci were found in the Japanese GWAS and cross-ethnic meta-analysis. Lead variants determined by a cross-ethnic meta-analysis at 135 previously reported loci (Nikpay, M. et al. Nat. Genet. 47, 1121-1130 (2015)., van der Harst, P. & Verweij, N. Circ. Res. 122, 433-443 (2018).) explained 11.7% of the heritability of CAD, and lead variants at 40 new loci explained an additional 1.12% of the heritability. Among these, we found a novel association on chromosome 5q13, which contains HMGCR, encoding 3-hydroxy-3-methylglutaryl coenzyme A (HMG-CoA) reductase, the rate-limiting enzyme in endogenous cholesterol synthesis. HMG-CoA reductase is the target enzyme of statins, the most common lipid-lowering drugs. The lead variant rs13354746 is located 13 kb upstream of HMGCR, and the risk allele of this variant is highly associated with elevated serum TC levels in the Japanese population (variant allele frequency (AAF) = 0.27, β TC(s.e.m.TC) =0.052(0.004), P TC =5.5×10 -34 ). Previous studies (Lu, X. et al. Nat. Genet. 49, 1722-1730 (2017)., Lu, X. et al. Hum. Mol. Genet. 25, 4107-4116 (2016).) have shown that rs191835914 (HMGCR p.Tyr311Ser, not in linkage disequilibrium (LD) with rs13354746: 2 =0.004) has been identified as an East Asian-specific nonsynonymous variant associated with serum lipid levels. This variant was significantly associated with serum TC levels in the Japanese population, but not with CAD (AAF=0.013, β TC (sem TC ) = ‐0.116 (0.017), P TC =4.8×10 -12 , β CAD (sem CAD ) = −0.021 (0.430), P CAD =0.623). In silico tissue enrichment analysis (log 10 BF > 5, 19,348 variants, 660 loci), these signals were found in 7 of 48 organs or tissues and 44 of 1,157 gene sets (P < 1.0 × 10 -3 and P < 4.3 × 10 -5 ) were found to be significantly enriched in the cardiac, fibroblast, and 24 pathways, which were only significant in the current cross-ethnic meta-analysis. [Table 3-1] [Table 3-2]
[0115] Credible set analysis We performed a credible set analysis to construct a set of variants likely to contain causal variants in each significant locus identified by the cross-ethnic meta-analysis. For each genome-wide associated locus, the PPA(π) for all variants was calculated according to formula (1):
number
[0116] To assess the contribution of cross-ethnic meta-analyses to the credible set size, we compared the number of variants included in the 99% credible sets from two previous meta-analyses of European studies (C4D and UKBB) with those from two cross-ethnic meta-analyses (C4D and BBJ, and C4D, BBJ, and UKBB). The 99% confidence set sizes for previously established loci from cross-ethnic meta-analyses (C4D, BBJ, and UKBB) were significantly reduced compared to those from European-only analyses (median variants (1st to 3rd quartiles) = 7 (2-17) and 12 (5-27), respectively, P = 1.9 × 10 -6 , Paired Wilcoxon rank sum test). To gain insight into cross-ethnicity specific effects in fine-mapping analyses, we compared results from the C4D and BBJ cross-ethnic meta-analysis with those from the C4D and UKBB European meta-analysis. Despite reduced sample sizes in the cross-ethnic meta-analyses (86,693 cases and 265,840 controls in C4D and BBJ, and 95,342 cases and 385,488 controls in C4D and UKBB), reliable set sizes were comparable (P = 0.98). We found three loci, including the TARID locus on chromosome 6, where single variants reduce reliable set size, particularly in cross-ethnic analyses. In this study, the lead variant rs2327429 had a posterior probability of 1 in cross-ethnic meta-analyses and is a powerful expression quantitative trait locus for TCF21, an important transcription factor that plays multiple roles in atherosclerotic lesions. CAD risk alleles at rs2327429 (T allele) reduced TCF21 expression in relevant tissues (tibial, coronary, and aortic arteries and cultured fibroblasts), consistent with the finding that reduced TCF21 expression in these tissues is atherogenic (Iyer, D. et al. PLoS Genet. 14, e1007681 (2018). Wirka, RC et al. Nat. Med. 25, 1280-1289 (2019).).
[0117] We then extended this analysis to all loci detected in the current cross-ethnic meta-analysis (175 loci) and identified 27 lead variants with a posterior probability of association (PPA) >80%, including four coding variants: rs11556924 (ZC3HC1 p.Arg363His); rs11601507 (TRIM5 p.Val112Phe); rs3741380 (EHBP1L1 p.Arg307Gln); and rs1169288 (HNF1A p.Ile27Leu).
[0118] PRS When comparing AAF between a Japanese study (BBJ) and two European studies (C4D and UKBB), very different allele frequencies were observed for 175 lead variants detected in the cross-ethnic meta-analysis. Nevertheless, we found significant positive correlations and concordant allele effects for these variants (Spearman's rho). BBJ / C4D =0.82, P BBJ / C4D =1.9×10 -43 ;P BBJ / UKBB =0.80, P BBJ / UKBB =2.5×10 -40 Directional agreement was 96% for BBJ vs. C4D and for BBJ vs. UKBB.
[0119] Cross-ethnic genetic correlations To examine the consistency of allelic effects, cross-ethnic genetic correlation analyses were performed. Popcorn (v.0.9.9) analysis was applied to cross-ethnic genetic correlation analysis. 1KG EUR and 1KG EASPopulations were used according to the software instructions. Variants in the major histocompatibility complex region (chromosome 6; 25,000,000-35,000,000) were excluded, and the analysis was restricted to variants at only HapMap3 SNPs. For direct comparison of allele frequencies, summary statistics from the BBJ and UKBB studies were merged, and then variants at various P-value thresholds (1, 0.1, 0.05) were pruned. P-values and 1KG thresholds reported in a previous study (Nikpay, M. et al. Nat. Genet. 47, 1121-1130 (2015)) were used. EUR was used as the reference LD for pruning. The genetic correlation between BBJ and C4D or UKBB was determined by the popcorn algorithm, and the genetic correlation between C4D and UKBB was determined by LD score regression.
[0120] We found a strong correlation between the current study (BBJ) and European studies (C4D or UKBB). Furthermore, we directly compared allele effects between the current study and a previously reported European study (Brown, BC & et al. Am. J. Hum. Genet. 99, 76-88 (2016)). We found a significant positive correlation. This positive association became even more significant when we filtered variants according to P values from the GWAS (P<0.10 and P<0.05) in the previous study (van der Harst, P. & Verweij, N. Circ. Res. 122, 433-443 (2018)).
[0121] PRS derivation and parameter tuning Cross-ethnic meta-analysis substantially increased the number of significant associations, suggesting common allelic effects within ethnic groups. These results indicate that cross-ethnic meta-analysis can improve the performance of PRSs. To determine the best PRS under these circumstances, we performed an exhaustive derivation of PRSs. We constructed 875 combinations of summary statistics, derivation methods, and parameters for PRS derivation (Figure 6).
[0122] To ensure independence between the derivation, validation, and test cohorts, a K-fold cross-validation approach was applied. K = 10 was used. First, individuals in the BBJ dataset were randomly divided into 10 groups, and nine of the 10 were used as the derivation cohort and one of the 10 as the validation cohort. The BBJ GWAS was then performed using the derivation cohort (derived GWAS). Using the results of the derived GWAS, weights were determined for the Japanese PRS. Further, a meta-analysis was performed using the derived GWAS, and weights were determined for the cross-ethnic PRS. Using these weights, the PRS was calculated for the withheld validation cohort, and its performance was verified. These procedures were repeated 10 times with different holdout validation cohorts to obtain performance values for each study. Note that the validation cohort was always excluded from the derived GWAS. For European studies (C4D, UKBB), weights were determined only once, then PRSs were calculated, and their performance was validated 10 times in the validation cohorts.
[0123] The performance of the PRS was evaluated using (1) the Nagelkerke pseudo-R obtained by modeling age, sex, and normalized PRS. 2 (2) Nagelkerke's pseudo-R 2 The area under the receiver operating characteristic curve within the model was measured and reported as (3) the OR for 1SD-PRS, and (4) the OR for individuals with PRS in the top 10 percentile versus the remaining individuals. The best model / parameter set for each discovery group (UKBB, C4D, BBJ, UKBB and C4D, UKBB and BBJ, C4D and BBJ, UKBB, C4D and BBJ) was calculated using Nagelkerke's pseudo-R 2 was determined by averaging, using three derivation methods: pruning and thresholding, the LDpred algorithm, and the metaGRS method. For single datasets (BBJ, C4D, UKBB), we applied the pruning and thresholding method and LDpred (https: / / github.com / bvilhjal / ldpred). For combined datasets (BBJ and C4D, BBJ and UKBB, C4D and UKBB, BBJ, C4D and UKBB), we applied two different methods to combine the data from these studies.
[0124] Meta-analysis was performed, and then pruning and thresholding methods were applied to determine weights for PRS from metaGWAS summary statistics. Weights were then determined for PRS from each GWAS separately using pruning and thresholding methods, and then combined using the metaGRS method. For P-value thresholding, P-value thresholds were set to 1.0, 0.5, 0.1, 0.05, and 5 × 10. -4 , 5×10 -6 and 5 x 10 -8 Apply as, r 2 Thresholds were applied as 1.0, 0.8, 0.6, 0.4, 0.2, and 0 (i.e., no pruning). For LDpred, variants were restricted to HapMap3 SNPs. For metaGRS, the mean correlation coefficient and mean β of the scores were used across the remaining nine tests for 1SD-PRS.
[0125] PRS performance test Using the best model and parameters determined in the derivation and validation process, weights were determined for each discovery group using all BBJ and European GWAS. Then, using the weights, PRSs were calculated for independent test cohorts (1,827 CAD case samples and 9,172 control samples from JPHC and J-MICC) to test their performance. Bootstrapping was applied to evaluate the distribution of PRS performance. Samples in the test cohort were randomly selected with replacement, and performance measures were calculated. This procedure was repeated 10 times. 5 The distribution of performance measures was estimated by repeating the test ≈ 0 times. Also, for each test, the Nagelkerke pseudo-R between each pair of scores (2 out of 7; 21 combinations) was calculated.2 (△R 2 ) and the pairwise differences were calculated. 2 ≦0 and △R 2 The number of >0s was counted to obtain a two-sided P value, and then the lower value was 2 × 10 -5 The significance was P = 2.3 × 10 -3 The results of the performance test are shown in Figures 7 to 9.
[0126] As a result, the PRS derived from the Japanese data performed better than the PRS derived from the European data, despite the small sample size (UKBB vs. BBJ, P = 0.009; C4D vs. BBJ, P = 0.016). Interestingly, the PRS derived from the combined Japanese and European data performed better than the Japanese- and European-specific PRSs (P < 2 × 10 -5 ) significantly outperformed the best PRS obtained from a meta-analysis of three studies (BBJ, C4D, and UKBB, including 75,028 variants, pseudo-R 2 = 0.087 (0.074-0.101), area under the receiver operating characteristic curve = 0.674 (0.661-0.687), OR for 1SD-PRS = 1.840 (1.744-1.943), OR for individuals with PRS in the top 10th percentile vs. the remaining individuals in the study population = 2.649 (2.295-3.046). PRS derived from meta-analyses of C4D and BBJ included more cases and controls, and were significantly lower for C4D and UKBB (P < 2 × 10 -5 ) showed significantly better performance than
[0127] PRS-related phenotypes To determine the clinical pathways that explain the contribution of CAD-PRS to CAD pathophysiology, we evaluated the correlation between the CAD and PRS that achieved the best performance in the test dataset (CAD and PRS derived from the BBJ, C4D, and UKBB meta-analyses) and various phenotypes. Half of the control sample was randomly assigned and withheld for clinical evaluation of the CAD-PRS. The remaining control sample and all case samples were used to conduct a Japanese GWAS, cross-ethnic meta-analysis, and PRS derivation. Using the derived model, the PRS was calculated for the withheld control sample, and the relationship between the CAD-PRS and clinical indicators, including 32 numerical clinical traits and two binary lifestyle traits (i.e., smoking and alcohol consumption), was evaluated. For numerical traits, Spearman's correlation coefficients, associated confidence intervals, and P values were calculated. For dichotomous traits, logistic regression analyses were performed for individuals recruited to BBJ, adjusting for sex, age, age, top 10 principal components, and disease status. The significance threshold was set at P<0.05 / 34. As a result, we observed significant correlations between 13 clinical factors and the CAD-PRS among the 34 variables tested (Figure 10).
[0128] Next, we evaluated the impact of CAD-PRS on mortality in long-term follow-up data. For survival analysis, survival follow-up data, including causes of death under the International Statistical Classification of Diseases and Related Health Problems, 10th Revision, ICD-10 codes (diseases of the circulatory system (I00–I99)), were obtained for 132,737 individuals from the BBJ project. Vital status was collected from medical records or resident cards. Vital statistics were then obtained from the Statistics and Information Department of the Ministry of Health, Labor, and Welfare, and causes of death were identified according to ICD-10. The follow-up rate was 97%, and the median follow-up period was 7.7 years. Causes of death were divided based on ICD-10 classification, and categories with fewer than 100 events were excluded. HRs and associated P values were calculated for genotype dosage or PRS using Cox proportional hazards models adjusted for sex, age, age, the top 10 principal components, and disease status. Analyses were performed using the R package survival v.2.44, and survival curves were estimated using the R package survminer v.0.4.6 with modifications.
[0129] As a result, as illustrated in Figure 11, individuals with high CAD-PRS showed a significant increase in mortality from diseases classified as circulatory system-related (ICD-10 (I00-I99) diseases) (HR for a 1 standard deviation (95% CI) increase in CAD-PRS = 1.10 (1.06-1.15), P = 7.4 × 10 -6 This association was highly specific for cardiovascular disease; other causes of death were not associated with CAD-PRS. A significant association was found between CAD-PRS and all-cause mortality (HR (95% CI) = 1.03 (1.01-1.05), P = 2.2 × 10 -3 However, when ICD-10 (I00-I99) diseases were excluded from all-cause mortality, no significance was observed (HR (95% CI) = 1.01 (0.99-1.04), P = 0.31). Furthermore, individuals who died from ICD-10 (I00-I99) diseases were divided into three subcategories: ischemic heart disease (ICD-10 (I20-I25)), heart failure (ICD-10 I50), and stroke (ICD-10 (I60-I69)). Two of these showed significant associations with CAD-PRS (Figure 11).
[0130] Genetic basis of CAD-PRS-associated traits To further characterize the relationship between CAD-associated loci and pleiotropic effects on CAD risk factors, unsupervised clustering was performed on the Z-score matrix of these variants (175 genome-wide significant loci), and k-means clustering of the Z-scores revealed distinct functional clusters.
[0131] Loci with positive effects on serum TC or LDL-C levels were separated into cluster 1. This cluster includes loci associated with serum lipid profiles, including PCSK9, APOB, HMGCR, and LDLR. Cluster 2 contains loci associated with glycemic traits. Loci in cluster 3 positively affect white blood cells, and the 9p21 (CDKN2B-AS1) locus is included in this cluster. Loci in cluster 4 positively affect serum triglyceride levels and negatively affect HDL-C levels. This cluster also includes the APOA5 and LPL loci, which are very important genes for lipoprotein metabolism. Cluster 5 contains lead variants that affect body mass index (BMI). Cluster 6 is the second largest cluster and includes loci associated with blood pressure.
[0132] Candidate causal genes at 175 genome-wide significant loci were prioritized by combining sets of observations to identify the most likely candidate causal gene for each locus.
[0133] As a result, we identified several genes with high predictive scores in the new loci. LOXL1 and MYH11 were prioritized genes in loci clustered in the "blood pressure" cluster (cluster 6). These two genes, highly expressed in arterial tissue, support their effect on blood pressure (https: / / gtexportal.org / home). MYH11 is an established marker of smooth muscle cells and is also expressed in related tissues (colon, esophagus). LOXL1 is not expressed in these tissues but is expressed in fibroblasts. Furthermore, MERTK achieved the highest score in cluster 3, which is related to immune response. Additionally, within lipid-related cluster 1, SCD was implicated as a causative gene in a new genome-wide significant locus on chromosome 10q24. SCD plays an important role in lipid metabolism. The lead variant of this locus, rs603424, is located in an intron of PKD2L1 and is an expression quantitative trait locus variant for SCD in adipose tissue. Furthermore, we implicated WT1 as a causative gene at a new locus on chromosome 11p13 in blood glucose trait-related cluster 2. WT1 is known as a marker for progenitor cells of visceral adipose tissue and plays an important role in glucose tolerance.
Claims
1. A method for determining a subject's risk of developing ischemic heart disease, comprising: preparing a data list including effect alleles and effect sizes of genetic variants associated with ischemic heart disease; obtaining genetic information of a subject; A step of calculating a risk score for developing ischemic heart disease from the genetic information of the subject based on information regarding the effect allele and effect size of the ischemic heart disease-related genetic variant in the data list; and determining the risk of developing ischemic heart disease based on the risk score; Including, The data list is a data list obtained by integrating the genome analysis results of European and American populations and the genome analysis results of populations other than European and American populations through meta-analysis, The genetic variants included in the data list are rs62190384, rs56171536, rs2380472, rs35093463, rs73386640, rs2607903, rs112735431, rs9951447, rs11552449, rs1481345, rs7420881, rs72822411, rs7604403, rs12469628, and rs2111485 , rs7566501, rs12714757, rs11133381, rs13105983, rs7680806, rs13354746, rs9351209, rs2744427, rs47 23406, rs6593297, rs9693598, rs2623168, rs13255004, rs35093463, rs11000448, rs603424, rs73392700, the gene is characterized by containing one or more single nucleotide polymorphisms selected from rs7115190, rs11230728, rs1307145, rs2607903, rs9556903, rs35224956, rs28522673, rs1879454, rs2076438, rs216158, rs12444314, rs4794213, rs2952286, rs948386, rs12625329, and rs6006426 (note that rs numbers indicate registration numbers in the dbSNP database of the National Center for Biotechnology Information); method.
2. 2. The method of claim 1, wherein the non-Western population is an Asian population.
3. The method of claim 1 or 2, wherein the risk score is calculated using a pruning and thresholding method.
4. The meta-analysis according to any one of claims 1 to 3 is performed using a random effects model. How to do it.
5. The method according to any one of claims 1 to 4, wherein the ischemic heart disease is myocardial infarction or angina pectoris.
6. A system for determining a subject's risk of developing ischemic heart disease, comprising: means for storing data including a data list including effect alleles and effect sizes of genetic variants associated with ischemic heart disease; means for receiving the subject's genetic information; A means for calculating a risk score for developing ischemic heart disease from the genetic information of a subject based on data on the effect allele and effect size of the ischemic heart disease-related genetic variant in the data list; and a means for presenting a risk of developing ischemic heart disease based on the risk score; The data list is a data list obtained by integrating the genome analysis results of European and American populations and the genome analysis results of populations other than European and American populations through meta-analysis, The genetic variants included in the data list are rs62190384, rs56171536, rs2380472, rs35093463, rs73386640, rs2607903, rs112735431, rs9951447, rs11552449, rs1481345, rs7420881, rs72822411, rs7604403, rs12469628, and rs2111485 , rs7566501, rs12714757, rs11133381, rs13105983, rs7680806, rs13354746, rs9351209, rs2744427, rs47 23406, rs6593297, rs9693598, rs2623168, rs13255004, rs35093463, rs11000448, rs603424, rs73392700, the gene is characterized by containing one or more single nucleotide polymorphisms selected from rs7115190, rs11230728, rs1307145, rs2607903, rs9556903, rs35224956, rs28522673, rs1879454, rs2076438, rs216158, rs12444314, rs4794213, rs2952286, rs948386, rs12625329, and rs6006426 (note that rs numbers indicate registration numbers in the dbSNP database of the National Center for Biotechnology Information); system.
7. The system according to claim 6 , wherein the ischemic heart disease is myocardial infarction or angina pectoris.
8. A method for examining ischemic heart disease, comprising: rs62190384, rs56171536, rs2380472, rs35093463, rs73386640, rs2607903, rs112735431, rs995144 7, rs11552449, rs1481345, rs7420881, rs72822411, rs7604403, rs12469628, rs2111485, rs7566501 , rs12714757, rs11133381, rs13105983, rs7680806, rs13354746, rs9351209, rs2744427, rs4723406 , rs6593297, rs9693598, rs2623168, rs13255004, rs35093463, rs11000448, rs603424, rs73392700, rs7115190, rs11230728, rs1307145, rs2607903, rs9556903, rs35224956, rs28522673, rs1879454, rs2076438, rs216158, rs12444314, rs4794213, rs2952286, rs948386, rs12625329, and rs6006426 (Note that rs numbers indicate registration numbers in the dbSNP database of the National Center for Biotechnology Information) and analyze one or more single nucleotide polymorphisms selected from the above.
1. A method for testing a subject for ischemic heart disease, comprising: (1) For rs62190384, if the base corresponding to position 230007146 of chromosome 2 in the reference sequence hg19 (GRCh37) is C; (2) rs56171536 corresponds to chromosome 6, position 74415868, of the reference sequence hg19 (GRCh37) When the base is C: (3) For rs2380472, if the base corresponding to position 69431711 of chromosome 8 in the reference sequence hg19 (GRCh37) is G; (4) For rs35093463, if the base corresponding to position 107586238 of chromosome 9 in the reference sequence hg19 (GRCh37) is A; (5) rs73386640 corresponds to base 203235 of chromosome 11 in the reference sequence hg19 (GRCh37) When the base is G: (6) rs2607903 corresponds to chromosome 12, position 10876573, of the reference sequence hg19 (GRCh37) When the base is C: (7) For rs112735431, when the base corresponding to position 78358945 of chromosome 17 of the reference sequence hg19 (GRCh37) is A; (8) For rs9951447, when the base corresponding to chromosome 18 position 20009691 of the reference sequence hg19 (GRCh37) is C; (9) For rs11552449, if the base corresponding to chromosome 1 position 114448389 of the reference sequence hg19 (GRCh37) is C; (10) rs1481345 corresponds to chromosome 1 position 218827855 of the reference sequence hg19 (GRCh37) When the base is C: (11) For rs7420881, if the base corresponding to position 62878928 of chromosome 2 in the reference sequence hg19 (GRCh37) is C; (12) rs72822411 corresponds to chromosome 2, position 65499468, of the reference sequence hg19 (GRCh37) When the base is A: (13) rs7604403 corresponds to chromosome 2, position 112656652, of the reference sequence hg19 (GRCh37) When the base is A: (14) For rs12469628, if the base corresponding to chromosome 2 position 144158418 of the reference sequence hg19 (GRCh37) is C; (15) rs2111485 corresponds to chromosome 2, position 163110536, of the reference sequence hg19 (GRCh37) When the base is G: (16) rs7566501 corresponds to chromosome 2, position 230009317, of the reference sequence hg19 (GRCh37) When the base is T: (17) rs12714757 corresponds to chromosome 3 position 69820782 of the reference sequence hg19 (GRCh37) When the base is C: (18) rs11133381 corresponds to chromosome 4, position 56316979, of the reference sequence hg19 (GRCh37). When the base is C: (19) rs13105983 corresponds to chromosome 4, position 73420634, of the reference sequence hg19 (GRCh37) When the base is A: (20) rs7680806 corresponds to chromosome 4, position 186692853, of the reference sequence hg19 (GRCh37). When the base is C: (21) rs13354746 corresponds to position 74619132 on chromosome 5 of the reference sequence hg19 (GRCh37). When the base is T: (22) For rs9351209, when the base corresponding to position 90314917 of chromosome 6 in the reference sequence hg19 (GRCh37) is G; (23) rs2744427 corresponds to chromosome 6, position 149714790, of the reference sequence hg19 (GRCh37). When the base is T: (24) For rs4723406, if the base corresponding to position 35286471 of chromosome 7 in the reference sequence hg19 (GRCh37) is A; (25) For rs6593297, if the base corresponding to position 56122058 of chromosome 7 in the reference sequence hg19 (GRCh37) is T; (26) For rs9693598, when the base corresponding to position 25064984 of chromosome 8 of the reference sequence hg19 (GRCh37) is A; (27) For rs2623168, if the base corresponding to position 95260225 of chromosome 8 in the reference sequence hg19 (GRCh37) is G; (28) For rs13255004, when the base corresponding to position 102832405 of chromosome 8 of the reference sequence hg19 (GRCh37) is T; (29) For rs35093463, if the base corresponding to position 107586238 of chromosome 9 in the reference sequence hg19 (GRCh37) is A; (30) rs11000448 corresponds to chromosome 10, position 74682633, of the reference sequence hg19 (GRCh37). When the base is G: (31) rs603424 corresponds to chromosome 10, position 102075479, of the reference sequence hg19 (GRCh37). When the base is A: (32) rs73392700 corresponds to chromosome 11, position 224845, of the reference sequence hg19 (GRCh37). When the base is C: (33) For rs7115190, when the base corresponding to position 32441377 of chromosome 11 of the reference sequence hg19 (GRCh37) is T; (34) rs11230728 corresponds to chromosome 11, position 61277698, of the reference sequence hg19 (GRCh37). When the base is G: (35) rs1307145 corresponds to chromosome 11, position 118950217, of the reference sequence hg19 (GRCh37). When the base is G: (36) For rs2607903, when the base corresponding to position 10876573 of chromosome 12 of the reference sequence hg19 (GRCh37) is A; (37) For rs9556903, when the base corresponding to position 98859335 of chromosome 13 of the reference sequence hg19 (GRCh37) is A; (38) For rs35224956, when the base corresponding to position 103900481 of chromosome 14 of the reference sequence hg19 (GRCh37) is G; (39) rs28522673 corresponds to the 74223716th position on chromosome 15 of the reference sequence hg19 (GRCh37). When the base is C: (40) For rs1879454, when the base corresponding to position 81377717 of chromosome 15 of the reference sequence hg19 (GRCh37) is A; (41) rs2076438 corresponds to chromosome 16, position 1584618, of the reference sequence hg19 (GRCh37). When the base is C: (42) rs216158 corresponds to chromosome 16, position 15917838, of the reference sequence hg19 (GRCh37). When the base is G: (43) rs12444314 corresponds to chromosome 16, position 86699163, of the reference sequence hg19 (GRCh37). When the base is G: (44) For rs4794213, if the base corresponding to position 49308707 of chromosome 17 of the reference sequence hg19 (GRCh37) is T; (45) For rs2952286, if the base corresponding to position 66469400 of chromosome 17 of the reference sequence hg19 (GRCh37) is G; (46) rs948386 corresponds to chromosome 18, position 19998810, of the reference sequence hg19 (GRCh37) When the base is C: (47) rs12625329 corresponds to chromosome 20, position 62709274, of the reference sequence hg19 (GRCh37). When the base is A; or (48) For rs6006426, if the base corresponding to the 30669883 base on chromosome 22 of the reference sequence hg19 (GRCh37) is A; A method for determining whether a person is at high risk of developing ischemic heart disease.
9. The method according to claim 8, wherein the ischemic heart disease is myocardial infarction or angina pectoris.
Citation Information
Patent Citations
Method for examining arteriosclerotic disease based on single nucleotide polymorphism on human chromosome 5p15.3
JP2011167124A