Method and system for determining risk of heart failure
A cross-ethnic GWAS and meta-analysis identify novel heart failure susceptibility loci, enabling the development of a polygenic risk score that accurately predicts heart failure risk and prognosis across diverse ethnic groups, addressing the limitations of previous ethnic-specific analyses.
Patent Information
- Application Number
- PCT/JP2025/012589
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-27
- Filing Date
- 2025-03-27
- Publication Date
- 2025-10-02
AI Technical Summary
Existing genetic analyses for heart failure have focused primarily on specific ethnic groups, making it difficult to predict heart failure risk universally across different races, and there is a lack of polygenic risk scores optimized for non-Western populations like the Japanese.
Conduct a genome-wide association study (GWAS) across diverse ethnicities, including Japanese and European populations, to identify disease susceptibility loci for heart failure subtypes, followed by a meta-analysis to integrate these findings, and develop a polygenic risk score (PRS) that combines genetic polymorphisms from various ethnic groups to predict heart failure risk accurately.
The PRS significantly outperforms previous scores, enabling accurate prediction of heart failure risk and prognosis, laying the foundation for precision medicine by providing guidelines for prevention and treatment across diverse ethnicities.
Smart Images

Figure JP2025012589_02102025_PF_FP_ABST
Abstract
Description
Method and system for assessing risk of heart failure
[0001] The present invention relates to a method and system for assessing the risk of heart failure using gene sequence analysis.
[0002] Heart failure (HF) has become a global pandemic with increasing prevalence and incidence, affecting approximately 26 million people worldwide. Previous genome-wide association studies (GWAS) of all-cause HF analyzed 115,150 cases and 1,550,331 controls and identified 47 disease susceptibility loci. The number of disease susceptibility loci identified for HF is smaller than the 175 disease susceptibility loci identified in a GWAS of ischemic heart disease (CAD) (121,234 cases) and the 150 disease susceptibility loci identified in a GWAS of atrial fibrillation (AF) (77,960 cases). This is due in part to the heterogeneity of HF syndromes resulting from all types of cardiovascular disease. To address this heterogeneity, GWASs of heart failure with reduced left ventricular ejection fraction (HFrEF) and heart failure with preserved left ventricular ejection fraction (HFpEF) were conducted (Non-Patent Documents 1-5). This GWAS uncovered additional disease susceptibility loci not found in GWAS for all-cause heart failure. Analysis of heart failure subtypes is useful for better understanding the genetic architecture of heart failure. However, few polygenic risk scores (PRS) for heart failure have demonstrated their usefulness. Furthermore, because most HF-GWAS studies have focused on European populations, no large-scale genomic studies of heart failure have been conducted in non-Western populations, such as Japanese populations. Therefore, no PRS for heart failure optimized for Japanese individuals existed.
[0003] Ponikowski P. et al. Heart failure: preventing disease and death worldwide. ESC Heart Fail. 1, 4-25 (2014).Levin MG. et al. Genome-wide association and multi-trait analyses characterize the common genetic architecture of heart failure. Nat Commun. 13, 6914 (2022).Koyama S. et al. Population-specific and trans-ancestry genome-wide analyses identify distinct and shared genetic risk loci for coronary artery disease. Nat Genet. 52,1169-1177 (2020).Miyazawa K. et al. Cross-ancestry genome-wide analysis of atrial fibrillation unveils disease biology and enables cardioembolic risk prediction. Nat Genet. 55, 187-197 (2023).Joseph J. et al. Genetic architecture of heart failure with preserved versus reduced ejection fraction. Nat Commun. 13, 7753 (2022).
[0004] Predicting the risk and prognosis of heart failure based on genetic information is expected to play an important role in advancing the medical science and treatment of heart failure. However, previous genetic analyses of heart failure have focused on specific ethnic groups, making it difficult to predict risk universally across races. In other words, it is known that the distribution of genetic polymorphisms varies widely among ethnicities, and it was unclear whether the results of studies using European populations could be applied to non-Western populations, such as East Asian populations, including the Japanese.
[0005] Therefore, an object of the present invention is to provide a method and system that can be applied across a wide range of races and that can determine the risk of heart failure with high accuracy.
[0006] First, the inventors' research group conducted a GWAS to comprehensively detect genetic polymorphisms characteristic of heart failure patients by comparing the genome sequences of 16,251 patients with all-cause heart failure and three heart failure subtypes, HFrEF, HFpEF, and non-ischemic heart failure (NIHF), registered in BioBank Japan (BBJ), with 197,577 controls. As a result, 18 disease susceptibility loci (gene loci) related to heart failure were identified, five of which were novel disease susceptibility loci that had not been reported previously.
[0007] Next, we conducted a meta-analysis to integrate the results of the Japanese GWAS (approximately 210,000 individuals) with those of GWASs on all-cause HF populations in Europe, China, America, and Africa, GWASs on European HFrEF and HFpEF populations, and GWASs on the NIHF populations from the UK Biobank and Finnish Biobank (FinnGen). This resulted in the world's largest cross-ethnic GWAS on HF, covering approximately 1.67 million individuals with all-cause HF, 480,000 individuals with HFrEF, 480,000 individuals with HFpEF, and 880,000 individuals with NIHF. As a result, we identified 58 disease susceptibility loci associated with HF. Eight of these were novel disease susceptibility loci not previously reported, and one was a novel disease susceptibility locus identified in the Japanese GWAS. A total of 12 novel disease susceptibility loci were identified by the Japanese GWAS and the cross-ethnic meta-analysis.
[0008] Next, we performed a multi-trait analysis of genetically engineered heart failure (GWAS) (Nature Genetics, 2018, vol. 50, pp. 229-237) to evaluate genetic correlations between heart failure and left ventricular, left atrial, right ventricular, and right atrial parameters, left ventricular mass, and fibrosis. Because left ventricular mass was strongly correlated with all heart failure subtypes, combining the results of HF-GWAS and left ventricular mass identified nine novel genetic loci.
[0009] Next, we developed a polygenic risk score (HF-PRS) using genetic polymorphisms obtained through cross-ethnic GWAS and genomic analysis of heart failure-related phenotypes such as AF, CAD, and left ventricular mass. We then evaluated its predictive performance and found that the HF-PRS significantly outperformed PRSs based on Japanese and European populations. Furthermore, analysis of Japanese data (BBJ) using the HF-PRS revealed that it could efficiently predict the age at onset and prognosis of heart failure. Based on these findings, we have completed the present invention.
[0010] One aspect of the present invention relates to a method for assessing a subject's risk of heart failure, comprising the steps of: preparing a data list including effect alleles and effect sizes of genetic polymorphisms; obtaining the subject's genetic information; calculating a heart failure risk score from the subject's genetic information based on the information on the effect alleles and effect sizes of genetic polymorphisms in the data list; and assessing the risk of heart failure based on the risk score, wherein the data lists include a data list of genomic analysis results for heart failure in Caucasian populations, a data list of genomic analysis results for heart failure in populations other than Caucasian, a data list of genomic analysis results for heart failure subtypes, a data list of genomic analysis results for diseases related to heart failure, and a data list of genomic analysis results for measured values related to cardiac function, and the risk score is a risk score obtained by combining the risk scores calculated from these data lists.
[0011] Another aspect of the present invention relates to a system for determining a subject's risk of heart failure, comprising: means for storing data including a data list containing effect alleles and effect sizes of genetic polymorphisms; means for receiving the subject's genetic information; means for calculating a heart failure risk score from the subject's genetic information based on the data on the effect alleles and effect sizes of genetic polymorphisms in the data list; and means for presenting the risk of heart failure based on the risk score, wherein the data lists include a data list of genomic analysis results for heart failure in Caucasian populations, a data list of genomic analysis results for heart failure in populations other than Caucasian, a data list of genomic analysis results for heart failure subtypes, a data list of genomic analysis results for diseases related to heart failure, and a data list of genomic analysis results for measurement values related to cardiac function, and the risk score is a risk score obtained by combining the risk scores calculated from these data lists.
[0012] Another aspect of the present invention relates to a method for assessing the risk of heart failure, which comprises a step of analyzing a single nucleotide polymorphism (SNP) at rs1484116 (the rs number indicates the registration number in the dbSNP database of the National Center for Biotechnology Information) present in the TTN gene region.
[0013] According to the method and system of the present invention, a risk score can be calculated using a data list including genomic analysis results for heart failure in multiple ethnicities, genomic analysis results for heart failure subtypes, genomic analysis results for heart failure-related diseases, and genomic analysis results for cardiac function-related measurements, thereby enabling accurate assessment of the risk of heart failure. This provides guidelines for the prevention and treatment of heart failure and lays the foundation for the realization of precision medicine for heart failure.
[0014] This figure shows a flowchart of a study consisting of a Japanese GWAS using BBJ, a cross-ethnic meta-analysis using GWAS from European and other large populations, and downstream analyses. The odds ratios for HF onset for independent signals in the Japanese GWAS, cross-ethnic meta-analysis, and MTAG are shown. The color of each point indicates the corresponding HF subtype. The size of each point indicates -log10 (P value). The shape of each point indicates the analysis method. Markers within each point indicate novelty. The performance evaluation (Nagelkerke pseudo-coefficient of determination) of the HF-PRS for Japanese people for each cohort combination is shown. The association between HF-PRS and age at onset of heart failure is shown. The age at onset of heart failure for individuals with available data (n = 10,810) is shown based on tertiles of HF-PRS. The number of individuals in each tertile ranges from 3,603 to 3,604. The center line of the boxplot indicates the median, the borders indicate the first and third quartiles, and the whiskers indicate 1.5 times the interquartile range. Kaplan-Meier estimates of cumulative events for cardiovascular death (left) and heart failure death (right) are shown with 95% confidence intervals. Patients were classified into high (top tertile), intermediate (middle tertile), and low (bottom tertile) HF-PRS groups. P values for HF-PRS were calculated using Cox proportional hazards analysis, with statistical significance of P < 8.3 × 10 after correction for multiple testing. -3 The following conditions must be met. The effects of genetic polymorphisms in cardiomyopathy genes on left ventricular function and long-term HF mortality are shown. Figure 6a shows the minor allele frequencies of lead SNPs in cardiomyopathy genes in East Asian and European populations using gnomAD. Figure 6b shows the effect of lead variants in cardiomyopathy genes on left ventricular ejection fraction (LVEF). Data are presented as estimated coefficients and their 95% confidence intervals. Point color indicates HF phenotype. Filled points indicate statistical significance. Figure 6c shows the effect of lead variants in cardiomyopathy genes on HF mortality (HF patients (ICD-10 I50, 326 deaths out of 8,481 patients) and non-HF patients (ICD-10 I50, 749 deaths out of 121,070 patients)). Figure 6d shows the cumulative incidence curves of HF mortality by carriers of rs1484116(TTN) in HF patients.
[0015] <Method for assessing the risk of heart failure of the present invention> The method for assessing the risk of heart failure of the present invention includes the steps of: preparing a data list including effect alleles and effect amounts of genetic polymorphisms (data list preparation step); obtaining genetic information of a subject (genetic information acquisition step); calculating a heart failure risk score from the genetic information of the subject based on the genetic polymorphism data in the data list (risk score calculation step); and assessing the risk of heart failure based on the risk score (assessment step).
[0016] Heart failure is a clinical syndrome in which some kind of cardiac dysfunction occurs, resulting in a breakdown in the compensatory mechanism of the cardiac pumping function, resulting in dyspnea, fatigue, and edema, and a corresponding decrease in exercise tolerance.Heart failure can be classified into HFrEF, HFpEF, and NIHF subtypes.
[0017] Risk of heart failure includes the risk of developing heart failure and the risk of worsening heart failure prognosis (including risk of death).
[0018] Each step will be described below.
[0019] <Data List Preparation Step> The data list is a list of data on effect alleles (risk alleles) and their effect sizes (degree of association with disease or cardiac function) for a large number of variants (genetic polymorphisms) associated with heart failure, heart failure subtypes, heart failure-related diseases, or measurements related to cardiac function (heart failure-related variants, heart failure subtype-related variants, heart failure-related disease-related variants, cardiac function-related variants). Here, the genetic polymorphisms are not particularly limited in type, as long as they are mutations that differ from the wild-type sequence, such as SNPs, base deletions, and base insertions.
[0020] The effect size of a genetic polymorphism associated with heart failure, a heart failure subtype, a disease associated with heart failure, or cardiac function is set according to the degree of association of each allele with heart failure, a heart failure subtype, a disease associated with heart failure, or cardiac function, for example, its frequency of occurrence in heart failure patients.
[0021] The heart failure-associated variant, heart failure subtype-associated variant, heart failure-associated disease-associated variant, or cardiac function-associated variant is preferably a variant with a minor allele frequency of 1% or more.
[0022] In the method of the present invention, the data list is characterized by including genome analysis results for heart failure in a Western population, genome analysis results for heart failure in a population other than Westerners (non-Western populations), genome analysis results for heart failure subtypes, genome analysis results for diseases related to heart failure, and genome analysis results for measurements related to cardiac function. Here, Western populations include, for example, Caucasian populations, such as European and American populations. Furthermore, non-Western populations include Asian and African populations, and Asian populations include, for example, Japanese and Chinese populations.
[0023] Here, heart failure in Western populations and heart failure in non-Western populations refer to all-cause heart failure, including all subtypes of heart failure.
[0024] Examples of heart failure subtypes include HFrEF, HFpEF, and NIHF. Here, examples of races for the heart failure subtypes include the aforementioned Caucasian population.
[0025] Examples of diseases related to heart failure include CAD and AF. Here, examples of races related to diseases related to heart failure include the aforementioned Caucasian population.
[0026] The cardiac function-related measurement is preferably a measurement genetically associated with heart failure, and the genetic association with heart failure is evaluated, for example, by LD score regression between heart failure and the cardiac function-related measurement.
[0027] An example of a measurement value related to cardiac function is left ventricular mass. Here, examples of races related to left ventricular mass include the aforementioned Caucasian population.
[0028] Genomic analysis results are preferably large-scale, for example, results obtained from more than 10,000 subjects. Also, GWAS results are preferably used. Genomic analysis results for heart failure in Western populations, genomic analysis results for heart failure in populations other than Western populations, genomic analysis results for diseases related to heart failure, and genomic analysis results for heart failure subtypes may be genomic analysis results that have already been performed and are publicly available, or may be the results of newly performed genomic analysis. Examples of publicly available genomic analysis results include, for example, BBJ for Japanese people and China Kadoorie Biobank for Chinese people. For Western people, examples include analysis results from the Global Biobank Meta-analysis Initiative (GBMI), UK Biobank, and the FinnGen project.
[0029] The genome analysis results for cardiac function-related measurements may be those that have already been analyzed and are publicly available, or may be the results of new genome analysis. Examples of publicly available genome analysis results include datasets from the Cardiovascular Knowledge Portal (CVDKP) and Zenodo. Examples of CVDKP datasets include Left ventricular mass 2023 GWAS: trans-ancestry and Left ventricular parameters 2019 GWAS: European ancestry. Examples of Zenodo datasets include, but are not limited to, the dataset published by Gustav Ahlberg et al. in "Genome-wide association study identifies 18 novel loci associated with left atrial volume and function" (European Heart Journal, Volume 42, Issue 44, November 2021, pp. 4523-4534).
[0030] The data list may be obtained by meta-analysis. There are no particular limitations on the meta-analysis method, and general meta-analysis methods can be used. A fixed-effects model that does not take into account heterogeneity between studies may be used, but a random-effects model that takes into account heterogeneity between studies is preferred. Meta-analysis can be performed using analysis software such as METAL (https: / / genome.sph.umich.edu / wiki / METAL).
[0031] <Genetic Information Acquisition Step> In this step, genetic information of the subject, i.e., genome sequence information, is acquired. The genome sequence information may be any information that provides sequence information corresponding to each variant to be analyzed included in the data list, but whole genome sequence information is preferred.
[0032] Genetic information of a subject can be obtained by conventional gene sequence analysis. For example, genetic information of a subject can be obtained using an SNP genotyping array or a next-generation sequencer. Note that, in the method of the present invention, gene sequence analysis is not essential; it is sufficient to obtain data that has already been analyzed.
[0033] In the present invention, the subject of analysis is a human, for example, a human suspected of being at risk of heart failure. The subject's race is not particularly limited, but examples include Asians, including Japanese, and Caucasians, including Europeans. The sample used for genetic information analysis is not particularly limited as long as it contains chromosomal DNA derived from the subject. Examples include body fluids such as blood, urine, and cerebrospinal fluid, cells from the cervix or oral mucosa, and body hair. While these samples can be used directly, it is preferable to isolate chromosomal DNA from these samples by standard methods and analyze it.
[0034] <Risk Score Calculation Step> In this step, for each of one or more data lists of populations, including a data list of genome analysis results for heart failure in Western populations, a data list of genome analysis results for heart failure in non-Western populations, a data list of genome analysis results for heart failure subtypes, genome analysis results for heart failure-related diseases, and a data list of genome analysis results for cardiac function-related measurements, a risk score corresponding to each data list is calculated from the subject's genetic information based on the risk allele and effect size data for heart failure-related variants, heart failure subtype-related variants, heart failure-related disease-related variants, or cardiac function-related variants. Furthermore, the risk scores of each data list are integrated to calculate the subject's heart failure risk score. Here, the risk score refers to a score (value) indicating disease risk calculated from genetic information, and may refer to, for example, a score calculated for each individual by weighting the sum of multiple genetic polymorphisms associated with the disease. For example, the genetic information of a subject is compared with the variant information in each data list, and the effect size (disease or cardiac function association) of each variant is multiplied by the number of risk alleles in the subject to calculate the sum of these values to calculate a risk score corresponding to each data list.Furthermore, the risk scores of each data list can be integrated, for example, by linear combination, to determine the heart failure risk score for the subject.
[0035] The risk score (PRS) corresponding to each data list k ) is exemplified by the following formula: PRS k =β1x1+β2x2+. .. .. β k x k +β n x n Here, β k is a SNP k The log odds ratio (OR) of a risk allele for heart failure, heart failure subtype, cardiac function-related measure, or heart failure-related disease is x k is SNP kn is the number of risk alleles (0, 1, or 2) for each SNP, and n is the total number of SNPs analyzed. Here, examples of heart failure include all-cause heart failure, HFrEF, HFpEF, and NIHF. Examples of measurements related to cardiac function include left ventricular mass. Examples of heart failure-related diseases include AF and CAD. Alternatively, PRS k can also be calculated logarithmically as follows: log(PRS k )=x1*logβ1+x2*logβ2+. .. .. x k *logβ k +x n *logβ n
[0036] The heart failure risk score (HF-PRS) is, for example, as illustrated below, a risk score (PRS) corresponding to each data list. k ) can be calculated by linear combination of HF-PRS = ω1PRS1 + ω2PRS2 + . . . ω k PRS k +ω n PRS n where ω k PRS k is the log odds ratio (OR) for heart failure, and PRS k is the risk score corresponding to each data list, and n is the total number of data lists. Here, heart failure includes, for example, all-cause heart failure, HFrEF, HFpEF, and NIHF. HF-PRS can also be calculated logarithmically as follows: log(HF-PRS) = PRS1 * logω1 + PRS2 * logω2 + ... PRS k *logω k +PRS n *logω n
[0037] Preferably, the risk score is calculated using PRS-CS, a Bayesian statistical method for calculating PRS, the method of which is publicly available (https: / / github.com / getian107 / PRScs).
[0038] Preferably, the method of linear combination is performed using Lasso regression.
[0039] <Assessment Step> In this step, the risk of heart failure is assessed based on the calculated heart failure risk score. That is, this step is an assessment step based on an objective index, the heart failure risk score, and does not rely on the judgment of a doctor. For example, if the heart failure risk score exceeds a certain reference value, it can be determined that the risk of heart failure is high. As the reference value, for example, a cutoff value can be determined by calculating the heart failure risk scores in advance for a group of healthy subjects and a group of heart failure patients. In addition, the disease risk can be ranked, for example, on a 5- or 10-level scale, according to the magnitude of the heart failure risk score. Furthermore, ranking can also be performed in association with the severity and prognosis of heart failure.
[0040] The method of the present invention can be implemented using a computer. That is, one embodiment of the method of the present invention is a computer-implemented method for determining the risk of heart failure, operable in a computer system including a processor and a memory, comprising: receiving genetic information of a subject; processing the received genetic information to determine a heart failure risk score; and determining and presenting the risk of heart failure based on the heart failure risk score. Here, the processing includes integrating each risk score calculated by comparing the received genetic information with information on the effect allele and effect size of each variant in a data list including genomic analysis results for heart failure in Caucasian populations, genomic analysis results for heart failure in populations other than Caucasian, genomic analysis results for heart failure subtypes, genomic analysis results for diseases related to heart failure, and genomic analysis results for cardiac function-related measurements.
[0041] The method of the present invention may be implemented as a computer-implemented method. That is, one aspect of the present invention provides a computer system for determining a subject's risk of heart failure.
[0042] <System for assessing the risk of developing heart failure> The system for assessing the risk of heart failure of the present invention comprises: a data storage means including a data list containing effect alleles and effect amounts of genetic polymorphisms; a means for receiving the genetic information of a subject; a means for calculating a heart failure risk score from the genetic information of the subject based on the information on the effect alleles and effect amounts of genetic polymorphisms in the data list; and a means for presenting the risk of heart failure based on the risk score, wherein the data list includes genomic analysis results for heart failure in Caucasian populations, genomic analysis results for heart failure in populations other than Caucasian, genomic analysis results for heart failure subtypes, genomic analysis results for diseases related to heart failure, and genomic analysis results for measurement values related to cardiac function, and the risk score is a risk score obtained by combining the risk scores calculated from these data lists.
[0043] In one embodiment, the data storage means containing the data list including the effect alleles and effect sizes of genetic polymorphisms is stored in a memory area within the computer. In another embodiment, the data storage means may be stored outside the computer, such as an external server, and may be accessed by the computer when in use.
[0044] The subject's genetic information may be received by being entered directly into the computer, or may be received from a user interface coupled to the computer, or the subject's genetic information may be received from a remote device via a wireless communication network.
[0045] The risk score calculation means compares the received genetic information with the data list stored in the data storage means and calculates a risk score based on the comparison result. For example, the risk score is calculated by substituting the comparison result for the risk allele into a pre-stored formula for calculating the risk score. The presentation means presents the risk assessment result based on the risk score and by referring to pre-stored criteria. Here, presenting the assessment result includes outputting information to a user interface connected to the computer system and transmitting information to a remote device via a wireless communication network. These means can be programmed and installed in a computer as software or an application, causing the computer system to perform functions such as processing the received genetic information (data comparison, risk score calculation, etc.) and presenting the results.
[0046] <Method for assessing the risk of heart failure using novel heart failure-associated SNPs> Another embodiment of the method for assessing the risk of heart failure of the present invention comprises the steps of analyzing the following SNPs and assessing the risk of heart failure based on the analysis results.
[0047] As mentioned above, heart failure is a clinical syndrome in which some kind of cardiac dysfunction causes the cardiac pump function to fail, resulting in dyspnea, fatigue, and edema, accompanied by a decrease in exercise tolerance. Heart failure can be classified as HFrEF, HFpEF, or NIHF. Risk of heart failure includes the risk of developing heart failure and the risk of worsening heart failure prognosis (including the risk of death).
[0048] rs1484116 is a SNP located in an intron of the TTN gene on chromosome 2, and refers to a cytosine (C) / guanine (G) SNP at base 179595117 of chromosome 2 in the reference sequence hg19 (GRCh37). When this base is G, the risk of heart failure increases. Furthermore, when genotype is taken into account in the analysis, the order of increasing risk of heart failure is GG>CG>CC.
[0049] Furthermore, the SNPs analyzed in the present invention are not limited to those mentioned above, and SNPs in linkage disequilibrium with the above SNPs may also be analyzed. 2 > 0.5, preferably r 2 > 0.8, more preferably r 2 SNPs that satisfy a relationship of > 0.9. 2 is the linkage disequilibrium coefficient. SNPs in linkage disequilibrium can be identified using, for example, the HapMap database (http: / / www.hapmap.org / index.html.ja).
[0050] By examining the types of bases of SNPs as described above and correlating the results with heart failure using an index such as a risk score based on the results, the risk of heart failure can be determined. That is, the method of the present invention can provide data for diagnosing heart failure.
[0051] The results determined by the method of the present invention are provided to a physician or the like as needed. The physician or the like who receives the results can then diagnose heart failure after performing necessary tests such as electrocardiogram, blood tests, and ultrasound. If the physician or the like diagnoses a high risk of heart failure, appropriate preventive measures such as drug administration can be taken, and if the physician or the like diagnoses that the patient has developed heart failure, treatment such as drug administration or surgery can be performed.
[0052] SNP analysis can be performed by a conventional genetic polymorphism analysis method, including, but not limited to, SNP genotyping array, sequence analysis, PCR, hybridization, and the Invader method.
[0053] The present invention can also provide test reagents, such as primers and probes, for testing heart failure. Examples of such probes include probes that contain the SNP site and can determine the type of base at the SNP site based on the presence or absence of hybridization. Specifically, these probes include probes that are 15 bases or longer and have a base sequence containing the polymorphic site of each SNP or its complementary sequence, and probes that are 15 bases or longer and have a sequence containing a base in linkage disequilibrium with the base or its complementary sequence. The length of the probe is preferably 15 to 35 bases, more preferably 20 to 35 bases.
[0054] Primers include those that can be used in PCR to amplify the SNP site or those that can be used for sequence analysis (sequencing) of the SNP site. Specifically, primers that can amplify or sequence a region containing the base at each SNP site, or primers that can amplify or sequence a region containing a base in linkage disequilibrium with the base, are included. The length of such primers is preferably 15 to 50 bases, more preferably 15 to 35 bases, and even more preferably 20 to 35 bases. Primers for sequencing SNP sites include primers having a sequence complementary to the 5' region of the base, preferably 30 to 100 bases upstream, or the 3' region of the base, preferably 30 to 100 bases downstream. Primers used to determine polymorphisms based on the presence or absence of PCR amplification include primers having a sequence containing the base and including the base on the 3' side, and primers having a complementary sequence to a sequence containing the base and including the complementary base of the base on the 3' side.
[0055] The present invention will be specifically described below with reference to examples, but the present invention is not limited to the following embodiments.
[0056] Methods: Samples. All Japanese GWAS subjects were enrolled in the BioBank Japan (BBJ) project (https: / / biobankjp.org / ). BBJ is a hospital-based nationwide biobank project that collects DNA, serum samples, and clinical information from 12 participating medical institutions nationwide (Osaka Prefectural Adult Disease Center, Cancer Institute Hospital Ariake, Juntendo University, Tokyo Metropolitan Institute of Gerontology, Nippon Medical School, Nihon University School of Medicine, Iwate Medical University, Tokushukai Hospital, Shiga University of Medical Science, Fukujuen Hospital, National Hospital Organization Osaka Hospital, and Iizuka Hospital). Approximately 270,000 patients with one of 47 target diseases were enrolled between 2003 and 2007 (Phase 1) and between 2013 and 2018 (Phase 2). All participants were aged 18 years or older. Informed consent was obtained from all participants, and this study was approved by the relevant ethics committees at each institution. For quality control of the GWAS, samples with a call rate <0.98 were excluded. Samples with a heterozygosity rate >4 standard deviations were excluded. To clarify population stratification, principal component analysis was performed using PLINK 2.0, and outliers in the Japanese cluster were excluded. Case samples for the GWAS were selected from individuals diagnosed with HF by a physician based on general medical practice. HFrEF was defined as HF with a left ventricular ejection fraction of <40%, and HFpEF was defined as HF with a left ventricular ejection fraction of ≥50%. NIHF was defined as HF without CAD.
[0057] Genotyping, imputation, and quality control: GWAS subjects were genotyped using Illumina Human OmniExpress Genotyping BeadChips or a combination of Illumina HumanOmniExpress and HumanExome BeadChips (Illumina, San Diego, CA). Quality control of genotypes was performed based on the following criteria: (i) SNP call rate < 0.99, (ii) MAF < 0.01, and (iii) Hardy-Weinberg equilibrium P value ≤ 1.0 × 10. -6Imputation was performed using minimac4 (https: / / genome.sph.umich.edu / wiki / Minimac4) with a reference panel of 7825 Japanese individuals.
[0058] GWAS For the Japanese GWAS, logistic regression analysis was performed using REGENIE (https: / / rgcgithub.github.io / regenie / ), employing an additive model adjusted for age, age squared, sex, and the top 10 principal components. Variants with a Minimac4 imputation quality score of 0.3 or higher and a MAF of 0.01 or higher were selected. The genome-wide significance threshold was P<5.0 × 10 -8 The genome inflation factors (λGC) were 1.044, 1.080, 1.050, and 1.083 for all-cause HF, HFrEF, HFpEF, and NIHF, respectively. Adjacent genome-wide significant SNPs were grouped into one locus if they were within 1 Mb of each other. Loci were defined as follows: (1) genome-wide significant variants (P < 5.0 × 10 -8 ) were extracted from the association results, (2) 500 Mb were added on either side of these variants, and (3) overlapping regions were merged. Loci were identified as those with known genome-wide significant variants (i.e., P < 5.0 × 10 for known loci). -8 A region was defined as novel if any coordinates contained no variants (i.e., variants in the locus). To identify independent association signals at loci, a stepwise conditioning analysis was performed on genome-wide significant loci identified in GWAS. First, a logistic regression was performed conditional on the lead variants at each locus. Significance across loci was determined when P < 1.0 × 10 -5 , and this procedure was repeated until no variants at each locus reached locus-wide significance.
[0059] LD score regression and heritability. LD-score regression (version 1.0.0) was performed to estimate confounding bias, such as population stratification. SNPs with a MAF of ≥ 0.01 were selected, and variants within the major histocompatibility complex region were excluded. The LD score for East Asians was used for regression (https: / / github.com / bulik / ldsc / ). The prevalence of all-cause heart failure was set at 6.5% according to the Japan Medical Data Center and Medical Data Vision dataset (ESC Heart Fail. 2023;10:1996-2009). Because HFrEF, HFpEF, and NIHF account for 37.4%, 45.1%, and 26.6% of all-cause heart failure cases, the respective prevalences were set at 2.4% for HFrEF, 2.9% for HFpEF, and 1.7% for NIHF.
[0060] Cross-ethnic meta-analysis: Fixed-effect meta-analyses based on inverse variance weighting were performed for all HF phenotypes. Genome-wide significant loci were defined by iteratively spanning a ±500 kb region around the most significant variant and merging overlapping regions until no genome-wide significant variants were detected within ±1 Mb. Loci were classified as "known" if the merged region was within ±500 kb of a variant for the corresponding phenotype in known all-cause HF, HFrEF, HFpEF, or NIHF GWAS; "identified in MTAG" if reported in a previous all-cause HF MTAG; or "novel" otherwise. The most significant variant at each locus was selected as the index variant. Cochran's Q test for heterogeneity was performed to identify loci with index variants that differed in effect size across GWAS datasets. Strong heterogeneity (P < 0.05) was not observed. het <1.0×10 -4 ) or variants with MAF<1% were excluded.
[0061] Multitrait association study: To identify additional HF risk loci, we used the European MTAG, which includes cardiac MRI traits most strongly correlated with HF. Left ventricular mass was included because it was most strongly correlated with all HF phenotypes. We excluded variants with MAF <1% and defined genome-wide significant loci in the same way as in the cross-ethnic meta-analysis.
[0062] Derivation and Performance of Polygenic Risk Score: First, we divided the dataset into three groups: (i) a discovery group for PRS derivation (12,335 cases, 107,333 controls), (ii) a group for linear combination (1,837 cases, 11,320 controls), (iii) a test group for PRS performance evaluation (1,837 cases, 11,319 controls), and (v) a survival analysis group (67,991 controls). Next, we used the PRS-CS (Nat Commun. 2019;10:1776) to derive PRSs for BBJ, European, American, and African all-cause heart failure, European HFrEF, HFpEF, NIHF, AF, CAD, and left ventricular mass. For each cohort, we used LD reference panels constructed using 1000 Genomes Project phase 3 samples. We then calculated the PRS in the retained BBJ cohort for linear combinations, using logistic regression with L1 regularization and 10-fold cross-validation. Finally, we constructed the HF-PRS using BBJ all-cause HF, European all-cause HF, African all-cause HF, European HFrEF, European HFpEF, European CAD, European AF, and European LV mass. The performance of the HF-PRS was evaluated using (1) the Nagelkerke pseudo coefficient of determination (R) obtained by modeling age, sex, the top 10 principal components, and the normalized PRS. 2 , and (2) Nagelkerke's pseudo coefficient of determination R 2 The HF-PRS was calculated using the aforementioned model and its performance was evaluated on an independent test cohort.
[0063] Association between HF-PRS and age at onset of HF To evaluate the association between HF-PRS and age at onset of HF, we extracted HF case samples for which data on age at onset of HF were available (n = 10,810, median age at onset of HF was 68 years [IQR 60-76]). To estimate the effect of HF-PRS on age at onset of HF, we constructed a linear regression model of HF-PRS and age at onset of HF, including sex and the top 10 principal components.
[0064] Survival analysis: The association between HF-PRS and long-term mortality was assessed using a Cox proportional hazards model. Survival follow-up data for 132,737 individuals were obtained from the BBJ dataset using ICD-10 codes. Causes of death were classified into two categories according to ICD-10 codes: (i) cardiovascular death (I00-99) and (ii) heart failure death (I50). The median follow-up period was 8.4 years (IQR 6.8-9.9). The Cox proportional hazards model was adjusted for sex, age, the top 10 principal components, and disease symptoms. Analyses were performed using the R package survival v.2.44, and survival curves were estimated using a modified version of the R package survminer v.0.4.6.
[0065] Results: Five novel loci for heart failure identified in a Japanese genome-wide association study. An overview of the study design is shown in Figure 1. We conducted a Japanese case-control GWAS to examine the association between up to 7,974,473 common genetic polymorphisms (minor allele frequency (MAF) > 1%) present on autosomes and the risk of four HF phenotypes. A total of 16,251 HF cases were identified, including 4,254 HFrEF, 7,154 HFpEF, and 11,122 NIHF cases (Figure 1). All-cause HF, HFrEF, and HFpEF cases were compared with 197,577 controls. For NIHF, CAD cases were excluded from the control sample, leaving 171,995 controls. The GWAS identified 18 genome-wide significant loci, five of which have not been previously reported (Figure 2, Table 1). The proportion of HF phenotype variation (single nucleotide polymorphism (SNP) heritability; h 2Using linkage disequilibrium (LD)-score regression (LDSC), the prevalence of heart failure was estimated to be 1.5% (standard error of the mean 0.3%) in all-cause heart failure, 1.5% (0.3%) in HFrEF, 0.5% (0.3%) in HFpEF, and 1.1% (0.3%) in NIHF using LDSC. 2 The prevalence was estimated to be 4.8% (SD 1.0%) for all-cause heart failure, 12.6% (2.9%) for HFrEF, 2.8% (1.7%) for HFpEF, and 3.1% (1.0%) for NIHF, suggesting heterogeneous heritability within each HF phenotype.
[0066]
[0067] Table 1 continued
[0068] Eight novel loci for heart failure identified in a cross-ethnic meta-analysis To improve statistical power, we performed a cross-ethnic meta-analysis combining the current Japanese GWAS (BBJ) with all-cause heart failure GWASs in Europeans, Chinese, Americans, and Africans, the European HFrEF and HFpEF GWASs, and the NIHF GWASs from UK biobank and FinnGen. Combined, these datasets included 120,203 cases of all-cause HF (BBJ: 16,251, EUR: 95,524, KADOORIE: 1,467, AMR: 1,170, AFR: 5,791), 23,749 cases of HFrEF (BBJ: 4,254, EUR: 19,495), 26,743 cases of HFpEF (BBJ: 7,154, EUR: 19,589), and 22,864 cases of NIHF (BBJ: 11,122, UKBB: 1,816, FinnGen: 9,926). The control sample sizes were 1,502,645 for all-cause heart failure (BBJ: 197,577, EUR: 1,270,968, KADOORIE: 75,149, AMR: 13,217, AFR: 20,883), and 456,520 for HFrEF / HFpEF (BBJ: 197,577, EUR: 258,943) and 863,254 for HFrEF / HFpEF (BBJ: 171,995, UKBB: 387,652, FinnGen: 303,607). A total of 11,119,536 variants in all-cause HF, 4,800,324 in HFrEF, 4,807,066 in HFpEF, and 4,998,828 in NIHF, including variants with a MAF of 1% or greater, were examined. These analyses identified 58 HF-associated loci with genome-wide significance (P < 5.0 × 10 -8 ; Table 2). Eight of these loci have not been reported previously, including one novel locus (CACNA1D) detected in this Japanese GWAS (Figure 2). A total of 12 novel loci were identified through this Japanese GWAS and cross-ethnic meta-analysis.
[0069]
[0070] Table 2 continued
[0071] Table 2 continued
[0072] Table 2 continued
[0073] Multitrait analysis identified eight novel loci for heart failure. To further increase the statistical power of the HF GWAS, we performed multi-trait analysis (MTAG). First, we evaluated the genetic correlations between HF and left ventricular, left atrial, right ventricular, and right atrial parameters, as well as left ventricular mass and fibrosis. Because left ventricular mass was strongly correlated with all HF phenotypes, we integrated the HF GWAS and left ventricular mass GWAS using MTAG. As a result, we found eight additional novel loci (Figure 2, Table 3).
[0074]
[0075] Table 3 continued
[0076] Table 3 continued
[0077] Derivation of PRS and correlation of PRS with long-term mortality. The case-control sample was divided into a PRS derivation dataset, a linear combination dataset, a test dataset, and a survival analysis dataset. Using the PRS-CS, PRSs were derived for all-cause BBJ, all-cause HF in Europeans, Americans, and Africans, and HFrEF, HFpEF, CAD, AF, and left ventricular mass in Europeans. Linear combinations of these PRSs were performed using Lasso regression with 10-fold cross-validation, which excluded all-cause HF in Americans and NIHF in Europeans. Consistent with population specificity, PRSs derived from BBJ showed better performance (pseudo-R ) than PRSs derived from Europeans, despite the smaller sample size. 2 (Figure 3). As previously suggested, the cross-ethnic PRS performed better than the single-population PRS (Figure 3). Combining HF-related phenotypes further improved performance (Figure 3). In the following, we used the best model (Model J in Figure 3).
[0078] The impact of HF-PRS on HF phenotype and prognosis. To evaluate the potential clinical application of the PRS, we examined the association between the PRS and age at HF onset using the BBJ case sample (n = 10,810). We observed a decrease in age at onset as the PRS increased, and estimated that age at HF onset was approximately 2 years younger in the upper tertile of PRS compared with the lower tertile (Figure 4). To further explore the clinical utility of the HF-PRS, we evaluated the impact of the PRS on mortality using long-term follow-up data from the BBJ. Cox regression analysis showed that higher PRS scores were associated with higher cardiovascular and HF mortality (Figure 5).
[0079] The important role of TTN genetic polymorphisms common in Japanese people in the prognosis of HF. A Japanese GWAS and cross-ethnic meta-analysis identified lead variants associated with the cardiomyopathy genes TTN, NEBL, and BAG3. Considering the nature of cardiomyopathy genes, these variants may affect cardiac function and patient prognosis. Among these, TTN lead variants were significantly more prevalent in East Asians than in Europeans, according to gnomAD (Figure 6a). We first examined the association between these variants and cardiac function, and found that the TTN lead variant (rs1484116) and NEBL lead variant (rs7075837) were significantly associated with reduced left ventricular ejection fraction (TTN: effect size -1.16, 95% confidence interval (95% CI) -1.55 to -0.77, P = 7.80 × 10). -9 NEBL: effect size -0.65, 95% CI -1.03 to -0.26, P = 9.15 × 10 -4 ; Fig. 6b). Next, we repeated the same analysis in NIHF to exclude the influence of CAD, and similar results were obtained (Fig. 6b). We further evaluated the impact of these variants on long-term HF mortality in non-HF and HF subjects. Cox regression analysis revealed that the common Japanese variant rs1484116 (TTN) was significantly associated with an increased risk of HF mortality in HF patients (HR 1.24, 95% CI 1.06-1.46, P = 8.12 × 10-3; Fig. 6c and d). These results indicate that the TTN genetic polymorphism independently determines the severity of HF in Japanese HF patients and can be used as a prognostic predictor.
Claims
1. A method for determining a subject's risk of developing heart failure, comprising the steps of: preparing a data list including effect alleles and effect sizes of genetic polymorphisms; obtaining the subject's genetic information; calculating a heart failure risk score from the subject's genetic information based on the information on the effect alleles and effect sizes of genetic polymorphisms in the data list; and determining the risk of heart failure based on the risk score, wherein the data lists include a data list of genomic analysis results for heart failure in Caucasian populations, a data list of genomic analysis results for heart failure in populations other than Caucasian, a data list of genomic analysis results for heart failure subtypes, a data list of genomic analysis results for diseases related to heart failure, and a data list of genomic analysis results for measurements related to cardiac function, and the risk score is a risk score obtained by combining the risk scores calculated from these data lists.
2. The method of claim 1, wherein the non-Western population is an Asian population.
3. The method of claim 1, wherein the heart failure subtype is selected from heart failure with reduced left ventricular ejection fraction (HFrEF), heart failure with preserved left ventricular ejection fraction (HFpEF), and non-ischemic heart failure (NIHF).
4. The method of claim 1, wherein the disease associated with heart failure is selected from coronary artery disease and atrial fibrillation.
5. The method of claim 1, wherein said cardiac function related measurement is left ventricular mass.
6. The method of claim 1, wherein the risk score is calculated by linearly combining a risk score calculated from the results of genomic analysis of heart failure in the Caucasian population, a risk score calculated from the results of genomic analysis of heart failure in a population other than the Caucasian population, a risk score calculated from the results of genomic analysis of the heart failure subtype, a risk score calculated from the results of genomic analysis of a disease related to heart failure, and a risk score calculated from the results of genomic analysis of a measurement value related to cardiac function.
7. The method of claim 6, wherein the risk score is calculated using a Bayesian method (PRS-CS).
8. A system for determining a subject's risk of heart failure, comprising: means for storing data including a data list containing effect alleles and effect sizes of genetic polymorphisms; means for receiving the subject's genetic information; means for calculating a heart failure risk score from the subject's genetic information based on the data on the effect alleles and effect sizes of genetic polymorphisms in the data list; and means for presenting the risk of heart failure based on the risk score, wherein the data list includes genomic analysis results for heart failure in Caucasian populations, genomic analysis results for heart failure in populations other than Caucasian, genomic analysis results for heart failure subtypes, genomic analysis results for measurements related to cardiac function, and genomic analysis results for diseases related to heart failure, and the risk score is a risk score obtained by combining the risk scores calculated from these data lists.
9. A method for assessing the risk of heart failure, comprising the step of analyzing a single nucleotide polymorphism (SNP) of rs1484116 (note that the rs number indicates the registration number in the dbSNP database of the National Center for Biotechnology Information) present in the TTN gene region.
10. The method of claim 9, wherein the heart failure includes heart failure with reduced ejection fraction (HFrEF), heart failure with preserved ejection fraction (HFpEF), and non-ischemic heart failure (NIHF).