Method for predicting healthy life expectancy, and biomarker

A polygenic score based on 53,760 SNVs from centenarian genome-wide analysis predicts healthy lifespan, addressing limitations in existing methods by providing accurate genetic insights for longevity and healthspan prediction.

JP2025107092APending Publication Date: 2025-07-17KEIO UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024000861
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-05
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Existing methods for predicting healthy lifespan are limited in accuracy and scope, failing to effectively utilize the genetic factors associated with longevity and healthspan, particularly in populations like centenarians.

Method used

A prediction method using a polygenic score calculated from 53,760 single nucleotide polymorphisms (SNVs) with a P-value of less than 0.01, derived from genome-wide association analysis of centenarians, to evaluate healthy life expectancy, incorporating additional phenotypic data and biomarkers for a comprehensive assessment.

Benefits of technology

The method provides a robust prediction of healthy life expectancy by leveraging genetic factors, enhancing the accuracy of lifespan prediction and identifying biomarkers associated with centenarians, thereby facilitating insights into healthspan and potential interventions for healthy aging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025107092000001_ABST
    Figure 2025107092000001_ABST
Patent Text Reader

Abstract

To provide a prediction method and a biomarker which enable prediction of healthy life expectancy.SOLUTION: The present invention provides a prediction method for predicting healthy life expectancy. This method includes: a step in which a polygenic score of a subject is calculated, using the following formula, based on 53,760 single nucleotide variants (SNVs) associated with centenarians at a P-value of less than 0.01 by a genome-wide association analysis for centenarians; and a step in which healthy life expectancy is evaluated based on the calculated polygenic score of the subject. The number n of SNVs used for calculation of the polygenic score is 100 or more. In the formula, i is a positive integer from 1 to n; n is the number of SNVs used for calculation of the polygenic score; SNVi is the genotype at the i-th SNV; and βi is the β value by the genome-wide association analysis for the i-th SNV.SELECTED DRAWING: Figure 21
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a prediction method for predicting a healthy life span and a biomarker.

Background Art

[0002] Human longevity and healthy aging are complex phenotypes formed by multiple factors such as genetic factors, lifestyle, and social environment, not just the absence of disease. Centenarians (persons aged 100 or older) and super centenarians (persons aged 110 or older) generally avoid or delay the onset of age-related diseases, maintain physical and cognitive independence even at extremely old ages, preserve organ functions such as those of the heart, kidneys, and liver, and have genetic factors. Therefore, they are excellent models for studying healthy longevity. Elucidating their genetic factors is important for understanding the human aging mechanism. Also, by understanding the biological mechanisms underlying the gene expression phenotypes of super centenarians, it may be possible to extend this characteristic to the general population.

[0003] Genome-wide association study (GWAS) is a powerful tool for identifying important genes and molecular signatures related to complex human traits, and the same applies to aging. Many GWAS, meta-GWAS, and multivariate GWAS have been reported for multiple age-related traits such as longevity of 90% or 99% survival in a population, healthy life span, parental life span, and multivariate analysis of these aging factors. Among these, several loci have been identified by multivariate GWAS targeting healthy life span, parental life span, etc. However, for longevity GWAS, only the APOE locus commonly exceeds the GWAS significance level, suggesting that these age-related traits share a genetic structure but are not identical (see, for example, Non-Patent Document 1).

Prior Art Documents

Non-Patent Documents

[0004] [Non-Patent Document 1] Deelen, J. et al. A meta-analysis of genome-wide association studies identifies multiple longevity genes. Nat Commun 10, 3669 (2019). [Summary of the Invention] [Problems to be Solved by the Invention]

[0005] An object of the present invention is to solve the above-mentioned conventional problems and achieve the following object. That is, an object of the present invention is to provide a prediction method and a biomarker that can predict a healthy life span. [Means for Solving the Problems]

[0006] Means for solving the above problems are as follows. That is, <1> A prediction method for predicting a healthy life span, comprising: calculating a polygenic score of a subject by the following formula (1) based on 53,760 single nucleotide polymorphisms (SNVs) related to centenarians with a P-value of less than 0.01 by genome-wide association analysis for centenarians; [Number] In the formula (1), i is a positive number from 1 to n, n is the number of SNVs used for calculating the polygenic score, and SNV i represents the genotype at the i-th SNV, and β i is the β value by the genome-wide association analysis of the i-th SNV. evaluating the healthy life span based on the calculated polygenic score of the subject, and the 53,760 SNVs are the SNVs shown in Tables 4 to 451 of this specification, and a prediction method characterized in that the number n of SNVs used for calculating the polygenic score is 100 or more. <2> The prediction method according to <1>, wherein the number n of SNVs used for calculating the polygenic score is 1,000 or more. <3> The prediction method according to <1>, wherein the number n of SNVs used for calculating the polygenic score is 10,000 or more. <4> The prediction method according to any one of <1> to <3>, wherein the SNV used for calculating the polygenic score includes at least any one of APOE4 represented by rs429358, APOE2 represented by rs7412, GRM7 represented by rs73116078, and EYS represented by rs75571981. <5> The prediction method according to any one of <1> to <4>, wherein the subject is of Asian ethnicity. <6> The prediction method according to any one of <1> to <5>, wherein the subject is Japanese. <7> The prediction method according to any one of <1> to <6>, further comprising a step of evaluating at least any one of phenotypic data, biomarkers, multi - layer omics, and analysis results of biological samples for the subject. <8> The prediction method according to <7>, further comprising a step of evaluating at least any one of the analysis results of the polygenic risk score (PRS) of aspartate aminotransferase (AST) and the polygenic risk score (PRS) of fibrinogen (Fbg) for the subject. <9> A biomarker for predicting healthspan, It is a single - nucleotide polymorphism (SNV) with a β - value of 0.21 or more obtained by genome - wide association analysis for centenarians and showing a positive correlation with centenarians. <10> A biomarker for predicting healthspan, It is a single - nucleotide polymorphism (SNV) with a β - value of - 0.20 or less obtained by genome - wide association analysis for centenarians and showing a negative correlation with centenarians.

Advantages of the Invention

[0007] According to the present invention, the above-described various problems in the prior art can be solved, the above object can be achieved, and a prediction method and a biomarker capable of predicting a healthy life span can be provided.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34

Modes for Carrying Out the Invention

[0009] (Prediction method) The prediction method of this embodiment is a prediction method for predicting a healthy life span, and is based on 53,760 single nucleotide polymorphisms (SNVs) related to centenarians with a P-value of less than 0.01 by genome-wide association analysis for centenarians. According to the following formula (1), a step of calculating a polygenic score of a subject, and a step of evaluating a healthy life span based on the calculated polygenic score of the subject are included, and further includes other steps as necessary. [Number] In the formula (1), i is a positive number from 1 to n, n is the number of SNVs used for calculating the polygenic score, and SNV i represents the genotype at the i-th SNV, and β i is the β value by the genome-wide association analysis of the i-th SNV. The 53,760 SNVs are the SNVs shown in Tables 4 to 451 of this specification, and the number n of SNVs used for calculating the polygenic score is 100 or more.

[0010] (Biomarker) The biomarker of this embodiment is a biomarker for predicting a healthy life span, and (1) a single nucleotide polymorphism (SNV) in which the β value by genome-wide association analysis for centenarians is 0.21 or more and shows a positive correlation with centenarians; or (2) a single nucleotide polymorphism (SNV) in which the β value by genome-wide association analysis for centenarians is -0.20 or less and shows a negative correlation with centenarians. The biomarker may be a single type alone or a combination of two or more types.

[0011] In GWAS, generally, in order to accurately obtain the P-values of loci with small effect sizes, 100,000 or more cases and controls are required. However, to improve longevity GWAS, it is necessary to increase the sample size or adopt a population with a high genetic factor. Against this background, in order to clarify the genetic structure of extreme human longevity, it is inevitable to emphasize the analysis of centenarians, especially supercentenarians, who are of advanced age. Furthermore, due to recent advances in analysis techniques, it has become possible to analyze the correlation between GWAS summary statistics and a polygenic risk score (PRS), which is a quantitative score of an individual's genetic factors based on the GWAS summary statistics. In fact, these techniques have begun to be used in aging research as well.

[0012] As a result of intensive studies to solve the above object, the present inventors have found the following findings and completed the present invention. That is, as shown in the examples described later, a genome-wide association study (GWAS) was conducted on 964 Japanese centenarians (including 173 supercentenarians and 7,306 controls) to search for phenotypes correlated with the genetic components of longevity. As a result of the GWAS summary statistical analysis of Japanese centenarians, the heritability based on SNPs (single nucleotide polymorphisms) was 0.204, and it was shown that the genetic components were shared with those of a low risk of multiple age-related diseases and biomarkers. As a result of the survival analysis, it was shown that the polygenic score based on centenarian analysis (also referred to as Centenarians-based Polygenic score; Cent.PGS) calculated from the GWAS summary statistics was associated with the healthy life expectancy in a healthy elderly cohort independently of the APOE4 genotype and sex. This suggests that Cent.PGS is a novel polygenic score related to life expectancy and healthy life expectancy in late-life elderly people, and based on this polygenic score Cent.PGS, it was found that the healthy life expectancy can be predicted, leading to the completion of the present invention. In addition, it has been found that single nucleotide polymorphisms (SNVs) that are positively strongly associated with centenarians and single nucleotide polymorphisms (SNVs) that are negatively strongly associated with centenarians can be used as biomarkers for predicting healthspan, leading to the completion of the present invention.

[0013] Among 964 Japanese centenarians, 35.2% of the centenarians are survivors of a group at the 99.99% tile or higher (1 in 10,000 people), and 77.7% of the centenarians are survivors of a group at the 99.9th percentile or higher (1 in 1,000 people). Comparing with previous longevity GWAS studies, this analysis shows that it is the best longevity GWAS.

[0014] To clarify the genetic structure of human ultimate centenarians, GWAS summary statistics and polygenic risk score analysis were applied to centenarians at the 99.99% tile or higher among 964 Japanese centenarians, and the association with the observed scores including lifespan was analyzed. Furthermore, since the differences in longevity genetic components are formed by the age of 100, and the phenotypes correlated with longevity genetic components are mainly observed in this age group, healthy individuals aged 85 to 89 were also included in the analysis. These analyses in both centenarians and healthy elderly individuals clarify the genetic structure of human ultimate centenarians. According to the prediction method and biomarkers of the embodiment, the healthspan of an individual in the general population can be predicted, and by evaluating in combination with other aging-related traits, more insights into healthspan can be obtained.

[0015] Hereinafter, the prediction method of the present embodiment will be described in detail. The prediction method of the present embodiment is a prediction method for predicting healthspan, including a calculation step and an evaluation step, and preferably further includes an additional evaluation step, and further includes other steps such as a genomic DNA extraction step, a genomic sequencing step, and a genotyping step as needed.

[0016] Here, "healthy life expectancy" refers to the period during which a person can live without their daily life being restricted due to health problems. For example, it is the period during which a person maintains their independence. Specifically, in terms of the level of care in Japan, it includes the period of "preventive care level 1" or the period during which a person maintains their independence without the need for support.

[0017] (Calculation step) The above-mentioned calculation step is a step of calculating the polygenic score of the subject according to the following formula (1) based on 53,760 single nucleotide polymorphisms (SNVs) related to centenarians with a P-value of less than 0.01 by genome-wide association study (GWAS) for centenarians. The polygenic score based on the centenarian GWAS analysis may be referred to as "Cent.PGS" (Centenarians-based Polygenic score).

[0018] [Number] In the above formula (1), i is a positive number from 1 to n, n is the number of SNVs used for calculating the polygenic score, and SNV i represents the genotype at the i-th SNV, and β i is the β value of the i-th SNV.

[0019] The single nucleotide variant (SNV) is a substitution mutation of a single base in a DNA sequence. On the other hand, a single nucleotide polymorphism (SNP) is a type of SNV, which means a mutation caused by a substitution mutation of a single base in the genomic base sequence of a certain biological species population. Sometimes, an SNV with a mutation frequency of 1% or more in the population is referred to as an SNP. In this embodiment, the single nucleotide polymorphism (SNV) is a SNV specified by a "single base" mutation and specified by an rs number. In addition, the 53,760 SNVs of the present embodiment specifically refer to the 53,760 SNVs shown in Tables 4 to 451 listed at the end of this specification, which were identified as being related to centenarians with a P-value of less than 0.01 through a genome-wide association analysis of centenarians in the present example described below.

[0020] Here, the "rs number" is a number (Reference SNP ID number) defined by the National Center for Biotechnology Information (NCBI) for each SNP and is uniformly defined worldwide. The "rs number" can be searched for in NCBI's dbSNP (https: / / www.ncbi.nlm.nih.gov / snp / ), jMorp (https: / / jmorp.megabank.tohoku.ac.jp / 201909 / ), etc.

[0021] As the number n of SNVs used for calculating the polygenic score, if it is 100 or more, compared to the evaluation of aging or healthy lifespan based on APOE2 or APOE4 of the known APOE gene, an evaluation can be performed using a large genetic component or different aspects of genetic components, and it can be appropriately selected according to the purpose. However, in terms of being able to perform a more comprehensive evaluation, any one of 250 or more, 500 or more, 750 or more, 1,000 or more, 2,500 or more, 5,000 or more, 7,500 or more, 10,000 or more, 15,000 or more, and 20,000 or more is more preferable.

[0022] For the number n of SNVs, any one can be selected from the 53,760 SNVs shown in Tables 4 to 451 listed at the end of this specification, but it is preferably included at least one of APOE4 represented by rs429358, APOE2 represented by rs7412, GRM7 represented by rs73116078, and EYS represented by rs75571981.

[0023] -APOE- The APOE gene is the abbreviation of apolipoprotein E, which functions as a ligand of the LDL receptor family present on the cell surface and is considered to play an important role in the uptake of lipoproteins into cells and their removal from plasma. Polymorphisms involving amino acid sequence substitutions are known in APOE. Among them, APOE4 with the 158th Arg substituted by Cys is the largest genetic factor in Alzheimer's disease, and it is also known that the carrier rate of APOE4 is low among centenarians.

[0024] -GRM7- GRM7 is the gene symbol for metabotropic glutamate receptor 7 and is the causative gene for neurodevelopmental disorders (NEDSHBA) accompanied by seizures, reduced muscle tone, and brain imaging abnormalities. Furthermore, since the deficiency of metabotropic glutamate receptors in Drosophila causes age-related sleep disorders and short lifespan, it is suggested that the GRM7 locus may affect human lifespan and healthspan through brain dysfunction.

[0025] -EYS- EYS is the symbol gene of eyes shut homolog that encodes an ortholog of the spacemaker in Drosophila and is mutated in autosomal recessive retinitis pigmentosa. Although the association between EYS and lifespan has not been reported, the EYS locus is highly likely to be associated with visual function even in the elderly.

[0026] In another aspect, the n SNVs are (1) One or more single nucleotide polymorphisms (SNVs) with a β value of 0.21 or more by genome-wide association analysis of centenarians and showing a positive correlation with centenarians; and / or (2) One or more single nucleotide polymorphisms (SNVs) with a β value of -0.20 or less by genome-wide association analysis of centenarians and showing a negative correlation with centenarians; preferably include.

[0027] In the formula (1), i is a positive number from 1 to n, n is the number of SNVs used for calculating the polygenic score, and SNV irepresents the genotype at the i-th SNV, and β i is the β value of the i-th SNV.

[0028] Here, for "SNV" i the genotype at the i-th SNV indicates whether the SNV is the reference allele (corresponding to the "A2" column in the table) or the alternative allele (also referred to as the minor allele; corresponding to the "A1" column in the table). When it is the reference allele, "zero" is substituted into "SNV" in the formula (1), and when it is the alternative allele, "1" is substituted. i This means substituting into the "SNV" in the formula (1). For example, GRM7 represented by rs73116078 indicates an SNV present in the intron of the GRM7 gene (reference allele G > alternative allele A). As shown in No. 8545 of Table 75, the A1 column indicates the alternative allele (adenine A), and the A2 column indicates the reference allele (guanine G). The β value is -0.0923 and the P value is 1.59E - 07.

[0029] There is no particular restriction on the individual as the target for predicting healthy life expectancy, and it can be appropriately selected according to the purpose. It may be an individual in the European - Asian population, preferably an Asian person, and more preferably a Japanese person.

[0030] <Evaluation step> The evaluation step is a step of evaluating the healthy life expectancy based on the calculated polygenic score (Cent.PGS) of the subject. Here, as a method of evaluating the healthy life expectancy based on the Cent.PGS, for example, (1) When the Cent.PGS is greater than zero or equal to or greater than a positive threshold, it is evaluated that the healthy life expectancy based on the genetic component may be long; (2) When the Cent.PGS is greater than zero or greater than a negative threshold and less than a positive threshold, it is evaluated that the healthy life expectancy based on the genetic component is equivalent to the control; (3) When the Cent.PGS is greater than zero or greater than or equal to a negative threshold, it can be evaluated that there is a risk of a short health span based on genetic components; a method can be mentioned. Also, as another aspect, a method of evaluating the degree of possibility / risk of a health span according to the magnitude of the value of the Cent.PGS can be mentioned.

[0031] The positive threshold is not particularly limited and can be appropriately selected according to the purpose. For example, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.5, 2.0, 2.5, 3.0, etc. can be mentioned. The negative threshold is not particularly limited and can be appropriately selected according to the purpose. For example, -0.1, -0.2, -0.3, -0.4, -0.5, -0.6, -0.7, -0.8, -0.9, -1.0, etc. can be mentioned.

[0032] <Additional evaluation step> The prediction method of this embodiment preferably further includes an additional evaluation step. The additional evaluation step is a step of evaluating the analysis results of at least any one of phenotypic data, biomarkers, multi - layer omics, and biological samples for the subject, and there is no particular limitation. Known aging - related analysis results can be appropriately selected according to the purpose. Thereby, it can be evaluated in combination with other aging - related traits and analyses, and more findings regarding the health span can be obtained.

[0033] Examples of the phenotypic data include investigations such as physical function, cognitive function, and blood tests; questionnaires and data such as medical history, social history, lifespan, health span, motor function analysis, and head MRI.

[0034] Examples of the biomarkers include plasma biomarkers and urine biomarkers related to the circulatory system, cognitive function, internal organs, inflammation, telomeres, etc.; biomarkers related to genomic DNA such as DNA methylation and somatic mutations.

[0035] Examples of multi-omics include genomic analysis such as genomic DNA (gDNA), other genome-wide association studies (GWAS), polygenic risk scores (PRS), etc.; transcriptome analysis such as blood cells, differential gene expression analysis (DGE), etc.; single-cell analysis; microbiota analysis, and the like.

[0036] Examples of the biological sample include, for example, blood, plasma, cells, mRNA, gDNA, urine, autopsy samples, and the like.

[0037] <Other processes> The prediction method may further include other processes such as a genomic DNA extraction process, a genome sequencing process, a genotyping process, etc. as required. Alternatively, at least any one of the genomic DNA extracted from the subject, the genomic DNA sequence of the subject, and the determined genotype of the subject may be used.

[0038] <<Genomic DNA extraction process>> The genomic DNA extraction process is a process of extracting genomic DNA from the subject. The method for extracting the genomic DNA is not particularly limited, and a known method can be appropriately selected according to the purpose. Examples thereof include a method for extracting genomic DNA from a sample such as whole blood.

[0039] <<Genome sequencing process>> The genome sequencing process is a process of determining the sequence of the genomic DNA extracted from the subject. As a method for determining the sequence of the genomic DNA, there are no particular limitations, and known methods can be appropriately selected according to the purpose. For example, a method for determining the whole-genome DNA sequence using a sequencing system such as HiSeq2500, HiSeqX, NovaSeq 6000, etc.; in addition to the above method, a method for performing re-sequencing analysis by adopting a workflow known as the GATK Best Practices workflow, which is becoming the global standard procedure for whole-genome sequencing analysis; a method for applying base quality score recalibration (BQSR), which is effective in reducing sequencer-specific biases, in addition to the re-sequencing analysis, etc. can be mentioned.

[0040] <<Genotype determination step>> The genotype determination step is a step of determining a genotype based on the genomic sequence of the subject. As a method for determining the sequence of the genomic DNA, there are no particular limitations, and known methods can be appropriately selected according to the purpose. For example, a method for determining the genotype of a target SNV using a gene analysis tool such as Axiom Japonica Array NEO, Infinium Asian Screening Array-24 v1.0 BeadChip kit, etc. can be mentioned.

[0041] (Biomarker) The biomarker of the present embodiment will be described below. The biomarker of the present embodiment is a biomarker for predicting a healthy lifespan, and (1) It is a single nucleotide polymorphism (SNV) with a β value of 0.21 or more by genome-wide association analysis of centenarians and showing a positive correlation with centenarians; or (2) It is a single nucleotide polymorphism (SNV) with a β value of -0.20 or less by genome-wide association analysis of centenarians and showing a negative correlation with centenarians. The biomarker may be a single type alone or two or more types may be used in combination.

[0042] In the genome-wide association analysis of the above-mentioned (1) centenarians, the β value is 0.21 or more, and the single nucleotide polymorphisms (SNVs) showing a positive correlation with centenarians specifically include 100 SNVs shown in Table 1 below.

[0043]

Table 1

[0044] In the genome-wide association analysis of the above-mentioned (2) centenarians, the β value is -0.20 or less, and the single nucleotide polymorphisms (SNVs) showing a negative correlation with centenarians specifically include 100 SNVs shown in Table 2 below.

[0045]

Table 2

[0046] According to the biomarker in this embodiment, it can be used as a biomarker for predicting a healthy life span and can also be preferably used as a single nucleotide polymorphism (SNV) of the prediction method of this embodiment.

Example

[0047] Hereinafter, the present invention will be described more specifically based on examples, but the present invention is not limited to the following examples.

[0048] [Method] <Research Group of Japanese Centenarians (Cent) and Japanese Healthy Longevity Persons (HA)> As Japanese centenarians, data from two prospective cohort studies of Japanese elderly people, the Tokyo Centenarian Study (TCS) and the Japanese Semi-Supercentenarian Study (JSS), were used. All centenarian data were frozen in May 2022. Among 967 centenarians with genomic DNA sequences (hereinafter sometimes referred to as gDNA sequences), 964 were registered as centenarians (hereinafter referred to as Cent) after excluding 3 centenarians with inconsistent gender information and 3 centenarians regarded as close relatives with a PI HAT of 0.1875 or more: 144 males and 820 females (female rate: 0.851), median age 106.0 [interquartile range (IQR): 103.9 - 107.1], number of deaths 915 (mortality rate: 0.969), median age at death 107.5 [IQR: 105.7 - 109.2] (Table 3).

[0049]

Table 3

[0050] The Kawasaki Aging and Wellbeing Project (KAWP) is a community-based prospective cohort study targeting people aged 85 - 89 years who maintained independence (i.e., "prevention of care level 1" or no need for support in the Japanese care level) as representatives of healthy elderly people with a long healthy lifespan during the registration review from March 2017 to December 2019. All KAWP data were frozen in September 2022. Among 1,026 people who participated in KAWP, 2 were excluded because permission for gDNA sequencing could not be obtained, and 8 were excluded because there were close relatives with a PI HAT of 0.1875 or more. Therefore, 1,016 people were registered as healthy longevous people (hereinafter referred to as HA): 513 males and 503 females (female rate: 0.495), median age 86.8 [IQR: 85.9 - 88.2], 168 deaths (mortality rate: 0.169), median age at death 90.4 [IQR: 89.0 - 91.8].

[0051] <Research Group of Japanese Controls (Cont)> The Tohoku Medical Megabank (TMM) Project is a population-based prospective cohort study that includes an adult cohort, a birth cohort, and a three-generation cohort based on the general population in the Tohoku region. Of the 30,081 participants, 11,519 from the three-generation cohort were excluded because they had a PI-HAT of 0.1875 or higher with a close relative, 1,237 were excluded due to lack of phenotype, and 9 were excluded due to PCA outliers from the Japanese cluster. Therefore, 7,306 individuals were registered as controls (hereinafter referred to as "Cont"): 2,578 males and 4,728 females (female ratio: 0.647), with a median age of 45 [IQR: 32 - 63] (Table 3). The TMM Project is entirely operated by the Tohoku Medical Megabank Organization (ToMMo) of Tohoku University. The allele frequency data of ToMMo38k were downloaded from jMorp (https: / / jmorp.megabank.tohoku.ac.jp; Known reference 1: Tadaka, S. et al., Nucleic Acids Res 49, D536 - D544 (2021).).

[0052] <Genomic DNA extraction> For Cent and KAWP, genomic DNA was extracted from whole blood using the FlexGene DNA Kit (Qiagen, Hilden, Germany). For Cont, genomic DNA was extracted from whole blood by the method described above.

[0053] <Whole-genome DNA sequencing> For Cent and KAWP, the whole-genome DNA sequences (hereinafter sometimes referred to as "WGS") of 526 centenarians were determined using whole-genome DNA sequencing with HiSeq2500, HiSeqX, or NovaSeq 6000 according to the protocol described above. For Cont, the whole-genome DNA sequences of 3,340 Japanese controls were determined by the whole-genome DNA sequencing method using HiSeq2500 or NovaSeq 6000 as described above.

[0054] <Re-sequence analysis of whole-genome sequences> The re-sequence analysis of the sequences obtained by the sequencer was performed with some modifications to the description in the known document 1: Tadaka et al., Nucleic Acids Res 49, D536-D544 (2021). Briefly, the workflow known as the GATK Best Practices workflow, which is becoming the global standard procedure for whole-genome sequence analysis, was adopted. And base quality score recalibration (BQSR), which is effective in reducing sequencer-specific biases, was applied.

[0055] <Genotyping and Imputation Using DNA Microarray> Using the Axiom Japonica Array NEO, genotypes of 650,000 SNVs of 441 centenarians and 26,741 Japanese controls were determined according to the manufacturer's protocol. Genotypes of 650,000 SNVs of 1,015 people in KAWP were determined using the Infinium Asian Screening Array-24 v1.0 BeadChip kit according to the manufacturer's protocol. All DNA microarray scan images were analyzed using the protocol described above.

[0056] <Principal Component Analysis (PCA)> To identify outliers in the Japanese population, all samples including Cent, HA, and Cont were clustered for each gDNA sequence using PCA with 2,504 gDNA sequences determined by the 1000 Genomes (known document 2: Fairley, S. et al., Nucleic Acids Res 48, D941-D947 (2020)). Common SNVs between our samples and the 1000 Genomes samples were extracted, and SNVs were trimmed using PLINK (version 1.90; known document 3: Purcell, S. et al., Am J Hum Genet 81, 559-75 (2007)). Finally, PC 1-20 were calculated using the pca command of PLINK 1.90.

[0057] <GWAS and Meta-GWAS Analysis> To identify SNVs related to Japanese centenarians, the gDNA sequences of Cent and Cont were compared by GWAS. WGS and DNA microarray were separately subjected to GWAS analysis using PLINK 1.90 and integrated as a meta-GWAS using N-weighted multivariate GWAMA (Reference 4: Baselmans, B.M.L. et al., Nat Genet 51, 445-451 (2019)). In N-weighted multivariate GWAMA, summary statistics of 5.98 million SNVs commonly found between the WGS samples and the DNA microarray input samples were extracted. The cross-trait intercept between the WGS samples and the DNA microarray input samples was 0.0606, and the SNP heritabilities of the WGS samples and the DNA microarray input samples were 0.3211 and 0.2757, respectively. Manhattan plots were created using the "qqman" package (version 0.1.8) of the program R (Reference 5: Turner, S.D., Journal of Open Source Software 3, 1-2 (2018)). An enlarged view of the Manhattan plot including recombination rate information was created using LocusZoom (version 1.3) (Reference 6: Pruim, R.J. et al., Bioinformatics 26, 2336-7 (2010)).

[0058] <Calculation of SNPh2, λ using LDSC and GWAS summary statistics GC , genetic correlation> SNP heritability, λ GCThe genetic correlation was calculated using LDSC (version 1.0.1; reference 7: Bulik-Sullivan, B. et al., Nat Genet 47, 1236-41 (2015)). For the GWAS summary statistics related to longevity, the European 90 / 99 survival percentiles and parental lifespan (PLS, reference 8: Timmers, P.R. et al., Elife 8(2019)) were used. The Japanese disease and quantitative GWAS summary statistics were downloaded from the RIKEN JENGER server (http: / / jenger.riken.jp; reference 9: Ishigaki, K. et al., Nat Genet 52, 669-679 (2020)), and the Japanese GWAS summary statistics related to Alzheimer's disease were downloaded from NBDC (https: / / humandbs.biosciencedbc.jp / , Research ID: hum0237.v1; reference 10: Shigemizu, D. et al., Transl Psychiatry 11, 151 (2021)).

[0059] <Transcriptional expression analysis by GTEx and promoter prediction analysis by Enformer> The SNV-related expressed genes of eQTLs, sQTLs, ieQTLs, and isQTLs were analyzed using the GTEx portal (V8, https: / / www.gtexportal.org / home / ). Enformer (reference 11: Avsec, Z. et al. Nat Methods 18, 1196-1203 (2021)) was used for promoter prediction.

[0060] <Generalized gene-based test for GWAS of Japanese centenarians by MAGMA> Using the FUMA web application (https: / / fuma.ctglab.nl / ; Known Document 12: Watanabe, K. et al., Nat Commun 8, 1826 (2017)), a gene-based test for the Japanese centenarian GWAS was performed. Independent significant SNPs in the GWAS summary statistics were identified based on the P-value (P < 1.0 × 10-2) within a 250 kb window and their mutual independence (r2 < 0.6 in the 1000 Genomes Phase 3 ALL of the reference panel population). The results of the MAGMA gene set analysis (Known Document 13: de Leeuw, C.A. et al., PLoS Comput Biol 11, e1004219 (2015)) were evaluated for enrichment analysis for both gene ontology and KEGG pathways using the "clusterProfiler" package in R. The MAGMA tissue expression analysis (Known Document 13: de Leeuw, C.A. et al., PLoS Comput Biol 11, e1004219 (2015)) was performed using eQTL data from the Genotype-Tissue Expression project (GTEx v8) (Known Document 14: Consortium, G.T., Nat Genet 45, 580-5 (2013)).

[0061] The gene lists related to aging of Longevity map (build 3), CellAge (build 3), and GenAge (build 21) were downloaded from the Human Ageing Genomic Resource (https: / / genomics.senescence.info; Known reference 15: Tacutu, R. et al., Nucleic Acids Res 46, D1083-D1090 (2018)). The gene list of longevity regulatory pathway genes including the IGF / insulin, sirtuin, AMP-AMPK, and TOR pathways (hsa04211) in the KEGG pathway database (https: / / www.genome.jp / ; Known reference 16: Kanehisa, M. et al., Nucleic Acids Res 51, D587-D592 (2023)) was used. For the type 2 diabetes risk loci in the Japanese population, the gene list of Known reference 17: Imamura et al., Nat Commun 7, 10531 (2016) was used. The gene symbols in the gene list were converted to Entrez ID using the clusterProfiler package in R, and the duplicate Entrez ID was counted in R.

[0062] <Heritability enrichment analysis for tissue-specific expression genes or enhancer regions using LDSC> The heritability enrichment analysis was performed according to the description of Cell type specific analyses on the github website for ldsc (https: / / github.com / bulik / ldsc / ; Known reference 18: Finucane, H.K. et al., Nat Genet 47, 1228-35 (2015)). The LD score data of the East Asian population containing 10 cell type group-specific annotations and 220 cell type-specific annotations were downloaded from the RIKEN JENGER server (http: / / jenger.riken.jp; Known reference 19: Kanai, M. et al, Nat Genet 50, 390-400 (2018)).

[0063] <PRS, Z-score of PRS, lsmean of PRS Z-score> Using PRS-CS (version 1.0.0; reference: Ge, T. et al., Nat Commun 10, 1776 (2019)), a polygenic risk score (PRS) was calculated using Japanese centenarians, Japanese diseases, and Japanese quantitative GWAS summary statistics. SNVs in the GWAS summary statistics were screened at p = 0.01 or greater, the phi parameter was 0.01, and the LD reference panel was the EAS reference panel constructed using the 1000 Genomes Project Phase 3 samples. Other parameters were set to default.

[0064] All PRSs were normalized to Z-scores using the mean and standard deviation of Cont.(PRS). The mean value of PRS for each group was calculated using least-squares means (lsmeans) adjusted for sex because the female ratios in the cohorts differed among Cent, HA, and Cont.

[0065] <Multiple logistic regression analysis and ROC analysis of PRS> In the multiple logistic regression analysis using genetic factors, 62 PRS scores, 4 SNV genotypes, and Cent.PGS (only for the logistic regression analysis between HA and Cont.(PRS)) were evaluated by univariate logistic regression adjusted for sex, and the least absolute shrinkage and selection operator (LASSO) was used to select covariates for the multiple logistic regression analysis among Cent, HA, and Cont.(PRS). After selecting the covariates, the difference in genetic factors and the contribution of PRS scores and SNV genotypes among Cent, HA, and Cont(PRS) were analyzed by multiple logistic regression analysis. ROC plots and AUC were calculated using the "pROC" package in R under default conditions.

[0066] <Kaplan-Meier and Cox regression, regularized multiple Cox regression survival analysis> In the univariate survival analysis, Kaplan-Meier analysis was performed using the "survival" package in R. In the multivariate survival analysis, Cox regression or regularized multiple Cox regression survival analysis was performed using the "survival" package in R. For the evaluation items of survival analysis, telephone surveys were used for the age at death of Cent, both telephone surveys and medical insurance claims were used for the age at death of HA, long-term care insurance claims were used for the age at the end of healthy life including the certified age of level 2 or higher long-term care, and medical insurance claims were used for the age at death of medical insurance claims.

[0067] <Baseline tests in Cent and HA> The methods for all observed scores in the baseline tests except for telomere length were previously described (Known literature 21: Sasaki, T. et al., Elife 12(2023).).

[0068] Briefly, both the Instrumental Activities of Daily Living (IADL) and the Activities of Daily Living (ADL) are questionnaires for evaluating the functions necessary for daily life.

[0069] TUG (Time Up and GO) is the time it takes for a person to stand up from a chair, walk 3 m, turn 180 degrees, return to the chair, and sit down while turning 180 degrees, and was used to evaluate motor ability.

[0070] The body mass index (BMI), systolic blood pressure (SBP), and diastolic blood pressure (DBP) were measured in the baseline health check. The Mini-Mental State Examination (MMSE; 0 - 30 points) is a questionnaire for evaluating cognitive function. The extended clinical dementia rating scale (ex.CDR) is a method for evaluating the late stage of dementia.

[0071] The geriatric depression scale (GDS) is a questionnaire for evaluating the depression level of the elderly.

[0072] The frailty index was calculated using the deficit accumulation model proposed by Rockwood (known reference 22: Rockwood, K. & Mitnitski, J Gerontol A Biol Sci Med Sci 62, 722-7 (2007).).

[0073] Blood biomarker concentrations such as NTproBNP, cystatin C, and interleukin-6 (IL-6) were measured by ELISA method. Blood tests for triglyceride (TG), high-density lipoprotein cholesterol (HDLC), low-density lipoprotein cholesterol (LDLC), cholinesterase (CHE), aspartate aminotransferase (AST), hemoglobin A1c (HbA1c), C-reactive protein (CRP), and albumin (ALB) were measured by a clinical laboratory company.

[0074] Medical histories were collected through interviews with physicians.

[0075] <Phenotype-genotype correlation analysis> To analyze the correlation between phenotype and genotype, all scores observed in the baseline tests were analyzed using a generalized linear model with age and gender at enrollment as moderator variables. All observed scores were standardized to compare the coefficients between the observed scores.

[0076] <Percentile analysis of Japanese centenarians> To correct the age distribution due to differences in the number of births in each year, the percentile corresponding to which percentile of the population born in the same year the age at death corresponded to was calculated in percentiles. Demographic data on the number of births and the population by age extracted from the census data of 2000, 2005, 2010, 2015, and 2020 were downloaded from the government statistics portal site e-Stat (https: / / www.e-stat.go.jp / en / ). Finally, for each birth year, ages at the 90th percentile, 99th percentile, 99.9th percentile, 99.9th percentile, and 99.9th percentile and above were calculated, and the boundaries of the 90th percentile, 99th percentile, 99.9th percentile, and 99.9th percentile were determined. Among Japanese centenarians, 77.7% (749 / 964 people) were at the 99.9th percentile and above.

[0077] <Statistical analysis> Baseline characteristics, biomarkers, and medical histories were expressed as median or numerical values with percentages or interquartile range (IQR). Differences in baseline data were evaluated using the Wilcoxon rank-sum test, chi-square test, and Fisher's Exact test. All statistical analyses were performed using the program R (version 4.2.2) and exactRankTests (wilcox.exact, Wilcoxon rank-sum test [version 0.8-31]), glmnet (LASSO and multivariate analysis [version 4.1-8]), survival (survival analysis (survfit, coxph, and cox.zph) [version 3.2-13]), lsmeans (lsmeans [version 2.30-0]), stats (glm [version 4.2.2]), pROC (roc, [version 1.18.4]), and default packages.

[0078] [Results] <Genome-wide association study of Japanese centenarians and post-GWAS analysis> To extract the genetic elements concentrated in centenarians, first, the genomic DNA (gDNA) sequences of 967 Japanese centenarians (Cent), including 30,081 Japanese control data (Table 3), 173 super centenarians (SC; people who lived up to 110 years old) and 614 semi-centenarians (SSC; people who lived up to 105 - 109 years old), were determined.

[0079] Among these samples, the gDNA sequences of 526 centenarians and 3,340 Japanese controls were determined by whole-genome sequencing, and the gDNA sequences of 441 centenarians and 26,741 Japanese controls were determined by DNA microarray using imputation. After removing sex-discrepant samples, close relatives by PI HAT analysis, and outliers of Japanese by principal component analysis (PCA) of gDNA sequences using 1000 Genomes data, 964 centenarians and 7,304 controls (Cont(GWAS)) were used for GWAS, and 10,000 controls (Cont(PRS)) were used for polygenic risk score analysis. GWAS was analyzed using PLINK1.90 separately for WGS data and DNA microarray data with 21 covariates (sex and PC1 - 20 in PCA analysis), and these GWAS statistics were combined using N-weighted multivariate GWAMA. As a result, four lead SNVs were isolated in three genes (APOE, EYS, GRM7, Figure 1). For the APOE locus, two different signals corresponding to the APOE4 missense mutation (rs429358, Figure 2) and the APOE2 missense mutation (rs7412, Figure 2) were found. For both the EYS and GRM7 loci, the lead SNVs (rs75571981 for EYS, Figure 3; rs73116078 for GRM7, Figure 4) were located in the intron region.

[0080] Here, Figure 1 is a diagram showing the results of the meta-GWAS of 964 Japanese centenarians and 7,306 Japanese controls, and the vertical axis is -log of the p-value 10is shown, where the horizontal axis represents each SNV present in each chromosome. Among the four lead SNVs, the effect sizes of APOE4 and GRM7 are negative, while the effect sizes of APOE2 and EYS are positive.

[0081] Figures 2 - 4 are enlarged views of three genes (APOE, GRM, and EYS) in the centenarian GWAS shown in Figure 1. Two lead SNVs, rs429358 (APOE4, p = 4.88×10-10) and rs7412 (APOE2, p = 2.21×10-7) are located in APOE (Figure 2), and rs75571981 (p = 3.48×10-8) and rs73116078 (p = 1.59×10-7) are located in EYS and GRM, respectively (Figures 3 and 4).

[0082] Next, the GTEx database was used to analyze the association of the four SNVs with gene expression. However, since expression data were not obtained for rs75571981 of EYS and rs73116078 of GRM7, additional analysis using Enformer, which integrates long-range interactions to predict gene expression from sequences, was performed. As a result, it was shown that the two missense mutations on APOE might affect gene expression, while rs75571981 of EYS and rs73116078 of GRM7 were less likely to be involved in gene expression (not shown).

[0083] Next, the minor allele frequencies (MAFs) of these SNVs were compared among Japanese controls (Cont(GWAS)), Japanese healthy elderly (HA), and centenarians (Cent). Figures 5 - 8 show the minor allele frequencies (MAF) of rs75571981 of APOE4, APOE2, EYS, and rs73116078 of GRM7 in the Japanese control database (ToMMo38k), Japanese controls (Cont(GWAS)), healthy elderly (HA), and centenarians (Cent).

[0084] HA, as a representative of the elderly with a long healthy life expectancy, is an elderly person who was 85 - 89 years old at the time of registration and maintained independence (''Prevention 1'' or no need for support in the Japanese care level). HA was involved in 54% of the 85 - 89-year-old population in Kawasaki City, and the remaining 46% were elderly people certified as ''Prevention 2'' or any care level. As a result of comparing the MAF, the MAF of the lead SNV of rs73116078 of APOE4 and GRM7 gradually decreased with age in Cent and HA, but the MAF of rs75571981 of APOE2 and EYS increased only in centenarians (Figs. 5 - 8).

[0085] To analyze the genetic components including low p-value regions, the GWAS summary statistics of Japanese centenarians were analyzed using GWAS gene set analysis by the MAGMA (Multi-marker analysis of genomic annotation) method. As a result of this GWAS gene set analysis, only 3 genes (APOE, NARS, FAM188B) were significant (P ≤ 0.05 / 17990) in the Bonferroni multiple comparison test, with P < 1.0x10 -4 、P < 1.0x10 -3 、13, 45, 284, and 1,099 genes were detected under the conditions of P < 0.01, P < 0.05. As a result of the MAGMA tissue expression analysis using these gene sets, the expression of the gene set (Cent.gene) with P = 0.01 or less was significant in 10 tissues including the heart, muscle, and pancreas, indicating that the gene set with P < 0.01 or less reflects the polygenic components of Japanese centenarians (not shown in the figure).

[0086] In addition, to understand the biological characteristics of the Cent. gene, the Cent. gene was analyzed using gene ontology (GO) and KEGG pathway enrichment analysis, but no significant GO and KEGG pathways were isolated.

[0087] To evaluate the overlap with known lists of aging-related genes, the ENTREZ IDs of genes in Cent.genes (270 genes), Longevity map (341 genes), CellAge (866 genes), GenAge (307 genes), KEGG longevity control pathway genes (hsa04211) (89 genes) including the IGF / insulin, sirtuin, AMP-AMPK, and TOR pathways from the Human Ageing Genomic Resources were compared. Also, the Japanese type 2 diabetes (T2D) risk loci (286 genes) were used as a reference for the analysis. Among these, the highest proportion of overlapping genes was 4.55% for T2D vs Cent.gene, and no obvious overlap was observed between Cent.gene and the known lists of aging-related genes (not shown).

[0088] SNP-based heritability (SNPh2), λ GC 、and to estimate the genetic correlations, the summary statistics of the Japanese centenarian GWAS18 were analyzed. Also, to evaluate the effects derived from the SNVs of these three genes (APOE, EYS, GRM7), the summary statistics excluding the loci of these three genes were also analyzed.

[0089] As a result, the SNPh2 of the Japanese centenarian GWAS was 0.204, which was similar to that of the Japanese centenarian GWAS without the three genes and the European 90th percentile GWAS, but lower than that of the European 99th percentile GWAS (Figure 9). The λ GC of the Japanese centenarian GWAS was 1.04, which was almost the same as that of the Japanese centenarian GWAS without the three genes, the European 90th percentile GWAS, and the European 99th percentile GWAS. This indicates that no obvious founder effect was observed in the genetic components of Japanese centenarians (Figure 10).

[0090] Here, Figure 9 shows the SNP-based heritability (SNPh2) calculated using the summary statistics of GWAS related to longevity in Japanese centenarians (Jpn.Cent.), Japanese centenarians excluding three loci (APOE, EYS, GRM7) (Jpn.Cent.-3genes), European 90%tile GWAS (Eur. 90%tile), European 99%tile GWAS (Eur. 99%tile), and European parental lifespan (Eur. PLS). Figure 10 shows λ GC calculated using the summary statistics of GWAS related to longevity.

[0091] As a result of the genetic correlation analysis using GWAS summary statistics, the GWAS summary statistics of Japanese centenarians showed positive correlations with the summary statistics of the 90%tile, 99%tile of Europeans, and parental lifespan, but the correlations were less than 0.331 and 0.191, respectively (Figure 11). The genetic correlation between the GWAS related to longevity and the summary statistics of other reported Japanese disease GWAS showed that the GWAS summary statistics of Japanese and European longevity were negatively correlated with the summary statistics of age-related diseases including Alzheimer's disease (AD), some cardiovascular diseases, and type 2 diabetes (T2D) (Figure 11).

[0092] Figure 11 shows the genetic correlation analysis among the GWAS related to longevity, Japanese disease GWAS, and Japanese quantitative GWAS using LD score regression. In Fig. 11, AD: Alzheimer's disease, CHF: congestive heart failure, PAD: peripheral arterial disease, IS: ischemic stroke, Atopy: atopic dermatitis, RA: rheumatoid arthritis, GaCa: gastric cancer, CoCa: colorectal cancer, T2D: type 2 diabetes, COPD: chronic obstructive pulmonary disease, BMI: body mass index, SBP: systolic blood pressure, DBP: diastolic blood pressure, PP: pulse pressure, MAP: mean arterial pressure, TC: total cholesterol, HDLC: high-density lipoprotein cholesterol, LDLC: low-density lipoprotein cholesterol, TG: triglyceride, BS: blood sugar, HbA1c: hemoglobin A1c, TP: total protein, Alb: albumin, NAP: non-albumin protein, A / G: albumin / globulin ratio, BUN: blood urea nitrogen, sCr: serum creatinine, eGFR: estimated glomerular filtration rate, UA: uric acid, Na: sodium ion, K: potassium ion, Cl: chloride ion, Ca: calcium ion, P: phosphorus ion, TBil: total bilirubin, ZTT: zinc sulfate turbidity test, AST: aspartate aminotransferase, ALT: alanine aminotransferase, ALP: alkaline phosphatase, GGT: γ-glutamyltransferase, APTT: activated partial thromboplastin time, PT: prothrombin time, Fbg: fibrinogen, CK: creatine kinase, LDH: lactate dehydrogenase, CRP: C-reactive protein, WBC: white blood cell count, Neutro: neutrophil count, Eosino: eosinophil count, Baso: basophil count, Mono: monocyte count, Lym: lymphocyte count, RBC: red blood cell count, Hb: hemoglobin, Ht: hematocrit, MCV: mean corpuscular volume, MCH: mean corpuscular hemoglobin, MCHC: mean corpuscular hemoglobin concentration, Plt: platelet count are shown.

[0093] Furthermore, when examining the genetic correlation between the longevity-related GWAS and the summary statistics of quantitative GWAS in Japanese, the summary statistics of GWAS for several biological indicators including blood pressure, glucose metabolism, and liver function showed a negative correlation with the summary statistics of GWAS for longevity including the GWAS of Japanese centenarians (Fig. 12).

[0094] Figure 12 shows the distribution of least squares means (lsmeans), which are sex-adjusted Z-scores of polygenic risk scores (PRS), calculated using GWAS summary statistics of Japanese controls (Cont(PRS)), healthy elderly (HA), and centenarians (Cent). Four lsmeans (SBP, CRP, T2D, CHF) are shown as representatives. Vertical bars indicate the range of the 95% confidence interval.

[0095] On the other hand, when we performed heritability enrichment analysis for tissue-specifically expressed genes and enhancer regions using the summary statistics of the Japanese centenarian GWAS, no specific enrichment was observed in any tissue, indicating that the genetic components of the Japanese centenarian GWAS are not concentrated in a small number of specific tissues, as is the case with the summary statistics of the disease GWAS (not shown).

[0096] Taken together, these results suggest that the genetic components of Japanese longevity analyzed in the Japanese centenarian GWAS include genetic components that show negative correlations with genetic components of multiple age-related diseases and biomarkers.

[0097] <Genetic component analysis of Japanese centenarians, healthy individuals, and controls using a series of PRS> To understand the details of genetic factors involved in longevity in Japanese people, we next compared the distribution of 62 polygenic risk scores (PRS) calculated by the Japanese GWAS summary statistics between centenarians and an additional 10,000 controls (Cont(PRS)). PRS was calculated for SNVs with p=0.01 or more using PRS-CS32 and output as PRS least squares mean (lsmeans), which is the sex-adjusted Z-score standardized by the PRS distribution of 10,000 Cont(PRS).

[0098] As a result of the lsmeans distribution analysis of PRS, the lsmeans distributions of PRS in seven phenotypes (mean arterial pressure (MAP), systolic blood pressure (SBP), diastolic blood pressure (DBP), γ-glutamyl transferase (GGT), alanine aminotransferase (ALT), congestive heart failure (CHF), basophil count (Baso)) were significantly different in both centenarians and HA compared to Cont(PRS).

[0099] For the additional seven phenotypes (C-reactive protein (CRP), blood sugar (BS), pulse pressure (PP), aspartate aminotransferase (AST), arrhythmia, type 2 diabetes (T2D), total cholesterol (TC)), there were significant differences in Cent(PRS) compared to Cont(PRS) by Bonferroni's multiple comparison test, and for the asthma phenotype, there was a significant difference in HA compared to Cont(PRS) by Bonferroni's multiple comparison test (p < 0.00079, Figures 12 and 13).

[0100] Figure 13 is a plot of the mean values of PRS in HA and Cent. For the multiple testing, Bonferroni's multiple comparison test was used (62 PRSs). The black circles indicate that the distribution of PRS lsmeans is significantly different from the control in both HA and centenarians. The dark gray circles and white circles indicate that the distribution of PRS lsmeans is significantly different from the control in Cent and HA, respectively. The gray circles indicate that the distribution of PRS lsmeans has no significant difference from the control in both Cent and HA.

[0101] To estimate the differences in the characteristics and amounts of genetic components between centenarians and controls, the genotypes of 62 PRSs and 4 SNVs (APOE4, APOE2, EYS, GRM7) were used, and a series of analyses of least absolute shrinkage and selection operator (LASSO) and multiple logistic regression were performed. Finally, 12 PRSs and 3 SNVs that significantly contributed to the differences in genetic components between Cont(PRS) and centenarians were identified (Figure 14).

[0102] Figure 14 shows a multiple logistic regression analysis between PRS and Cont(PRS) using the genotypes of 62 PRSs and 4 SNVs. For variable selection, univariate logistic regression and the least absolute shrinkage and selection operator (LASSO) were used. The genotypes of 12 PRSs and 3 SNVs significantly contributed to the difference in the genetic components between Cont(PRS) and Cent(PRS).

[0103] As a result of receiver operating characteristic curve (ROC) analysis, the area under the curve (AUC) by 12 PRSs and 3 SNVs was 0.694. The AUC of 12 PRSs was larger than that of 3 SNVs, suggesting that the polygenic components related to longevity in Japanese are larger than the genetic components derived from 4 SNVs including APOE4 and APOE2 (Figures 15 and 16).

[0104] Figure 15 shows the results of ROC curve analysis using the genotypes of 12 PRSs and 3 SNVs. Figure 16 shows the AUC calculated using only the genotypes of 3 SNVs, only 12 PRSs, and both the genotypes of 3 SNVs and 12 PRSs.

[0105] Next, the genetic differences between HA and the control were evaluated by a series of analyses of LASSO and multiple logistic regression. For this analysis, "Cent.PGS", a polygenic score based on the GWAS summary statistics of Japanese centenarians, was additionally introduced to evaluate the contribution of Cent.PGS to the genetic components between HA and the control (Figure 17).

[0106] Figure 17 shows the distribution of Cent.PGS in centenarians (Cent), healthy elderly (HA), Japanese controls of PRS (Cont(PRS)), and Japanese controls (Cont(GWAS)). In Figure 17, the horizontal axis represents Cent.PGS. It should be noted that the Cent.PGS of both Cent and Cont(GWAS) used to calculate the GWAS statistic is overestimated.

[0107] When comparing the Cent.PGS distribution of PGS, the Cent.PGS of HA was significantly higher than that of Cont(PRS) (P<0.001, Figure 17).

[0108] Next, a series of analyses of LASSO and multiple logistic regression were performed using the genotypes of 62 PRSs, Cent.PGS, and 4 SNVs (APOE4, APOE2, EYS, GRM7) (Figure 18). Figure 18 shows a diagram of the multiple logistic regression analysis between HA and Cont(PRS) using the genotypes of Cent.PGS, 62 PRSs, and 4 SNVs. Univariate logistic regression and LASSO were used for variable selection. PGS, 8 PRSs, and the genotype of APOE4 significantly contributed to the difference in genetic factors between HA and Cont(PRS).

[0109] From the results shown in Figure 18, it was shown that 8 PRSs (Baso, Eosinophil count (Eosino), Potassium (K), GGT, Monocyte count (Mono), CHF, Asthma, MAP), Cent.PGS, and 1 SNV (APOE4, APOE2, EYS, GRM7) were significantly higher than the PRS of HA.

[0110] PGS and 1 SNV (APOE4) are important genetic factors between HA and Cont.(PRS) (Figure 18). In particular, Cent.PGS showed the highest effect among these genetic factors (OR(95%CI): 1.33(1.24 - 1.42)), and it is known that Cont.PGS is the largest genetic factor in HA rather than APOE4 (OR(95%CI): 0.84(0.78 - 0.91)).

[0111] <According to the survival time analysis, PGS was associated with healthy life expectancy in HA> If the genetic factors identified from the GWAS of Japanese centenarians are truly longevity genetic factors, they should be associated with the survival of future generations. To evaluate this prediction, a survival rate analysis stratified by each genetic factor including 4 SNVs and Cent.PGS was performed in the HA cohort.

[0112] APOE was stratified into three groups (APOE e2 pos, APOE e3 / e3, APOE e4 pos), EYS and GRM7 were stratified into three groups according to genotype, and Cent.PGS was stratified into two groups: less than 0 (low Cent.PGS HA) and 0 or more (high Cent.PGS HA). For this survival analysis, two different endpoints, lifespan and healthspan, were used. In addition, since gender has already been shown to be a strong determinant for both lifespan and healthspan, these survival analyses were performed separately for men and women.

[0113] Figures 19 - 20 are diagrams showing Kaplan - Meier survival analysis of lifespan in healthy elderly (HA). Figure 19 shows the analysis results for HA men, and Figure 20 shows the analysis results for HA women. Figures 21 - 22 are diagrams showing Kaplan - Meier survival analysis of healthspan in healthy elderly (HA). Figure 21 shows the analysis results for HA men, and Figure 22 shows the analysis results for HA women.

[0114] According to the Kaplan - Meier survival analysis for lifespan, the survival probability of high Cent.PGS HA men was significantly higher than that of low Cent.PGS HA men, but this was not the case for HA women (Figures 19 - 20). In the Kaplan - Meier survival analysis of healthspan, it was shown that for both men and women, high Cent.PGS HA had a significantly higher survival probability than low Cent.PGS HA (Figures 21 - 22).

[0115] Furthermore, since the genetic elements of Cent.PGS partially overlap with the genetic elements of APOE, EYS, and GRM7, when survival analysis using regularized multiple Cox regression was performed, Cent.PGS and gender were independently associated with lifespan, and APOE4, Cent.PGS, and gender were independently associated with healthspan (Figures 23 - 24).

[0116] Figures 23 to 24 are diagrams showing survival analysis using regularized multiple Cox regression with genetic factors including Cent.PGS, genotypes of four SNVs, sex, and educational history for both lifespan and healthspan. Figure 23 shows the analysis results for lifespan, and Figure 24 shows the analysis results for healthspan.

[0117] In the Kaplan-Meier analysis of the lifespan-healthspan gap, sex was also a significant genetic factor for the lifespan-healthspan gap. On the other hand, in the regularized multiple Cox regression analysis, sex and age at the end of healthspan were significantly associated with the lifespan-healthspan gap, while Cent.PGS and the four SNVs were not associated (Figures 25 to 26). In contrast, the four SNVs and Cent.PGS were not significantly associated with either the number of survival days since the survey or the remaining days of healthspan (not shown).

[0118] Here, Figure 25 is a diagram showing the Kaplan-Meier survival analysis for the lifespan-healthspan gap, and Figure 26 is a diagram showing the survival analysis using regularized multiple Cox regression analysis for the lifespan-healthspan gap.

[0119] Next, the relationship between genetic factors (four SNVs and Cent.PGS) and the standardized observable score in HA was analyzed using generalized linear regression adjusted for sex and age at registration (Figure 27). Figure 27 is a diagram showing the correlation analysis between genetic factors and the standardized observable score. It is a generalized linear regression adjusted for sex and age at registration. In Figure 27, the circles represent the β coefficients, and the range of the bar graph represents the 95% confidence interval. The Bonferroni multiple comparison test (q = 0.05 / 26) was used for multiple testing. Open circles indicate p < 0.05, and gray circles indicate p < 0.05 and q < 0.05.

[0120] The results in Figure 27 show that known associations were found between the APOE4 genotype and a history of dementia, and between the APOE2 genotype and LDLC concentration, but no significant associations were found between Cent.PGS and known scores related to age and survival, including leukocyte telomere length, frailty index, ALB, and BMI (Figure 27).

[0121] These survival analyses showed that Cent.PGS was associated with healthy lifespan. Taken together, these results suggest that Cent.PGS, APOE4, and sex are genetic factors independently associated with healthy lifespan, and that Cent.PGS based on the GWAS analysis of centenarians in this example can be used as a novel genetic indicator associated with healthy lifespan.

[0122] <Genetic components that persist in the oldest age group of humans, exceeding the 99.99% survival rate of the population> In this example, the genome analysis of 964 centenarians, including 173 supercentenarians, is enabling the analysis of genetic factors required for exceptional longevity, which is approaching the limit of current human lifespan. To understand the genetic factors associated with the limit of lifespan of centenarians, we performed survival analysis of centenarians using Cent.PGS, four SNVs, and gender (Figures 28 to 30). 28 and 29 are diagrams showing Kaplan-Meier survival analyses of the lifespan of Japanese centenarians, with Fig. 28 showing the analysis results for men and Fig. 29 showing the analysis results for women. FIG. 30 shows survival analysis of life span of Japanese centenarians using regularized multiple Cox regression with genetic factors including Cent.PGS, genotypes of four SNVs, and gender.

[0123] Survival analyses using both Kaplan-Meier and regularized multiple Cox regression analyses showed that the only genetic factor associated with age at death among centenarians was sex (Figures 28-30). Furthermore, these genetic factors were not associated with survival days after the survey (not shown). From these results, it was shown that Cent.PGS concentrates genetic factors up to 100 years old, but is not associated with lifespan beyond 100 years old.

[0124] Next, the relationship between genetic factors (4 SNVs and Cent.PGS) of Japanese centenarians and the standardized observable score was analyzed using generalized linear regression adjusted for gender and age at registration. As a result, a known relationship was observed between the APOE4 genotype and the extended clinical dementia rating scale (ex.CDR), and between the APOE2 genotype and LDL-C concentration. However, no significant relationship was observed between Cent.PGS and the known aging and survival-related score, which was the same as that of healthy individuals (not shown in the figure).

[0125] To extract genetic factors that enable humans to reach the maximum age studied so far, genetic factors were compared between two groups of centenarians stratified by lifespan. To more accurately stratify centenarians, centenarians were first divided into men and women, and then among people of the same birth year, they were divided into those above and below the 99.99% tile (the oldest old (HHA) and ordinary centenarians (nCent), Figure 31). Figure 31 shows the distribution of the death year and age at death of Japanese centenarians. The boundary ages of survivors at the 90%, 99%, 99.9%, and 99.99% tiles were calculated based on the total population number of the birth year. The Highest Human Agers (HHA) at the 99.99% tile were shown in black, and the remaining centenarians were considered normal centenarians (nCent).

[0126] As a result of the PRS lsmeans distribution analysis, the PRS lsmeans distributions for 8 phenotypes (DBP, SBP, MAP, BS, T2D, PP, CRP, Baso) were significantly different for Cont.(PRS) in both HHA and nCent, there was a significant difference between HHA and Cont.(PRS) for one phenotype (AST), and there was a significant difference between nCent and Cont.(PRS) for arrhythmia (Bonferroni's multiple comparison test, p < 0.00079, Figures 32 - 33). The lsmeans distribution of PRS in fibrinogen (Fbg) showed no significant difference between nCent and Cont. (PRS), or between HHA and Cont. (PRS). However, a significant difference in the PRS distribution in Fbg was observed between nCent and HHA (Figs. 32 - 33).

[0127] Here, Fig. 32 shows the lsmeans distribution of the sex - adjusted Z - scores of PRS calculated using Japanese GWAS summary statistics for Japanese controls of nCent and HHA. Four lsmeans (AST, SBP, CRP, Fbg) are shown as representatives. Fig. 33 shows the mean values of PRS in nCent and HHA. The Bonferroni multiple - comparison test was used for multiple testing (62 PRS). In Fig. 33, the black circles indicate that the distribution of PRS lsmeans is significantly different from the control in both HHA and nCent. The dark - gray circles and white - circles indicate that the distribution of PRS lsmeans is significantly different from the control in HHA and nCent, respectively.

[0128] To estimate the differences in the characteristics and amounts of genetic components between HHA and nCent, a series of analyses of LASSO and multiple logistic regression were performed using the genotypes of 62 PRS and 4 SNVs (APOE4, APOE2, EYS, GRM7) (Fig. 34). Fig. 34 shows the multiple logistic regression analysis between PRS and Cont(PRS) using the genotypes of 62 PRS and 4 SNVs. Univariate logistic regression and LASSO were used for variable selection.

[0129] Finally, two PRS (Fbg and AST) and one SNV (APOE4) were identified as being significantly associated with the differences in genetic elements between HHA and nCent (Fig. 34).

[0130] [Discussion] In this example, genome-wide association studies (GWAS) were used to identify genetic components of exceptional longevity in 964 Japanese centenarians, including 173 supercentenarians (SCs). Although the sample size is relatively small compared to recent GWAS, the characteristics of the Japanese centenarian GWAS in this example are: 1) it is a longevity GWAS that includes 35.2% of the 99.99% tile survivors and 42.5% of the 99.9% tile survivors in the Japanese population; 2) it is a non-European longevity GWAS, and it is possible to understand the genetic elements of longevity that are common to European-Asian populations.

[0131] Survival analysis by regularized regression using sex, four representative single nucleotide variants (SNVs) identified in the Japanese centenarian GWAS, and Cent.PGS calculated based on the summary statistics of the Japanese centenarian GWAS showed that sex, APOE4 genotype, and Cent.PGS were independently associated with healthspan in a healthy elderly cohort but not in centenarians.

[0132] The genetic components of the APOE4 genotype and Fbg and AST were significantly enriched in the Highest Human Agers (HHAs), who are 99.99% tile survivors, compared to other Japanese normal centenarians (nCent) (men 106 years and older, women 108 years and older, respectively). In summary, the 53,760 single nucleotide polymorphisms (SNVs) shown in Tables 4 to 451, which are the genetic components of longevity revealed by the Japanese centenarian GWAS, are associated with healthspan in later life.

[0133] In the Japanese centenarian GWAS of this example, four representative SNVs, including APOE4 and APOE2, were observed at three loci (APOE, EYS, GRM7), and the allele frequency of the APOE4 allele decreasing with age was a characteristic of the genetic components of longevity common to both European and Asian populations.

[0134] As a result of the genetic correlation analysis, it was revealed that the GWAS summary statistics of longevity in Japanese are positively correlated with those of longevity in Europeans and parental lifespan, and negatively correlated with those of age-related diseases and biomarkers in Japanese such as T2D, CVD, blood pressure, and blood glucose levels. This is consistent with the reported results of genetic correlation analysis of healthspan, lifespan, and longevity, suggesting that the genetic structure related to these age-related phenotypes is likely to be conserved between European and Asian populations. In contrast, both genetic correlation analysis and PRS analysis showed that the genetic components of AST are included in the genetic components of Japanese longevity but not in those of European longevity and parental lifespan, suggesting that AST is a longevity-related phenotype characteristic of Asian populations.

[0135] It is difficult to say that the content analysis method of genetic elements in the GWAS summary statistics has been established. The lsmean of the Z-scored PRS performed in this example was able to evaluate the differences between groups of genetic elements for the target phenotype. In this example, by multiple logistic regression, diseases and quantitative biomarkers with significantly different Z-score PRS distributions between centenarians and HA were identified, and it was clarified that they are consistent with the diseases and quantitative biomarkers identified by genetic correlation analysis.

[0136] One of the important findings of this example is Cent.PGS. Cent.PGS is based on 53,760 single nucleotide polymorphisms (SNVs) related to centenarians with a P-value of less than 0.01 by genome-wide association analysis for centenarians, and can quantify the genetic components of longevity in each individual. It was independently associated with healthspan, along with gender and APOE4 genotype. This Cent.PGS is expected to be used in gene-environment interaction analysis that numerically evaluates the interaction between environmental and behavioral factors on healthspan, and may lead to the identification of preventive interventions that promote healthy aging in more people.

[0137] Among the 53,760 single nucleotide polymorphisms (SNVs), SNVs with large absolute values of β-values and large contributions to Cent.PGS can be used as indicators and biomarkers for predicting healthspan and calculating Cent.PGS. Specifically, SNVs shown in Table 1 with β-values of 0.21 or more from genome-wide association studies of centenarians and showing a positive correlation with centenarians; SNVs shown in Table 2 with β-values of -0.20 or less from genome-wide association studies of centenarians and showing a negative correlation with centenarians.

[0138] Furthermore, by comparing gene expressions related to Cent.PGS through a combined analysis of Cent.PGS and cell experiments, especially those using induced pluripotent stem cells (iPSCs), it becomes possible to analyze genes and pathways related to healthspan, and in recent years, it has been possible to find new pharmacological intervention targets for improving healthy aging and resilience to aging, which have been attracting attention as concepts of the aging process.

[0139] Another unique aspect of this example is that a genetic analysis was conducted on the limits of human lifespan between the population of 99.99% tile survivors and other centenarians, and the APOE4 genotype, which is a multi-gene component of AST and Fbg, was identified. The multi-gene component of AST was identified in Japanese centenarians through genetic correlation and PRS analysis, but not in European centenarians, suggesting that it is not only unique to Asian centenarians but also effective in the limits of human lifespan. Also, survival analysis using the scores observed in Japanese centenarians showed that AST as a phenotype was associated with survival from the study subjects. This indicates that hepatocytes and muscle cells are less likely to lose their functions and be destroyed, which is considered to be one of the factors related to the limits of human lifespan.

[0140] In this example, multiple genetic factors of Fbg were also identified as genetic factors related to the human lifespan limit. In fact, a hypercoagulability including a significant increase in plasma fibrinogen has been reported in a strictly selected group of healthy centenarians. Although the clinical significance of fibrinogen in human longevity has not yet been elucidated, recent genome-wide analysis of plasma fibrinogen has identified multi-gene components shared with liver enzymes and liver regulators, suggesting that liver homeostasis may play an important role in achieving the human lifespan limit.

[0141] In conclusion, through the comprehensive analysis of genomic data and detailed phenotypic data of centenarians and healthy elderly individuals, this example established the concept of Cent.PGS, which is a genetic component extracted from centenarian GWAS and is related to the healthy lifespan of elderly individuals aged 85 to 89 years. This example anticipates that the genetic components represented by Cent.PGS will be related not only to the healthy lifespan but also to the resilience in age-related pathologies. It will be important to analyze how genetic factors, observable biomarkers, and environmental factors interact and are related to the dynamic changes in healthy lifespan and resilience. This example is convinced that Cent.PGS, which quantifies the genetic factors of longevity for each individual, will play an important role in future aging research.

[0142] [List of known literature] Known literature 1: Tadaka, S. et al., Nucleic Acids Res 49, D536-D544 (2021). Known literature 2: Fairley, S. et al., Nucleic Acids Res 48, D941-D947 (2020). Known literature 3: Purcell, S. et al., Am J Hum Genet 81, 559-75 (2007). Known literature 4: Baselmans, B.M.L. et al., Nat Genet 51, 445-451 (2019). Prior art document 5: Turner, S.D., Journal of Open Source Software 3, 1-2 (2018). Prior art document 6: Pruim, R.J. et al., Bioinformatics 26, 2336-7 (2010). Prior art document 7: Bulik-Sullivan, B. et al., Nat Genet 47, 1236-41 (2015). Prior art document 8: Timmers, P.R. et al., Elife 8(2019). Prior art document 9: Ishigaki, K. et al., Nat Genet 52, 669-679 (2020). Prior art document 10: Shigemizu, D. et al., Transl Psychiatry 11, 151 (2021). Prior art document 11: Avsec, Z. et al. Nat Methods 18, 1196-1203 (2021). Prior art document 12: Watanabe, K. et al., Nat Commun 8, 1826 (2017). Prior art document 13: de Leeuw, C.A. et al., PLoS Comput Biol 11, e1004219 (2015). Prior art document 14: Consortium, G.T., Nat Genet 45, 580-5 (2013). Prior art document 15: Tacutu, R. et al., Nucleic Acids Res 46, D1083-D1090 (2018). Prior art document 16: Kanehisa, M. et al., Nucleic Acids Res 51, D587-D592 (2023). Prior art document 17: Imamura et al., Nat Commun 7, 10531 (2016). Prior art document 18: Finucane, H.K. et al., Nat Genet 47, 1228-35 (2015). Prior art document 19: Kanai, M. et al, Nat Genet 50, 390-400 (2018). Prior art document 20: Ge, T. et al., Nat Commun 10, 1776 (2019). Prior art document 21: Sasaki, T. et al., Elife 12(2023). Prior art document 22: Rockwood, K. & Mitnitski, J Gerontol A Biol Sci Med Sci 62, 722-7 (2007).

[0143] Tables 4 to 451 below show 53,760 single nucleotide polymorphisms (SNVs) associated with centenarians with a P-value of less than 0.01 by genome-wide association analysis in the prediction method of this embodiment for centenarians. [Table 4] [Table 5] [Table 6] [Table 7] [Table 8] [Table 9] [Table 10] [Table 11] [Table 12] [Table 13] [Table 14]

Table 15

Table 16

Table 17

Table 18

Table 19

Table 20

Table 21

Table 22

Table 23

Table 24

Table 25

Table 26

Table 27

Table 28

Table 29

Table 30

Table 31

Table 32

Table 33

Table 34

Table 35

Table 36

Table 37

Table 38

Table 39

Table 40

Table 41

Table 42

Table 43

Table 44

Table 45

Table 46

Table 47

Table 48

Table 49

Table 50

Table 51

Table 52

Table 53

Table 54

Table 55

Table 56

Table 57

Table 58

Table 59

Table 60

Table 61

Table 62

Table 63

Table 64

Table 65

Table 66

Table 67

Table 68

Table 69

Table 70

Table 71

Table 72

Table 73

Table 74

Table 75

Table 76

Table 77

Table 78

Table 79

Table 80

Table 81

Table 82

Table 83

Table 84

Table 85

Table 86

Table 87

Table 88

Table 89

Table 90

Table 91

Table 92

Table 93

Table 94

Table 95

Table 96

Table 97

Table 98

Table 99

Table 100

Table 101

Table 102

Table 103

Table 104

Table 105

Table 106

Table 107

Table 108

Table 109

Table 110

Table 111

Table 112

Table 113

Table 114

Table 115

Table 116

Table 117

Table 118

Table 119

Table 120

Table 121

Table 122

Table 123

Table 124

Table 125

Table 126

Table 127

Table 128

Table 129

Table 130

Table 131

Table 132

Table 133

Table 134

Table 135

Table 136

Table 137

Table 138

Table 139

Table 140

Table 141

Table 142

Table 143

Table 144

Table 145

Table 146

Table 147

Table 148

Table 149

Table 150

Table 151

Table 152

Table 153

Table 154

Table 155

Table 156

Table 157

Table 158

Table 159

Table 160

Table 161

Table 162

Table 163

Table 164

Table 165

Table 166

Table 167

Table 168

Table 169

Table 170

Table 171

Table 172

Table 173

Table 174

Table 175

Table 176

Table 177

Table 178

Table 179

Table 180

Table 181

Table 182

Table 183

Table 184

Table 185

Table 186

Table 187

Table 188

Table 189

Table 190

Table 191

Table 192

Table 193

Table 194

Table 195

Table 196

Table 197

Table 198

Table 199

Table 200

Table 201

Table 202

Table 203

Table 204

Table 205

Table 206

Table 207

Table 208

Table 209

Table 210

Table 211

Table 212

Table 213

Table 214

Table 215

Table 216

Table 217

Table 218

Table 219

Table 220

Table 221

Table 222

Table 223

Table 224

Table 225

Table 226

Table 227

Table 228

Table 229

Table 230

Table 231

Table 232

Table 233

Table 234

Table 235

Table 236

Table 237

Table 238

Table 239

Table 240

Table 241

Table 242

Table 243

Table 244

Table 245

Table 246

Table 247

Table 248

Table 249

Table 250

Table 251

Table 252

Table 253

Table 254

Table 255

Table 256

Table 257

Table 258

Table 259

Table 260

Table 261

Table 262

Table 263

Table 264

Table 265

Table 266

Table 267

Table 268

Table 269

Table 270

Table 271

Table 272

Table 273

Table 274

Table 275

Table 276

Table 277

Table 278

Table 279

Table 280

Table 281

Table 282

Table 283

Table 284

Table 285

Table 286

Table 287

Table 288

Table 289

Table 290

Table 291

Table 292

Table 293

Table 294

Table 295

Table 296

Table 297

Table 298

Table 299

Table 300

Table 301

Table 302

Table 303

Table 304

Table 305

Table 306

Table 307

Table 308

Table 309

Table 310

Table 311

Table 312

Table 313

Table 314

Table 315

Table 316

Table 317

Table 318

Table 319

Table 320

Table 321

Table 322

Table 323

Table 324

Table 325

Table 326

Table 327

Table 328

Table 329

Table 330

Table 331

Table 332

Table 333

Table 334

Table 335

Table 336

Table 337

Table 338

Table 339

Table 340

Table 341

Table 342

Table 343

Table 344

Table 345

Table 346

Table 347

Table 348

Table 349

Table 350

Table 351

Table 352

Table 353

Table 354

Table 355

Table 356

Table 357

Table 358

Table 359

Table 360

Table 361

Table 362

Table 363

Table 364

Table 365

Table 366

Table 367

Table 368

Table 369

Table 370

Table 371

Table 372

Table 373

Table 374

Table 375

Table 376

Table 377

Table 378

Table 379

Table 380

Table 381

Table 382

Table 383

Table 384

Table 385

Table 386

Table 387

Table 388

Table 389

Table 390

Table 391

Table 392

Table 393

Table 394

Table 395

Table 396

Table 397

Table 398

Table 399

Table 400

Table 401

Table 402

Table 403

Table 404

Table 405

Table 406

Table 407

Table 408

Table 409

Table 410

Table 411

Table 412

Table 413

Table 414

Table 415

Table 416

Table 417

Table 418

Table 419

Table 420

Table 421

Table 422

Table 423

Table 424

Table 425

Table 426

Table 427

Table 428

Table 429

Table 430

Table 431

Table 432

Table 433

Table 434

Table 435

Table 436

Table 437

Table 438

Table 439

Table 440

Table 441

Table 442

Table 443

Table 444

Table 445

Table 446

Table 447

Table 448

Table 449

Table 450

Table 451

Claims

**Claim 1** A prediction method for predicting healthy life expectancy, comprising: a step of calculating a polygenic score of a subject by the following formula (1) based on 53,760 single nucleotide polymorphisms (SNVs) related to centenarians with a P-value of less than 0.01 by genome-wide association analysis for centenarians; 【Number 1】 In the formula (1), i is a positive number from 1 to n, n is the number of SNVs used for calculating the polygenic score, and SNV i represents the genotype at the i-th SNV, and β i is the β value of the SNV by the i-th genome-wide association analysis. a step of evaluating healthy life expectancy based on the calculated polygenic score of the subject, wherein the 53,760 SNVs are the SNVs shown in Tables 4 to 451 of this specification, and a prediction method characterized in that the number n of SNVs used for calculating the polygenic score is 100 or more. **Claim 2** The prediction method according to claim 1, wherein the number n of SNVs used for calculating the polygenic score is 1,000 or more. **Claim 3** The prediction method according to claim 1, wherein the number n of SNVs used for calculating the polygenic score is 10,000 or more. **Claim 4** The prediction method according to claim 1, wherein the SNVs used for calculating the polygenic score include at least any one of APOE4 represented by rs429358, APOE2 represented by rs7412, GRM7 represented by rs73116078, and EYS represented by rs75571981. **Claim 5** The prediction method according to claim 1, wherein the subject is an Asian person. **Claim 6** The prediction method according to claim 1, wherein the subject is a Japanese person. **Claim 7** The prediction method according to claim 1, further comprising a step of evaluating at least any one of phenotypic data, biomarkers, multi-omics, and analysis results of biological samples for the subject. **Claim 8** The prediction method according to claim 7, further comprising a step of evaluating at least any one of the analysis results of the polygenic risk score (PRS) of aspartate aminotransferase (AST) and the polygenic risk score (PRS) of fibrinogen (Fbg) for the subject. **Claim 9** A biomarker for predicting healthy life expectancy, characterized in that it is a single nucleotide polymorphism (SNV) with a β value of 0.21 or more by genome-wide association analysis of centenarians and showing a positive correlation with centenarians. **Claim 10** A biomarker for predicting healthy life expectancy, characterized in that it is a single nucleotide polymorphism (SNV) with a β value of -0.20 or less by genome-wide association analysis of centenarians and showing a negative correlation with centenarians.