Longevity-related protein modification and use thereof

By analyzing proteomic databases across species to identify amino acid residues associated with longevity, the method addresses the lack of understanding in PTMs for longevity regulation, enabling the extension of lifespan and treatment of age-related diseases.

WO2026062653A2PCT designated stage Publication Date: 2026-03-26BAR ILAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-17
Publication Date
2026-03-26

Smart Images

  • Figure IL2025050823_26032026_PF_FP_ABST
    Figure IL2025050823_26032026_PF_FP_ABST
Patent Text Reader

Abstract

Methods for identifying an amino acid alteration associated with a biological trait are provided. Methods of extending the longevity of a human subject and methods of extending the longevity of a non-human subject are also provided as are methods of determining is a subject is suitable to have their longevity extended.
Need to check novelty before this filing date? Find Prior Art

Description

LONGEVITY-RELATED PROTEIN MODIFICATION AND USE THEREOFREFERENCE TO AN ELECTRONIC SEQUENCE LISTING

[0001] The contents of the electronic sequence listing (BIU-P-050-PCT.xml; Size: 2,652 bytes; and Date of Creation: September 15, 2025) is herein incorporated by reference in its entirety.CROSS REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit of priority of U.S. Provisional Patent Application No. 63 / 695,424, filed on September 17, 2024, and U.S. Provisional Patent Application No. 63 / 747,097, filed on January 20, 2025 the contents of which are all incorporated herein by reference in their entirety.FIELD OF INVENTION

[0003] The present invention is in the field of mammalian longevity.BACKGROUND OF THE INVENTION

[0004] The significant increase in human longevity and the growing proportion of elderly in society over the past century are accompanied by an exponential rise in age-related diseases. This led to an urgent need for new approaches for extending human health span. However, to this end, a deeper understanding of the underlying mechanisms of healthy ageing is required. Armed with such knowledge, it will be possible to develop interventions that will help alleviate the negative impact of ageing and convert the elderly population from dependent to contributing individuals. One intriguing option to explore how healthy ageing may be achieved is by analyzing nature's largest on-going biological experiment, namely, evolution, and in particular, the development of long-lived animals.

[0005] To date, few studies have sought to understand the mechanisms underlying the increased lifespans of long-lived organisms. These studies suggested that such species develop and enhance specific longevity-favoring characteristics such as body size, brain development, sociality, increased DNA repair and protection against tumor formation.Specifically, a number of proteins / pathways shown to control lifespan in short-lived model organisms may have life extending functions in long-lived ones. For example, the expression of the well-known tumor suppressor protein p53, that was shown to control the lifespan of short-lived organisms, is significantly increased in elephants. This finding was suggested to explain the low incidence of cancer and increased survival in these long-lived mammals.

[0006] Various attempts were also performed at the -omics level to identify the regulation of lifespan in long-lived animals. Meta-analysis of age-related gene expression profiles in mice, rats, and humans identified common signatures of ageing, including inflammation, immune response and energy metabolism-related pathways. Wider meta RNA-seq analyses across databases of 26 or 41 mammalian species also identified transcriptional signatures of known longevity-related pathways. In addition to the above-mentioned pathways, these analyses identified DNA repair, IGF1 expression and mitochondrial translation. Interestingly, the expression of these genes was modulated by pro-longevity interventions, including mTOR inhibition and caloric restriction (CR). In contrast to transcriptome analyses, most attempts to identify longevity-associated proteomic signatures were performed in humans. For example, a recent study identified 754 human plasma proteins that are associated with chronological age, and are mostly related to inflammatory response, organismal injury, cell and organismal survival, and cell death pathways. In addition, Coenen et al. (Coenen, et al., “Markers of aging: Unsupervised integrated analyses of the human plasma proteome”, Frontiers in Aging 4, 1112109 (2023)) identified 273 plasma proteins significantly associated with ageing. Pathway and network analyses of these proteins highlighted several pathways including those that were shown to regulate longevity, such as IGF signaling, sulfur binding, TNFa (inflammation) and metabolic diseases. Likewise, among the pathways of ageing-related proteins identified by others were sulfur binding and IGF 1 -related hormonal signaling. Beside transcriptome and proteome analyses, several studies attempted to characterize the human age-related metabolome and found several known pathways that were shown to regulate longevity such as kynurenine NAD+ biosynthesis, one carbon / transsulfuration, inflammation, oxidative stress, and lipid metabolism pathways.

[0007] Reversible post-translational modifications (PTMs) refer to the posttranslational addition of any of a plethora of modifying molecules, such as acetyl, phosphoryl and methyl groups, to one or more specific amino acid residues of a target protein (PTM sites).

[0008] Surprisingly, despite the extensive studies on various -omics datasets, to the best of the inventor’s knowledge, no in-depth study was performed to explore the role of reversibleprotein PTMs in longevity regulation. Yet, some PTMs were found to be associated with age-related diseases, and to accumulate with age (Santos, et al., “Protein Posttranslational Modifications: Roles in Aging and Age-Related Disease”, Oxid Med Cell Longev 2017, 5716409 (2017)). The attachment of specific moi eties to a protein can affect a wide range of its characteristics, including stability, enzymatic activity, localization, and interactions. Thus, PTMs regulate various biological processes, such as signal transduction, gene expression, DNA repair and metabolism. As a result, in addition to the fundamental cellular roles, PTMs are a key factor in regulating systemic processes, including adaptation to environmental changes encountered by the organism over a lifetime. The plasticity of such reversible PTMs can maintain homeostasis to preserve healthy lifespan and thus provides a specific rationale for their regulation of longevity. Moreover, it can be hypothesized that over the course of evolution, specific PTM sites evolved or mutated to support the lifespan extension occurring in long-lived organisms. Accordingly, exploring the regulation of lifespan by PTMs could offer major new insight into the ageing process.

[0009] Protein acetylation on lysine s-amino group is one of the most extensively studied PTMs in eukaryotes. Lysine acetyltransferases / deacetylases (KATs / KDACs) control the acetylation status of thousands of proteins. Still, although acetylation has been connected to the regulation of lifespan (Lu, et al., “Protein acetylation and aging”, Aging vol. 3 911-912 (2011)), it remains largely unknown which of these proteins / pathways directly regulate ageing. Significant results have demonstrated that KAT / KDAC activities directly control yeast, nematode, fly and mammalian lifespan. For example, a transgenic (TG) mouse model over-expressing the SIRT6 deacetylase lives longer and with improved health compared to wild type (WT) littermates. Yet, acetylation sites that control healthy lifespan remain to be specifically explored. Traditionally, a point mutation of acetylated lysine (K) to glutamine (Q) is used in molecular biology to mimic constant acetylation, while the exchange from K to arginine (R) is used to mimic the fixation of the non-acetylated state of the protein. An algorithm for determining sites that regulate longevity as well as methods of extending lifespan are greatly needed.SUMMARY OF THE INVENTION

[0010] The present invention provides methods for identifying an amino acid alteration associated with a biological trait are provided. Methods of extending the longevity of a human subject and methods of extending the longevity of a non -human subject are alsoprovided as are methods of determining is a subject is suitable to have their longevity extended.[Oi l] According to a first aspect, there is provided a method for identifying an amino acid residue associated with a biological trait, the method comprising: a. receiving a proteomic database for a plurality of species of mammals; b. receiving a database comprising a quantifiable measure of the biological trait for the plurality of species of mammals; c. selecting a first amino acid to test and an organism from the plurality of species of mammals; d. for sites of a residue of the first amino acid in a proteome of the organism, establishing a first group of species of mammals of the plurality that also have the first amino acid at the residue and a second group of species of mammals of the plurality that have a second amino acid at the residue, wherein the first and second amino acid are not the same amino acid; e. performing a statistical test to determine i. sites where the quantifiable measure of the biological trait in the first group is significantly higher than the quantifiable measure of the biological trait in the second group to produce a first list of sites of the first amino acid potentially associated with the biological trait, and ii. sites wherein the quantifiable measure of the biological trait in the second group is significantly higher than the quantifiable measure of the biological trait in the first group to produce a second list of sites of the second amino acid potentially associated with the biological trait; f. for sites of the first list and the second list determining if the potential association is due to phylogenetic distance; andg. selecting a site and amino acid residue from the first list or the second list whose association with the biological trait is not due to phylogenetic distance; thereby identifying an amino acid residue associated with a biological trait.

[0012] According to some embodiments, the first amino acid is an amino acid that can be post-translationally modified and the second amino acid is an amino acid that mimics the modified or unmodified state of the first amino acid.

[0013] According to some embodiments, the method further comprises before step (d) receiving an atlas of marks of the post-translation modification on the first amino acid for the organism and identifying is the received atlas residues of the first amino acid that can be modified in the organism, wherein step (d) is performed for each identified modifiable residue.

[0014] According to some embodiments, the first amino acid is lysine and the second amino acid is arginine or glutamine

[0015] According to some embodiments, the method further comprises before step (d) receiving an acetylome for the organism and identifying is the received acetylome lysine residues that can be acetylated in the organism, wherein step (d) is performed for each identified acetylatable residue.

[0016] According to some embodiments, the statistical test is performed on the average of the quantifiable measure in the first group and the average of the quantifiable measure in the second group.

[0017] According to some embodiments, determining if the potential association is due to phylogenetic distance comprises producing a site score for sites, performing a permutation test to produce a random distribution of site scores for the sites and wherein a statistically significant site score as compared to the random distribution indicates the potential association is not due to phylogenetic distance, wherein the producing a site score comprises: a. for every pair of species in the plurality of species determining the difference between the quantifiable measure in the pair of species and multiplying the difference by the phylogenetic distance between the pair species to produce a product;b. multiplying the product by a weight to produce a weighted product, wherein the weight is representative of a site’s contribution to the hypothesis that the first amino acid or the second amino acid at the site is associated with the biological trait; and c. summing all the weighted products for every pair of species in the plurality of species to produce a site score; and wherein the permutation test produces a random distribution of site scores by assigning random weights to produce random site scores.

[0018] According to some embodiments, the weight is the probability that regardless of amino acid at a site the pair of species would have a difference in the trait that is greater than average and have a phylogenetic distance that is smaller than average or have a difference in the trait that is smaller than average and have a phylogenetic distance that is greater than average.

[0019] According to some embodiments, the weight is a. positive for pairs i. with the same amino acid at the position, a difference in the trait that is smaller than average and a phylogenetic distance that is greater than average; and ii. with different amino acids at the position, a difference in the trait that is larger than average and a phylogenetic distance that is smaller than average; and b. negative for pairs i. with the same amino acid at the position, a difference in the trait that is greater than average and a phylogenetic distance that is greater than average; ii. with the same amino acid at the position, a difference in the trait that is greater than average and a phylogenetic distance that is smaller than average;iii. with different amino acids at the position, a difference in the trait that is smaller than average and a phylogenetic distance that is greater than average; and iv. with different amino acids at the position, a difference in the trait that is smaller than average and a phylogenetic distance that is smaller than average.

[0020] According to some embodiments, the weights are selected from the weights provided in Figure 7B.

[0021] According to some embodiments, the method further comprises receiving a cell of the organism in culture, producing the selected second amino acid at the site in place of the first amino acid, measuring an output indicative of the trait and confirming the second amino acid at the site is associated with the trait.

[0022] According to some embodiments, the biological trait is selected from lifespan, blood pressure, brain size, rate of memory decline, resting heart rate, basal metabolic rate, insulin sensitivity, glucose tolerance, incidence of neurodegenerative disease, incidence of cancer, cancer latency, DNA repair efficiency, infection resistance, telomere length, telomere attrition rate, protein turnover rate, proteostasis capacity, reactive oxidation species (ROS) production, mitochondrial DNA copy number, mitochondrial membrane potential, epigenetic clock, age of sexual maturity, reproductive span, reproductive output, post- reproductive lifespan, extrinsic mortality, longevity quotient, and metabolite levels.

[0023] According to another aspect, there is provided a84 method of extending the longevity of a human subject, the method comprising converting a lysine to an arginine at a position selected from those provided in Table 5 in the subject thereby extending the longevity of the human subject.

[0024] According to another aspect, there is provided a method of extending the longevity of a human subject, the method comprising: a. detecting a mutation in the human subject wherein the mutation produces a lysine at a position selected from those provided in Table 1 or Table 2, an arginine at a position selected from those provided in Table 3 or a glutamine at a position selected from those provided in Table 4; andb. converting a lysine at a position provided in Table 1 to an arginine, a lysine at a position provided in Table 2 to a glutamine, an arginine at a position provided in Table 3 to a lysine or a glutamine at a position provided in Table 4 to a lysine in the subject; thereby extending the longevity of a non-human subject.

[0025] According to another aspect, there is provided a method of extending the longevity of a subject, the method comprising, exogenously expressing a protein selected from those provided in Tables 1, 2, 3, 4 and 5 in the subject, thereby extending the longevity of a subject.

[0026] According to another aspect, there is provided a method of extending the longevity of a non-human subject, the method comprising: a. receiving DNA from the non-human subject and determining the DNA encodes a protein that comprises a lysine at a position selected from those provided in Table 1 or Table 2, comprises an arginine at a position selected from those provided in Table 3 or comprises a glutamine at a position selected from those provided in Table 4; and b. converting a lysine at a position selected from Table 1 to an arginine, a position selected from Table 2 to a glutamine, or a position selected from Table 3 or 4 to lysine, in the non-human subject; thereby extending the longevity of the non-human subject.

[0027] According to some embodiments, the extending longevity comprises treating an aging related disease.

[0028] According to some embodiments, an aging related disease is selected from a metabolic disease, a liver disease, a respiratory disease, a muscle disease, a proliferative disease and a neurological disease.

[0029] According to some embodiments, a. the metabolic or liver disease is selected from hypoglycemia, liver disease, hepatitis, hyperammonemia, hepatomegaly, diabetes, fatty liver disease, and hyperlipidemia; b. the muscle disease is myopathy;c. the respiratory disease is respiratory insufficiency; d. the proliferative disease is cancer; or e. the neurological disease is selected from polyneuropathy, and Alzheimer’s disease.

[0030] According to some embodiments, the converting comprises administering to the subject a genome editing compound targeted to the lysine residue and which edits the subject’s genome to have an arginine or glutamine at the location or targeted to the arginine or glutamine residue and which edits the subject’s genome to have a lysine residue at the location.

[0031] According to some embodiments, the genome editing compound is a CRISPR-CAS complex comprising a guide RNA (gRNA) complementary to the subject’s genomic sequence comprising the lysine, arginine or glutamine to be converted.

[0032] According to some embodiments, the non-human animal is an animal with a short life span, optionally wherein a short life span is below 30 years on average.

[0033] According to some embodiments, the location in Table 5 is selected from: Cox4il amino acid 164, Amacr amino acid 134, Fasn amino acid 1516, Fasn amino acid 1920, Shmtl amino acid 319, Bdhl amino acid 280, Bdhl amino acid 212, Hadha amino acid 406, Acatl amino acid 260, Shmt2 amino acid 464, Acads amino acid 339, Etfdh amino acid 257, Cptlb amino acid 41, and Decrl amino acid 319.

[0034] According to another aspect, there is provided a method for determining suitability of a human subject to have their longevity extended by a method of the inveniton, the method comprising receiving DNA from the human subject and determining the DNA encodes a protein that comprises a lysine at a position provided in Table 1, 2 or 5, an arginine at a position provided in Table 3 or a glutamine at a position provided in Table 4, wherein the presence of a lysine, arginine or glutamine at the position indicates the subject is suitable to have their longevity extended, thereby determining suitability.

[0035] Further embodiments and the full scope of applicability of the present invention will become apparent from the detailed description given hereinafter. However, it should be understood that the detailed description and specific examples, while indicating preferred embodiments of the invention, are given by way of illustration only, since various changes and modifications within the spirit and scope of the invention will become apparent to those skilled in the art from this detailed description. yBRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figures 1A-1C: General overview of PHARAOH computational tool. PHARAOH was designed to assess and identify correlations between amino acid (AA) changes and longer lifespan. (1A) The tool uses three types of data sets, mouse / human acetylomes (left box), a collection of mammalian orthologue proteins, and the phylogenetic tree of each species based on its orthologs set (middle box), and mammalian maximal lifespans (right box). (IB) The tool analyses change in amino acids at each acetylation site using pairwise sequence alignment between the acetylated peptide and the orthologous protein of each mammal (upper panel), generating a replacement matrix (middle panel). Subsequently, statistical analysis and validation against the phylogenetic tree are employed to examine the correlation between amino acid replacements and longer lifespan (lower panel). (1C) An example of a significant acetylation site on Bphl KI 88 from K-to-R analysis - the central portion of the figure showcases the OrthoFinder phylogenetic tree. Surrounding the tree, the inner circle represents the lifespan of each mammal, while the outer circle displays the amino acid found at that site in each mammal. Lysine (K) in blue, arginine (R) in red, glutamine (Q) in yellow, other AA in grey, and in black sites with no ortholog found. Color coding is shown on the right.

[0037] Figures 2A-2H: Enrichment analysis of mouse longevity-associated sites. (2A) Pathway enrichment analysis of K-to-R replacement sites (hypergeometric test and Benjamini-Hochberg p value correction by Metascape), and (2B) their corresponding MCODE clusters identified from PPI networks. (2C-2D) Similar analysis for K-to-Q replacements. (2E) CC analysis of K-to-R (green) and K-to-Q (blue) replacements. pLOGO of 7 amino acids flanking the mouse total acetylome (2F), the longevity associated acetylation sites replaced to (2G) R, or (2H) Q sites in the mammalian ortholog set. K-to-R = lysine to arginine replacement, K-to-Q = lysine to glutamine replacement. For f-h, significantly enriched AA are seen between the red lines.

[0038] Figures 3A-3D: CBS K386R enhances its H2S production activity. (3 A) Maximal lifespan of animals with conserved K vs. animals with K-to-R conversion at cystathionine beta synthase (CBS) lysine 386 (K386) (FDR<0.1). Dot color represents the maximal lifespan of a given species. Schematic representation of the various maximal lifespans of species with K vs. R is shown on the left. Species are ordered by lifespan along the Y axis. (3B) Sequence alignment of CBS K386 in mouse and R389 in human using Clustal Omega.(3C) Hydrogen sulfide (H2S) production capacity in HEK 293T cells lysates overexpressing flag-tagged mouse wild-type (WT), K386R, K386Q CBS or GFP control, using lead acetate strips. Tubulin was used as loading control. (3D) H2S production capacity of flag tagged human WT, R389K or R389Q CBS immunoprecipitated from HEK 293T cells. ImageJ quantification of c and d are shown on the right. Number of biological replicates: (3C) n= 3 and (3D) 4 independent experiments. H2S production was corrected to CBS expression levels. For c and d, one-way ANOVA with Bonferroni post hoc was used. For all graphs, bars represent mean ± SEM. *p<0.05, *** / ?<0.001. Source data are provided as a Source Data file.

[0039] Figures 4A-4F: Enrichment analysis of human longevity associated sites. (4A) Pathway enrichment analysis of R-to-K replacement sites (hypergeometric test and Benjamini-Hochberg p value correction by Metascape) and (4B) their DNA repair corresponding MCODE clusters identified from PPI networks. (4C) Pathway enrichment analysis of Q-to-K replacement sites using Metascape. (4D) CC analysis of R-to-K (green) and Q-to-K (blue), based on Gene Ontology (GO) annotations. (4E) Pathway enrichment analysis of R-to-K DNA metabolic processes and (4F) their corresponding DNA repair MCODE clusters identified from PPI networks.

[0040] Figures 5A-5C: Acquisition of acetylated K714 on hUSPIO regulates PCNA levels. (5A) Significant acetylation site on Ubiquitin Specific Peptidase 10 (USP10) lysine 714 (K714) from R-to-K analysis. The central portion of the left panel shows the OrthoFinder phylogenetic tree. Surrounding the tree, the inner circle represents the lifespan of each mammal, while the outer circle displays the amino acid found at the site in each mammal (K in blue, R in red, Q in yellow, other AA in grey, and in black sites with no ortholog found). Color code is shown on the right (left panel). Maximal lifespan of animals with conserved K vs. animals with R / Q / other AA at USP10 K714 (FDR<0.1). Dot color represents the maximal lifespan of a given species. (5B) Pairwise pearson correlation between USP10 and PCNA protein expression levels in lung (left), glioblastoma (central) and breast (right) cancers. Each dot represents a single patient, color code indicating tumor stage is shown on the right. (5C) Biological triplicates of PCNA levels in HCT116 cells overexpressing GFP, WT, K714R and K714Q USP10 with or without cycloheximide (CHX) treatment for 8 hours. Each condition was tested using biological triplicates. n= 3 independent experiments. Tubulin was used as loading control. ImageJ quantification is shown below. One-way ANOVA with Bonferroni post hoc was used. For all graphs, barsrepresent mean ± SEM. ** / ?<0.01, *** / ?<0.001, **** / ?<0.0001. Source data are provided as a Source Data file.

[0041] Figures 6A-6B: Overview of acetylation replacements. (6A) The molecular structure of acetylated lysine (acK) vs. arginine (R) and glutamine (Q). Created in BioRender. (6B) Percentage of conservation of mouse lysine residues, with (blue) or without (black) acetylation, relative to the human proteome. X axis shows the AA found in humans, and the Y axis shows the percentage of the exchange of lysine residue within the mouse proteome to the listed AA in the human proteome out of total lysine replacements.

[0042] Figures 7A-7C: General overview of statistical validation pipeline. To calculate the significance of acetylated lysine replacements, the validation focuses on a single row from the replacement matrix, representing a specific acetylation site. (7A) Two representative rows are shown on the top. The first step of the validation involves assessing the correlation between AA changes and lifespan. The row is divided into two groups based on the AA found at the acetylation site. A Mann-Whitney U test is then conducted to compare the lifespans of the two groups. (7B) To correct the influence of the phylogenetic tree, a normalization process is applied. All mammals are paired, and their relationships are categorized as described in the table. The statistical score for specific acetylation site for each pair is given based on the differences in their lifespans, tree distance and AA in at the acetylation site. The final score is computed as the sum of scores for all mammalian pairs associated with that site. To determine the significance of each site's score, a permutation test is performed. (7C) Significance is determined by selecting the sites that are significant in both the correlation analysis, and the normalization to the phylogenetic tree.

[0043] Figures 8A-8E: Comparative analysis of Zoonomia and OrthoFinder trees. (8A) Zoonomia (red and purple, 241 mammals) and OrthoFinder (green and yellow, 107 mammals) trees have an overlap of 62 mammals. Two phylogenetic trees were generated based on the overlapping 62-mammal and 107-mammal data. Analyses were conducted on the two different trees comparing (8B) mouse lysine to arginine replacement (K-to-R), (8C) mouse lysine to glutamine replacement (K-to-Q), (8D) human arginine to lysine replacement (Rto- K), and (8E) human glutamine to lysine replacement (Q-to-K), and the intersections are shown in pink.

[0044] Figures 9: Longevity associated mouse K-to-R / Q conversions across 107 mammals. Heatmap overview of significant acetylation sites identified by PHARAOH. Each row corresponds to a specific site. Each column represents one mammal out of the 107mammals’ database. K-to-R is shown on the top and K-to-Q on the bottom. A bar plot depicted above represents the maximal lifespan in captivity as provided by AnAge. The number of K residues in each column is indicated above each heatmap. A corresponding trendline is shown as well. Lysine (K) in blue, arginine (R) in red, glutamine (Q) in yellow, other amino acids (AA) in grey, sites with no ortholog found are in black.

[0045] Figures 10A-10B: Human orthologs of longevity-associated mouse acetylome pathway enrichment analysis. Pathway enrichment analysis (left panels) and PPIs (right panels) of the human orthologs of the longevity associated mouse (10A) K-to-R and (10B) K-to-Q acetylomes using Metascape.

[0046] Figures 11A-11C: pLOGO analysis of different cellular components in total mouse acetylome. pLOGO of 7 amino acids flanking around the total mouse acetylation found in (11 A) cytosol (11B) mitochondria and (11C) nucleus. Significantly enriched amino acids are seen between the red lines.

[0047] Figures 12A-12D: Human orthologs of longevity-associated mouse acetylome regulator and disease enrichment analyses. (12A-12B) Regulator enrichment analysis using TRRUST of (12A) lysine to arginine replacement (K-to-R), and (12B) lysine to glutamine replacement (K-to-Q) orthologous sites (green). (12C-12D) Summary of genedisease associations enrichment analysis using DisGeNET of (12C) K-to-R, and (12D) K-to-Q orthologous sites (purple).

[0048] Figure 13: CBS K386 replacements. An illustration of the variation in amino acids at the mouse CBS K386 from the K-to-R analysis - the central portion of the left panel shows the OrthoFinder phylogenetic tree. Surrounding the tree, the inner circle represents the lifespan of each mammal, while the outer circle displays the amino acid found at the site in each mammal: lysine (K) in pink, arginine (R) in purple, other amino acids in blue and no ortholog found in green. Color code is shown on the right.

[0049] Figures 14A-14J: Human longevity associated acetylome enrichment analysis. PPI networks of significant human longevity associated K-to-R (14A) and K-to-Q (14B) sites. pLOGO of 7 amino acids flanking around the total human acetylation (14C) and in cytosol (14D), mitochondria, (14E) and nucleus (14F), longevity associated acetylation sites replaced by R (14G), or Q (14H). (14I-14J) The corresponding TRRUST transcription factor analyses. For 14C-14H, significantly enriched AA are seen between the red lines.

[0050] Figure 15: Longevity associated human R / Q-to-K conversions across 107 mammals. Heatmap overview of significant acetylation sites identified by PHARAOH.Each row corresponds to a specific site. Each column represents one mammal out of the 107 mammals database. R-to-K is shown on the top and Q-to-K on the bottom. A bar plot depicted above represents the maximal lifespan in captivity as provided by AnAge. The number of K residues in each column is indicated above each heatmap. A corresponding trendline is shown as well. K in blue, R in red, Q in yellow, other AA in grey, and sites with no ortholog found are in black.

[0051] Figure 16: Loss of acetylated K714 on hUSPIO decreases total ubiquitination. Total ubiquitination levels in HCT116 cells overexpressing GFP, WT, K714R and K714Q USP10. Ponceau was used as a loading control. An ImageJ quantification is shown on the right, n = 3 independent experiments.DETAILED DESCRIPTION OF THE INVENTION

[0052] The present invention, in some embodiments, provides methods for identifying an amino acid alteration associated with a biological trait are provided. Methods of extending the longevity of a human subject and methods of extending the longevity of a non-human subject are also provided as are methods of determining is a subject is suitable to have their longevity extended.

[0053] By a first aspect, there is provided a method for identifying an amino acid associated with a biological trait, the method comprising: a. receiving a proteomic database for a plurality of species of mammals; b. receiving a database comprising a measure of the biological trait for the plurality of species of mammals; c. selecting a first amino acid and a species from the plurality of species of mammals; d. for sites of a residue of the first amino acid in a proteome of the selected species, establish a first group of species of mammals of the plurality that also have the first amino acid at the residue and a second group of species of mammals of the plurality that have a second amino acid at the residue; e. performing a statistical test to determine sites where in the measure of the biological trait in the first group and the second group are significantly different to produce a list of sites of the first or second amino acid potentially associated with the biological trait;f. for sites of the list determining if the potential association is due to phylogeny; and g. selecting a site and amino acid residue from the list whose association with the biological trait is not due to phylogeny; thereby identifying an amino acid associated with a biological trait.

[0054] In some embodiments, the method is a diagnostic method. In some embodiments, the method is a computational method. In some embodiments, the method is an in vitro method. In some embodiments, the method is an ex vivo method. In some embodiments, the method cannot be done in a human mind. In some embodiments, the method is a method for identifying an amino acid alteration associated with a biological trait. In some embodiments, the method is a method for identifying an amino acid mutation associated with a biological trait. In some embodiments, an amino acid is a plurality of amino acids. In some embodiments, an amino acid is a post-translational modification to an amino acid. I some embodiments, the amino acid is an amino acid that can be post-translationally modified.

[0055] As used herein, the term “biological trait” refers to a measurable aspect of mammalian biology. This can be any trait which can be quantified numerically. In some embodiments, the database comprises a quantifiable measure of the biological trait. In some embodiments, the measure is a numeric measure. Examples include but are not limited to lifespan, blood pressure, brain size, metabolic rate, and occurrence of cancer. Further, broad categories of traits include, for example, 1) lifespan and mortality measures (e.g., maximum lifespan, median lifespan, rate of senescence); 2) cardiovascular and metabolic traits (e.g., blood pressure, resting heart rate, basal metabolic rate, insulin sensitivity, glucose tolerance, lipid profile (HDL, LDL, triglycerides, etc.)); 3) neurological and cognitive traits (e.g., absolute brain size, relative brain size, cognitive performance, rate of memory decline, incidence of neurodegenerative disease); 4) cancer and disease resistance (e.g., cancer incidence, cancer latency, DNA repair efficiency, immune system robustness, immunosenescence, infection resistance); 5) cellular and molecular traits (e.g., telomere length, telomere attrition rate, protein turnover, proteostasis capacity, mitochondrial function (reactive oxidation species (ROS) production, copy number, membrane potential, etc.)); 6) epigenetic clock (e.g., epigenetic age estimated from DNA methylation, epigenetic age acceleration, Horvath clock, Hannum clock, PhenoAge clock, GrimAge clock, tissuespecific epigenetic clocks); 7) reproductive traits (e.g., age at sexual maturity, reproductive span, reproductive output, post-reproductive lifespan, menopause age, survival after fertilityends); 8) ecological and life-history traits (e.g., predation pressure, extrinsic mortality, longevity quotient, observed vs. expected lifespan for body size, sociality and cooperative behaviors); and 9) metabolites and small molecules associated with longevity (e.g., NAD+, glutathione, hydrogen sulfide, ketone bodies, polyamines, a-ketoglutarate, lactate, methionine, homocysteine). It will be understood that the method of the invention is sufficiently robust that any trait can be measured so long as it can be expressed numerically. In order to produce a score there must be a way to quantifiably compare the trait in various species. In some embodiments, the trait is longevity. It will be understood that for each trait one can look at it as a positive trait or a negative trait. Thus, the amino acid could be associated with long life or with short life. In both cases the amino acid is associated with life span. This will be true for any trait, as the amino acid could be associated with a large brain or a small brain, high blood pressure or low, etc.

[0056] In some embodiments, a plurality of species is at least two species. In some embodiments, a plurality is at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 107 or 110. Each possibility represents a separate embodiment of the invention. In some embodiments, a plurality is at least 10. In some embodiments, a plurality is at least 50. In some embodiments, a plurality is at least 100. In some embodiments, a plurality of species comprises at least 1 species with the trait and at least one species without the trait. In some embodiments, a plurality of species comprises a plurality of species with the trait and a plurality of species without the trait. In some embodiments, a plurality comprises at least 5 species with the trait and at least 5 species without the trait. In some embodiments, a plurality comprises at least 10 species with the trait and at least 10 species without the trait. In some embodiments, a plurality comprises at least 20 species with the trait and at least 20 species without the trait. In some embodiments, a plurality comprises at least 30 species with the trait and at least 30 species without the trait.

[0057] In some embodiments, the plurality of species comprises humans. In some embodiments, humans are Homo sapiens. In some embodiments, selected species is Homo sapiens. In some embodiments, the selected species is an organism. In some embodiments, the selected species is a species with the trait. In some embodiments, the selected species is a selected species without the trait. In some embodiments, the selected species is a rodent. In some embodiments, the selected species is a mouse. In some embodiments, the mouse is selected from Mus spicilegus and Mus musculus. In some embodiments, the plurality of species is selected from Mus spicilegus, Rattus norvegicus, Mesocricetus auratus, Mus musculus, Monodelphis domestica, Cricetulus griseus, Microtus ochrogaster, Molossusmolossus, Jaculus jaculus, Ictidomys tridecemlineatus, Pipistrellus kuhlii, Peromyscus maniculatus bairdii, Phyllostomus discolor, Dipodomys ordii, Neotoma lepida, Mustela putorius furo, Neovison vison, Erinaceus europaeus, Cavia porcellus, Tupaia chinensis, Oryctolagus cuniculus, Sarcophilus harrisii, Marmota monax, Octodon degus, Sciurus vulgaris, Fukomys damarensis, Carlito syrichta, Nyctereutes procyonoides, Chinchilla lanigera, Marmota marmota marmota, Prolemur simus, Microcebus murinus, Bos taurus, Otolemur garnettii, Pteropus alecto, Acinonyx jubatus, Canis lupus familiaris, Suricata suricatta, Capra hircus, Pteropus vampyrus, Vulpes vulpes, Phascolarctos cinereus, Omithorhynchus anatinus, Ovis aries, Callithrix jacchus, Rousettus aegyptiacus, Odocoileus virginianus texanus, Muntiacus reevesi, Castor canadensis, Lipotes vexillifer, Vicugna pacos, Bos mutus, Panthera tigris altaica, Lynx canadensis, Sus scrofa, Enhydra lutris kenyoni, Puma concolor, Panthera leo, Camelus dromedarius, Rhinopithecus roxellana, Orycteropus afer afer, Felis catus, Aotus nancymaae, Vombatus ursinus, Saimiri boliviensis boliviensis, Rhinolophus ferrumequinum, Heterocephalus glaber, Propithecus coquereli, Cervus elaphus hippelaphus, Chlorocebus sabaeus, Bison bison bison, Myotis lucifugus, Ursus americanus, Colobus angolensis palliatus, Theropithecus gelada, Ailuropoda melanoleuca, Myotis myotis, Papio anubis, Macaca nemestrina, Macaca fascicularis, Mandrillus leucophaeus, Macaca mulatta, Delphinapterus leucas, Myotis brandtii, Crocuta crocuta, Ursus maritimus, Nomascus leucogenys, Cercocebus atys, Sapajus apella, Equus asinus asinus, Balaenoptera acutorostrata scammoni, Monodon monoceros, Tursiops truncatus, Cebus imitator, Pan paniscus, Equus caballus, Diceros bicomis minor, Pongo abelii, Gorilla gorilla gorilla, Loxodonta africana, Pan troglodytes, Trichechus manatus latirostris, Eschrichtius robustus, Physeter macrocephalus, Balaenoptera musculus, Balaenoptera physalus, and Homo sapiens. These are the 107 species exemplified in the Examples provided herein below. In some embodiments, the plurality comprises all 107 species listed hereinabove.

[0058] Databases with consensus protein sequences, biological traits and phylogenetic distances for a variety of organisms / species are well known in the art and any such database may be used. Examples include, but are not limited to the Uniprot protein database, the RefSeq protein database, the Animal ageing and longevity database, phylogenetic trees the Zoonomia project tree and the OrthoFinder-based tree. It will be understood that the appropriate database can be selected based on the trait being investigated. In some embodiments, the method further comprises aligning orthologous proteins from different species. In some embodiments, the residue in a first species is its corresponding alignedresidue in another species. In some embodiments, the alignment is a pairwise sequence alignment. In some embodiments, the alignment is a pairwise sequence alignment of orthologs. It will be understood that the amino acid numbering may not be conserved between species, but by aligning (i.e., by OrthoFinder) one can determine the corresponding residue in another species. In some embodiments, the method further comprises producing a multisequence alignment to determine the same residue in each species. In some embodiments, the method further comprises producing a multisequence alignment to determine the same residue in the selected species and another species of the plurality.

[0059] In some embodiments, the first amino acid is an amino acid to test. In some embodiments, the first and second amino acid are different amino acids. In some embodiments, the first and second amino acid are not the same amino acid. In some embodiments, the first amino acid is an amino acid that can be post-translationally modified. In some embodiments, the first amino acid is the amino acid that is altered. In some embodiments, the first amino acid is mutated into the second amino acid in some species. In some embodiments, the first amino acid is a canonical amino acid. In some embodiments, the first amino acid is selected from Alanine, Arginine, Asparagine, Aspartic Acid, Cysteine, Glutamic Acid, Glutamine, Glycine, Histidine, Isoleucine, Leucine, Lysine, Methionine, Phenylalanine, Proline, Serine, Threonine, Tryptophan, Tyrosine, and Valine. In some embodiments, the first amino acid is lysine. In some embodiments, the first amino acid is selected from serine, threonine and tyrosine and the post-translational modification (PTM) is phosphorylation. In some embodiments, the first amino acid is selected from lysine and arginine and the PTM is acetylation. In some embodiments, the first amino acid is lysine and the PTM is acetylation. In some embodiments, the first amino acid is selected from lysine and arginine and the PTM is methylation. In some embodiments, the first amino acid is selected from lysine and arginine and the PTM is ubiquitination. In some embodiments, the first amino acid is cysteine and the PTM is palmitoylation or S-nitrosylation. In some embodiments, the first amino acid is selected from asparagine and glutamine and the PTM is glycosylation. Other examples of PTM mimics include, for example, aspartic acid or glutamic acid mimicking phosphorylated serine, alanine mimicking unphosphorylated serine. In some embodiments, the first amino acid is lysine and the second amino acid is arginine. In some embodiments, the first amino acid is lysine and the second amino acid is glutamine.

[0060] In some embodiments, the first amino acid can be post-translationally modified and the second amino acid mimics the modified or unmodified state of the first amino acid. Insome embodiments, the second amino acid cannot be post-translationally modified. In some embodiments, the second amino acid cannot be post-translationally modified with the modification that is made to the first amino acid. In some embodiments, the first amino acid is lysine, the PTM is acetylation and arginine mimics the unacetylated lysine. In some embodiments, the first amino acid is lysine, the PTM is acetylation and glutamine mimics the acetylated lysine. In some embodiments, the first amino acid is serine, the PTM is phosphorylation and aspartic acid or glutamic acid mimics phosphorylated serine. In some embodiments, the first amino acid is serine, the PTM is phosphorylation and alanine mimics unphosphorylated serine. In some embodiments, the first amino acid is threonine, the PTM is phosphorylation and aspartic acid or glutamic acid mimics phosphorylated threonine. In some embodiments, the first amino acid is threonine, the PTM is phosphorylation and glutamic acid mimics phosphorylated threonine. In some embodiments, the first amino acid is threonine, the PTM is phosphorylation and asparagine mimics unphosphorylated threonine. In some embodiments, the first amino acid is threonine, the PTM is phosphorylation and glutamate or aspartate mimics unphosphorylated threonine. In some embodiments, the first amino acid is tyrosine, the PTM is phosphorylation and aspartic acid or glutamic acid mimics phosphorylated tyrosine. In some embodiments, the first amino acid is tyrosine, the PTM is phosphorylation and phenylalanine mimics unphosphorylated tyrosine.

[0061] In some embodiments, the method further comprises receiving an atlas of marks of the PTM on the first amino acid for the mammal. In some embodiments, the mammal is the selected mammal. In some embodiments, the atlas is of marks of the PTM on the first amino acid for the organism. In some embodiments, the atlas is an acetylome atlas. In some embodiments, the atlas is a methylome atlas. In some embodiments, the atlas is a phosphorylome atlas. In some embodiments, the atlas is received before step (d). In some embodiments, the atlas is received before said calculating. In some embodiments, the method further comprises identifying in the atlas residues of the first amino acid that can be modified in the selected species. In some embodiments, the method further comprises identifying in the atlas modifiable residues of the first amino acid in the selected species. In some embodiments, the calculating is performed for each identified modifiable residue. In some embodiments, the first amino acid is lysine and the calculating is performed for each aceylatable residue.

[0062] In some embodiments, the group of species of mammals is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45 or 50 species of mammals. Each possibility represents aseparate embodiment of the invention. In some embodiments, the group of species of mammals is at least 2 species. In some embodiments, the group of species of mammals is at least 5 species. In some embodiments, the group of species of mammals is at least 10 species. In some embodiments, the group of species of mammals is at least 20 species. In some embodiments, the group of species of mammals is at least 50 species. In some embodiments, a first group of species has the first amino acid at the residue / site. In some embodiments, a first group of species is all the species of the plurality that has the first amino acid at the residue / site. In some embodiments, a second group of species does not have the first amino acid at the residue / site. In some embodiments, a second group of species is all the species of the plurality that does not have the first amino acid at the residue / site. In some embodiments, a second group of species has the second amino acid at the residue / site. In some embodiments, a second group of species is all the species of the plurality that has the second amino acid at the residue / site. It will be understood that determining the same residue / site is done via sequence alignment of the orthologs of the protein from different species. Thus, the same site is not the same numeric site, but rather its equivalent site by alignment. Websites and programs for calculating sequence alignment are well known in the art and any may be used.

[0063] In some embodiments, a first group and a second group are established for a plurality sites of the first amino acid. In some embodiments, a plurality of sites is at least 2, 5, 10, 15, 20, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 5000, 10000, 50000, 100000, 500000, 1000000, 5000000, 10000000 or 50000000 sites. Each possibility represents a separate embodiment of the invention. In some embodiments, for each site of the first amino acid a first group and a second group is established. In some embodiments, establishing is identifying. In some embodiments, establishing is determining. In some embodiments, establishing is producing.

[0064] In some embodiments, the statistical test is a T-test. In some embodiments, the statistical test is a Mann-Whitney test. Tests for determining statistical significance are well known in the art and any such test may be used to determine if the differences in the biological trait are statistically significant. In some embodiments, the statistical test is performed on the average of the quantifiable measure in the first group and average of the quantifiable measure in the second group. In some embodiments, the statistical test is performed on the set of quantifiable measures in the first group and the set of quantifiable measures in the second group.

[0065] In some embodiments, sites wherein the quantifiable measure of the biological trait in the first group is significantly higher than the quantifiable measure of the biological trait in the second group are sites of the first amino acid potentially associated with the biological trait. In some embodiments, all the sites wherein the quantifiable measure of the biological trait in the first group is significantly higher than the quantifiable measure of the biological train in the second group are a list of sites of the first amino acid potentially associated with the biological trait. In some embodiments, potentially associated is associated. In some embodiments, performing the statistical test produces the list. In some embodiments, the list of first amino acids is a first list. In some embodiments, sites wherein the quantifiable measure of the biological trait in the second group is significantly higher than the quantifiable measure of the biological trait in the first group are sites of the second amino acid potentially associated with the biological trait. In some embodiments, all the sites wherein the quantifiable measure of the biological trait in the second group is significantly higher than the quantifiable measure of the biological train in the first group are a list of sites of the second amino acid potentially associated with the biological trait. In some embodiments, potentially associated is associated. In some embodiments, performing the statistical test produces the list. In some embodiments, the list of second amino acids is a second list. In some embodiments, a first and second list are produced.

[0066] In some embodiments, the method comprises for a plurality of sites of the first list determining if the potential association is due to phylogeny. In some embodiments, the method comprises for each site of the first list determining if the potential association is due to phylogeny. In some embodiments, the method comprises for a plurality of sites of the second list determining if the potential association is due to phylogeny. In some embodiments, the method comprises for each site of the second list determining if the potential association is due to phylogeny. In some embodiments, the method comprises for a plurality of sites of the first list and a plurality of sites of the second list determining if the potential association is due to phylogeny In some embodiments, the method comprises for each site of the first list and each site of the second list determining if the potential association is due to phylogeny. In some embodiments, due to phylogeny is due to phylogenetic distance.

[0067] In some embodiments, the method comprises selecting at least one site from the first list whose association of the first amino acid at the site with the biological trait is not due to phylogeny. In some embodiments, the method comprises selecting at least one site from the second list whose association of the second amino acid at the site with the biological trait isnot due to phylogeny. In some embodiments, the method comprises selecting at least one site from either the first or second list whose association of the amino acid at that site with the biological trait is not due to phylogeny. In some embodiments, the method identifies an amino acid associated with the biological trait. In some embodiments, the method identifies an amino acid residue with the biological trait. In some embodiments, the method identifies an amino acid at a specific site associated with the biological trait.

[0068] In some embodiments, determining if the potential association is due to phylogeny comprises producing a site score for a site. In some embodiments, the site score is produced for a plurality of sites. In some embodiments, the site score is produced for each site. In some embodiments, each site is every site. In some embodiments, the method of producing the site score is the method provided in Figure 7B. In some embodiments, the site score is determined by the formula: Site score = ^each pair of mammals (weight * [measure 1 - measure 2] * tree distance).

[0069] In some embodiments, producing a site score comprises, for a pair of species of the plurality determining the difference between the quantifiable measure. In some embodiments, the difference is the quantifiable measure in the first species of the pair minus the quantifiable measure in the second species of the pair. In some embodiments, the quantifiable measure is taken from the database. In some embodiments, the difference is determined for a plurality of pairs of species. In some embodiments, the difference is determined for each pair of species in the plurality. It will be understood that each pair means every permutation of two different species paired together. In some embodiments, producing a site score comprises, for a pair of species of the plurality determining the phylogenetic distance between the two species of the pair. In some embodiments, the phylogenetic distance is determined for a plurality of pairs. In some embodiments, the phylogenetic distance is determined for each pair. In some embodiments, each pair is all pairs. In some embodiments, producing a site score comprises multiplying the difference in quantifiable measure by the phylogenetic distance to produce a product. In some embodiments, product is produced for a plurality of pairs. In some embodiments, a product is produced for each pair.

[0070] In some embodiments, producing a site score further comprises multiplying the product by a weight to produce a weighted product. In some embodiments, the weight is representative of a site’s contribution to the hypothesis that the first amino acid at the site is associated with the trait. In some embodiments, the weight is representative of a site’s contribution to the hypothesis that the second amino acid at the site is associated with thetrait. In some embodiments, the weight is representative of a site’s contribution to the hypothesis that the first amino acid or the second amino acid at the site is associated with the trait. In some embodiments, a positive weight is given to sites that strengthen the hypothesis and a negative weight is given to sites that do not weaken the hypothesis. In some embodiments, the weight is representative of how much a site supports the first amino acid being associated with the trait. In some embodiments, the weight is representative of how much a site supports the second amino being associated with the trait. In some embodiments, a positive weight is given to sites that support the second amino being associated with the trait. In some embodiments, a negative weight is given to sites that do not support the second amino being associated with the trait. In some embodiments, do not support is go against.

[0071] In some embodiments, the weight is positive for pairs with the same amino acid at the position, a difference in the trait that is smaller than average and a phylogenetic distance that is greater than average. In some embodiments, greater is significantly greater. In some embodiments, smaller is significantly smaller. In some embodiments, significantly is statistically significantly. In some embodiments, the weight is positive for pairs with different amino acids at the position, a difference in the trait that is larger than average and a phylogenetic distance that is smaller than average. In some embodiments, the positive weights are taken from the table in Figure 7B. In some embodiments, the weight for the same amino acid at the position, a difference in the trait that is smaller than average and a phylogenetic distance that is greater than average is 0.42. In some embodiments, the weight for different amino acids at the position, a difference in the trait that is larger than average and a phylogenetic distance that is smaller than average is 0.03.

[0072] In some embodiments, a weight of zero is given to sites that do not contribute to the hypothesis. In some embodiments, a weight of zero is given to sites that do not support the first amino acid being associated with the trait. In some embodiments, a weight of zero is given to sites that do not support the second amino acid being associated with the trait. In some embodiments, the weight is zero for pairs with the same amino acid at the position, a difference in the trait that is smaller than average and a phylogenetic distance that is smaller than average. In some embodiments, the weight is zero for pairs with different amino acids at the position, a difference in the trait that is larger than average and a phylogenetic distance that is larger than average.

[0073] In some embodiments, the weight is negative for pairs with the same amino acid at the position, a difference in the trait that is greater than average and a phylogenetic distance that is greater than average. In some embodiments, the weight is negative for pairs with thesame amino acid at the position, a difference in the trait that is greater than average and a phylogenetic distance that is smaller than average. In some embodiments, greater is significantly greater. In some embodiments, smaller is significantly smaller. In some embodiments, significantly is statistically significantly. In some embodiments, the weight is negative for pairs with different amino acids at the position, a difference in the trait that is smaller than average and a phylogenetic distance that is smaller than average. In some embodiments, the weight is negative for pairs with different amino acids at the position, a difference in the trait that is smaller than average and a phylogenetic distance that is greater than average. In some embodiments, the negative weights are taken from the table in Figure 7B. In some embodiments, the weight for the same amino acid at the position, a difference in the trait that is greater than average and a phylogenetic distance that is greater than average is -0.144. In some embodiments, the weight for the same amino acid at the position, a difference in the trait that is greater than average and a phylogenetic distance that is smaller than average is -0.03. In some embodiments, the weight for different amino acids at the position, a difference in the trait that is smaller than average and a phylogenetic distance that is smaller than average is -0.32. In some embodiments, the weight for different amino acids at the position, a difference in the trait that is smaller than average and a phylogenetic distance that is larger than average is -0.42.

[0074] In some embodiments, determining if the potential association is due to phylogeny further comprises performing a permutation test to produce random site scores for the site. In some embodiments, the permutation test produces a random distribution of site scores. In some embodiments, random site scores are produced and placed into a standard distribution. In some embodiments, random site scores are produced by assigning random weights. In some embodiments, all possible random site scores are produced by assigning random weights. In some embodiments, a standard distribution of scores is produced by assigning random weights to produce random scores. In some embodiments, a site score is considered statistically significant if it is more than 1 standard deviation from the mean of the standard deviation. In some embodiments, a site score is considered statistically significant if it is more than 2 standard deviations from the mean of the standard deviation. The use of a standard distribution to determine a P-value and thereby statistical significance is well known in the art and any such method may be used. In some embodiments, a site score with a p-value of less than 0.05 is considered statistically significant. In some embodiments, a statistically significant score indicates the potential association is not due to phylogeny. In some embodiments, statistical significance is as compared to the random / standarddistribution. In some embodiments, a significant site score indicates the amino acid at the site is associated with the trait. In some embodiments, a significant site score indicates the second amino acid at the site is associated with the trait. In some embodiments, a significant site score indicates the second amino acid at the site is associated with the trait not due to phylogeny. In some embodiments, a site from the list with a significant site score is selected. In some embodiments, an amino acid residue at a site from the first list with a significant site score is selected. In some embodiments, an amino acid residue at a site from the second list with a significant site score is selected.

[0075] In some embodiments, the site score is a numeric representation of the biological trait in the group. In some embodiments, the first site score is the numeric representation of the biological trait in the first group and the second site score is the numeric representation of the biological trait in the second group. In some embodiments, the numeric representation is an average for the group. In some embodiments, the numeric representation is a sum for the group normalized to the number of species in the group. In some embodiments, if the second site score is greater than the first site score then the second amino acid is associated with the biological trait. In some embodiments, if the first site score is greater than the second site score then the first amino acid is associated with the biological trait. In some embodiments, greater than is significantly greater than. In some embodiments, significance is based on statistical test. In some embodiments, the test is the Mann-Whitney test. The Mann-Whitney test was used in the Examples hereinbelow, but any appropriate statistical test selected by a skilled artisan can be used to determine significance.

[0076] In some embodiments, the site score is a numeric representation of the similarity in the trait between the selected species and the group of species. In some embodiments, the first site score is a numeric representation of the similarity in the trait between the selected species and the first group of species. In some embodiments, the second site score is a numeric representation of the similarity in the trait between the selected species and the second group of species. In some embodiments, the similarity is similarity in the quantifiable measure. In some embodiments, the score is produced based on the database of biological trait in the species of mammals. In some embodiments, the score is produced based on the quantifiable measure of the trait for each species in the database. In some embodiments, the similarity is based on the difference between the quantifiable measure for the selected species and the quantifiable measure for the group of species. In some embodiments, the difference is the numeric difference between the measure for the selected species and the average of the measure in the group. In some embodiments, the difference is the sum of thedifferences between the measure for the selected species and the measure for each species of the group.

[0077] In some embodiments, the difference in the measure is multiplied by the phylogenetic distance between the selected species and the species of the group. In some embodiments, the phylogenetic distance is an inverse of the distance. It will be understood that more weight is given to differences that are close in phylogenetic distance and less weight to farther distances. This is because the greater the distance the greater the chance the amino acid difference is by chance or drift. When the two species are close in distance but different in the trait the amino acid change is more likely to be responsible. In some embodiments, for each species of the group the difference is multiplied by the phylogenetic distance between the selected species and that species of the group. In some embodiments, the sum is the sum of the differences multiplied by the distance. Thus, for each species of the group a difference in the measure of the trait is calculated and a distance is calculated. The difference and the distance are multiplied and all the products are added together to get the score. In some embodiments, added together is summed.

[0078] In some embodiments, the selecting comprises selecting the second amino acid at a site as being associated with the biological trait when the second site score is greater than the first site score. In some embodiments, greater is significantly greater. In some embodiments, a statistical test is used. In some embodiments, the selecting comprises selecting the first amino acid at a site as being associated with the biological trait when the first site score is greater than the second site score. In some embodiments, the method comprises determining the second site score is greater than the first site score and selecting the second amino acid at that site as being associated with the biological trait. In some embodiments, the method comprises determining the first site score is greater than the second site score and selecting the first amino acid at that site as being associated with the biological trait.

[0079] In some embodiments, the method further comprises selecting from the selected residues residues for which the change in amino acids is not due to phylogenetic distance. In some embodiments, the method further comprises selecting from the second amino acid at a specific site, residues for which the change in amino acids is not due to phylogenetic distance. In some embodiments, residues are a second amino acid at a specific site. In some embodiments, the method comprises normalizing each site score to phylogenetic distance. In some embodiments, a permutation test is applied to each site to obtain a p value. In some embodiments, the permutation test chooses random weights. In some embodiments, the pvalue is used to determine the significance of each site’s score. In some embodiments, a statistical test is performed to determine if the site change is due to phylogenetic distance. In some embodiments, a statistical test is performed to determine if the difference in biological trait between the two groups is due to phylogenetic distance. In some embodiments, the second selecting is sub-selecting. In some embodiments, the method further comprises subselecting a second amino acid at a site where the difference in biological trait between the first group and the second group is not due to phylogenetic distance. In some embodiments, a significant p value indicates the selected site is selected not due to phylogenetic distance.

[0080] In some embodiments, the method further comprises determining a weight for each pair of selected species and species of the plurality of species. In some embodiments, the weight is for each pair of organism and species from the plurality of species. In some embodiments, the weight is applied to the score. In some embodiments, the weight is applied to the difference. In some embodiments, the weight is applied to the product of the difference and distance. In some embodiments, the weight is the probability that the selected species and the species of the plurality would have a difference in the trait that is greater than the average. In some embodiments, the weight is the probability that the selected species and the species of the plurality would have a difference in the trait that is greater than the average and a phylogenetic distance that is smaller than the average. In some embodiments, the weight is the probability that the selected species and the species of the plurality would have a difference in the trait that is smaller than the average. In some embodiments, the weight is the probability that the selected species and the species of the plurality would have a difference in the trait that is smaller than the average and a phylogenetic distance that is greater than the average. In some embodiments, the probability is the probability regardless of the amino acid at the site. In some embodiments, the weight is calculated based on Figure 7B. In some embodiments, the weight is selected from the weights provided in Figure 7B. In some embodiments, a positive weight is given when the amino acid at the site is the same, the difference in trait is smaller than average and the phylogenetic distance is larger than average. In some embodiments, the positive weight is about 0.42. In some embodiments, a positive weight is given when the amino acid at the site is different (the first and second amino acid are different), the difference in trait is larger than average and the phylogenetic distance is smaller than average. In some embodiments, the positive weight is about 0.03. In some embodiments, a negative weight is given when the amino acid at the site is different (the first and second amino acid are different), the difference in trait is smaller than average and the phylogenetic distance is smaller than average. In some embodiments, the negativeweight is about -0.32. In some embodiments, a negative weight is given when the amino acid at the site is different (the first and second amino acid are different), the difference in trait is smaller than average and the phylogenetic distance is larger than average. In some embodiments, the negative weight is about -0.42. In some embodiments, a negative weight is given when the amino acid at the site is the same, the difference in trait is bigger than average and the phylogenetic distance is larger than average. In some embodiments, the negative weight is about -0.144. In some embodiments, a negative weight is given when the amino acid at the site is the same, the difference in trait is bigger than average and the phylogenetic distance is smaller than average. In some embodiments, the negative weight is about -0.03. In some embodiments, the weight is zero when the amino acid at the site is the same, the difference in trait is smaller than average and the phylogenetic distance is smaller than average. In some embodiments, the weight is zero when the amino acid at the site is different (the first and second amino acid are different), the difference in trait is bigger than average and the phylogenetic distance is bigger than average.

[0081] In some embodiments, the method further comprises testing a selected site and confirming it is associated with the trait. In some embodiments, the testing comprises receiving a cell from the selected species. In some embodiments, the cell is in culture. In some embodiments, the cell is in vitro. In some embodiments, the cell is ex vivo. In some embodiments, the cell is in vivo. In some embodiments, the testing comprises mutating the first amino acid at the selected site to be the second amino acid. In some embodiments, the testing comprises producing the selected second amino acid at the site in place of the first amino acid. Methods of mutagenesis and in particular site directed point mutagenesis are well known. Further, as the genetic code is known any amino acid can be converted to any other amino acid though a 1, 2 or 3 nucleotide change. In some embodiments, the mutagenesis is my genome editing. In some embodiments, genome editing is CRISPR-CAS editing. In some embodiments, the testing comprises measuring output in the cell with and without mutagenesis. In some embodiments, the testing comprises measuring output in the cell with the first amino acid at the site and the cell with the second amino acid at the site. In some embodiments, the output is an output indicate of the trait. In some embodiments, the second amino acid is associated with the trait and the method comprises confirming that the second amino acid at the site produced the output indicative of the trait. In some embodiments, the first amino acid is associated with the trait and the method comprises confirming that the first amino acid at the site produced the output indicative of the trait. In some embodiments, produced the output indicative of the trait is produced the output morehighly. In some embodiments, the first amino acid is associated with the trait and the method comprises confirming that the cell with the first amino acid at the site produced the output indicative of the trait at a higher level than the cell with the second amino acid at the site. In some embodiments, the second amino acid is associated with the trait and the method comprises confirming that the cell with the second amino acid at the site produced the output indicative of the trait at a higher level than the cell with the first amino acid at the site. In some embodiments, the method comprises selecting a second amino acid and site that is confirmed.

[0082] By another aspect, there is provided a method of extending the longevity of a human subject, the method comprising converting a lysine to an arginine at a position selected from those provided in Table 5 in the subject, thereby extending the longevity of a human subject.

[0083] By another aspect, there is provided a method of extending the longevity of a human subject, the method comprising: a. detecting in the subj ect a lysine at a position selected from those provided in Table 1; and b. converting the lysine at a position provided in Table 1 to an arginine in the subject, thereby extending the longevity of a non-human subject.

[0084] By another aspect, there is provided a method of extending the longevity of a human subject, the method comprising: a. detecting in the subj ect a lysine at a position selected from those provided in Table 2; and b. converting the lysine at a position provided in Table 2 to a glutamine in the subject, thereby extending the longevity of a non-human subject.

[0085] By another aspect, there is provided a method of extending the longevity of a human subject, the method comprising: a. detecting in the subject an arginine at a position selected from those provided in Table 3; and b. converting the arginine at a position provided in Table 3 to a lysine in the subject,thereby extending the longevity of a non-human subject.

[0086] By another aspect, there is provided a method of extending the longevity of a human subject, the method comprising: a. detecting in the subject a glutamine at a position selected from those provided in Table 4; and b. converting the glutamine at a position provided in Table 4 to a lysine in the subject, thereby extending the longevity of a non-human subject.

[0087] It will be understood that the sites in Table 5 are sites where the residue in humans is a lysine. In some embodiments, the method comprises confirming the a lysine at a site selected from those provided in Table 5 and converting the confirmed lysine to arginine. In some embodiments, the conversion is in a cell from the subject. In some embodiments, the conversion is in cells of the subject. In some embodiments, the sites in Table 5 are in proteins involved in fat metabolism. In some embodiments, the sites in Table 5 are selected from: Cox4il amino acid 164, Amacr amino acid 134, Fasn amino acid 1516, Fasn amino acid 1920, Shmtl amino acid 319, Bdhl amino acid 280, Bdhl amino acid 212, Hadha amino acid 406, Acatl amino acid 260, Shmt2 amino acid 464, Acads amino acid 339, Etfdh amino acid 257, Cptlb amino acid 41, and Decrl amino acid 319. In some embodiments, the site in Table 5 is Cox4il amino acid 164. In some embodiments, the site in Table 5 is Amacr amino acid 134. In some embodiments, the site in Table 5 is Fasn amino acid 1516. In some embodiments, the site in Table 5 is Fasn amino acid 1920. In some embodiments, the site in Table 5 is Shmtl amino acid 319. In some embodiments, the site in Table 5 is Bdhl amino acid 280. In some embodiments, the site in Table 5 is Bdhl amino acid 212. In some embodiments, the site in Table 5 is Hadha amino acid 406. In some embodiments, the site in Table 5 is Acatl amino acid 260. In some embodiments, the site in Table 5 is Shmt2 amino acid 464. In some embodiments, the site in Table 5 is Acads amino acid 339. In some embodiments, the site in Table 5 is Etfdh amino acid 257. In some embodiments, the site in Table 5 is Cptlb amino acid 41. In some embodiments, the site in Table 5 is Decrl amino acid 319. It will be understood that these names refer to well known proteins / genes who can be found in the Uniprot of Reference gene databases. The full names of all the proteins provided in all the tables can be found in such databases, the contents of which are hereby incorporated by reference in their entirety. The full names of these proteins are: Cox4il - Cytochrome c oxidase subunit 4 isoform 1, Amacr - Alpha-methylacyl-CoA racemase, Fasn- Fatty acid synthase, Shmtl - Serine hydroxymethyltransferase, cytosolic (SHMT1), Bdhl- D-P-hydroxybutyrate dehydrogenase, mitochondrial, Hadha - Trifunctional enzyme subunit alpha, mitochondrial (hydroxy acyl -Co A dehydrogenase / 3-ketoacyl-CoA thiolase / enoyl-CoA hydratase alpha subunit), Acatl - Acetyl-CoA acetyltransferase, mitochondrial, Shmt2 - Serine hydroxymethyltransferase, mitochondrial (SHMT2), Acads- Short-chain specific acyl-CoA dehydrogenase, mitochondrial, Etfdh - Electron transfer flavoprotein-ubiquinone oxidoreductase (ETF dehydrogenase), Cptlb - Carnitine O- palmitoyltransferase 1, muscle isoform (mitochondrial outer membrane), and Decrl - 2,4- dienoyl-CoA reductase, mitochondria.

[0088] The proteins in Table 1 are as follows: Aadat - Aminoadipate aminotransferase, Aass- Alpha-aminoadipic semialdehyde synthase, Abcd3 - ATP -binding cassette sub-family D member 3, Acad 12 - Acyl-CoA dehydrogenase family member 12, Acad 10 - Acyl-CoA dehydrogenase family member 10, Acad8 - Short / branched-chain specific acyl-CoA dehydrogenase, Acadm - Medium-chain specific acyl-CoA dehydrogenase, Acads - Shortchain specific acyl-CoA dehydrogenase, Acadvl - Very long-chain specific acyl-CoA dehydrogenase, Acatl - Acetyl-CoA acetyltransferase, mitochondrial, Acnatl - Acyl-CoA N-acyltransferase 1, Acnat2 - Acyl-CoA N-acyltransferase 2, Acol - Cytoplasmic aconitate hydratase, Aco2 - Aconitate hydratase, mitochondrial, Acox - Peroxisomal acyl-coenzyme A oxidase, Acp5 - Tartrate-resistant acid phosphatase type 5, Acsf2 - Acyl-CoA synthetase family member 2, Acsf3 - Acyl-CoA synthetase family member 3, Acsf - Acyl-CoA synthetase family protein, Acsll - Long-chain-fatty-acid — CoA ligase 1, Acsl5 - Long- chain-fatty-acid — CoA ligase 5, Adck3 - AarF domain-containing kinase 3, Adsl4 - Adenylosuccinate lyase, Agll - Glycogen debranching enzyme, Agpat - 1-acyl-sn-glycerol- 3 -phosphate acyltransferase, Agxt - Alanine-glyoxylate aminotransferase, Aimpl - Aminoacyl tRNA synthase complex -interacting multifunctional protein 1, Akl - Adenylate kinase isoenzyme 1, Akrlal - Aldo-keto reductase family 1 member Al, Aldhlal - Retinal dehydrogenase 1, Aidhill - Cytosolic 10-formyltetrahydrofolate dehydrogenase, Aldh5al- Succinate-semialdehyde dehydrogenase, Als21 - Alpha-L-arabinofuranosidase 21 (putative), Amt - Aminomethyltransferase, Ankrdl l - Ankyrin repeat domain-containing protein 11, Anpep - Aminopeptidase N, Anxa6 - Annexin A6, Aplg2 - AP-1 complex subunit gamma-2, Aplgl - AP-1 complex subunit gamma-1, Aridlbl - AT -rich interactive domain-containing protein IB, Arid5a - AT -rich interactive domain-containing protein 5A, Ascc3 - Activating signal cointegrator 1 complex subunit 3, Assl - Argininosuccinate synthase 1, Atplal - Sodium / potassium -transporting ATPase subunit alpha- 1, Atp2al -Sarcoplasmic / endoplasmic reticulum calcium ATPase 1, Atp4a - Potassium-dependent sodium / calcium ATPase subunit alpha, Atp5cl - ATP synthase Fl subunit gamma, Atp5o - ATP synthase subunit O, mitochondrial, Atp5pb - ATP synthase protein b, mitochondrial, Baat4 - Bile acid-CoA: amino acid N-acyltransf erase family member 4, Bclaf3 - BCL2- associated transcription factor 3, Bdhl - D-beta-hydroxybutyrate dehydrogenase, mitochondrial, Bphl - Biphenyl hydrolase-like protein, Brd8638 - Bromodomain-containing protein 8638 (putative), C3466 - Chromosome 3 open reading frame 466 protein (putative), C390 - Chromosome 3 open reading frame 90 protein (putative), Cacnalsl - Voltagedependent L-type calcium channel subunit alpha-1 S, Calr - Calreticulin, Cbs - Cystathionine beta-synthase, Ccbl2 - Cysteine-S-conjugate beta-lyase 2, Ccdcl30 - Coiled- coil domain-containing protein 130, Ccdcl36 - Coiled-coil domain-containing protein 136, Ccdcl37 - Coiled-coil domain-containing protein 137, Ccdc39386 - Coiled-coil domaincontaining protein 39386 (putative), Cell 992 - C-C motif chemokine ligand 1992 (putative), Cct5 - T-complex protein 1 subunit epsilon, Cdc42 - Cell division control protein 42 homolog, Cdkall - CDK5 regulatory subunit-associated protein 1-like 1, Cesld - Carboxylesterase ID, Chat - Choline O-acetyltransferase, Chchd3 - MICOS complex subunit Micl9, Cisd - CDGSH iron-sulfur domain-containing protein, Clip2 - CAP-Gly domain-containing linker protein 2, Cmbl - Carboxymethylenebutenolidase homolog, Cndp- Carnosine dipeptidase, Cox4il - Cytochrome c oxidase subunit 4 isoform 1, Cox5b - Cytochrome c oxidase subunit 5B, Cpsl l - Carbamoyl-phosphate synthase 1-like protein, Cpsl - Carbamoyl-phosphate synthase 1, Cptla - Carnitine O-palmitoyltransferase 1, liver isoform, Cptlb - Carnitine O-palmitoyltransferase 1, muscle isoform, Crat - Carnitine O- acetyltransferase, Crip - Cysteine-rich intestinal protein, Ctsz - Cathepsin Z, Cyb5r - NADH-cytochrome b5 reductase, Cyp2c44 - Cytochrome P4502C44, Cyp2e - Cytochrome P450 2E1, Cyp2ul - Cytochrome P450 2U1, Cytsb - Cystatin B, Dbi - Acyl-CoA-binding protein, Dbt - Lipoamide acyltransferase component of branched-chain alpha-keto acid dehydrogenase complex, Dcafl21 - DDB1- and CUL4-associated factor 12-like protein, Ddx58 - Probable ATP-dependent RNA helicase RIG-I, Dennd - DENN domain-containing protein (family member, putative), Dhfr - Dihydrofolate reductase, Dis31 - Exosome complex exonuclease RRP44-like protein, Dlat - Dihydrolipoamide acetyltransferase, Dnahl2 - Dynein axonemal heavy chain 12, Dnahl7 - Dynein axonemal heavy chain 17, Dpp - Dipeptidyl peptidase, Dpys - Dihydropyrimidinase, Ecil - Enoyl-CoA delta isomerase 1, Eeflal - Elongation factor 1-alpha 1, Eefla - Elongation factor 1-alpha, Eeflg- Elongation factor 1 -gamma, Ehhadh - Enoyl-CoA hydratase and 3-hydroxyacyl CoA dehydrogenase, Eif3c - Eukaryotic translation initiation factor 3 subunit C, Entpd5 -Ectonucleoside triphosphate diphosphohydrolase 5, Eppkl6 - Epiplakin 16, Erp29 - Endoplasmic reticulum resident protein 29, Etfa - Electron transfer flavoprotein subunit alpha, Fah - Fumarylacetoacetase, Fahd - Fumarylacetoacetate hydrolase domaincontaining protein, Faml 90a -Family with sequence similarity 190 member A, Fasn - Fatty acid synthase, Fdps - Farnesyl pyrophosphate synthase, Fdxr - Ferredoxin reductase, Fgf7 - Fibroblast growth factor 7, Fkbp476 - FK506-binding protein 4-like 76 (putative), Fnl - Fibronectin, Ftcd - Formimidoyltransferase-cyclodeaminase, Ftsj3 - rRNA methyltransferase FTSJ3, Gale - UDP-galactose-4-epimerase, Galkl - Galactokinase 1, Gchfir - GTP cyclohydrolase 1 feedback regulator, Gldc - Glycine decarboxylase, Glul - Glutamine synthetase, Glyctk - Glycerate kinase, Gm20425 - Predicted gene 20425, Trf - Transferrin, Tfl34 - Transferrin-like protein 134 (putative), Got - Aspartate aminotransferase (glutamate oxaloacetate transaminase), Gpcl - Glypican-1, Gpi - Glucose- 6-phosphate isomerase, Gml840 - Predicted gene 1840, Gsta3 - Glutathione S-transferase alpha-3, Gstkl - Glutathione S-transferase kappa 1, Gtf2f- General transcription factor IIF, Gys - Glycogen synthase, Hl-2 - Histone Hl.2, Hlf2 - Histone Hl.2 (alternative nomenclature), H6pd - Hexose-6-phosphate dehydrogenase, Hacd3 - 3-hydroxyacyl-CoA dehydratase 3, Hadh - Hydroxyacyl-CoA dehydrogenase, Hadha - Trifunctional enzyme subunit alpha, Hadhb - Trifunctional enzyme subunit beta, Hagh - Hydroxyacylglutathione hydrolase, Hdhd5 - Haloacid dehalogenase-like hydrolase domain-containing protein 5, Hibch - 3-hydroxyisobutyryl-CoA hydrolase, Histlhlc - Histone Hl.2, Histlhld - Histone Hl.3, Histlhle - Histone Hl.4, Histlhlb - Histone Hl.5, Hivep31 - Transcription factor HIVEP3 isoform 1, Hkl - Hexokinase-1, Hrsp - Histidine-rich glycoprotein serine protease (putative), Hsdl7bl - Estradiol 17-beta-dehydrogenase 1, Hsdl - Hydroxy steroid dehydrogenase-like protein, Hsp90aal - Heat shock protein HSP 90-alpha, Hspa9 - Stress- 70 protein, mitochondrial (mortalin / GRP75), Htra3 - Serine protease HTRA3, Idh2 - Isocitrate dehydrogenase [NADP], mitochondrial, IllOrb - Interleukin- 10 receptor subunit beta, Immt - MICOS complex subunit Mic60 (mitofilin), Isoc2a - Isochori smatase domaincontaining protein 2 A, Kegl - Kinase expressed in germline 1 (putative), Kif7 - Kinesin family member 7, Kmo - Kynurenine 3 -monooxygenase, Krtl4 - Keratin, type I cytoskeletal 14, Krt8 - Keratin, type II cytoskeletal 8, Krt5 - Keratin, type II cytoskeletal 5, Krt79 - Keratin, type II cytoskeletal 79, Krt77 - Keratin, type II cytoskeletal 77, L2hgdh - L-2- hydroxyglutarate dehydrogenase, Lgals - Galectin family proteins, Lrpprc - Leucine-rich PPR motif-containing protein, Lxn - Latexin, Matla - S-adenosylmethionine synthase isoform type-1, Meat - Malonyl-CoA-acyl carrier protein transacylase, Mcccl - Methyl crotonoyl -Co A carboxylase subunit alpha, Me3 - Malic enzyme 3 (NADP(+)-dependent, mitochondrial), Mettl7al - Methyltransferase-like protein 7A1, UbiE2 - Ubiquitin-conjugating enzyme E2 (putative), Mettl7a2 - Methyltransferase-like protein 7A2, Mettl7a3 - Methyltransferase-like protein 7A3, Methigl - Methionine synthase-like protein (putative), Mia31 - Mitochondrial import inner membrane translocase subunit TIM31 (MIA31), Morc2a - MORC family CW-type zinc finger protein 2 A, MovlO - Putative helicase MOVIO, Mpdz - Multiple PDZ domain protein, Mpst - 3- mercaptopyruvate sulfurtransferase, Mrpll l - 39S ribosomal protein Li l, mitochondrial, Mrpll6 - 39S ribosomal protein L16, mitochondrial, Mrpl4 - 39S ribosomal protein L4, mitochondrial, Mrps22 - 28 S ribosomal protein S22, mitochondrial, Mrps2 - 28S ribosomal protein S2, mitochondrial, Myl - Myosin light chain, Myl3 - Myosin light chain 3, Nadkdl- NAD kinase domain-containing protein 1, Naprt - Nicotinate phosphoribosyltransferase, Nat6 - N-acetyltransferase 6, NdufalO - NADH dehydrogenase [ubiquinone] 1 alpha subcomplex subunit 10, Ndufsl - NADH dehydrogenase [ubiquinone] iron-sulfur protein 1, Ndufs - NADH dehydrogenase [ubiquinone] core subunits, Nebl - Nuclear export mediator factor NEMF1 (putative NEB1), Nhlrc2 - NHL repeat-containing protein 2, Nipsnapl - Nipsnap homolog 1, Nopl - Nucleolar protein 1, Notchl2 - Neurogenic locus notch homolog protein 1 / 2 (putative combined), Nqo - NAD(P)H dehydrogenase [quinone], Nsf- Vesicle-fusing ATPase, Nuggc - Nuclear GTP -binding protein, Nupl551 - Nuclear pore complex protein Nupl55, Oat - Ornithine aminotransferase, Ogdhl - 2-oxoglutarate dehydrogenase-like protein, Oxldl - Oxidoreductase-like protein 1, P4hb - Protein disulfide-isomerase, Pah - Phenylalanine-4-hydroxylase, Papss - Bifunctional 3'- phosphoadenosine 5 '-phosphosulfate synthase, Pbldl -Phenazine biosynthesis-like domaincontaining protein 1, Pbxipl - Pre-B-cell leukemia transcription factor-interacting protein1, Pcx - Pyruvate carboxylase, Pc3 - Pyruvate carboxylase isoform 3 (putative), Pdcd - Programmed cell death protein, Pdelc - Calcium / calmodulin-dependent 3',5'-cyclic nucleotide phosphodiesterase 1C, Pdhal - Pyruvate dehydrogenase El component subunit alpha, Pdia3 - Protein disulfide-isomerase A3, Pdia4 - Protein disulfide-isomerase A4, Pdia5 - Protein disulfide-isomerase A5, Pdzd22 - PDZ domain-containing protein 22, Pgk- Phosphoglycerate kinase, Pgml - Phosphoglucomutase- 1, Pgm2 - Phosphoglucomutase-2, Pipox - Pipecolate oxidase, Pitrm 1 - Presequence protease, mitochondrial, Pklrl - Pyruvate kinase isozymes R / L, Plec4 - Plectin isoform 4, Polal - DNA polymerase alpha catalytic subunit, Polg2 - DNA polymerase subunit gamma-2, Por - NADPH-cytochrome P450 reductase, Ppig - Peptidyl-prolyl cis-trans isomerase G, Ppmle - Protein phosphatase IE, Prkcsh - Glucosidase 2 subunit beta, Prkdc3 - DNA-dependent protein kinase catalytic subunit 3 (putative), Prodh2 - Proline dehydrogenase 2, Ptrh2 - Peptidyl-tRNA hydrolase 2,Pygb - Glycogen phosphorylase, brain form, Qdpr - Dihydropteridine reductase, Rblccl l- RBI -inducible coiled-coil protein 1-like 1, Rheb - GTP -binding protein Rheb, Rpll3a - 60S ribosomal protein L13a, Rpl4 - 60S ribosomal protein L4, Rpl6 - 60S ribosomal protein L6, Rpusd - RNA pseudouridine synthase domain-containing protein, Rrbpl 1 - Ribosomebinding protein 1-like 1, Satll - Spermatogenesis-associated protein-like 1, Secl3 - Protein transport protein Secl3, Seel - Syntaxin-binding protein 1, Sfrl - Swi 5 -dependent recombination repair protein 1, Sfxn - Sideroflexin, Shmt2 - Serine hydroxymethyltransferase, mitochondrial, Slc25al - Mitochondrial citrate transporter, Slc25a25 - Calcium-binding mitochondrial carrier protein SCaMC-2, Slc25a5 - ADP / ATP translocase 2, Slc27a2 - Very long-chain acyl-CoA synthetase, Slc27a5 - Bile acyl-CoA synthetase, Slc2a2 - Solute carrier family 2, facilitated glucose transporter member 2 (GLUT2), Slc9a3rl - Na(+) / H(+) exchange regulatory cofactor NHE-RF1, Spg7 - Paraplegin, Sphkapl - Sphingosine kinase type 1 -interacting protein, Sqor- Sulfide:quinone oxidoreductase, mitochondrial, Sultldl - Sulfotransferase family ID member 1, Surf6 - Surfeit locus protein 6, Taldol - Transaldolase, Tesk2 - Dual specificity testis-associated serine / threonine protein kinase 2, Tinag - Tubulointerstitial nephritis antigen, Tkt - Transketolase, Tmem9 - Transmembrane protein 9, Tptl - Translationally-controlled tumor protein, Trabd - TraB domain-containing protein, Trafiip l - TRAF3 -interacting protein 1, Trdmtl - tRNA aspartic acid methyltransferase 1, Trim42 - Tripartite motif-containing protein 42, Ttn27 - Titin isoform 27, Ttn9 - Titin isoform 9, Txndc5 - Thioredoxin domaincontaining protein 5, Txnrd2 - Thioredoxin reductase 2, mitochondrial, Uaca - Uveal autoantigen with coiled-coil domains and ankyrin repeats, Ugdh - UDP-glucose 6- dehydrogenase, Uggtl - UDP-glucose:gly coprotein glucosyltransferase 1, Uox - Uricase, Uqcrc2 - Cytochrome b-cl complex subunit 2, Urocl - Uronate oxidoreductase 1, Wdrl 11- WD repeat-containing protein 111, Wdr - WD repeat-containing protein, Xdh - Xanthine dehydrogenase / oxidase, Zfr - Zinc finger RNA-binding protein, Strbp - Spermatid perinuclear RNA-binding protein, Ilf3 - Interleukin enhancer-binding factor 3, Zfr2 - Zinc finger RNA-binding protein 2, Znf541 - Zinc finger protein 541, and Zwint - ZwlO- interacting protein.

[0089] The proteins in Table 2 are as follows: Aass - Alpha-aminoadipic semialdehyde synthase, Acadl - Acyl-CoA dehydrogenase family member 1, Acadsb - Short / branched chain specific acyl-CoA dehydrogenase, Acadvl - Very long-chain specific acyl-CoA dehydrogenase, Acatl - Acetyl-CoA acetyltransferase 1 (mitochondrial), Acly - ATP-citrate synthase, Acnatl - Acyl-CoA N-acyltransferase 1, Acnat2 - Acyl-CoA N-acyltransferase 2,Acp6 - Lysophosphatidic acid phosphatase type 6, Acsf - Acyl-CoA synthetase family member, Acsll - Long-chain-fatty-acid-CoA ligase 1, Acsl5 - Long-chain-fatty-acid-CoA ligase 5, Acy3 - Aspartoacylase-3, Ada - Adenosine deaminase, Adhfel - Alcohol dehydrogenase, iron-containing 1, Agxt2 - Alanine-glyoxylate aminotransferase 2, Agxt - Alanine-glyoxylate aminotransferase, Akrlal - Aldo-keto reductase family 1 member Al, Akrlc6 - Aldo-keto reductase family 1 member C6, Alb - Serum albumin, Aldh2 - Aldehyde dehydrogenase, mitochondrial, Apoal - Apolipoprotein A-I, Atic - Bifunctional purine biosynthesis protein PURH, Atplal - Sodium / potassium-transporting ATPase subunit alpha- 1, Atp5fl - ATP synthase F(0) complex subunit Bl, Atp51 - ATP synthase subunit g, mitochondrial, Gm542 - Predicted gene 542, BC026585 - cDNA sequence BC026585 (uncharacterized protein), Bpgm - 2,3-bisphosphoglycerate mutase, Cbr - Carbonyl reductase, Ccntl - Cyclin-Tl, Cct5 - T-complex protein 1 subunit epsilon, Cdhr4 - Cadherin-related family member 4, Chchd3 - MICOS complex subunit Micl9, Clybl - Citrate lyase beta-like, Cmpkl - UMP-CMP kinase, Cox6b - Cytochrome c oxidase subunit 6B1, Cpsl - Carbamoyl-phosphate synthase 1, Cpvl - Probable serine carboxypeptidase CPVL, Cyp2el - Cytochrome P450 2E1, Dcxr - Dicarbonyl / L-xylulose reductase, Decrl - 2,4-dienoyl-CoA reductase [NADPH], Dmgdh - Dimethylglycine dehydrogenase, Dnajb8 - DnaJ homolog subfamily B member 8, Ecil - Enoyl-CoA delta isomerase 1, Ehhadh - Enoyl-CoA hydratase and 3 -hydroxy acyl -Co A dehydrogenase, Eif6 - Eukaryotic translation initiation factor 6, Ephx2 - Epoxide hydrolase 2, Fah - Fumarylacetoacetase, Fam21 - WASH complex subunit FAM21, Fgll - Fibrinogen -like protein 1, Fmo3 - Dimethylaniline monooxygenase [N-oxide-forming] 3, Fmo5 - Dimethylaniline monooxygenase [N-oxide- forming] 5, Gm4846 - Predicted gene 4846, Ftcd - Formimidoyltransferasecyclodeaminase, Ganab - Neutral alpha-glucosidase AB, Gcdh - Glutaryl-CoA dehydrogenase, Gclc - Glutamate-cysteine ligase catalytic subunit, Gpcpdl - Glycerophosphocholine phosphodiesterase 1, Gpt2 - Alanine aminotransferase 2, Gsta3 - Glutathione S-transferase alpha-3, Gstkl - Glutathione S-transferase kappa 1, Gstpl - Glutathione S-transferase Pl, Gstp2 - Glutathione S-transferase P2, Gusb3 - Beta- glucuronidase-like protein 3, Haao - 3-hydroxyanthranilate 3, 4-di oxygenase, Hadha - Trifunctional enzyme subunit alpha (mitochondrial), Hadhb - Trifunctional enzyme subunit beta (mitochondrial), Hmgcs2 - Hydroxymethylglutaryl-CoA synthase, mitochondrial, Hnmpa3 - Heterogeneous nuclear ribonucleoprotein A3, Hsdl Ibl - Corticosteroid 11-beta- dehydrogenase isozyme 1, Hsdl2 - Hydroxy steroid dehydrogenase-like protein 2, lahl - Isoamyl acetate-hydrolyzing esterase 1 homolog, Idh3b - Isocitrate dehydrogenase [NAD] subunit beta, Igbplb - Immunoglobulin-binding protein IB, Isoc2a - Isochori smatasedomain-containing protein 2 A, Itfg - Integrin-alpha FG-GAP repeat-containing protein, Keg- Predicted kelch-like protein EG, Klraq - Kelch-like Raq protein, Krt80 - Keratin, type II cytoskeletal 80, Maob - Amine oxidase [flavin-containing] B, Map7d2 - MAP7 domaincontaining protein 2, Matla - S-adenosylmethionine synthetase isoform type-1, Mcurl - Mitochondrial calcium uniporter regulator 1, Mdh2 - Malate dehydrogenase, mitochondrial, Mel - NADP-dependent malic enzyme, Meer - Mitochondrial trans-2-enoyl-CoA reductase, Mettl7al - Methyltransferase-like protein 7A1, UbiE2 - Ubiquitin-conjugating enzyme E2 (putative), Mettl7a2 - Methyltransferase-like protein 7A2, Mettl7a3 - Methyltransferase-like protein 7 A3, Methigl - Methionine synthase-like protein, Mlycd - Malonyl-CoA decarboxylase, Mrpll l - 39S ribosomal protein Li l, mitochondrial, Mrpll6- 39S ribosomal protein L16, mitochondrial, Mrpl38 - 39S ribosomal protein L38, mitochondrial, Mrpl5 - 39S ribosomal protein L5, mitochondrial, Mrps22 - 28S ribosomal protein S22, mitochondrial, Mrps26 - 28S ribosomal protein S26, mitochondrial, Mrps31 - 28S ribosomal protein S31, mitochondrial, Mybbpla - Myb-binding protein 1A, Myom31 - Myomesin 3 isoform 1, Naprt - Nicotinate phosphoribosyltransferase, Nel - Nucleolin, Ndufab - NADH dehydrogenase [ubiquinone] 1 alpha / beta subcomplex protein, NdufblO - NADH dehydrogenase [ubiquinone] subunit BIO, Nebl - Nuclear export mediator factor 1, Nepro - Nephronophthisis-related protein, Nqo - NAD(P)H dehydrogenase [quinone], Olfr827 - Olfactory receptor 827, Oxnad - Oxidoreductase NAD-binding domaincontaining protein, Oxsm - 3-oxoacyl-[acyl-carrier-protein] synthase, mitochondrial, Pcx - Pyruvate carboxylase, Pc9 - Pyruvate carboxylase isoform 9, Pdelc - Calcium / calmodulin- dependent 3 ',5 '-cyclic nucleotide phosphodiesterase 1C, Pde6c - Cone photoreceptor cGMP-specific 3 ',5 '-cyclic phosphodiesterase subunit alpha', Peer - Peroxisomal trans-2- enoyl-CoA reductase, Pgls - 6-phosphogluconolactonase, Phkb - Phosphorylase b kinase regulatory subunit beta, Plecl l - Plectin isoform 11, Plec3 - Plectin isoform 3, Plec - Plectin, Pmm2 - Phosphomannomutase 2, Pnp - Purine nucleoside phosphorylase, Pnpo - Pyridoxamine 5 '-phosphate oxidase, Por - NADPH-cytochrome P450 reductase, Ppmle - Protein phosphatase IE, Psmb8 - Proteasome subunit beta type-8, Ptpmt - Protein tyrosine phosphatase, mitochondrial, Pygl - Glycogen phosphorylase, liver form, Rasgrf21 - RAS protein-specific guanine nucleotide-releasing factor 2-like 1, Rpll4 - 60S ribosomal protein L14, Rpll - 60S ribosomal protein LI, Rrbpl -Ribosome-binding protein 1, Sat2 -Diamine acetyltransferase 2, Sdha - Succinate dehydrogenase [ubiquinone] flavoprotein subunit, Slc27a5 - Bile acyl-CoA synthetase, Sod - Superoxide dismutase, Strap - Serine-threonine kinase receptor-associated protein, Sucla - Succinyl-CoA ligase [ADP-forming] subunit alpha, Taldol - Transaldolase, Tars2 - Threonyl-tRNA synthetase, mitochondrial, Tatdn2 -TatD DNase domain-containing protein 2, Tgml - Protein-glutamine gammaglutamyltransferase 1, Tinag - Tubulointerstitial nephritis antigen, Tmsb4x - Thymosin beta-4, Triml 1 - Tripartite motif family-like protein 1, Uaca- Uveal autoantigen with coiled- coil domains and ankyrin repeats, Uggtl 1 - UDP-glucose:glycoprotein glucosyltransferase- like 1, Uox - Uricase, Urocl - Uronate oxidoreductase 1, Vwa8 - von Willebrand factor A domain-containing protein 8, Xrral - X-ray radiation resistance-associated protein 1, Xylb- Xylulose kinase, Zcwpwl - Zinc finger CW-type and PWWP domain-containing protein 1, and Zfp62 - Zinc finger protein 62.

[0090] The proteins in Table 3 are as follows: A2ML1 - Alpha-2-macroglobulin-like protein 1, AACS - Acetoacetyl-CoA synthetase, AADAT - Aminoadipate aminotransferase, AASS- Alpha-aminoadipic semialdehyde synthase, ACAA2 - 3 -ketoacyl -Co A thiolase (mitochondrial), ACAD8 - Isobutyryl-CoA dehydrogenase, ACADM - Medium-chain specific acyl-CoA dehydrogenase (mitochondrial), ACSF3 - Malonyl-CoA / methylmalonyl- CoA synthetase (mitochondrial), ADGRV1 - Adhesion G protein-coupled receptor VI (GPR98), AFAP1 - Actin filament-associated protein 1, AFF1 - AF4 / FMR2 family member1, AFF - AF4 / FMR2 family protein (unspecified member), AGT - Angiotensinogen, ALDH9A1 - 4-trimethylaminobutyraldehyde dehydrogenase, ANGEL2 - Angel homolog2, ANKRD11 - Ankyrin repeat domain-containing protein 11, ANKRD12 - Ankyrin repeat domain-containing protein 12, ANKRD311 - Ankyrin repeat domain-containing protein 31- like, ANXA4 - Annexin A4, ANXA - Annexin family protein, APO Al - Apolipoprotein A-I, AREG - Amphiregulin, ARHGAP 18 - Rho GTPase-activating protein 18, ARHGAP26- Rho GTPase-activating protein 26, ARID IB 1 - AT-rich interactive domain-containing protein IB, ARID5B - AT-rich interactive domain-containing protein 5B, ARMCX2 - Armadillo repeat-containing X-linked protein 2, ARNTL2 - Aryl hydrocarbon receptor nuclear translocator-like protein 2, ARRB2 - Beta-arrestin-2, ARVCF - Armadillo repeat gene deleted in velocardiofacial syndrome protein, ASAP3 - ArfGAP with SH3 domain, ankyrin repeat, and PH domain-containing protein 3, ASPM1 - Abnormal spindle-like microcephaly -associated protein 1, ASPM - Abnormal spindle-like microcephaly- associated protein, ASPN - Asporin, ASXL1 - Additional sex combs-like protein 1, ATL3- Atlastin-3, ATP1B3 - Sodium / potassium-transporting ATPase subunit beta-3, ATP5F1C- ATP synthase Fl subunit gamma (mitochondrial), ATP5IF - ATP synthase inhibitory factor subunit 1 (mitochondrial), ATRX - ATP-dependent helicase ATRX, BAZ1A1 - Bromodomain adjacent to zinc finger domain protein 1 A, BBOF - Bardet-Biedl syndrome 9 associated factor (BB Some-related protein), BBS9 - Bardet-Biedl syndrome 9 protein,BCAT2 - Branched-chain-amino-acid aminotransferase, mitochondrial, BCL - B-cell lymphoma-associated protein (unspecified family member), BCO2 - Beta-carotene oxygenase 2, BDH - D-beta-hydroxybutyrate dehydrogenase (mitochondrial), BHMT - Betaine-homocysteine S-methyltransferase 1, BIRC5 - Baculoviral IAP repeat-containing protein 5 (Survivin), BLM1 - Bloom syndrome protein homolog 1, BLOC1S6 - Biogenesis of lysosome-related organelles complex 1 subunit 6, BMX - Non-receptor tyrosine-protein kinase BMX, BOD IL 1 - Biorientation of chromosomes in cell division protein 1-like 1, BRAP - BRCA1 -associated protein, BUB IB - Mitotic checkpoint serine / threonine-protein kinase BUB1 beta, BVES - Blood vessel epicardial substance, C2CD2 - C2 domaincontaining protein 2, C6orfl l8 - Uncharacterized protein C6orfl l8, C9orfl31 - Uncharacterized protein C9orfl31, CACYBP - Calcyclin-binding protein, CAND - Cullin- associated NEDD8-dissociated protein, CAPN11 - Calpain-11, CASQ - Calsequestrin, CBLL - Cbl proto-oncogene-like protein, CCDC127 - Coiled-coil domain-containing protein 127, CCDC148 - Coiled-coil domain-containing protein 148, CCDC178 - Coiled- coil domain-containing protein 178, CCNA2 - Cyclin-A2, CCT2 - T-complex protein 1 subunit beta, CDC6 - Cell division control protein 6, CDCA2 - Cell division cycle- associated protein 2, CENPJ - Centromere protein J, CEP 13 - Centrosomal protein of 13 kDa, CEP295 - Centrosomal protein of 295 kDa, CEP3501 - Centrosomal protein 350-like protein, CES1 - Carboxylesterase 1, CFAP2511 - Cilia- and flagella-associated protein 251, CFAP542 - Cilia- and flagella-associated protein 54, CHCHD - Coiled-coil-helix-coiled- coil-helix domain-containing protein, CHRNA1 - Nicotinic acetylcholine receptor subunit alpha-1, CIITA - MHC class II transactivator, CIPC - CLOCK-interacting pacemaker protein, CIT - Citron Rho-interacting kinase, CLIC5 - Chloride intracellular channel protein 5, CLSPN - Claspin, CMBL - Carboxymethylenebutenolidase homolog, CNMD - Chondromodulin-1, CNTN5 - Contactin-5, COIL - Coil protein, COL28A1 - Collagen alpha- 1 (XXVIII) chain, COL5A3 - Collagen alpha-3 (V) chain, COL6A31 - Collagen alpha- 1(VI) chain isoform 31, COMT - Catechol O-m ethyltransferase, COTL1 - Coactosin-like protein, CPS11 - Carbamoyl-phosphate synthase 1-like protein, CRACD - Capping protein regulator and myosin 1 linker, CRAT - Carnitine O-acetyltransferase, CRIP2 - Cysteine- rich protein 2, CRISP - Cysteine-rich secretory protein, CRYB Al - Beta-crystallin Al, CSTB - Cystatin-B, CTSE - Cathepsin E, CTSH - Cathepsin H, CYP2U1 - Cytochrome P450 2U1, DARS2 - Aspartyl-tRNA synthetase, mitochondrial, DBN1 - Drebrin, DCN - Decorin, DCPS - m7GpppX diphosphatase, DCTN4 - Dynactin subunit 4, DDB2 - DNA damage-binding protein 2, DDHD1 - DDHD domain-containing protein 1, DDIAS - DNA damage-induced apoptosis suppressor, DDX19A - ATP-dependent RNA helicase DDX19A,DDX19B - ATP-dependent RNA helicase DDX19B, DDX23 - Probable ATP-dependent RNA helicase DDX23, DDX39A - ATP-dependent RNA helicase DDX39A, DDX60L - Probable ATP-dependent RNA helicase DDX60-like, DGUOK - Deoxyguanosine kinase (mitochondrial), DIRAS3 - GTP -binding protein Di-Ras3, DMGDH - Dimethylglycine dehydrogenase, DMXL2 - DmX-like protein 2, DNAAF4 - Dynein axonemal assembly factor 4, DNAH5 - Dynein heavy chain 5 (axonemal), DNAH7 - Dynein heavy chain 7 (axonemal), DNAH9 - Dynein heavy chain 9 (axonemal), DNAJC2 - DnaJ homolog subfamily C member 2, DYNC2H1 - Cytoplasmic dynein 2 heavy chain 1, DYRK3 - Dual specificity tyrosine-phosphorylation-regulated kinase 3, E2F8 - Transcription factor E2F8, EBNA1BP - EBNA1 -binding protein 2, ECHS1 - Enoyl-CoA hydratase, short chain 1 (mitochondrial), ECU - Enoyl-CoA delta isomerase 1 (mitochondrial), EEF2 - Elongation factor 2, EFCAB6 - EF-hand calcium-binding domain-containing protein 6, EGFL6 - EGF- like domain multiple 6, EGFR - Epidermal growth factor receptor, EIF1AX - Eukaryotic translation initiation factor 1 A, X-linked, EIF1 AY - Eukaryotic translation initiation factor 1 A, Y-linked, ELAC1 - Zinc phosphodiesterase EL AC protein 1, EL0VL2 - Elongation of very long-chain fatty acids protein 2, EPB41 - Protein 4.1 (erythrocyte membrane protein band 4.1), EPHX1 - Epoxide hydrolase 1, ESD - S-formylglutathione hydrolase (esterase D), ETHE1 - Sulfur dioxygenase ETHE1, EVC2 - Ellis-van Creveld syndrome protein 2, EX0C7 - Exocyst complex component 7, EXOG - Exonuclease G, mitochondrial, EXPH5 - Exophilin-5, F13A1 - Coagulation factor XIII A chain, FAAH2 - Fatty acid amide hydrolase 2, FAAP24 - Fanconi anemia-associated protein 24, FABP - Fatty acid-binding protein (family member), FAM210B - Protein FAM210B, FAM21A - WASH complex subunit FAM21A, FAM65C - Protein FAM65C, FAM83E - Protein FAM83E, FASN - Fatty acid synthase, FAT1 - Protocadherin Fat 1, FBN1 - Fibrillin-1, FBX04 - F-box protein 4, FEN1 - Flap endonuclease 1, FGG - Fibrinogen gamma chain, FHAD1 - Forkhead-associated domain-containing protein 1, FN1 - Fibronectin, FNIP - Folliculin- interacting protein, FRA10AC - Fragile site-associated C protein, FRMPD3 - FERM and PDZ domain-containing protein 3, FSIP2 - Fibrous sheath-interacting protein 2, FURIN - Furin, G6PD - Glucose-6-phosphate 1 -dehydrogenase, GARS - Glycyl-tRNA synthetase, GART - Trifunctional purine biosynthetic protein adenosine-3, GCM1 - Chorion-specific transcription factor GCM1, GCN - General control non-derepressible homolog, GLDC - Glycine dehydrogenase (decarboxylating), GLT8D1 - Glycosyltransferase 8 domaincontaining protein 1, GPATCH11 - G-patch domain-containing protein 11, GPATCH1 - G- patch domain-containing protein 1, GPATCH8 - G-patch domain-containing protein 8, GPR19 - G-protein coupled receptor 19, GPRIN3 - G-protein-regulated inducer of neuriteoutgrowth 3, GRB10 - Growth factor receptor-bound protein 10, GRIA1 - Glutamate receptor 1 (AMPA subtype), GRK - G-protein coupled receptor kinase (unspecified isoform), GRK7 - G-protein coupled receptor kinase 7, GRPEL1 - GrpE protein homolog 1 (mitochondrial), GSE1 - Genetic suppressor element 1, GSTK1 - Glutathione S- transferase kappa 1, GSTO1 - Glutathione S-transferase omega-1, GSTT1 - Glutathione S- transferase theta-1, GSTT2B - Glutathione S-transferase theta-2B, H2AJ - Histone H2A.J, HACD - Very-long-chain (3R)-3-hydroxyacyl-CoA dehydratase (family member), HADH- Hydroxyacyl-CoA dehydrogenase (mitochondrial), HADHA - Trifunctional enzyme subunit alpha (mitochondrial), HAL - Histidine ammonia-lyase, HBE - Hemoglobin subunit epsilon, HEBP1 - Heme-binding protein 1, HEXB - Beta-hexosaminidase subunit beta, HNRNPM - Heterogeneous nuclear ribonucleoprotein M, HSD17B4 - Peroxisomal multifunctional enzyme type 2 (17P-hydroxysteroid dehydrogenase type 4), IBTK - Inhibitor of Bruton’s tyrosine kinase, IFIT2 - Interferon-induced protein with tetratricopeptide repeats 2, IFT140 - Intraflagellar transport protein 140 homolog, IL1RAP- Interleukin- 1 receptor accessory protein, INCA - Inhibitor of nuclear factor kappa-B kinase complex-associated protein, ING5 - Inhibitor of growth protein 5, INTS7 - Integrator complex subunit 7, IQGAP3 - Ras GTPase-activating-like protein IQGAP3, IRAK2 - Interleukin-1 receptor-associated kinase 2, ISG20L - Interferon-stimulated gene 20 kDa protein-like, ITPR2 - Inositol 1,4,5-trisphosphate receptor type 2, ITSN1 - Intersectin- 1, KAT14 - Lysine acetyltransferase 14, KAT2B - Histone acetyltransferase KAT2B (PCAF), KDM3B - Lysine-specific demethylase 3B, KDM4B - Lysine-specific demethylase 4B, KIAA0319 - Dyslexia susceptibility protein KIAA0319, KIF14 - Kinesin-like protein KIF14, KIF27 - Kinesin-like protein KIF27, KIN - Kinase suppressor of Ras 1-like, KL - KI otho, KNTC1 - Kinetochore-associated protein 1, KRT18 - Keratin, type I cytoskeletal 18, KRT23 - Keratin, type I cytoskeletal 23, KSR2 - Kinase suppressor of Ras 2, L2HGDH- L-2-hydroxyglutarate dehydrogenase, LACTB - Serine beta-lactamase-like protein, LANCL1 - LanC-like protein 1, LATS2 - Large tumor suppressor kinase 2, LBP - Lipopolysaccharide-binding protein, LDHB - L-lactate dehydrogenase B chain, LEMD2 - LEM domain-containing protein 2, LIFR - Leukemia inhibitory factor receptor, LMNB2 - Lamin-B2, LONP1 - Lon protease homolog, mitochondrial, LRPPRC - Leucine-rich PPR motif-containing protein, LRRC41 - Leucine-rich repeat-containing protein 41, LTO1 - Photosystem II biogenesis protein LTO1 homolog, LUZP1 - Leucine zipper protein 1, MACR0H2A1 - Core histone macro-H2A.1, MAGI3 - Membrane-associated guanylate kinase inverted 3, MAPI A - Microtubule-associated protein 1A, MAP4 - Microtubule- associated protein 4, MAP7D3 - MAP7 domain-containing protein 3, MAST2 -Microtubule-associated serine / threonine-protein kinase 2, MBTPS1 - Membrane-bound transcription factor site-1 protease, MCCC2 - Methylcrotonoyl-CoA carboxylase beta chain, MCU - Calcium uniporter protein, MDFI - MyoD family inhibitor, MDH2 - Malate dehydrogenase (mitochondrial), MDM2 - E3 ubi quitin-protein ligase Mdm2, ME2 - Malic enzyme 2 (mitochondrial), MED 13 - Mediator complex subunit 13, MEFV - Pyrin (Mediterranean fever protein), MGAT3 - Beta- 1,4-mannosyl-gly coprotein beta-l,4-N- acetylglucosaminyltransferase 3, MLF1 - Myeloid leukemia factor 1, MLH1 - DNA mismatch repair protein MLH1, MMP1 - Interstitial collagenase, MMP - Matrix metalloproteinase (unspecified family member), M0RC3 - MORC family CW-type zinc finger protein 3, MPH0SPH8 - M-phase phosphoprotein 8, MRM - Mitochondrial rRNA methyltransferase, MRPS22 - 28S ribosomal protein S22 (mitochondrial), MSL1 - Malespecific lethal 1 homolog, MTHFD1 - C-l -tetrahydrofolate synthase, MTMR8 - Myotubularin-related protein 8, MYBBP1 A - Myb-binding protein 1 A, MYH10 - Myosin- 10, MYH14 - Myosin-14, MYH9 - Myosin-9, MYO1F - Unconventional myosin-Ib family member F, MY05C - Unconventional myosin-Vc, N4BP2 - NEDD4-binding protein 2, NADK2 - NAD kinase 2 (mitochondrial), NAP1L2 - Nucleosome assembly protein 1 -like 2, NARS2 - Aspartyl-tRNA synthetase 2 (mitochondrial), NCKAP5 - NCK-associated protein 5, NDUFAF7 - NADH dehydrogenase [ubiquinone] complex I assembly factor 7, NDUFS3 - NADH dehydrogenase [ubiquinone] iron-sulfur protein 3, NDUFV2 - NADH dehydrogenase [ubiquinone] flavoprotein 2, NFE2L2 - Nuclear factor erythroid 2-related factor 2 (NRF2), NLRX1 - NLR family member XI, NNT - NAD(P) transhydrogenase, N0D2 - Nucleotide-binding oligomerization domain-containing protein 2, NOL11 - Nucleolar protein 11, NOP53 - Nucleolar protein 53, NPEPPS - Puromycin-sensitive aminopeptidase, NSD1 -Histone-lysineN-methyltransferaseNSDl, NSD2 -Histone-lysine N-m ethyltransferase NSD2, NSFL1C - NSFL1 cofactor p47, NSRP1 - Nuclear speckle splicing regulatory protein 1, NUDC - Nuclear distribution protein C, NUDT - Nudix hydrolase family member, NUP50 - Nuclear pore complex protein Nup50, NUSAP1 - Nucleolar and spindle-associated protein 1, OAS - 2'-5'-oligoadenylate synthase, ODAD3 - Outer dynein arm docking complex subunit 3, ODF3L - Outer dense fiber of sperm tails 3- like protein, OGFR - Opioid growth factor receptor, OR8G5 - Olfactory receptor 8G5, OSBPL9 - Oxysterol -binding protein-related protein 9, OXSR- Oxidative stress-responsive protein, PAICS - Multifunctional protein ADE2, PARP9 - Poly [ADP-ribose] polymerase 9, PBK - Serine / threonine-protein kinase PBK, PCCA - Propionyl-CoA carboxylase alpha chain, PCGF3 - Polycomb group RING finger protein 3, PCK1 - Phosphoenol pyruvate carboxykinase, cytosolic, PCNP - PEST-containing nuclear protein, PCYT1 A - Phosphatecytidylyltransferase 1A, choline, PDCD11 - Programmed cell death protein 11, PDE6C - Cone cGMP-specific phosphodiesterase subunit alpha', PECR - Peroxisomal trans-2-enoyl- CoA reductase, PEX - Peroxisomal biogenesis factor (family), PFKM - Phosphofructokinase, muscle type, PGAM - Phosphoglycerate mutase (family), PHF20L1- PHD finger protein 20-like 1, PHLDB1 - Pleckstrin homology-like domain family B member 1, PIK3C2A - Phosphatidylinositol-4-phosphate 3 -kinase C2 domain-containing alpha, PINK1 - PTEN-induced putative kinase 1, PLIN - Perilipin, PLIN5 - Perilipin 5, PMS1 - Postmeiotic segregation increased 1, PMS2 - Postmeiotic segregation increased 2, PNPT1 - Polyribonucleotide nucleotidyltransferase 1, POF1B - Premature ovarian failure protein IB, P0LD3 - DNA polymerase delta subunit 3, POLQ - DNA polymerase theta, POP - Prolyl oligopeptidase, PPA2 - Inorganic pyrophosphatase 2, PPAT - Amidophosphoribosyltransferase, PPP5C - Serine / threonine-protein phosphatase 5, PRDX- Peroxiredoxin (family), PRKCSH - Glucosidase 2 subunit beta, PRKDC - DNA- dependent protein kinase catalytic subunit, PRORP - Ribonuclease P protein subunit, PRPF4B - Pre-mRNA-processing factor 4B, PTPRD - Protein tyrosine phosphatase receptor type D, PUS7L - Pseudouridylate synthase 7-like, PYCR1 - Pyrroline-5- carboxylate reductase 1, PYGB - Glycogen phosphorylase, brain form, RAB22A - Ras- related protein Rab-22A, RABGAP1 - Rab GTPase-activating protein 1, RABL2 - RAB- like protein 2, RAD52 - DNA repair protein RAD52, RAD54L - DNA repair and recombination protein RAD54-like, RALB - Ras-related protein Ral-B, RAPGEF6 - Rap guanine nucleotide exchange factor 6, RARS - Arginyl-tRNA synthetase, RCC2 - Regulator of chromosome condensation 2, REPIN1 - Replication initiator 1, RFESD - Iron-sulfur cluster-binding protein RFESD, RIPK1 - Receptor-interacting serine / threonine-protein kinase 1, R0CK1 - Rho-associated protein kinase 1, ROS1 - Proto-oncogene tyrosineprotein kinase ROS, RPA2 - Replication protein A 32 kDa subunit, RPL7L1 - 60S ribosomal protein L7-like 1, RPN2 - Dolichyl-diphosphooligosaccharide-protein glycosyltransferase subunit 2, RPS2 - 40S ribosomal protein S2, RPS4Y - 40S ribosomal protein S4, Y-linked, RPS6KA6 - Ribosomal protein S6 kinase alpha-6, RRBP1 - Ribosome-binding protein 1, RRP15 - Ribosomal RNA-processing protein 15, RSPH3 - Radial spoke head protein 3 homolog, RTFDC1 - Regulatory factor of telomere elongation 1, S100A1 - Protein S100-A1, SAFB - Scaffold attachment factor B, SAR1A - Secretion- associated Ras-related protein SAR1A, SCLT1 - Sodium channel and clathrin linker 1, SCN4B - Sodium channel subunit beta-4, SELENOS - Selenoprotein S, SERPINI1 - Neuroserpin, SETX - Senataxin, SFMBT2 - Scm-like with four MBT domains protein 2, SHMT1 - Serine hydroxymethyltransferase 1, SHMT2 - Serine hydroxymethyltransferase2, SLC12A6 - Solute carrier family 12 member 6, SLC25A24 - Mitochondrial calcium- binding carrier protein SCaMC-1, SLC25A6 - ADP / ATP translocase 3, SLC35G1 - Solute carrier family 35 member Gl, SLIRP - SRA stem -loop interacting RNA-binding protein, SPEN - Spen family transcriptional repressor, SPIN2B - Spindlin-2B, SPTA1 - Spectrin alpha chain, erythrocytic 1, SQOR - Sulfide quinone oxidoreductase, SRD5A3 - Steroid 5- alpha-reductase 3, SRP - Signal recognition particle (complex), SSB - Single-stranded DNA-binding protein, SSH2 - Slingshot protein phosphatase 2, STAU1 - Double-stranded RNA-binding protein Staufen homolog 1, STX - Syntaxin (family), STXBP3 - Syntaxinbinding protein 3, SUCLG2 - Succinate-CoA ligase [GDP-forming] subunit beta, SUGP2 - SURP and G-patch domain-containing protein 2, SUPV3L1 - ATP-dependent RNA helicase SUPV3L1, SUZ12 - Polycomb repressive complex 2 subunit SUZ12, SYNE1 - Nesprin-1, SYNM - Synemin, TAF3 - TATA-box-binding protein-associated factor 3, TBC1D14 - TBC1 domain family member 14, TCOF1 - Treacle protein, TDRD6 - Tudor domaincontaining protein 6, TEP1 - Telomerase-associated protein 1, TES - Testin, TFPI - Tissue factor pathway inhibitor, THG1L - tRNA-histidine guanylyltransferase 1-like, THYN1 - Thymocyte nuclear protein 1, TNC - Tenascin-C, TOPI - DNA topoisomerase 1, TOP2B - DNA topoisomerase 2-beta, TP53 - Cellular tumor antigen p53, TP53BP1 - Tumor protein p53-binding protein 1, TPPP - Tubulin polymerization-promoting protein, TPX2 - Targeting protein for Xklp2, TRAP1 - Heat shock protein TRAP1, TRIM29 - Tripartite motif-containing protein 29, TRIP 11 - Thyroid hormone receptor interactor 11, TST - Thiosulfate sulfurtransferase, TTN - Titin, TULP1 - Tubby-like protein 1, UBAP2 - Ubiquitin-associated protein 2, UBE3B - E3 ubiquitin-protein ligase UBE3B, UFL1 - UFM1 -specific ligase 1, UPF3B - Regulator of nonsense transcripts 3B, UQCRB - Ubiquinol-cytochrome c reductase binding protein, USP10 - Ubiquitin carboxyl-terminal hydrolase 10, USP5 - Ubiquitin carboxyl-terminal hydrolase 5, UTP11 - U3 small nucleolar RNA-associated protein 11 homolog, VCL - Vinculin, VDAC3 - Voltage-dependent anionselective channel protein 3, VIT - Vitrin, VPS41 - Vacuolar protein sorting-associated protein 41, VRK1 - Serine / threonine-protein kinase VRK1, WASHC2C - WASH complex subunit 2C, WASHC3 - WASH complex subunit 3, WDR1 - WD repeat-containing protein 1, WDR75 - WD repeat-containing protein 75, XPOT - Exportin-T, XRCC6 - X-ray repair cross-complementing protein 6, YLPM1 - YLP motif-containing protein 1, ZC3H14 - Zinc finger CCCH domain-containing protein 14, ZKSCAN1 - Zinc finger protein with KRAB and SCAN domains 1, ZMAT1 - Zinc finger matrin-type protein 1, ZNF182 - Zinc finger protein 182, ZNF197 - Zinc finger protein 197, ZNF296 - Zinc finger protein 296, ZNF512B - Zinc finger protein 512B, ZNF534 - Zinc finger protein 534, ZNF598 - Zincfinger protein 598, ZNF644 - Zinc finger protein 644, ZNF7 - Zinc finger protein 7, and ZP2 - Zona pellucida sperm-binding protein 2.

[0091] The proteins in Table 4 are as follows: AACS - Acetoacetyl -Co A synthetase, ABHD10 - Ab hydrolase domain-containing protein 10, ACAD8 - Isobutyryl-CoA dehydrogenase, AC ATI - Acetyl-CoA acetyltransferase, ACINI - Apoptotic chromatin condensation inducer in the nucleus, ADGRG3 - Adhesion G-protein coupled receptor G3, AFG3L2 - AFG3-like protein 2, AIMP2 - Aminoacyl tRNA synthase complex -interacting multifunctional protein 2, ALB - Serum albumin, ALDH18A1 - Delta- l-pyrroline-5- carboxylate synthase, ALDH5A1 - Succinate-semialdehyde dehydrogenase, ALDH6A1 - Methylmalonate semialdehyde dehydrogenase, ANXA4 - Annexin A4, APOC - Apolipoprotein C (family), ASPM - Abnormal spindle-like microcephaly-associated protein, ATP5F1C - ATP synthase subunit gamma, mitochondrial, ATP5IF1 - ATPase inhibitory factor 1, ATP5PD - ATP synthase subunit d, mitochondrial, AXDND1 - AXDND domain-containing protein 1, BPHL - Biphenyl hydrolase-like protein, BRCA2 - Breast cancer type 2 susceptibility protein, CCDC38 - Coiled-coil domain-containing protein 38, CD101 - CD101 antigen, CFAP77 - Cilia and flagella-associated protein 77, CHAMP 1 - Chromosome alignment-maintaining phosphoprotein 1, CIAPIN1 - Cytokine-induced apoptosis inhibitor 1, CKAP4 - Cytoskeleton-associated protein 4, COL6A3 - Collagen alpha-3(VI) chain, COX5B - Cytochrome c oxidase subunit 5B, CPS1 - Carbamoylphosphate synthase [ammonia], CXCL1 - C-X-C motif chemokine 1, DARS2 - Aspartyl- tRNA synthetase 2, mitochondrial, DDAH1 - Dimethylarginine dimethylaminohydrolase 1, DIAPH1 - Diaphanous-related formin-1, DKC1 - Dyskerin, DLD - Dihydrolipoamide dehydrogenase, DNAAF4 - Dynein axonemal assembly factor 4, DUSP - Dual specificity protein phosphatase (family), ECU - Enoyl-CoA delta isomerase 1, EGFR - Epidermal growth factor receptor, EMG1 - EMG1 nucleolar protein, ENO1 - Alpha-enolase, EP300 - ElA-binding protein p300, EPHX1 - Epoxide hydrolase 1, ERC1 - ELKS / RAB6- interacting / CAST family member 1, ESD - S-formylglutathione hydrolase, FABP - Fatty acid-binding protein (family), FAM65C - Protein FAM65C, FECH - Ferrochelatase, FOXF1 - Forkhead box protein Fl, GNPDA - Glucosamine-6-phosphate deaminase 1, GNPNAT1 - Glucosamine-phosphate N-acetyltransferase 1, GPD1 - Glycerol-3 -phosphate dehydrogenase 1, Hl-2 - Histone Hl.2, HBA - Hemoglobin subunit alpha, HEXB - Betahexosaminidase subunit beta, HNRNPCL1 - Heterogeneous nuclear ribonucleoprotein C- like 1, IARS2 - Isoleucyl -tRNA synthetase 2, mitochondrial, IFIT3 - Interferon-induced protein with tetratricopeptide repeats 3, IL 18 - Interleukin- 18, IL18RAP - Interleukin- 18receptor accessory protein, KEAP1 - Kelch-like ECH-associated protein 1, KIAA1328 - Protein KIAA1328, KIF16B - Kinesin-like protein KIF16B, KRT23 - Keratin, type I cytoskeletal 23, LACTB2 - Serine beta-lactamase-like protein 2, LAMC2 - Laminin subunit gamma-2, LARP7 - La-related protein 7, LRPPRC - Leucine-rich PPR motif-containing protein, MAML3 - Mastermind-like protein 3, MAP10 - Microtubule-associated protein 10, MCM5 - DNA replication licensing factor MCM5, METTL25 - Methyltransferase-like protein 25, MRM - Mitochondrial rRNA methyltransferase, MRPL46 - 39S ribosomal protein L46, mitochondrial, MRPS22 - 28S ribosomal protein S22, mitochondrial, MTHFD1 - C-l -tetrahydrofolate synthase, MYL7 - Myosin light chain 7, MYLK - Myosin light chain kinase, NIBAN1 - Niban apoptosis regulator 1, NUFIP1 - Nuclear fragile X mental retardation-interacting protein 1, NUMA1 - Nuclear mitotic apparatus protein 1, NUP214 - Nuclear pore complex protein Nup214, OAT - Ornithine aminotransferase, OBSCN - Obscurin, ORC - Origin recognition complex subunit (family), PARP10 - Poly [ADP-ribose] polymerase 10, PASD1 - PAS domain-containing protein 1, PAXIP1 - PAX interacting protein 1, PBXIP1 - Pre-B-cell leukemia homeobox-interacting protein 1, PCCA- Propionyl-CoA carboxylase alpha chain, PCK1 - Phosphoenol pyruvate carboxykinase, cytosolic, PEBP1 - Phosphatidylethanolamine-binding protein 1, PEPD - Peptidase D, PEX- Peroxisomal biogenesis factor (family), PHC3 - Polyhomeotic-like protein 3, POLRMT - DNA-directed RNA polymerase mitochondrial, POP - Prolyl oligopeptidase, PPHLN1 - Periphilin-1, PPIA - Peptidyl -prolyl cis-trans isomerase A (Cyclophilin A), PPIF - Peptidyl- prolyl cis-trans isomerase F (Cyclophilin D), PPL - Periplakin, PPP1R12C - Protein phosphatase 1 regulatory subunit 12C, PRKCSH - Glucosidase 2 subunit beta, PSMB10 - Proteasome subunit beta type-10, RAD52 - DNA repair protein RAD52, RP1L1 - Retinitis pigmentosa 1 -like protein 1, RPF - Ribosome production factor, RRP - Ribosomal RNA- processing protein (family), RUFY2 - RUN and FYVE domain-containing protein 2, SAFB2 - Scaffold attachment factor B2, SDHA - Succinate dehydrogenase [ubiquinone] flavoprotein subunit, SERPING1 - Plasma protease Cl inhibitor, SETMAR-Histone-lysine N-methyltransferase SETMAR, SHMT2 - Serine hydroxymethyltransferase 2, SHPRH - E3 ubiquitin-protein ligase SHPRH, SKIL - SKI-like protein, SLC25A3 - Mitochondrial phosphate carrier protein, SLC44A2 - Choline transporter-like protein 2, SORCS3 - VPS 10 domain-containing receptor SorCS3, SPANXN - Sperm protein associated with the nucleus on the X chromosome (family), SPEF2 - Sperm flagellar protein 2, SQOR - Sulfide quinone oxidoreductase, STAU1 - Double-stranded RNA-binding protein Staufen homolog 1, STXBP2 - Syntaxin-binding protein 2, SUGP2 - SURP and G-patch domain-containing protein 2, TBL2 - Transducin beta-like protein 2, TDRD6 - Tudor domain-containingprotein 6, TLX1 - T-cell leukemia homeobox protein 1, TPK1 - Thiamine pyrophosphokinase 1, TRIM59 - Tripartite motif-containing protein 59, TTC3 - Tetratri copeptide repeat protein 3, TTC4 - Tetratricopeptide repeat protein 4, TTF1 - Transcription termination factor 1, TXNDC12 - Thioredoxin domain-containing protein 12, UHRF1 - Ubiquitin-like with PHD and ring finger domains 1, VC AN - Versican, XIRP2 - Xin actin-binding repeat-containing protein 2, XRCC5 - X-ray repair cross-complementing protein 5 (Ku80), and ZNF34 - Zinc finger protein 34.

[0092] It will be understood that the consensus sequence for humans for the sites in Tables 1-2 are not lysine (arginine for Tables 1 and glutamine for Table 2) and are lysine for the sites in Tables 3-4, but a given specific human subject may instead have a different base. In some embodiments, the method comprises detecting a mutation in the subject that produces a lysine in a position selected from those provided in Tables 1-2. In some embodiments, the method comprises detecting a mutation in the subject that produces an arginine in a position selected from those provided in Table 3. In some embodiments, the method comprises detecting a mutation in the subject that produces a glutamine in a position selected from those provided in Table4. In some embodiments, the lysine / arginine / glutamine is detected in a cell of the subject. In some embodiments, the conversion is performed in that cell. In some embodiments, the lysine / arginine / glutamine is detected in a tissue or organ of the subject and the conversion is performed in that tissue or organ. In some embodiments, the lysine / arginine / glutamine is detected in the subject and the conversion is performed throughout the subject.

[0093] By another aspect, there is provided a method of extending the longevity of a nonhuman subject, the method comprising: a. detecting in the non -human subject a lysine at a position selected from those provided in Table 1; and b. converting the lysine at a position provided in Table 1 to an arginine in the non-human subject, thereby extending the longevity of a non-human subject.

[0094] By another aspect, there is provided a method of extending the longevity of a non- human subject, the method comprising: a. detecting in the non-human subject a lysine at a position selected from those provided in Table 2; andb. converting the lysine at a position provided in Table 2 to a glutamine in the non-human subject, thereby extending the longevity of a non-human subject.

[0095] By another aspect, there is provided a method of extending the longevity of a non- human subject, the method comprising: a. detecting in the non-human subj ect an arginine at a position selected from those provided in Table 3; and b. converting the arginine at a position provided in Table 3 to a lysine in the non-human subject, thereby extending the longevity of a non-human subject.

[0096] By another aspect, there is provided a method of extending the longevity of a non- human subject, the method comprising: a. detecting in the non-human subj ect a glutamine at a position selected from those provided in Table 4; and b. converting the glutamine at a position provided in Table 4 to a lysine in the non-human subject, thereby extending the longevity of a non-human subject.

[0097] It will be understood that the consensus sequence for non-humans for the sites in Tables 1-2 are lysine and are not lysine for the sites in Tables 3-4 (arginine for Tables 3 and glutamine for Table 4). These non-human subjects can thus have their longevity extended by replacing the base with the base found in humans (and other long-lived animals). In some embodiments, the lysine / arginine / glutamine is detected in a cell of the subject. In some embodiments, the conversion is performed in that cell. In some embodiments, the lysine / arginine / glutamine is detected in a tissue or organ of the subject and the conversion is performed in that tissue or organ. In some embodiments, the lysine / arginine / glutamine is detected in the subject and the conversion is performed throughout the subject. In some embodiments, the non-human subject is a veterinary animal. In some embodiments, the non- human subject is a pet. In some embodiments, the non-human animal is a domesticated animal. In some embodiments, the non-human animal is a farm animal. In some embodiments, the veterinary animal is selected from, a cat, a dog, a horse, a cow, a pig, a sheep and a goat. In some embodiments, the non-human animal is a short lived animal. In some embodiments, a short lived animal is an animal with a short life span. In someembodiments, life span is average life span. In some embodiments, life span is maximum life span. In some embodiments, a short life span is below 5, 10, 15, 20, 25, 30, 35, 40, 45 or 50 years. Each possibility represents a separate embodiment of the invention. In some embodiments, a short life span is below 10 years. In some embodiments, a short life span is below 20 years. In some embodiments, a short life span is below 30 years.

[0098] In some embodiments, the method further comprises receiving DNA from the subject. In some embodiments, the detecting is detecting in the DNA. In some embodiments, the method further comprises receiving cells from the subject and the detecting is detecting in the cells. In some embodiments, DNA is extracted / isolated form the cells and the detecting is in the extracted / isolated DNA. In some embodiments, detecting is determining the DNA comprises the lysine / arginine / glutamine. In some embodiments, detecting is determining the DNA encodes a protein that comprises the lysine / arginine / glutamine. In some embodiments, a position is a location. In some embodiments, a position is within a protein. It will be understood that the presence of an amino acid residue can still be determined by examining (e.g., sequencing) DNA. One would look for the codon encoding the amino acid being investigated and if the codon is present it indicates the amino acid is present. The genetic code of codons and their corresponding amino acids is well known in the art. In some embodiments, the method comprises receiving protein from the subject and sequencing the received protein. Methods of protein sequence and DNA sequencing are known in the art and either can be used to determine the presence of the recited amino acid at the specific site present in any one of the tables provided herein.

[0099] In some embodiments, the conversion comprises genome editing. In some embodiments, the genome editing is CRISPR-CAS genome editing complex. In some embodiments, the complex comprises a guide RNA (gRNA). In some embodiments, the gRNA is complementary to the subject’s genomic sequence comprising the lysine, arginine or glutamine to be converted. The design of gRNAs and programs for designing are well known in the art and any may be used for the purpose of designing a gRNA to target the genomic location to be converted. In some embodiments, converting comprises administering to the subject a genome editing compound. In some embodiments, a compound is a complex. In some embodiments, the compound is targeted to the residue. In some embodiments, the compound is targeted to the lysine. In some embodiments, the compound edits the subject’s genome to have an arginine instead of the lysine. In some embodiments, the compound edits the subject’s genome to have a glutamine instead of the lysine. In some embodiments, the compound is targeted to the arginine. In someembodiments, the compound edits the subject’s genome to have a lysine instead of the arginine. In some embodiments, the compound is targeted to the glutamine. In some embodiments, the compound edits the subject’s genome to have a lysine instead of the glutamine.

[0100] By another aspect, there is provided a method of extending the longevity of a subject, the method comprising exogenously expressing a protein in the subject, wherein the protein is selected from those provided in Table 1, thereby extending the longevity of a subject.

[0101] By another aspect, there is provided a method of extending the longevity of a subject, the method comprising exogenously expressing a protein in the subject, wherein the protein is selected from those provided in Table 5, thereby extending the longevity of a subject.

[0102] By another aspect, there is provided a method of extending the longevity of a subject, the method comprising exogenously expressing a protein in the subject, wherein the protein is selected from those provided in Table 2, thereby extending the longevity of a subject.

[0103] By another aspect, there is provided a method of extending the longevity of a subject, the method comprising exogenously expressing a protein in the subject, wherein the protein is selected from those provided in Table 3, thereby extending the longevity of a subject.

[0104] By another aspect, there is provided a method of extending the longevity of a subject, the method comprising exogenously expressing a protein in the subject, wherein the protein is selected from those provided in Table 4, thereby extending the longevity of a subject.

[0105] In some embodiments, expressing a protein comprises administering the protein to the subject. In some embodiments, expressing a protein comprises expressing a nucleic acid molecule encoding the protein. In some embodiments, expressing comprises administering a nucleic acid molecule encoding the protein. In some embodiments, the nucleic acid molecule is an RNA. In some embodiments, the RNA is an mRNA. In some embodiments, the nucleic acid molecule is a DNA. In some embodiments, the nucleic acid molecule is a vector. In some embodiments, the vector is an expression vector.

[0106] Expression of a gene / protein within a cell is well known to one skilled in the art. It can be carried out by, among many methods, transfection, viral infection, or direct alteration of the cell’s genome. In some embodiments, the gene is in an expression vector such as plasmid or viral vector.

[0107] A vector nucleic acid sequence generally contains at least an origin of replication for propagation in a cell and optionally additional elements, such as a heterologouspolynucleotide sequence, expression control element (e.g., a promoter, enhancer), selectable marker (e.g., antibiotic resistance), poly-Adenine sequence.

[0108] The vector may be a DNA plasmid delivered via non -viral methods or via viral methods. The viral vector may be a retroviral vector, a herpesviral vector, an adenoviral vector, an adeno-associated viral vector or a poxviral vector. The promoters may be active in mammalian cells. The promoters may be a viral promoter.

[0109] In some embodiments, a coding sequence encoding the protein is operably linked to a promoter. The term “operably linked” is intended to mean that the nucleotide sequence of interest is linked to the regulatory element or elements in a manner that allows for expression of the nucleotide sequence (e.g. in an in vitro transcription / translation system or in a host cell when the vector is introduced into the host cell).

[0110] In some embodiments, the vector is introduced into the cell by standard methods including electroporation (e.g., as described in From et al., Proc. Natl. Acad. Sci. USA 82, 5824 (1985)), Heat shock, infection by viral vectors, high velocity ballistic penetration by small particles with the nucleic acid either within the matrix of small beads or particles, or on the surface (Klein et al., Nature 327. 70-73 (1987)), and / or the like.

[0111] The term "promoter" as used herein refers to a group of transcriptional control modules that are clustered around the initiation site for an RNA polymerase i.e., RNA polymerase II. Promoters are composed of discrete functional modules, each consisting of approximately 7-20 bp of DNA, and containing one or more recognition sites for transcriptional activator or repressor proteins. In some embodiments, the promoter is a constitutive promoter. In some embodiments, the promoter is an inducible promoter. In some embodiments, the promoter is a tissue specific promoter. In some embodiments, the promoter is a cell or cell type specific promoter. In some embodiments, the cell is the target cell.

[0112] In some embodiments, nucleic acid sequences are transcribed by RNA polymerase II (RNAP II and Pol II). RNAP II is an enzyme found in eukaryotic cells. It catalyzes the transcription of DNA to synthesize precursors of mRNA and most snRNA and microRNA.

[0113] In some embodiments, mammalian expression vectors include, but are not limited to, pcDNA3, pcDNA3.1 (±), pGL3, pZeoSV2(±), pSecTag2, pDisplay, pEF / myc / cyto, pCMV / myc / cyto, pCR3.1, pSinRep5, DH26S, DHBB, pNMTl, pNMT41, pNMT81, which are available from Invitrogen, pCI which is available from Promega, pMbac, pPbac, pBK-RSV and pBK-CMV which are available from Strategene, pTRES which is available from Clontech, and their derivatives.

[0114] In some embodiments, expression vectors containing regulatory elements from eukaryotic viruses such as retroviruses are used by the present invention. SV40 vectors include pSVT7 and pMT2. In some embodiments, vectors derived from bovine papilloma virus include pBV-lMTHA, and vectors derived from Epstein Bar virus include pHEBO, and p2O5. Other exemplary vectors include pMSG, pAV009 / A+, pMTO10 / A+, pMAMneo- 5, baculovirus pDSVE, and any other vector allowing expression of proteins under the direction of the SV-40 early promoter, SV-40 later promoter, metallothionein promoter, murine mammary tumor virus promoter, Rous sarcoma virus promoter, polyhedrin promoter, or other promoters shown effective for expression in eukaryotic cells.

[0115] In some embodiments, recombinant viral vectors, which offer advantages such as lateral infection and targeting specificity, are used for in vivo expression. In one embodiment, lateral infection is inherent in the life cycle of, for example, retrovirus and is the process by which a single infected cell produces many progeny virions that bud off and infect neighboring cells. In one embodiment, the result is that a large area becomes rapidly infected, most of which was not initially infected by the original viral particles. In one embodiment, viral vectors are produced that are unable to spread laterally. In one embodiment, this characteristic can be useful if the desired purpose is to introduce a specified gene into only a localized number of targeted cells.

[0116] Various methods can be used to introduce the expression vector of the present invention into cells. Such methods are generally described in Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Springs Harbor Laboratory, New York (1989, 1992), in Ausubel et al., Current Protocols in Molecular Biology, John Wiley and Sons, Baltimore, Md. (1989), Chang et al., Somatic Gene Therapy, CRC Press, Ann Arbor, Mich. (1995), Vega et al., Gene Targeting, CRC Press, Ann Arbor Mich. (1995), Vectors: A Survey of Molecular Cloning Vectors and Their Uses, Butterworths, Boston Mass. (1988) and Gilboa et at. [Biotechniques 4 (6): 504-512, 1986] and include, for example, stable or transient transfection, lipofection, electroporation and infection with recombinant viral vectors. In addition, see U.S. Pat. Nos. 5,464,764 and 5,487,992 for positive-negative selection methods.

[0117] It will be appreciated that other than containing the necessary elements for the transcription and translation of the inserted coding sequence (encoding the polypeptide), theexpression construct of the present invention can also include sequences engineered to optimize stability, production, purification, yield or activity of the expressed polypeptide.

[0118] A person with skill in the art will appreciate that a gene can also be expressed from a nucleic acid construct administered to the individual employing any suitable mode of administration, described hereinabove (i.e., in vivo gene therapy). In one embodiment, the nucleic acid construct is introduced into a suitable cell via an appropriate gene delivery vehicle / method (transfection, transduction, homologous recombination, etc.) and an expression system as needed and then the modified cells are expanded in culture and returned to the individual (i.e., ex vivo gene therapy).

[0119] As used herein, the terms “administering,” “administration,” and like terms refer to any method which, in sound medical practice, delivers a composition containing an active agent to a subject in such a manner as to provide a therapeutic effect. Routes of administration include, but are not limited to parenteral, subcutaneous, intravenous, oral, intramuscular, and intraperitoneal.

[0120] The dosage administered will be dependent upon the age, health, and weight of the recipient, kind of concurrent treatment, if any, frequency of treatment, and the nature of the effect desired. In some embodiments, a therapeutically effective amount of the protein or nucleic acid encoding the protein is administered. The term "therapeutically effective amount" refers to an amount of a compound effective to treat a disease or disorder in a mammal. The term “a therapeutically effective amount” refers to an amount effective, at dosages and for periods of time necessary, to achieve the desired therapeutic or prophylactic result. The exact dosage form and regimen would be determined by the physician according to the patient's condition.

[0121] In some embodiments, extending longevity comprises extending the life span of the subject. In some embodiments, extending longevity comprises extending the life of the subject. In some embodiments, extending longevity comprises treating aging. In some embodiments, extending longevity comprises treating an aging related disease or condition. In some embodiments, an aging related disease is a metabolic disease. In some embodiments, an aging related disease is a liver disease. In some embodiments, an aging related disease is a respiratory disease. In some embodiments, an aging related disease is a muscle disease. In some embodiments, an aging related disease is a proliferative disease. In some embodiments, an aging related disease is a neurological disease. In some embodiments, an aging related disease is a joint disease.

[0122] As used herein the term “aging” refers to the natural deterioration over time of an organism, and specifically the cells of an organism. In some embodiments, aging comprises a diminished capacity of stem cells to produce differentiated cells. In some embodiments, aging comprises a diminished capacity of stem cells to self-renew. In some embodiments, aging comprises cells entering senescence. In some embodiments, aging comprises increased cell death. In some embodiments, aging comprises decreased cellular respiration. In some embodiments, aging comprises increased cellular reactive oxidation species (ROS). In some embodiments, aging comprises increased inflammation. In some embodiments, aging comprises increased fibrosis. In some embodiments, aging comprises increased scar tissue. In some embodiments, aging comprises cardiac heterotrophy. In some embodiments, aging comprises impaired glucose homeostasis.some embodiments, aging comprises reduced cognitive function.some embodiments, aging comprises reduced or impaired memory.some embodiments, aging comprises reduced chondrocyte survival.

[0123] In some embodiments, aging is skin aging. In some embodiments, aging is muscle aging. In some embodiments, aging is not muscle aging. In some embodiments, aging is any aging except skin and muscle aging. In some embodiments, aging is neuronal aging. In some embodiments, aging is pancreatic aging. In some embodiments, aging is joint aging. In some embodiments, aging is brain aging. In some embodiments, aging is selected from muscle, neuronal, pancreatic and joint aging. In some embodiments, aging is selected from neuronal, pancreatic and joint aging.

[0124] In some embodiments, aging comprises decreased cognitive function. In some embodiments, aging comprises decreased muscle mass. In some embodiments, aging comprises decreased hormone production. In some embodiments, aging comprises impaired reflexes. In some embodiments, aging comprises impaired function of one of the systems of the body, including, but not limited to, the circulatory system, the muscular-skeletal system, the immune system, the respiratory system, the nervous system, the digestive system, the limbic system, glucose homeostatic system, neuro-muscular systemjoint system and the renal system.

[0125] As used herein, an “aging-associated disease” refers to a disease of old age. In some embodiments, an aging-associated disease refers to a condition or disease whose prevalence increases with age. In some embodiments, an aging-associate disease is a disease that occurs with increasing frequency when there is increasing or increased cellular senescence.Examples of aging-associated diseases include, but are not limited to arthritis, cardiovascular disease, cancer, diabetes, dementia, Alzheimer’s disease, Parkinson’s disease, hypertension, stroke, osteoporosis, chronic respiratory disease (COPD), hearing loss, vision loss, cachexia and sarcopenia, to name but a few.

[0126] In some embodiments, a metabolic disease is selected from hypoglycemia, liver disease, hepatitis, hyperammonemia, hepatomegaly, diabetes, fatty liver disease, cardiovascular disease, and hyperlipidemia. In some embodiments, a liver disease is selected from hypoglycemia, liver disease, hepatitis, hyperammonemia, hepatomegaly, diabetes, fatty liver disease, and hyperlipidemia. In some embodiments, a muscle disease is selected from sarcopenia, cachexia, muscle fibrosis and myopathy. In some embodiments, a muscle disease is myopathy. In some embodiments, a respiratory disease is respiratory insufficiency. In some embodiments, a proliferative disease is cancer. In some embodiments, neurological disease is a neurodegenerative disease. In some embodiments, the neurological disease is selected from: impaired memory, impaired cognitive function, dementia, stroke-related brain damage, polyneuropathy and Alzheimer's disease. In some embodiments, neurological disease is selected from polyneuropathy and Alzheimer’s disease. In some embodiments, a joint disease is arthritis. In some embodiments, an aging related disease is Hutchinson- Gilford Progeria Syndrome (HGPS).

[0127] In some embodiments, related is associated. In some embodiments, an aging related disease is an aging associated impairment. The aging-associated impairment may manifest in a number of different ways, e.g., as aging-associated cognitive impairment and / or physiological impairment, e.g., in the form of damage to central or peripheral organs of the body, such as but not limited to: cell injury, tissue damage, organ dysfunction, aging associated lifespan shortening and carcinogenesis, where specific organs and tissues of interest include, but are not limited to skin, neuron, muscle, pancreas, brain, kidney, lung, stomach, intestine, spleen, heart, adipose tissue, testes, ovary, uterus, liver and bone; in the form of decreased neurogenesis, etc. In some embodiments, the aging-associated impairment is an aging-associated impairment in cognitive ability in an individual, i.e., an aging-associated cognitive impairment. By cognitive ability, or "cognition", it is meant the mental processes that include attention and concentration, learning complex tasks and concepts, memory (acquiring, retaining, and retrieving new information in the short and / or long term), information processing (dealing with information gathered by the five senses), visuospatial function (visual perception, depth perception, using mental imagery, copying drawings, constructing objects or shapes), producing and understanding language, verbalfluency (word-finding), solving problems, making decisions, and executive functions (planning and prioritizing). By "cognitive decline", it is meant a progressive decrease in one or more of these abilities, e.g., a decline in memory, language, thinking, judgment, etc. By "an impairment in cognitive ability" and "cognitive impairment", it is meant a reduction in cognitive ability relative to a healthy individual, e.g., an age-matched healthy individual, or relative to the ability of the individual at an earlier point in time, e.g., 2 weeks, 1 month, 2 months, 3 months, 6 months, 1 year, 2 years, 5 years, or 10 years or more previously. Aging- associated cognitive impairments include impairments in cognitive ability that are typically associated with aging, including, for example, cognitive impairment associated with the natural aging process, e.g., mild cognitive impairment (M.C.I.); and cognitive impairment associated with an aging associated disorder, that is, a disorder that is seen with increasing frequency with increasing senescence, e.g., a neurodegenerative condition such as Alzheimer's disease, Parkinson's disease, frontotemporal dementia, Huntington's disease, amyotrophic lateral sclerosis, multiple sclerosis, glaucoma, myotonic dystrophy, vascular dementia, and the like. By "treatment" it is meant that at least an amelioration of one or more symptoms associated with an aging-associated impairment afflicting the adult mammal is achieved, where amelioration is used in a broad sense to refer to at least a reduction in the magnitude of a parameter, e.g., a symptom associated with the impairment being treated. As such, treatment also includes situations where a pathological condition, or at least symptoms associated therewith, are completely inhibited, e.g., prevented from happening, or stopped, e.g., terminated, such that the adult mammal no longer suffers from the impairment, or at least the symptoms that characterize the impairment. In some instances, "treatment", "treating" and the like refer to obtaining a desired pharmacologic and / or physiologic effect. The effect may be prophylactic in terms of completely or partially preventing a disease or symptom thereof and / or may be therapeutic in terms of a partial or complete cure for a disease and / or adverse effect attributable to the disease. "Treatment" may be any treatment of a disease in a mammal, and includes: (a) preventing the disease from occurring in a subject which may be predisposed to the disease but has not yet been diagnosed as having it; (b) inhibiting the disease, i.e., arresting its development; or (c) relieving the disease, i.e., causing regression of the disease. Treatment may result in a variety of different physical manifestations, e.g., modulation in gene expression, increased neurogenesis, rejuvenation of tissue or organs, etc. Treatment of ongoing disease, where the treatment stabilizes or reduces the undesirable clinical symptoms of the patient, occurs in some embodiments. Such treatment may be performed prior to complete loss of function in the affected tissues. The subject therapy may be administered during the symptomatic stage of the disease, and insome cases after the symptomatic stage of the disease. In some instances where the aging- associated impairment is aging-associated cognitive decline, treatment by methods of the present disclosure slows, or reduces, the progression of aging-associated cognitive decline. In other words, cognitive abilities in the individual decline more slowly, if at all, following treatment by the disclosed methods than prior to or in the absence of treatment by the disclosed methods. In some instances, treatment by methods of the present disclosure stabilizes the cognitive abilities of an individual. For example, the progression of cognitive decline in an individual suffering from aging-associated cognitive decline is halted following treatment by the disclosed methods.

[0128] As used herein, the terms “treatment” or “treating” of a disease, disorder, or condition encompasses alleviation of at least one symptom thereof, a reduction in the severity thereof, or inhibition of the progression thereof. Treatment need not mean that the disease, disorder, or condition is totally cured. To be an effective treatment, a useful composition or method herein needs only to reduce the severity of a disease, disorder, or condition, reduce the severity of symptoms associated therewith, or provide improvement to a patient or subject’s quality of life.

[0129] By another aspect, there is provided a method of determining suitability of a human subject to have their longevity extended by a method of the invention, the method comprising receiving DNA from the subject and determine the DNA encodes a protein that comprises a lysine at a position provided in Table 1, 2 or 5, wherein a presence of a lysine at the position indicates the subject is suitable, thereby determining suitability.

[0130] By another aspect, there is provided a method of determining suitability of a subject to have their longevity extended by a method of the invention, the method comprising receiving DNA from the subject and determine the DNA encodes a protein that comprises an arginine at a position provided in Table 3, wherein a presence of an arginine at the position indicates the subject is suitable, thereby determining suitability.

[0131] By another aspect, there is provided a method of determining suitability of a subject to have their longevity extended by a method of the invention, the method comprising receiving DNA from the subject and determine the DNA encodes a protein that comprises a glutamine at a position provided in Table 4, wherein a presence of a glutamine at the position indicates the subject is suitable, thereby determining suitability.

[0132] As used herein, the term "about" when combined with a value refers to plus and minus 10% of the reference value. For example, a length of about 1000 nanometers (nm) refers to a length of 1000 nm+- 100 nm.

[0133] It is noted that as used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a polynucleotide" includes a plurality of such polynucleotides and reference to "the polypeptide" includes reference to one or more polypeptides and equivalents thereof known to those skilled in the art, and so forth. It is further noted that the claims may be drafted to exclude any optional element. As such, this statement is intended to serve as antecedent basis for use of such exclusive terminology as "solely," "only" and the like in connection with the recitation of claim elements, or use of a "negative" limitation.

[0134] In those instances where a convention analogous to "at least one of A, B, and C, etc." is used, in general such a construction is intended in the sense one having skill in the art would understand the convention (e.g., "a system having at least one of A, B, and C" would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). It will be further understood by those within the art that virtually any disjunctive word and / or phrase presenting two or more alternative terms, whether in the description, claims, or drawings, should be understood to contemplate the possibilities of including one of the terms, either of the terms, or both terms. For example, the phrase "A or B" will be understood to include the possibilities of "A" or "B" or "A and B."

[0135] It is appreciated that certain features of the invention, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the invention, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable sub-combination. All combinations of the embodiments pertaining to the invention are specifically embraced by the present invention and are disclosed herein just as if each and every combination was individually and explicitly disclosed. In addition, all subcombinations of the various embodiments and elements thereof are also specifically embraced by the present invention and are disclosed herein just as if each and every such sub-combination was individually and explicitly disclosed herein.

[0136] As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include plural referents, unless the context clearly dictates otherwise. The terms“a” (or “an”) as well as the terms “one or more” and “at least one” can be used interchangeably.

[0137] Furthermore, “and / or” is to be taken as specific disclosure of each of the two specified features or components with or without the other. Thus, the term “and / or” as used in a phrase such as “A and / or B” is intended to include A and B, A or B, A (alone), and B (alone). Likewise, the term “and / or” as used in a phrase such as “A, B, and / or C” is intended to include A, B, and C; A, B, or C; A or B; A or C; B or C; A and B; A and C; B and C; A (alone); B (alone); and C (alone).

[0138] Wherever embodiments are described with the language “comprising,” otherwise analogous embodiments described in terms of “consisting of’ and / or “consisting essentially of’ are included.

[0139] Additional objects, advantages, and novel features of the present invention will become apparent to one ordinarily skilled in the art upon examination of the following examples, which are not intended to be limiting. Additionally, each of the various embodiments and aspects of the present invention as delineated hereinabove and as claimed in the claims section below finds experimental support in the following examples.

[0140] Various embodiments and aspects of the present invention as delineated hereinabove and as claimed in the claims section below find experimental support in the following examples.EXAMPLES

[0141] Generally, the nomenclature used herein and the laboratory procedures utilized in the present invention include molecular, biochemical, microbiological and recombinant DNA techniques. Such techniques are thoroughly explained in the literature. See, for example, "Molecular Cloning: A laboratory Manual" Sambrook et al., (1989); "Current Protocols in Molecular Biology" Volumes I-III Ausubel, R. M., ed. (1994); Ausubel et al., "Current Protocols in Molecular Biology", John Wiley and Sons, Baltimore, Maryland (1989); Perbal, "A Practical Guide to Molecular Cloning", John Wiley & Sons, New York (1988); Watson et al., "Recombinant DNA", Scientific American Books, New York; Birren et al. (eds) "Genome Analysis: A Laboratory Manual Series", Vols. 1-4, Cold Spring Harbor Laboratory Press, New York (1998); methodologies as set forth in U.S. Pat. Nos. 4,666,828; 4,683,202; 4,801,531; 5,192,659 and 5,272,057; "Cell Biology: A Laboratory Handbook", Volumes I-Ill Cellis, J. E., ed. (1994); "Culture of Animal Cells - A Manual of Basic Technique" by Freshney, Wiley-Liss, N. Y. (1994), Third Edition; "Current Protocols in Immunology" Volumes I-III Coligan J. E., ed. (1994); Stites et al. (eds), "Basic and Clinical Immunology" (8th Edition), Appleton & Lange, Norwalk, CT (1994); Mishell and Shiigi (eds), "Strategies for Protein Purification and Characterization - A Laboratory Course Manual" CSHL Press (1996); all of which are incorporated by reference. Other general references are provided throughout this document.Materials and Methods

[0142] The research conducted in this study complies with all relevant ethical regulations. All procedures involving mice and experimental protocols were approved by the Institutional Animal Care and Use Committee of Bar Ilan University and by the Ministry of Health of Israel.

[0143] Animal experimentation and acetylome generation: Mice were housed on a 12h light / dark cycle at room temperature (22-24°C) with a relative humidity within 45-65% in the Bar Ilan animal facilities. Mice were maintained on a standard rodent chow diet with ad libitum access to food and water. All mice used in these experiments were in good health. Mice were kept under specific pathogen-free conditions in IVC cages that were routinely screened and found negative for viral serology and both endo and ectoparasites. Male littermates were used for all experiments.

[0144] To generate mouse acetylome in the lab, hepatocytes or liver tissue of ten young (6 months old) and eighteen old (25 months old) C57BL / J6 mice were taken and processed. Two pan-acetyl lysine antibodies (Cell Signaling, 13416 and Immunochem, ICP0388-2MG) were used for acetylated peptides enrichment, following Cell Signaling protocol.

[0145] Orthologous proteins and phylogenetic tree prediction: OrthoFinder 2.5.4 was applied to compute the orthologs of mouse or human proteins in 107 mammals. To reduce computation time, the proteomes of human / mouse were grouped with those of other mammals within separate directories, allowing the algorithm to be executed in parallel across these directories. Next, a phylogenetic tree comprising all 107 mammals was constructed based on the calculation of all versus all orthologs.

[0146] PHARAOH analyzed two phylogenetic trees, the Zoonomia project tree (Vignieri, S., “ZOONOMIA”, Science (1979) 380, 356-357 (2023), the contents of which are hereby incorporated by reference in their entirety), containing 62 animals overlapping with the dataset, and the OrthoFinder-based tree (Emms, et al., “OrthoFinder: Phylogenetic orthologyinference for comparative genomics”, Genome Biol 20, 1-14 (2019), the contents of which are hereby incorporated by reference in their entirety) using the 107-mammals’ dataset (Fig. 8A). Final outcomes were derived from the intersection of significant findings obtained from both analyses. The distribution of these results is presented in Figures 8B-8E.

[0147] Phylogenies and comparative methods

[0148] Statistical validation: Since during evolution different organisms evolved from common ancestors, it introduces a level of dependence between animals. Thus, the standard assumption of variable independence cannot be used for the statistical analysis of the PHARAOH. To acknowledge the fact that species are not fully independent, a novel statistical test was devised that incorporates the evolutionary context. This allows the examination of the relationship between amino acid (AA) changes and longevity, regardless of phylogenetic distances. The statistical test consisted of two parts. First, the correlation between maximal lifespan and AA replacement was examined; then, a correction was introduced for the distances between the animals on the evolutionary tree.

[0149] Part I (Fig. 7A) - Correlation between amino acid exchange and longevity-

[0150] Data Preparation: The acetylome datasets used for the replacement matrix are based on the 14 AA flanking the amino acid of interest. An entry in the matrix is denoted as: Mi,j = the amino acid found in mammal j, at the orthologous site to the i acetylation site in human / mouse.

[0151] For each row in the replacement matrix, the mammals were divided into two groups based on the amino acid found at the specific acetylated site: lysine site of reference (K) (includes either mouse or human), or arginine (R) / glutamine (Q) / other. Each group contains the maximal lifespan in captivity for each member.

[0152] Correlation Test: To identify AA substitutions (K-to-R / Q) that significantly correlate with longevity, the Mann-Whitney test was employed to compare the lifespans between the two above-mentioned groups that were divided based on their amino acid found at the acetylation site. A minimum of three mammals was required in each group to ensure significant findings. False Discovery Rate (FDR) correction was applied to account for multiple testing, and a significance level of q-value < 0.1 was adopted.

[0153] Part II (Fig. 7B-7C) - Phylogenetic Correction.

[0154] The assumption that lifespan difference between the two mammals in each comparison is derived from an exchange of the examined acetylated lysine, without contributions from phylogeny was taken as the Ho hypothesis.

[0155] To normalize the impact of the phylogenetic tree on the statistical validation process, each acetylation site (represented by a row in the replacement matrix) was further divided into a set of pairs derived from the database. The relationship between a pair of mammals is represented by three factors (Fig. 7B):• Lifespan difference.• Phylogenetic (tree) distance.• The amino acid (AA) found in the examined site.

[0156] Importantly, this part is independent of Part I, and aims to examine if the difference in lifespan is due to changes in amino acids or their relative position on the evolutionary tree.

[0157] Strengthening, weakening, and non -informative categories: The entries in the table shown in Figure 7B are categorized into three distinct groups based on their impact on the hypothesis: strengthening, weakening, or not informative.• Claims that are not informative with regard to the hypothesis (colored in gray in the table): o Pairs that have a small difference in lifespan, small distance in the phylogenetic tree and have the same amino acid at the acetylation site (e.g. human and chimp sharing the same AA at the specific site). o Pairs that have a large difference in lifespan, large distance in the phylogenetic tree and have different amino acids at the acetylation site (e.g. mouse and human that have a different AAs at a specific site).• Claims that strengthen the hypothesis (colored in green in the table): o Pairs that have a small difference in lifespan, large distance in the phylogenetic tree and have the same amino acid at the acetylation site (e.g. human and whale sharing the same AA at a specific site). o Pairs that have a large difference in lifespan, small distance in the phylogenetic tree and have different amino acids at the acetylation site (e.g. mouse and naked mole rat, having different AAs at a specific site).• Claims that weaken the hypothesis (colored in red in the table):o Pairs that have a small difference in lifespan, small distance in the phylogenetic tree and have different amino acids at the acetylation site (e.g. human and chimp having a different AA at a specific site). o Pairs that have a small difference in lifespan, large distance in the phylogenetic tree, and have different amino acids at the acetylation site (e.g. human and whale having different AAs at a specific site). o Pairs that have a large difference in lifespan, small distance in the phylogenetic tree, and have the same amino acid in the acetylation site (e.g. mouse and naked mole rat having the same AA at a specific site). o Pairs that have a large difference in lifespan, large distance in the phylogenetic tree, and have the same amino acid at the acetylation site (e.g. mouse and human having the same AA at a specific site).

[0158] Weight Calculation: Weights were assigned to each comparison according to the following equation:(1) Weight = ±P (\lifespanl — Ufespan2\ A tree_distance)Where P is the probability of the mammalian pair occupying a specific position in the table (Fig. 7B, right column), regardless of the amino acid found at the acetylation site. For instance, the probability of a pair from the data falling into the first strengthening row is determined by their lifespan difference being smaller than average, AND their distance in the phylogenetic tree being larger than average. The strengthening or weakening category is assigned positive or negative weights, respectively.

[0159] Score Calculation: Each acetylation site was addressed as a collection of mammalian pairs, with each pair assigned a position / category in the options table.

[0160] The overall score for an acetylation site was computed as shown in the following equation:In which the score for each site is calculated as the sum of the values received from all two- mammal comparisons: weight multiplied by the difference in lifespans and the distance in the phylogenetic tree (both normalized to numbers between 0-1, see Fig. 7B).

[0161] Permutation Test: To determine a p-value for each site's score, 1,000 permutations of the non-zero weights were performed, effectively randomizing the positions where pairs are placed in the tree. FDR correction was applied to adjust for multiple testing.

[0162] Changed acetylation sites are considered longevity-associated only in case they are significant in both Mann-Whitney test and the phylogenetic tree validation described here (Fig. 7C).

[0163] PHARAOH implementation: The PHARAOH tool was fully implemented in Python3. The replacement matrix algorithm uses SQLite3 tables as the database. Pandas, more i tertools, and NumPy python libraries were used throughout creation of the matrix. In the statistical validation Bio (biopython) was used for working with phylogenetic data (Phylo), as well as statsmodels and scipy. stats for the Mann-Whitney and FDR tests. Matplotlib was used for basic visualization of the outputs.

[0164] Lysine residue conservation: PHARAOH was employed to investigate the conservation of lysine residues. Matplotlib was used for basic visualization, with or without acetylation, between mouse and human, by comparing proteomics / acetylome data. The conservation patterns within the proteomes were independently assessed. Orthologous proteins were obtained through OrthoFinder, and the alignment of these orthologous protein sequences was conducted using the Linux version of bl2seq. To identify unacetylated lysines, the acetylome sites were subtracted from the proteome outcomes. Subsequently, the percentage of each amino acid replacement was calculated in relation to the total count of lysines in the proteome, or the acetylated lysines present within the acetylome.

[0165] Analysis of the acetylome results: Pathway enrichment analyses were performed via the Metascape site, which uses MCODE and various interaction tools to create the proteinprotein interaction network. Cytoscape was used to edit and the visualize the PPI networks.

[0166] Gene ontology was used for cell component (CC) analysis, and the results were manually divided into categories of extracellular, nucleus, cytoplasm, plasma membrane, mitochondrion, endoplasmic reticulum, and others.

[0167] pLogo (probability Logo generator) was used with the appropriate background (mouse background for the mouse acetylome and human for the human acetylome).

[0168] TRRUST database was used for human transcription factor analysis. Gene-disease association analysis was performed using DisGeNET.

[0169] R was used via RStudio to enable visualization of the results, ggtree was used for visualization of the phylogenetic tree, and gheatmap was used to append the heatmaps to the phylogenetic tree. Heatmap of the replacement matrix data was created by ComplexHeatmap, filtered only to the significant results and the relevant changes of AAs. Proteomics data from the NIH Cancer Institute (Thangudu, et al., “Abstract LB-242: Proteomic Data Commons: A resource for proteogenomic analysis”, Cancer Res 80, LB-242 (2020), the contents of which are hereby incorporated by reference in its entirety) was used for Pearson correlation between USP10 and PCNA in cancer data, visualized and calculated using R. Venn diagrams for the methods visualizations were also created in R using eulerr.

[0170] Cloning and mutagenesis: The human and mouse CBS and the human USP10 cDNA sequences were cloned to a pcDNA3.1+ vector using EcoRI and Xhol (CBS) or Xhol and Xbal (USP10) restriction enzymes. All genes were tagged with a flag tag at the C-terminal region. Site directed mutagenesis was done with PfuUltra II Fusion HS DNA Polymerase kit (Agilent), following the manufacturer's protocol.

[0171] Cell culture and treatments: HEK293T and HCT116 cell lines were purchased from ATCC (cat. CRL-3216 and CCL-247 respectively) and grown in Dulbecco's Modified Eagle Medium (DMEM) supplemented with 10% fetal bovine serum (FBS), 1% glutamine and 1% penicillin / streptomycin in 5% CO2. For PCNA ubiquitin-related degradation, HCT116 cells were treated with lOOpg / ml cycloheximide for 8 h, and then harvested with PBSxl and lysed in urea buffer for western blot analysis.

[0172] Transfection: HEK293T cells were seeded in 10cm plates. The cells were grown for 48 h before harvesting in cold PBS xl. HCT116 cells were transfected using Lipofectamine 3000 (Thermo L3000008) following the manufacturer's protocol.

[0173] Flag immunoprecipitation (IP-flag): HEK293T transfected with CBS-flag were harvested in cold PBS xl and lysed on ice with lysis buffer (50mM Tris pH 7.4, 150mM NaCl, 1% Triton, 0.5% NP40, 10% glycerol) for 30 min. CBS-flag was immunoprecipitated from Img lysate and used for H2S production assay.

[0174] H2S production capacity assay: H2S capacity was measures using 200-500pg cell lysate or 5 pl eluate after IP-flag of CBS, supplemented with lOmM L-cysteine, lOpM PLP and lOmM L-homocysteine. Lead acetate papers were purchased from Sigma (37104-1EA). Measurements were quantified using Imaged.

[0175] Western blot and antibodies: Western blot analyses were done using antibodies for PCNA (Cell Signaling 13110S) or USP10 (Cell Signaling CST-8501S) as primary antibodyand the appropriated HRP-conjugated secondary antibodies (ENCO). HRP -conjugated primary antibodies were used against Ubiquitin (Santa Cruz 8017), Flag (Proteintech HRP- 66008) and Tubulin (Proteintech HRP-66240).

[0176] Statistical analysis for H2S production and western blot analyses: For the analysis of H2S production capacity assay and western blot measurements, statistical significance was tested using one-way ANOVA.

[0177] Data and code availability: Data and code details are available in Feldman-Trabelsi et al., “The mammalian longevity associated acetylome”, Nat Commun. 2025 Apr 22;16(1):3749 by the inventors, the contents of which are hereby incorporated by reference in their entirety as well as at github.com / hcohenlab / PHARAOH-tool.Example 1: Identifying longevity-associated acetylation sites

[0178] To identify PTMs that potentially regulate the extension in lifespan in long-lived organisms, the inventors employed the PHARAOH computational tool. PHARAOH compares the conservation of PTM sites, and whether each site was replaced by another amino acid in a set of all identified mammalian orthologous proteins. Next, PHARAOH uses sequence and lifespan data to examine whether any PTM or specific amino acid (AA) replacement is associated with longer lifespan. Thus, PHARAOH is utilized to search for specific lysine acetylation sites that regulate longevity. Global mouse and human acetylomes were created based on identified acetylation data found in the PHOSIDA and PhosphoSite databases together with the mouse acetylome generated by the inventors, and consisted of 25,959 and 22,849 mouse and human acetylation sites, respectively (Fig. 1A, left panel). For this study both high and low throughput PhosphoSite analyses were used. To identify the acetylation sites that are significantly associated with longevity, three additional datasets were created. An orthologues protein dataset and phylogenetic tree were created using the OrthoFinder tool based on 107 mammalian proteomes from Uniprot (Fig. 1A, middle panel). In addition, a maximum lifespan dataset for these animals was generated, based on The Animal Ageing and Longevity Database (AnAge, Tacutu, et al., “Human Ageing Genomic Resources: new and updated databases”, Nucleic Acids Res 46, D1083-D1090 (2018) the contents of which are hereby incorporated by reference in their entirety) (Fig. 1A, right panel). Specifically, during evolution, an acetylated / deacetylated lysine (K) can be converted into arginine (K-to-R) or glutamine (K-to-Q), mimicking a permanently deacetylated or acetylated lysine, respectively (Fig. 6A). Interestingly, acetylated lysines tend to be more conserved than unacetylated lysines (%2, =0.08, Fig. 6B) in comparison between mouse and human. Thus, the PHARAOH tool calculates the significance of correlation betweenconservation of a given acetylated site versus conversion to R or Q, and maximal lifespan (Fig. IB)

[0179] A replacement matrix was built based on pairwise sequence alignment between each acetylated peptide and the orthologous mammalian protein. Using the replacement matrix, maximal lifespan data and the phylogenetic tree, PHARAOH calculates the statistical significance of the correlation between an acetylation / replacement site and longevity (Fig. IB and Fig. 7A-7C, see Materials and Methods for a more detailed description). The OrthoFinder phylogenetic tree of 107 mammals was validated with the recently published Zoonomia project phylogenetic tree of 62 overlapping mammals (Fig. 8A). Importantly, using a correction based on the phylogenetic tree enabled the inventors to eliminate the influence of evolutionary distances between different mammals. This ensures that any bias introduced by other evolutionary factors is neutralized. A representative output of the analysis is shown in Figure 1C. The average age of the 107 mammals was 30.31 years.Example 2: R / Q conversion of acetylation sites in longevity

[0180] Comparison of the acetylation sites contained in the mouse and human acetylomes with the orthologues' dataset revealed 321 lysine to arginine (K-to-R, Table 1) and 161 lysine to glutamine (K-to-Q, Table 2) substitutions that were significantly associated with longer lifespan (Mann Whitney FDR < 0.1; phylogenetic tree statistical correction FDR < 0.05) (Fig. 8B-8C). Importantly, the same analysis with K-to-L, as a random control, identified only 8 longevity associated substitutions. These sites are only 3% of the total K-to-L replacements, between mouse and human, demonstrating the insignificance of this process in comparison to K-to-R / Q that are consisting of -40% of such substitutes. In addition, in comparison to the global acetylome, K-to-R / Q replacements did not show significant enrichment in any specific protein domains or non-domain regions across the 9,251 known domains in the mouse proteome (based on UniProt). Similar to the global acetylome, the largest group was mapped to disordered regions, suggesting that acetylation may play a role in stabilizing undefined structures. Interestingly, as seen in Figure 9, with the increase in lifespan during evolution, the conversion from K to either R or Q happened gradually, rather than after a specific lifespan threshold. Pathway analyses and protein-protein interaction (PPI) of the K-to-R sites assigned these changes to proteins associated with protein translation and folding, as well with many metabolic pathways, including those of fatty acid metabolism, PPAR signaling, the TCA cycle, amino acid biosynthesis, and the one-carbon / TSP cycle. (Fig. 2A-2B). Pathway analyses of the K-to-Q sites identified cytochrome p450, peroxisome, mitochondrial translation, taurine metabolism, fatty acid P-oxidation, andothers (Fig. 2C-2D). Importantly, pathway analyses of the human orthologs of these mouse proteins revealed a similar set of pathways, demonstrating that their role in longevity is potentially conserved (Fig. 10A-10B).

[0181] While maximal and average lifespan are correlated, their regulatory pathways may differ. The inventors analyzed K-to-R / Q replacements using an average lifespan dataset of 93 mammals, based on Animal Diversity, Max Planck Longevity Records, and other datasets. Pathways unique to average lifespan included beta oxidation of long fatty acids, NAD / NADH metabolism, and glycolysis / gluconeogenesis. For maximal lifespan, the Alzheimer’s disease pathway was prominent. Despite these differences, over 80% of the pathways were common, suggesting that analyses based on maximal lifespan provide insights into both median and maximal lifespan. Interestingly, major members of these pathways were previously found to be associated with longevity. For example, the ratelimiting enzyme of the TSP, CBS, is a main producer of hydrogen sulfide (H2S) which was found to mediate the effects of a CR diet on longevity. Cellular component (CC) analyses, based on Gene Ontology (GO) annotations, of the identified sites revealed that these proteins are localized mainly to the cytosol and the mitochondria (Fig. 2E). No K-to-R sites were found in the nuclear proteins, nor K-to-Q sites in proteins of the plasma membrane and the ER.

[0182] In order to identify the specific acetylation consensus sequence, pLogo analysis was performed on the whole mouse acetylome, and no consensus sequence was found. Yet, an enrichment for lysine residues flanking the acetylated K was found, probably due to other acetylation sites nearby (Fig. 2F). Importantly, these positive K’s and R were eliminated on positions -1,-2 and were replaced with DZE on these positions. F / Y were enriched on position +1. These findings might have implications for the enzymatic activity of KATs / KDACs. pLogo analyses of main CC as cytosol, mitochondria and nucleus showed that DZE at positions -1,-2 are mostly from mitochondrial and cytosolic acetylated proteins. Whereas the G / A at position -1 originated form nuclear proteins (Fig. 11A-11C). However, pLogo analysis on K-to-R sites identified overrepresented DZEAV at positions -3, -2, -1 and an enrichment for K downstream to the acetylation site. Additionally, R / YK / S were underrepresented at positions -2, -1, and T on position +1 (Fig. 2G). Interestingly, tyrosine, serine and threonine (Y, S and T, respectively) can potentially be phosphorylated, suggesting that additional PTMs flanking the acetylation sites may also be associated with lifespan. For K-to-Q sites, pLogo analysis identified hydrophobic amino acids VVL at positions -3, -2, -1 and DGK at positions +1+2+3 (Fig. 2H). This finding suggests that specific KATs / KDACs may regulate the acetylation status of longevity-associated acetylation sites.

[0183] Since the effect of these conversions is expressed in long-lived animals, the inventors searched for transcription factors that are known to regulate the enriched pathways in humans. Such factors can provide an additional regulatory layer on ageing. SP1, PPARG, PPARA, SREBF1, ATF2, RelA, VDR and NFKB1 regulators were found within K-to-R longevity-associated sites (Fig. 12A). Importantly, these regulators are significantly associated with ageing mechanisms such as inflammation (RelA of NFKB1), DNA repair (ATF2 and SP1), and metabolism (PPARG, PPARA and VDR). The same analysis identified SP1 and NFE2L2 / NRF2 transcription factors in the K-to-Q set (Fig. 12B). Similarly, NFE2L2 / NRF2 is involved in ageing-related pathways, such as protection against oxidative stress, and likely mediates CR protection against carcinogenesis. In addition, gene-disease association (GO DisGeNET) analysis was performed in order to examine which diseases are associated with longevity related- K-to-R or K-to-Q sites in humans. Interestingly, the majority of the most highly significant results were of metabolic diseases, particularly liver- related, and other diseases such as diabetes and other metabolic related pathologies such as myopathy and lethargy, all known to be ageing-related (Fig. 12C-12D).

[0184] Table 1 : 321 mouse K to long living mammals (e.g., human) R sites. The number given after the gene name is the amino acid position of the K in the mouse protein.0185] Table 2: 161 mouse K to long living mammals (e.g., human) Q sites. The number given after the gene name is the amino acid position of the K in the mouse protein.Example 3: CBS K386R enhances its pro-longevity activity

[0186] Next, the inventors further explored the role of a longevity-associated acetylation site on CBS, the mediator of the CR response. CBS K386 was acetylated in the mouse reference dataset and exchanged with R in long-lived mammals. The average lifespan of mammals with K-to-R conversion, such as humans, was significantly higher than that of animals with conserved K at this site (Mann Whitney FDR < 0.1, q value=0.00014; tree statistical correction FDR < 0.001) (Fig. 3A-3B and 13). Thus, the inventors next followed the effect of K386R replacement on CBS FFS production activity. Human embryonic kidney (HEK) 293T cells overexpressing either WT, K386Q, or K386R mouse CBS were examined for their H2S production capacity using the lead acetate method. As seen in Figure 3C, in comparison to cells overexpressing the WT protein, overexpression of K386R CBS resulted in significantly higher H2S production capacity (p<0.05). Cells expressing K386Q CBS had similar H2S production capacity as cells expressing the WT protein. In the reverse experiment, H2S production activity of immunoprecipitated WT, R389K or R389Q human CBS was tested. While activity of the R389K mutant was similar to the activity of the WT protein, the acetylation-mimicking mutant R389Q exhibited a significantly reduced H2S production capacity (p<0.001, Fig. 3D). The non-significant trend of lower H2S production of R389K compared to the WT CBS is most likely due to low percentages of acetylation on the mutated lysine. All together, these findings show that constitutive deacetylation of the CBS K386 residue, as found in long-lived animals, promotes H2S production and potentially contributes to their longer lifespan.Example 4: The long-lived acquired acetylome

[0187] During evolution, the specific generation of new acetylation sites would be a complementary set of events to the above-mentioned replacements. Specifically, this would involve the replacement of R and Q residues present in short-lived animals with acetylated K residues in long-lived mammals (R-to-K and Q-to-K, respectively). Therefore, using PHARAOH, the conservation of human acetylated K sites among all mammalian orthologs was examined. Particularly, the data was searched for R-to-K and Q-to-K sites within the 22,849 human acetylation sites. The inventors identified 495 R-to-K (Table 3) and 150 Q- to-K (Table 4) sites that were significantly associated with longer lifespan (Mann Whitney FDR < 0.1, tree statistical correction FDR < 0.05). In comparison to the global human acetylome, R-to-K replacements did not show significant enrichment in any specific protein domains or non-domain regions across the 9,767 known domains in the human proteome (based on UniProt). However, a comparison of the total number of acetylation sites mapped to domains between the global human acetylome and R-to-K replacements revealed asignificant association, which was not observed for Q-to-K replacements. Similar to mouse data, the largest group was mapped to disordered regions. Enrichment analyses of the R-to- K sites assigned these changes to various pathways, including chromosome organization, amino acid metabolism, and cell cycle (Fig. 4A). Strikingly, PPI analysis identified many ageing-related pathways in this dataset, including cellular respiration, ribosomal biogenesis, regulation of protein translation, response to stress, and mismatch DNA repair and diseases of DNA repair pathways (Fig. 4B and 14A). Analyses of Q-to-K events also identified many ageing-related pathways, such as response to stress, the TCA cycle, sulfur compound metabolic process (including proteins from the TSP related pathway, one-carbon pool by folate cycle), and response to starvation (Fig. 4C). PPI analysis also identified ribosomal biogenesis and DNA repair pathways (Fig. 14B). Interestingly, CC analyses of all identified R / Q-to-K sites revealed that these proteins localized to most cell components. However, no R-to-K sites were found in proteins expressed in the endoplasmic reticulum (ER), and no Q- to-K sites were found in the extracellular space or cytoplasmic proteins (Fig. 4D).

[0188] pLogo analysis of the global human acetylome revealed an enrichment for lysine and an under-representation of C in the flanking sequence. As suggested above for the mouse acetylome these lysines are the adjacent acetylated lysines (Fig. 14C). Interestingly, there was an overrepresentation for G / A / D in position -1 suggesting that in human there is a G / A / D (-i)K enrichment. Further pLogo analyses of CC, showed that G / A / D in position -1 stemmed from the mitochondrial and cytosolic acetylated proteins, whereas nuclear protein showed only G at -1 (Fig. 14D-14F). Likewise, a pLogo analysis was performed using the R-to-K sites (Fig. 14G) and showed an enrichment for K downstream of the acetylated site, E on position -7 and SS / AK on positions -1 and -2. For Q-to-K sites, the pLogo analysis did not identify a specific logo besides a tendency for F on position -2 and a higher probability for A between -7 and +1 (Fig. 14H). Further searches for regulators of the enriched pathways in identified R-to-K proteins showed that the vast majority of regulators have a negative / positive role in tumorigenesis, such as STAT1 / 3, P53, BRCA1, PARP1, MYCN, and E2F1 / 4 (Fig. 141). Likewise, out of the seven identified regulators of Q-to-K proteins, TP53, STAT1, BRCA1, and SPI1 have a direct enhancing / preventing role in tumorigenesis, whereas manipulation of the others was also suggested to affect cancer (Fig. 14J). In addition, GO DisGeNET analyses of R-to-K and Q-to-K revealed several human ageing- associated diseases, including myocardial ischemia, inflammatory disorders, multiple types of cancer, ataxia telangiectasia, and Werner syndrome.

[0189] Remarkably, as seen in Figure 15, the initial conversions from R / Q-to-K occurred early in evolution and more sites accumulated as it progressed. Interestingly, as seen in Figures 4E and 4F, pathway enrichment analyses of the long-lived associated acetylome also highlighted various DNA repair pathways. These include non-homologous end joining (NHEJ), homology directed repair of DNA double-strand break, and DNA mismatch repair. Importantly, substantial evidence suggests DNA repair as a key mechanism of longevity, mostly via protection against the two major threats to long survival: cancer and cognitive decline. Thus, acquiring acetylation on DNA repair proteins during evolution can support healthy longevity. This suggests that accumulating new acetylation sites helps in addressing the gradual increase in body size associated with longevity.

[0190] Table 3: 495 human K to short living mammals (e.g., mouse) R sites. The number given after the gene name is the amino acid position of the base in the mouse protein.

[0191] Table 4: 150 human K to short living mammals (e.g., mouse) Q sites. The number given after the gene name is the amino acid position of the base in the mouse protein.Example 5: USP10 K714 Acetylation controls PCNA stability

[0192] Next, the inventors aimed to further elucidate the effect of acquired acetylation sites in longevity. To this end, the effect of K714 acetylation of a DNA repair pathway protein (Fig. 4B) ubiquitin specific peptidase 10 (USP10) (Mann Whitney FDR < 0.05; tree statistical correction FDR < 0.05) was examined (Fig. 5A). In short-lived mammals, there is a R instead of the K714 residue of human USP10. USP10 catalyzes a hydrolase activity, which removes conjugated ubiquitin from its target proteins, such as Proliferating Cell Nuclear Antigen (PCNA), thereby stabilizing them. Increased PCNA levels are associated with poor prognosis of various tumors. Indeed, a correlation test showed a significant positive correlation between PCNA and USP10 levels in lung, glioblastoma and breast cancers (P=0.001, 0.000001 and 0.002, respectively) (Fig. 5B), supporting the function of USP10 in stabilizing PCNA. Since long-lived animals tend to be physically larger, one of the major challenges facing such animals is the higher probability of cancer development with age. Thus, the acquisition of USP10 acetylation in long-lived animals might contribute to addressing this challenge. PCNA protein levels were examined in human colorectal carcinoma (HCT116) cells overexpressing either WT, K714Q or K714R human USP10. The translation inhibitor cycloheximide (CHX) was used to specifically examine the effect on protein stabilization. As seen in Figure 5C, in comparison to cells overexpressing the WT protein or the mutant K714R, overexpression of K714Q resulted in significantly higher PCNA protein levels with or without CHX treatment. Besides PCNA, USP10 has a spectrum of targets that might affect tumorigenesis as well. Thus, the inventors followed the role of acetylated USP10 on its global deubiquitylation activity. As seen in Figure 16, incomparison to WT or K714Q, the expression of constitutively deacetylated USP10 mutant K714R results in a significantly lower global ubiquitylation levels. Therefore, acetyl ation / deacetylati on controls USP10 activity and it would be of interest to identify additional USP10 targets that affect longevity. This result shows a potential role of the acquired acetylation on USP10 in reducing cancer incidence, and hence its role in promoting longevity. Altogether, fixation of the acetylation status or acquiring new acetylation sites during evolution enables long-lived animals to enhance pro-longevity mechanisms. Specifically, this study demonstrated that acetylations play a crucial role in promoting H2S production and DNA repair, two key factors that play a significant role in enhancing a healthy and extended lifespan.Example 6: Locations targetable in humans

[0193] Next, the inventors looked for locations where the human amino acid matched shortlived animals (both K), but was R in other long-lived animals. 123 locations were found where the mismatch between humans and other long-lived animals was statistically significant. These locations are provided in Table 5. Of these 123 locations 14 were in genes (12 total genes) related to fat metabolism and thus were of particular interest.

[0194] Table 5: 123 human K and mouse K sites that are R in other long-lived mammals. The number given after the gene name is the amino acid position of the K in the mouse protein. Bolded locations are related to fat metabolism.

[0195] Although the invention has been described in conjunction with specific embodiments thereof, it is evident that many alternatives, modifications and variations will be apparent to those skilled in the art. Accordingly, it is intended to embrace all such alternatives, modifications and variations that fall within the spirit and broad scope of the appended claims.

Claims

CLAIMS:

1. A method for identifying an amino acid residue associated with a biological trait, the method comprising: a. receiving a proteomic database for a plurality of species of mammals; b. receiving a database comprising a quantifiable measure of said biological trait for said plurality of species of mammals; c. selecting a first amino acid to test and an organism from said plurality of species of mammals; d. for sites of a residue of said first amino acid in a proteome of said organism, establishing a first group of species of mammals of said plurality that also have said first amino acid at said residue and a second group of species of mammals of said plurality that have a second amino acid at said residue, wherein said first and second amino acid are not the same amino acid; e. performing a statistical test to determine i. sites where said quantifiable measure of said biological trait in said first group is significantly higher than said quantifiable measure of said biological trait in said second group to produce a first list of sites of said first amino acid potentially associated with said biological trait, and ii. sites wherein said quantifiable measure of said biological trait in said second group is significantly higher than said quantifiable measure of said biological trait in said first group to produce a second list of sites of said second amino acid potentially associated with said biological trait; f. for sites of said first list and said second list determining if said potential association is due to phylogenetic distance; and g. selecting a site and amino acid residue from said first list or said second list whose association with said biological trait is not due to phylogenetic distance;thereby identifying an amino acid residue associated with a biological trait.

2. The method of claim 1, wherein said first amino acid is an amino acid that can be post- translationally modified and said second amino acid is an amino acid that mimics the modified or unmodified state of said first amino acid.

3. The method of claim 2, further comprising before step (d) receiving an atlas of marks of said post-translation modification on said first amino acid for said organism and identifying is said received atlas residues of said first amino acid that can be modified in said organism, wherein step (d) is performed for each identified modifiable residue.

4. The method of any one of claims 1 to 3, wherein said first amino acid is lysine and said second amino acid is arginine or glutamine5. The method of claim 4, further comprising before step (d) receiving an acetylome for said organism and identifying is said received acetylome lysine residues that can be acetylated in said organism, wherein step (d) is performed for each identified acetylatable residue.

6. The method of any one of claims 1 to 5, wherein said statistical test is performed on the average of said quantifiable measure in said first group and the average of said quantifiable measure in said second group.

7. The method of any one of claims 1 to 6, wherein determining if said potential association is due to phylogenetic distance comprises producing a site score for sites, performing a permutation test to produce a random distribution of site scores for said sites and wherein a statistically significant site score as compared to said random distribution indicates said potential association is not due to phylogenetic distance, wherein said producing a site score comprises: a. for every pair of species in said plurality of species determining the difference between said quantifiable measure in said pair of species and multiplying said difference by the phylogenetic distance between said pair species to produce a product; b. multiplying said product by a weight to produce a weighted product, wherein said weight is representative of a site’s contribution to the hypothesis that saidfirst amino acid or said second amino acid at said site is associated with said biological trait; and c. summing all said weighted products for every pair of species in said plurality of species to produce a site score; and wherein said permutation test produces a random distribution of site scores by assigning random weights to produce random site scores.

8. The method of claim 7, wherein said weight is the probability that regardless of amino acid at a site the pair of species would have a difference in the trait that is greater than average and have a phylogenetic distance that is smaller than average or have a difference in the trait that is smaller than average and have a phylogenetic distance that is greater than average.

9. The method of claim 7 or 8, wherein said weight is a. positive for pairs i. with the same amino acid at said position, a difference in the trait that is smaller than average and a phylogenetic distance that is greater than average; and ii. with different amino acids at said position, a difference in the trait that is larger than average and a phylogenetic distance that is smaller than average; and b. negative for pairs i. with the same amino acid at said position, a difference in the trait that is greater than average and a phylogenetic distance that is greater than average; ii. with the same amino acid at said position, a difference in the trait that is greater than average and a phylogenetic distance that is smaller than average;iii. with different amino acids at said position, a difference in the trait that is smaller than average and a phylogenetic distance that is greater than average; and iv. with different amino acids at said position, a difference in the trait that is smaller than average and a phylogenetic distance that is smaller than average.

10. The method of claim 9, wherein said weights are selected from the weights provided in Figure 7B.

11. The method of any one of claims 1 to 10, further comprising receiving a cell of said organism in culture, producing said selected second amino acid at said site in place of said first amino acid, measuring an output indicative of said trait and confirming said second amino acid at said site is associated with said trait.

12. The method of any one of claims 1 to 11, wherein said biological trait is selected from lifespan, blood pressure, brain size, rate of memory decline, resting heart rate, basal metabolic rate, insulin sensitivity, glucose tolerance, incidence of neurodegenerative disease, incidence of cancer, cancer latency, DNA repair efficiency, infection resistance, telomere length, telomere attrition rate, protein turnover rate, proteostasis capacity, reactive oxidation species (ROS) production, mitochondrial DNA copy number, mitochondrial membrane potential, epigenetic clock, age of sexual maturity, reproductive span, reproductive output, post-reproductive lifespan, extrinsic mortality, longevity quotient, and metabolite levels.

13. A method of extending the longevity of a human subject, the method comprising: a. converting a lysine to an arginine at a position selected from those provided in Table 5 in said subject; thereby extending the longevity of said human subject.

14. A method of extending the longevity of a human subject, the method comprising: a. detecting a mutation in said human subject wherein said mutation produces a lysine at a position selected from those provided in Table 1 or Table 2, anarginine at a position selected from those provided in Table 3 or a glutamine at a position selected from those provided in Table 4; and b. converting a lysine at a position provided in Table 1 to an arginine, a lysine at a position provided in Table 2 to a glutamine, an arginine at a position provided in Table 3 to a lysine or a glutamine at a position provided in Table 4 to a lysine in said subject; thereby extending the longevity of a non-human subject.

15. A method of extending the longevity of a subj ect, the method comprising, exogenously expressing a protein selected from those provided in Tables 1, 2, 3, 4 and 5 in said subject, thereby extending the longevity of a subject.

16. A method of extending the longevity of a non-human subject, the method comprising: a. receiving DNA from said non-human subject and determining said DNA encodes a protein that comprises a lysine at a position selected from those provided in Table 1 or Table 2, comprises an arginine at a position selected from those provided in Table 3 or comprises a glutamine at a position selected from those provided in Table 4; and b. converting a lysine at a position selected from Table 1 to an arginine, a position selected from Table 2 to a glutamine, or a position selected from Table 3 or 4 to lysine, in said non-human subject; thereby extending the longevity of said non-human subject.

17. The method of any one of claims 13 to 16, wherein said extending longevity comprises treating an aging related disease.

18. The method of claim 17, wherein an aging related disease is selected from a metabolic disease, a liver disease, a respiratory disease, a muscle disease, a proliferative disease and a neurological disease.

19. The method of claim 18, whereina. said metabolic or liver disease is selected from hypoglycemia, liver disease, hepatitis, hyperammonemia, hepatomegaly, diabetes, fatty liver disease, and hyperlipidemia; b. said muscle disease is myopathy; c. said respiratory disease is respiratory insufficiency; d. said proliferative disease is cancer; or e. said neurological disease is selected from polyneuropathy, and Alzheimer’s disease.

20. The method of any one of claims 13 to 19, wherein said converting comprises administering to said subject a genome editing compound targeted to said lysine residue and which edits said subject’s genome to have an arginine or glutamine at said location or targeted to said arginine or glutamine residue and which edits said subject’s genome to have a lysine residue at said location.

21. The method of claim 20, wherein said genome editing compound is a CRISPR-CAS complex comprising a guide RNA (gRNA) complementary to said subject’s genomic sequence comprising said lysine, arginine or glutamine to be converted.

22. The method of any one of claims 16 to 21, wherein said non-human animal is an animal with a short life span, optionally wherein a short life span is below 30 years on average.

23. The method of any one of claims 13 and 17 to 22, wherein said location in Table 5 is selected from: Cox4il amino acid 164, Amacr amino acid 134, Fasn amino acid 1516, Fasn amino acid 1920, Shmtl amino acid 319, Bdhl amino acid 280, Bdhl amino acid 212, Hadha amino acid 406, Acatl amino acid 260, Shmt2 amino acid 464, Acads amino acid 339, Etfdh amino acid 257, Cptlb amino acid 41, and Decrl amino acid 319.

24. A method for determining suitability of a human subject to have their longevity extended by a method of any one of claims 13-15, 17 to 21 and 23, the method comprising receiving DNA from said human subject and determining said DNA encodes a protein that comprises a lysine at a position provided in Table 1, 2 or 5, an arginine at a position provided in Table 3 or a glutamine at a position provided in Table4, wherein the presence of a lysine, arginine or glutamine at said position indicates said subject is suitable to have their longevity extended, thereby determining suitability.