Methods for identifying disease-risk missense mutations

The method uses Mahalanobis distance calculations to predict the impact of missense mutations on protein function, enhancing diagnostic accuracy for rare genetic diseases by distinguishing between pathogenic and benign mutations.

JP2026136508APending Publication Date: 2026-08-26TOHOKU MEDICAL & PHARM UNIV +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025022050
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2026-08-26

AI Technical Summary

Technical Problem

Existing methods struggle to accurately distinguish between pathogenic and benign missense mutations in proteins, leading to undiagnosed rare genetic diseases.

Method used

A method using Mahalanobis distance calculations based on mutation energy, standardized solvent-accessible surface area, and predicted local distance difference test scores to predict the effect of missense mutations on protein function.

Benefits of technology

Enhances the ability to identify disease-causing missense mutations, improving diagnostic accuracy for rare genetic diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026136508000001_ABST
    Figure 2026136508000001_ABST
Patent Text Reader

Abstract

This invention provides a method for identifying or predicting missense mutations that pose a risk of causing disease. [Solution] A method for predicting whether a missense mutation is pathogenic or benign includes: (1) creating a pathogenic mutation population and a benign mutation population for the target protein; (2) calculating the mutation energy (△△Gmut), standardized solvent-accessible surface area (nSASA), and predicted local distance difference test (pLDDT) score for each individual mutation constituting each population, and calculating the average values ​​of △△Gmut, nSASA, and pLDDT score for each population; (3) calculating △△Gmut, nSASA, and pLDDT score for the missense mutation to be predicted; and (4) calculating the Mahalanobis distance of the missense mutation to be predicted from the pathogenic mutation population and the benign mutation population, respectively.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method for predicting the effects of missense mutations on protein function, and more specifically, to a method for identifying or predicting missense mutations that may cause disease. [Background technology]

[0002] Rare diseases are epidemiologically estimated to affect approximately 5-8% of the total population, or about 350 million people worldwide. Rare diseases are typically severe, chronic, and cause physical disability, often manifesting in childhood or early adulthood, with approximately 80% having a genetic origin. Advances in genome sequencing technology have led to the accumulation of vast amounts of data on human genetic variation. However, because distinguishing between pathogenic and benign variants remains challenging, many rare genetic diseases remain undiagnosed.

[0003] One approach to addressing this problem is the development of algorithms to evaluate missense mutants based on protein structure. Known methods (software) for scoring the impact of missense mutants on protein function include PolyPhen-2, SIFT, and CADD. PolyPhen-2 predicts functional changes based on human amino acid sequence homology, protein domains, and, where possible, three-dimensional structure, determining that amino acid substitutions that strongly affect the protein's three-dimensional structure and function have a greater impact. SIFT predicts the impact on protein function based on the interspecies conservation of amino acid sequences and the similarity of substituted amino acids, determining that substitutions of amino acids that are evolutionarily conserved within the protein family have a greater impact. CADD integrates multiple annotations, including sequence conservation, epigenetic modification, and amino acid mutation information, to make its determination.

[0004] The inventors have recently developed and reported a novel prediction method called VarMeter (Non-Patent Literature 1: Aoki et al., Mol Genet Metab Rep 37: 101016, 2023). VarMeter utilizes the AlphaFold2 model and is based on two parameters: the standardized solvent-accessible surface area (nSASA) of amino acid residues in the wild-type protein, and the mutational energy, which reflects the difference in Gibbs free energy (△△G) of protein folding between the wild-type protein and the mutant protein. [Prior art documents] [Non-patent literature]

[0005] [Non-Patent Document 1] Aoki et al., Mol Genet Metab Rep 37: 101016, 2023. [Overview of the project] [Problems that the invention aims to solve]

[0006] The object of the present invention is to provide a novel method for predicting the effects of missense mutations on protein function when missense mutations are present. Another object of the present invention is to identify or predict disease-causing missense mutations by predicting the effects of missense mutations on protein function. [Means for solving the problem]

[0007] The inventors of this invention performed an analysis using missense mutations registered in the ClinVar database, a database provided by the National Center for Biotechnology Information (NCBI) that compiles information on genetic mutations (variants) and related diseases. They discovered that the effect of missense mutations on protein function can be predicted by calculating the Mahalanobis distance using the mutation energy, standardized solvent-accessible surface area (nSASA), and predictive local distance difference test (pLDDT) score in the missense mutants, thus completing the present invention. This invention includes the following: [1] A method for predicting whether a missense mutation in a target protein will impair the function of the protein, (1) Using mutation data for an arbitrary protein, create a pathogenic mutation population and a benign mutation population, respectively. For each mutation constituting the population, determine the mutation energy (△△Gmut) which indicates the effect of the mutation on stability, the standardized solvent-accessible surface area (nSASA) of the mutation site, and the predicted local distance difference test (pLDDT) score. Then, calculate the average values ​​of the mutation energy, standardized solvent-accessible surface area, and predicted local distance difference test score for each population, and set the respective average values ​​obtained from these calculations. (2) For the missense mutation to be predicted, a process to calculate the mutation energy, the standardized solvent-accessible surface area of ​​the mutation site, and the predicted local distance difference test score. (3) The following equation:

[0008]

number

[0009]

number

[0010]

number

[0011]

number

[0012] [9] A system for predicting whether a missense mutation in a target protein will impair the function of the protein, This module sets the average values ​​of population variables by creating pathogenic and benign variant populations using mutation data for an arbitrary protein, calculating the mutation energy (△△Gmut) which indicates the effect of the mutation on stability, the standardized solvent-accessible surface area (nSASA) of the mutation site, and the predicted local distance difference test (pLDDT) score for each individual mutation constituting the population, and then calculating the average values ​​of the mutation energy, standardized solvent-accessible surface area, and predicted local distance difference test score for each population. A target variable calculation module calculates the mutation energy, the standardized solvent-accessible surface area of ​​the mutation site, and the predicted local distance difference test score for the missense mutation to be predicted. The following formula:

[0013]

number

[0014]

number

[0015]

number

[0016]

number

[10] The system according to [9] further comprises a population variable mean calculation module for calculating the mean values ​​of mutation energy, standardized solvent-accessible surface area, and predicted local distance difference test scores in the pathogenic mutation population and the benign mutation population.

[11] The system according to [9] or

[10] above, further comprising a display module configured to visually display the above prediction.

[0017]

[12] A method performed in a computer device equipped with a processor for predicting whether a missense mutation in a target protein will impair the function of the protein, A population variable mean setting step involves creating pathogenicity and benign mutation populations using mutation data for an arbitrary protein, determining the mutation energy (△△Gmut) indicating the effect of the mutation on stability, the standardized solvent-accessible surface area (nSASA) of the mutation site, and the predicted local distance difference test (pLDDT) score for each individual mutation constituting the population, and then calculating the mean values ​​of the mutation energy, standardized solvent-accessible surface area, and predicted local distance difference test score for each population, and setting the respective mean values ​​obtained from these calculations. For the missense mutation to be predicted, the target variable calculation step involves calculating the mutation energy, the standardized solvent-accessible surface area of ​​the mutation site, and the predicted local distance difference test score. The following formula:

[0018]

number

[0019]

number

[0020]

number

[0021]

number

[13] The method according to

[12] , further comprising a step of calculating population variable means for the mutation energy, standardized solvent-accessible surface area, and predicted local distance difference test score in the pathogenic mutation population and the benign mutation population. [Effects of the Invention]

[0022] The method of the present invention makes it possible to predict the effect of missense mutations on the function of proteins that have missense mutations. [Brief explanation of the drawing]

[0023] [Figure 1] This is the result of classifying the pLDDT scores of 296 pathogenic variants and 240 benign variants into four confidence levels. [Figure 2] This is the result of analyzing the correlation between pathogenicity and allele frequency in 296 pathogenic variants and 240 benign variants. [Figure 3] This report presents the results of analyses of the correlation between pLDDT score and nSASA, the correlation between pLDDT score and mutational energy, and the correlation between mutational energy and nSASA for both 296 pathogenic variants and 240 benign variants. [Figure 4] This is a conceptual diagram of variant classification using Mahalanobis distance. Black circles represent individual pathogenic variants, and white circles represent individual benign variants. [Figure 5] This shows the predictive accuracy of VarMeter, VarMeter2, AlphaMissense, and CADD methods based on allele frequencies in the ClinVar dataset. (A) Predictive accuracy for pathogenic variants (n=296). (B) Predictive accuracy for benign variants (n=240). (C) Overall predictive accuracy for all variants (pathogenic and benign, n=536). [Figure 6]This shows a structural comparison and variant mapping of human SGSH. (A) The crystal structure of human SGSH dimers analyzed at 2 Å resolution and a 3D AlphaFold model of monomeric human SGSH are shown. For comparison, the AlphaFold model is superimposed on the A chain of the crystal structure. (B) The structure of the wild-type monomeric SGSH is shown with missense variants mapped onto the AlphaFold model. The locations of pathogenic and benign variants are shown. The active site residues of SGSH (D31, D32, C70, and D273) and N274 are shown as meshes. (C) A magnified view of the SGSH active site from the AlphaFold model (corresponding to the dotted line frame in B). The structure is depicted in cartoon form, and the figure was created using PyMOL software. [Figure 7] These are the results of predictions made using each prediction method for pathogenic (n=24) and benign (n=8) SGSH variants. The classification is as follows: P (pathogenic), PD (potentially damaging), B (benign), A (unknown). b is located at the dimer interface. [Figure 8] The results of the accuracy rate of predictions using each prediction method are shown. [Figure 9] The results of measuring the expression and enzymatic activity of wild-type and Q365P SGSH proteins are shown. The left figure shows the results of measuring protein expression by Western blotting. The right figure shows the results of measuring enzymatic activity. [Figure 10] This is the result of analyzing the intracellular localization of SGSH in HEK293-WT cells and HEK293-Q365P cells. [Figure 11] These are the results of a cycloheximide (CHX) tracking assay for wild-type and Q365P SGSH proteins. [Figure 12] This shows the results of measuring protease resistance of wild-type and Q365P SGSH proteins using a trypsin-sensitivity assay. [Modes for carrying out the invention]

[0024] The present invention will be described below using exemplary embodiments as examples, but the present invention or its application or use is not limited to the embodiments described below. Unless otherwise specified herein, all technical and scientific terms used herein have the same meaning as generally understood by those skilled in the art to which the present invention pertains. Furthermore, where there is a reference to the substitutability of any means and methods described herein, any means or methods equivalent or similar to those described herein may be used in the practice of the present invention. In addition, all publications and patents cited herein in connection with the present invention are cited herein and constitute part of this specification, for example, to indicate means, methods, and other things that can be used in the present invention.

[0025] In this specification, the terms “variant,” “genetic variant,” or “variant(s)” refer to alternative forms of genes, genome sequences, or parts thereof. In this specification, variants are referred to at the protein level and are identified as changes in amino acids in a protein sequence.

[0026] The present invention is a method for predicting the effect of a missense mutation on the function of a protein having a missense mutation, characterized by using the Gibbs energy (also called Gibbs free energy), the normalized solvent-accessible surface area, the predicted local distance difference test (pLDDT) score, and the Mahalanobis distance from the mean values ​​of these variables, respectively. More specifically, first, for each of the pathogenic mutation group (Group-pathogenic) in which the missense mutation affects the function of the protein and the benign mutation group (Group-benign) in which the missense mutation does not affect the function of the protein, the mean values ​​of the mutation energy (△△Gmut) indicating the effect of the mutation on stability, the normalized solvent-accessible surface area (normalized SASA:nSASA) of the mutation site, and the predicted local distance difference test (pLDDT) score (for the pathogenic mutation group (x - p ,y - p ,z - p), benign variant population (x - B ,y - B ,z - B )) and the mutation energy of the protein with the missense mutation to be predicted, the standardized solvent-accessible surface area, and the predicted local distance difference test score (x i ,y i ,z i ) and then calculate the Mahalanobis distance from the mean value of each population (Mahalanobis distance from the pathogenic variant population (D p 2 ) and the Mahalanobis distance (D) from the benign variant population. B 2 This method calculates the Mahalanobis distance and predicts whether the missense mutation is pathogenic, suspected pathogenic, or benign based on the relationship between each Mahalanobis distance.

[0027] Protein stability can be considered as one of the factors affecting protein function. Protein stability can be considered in terms of the thermodynamic stability of a protein that is unfolded and then rapidly and reversibly refolded. In other words, protein stability is the difference in Gibbs energy between the folded state and the unfolded state, and can be expressed by the following equation. △△Gfolding=△Gfolded-△Gunfolded Here, △△Gfolding corresponds to the difference in total Gibbs energy (△Gtotal) between the folded and unfolded states of the protein (△Gfolded - △Gunfolded). The greater the difference in Gibbs energy between the folded and unfolded states, the more stable the protein is against denaturation.

[0028] The inventors have found that, in determining the difference in Gibbs energy between the folded and unfolded states of a target protein, the total Gibbs energy (ΔGtotal) is best calculated as the sum of the values ​​obtained by multiplying each energy term by a coefficient, using the following formula. △Gtotal=a×△GvdW+b×△Gelec(T)+c×△Gentr+d×△Gnp Here, △Gtotal is the total Gibbs energy, △GvdW is the van der Waals interaction, △Gelec is the electrostatic interaction, △Gentr is the entropy contribution related to the mobility of the side chains, and △Gnp is the non-polar, surface-dependent solvation energy, with coefficients a=b=0.5, c=0.8, and d=0. Therefore, the above equation can be expressed as follows. △Gtotal=0.5×△GvdW+0.5×△Gelec(T)+0.8×△Gentr

[0029] The contributions of entropy related to van der Waals interactions (△GvdW) and side chain mobility (△Gentr) can be determined using tools included in various molecular dynamics simulations. Examples of molecular dynamics simulations, though not limited to these, include CHARMM, ABINIT, AMBER, Ascalaph, CASTEP, CPMD, DL_POLY, FIREBALL, GROMACS, GROMOS, LAMMPS, MDynaMix, MOLDY, MOSCITO, NAMD, Newton-X, ProtoMol, PWscf, SIESTA, VASP, TINKER, YASARA, ORAC, and XMD.

[0030] The electrostatic interaction (ΔGelec) can be determined using any tool capable of doing so, but it is preferable to use the Generalized Born implicit solvent model. Tools for determining this model, including electrostatic interactions, are included in various molecular dynamics simulations and can be used to determine the interaction. For example, but not limited to, Discovery Studio can be used.

[0031] In the prediction method of the present invention, the mutation energy (△△Gmut), which indicates the effect of mutations on protein stability, can be determined by the following formula. △△Gmut = △△Gfolding (mutant) - △△Gfolding (wild type) Here, △△Gfolding(wild-type) is the difference in Gibbs energy between the folded and unfolded states of the wild-type protein, △△Gfolding(mutant) is the difference in Gibbs energy between the folded and unfolded states of the mutant protein, and △△Gmut corresponds to the difference between them. △△Gfolding can be calculated using the formula (△Gtotal) described above. Mutation energy can be determined using the above formula and known molecular modeling simulations, protein stability and interaction prediction tools, and other tools, such as Discovery Studio, FoldX (https: / / foldxsuite.crg.eu / ), and mCSM (https: / / biosig.lab.uq.edu.au / mcsm / ).

[0032] As a factor affecting protein function, the inventors focused on the solvent-accessible surface area (SASA) of the protein mutation site and calculated the normalized solvent-accessible surface area (nSASA) of the protein mutation site using the following formula, which was then used in the prediction method of the present invention. nSASA = (SASA in folded state) / (SASA in mutated state) Here, the solvent-accessible surface area (SASA) was calculated assuming a water molecule radius of 1.4 Å. nSASA can be determined as follows. For example, suppose the Xth amino acid residue of protein A has mutated from serine to tyrosine. The solvent-accessible surface area (SASA) in the folded state is the degree of solvent exposure of that serine residue. Various data are available for obtaining the three-dimensional structural data of the folded state used to determine solvent exposure, and can be used to determine this. While not limited to these, for example, the three-dimensional structure registered in the RCSB's Protein Data Bank, the AlphaFold2 model, or homology models can be used to determine the solvent exposure of the folded state. The calculation of solvent exposure itself can be performed using any software (e.g., Discovery Studio or PyMOL). On the other hand, the SASA of the denatured state can be determined based on the amino acid residue at the mutation site. Here, since the mutation is from serine to tyrosine, the SASA of the denatured state is the SASA of serine. Since the value is almost fixed for each amino acid residue, values ​​reported in the literature can be used. While not limited to these, relevant literature includes, for example, Oobatake M, Ooi T., Hydration and heat stability effects on protein unfolding, Prog Biophys Mol Bio. 1993:59(3): 237-84, and Miller S, Janin J, Lesk AM, Chothia C. Interior and surface of monomeric proteins. J Mol Biol. 1987 Aug 5;196(3):641-56. The standardized solvent-accessible surface area (nSASA) can be determined using the above formula and known molecular modeling simulations, including but not limited to Discovery Studio or PyMOL. Alternatively, publicly available relative solvent exposure data (e.g., from a database such as jMJorp: https: / / jmorp.megabank.tohoku.ac.jp / ) can be used.

[0033] In the prediction method of the present invention, the predicted local distance difference test (pLDDT) score is a score that indicates the prediction accuracy of each amino acid residue in protein structure prediction. It is calculated by the structure prediction algorithm and used to evaluate the reliability of the protein structure. The pLDDT score is expressed in the range of 0 to 100, is calculated individually for each amino acid residue, and indicates the local reliability within the predicted structural model. A score of 90-100 indicates high reliability, 70-90 indicates moderate reliability, 50-70 indicates low reliability, and <50 indicates very low reliability. For example, a score of 70-90 suggests that secondary structures such as α-helices and β-sheets are accurately predicted. The pLDDT score can be used as an indicator to understand which parts of the predicted protein structure are reliable and which parts have room for improvement. For example, in AlphaFold, the loss function is set to reflect the difference between the predicted LDDT value and the true LDDT value for each residue in the predicted structural model, based on the difference in Cα interatomic distance between the "correct" structure and the predicted structural model. This predicted LDDT value is the residue-level prediction confidence score, the pLDDT score. In the prediction method of the present invention, any known database, but not limited to, can be used, for example, the predicted local distance difference test (pLDDT) score registered in the AlphaFold protein structure database (https: / / alphafold.ebi.ac.uk / ). Alternatively, it may be calculated using known molecular modeling simulations or protein structure analysis tools.

[0034] The method for predicting whether a missense mutation in the target protein of the present invention impairs the function of the protein includes the following steps. Step (1) involves creating pathogenic mutation populations and benign mutation populations using mutation data for an arbitrary protein, calculating the mutation energy (△△Gmut) which indicates the effect of the mutation on stability, the standardized solvent-accessible surface area (nSASA) of the mutation site, and the predicted local distance difference test (pLDDT) score for each individual mutation constituting the population, and setting the respective average values ​​obtained by calculating the average values ​​of the mutation energy, standardized solvent-accessible surface area, and predicted local distance difference test score for each population.

[0035] Step (1), which involves "creating a population of pathogenic and benign variants using mutation data for an arbitrary protein" (hereinafter referred to as Step (A)), extracts mutation data for an arbitrary protein from a database and creates a population of pathogenic and benign variants. While not limited to these, for example, a publicly known database, a privately owned database, or a combination thereof may be used to create a population of pathogenic and benign variants for an arbitrary protein. Examples of publicly known databases, though not limited to these, include UniProt, GlyGen, ClinVar, BioMuta, jMorp, dbSNP, dbVar, gnomAD, and TogoVar.

[0036] The types of proteins with mutation data may be the same or different, and may overlap, for each of the pathogenic and benign mutation populations. There may be one or more types of proteins with mutation data in each of the pathogenic and benign mutation populations. For example, there may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, 100 or more types of proteins. However, for example, 1 to 100 types, preferably 5 to 50 types, and more preferably 10 to 50 types of proteins may be selected. There are no particular restrictions on the individual mutations that make up a population, but preferably, each population comprises at least 5 different mutations, preferably at least 10, and more preferably at least 20 different mutations.

[0037] Step (1), which describes "determining the mutation energy (△△Gmut) indicating the effect of the mutation on stability, the standardized solvent-accessible surface area (nSASA) of the mutation site, and the predicted local distance difference test (pLDDT) score for each individual mutation constituting each population, and calculating the mean values ​​of the mutation energy, standardized solvent-accessible surface area, and predicted local distance difference test score for each population" (hereinafter referred to as step (B) in this specification), can be performed by determining the mutation energy, standardized solvent-accessible surface area, and predicted local distance difference test score (hereinafter sometimes referred to as "3 variables") for each individual mutation using the method described above, and calculating the mean values ​​of the mutation energy, standardized solvent-accessible surface area, and predicted local distance difference test score (hereinafter sometimes referred to as "mean values ​​of 3 variables") for each population. Preferably, the mean values ​​of 3 variables for each population can be calculated using the mutation energy, standardized solvent-accessible surface area, and predicted local distance difference test score described for each mutation in any known database, but not limited to, the ClinVar database.

[0038] In the prediction method of the present invention, steps (A) and (B) may be performed in advance, and in step (1), the average value of the three variables calculated from the previously performed steps (A) and (B) may be set in step (1). Alternatively, steps (A) and (B) may be performed in step (1).

[0039] In this specification, the mean value of each variable is (x - ,y - ,z - ) and or

[0040]

number

[0041]

number

[0042] Step (3) is a step of determining the mutation energy, the standardized solvent-accessible surface area, and the predicted local distance difference test score for the predicted missense mutation. In this specification, each variable for the predicted missense mutation is (x i ,y i ,z i ) and, or,

[0043]

number

[0044] Step (4) is calculated using the following formula to determine the Mahalanobis distance (Di) from each population of the missense variant to be predicted. 2 This is the process of calculating ).

[0045]

number

[0046]

number

[0047]

number

[0048]

number

[0049]

number

[0050]

number

[0051] Step (5) is the Mahalanobis distance (D) of the predicted missense mutation from the pathogenicity mutation population. p 2 ) and the Mahalanobis distance (D) from the benign variant population of the predicted missense mutation. B 2) and comparing to predict whether the missense mutation is a pathogenic mutation, a benign mutation, or a suspected pathogenic mutation. D P 2 <D B 2 In the case of, it is predicted to be a pathogenic mutation, D P 2 >D B 2 In the case of, it is predicted to be a benign mutation.

[0052] The various calculations performed in the method of the present invention can be performed using various calculation softwares or simulation softwares capable of performing molecular dynamics calculations. Without being limited thereto, for example, CHARMM, ABINIT, AMBER, Ascalaph, CASTEP, CPMD, DL_POLY, FIREBALL, GROMACS, GROMOS, LAMMPS, MDynaMix, MOLDY, MOSCITO, NAMD, Newton-X, ProtoMol, PWscf, SIESTA, VASP, TINKER, YASARA, ORAC, XMD, etc. can be mentioned. Also, it can be performed using modeling, calculation, and simulation packages such as Discovery Studio equipped with those calculation modules.

[0053] In the prediction method of the present invention, the function of a protein is not particularly limited as long as it is a function possessed by the target protein. Damage to the function of a protein can include, for example, changes in stability, changes in protease resistance, changes in biological activity including enzyme activity, changes in ligand and substrate binding ability, changes in binding ability to other substances (e.g., nucleic acids and inhibitors), changes in signal transduction activity, changes in localization (intracellular and extracellular), aggregation, changes in solubility, changes in three-dimensional structure and oligomeric state, changes in post-translational modification, and carcinogenicity. The method of the present invention determines whether at least one or more of these protein functions are affected. Preferably, the function of the target protein to which the method of the present invention is used is a function that is involved in physiological activity in vivo, for example, a function that has been shown to be associated with or is suspected to be associated with a specific disease. Therefore, the prediction method of the present invention is also a prediction method that associates the effects of missense mutations in target proteins with missense mutations with specific diseases.

[0054] The prediction method of the present invention can be used in combination with other prediction tools that predict the effects of missense mutations. Other prediction tools, though not limited to these, include, for example, SIFT (Ng PC, Henikoff S: SIFT: Predicting amino acid changes that affect protein function. Nucleic Acids Res 31: 3812-3814, 2003.), PolyPhen, PolyPhen-2 (Adzhubei IA, et al.: A method and server for predicting damaging missense mutations. Nat Methods 7: 248-249, 2010.), PROVEAN (Choi Y, et al.: Predicting the functional effect of amino acid substitutions and indels. PLoS One 7: e46688, 2012.), MutationTaster (Schwarz JM, et al.: Mutationtaster2: Mutation prediction for the deep-sequencing age. Nat Methods 11: 361-362, 2014.), and CADD (Rentzsch P, et al.: CADD: Predicting the deleteriousness of variants throughout the human genome. Nucleic Acids Res 47: D886-D894, 2019.), SAAPpred(Al-Numair NS, Martin AC. The SAAP pipeline and database: tools to analyze the impact and predict the pathogenicity of mutations. BMC Genomics 14: Suppl 3:S4, 2013.), AlphaMisssense (Cheng J, et al.Examples include: Accurate proteome-wide missense variant effect prediction with AlphaMissense (Science 381: eadg7492, 2023.), and VarMeter (Aoki E Manabe N, Ohno S, et al.: Predicting the pathogenicity of missense variants based on protein instability to support diagnosis of patients with novel variants of ARSL (Mol Genet Metab Rep 37: 101016, 2023.)). The mutation effect prediction is calculated from the mutation score obtained by applying one or more algorithms selected from this group to the amino acid sequence of the target protein. Preferably, SIFT, PolyPhen-2, or CADD are used. By using the prediction method of the present invention in combination with other prediction tools, the impact of missense mutations on proteins can be predicted more accurately.

[0055] The prediction method of the present invention allows for the association of a specific disease with a missense mutation based on the in vivo function of the protein, if the missense mutation in a protein is predicted to impair the protein's function, i.e., if it is predicted to be a pathogenic mutation. In addition to steps (1) to (5) above, the prediction method of the present invention may further include a step of associating the missense mutation with a specific disease. A prediction method for associating with a disease is also included in the present invention.

[0056] Specifically, if a missense mutation is identified in protein A, but the relationship between that missense mutation and the disease is unclear, the method of the present invention can predict the effect of the missense mutation, thereby linking the missense mutation to the disease. Alternatively, if multiple missense mutations are identified in protein A, the method of the present invention can predict which missense mutation is associated with the disease.

[0057] The protein targeted by the prediction method of the present invention is not particularly limited, and it is sufficient if the amino acid sequence of the protein and the missense mutations contained therein are identified. Information on proteins containing missense mutations is registered in various public databases, and is not limited to these, but examples include UniProt, GlyGen, ClinVar, BioMuta, and jMorp. From these databases, amino acid sequence information of proteins containing wild-type and missense mutations can be obtained and used in the prediction method of the present invention. For example, mutation information of the target protein, information on whether it is a pathogenic or benign mutation, disease information, and other information can be obtained from these databases and classified into a population of pathogenic mutations and a population of benign mutations in step (1) of the present invention. Furthermore, it becomes possible to associate the target missense mutation with disease. Therefore, in one embodiment, the method of the present invention is also a method for determining the clinical significance of a mutant.

[0058] Among naturally occurring proteins, there are those called intrinsically unfolded proteins (or intrinsically unstructured proteins), which do not form a specific three-dimensional structure under natural conditions or physiological conditions (not denaturing conditions), as well as proteins that have a partially intrinsically disordered region (called an intrinsically disordered region). The proteins to which the method of the present invention is applied are not particularly limited, but preferably proteins that form a specific three-dimensional structure in at least a part of their structure.

[0059] The proteins targeted by the method of the present invention are all proteins in the body, preferably all proteins that have physiological functions in the body, and more preferably proteins whose impairment of physiological function manifests as a disease. Examples include proteins associated with hereditary diseases and proteins associated with acquired gene mutations such as cancer. The proteins and diseases associated with them are not particularly limited, and any protein and disease can be mentioned. For example, arylsulfatase E as a protein and X-linked recessive chondrodysplasia punctate (CDPX1) as an associated disease; N-sulfoglucosamine sulfohydrolase as a protein and Sanfilippo syndrome or mucopolysaccharidosis type III as associated diseases; and proteins involved in the synthesis and modification of glycoproteins such as GALNT2 (polypeptide N-acetylgalactosaminyltransferase 2), MGAT2 (alpha-1,6-mannosyl-glycoprotein 2-beta-N-acetylglucosaminyltransferase), and ALG9 (α-1,2-mannosetransferase) as proteins and various types of congenital glucosylation disorders (CDG) as associated diseases. Other examples include, for instance, the proteins and associated diseases described in Paulina Sosicka, et. al., 5.19 - Congenital Disorders of Glycosylation, Editor(s): Joseph J. Barchi, Comprehensive Glycoscience (Second Edition), Elsevier, 2021, Pages 294-334, ISBN 9780128222447; https: / / www.sciencedirect.com / science / article / pii / B9780128194751000134), and these references are incorporated herein by reference.

[0060] The methods described herein can be performed on a computer system or network. In one embodiment, the prediction method of the present invention is a computer-based method for analyzing the clinical significance of a variant. In one embodiment, the present invention also includes a product comprising a non-transitory computer-readable medium containing computer-readable instructions that, when executed by a computer, cause a computer to carry out the method disclosed herein. In one embodiment, the present invention also includes a product comprising a non-transitory computer-readable storage medium for physically storing instructions for carrying out the method disclosed herein, which, when executed, causes one or more computers in a computer network to receive a request to display a report and display a report that displays a visual representation of the clinical significance of the variant.

[0061] In one embodiment, a preferred computer system may optionally include a computer-readable medium for storing computer code for execution by the processor, including at least a processor and memory. Once the code is executed, the computer system performs the method described herein. [Examples]

[0062] The present invention will be specifically described below with reference to examples, but the present invention is not limited to the following examples.

[0063] (Example 1) Comparison of pathogenic and benign mutants from the ClinVar dataset As the target protein, N-sulphoglucosamine sulphonohydrolase (SGSH) was selected. Missense variants were extracted from the ClinVar database as of July 6, 2023, and a dataset consisting of 536 missense variants (296 pathogenic variants consisting of 24 types of proteins and 240 benign variants consisting of 23 types of proteins) was obtained. The amino acid sequences of each protein were retrieved from the UniProt database.

[0064] The 3D model of the reference (wild-type) protein was obtained from the AlphaFold2 protein structure database. The 3D models of each missense variant were created using the "Calculate Mutation Energy / Stability" module of Discovery Studio 2021 (BIOVIA, Dassault Systemes, USA).

[0065] First, the correlation between the pathogenicity and the predicted local distance difference test (pLDDT) score, which reflects the per-residue structure accuracy provided by the AlphaFold structure database, was examined. The pLDDT scores of 296 pathogenic variants and 240 benign variants were classified into four confidence levels: very low (pLDDT score < 50), low (50 < pLDDT score < 70), high (70 < pLDDT score < 90), very high (pLDDT score > 90). Most of the pathogenic variants had a high pLDDT score (pLDDT score > 70), while a significant proportion of the benign variants had a low pLDDT score (pLDDT score < 50), and only a small fraction had a high pLDDT score (pLDDT score > 70). The results are shown in Figure 1.

[0066] Subsequently, the correlation between pathogenicity and allele frequency was analyzed for 296 pathogenic variants and 240 benign variants. The results are shown in Figure 2. Most pathogenic variants tended to have lower allele frequencies compared to benign variants. This result was consistent with previous research indicating a strong association between low allele frequency and pathogenicity. However, there was an overlap in allele frequencies between pathogenic and benign variants, ranging from 0.000001 to 0.0001. While allele frequency can be a useful indicator for distinguishing between pathogenic and benign variants, this overlap requires careful interpretation.

[0067] Essentially disordered regions are often associated with low pLDDT scores, suggesting that mutations in pathogenic variants are primarily located in well-structured (folded) regions of the protein. In contrast, in benign variants, mutations tend to occur in unstructured regions, with only a small amount present in folded regions. This trend highlights the importance of considering protein folding state when assessing the potential pathogenicity of missense variants.

[0068] Next, mutational energy, nSASA, and pLDDT scores were calculated for 240 benign mutants and 296 pathogenic mutants. The pLDDT score (mean ± SD) for benign mutants was 53.1 ± 27.9, while that for pathogenic mutants was significantly higher at 88.2 ± 14.2. The mean nSASA for benign mutants was 0.64 ± 0.29, compared to 0.21 ± 0.28 for pathogenic mutants. Mutational energy was 0.1 ± 2.0 kcal / mol for benign mutants, but significantly higher at 3.9 ± 9.6 kcal / mol for pathogenic mutants. These results are shown in Table 1.

[0069] [Table 1]

[0070] Furthermore, the correlation between pLDDT score and nSASA is shown in Figure 3 (left panel). A negative correlation was observed between pLDDT score and nSASA for both benign and pathogenic mutants. Mutants with low pLDDT scores (20-40) tended to have high nSASA values ​​(0.6-1.0), indicating that the mutant residues were generally exposed to the solvent, while mutants with high pLDDT scores (80-100) tended to have low nSASA values ​​(0-0.4). This reflects that the mutations are located in more structured and buried regions. This is consistent with the known behavior of endogenous disorder regions, which often exhibit low pLDDT scores and are typically exposed to the solvent.

[0071] The correlation between mutation energy and pLDDT score was analyzed for both benign and pathogenic mutants. The results are shown in Figure 3 (center panel). This trend confirmed that pathogenic mutants generally exhibit greater protein structural instability, which is reflected in higher mutation energy. However, some pathogenic mutants showed low mutation energy (Figure 3, center bottom panel). This suggests that mutations in these mutants occur in functionally important sites, such as residues involved in ligand or protein binding, and that even slight structural changes can impair function.

[0072] The correlation between mutational energy and nSASA was analyzed for both benign and pathogenic mutants. The results are shown in Figure 3 (right panel). A negative correlation was observed between nSASA and mutational energy. In benign mutants, high nSASA values ​​were typically associated with low mutational energy (Figure 3, upper right panel). Conversely, in pathogenic mutants, low nSASA values ​​tended to be associated with relatively high mutational energy (Figure 3, lower right panel). This suggests that pathogenic mutations are more likely to occur in structurally important buried regions, and even small changes can have a significant destabilizing effect.

[0073] (Example 2) Classification of missense mutants based on Mahalanobis distance The Mahalanobis distance, a statistical method used to determine group membership for various purposes, was applied to classify pathogenic and benign variants using the mutation energy, nSASA, and pLDDT scores. First, the mean values, variances, and covariances of the three variables were calculated for pathogenic and benign variants. The results are shown in Table 1 above.

[0074] A conceptual diagram of mutation classification using the Mahalanobis distance is shown in Figure 4. The mean values of the three variables (mutation energy, nSASA, and pLDDT score) were calculated for each of the pathogenic group (black circles) and the benign group (white circles). Let (x - p ,y - p ,z - p ) and (x - B ,y - B ,z - B ) be denoted. The Mahalanobis distances (D p 2 and D B 2 ) were calculated for each data point (x i ,y i ,z i ) using the inverse covariance matrix. A mutation is classified as pathogenic if D P 2 <D B 2 , and as benign if D P 2 >D B 2 .

[0075] The Mahalanobis distance (D​​​​​​​​​​

[0076]

number

[0077] [Table 2]

[0078] These Mahalanobis distances (D P and D B The inventors used this to improve a VarMeter prediction method they had already developed. Specifically, they integrated mutation energy, nSASA, and pLDDT score to create a new prediction method, VarMeter2.

[0079] To evaluate the performance of VarMeter2, the method of the present invention (VarMeter2), defined under condition (ii) below, was applied to the ClinVar dataset. Simultaneously, it was compared with the original VarMeter (Non-Patent Literature 1), the AlphaMissense method (Cheng J, Novati G, Pan J, Bycroft C, Zemgulyte A, et al. (2023) Accurate proteome-wide missense variant effect prediction with AlphaMissense. Science 381: eadg7492.), and the CADD method (Rentzsch P, Witten D, Cooper GM, Shendure J, Kircher M (2019) CADD: predicting the deleteriousness of variants throughout the human genome. Nucleic Acids Res 47: D886-D894.), defined under conditions (i), (iii), and (iv), respectively.

[0080] Condition (i), original VarMeter: Requirement #1: nSASA ≤ 0.11 Requirement #2: Mutation energy ≥ 0.88 kcal / mol Harmful (pathogenic) mutation (D): When both requirements #1 and #2 are met. Probably harmful (pathogenic) (PD): If either requirement #1 or #2 is met. Benign mutation (B): If neither requirement #1 nor #2 is met. Condition (ii), VarMeter2: Pathogenicity variant (P): D P 2 <D B 2 Benign mutation (B): D P 2 >D B 2 Condition (iii), AlphaMissense: Pathogenicity variant (P): 0.564 <AMスコア<1 Ambiguous mutation (A): 0.340 <AMスコア<0.564 Benign mutation (B): 0 <AMスコア<0.340 Condition (iv), CADD: Pathogenicity variant (P): CADD score ≥ 25.3 Ambiguous variation (A): 22.7 <CADDスコア<25.3 Benign mutation (B): CADD score ≤ 22.7

[0081] Table 3 summarizes the prediction accuracy of each method. VarMeter2 showed a prediction accuracy of 85% for pathogenic mutations and 78% for benign mutations, with an overall accuracy of 82%. This represents a significant improvement compared to the original VarMeter (74%). On the other hand, AlphaMissense, a deep learning method based on AlphaFold, achieved an overall accuracy of 91% on the ClinVar dataset. Similarly, CADD also showed an accuracy of 91%.

[0082] [Table 3]

[0083] To evaluate whether prediction accuracy depends on allele frequency, prediction accuracy was plotted against allele frequency. The results are shown in Figure 5. The results showed that the prediction accuracy of VarMeter2 was not significantly affected by allele frequency. This is likely because VarMeter2 is based solely on 3D structural parameters and does not incorporate allele frequency into its calculations. On the other hand, CADD showed higher prediction accuracy for mutations with low allele frequency. This is likely because CADD incorporates annotations that reflect population-level indicators as part of its scoring framework. Despite these advances, challenges remain in prediction based on 3D structure. Specifically, these include refining the mutant protein model, more accurate mutation energy calculations (△△G), and improving predictive power through optimization of structural parameters. Further development in these areas will be necessary to improve the accuracy of pathogenicity prediction.

[0084] (Example 3) Evaluation of VarMeter2 using SGSH mutant data SGSH is an enzyme that catalyzes the conversion of N-sulfo-D-glucosamine to D-glucosamine during the degradation of heparan sulfate. VarMeter2 was further validated using published data on SGSH variants. Numerous pathogenic variants of this enzyme associated with Sanfilippo syndrome A have been reported. The locations of these variants were mapped onto a 3D structural model of SGSH (Figure 6).

[0085] In Figure 6A, the root-mean-squared deviation (RMSD) between Cα atoms in the crystal structure and the corresponding AlphaFold model is 0.2 Å, indicating that the AlphaFold model of SGSH has sufficient accuracy for calculating mutation energy and standardized solvent-accessible surface area (nSASA). Next, the locations of pathogenic variants were mapped onto the 3D structural model of SGSH.

[0086] First, mutation energy, nSASA, and pLDDT scores were calculated for pathogenic (n=24) and benign (n=8) SGSH variants (Figure 7). Next, SGSH variants were classified as pathogenic or benign using VarMeter2. For comparison, predictions were also made using the original VarMeter method, AlphaMissense method, and CADD method. The accuracy rates of these prediction methods are shown in Figure 8. The overall accuracy of VarMeter2 (84%) in predicting SGSH variants was higher than both the original VarMeter (78%) and CADD (79%), but lower than AlphaMissense (89%). Furthermore, the accuracy of VarMeter2 in predicting pathogenic variants was 96%. This evaluation indicates that VarMeter2 can accurately classify SGSH variants and is equivalent to or an improved version of previous prediction methods.

[0087] (Example 4) Evaluation of novel SGSH variants To validate the capabilities of VarMeter2, it was applied to predict SGSH variants recorded in a database of 900 patients with rare or undiagnosed diseases, constructed by the inventors. As a result, a new Q365P variant was identified among the missense variants in the database, and its pathogenicity was predicted by VarMeter2 with the following parameters:D P 2 (0.7), D B 2 (14.8), nSASA (0.03), mutant energy (7.2 kcal / mol), pLDDT score (98.89). Furthermore, the AlphaMissense score for this mutation was 0.90, similarly predicting that this variant is pathogenic. In addition, based on the clinical symptoms observed in patients, this variant was suspected to be pathogenic. Therefore, to determine whether this newly identified Q365P mutation is directly involved in the observed symptoms and to support the diagnosis, the activity and expression of the Q365P SGSH protein were evaluated.

[0088] Wild-type and Q365P SGSH proteins were recombinantly expressed, purified, and subjected to enzyme assays using the artificial substrate 4-MU-GlcNS. Specifically, HEK293-WT cells and HEK293-Q365P cells were established that stably expressed wild-type and Q365P SGSH-FLAG fusion proteins, respectively. Expression vectors containing genes encoding wild-type or Q365P-mutated SGSH were created using the Gateway Cloning System (Thermo Fisher Scientific, USA). HEK293 cells were transfected with 30 μg of expression vector (wild-type or Q365P SGSH-FLAG fusion protein) using lipofectamine. Transformants: HEK293-WT cells and HEK293-Q365P cells were selected, and stably expressed wild-type and Q365P SGSH-FLAG fusion proteins, respectively.

[0089] Exogenous SGSH mRNA expression was approximately 300-fold higher in HEK293-WT cells and approximately 150-fold higher in HEK293-Q365P cells compared to endogenous SGSH mRNA expression in control HEK293 cells. Furthermore, the amount of Q365P protein in HEK293-Q365P cells was 66% of the wild-type protein in HEK293-WT cells. However, while wild-type SGSH-FLAG fusion protein was secreted, Q365P SGSH-FLAG fusion protein was not detected in the culture medium. Moreover, even when twice the amount of protein was loaded onto the gel, no secretion of the Q365P mutant protein was observed.

[0090] The enzymatic activity and intracellular localization of the Q365P variant were analyzed as follows. To verify the enzymatic activity of the Q365P variant, wild-type and Q365P SGSH proteins were purified from the respective cell lysates using anti-FLAG magnetic beads. The activity of the purified wild-type and Q365P SGSH proteins was measured by a two-step enzyme assay using fluorescence assay with 4-methylumbelliferyl 2-deoxy-2-sulfamino-α-D-glucopyranoside (4-MU-GlcNS) as the substrate. The enzyme activity measurements showed that no activity of the Q365P protein was detected, while the wild-type protein showed high activity. The results of the analysis are shown in Figure 9.

[0091] Next, the intracellular localization of Q365P and wild-type SGSH proteins was compared using HEK293-stable expressing cells by immunostaining. After fixing the cells according to standard methods, immunostaining was performed. For primary labeling, anti-FLAG mouse monoclonal antibody and anti-calnexin (CANX) rabbit polyclonal antibody or anti-GOLPH2 rabbit polyclonal antibody were used. Cells were stained with Alexa Fluor 647-conjugated anti-mouse IgG1 and Alexa Fluor 555-conjugated anti-rabbit IgG. Hoechst 33258 was used as nuclear counterstaining. Images were obtained using an LSM 700 confocal laser microscope. Three-dimensional images were constructed using IMARIS software. Co-localization analysis was performed using ImarisColoc software. As a result, wild-type SGSH co-localized with cis-Golgi markers, while co-localization of Q365P SGSH protein was hardly observed. This is consistent with the previous observation (no Q365P SGSH-FLAG fusion protein was secreted). However, both wild-type and Q365P SGSH proteins co-localized with endoplasmic reticulum (ER) markers. Quantitative analysis was also performed. The results are shown in Figure 10. The proportion of SGSH co-localized with the Golgi marker GOLPH2 relative to total SGSH was significantly lower in HEK293-Q365P cells (Figure 10A), and similarly, the proportion of GOLPH2 co-localized with SGSH relative to total GOLPH2 was also significantly lower in HEK293-Q365P cells (Figure 10B). On the other hand, the proportion of SGSH co-localized with the ER marker CANX relative to total SGSH, and the proportion of CANX co-localized with SGSH relative to total CANX, were almost the same in HEK293-WT cells and HEK293-Q365P cells (Figures 10C and D). From these results, it became clear that almost all Q365P SGSH proteins are localized in the endoplasmic reticulum (ER) rather than the Golgi apparatus.

[0092] Next, we investigated the structural instability and loss of function mechanisms of the Q365P variant SGSH protein. Previous results suggest that the Q365P variant SGSH protein cannot form a stable 3D structure and remains in the endoplasmic reticulum (ER) instead of being transported to the Golgi apparatus, which is considered the main cause of the loss of enzyme activity. Therefore, to evaluate the protein's stability, we performed a cycloheximide (CHX) tracking assay. 100 μg / mL of CHX was added to HEK293-WT cells and HEK293-Q365P cells in cell culture, and cells were harvested after 2, 4, and 8 hours. Western blot analysis was then performed using 10 μg of cell lysate protein, and the protein was detected using an HRP-labeled anti-FLAG mouse monoclonal antibody. The results are shown in Figure 11. The cycloheximide tracking assay confirmed that the stability of the Q365P SGSH protein was reduced compared to wild-type SGSH.

[0093] Generally, incompletely folded, unstable proteins are known to be more sensitive to protease digestion than intact proteins. Therefore, the protease resistance of purified wild-type and Q365P SGSH proteins was compared using a trypsin sensitivity assay. The concentrations of wild-type and mutant SGSH proteins were quantified and diluted with TBS to a concentration of 3 μg / μL. Sequencing-grade modified trypsin was added at a concentration of 10 ng / μL and mixed in an 8:1 ratio. A control sample was prepared in the same buffer without trypsin. After incubation at 37°C for 1 hour, the reaction was stopped by adding PMSF. Subsequently, SDS-PAGE was performed, the gel was stained, and the intensity of the resulting bands was quantified using ImageJ. The results are shown in Figure 12. The intensity of the trypsin resistance band after trypsin treatment was significantly lower in the Q365P protein compared to wild-type SGSH. This result further supports the prediction by VarMeter2 that the loss of Q365P SGSH activity is due to a decrease in protein stability.

[0094] Mutations that cause protein destabilization, such as Q365P SGSH, are not only the primary cause of disease due to reduced enzyme activity, but can also exacerbate the disease by causing the accumulation of destabilized proteins in the endoplasmic reticulum and inducing ER stress, as has been reported in diseases such as congenital insensitivity to pain and anhidrosis (CIPA).

[0095] (Example 5) Evaluation of clinical symptoms in SGSH variant Q365P patients As shown below, patients with the Q365P variant of SGSH were diagnosed with Sanfilippo syndrome or mucopolysaccharidosis type III, rare autosomal recessive lysosomal storage disorders caused by impaired heparan sulfate degradation. This diagnosis was consistent with predictions by VarMeter2, suggesting that the Q365P variant is associated with reduced protein stability and contributes to disease manifestation. Sanfilippo syndrome is characterized by severe degeneration of the central nervous system.

[0096] The patient was a 16-year-old female born at 28 weeks gestation with a birth weight of 1338g. She underwent surgery for a right inguinal hernia at 2 years and 1 month of age, and surgery for an abdominal wall incisional hernia 5 months later. She presented with hearing loss, hepatomegaly, hirsutism, and developmental delay. She had adenoids, which were removed. At 5 years and 3 months of age, increased urinary mucopolysaccharides and decreased heparan N-sulfatase activity in the blood were observed. At age 11, a homozygous pathogenic variant of the SGSH gene (p.Q365P) was identified in the patient by trio-whole exome sequencing. VarMeter2 predictions suggested that the Q365P mutation causes decreased protein stability and is involved in disease development. She gradually showed developmental regression, and at age 16, her digital quotient (DQ) score was less than 10.

[0097] The above detailed description merely illustrates the object and subject matter of the present invention and does not limit the scope of the appended claims. Various modifications and substitutions to the embodiments described without departing from the scope of the appended claims will be apparent to those skilled in the art from the teachings described herein. [Industrial applicability]

[0098] This invention provides a method for predicting the impact of missense mutations on protein function. The method is simple and rapid, and may be useful as a novel diagnostic support method for predicting clinically important missense variants that affect protein function. In particular, since the symptoms and signs of rare diseases are often nonspecific, the predictive method of this invention, being a non-symptomatic diagnostic method, may be useful.

Claims

1. A method for predicting whether a missense mutation in a target protein will impair the function of the protein, (1) Using mutation data for an arbitrary protein, create a pathogenic mutation population and a benign mutation population, respectively. For each mutation constituting the population, determine the mutation energy, which indicates the effect of the mutation on stability, the standardized solvent-accessible surface area of ​​the mutation site, and the predicted local distance difference test score. Then, calculate the average values ​​of the mutation energy, standardized solvent-accessible surface area, and predicted local distance difference test score for each population and set the respective average values ​​obtained from these calculations. (2) A process for calculating the mutation energy, the standardized solvent-accessible surface area of ​​the mutation site, and the predicted local distance difference test score for the missense mutation to be predicted. (3) The following formula: [Math 1] (Here, [Math 2] These represent the mean values ​​of mutation energy, standardized solvent-accessible surface area, and predicted local distance difference test score, respectively. [Math 3] These represent the variance of the mutation energy, the standardized solvent-accessible surface area, and the predicted local distance difference test score, respectively. [Math 4] These represent the covariances between mutational energy and standardized solvent-accessible surface area, between mutational energy and predicted local distance difference test score, and between standardized solvent-accessible surface area and predicted local distance difference test score, respectively. From this, the Mahalanobis distance (Di) of the predicted missense mutation from the pathogenic mutation population and the benign mutation population, respectively. 2 The process of calculating ) and (4) Mahalanobis distance (D) of the predicted missense mutation from the pathogenicity mutation population p 2 ) and the Mahalanobis distance (D) from the benign variant population of the predicted missense mutation. B 2 ) compared with, D P 2 <D B 2 In the case of, it is predicted to be a pathogenic mutation, and D P 2 >D B 2 In the case of, the step of predicting it to be a benign mutation, A prediction method that includes this.

2. The prediction method according to claim 1, wherein the step of obtaining the mean values ​​of mutation energy, standardized solvent-accessible surface area, and predicted local distance difference test score in step [1] is performed in advance, and step [1] is a step of setting the mean values ​​of mutation energy, standardized solvent-accessible surface area, and predicted local distance difference test score in each of the pathogenic mutation population and the benign mutation population, which have been performed and calculated in advance.

3. The prediction method according to claim 1, wherein the energy calculation in step (1) and / or (2) is performed using molecular dynamics simulation.

4. The prediction method according to claim 1, wherein the electrostatic energy calculation in the electrostatic interaction in step (1) and / or (2) is performed using a Generalized Born implicate solvento model.

5. The prediction method according to any one of claims 1 to 4, wherein the pathogenic mutation is at least one selected from the group consisting of changes in stability, changes in protease resistance, changes in enzyme activity, changes in binding ability to ligands, substrates, nucleic acids or inhibitors, changes in signal transduction activity, changes in intracellular and extracellular localization, increased aggregation, changes in solubility, changes in three-dimensional structure or oligomeric state, changes in post-translational modification, and increased carcinogenicity.

6. The prediction method according to any one of claims 1 to 4, characterized by being used in combination with other prediction tools for missense mutations.

7. The prediction method according to any one of claims 1 to 4, further comprising the step of associating the missense mutation with a specific disease based on the in vivo function of the target protein when a pathogenic mutation is predicted in step (4).

8. The prediction method according to claim 7, wherein the target protein is N-sulfoglucosamine sulfohydrolase, and the associated disease is Sanfilippo syndrome or mucopolysaccharidosis type III.

9. A system for predicting whether a missense mutation in a target protein will impair the function of the protein. A population variable mean setting module is used to create pathogenic and benign mutation populations using mutation data for an arbitrary protein. For each mutation constituting the population, the mutation energy (indicating the effect of the mutation on stability), the standardized solvent-accessible surface area of ​​the mutation site, and the predicted local distance difference test score are calculated. The mean values ​​of the mutation energy, standardized solvent-accessible surface area, and predicted local distance difference test score are then calculated for each population, and the resulting mean values ​​are set. A target variable calculation module calculates the mutation energy, the standardized solvent-accessible surface area of ​​the mutation site, and the predicted local distance difference test score for the missense mutation to be predicted. The following formula: [Math 5] (Here, [Math 6] These represent the mean values ​​of mutation energy, standardized solvent-accessible surface area, and predicted local distance difference test score, respectively. [Number 7] These represent the variance of the mutation energy, the standardized solvent-accessible surface area, and the predicted local distance difference test score, respectively. [Number 8] These represent the covariances between mutational energy and standardized solvent-accessible surface area, between mutational energy and predicted local distance difference test score, and between standardized solvent-accessible surface area and predicted local distance difference test score, respectively. From this, the Mahalanobis distance (Di) of the predicted missense mutation from the pathogenic mutation population and the benign mutation population, respectively. 2 A Mahalanobis distance calculation module that calculates ) and The Mahalanobis distance (D) of the predicted missense mutation from the pathogenicity mutation population. p 2 ) and the Mahalanobis distance (D) from the benign variant population of the predicted missense mutation. B 2 ) compared with D P 2 <D B 2 In this case, it is predicted to be a pathogenic mutation, D P 2 >D B 2 In this case, a judgment module predicts that it is a benign mutation. A system comprising a computer equipped with a processor that executes each of the above modules.

10. The system according to claim 9, further comprising a population variable mean calculation module that calculates the mean values ​​of mutation energy, standardized solvent-accessible surface area, and predicted local distance difference test score in the pathogenic mutation population and the benign mutation population.

11. The system according to claim 9 or 10, further comprising a display module configured to visually display the above prediction.

12. A method performed in a computer device equipped with a processor for predicting whether a missense mutation in a target protein will impair the function of the protein, A population variable mean setting step involves creating pathogenicity and benign mutation populations using mutation data for an arbitrary protein, determining the mutation energy (which indicates the effect of the mutation on stability), the standardized solvent-accessible surface area of ​​the mutation site, and the predicted local distance difference test score for each individual mutation constituting the respective population, and then calculating the mean values ​​of the mutation energy, standardized solvent-accessible surface area, and predicted local distance difference test score for each population, and setting the respective mean values ​​obtained from these calculations. For the missense mutation to be predicted, the target variable calculation step involves calculating the mutation energy, the standardized solvent-accessible surface area of ​​the mutation site, and the predicted local distance difference test score. The following formula: [Number 9] (Here, [Number 10] These represent the mean values ​​of mutation energy, standardized solvent-accessible surface area, and predicted local distance difference test score, respectively. [Math 11] These represent the variance of the mutation energy, the standardized solvent-accessible surface area, and the predicted local distance difference test score, respectively. [Math 12] These represent the covariances between mutational energy and standardized solvent-accessible surface area, between mutational energy and predicted local distance difference test score, and between standardized solvent-accessible surface area and predicted local distance difference test score, respectively. From this, the Mahalanobis distance (Di) of the predicted missense mutation from the pathogenic mutation population and the benign mutation population, respectively. 2 A Mahalanobis distance calculation step that calculates ) and The Mahalanobis distance (D) of the predicted missense mutation from the pathogenicity mutation population. p 2 ) and the Mahalanobis distance (D) from the benign variant population of the predicted missense mutation. B 2 ) compared with D P 2 <D B 2 In this case, it is predicted to be a pathogenic mutation, D P 2 >D B 2 In this case, the determination step predicts that it is a benign mutation. A method that includes this.

13. The method according to claim 12, further comprising a step of calculating the mean values ​​of the mutation energy, the standardized solvent-accessible surface area, and the predicted local distance difference test score in the pathogenic mutation population and the benign mutation population.