RSA-based methods, systems, and equipment for predicting mutation pathogenicity.

By fusing gene information and protein structure information, utilizing RSA to deeply mine variant features, and combining machine learning algorithms to quantify the exposure status of mutated residues, the problem of inaccurate prediction of pathogenicity of gene mutations in existing technologies has been solved, achieving higher precision in pathogenicity analysis and disease diagnosis.

CN115472220BActive Publication Date: 2026-03-10PEKING UNION MEDICAL COLLEGE HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-21
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing methods for predicting the pathogenicity of gene mutations are insufficient to accurately predict the impact of variations on protein structure and function, resulting in inaccurate pathogenicity assessments.

Method used

By fusing gene information and protein structure information, RSA is used to deeply mine variant features. Combined with machine learning algorithms, the exposure status of mutated residues is quantified. In particular, for ion channel and collagen gene families, AlphaFold2 is used to generate protein structures, and TBtools and HMM are used for gene family analysis. RSA is used to quantify the pathogenicity of mutations.

Benefits of technology

It significantly improves the accuracy and depth of gene mutation pathogenicity prediction, better reflects the spatial clustering of pathogenic variants, and enhances the accuracy of auxiliary diagnosis and treatment selection for disease development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115472220B_ABST
    Figure CN115472220B_ABST
Patent Text Reader

Abstract

This invention relates to a method, system, and device for predicting the pathogenicity of mutations based on RSA (Reactive Signal Absorption Spectrum). The method includes: acquiring the mutation of the gene to be predicted; mapping the mutation of the gene to be predicted onto the protein structure of the gene to obtain the mapped protein structure; extracting features from the mapped protein structure to obtain the structural features of the mutated protein; inputting the mutation of the gene to be predicted into a gene family analyzer to obtain the gene family to which the mutation belongs; when the gene family belongs to an ion channel, selecting RSA from the structural features of the mutated protein for prediction to obtain a classification result of whether the mutation is pathogenic or benign. This invention aims to predict the mutation of the gene to be predicted based on RSA from the structural features of the mutated protein, exploring its high predictive ability and potential application value for the pathogenicity of mutations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of gene data analysis, and more specifically, to an analytical method, apparatus, system, computer-readable storage medium, and application of RSA-based prediction of mutation pathogenicity. Background Technology

[0002] The human genome contains approximately 3.16 billion base pairs, encoding about 30,000 genes, with a variation occurring on average once every 500-1000 base pairs. Within the human genome, protein-coding regions alone contain enormous differences between individuals; to date, 6.5 million missense variants have been observed. In fact, aside from the potentially lethal protein locations in the genomes of the 8 billion people living on Earth, every protein location is potentially susceptible to variation. Most clinically relevant variants have highly variable effects on protein structure and function, with only a small fraction being pathogenic.

[0003] The AlphaFold2 protein structure generation model expands the structural coverage of the human proteome from 17% to 98.5%, and this extensive and accurate structural information has unprecedented potential to help predict variant effects. To assess the pathogenicity of gene mutations, many state-of-the-art variant effect predictors have been developed, such as SIFT, REVEL, and EVE. Given the close relationship between protein structure and function, the structural background of mutations represents promising information independent of mutation frequency and evolutionary conservation. The degree of residue burial or exposure in the 3D structure is crucial for protein folding and stability. Therefore, regardless of specific amino acid changes, the location of the mutation is a strong predictor of variant pathogenicity. Summary of the Invention

[0004] The method of this invention integrates the information of the gene to be predicted and the predicted protein structure information, and uses RSA to deeply mine the variation features in the gene data, thereby predicting the pathogenicity of the mutation of the gene to be predicted, exploring its high predictive ability and potential application value for the pathogenicity of the mutation, and solving related life science problems.

[0005] This application discloses an analytical method for predicting the pathogenicity of mutations based on RSA, including:

[0006] Obtain the mutation of the gene to be predicted, wherein the gene to be predicted includes wild-type genes;

[0007] The mutation of the gene to be predicted is mapped onto the protein structure of the gene to be predicted, and the protein structure mapped with the mutation is obtained.

[0008] Feature extraction is performed on the mapped mutated protein structure to obtain the mutated protein structure features;

[0009] The mutation of the gene to be predicted is input into the gene family analyzer to obtain the gene family to which the mutation of the gene to be predicted belongs.

[0010] When the gene family to which the mutation of the gene to be predicted belongs is an ion channel, the RSA in the structural features of the mutant protein is selected for prediction to obtain the classification result of whether the mutation is a pathogenic mutation or a benign mutation.

[0011] Furthermore, obtaining the mutation of the gene to be predicted includes preprocessing the mutation of the gene to be predicted and then mapping it onto the protein structure of the gene to be predicted.

[0012] Optionally, the preprocessing uses an algorithm to predict the effects of mutation splicing and exclude mutations with scores exceeding a threshold; preferably, the SpliceAI algorithm is used to predict the effects of mutation splicing and exclude mutations with scores greater than 0.5.

[0013] Optionally, the preprocessing includes excluding mutations with a residue count exceeding a threshold;

[0014] Optionally, the pretreatment includes excluding mutations that cannot be mapped to a structure due to isoform inconsistency; alternatively, the pretreatment includes excluding mutations with clinical significance that have "contradictory explanations for pathogenicity".

[0015] Furthermore, the protein structure of the gene to be predicted is obtained by inputting the gene to be predicted into a protein structure generation model; optionally, the protein structure generation model includes one or more of the following models: AlphaFold, AlphaFold2, ProteoGAN; preferably, the protein structure generation model is AlphaFold2.

[0016] Furthermore, the mutation of the gene to be predicted is input into a gene family analyzer to obtain the gene family to which the mutation of the gene to be predicted belongs; the gene family analyzer obtains the gene family to which the mutation belongs through conserved structural domain analysis. Optionally, the gene family analyzer includes any one or more of the following software: TBtools, HMM, Motif; preferably, the gene family analyzer is TBtools or HMM.

[0017] Furthermore, when the gene family to which the mutation of the gene to be predicted belongs is an ion channel, the classification result of the mutation as a pathogenic mutation or a benign mutation is obtained based on the exposure status between the mutation residues quantified by RSA. The exposure status includes hidden residues (low RSA residues), intermediate residues (medium RSA residues), and exposed residues (high RSA residues).

[0018] Furthermore, the mutant protein structural features are obtained by extracting the mutant protein structural features, which are reflected at three levels: protein level features, residue level features, and mutation level features. They mainly include thermodynamic features, protein volume features, secondary structure related features, relative solvent accessibility features of mutant residues, and empirical rule features that can determine whether mutations have a significant impact on protein structure.

[0019] Optionally, the protein level characteristics include thermodynamic characteristics and protein volume characteristics;

[0020] Optionally, the residue-level features include secondary structure-related features and RSA, wherein the secondary structure-related features include features generated using the DSSP program: DSSP(H), DSSP(E), DSSP(G), DSSP(I), DSSP(T), DSSP(S), DSSP(C);

[0021] Optionally, the mutation level characteristics include thermodynamic characteristics and empirical rules that can determine whether a mutation has a significant impact on protein structure.

[0022] Furthermore, when the gene family to which the mutation of the gene to be predicted belongs is the collagen gene family, thermodynamic features, protein volume features, secondary structure related features, and mutation level features among the structural features of the mutant protein are selected for prediction to obtain the classification result;

[0023] When the gene family to which the mutation of the gene to be predicted belongs is neither collagen nor ion channel, the structural features of the mutant protein are used for prediction to obtain a classification result; wherein, the structural features of the mutant protein include protein-level features, residue-level features, and mutation-level features.

[0024] An RSA-based device for predicting mutation pathogenicity analysis, the device comprising:

[0025] Memory and processor;

[0026] The memory is used to store program instructions;

[0027] The processor is used to call program instructions, which, when executed, are used to perform the aforementioned analysis method for predicting the pathogenicity of mutations based on RSA.

[0028] A system for predicting mutation pathogenicity based on RSA, comprising:

[0029] The acquisition module is used to acquire mutations in the gene to be predicted.

[0030] The structure generation module is used to map the obtained mutations of the gene to be predicted onto the protein structure of the gene to be predicted, so as to obtain the protein structure mapped with the mutations.

[0031] The feature processing module processes the mapped mutated protein structure to obtain the mutated protein structure features;

[0032] The gene family generation module inputs the mutation of the gene to be predicted into the predictor to obtain the gene family to which the mutation of the gene to be predicted belongs.

[0033] The classification module selects RSA from the structural features of the mutant protein when the gene family to which the mutation belongs is an ion channel, and obtains the classification result of the pathogenicity of the mutation.

[0034] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described analysis method for predicting the pathogenicity of mutations based on RSA.

[0035] The association between the aforementioned structural features and the pathogenicity of variants was assessed based on genotype data types. For dichotomous features, contingency tables were constructed, and the chi-square test was used to determine whether an association existed between the two dichotomous variables. For continuous features, logistic regression analysis was used to examine the association between the feature and the pathogenicity of variants, and the strength of the association between the feature and the pathogenicity of variants was quantified using the odds ratio (OR) and 95% confidence interval (CI). All statistical analyses and data visualizations were performed using R software packages: h2o, caret, pROC, forestplot, ggpubr, ggsci, viridis, and cutpointr. A p-value <0.05 was considered statistically significant.

[0036] The above-mentioned equipment is used in the selection of auxiliary mutation pathogenicity analysis schemes, the schemes of which include mutation pathogenicity analysis affected by changes in the type and number of gene mutation characteristics; optionally, the mutation pathogenicity analysis has a positive impact on research on diabetes, blood vessels, bones, brain function and anti-aging.

[0037] Application of the aforementioned equipment in predicting the pathogenicity of mutations;

[0038] The above-mentioned device is used in mutation classification or prediction of mutation attributes; optionally, the attributes include the gene family to which the mutation of the gene to be predicted belongs, the gene family including ion channels and collagen;

[0039] The above-mentioned system is applied in the diagnosis of the occurrence and development of mutational pathogenicity; optionally, the occurrence and development of mutational pathogenicity is related to changes in the type and number of gene mutation characteristics; optionally, the occurrence and development of mutational pathogenicity has a positive impact on research on diabetes, blood vessels, bones, brain function and anti-aging.

[0040] Application of the above system in predicting mutation pathogenicity;

[0041] The above-described system is applied to mutation classification or prediction of mutation attributes; optionally, the attributes include the gene family to which the mutation of the gene to be predicted belongs, the gene family including ion channels and collagen.

[0042] This invention trains high-quality mutation data with clinical significance using machine learning algorithms. By identifying key regions susceptible to genetic variations, the structural feature RSA calculated based on the structure of mutant proteins can better reflect the spatial clustering of pathogenic variations, significantly improving prediction performance. It explores and determines the significant location clustering effect of pathogenic variations based on the structural features of mutant proteins, which is highly innovative in the field of life sciences and will have a beneficial promoting effect on the pathogenicity analysis of gene data.

[0043] Advantages of this application:

[0044] 1. This application innovatively discloses a novel method for analyzing the pathogenicity of mutated genes. This method is based on RSA to deeply mine the life laws hidden behind gene data, and maps the mutation of the gene to be predicted to the protein structure of the gene to be predicted, thereby obtaining the protein structure mapped with the mutation. When the gene family to which the mutation of the gene to be predicted belongs is an ion channel, the characteristic RSA is extracted based on the protein structure mapped with the mutation to quantify the exposure status of the mutated residues. Through in-depth analysis of the pathogenicity of the mutation, the accuracy and depth of data analysis are improved.

[0045] 2. This application innovatively uses the gene family to which the mutation of the gene to be predicted belongs, and selects effective mutant protein structural features to predict the variation effects of multiple gene families such as collagen and ion channels. It has been determined that RSA is significantly effective in predicting ion channel gene families. Among them, thermodynamic features, protein volume features, secondary structure-related features, and mutation level features of mutant protein structural features are effective in predicting collagen gene families. In addition, protein level features, residue level features, and mutation level features are effective in predicting non-ion channel and collagen gene families.

[0046] 3. This application creatively discloses an RSA-based device and system for predicting the pathogenicity of mutations. By deeply interpreting gene data through structural feature RSA and combining it with other VEP analyses to predict the pathogenicity of mutations in the gene to be predicted, it can better reflect the spatial clustering of pathogenic variants, significantly improve performance, and make this application more accurately applied to the auxiliary diagnosis and treatment selection of diseases related to gene data. Attached Figure Description

[0047] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0048] Figure 1 This is a schematic flowchart of the mutation pathogenicity analysis based on RSA prediction of ion channel gene families provided in the embodiments of the present invention;

[0049] Figure 2 This is a schematic flowchart of RSA-based analysis of the pathogenicity of mutations in various gene families provided in this embodiment of the invention;

[0050] Figure 3 This is a protein structure feature map based on AlphaFold2 that maps to mutations, provided in an embodiment of the present invention.

[0051] Figure 4 This is an assessment chart of the importance of RSA in predicting the pathogenicity of variants, provided in an embodiment of the present invention.

[0052] Figure 5 This is an analysis chart of the effectiveness of RSA-based prediction of TSC2 protein mutations provided in an embodiment of the present invention;

[0053] Figure 6 This is an effect analysis diagram of SIGMA+-based prediction of mutation pathogenicity provided in an embodiment of the present invention;

[0054] Figure 7 This is a schematic diagram of the pathogenicity analysis device based on RSA prediction of mutations provided in an embodiment of the present invention. Detailed Implementation

[0055] To enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0056] In some of the processes described in the specification, claims, and accompanying drawings of this invention, multiple operations appearing in a specific order are included. However, it should be clearly understood that these operations may not be performed in the order they appear herein, or may be performed in parallel. The operation numbers, such as S101, S102, etc., are merely used to distinguish different operations and do not themselves represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be performed sequentially or in parallel.

[0057] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0058] Figure 1 This is a schematic flowchart of an embodiment of the present invention for predicting the pathogenicity of ion channel gene mutations based on RSA. Specifically, the method includes the following steps:

[0059] S101: Obtain the mutation of the gene to be predicted.

[0060] In one embodiment, the mutation information of the gene to be tested includes publicly available genetic data, including wild-type genes. Optionally, the mutation information of the gene to be tested also includes genetic information of the sample obtained by high-throughput sequencing. Optionally, the mutation of the gene to be predicted includes one or more of the following datasets: dbSNP, dbVar, gnomAD, RefSeq, and ExAC, where dbSNP, dbVar, and RefSeq are all from the ClinVar database. A total of 27,165 benign and 22,957 pathogenic missense variants were retrieved from the gnomAD and ClinVar databases. For ClinVar mutations, single nucleotide missense mutations with an audit status of at least one star were retained (expert panel audit: provided standard, multiple submitters, no conflict; provided standard, single submitter).

[0061] Furthermore, step S101 also includes preprocessing the mutation data; optionally, the preprocessing uses an algorithm to predict the impact of mutations on splicing and excludes mutations with scores exceeding a threshold; preferably, the SpliceAI algorithm is used to predict the impact of mutations on splicing and exclude mutations with scores greater than 0.5.

[0062] SpliceAI predicts splicing changes caused by single nucleotide variations, accurately predicting splicing sites (location and probability of abnormal splicing) from any mRNA precursor sequence, thus enabling the prediction of hidden splicing caused by mutations in non-coding RNA regions.

[0063] Optionally, the pretreatment also includes excluding mutations with more than a threshold number of residues; for example, proteins with more than 2,700 residues are excluded.

[0064] Optionally, preprocessing includes excluding mutations that cannot be mapped to the structure due to isomer inconsistencies;

[0065] Optionally, pretreatment may include excluding mutations with clinical significance that are “conflicting in their pathogenicity explanations”.

[0066] In one embodiment, the mutated gene comprised 20,047 pathogenic (red) and 20,148 benign (blue) variants, and also included 27,928 mutations in six proteins characterized by the DMS experimental system, using 27 computational VEPs, such as EVE screening and DEOGEN2. The EVE score for each mutation was obtained from https: / / evemodel.org / . Predictions from other VEPs for each mutation were retrieved from the dbNSFP database (version 4.1a). These VEPs were divided into two categories: single predictors independent of other VEPs (n=16) and meta-predictors that integrated the results of other VEPs into input features (n=11).

[0067] S102: Map the mutation of the gene to be predicted onto the protein structure of the gene to be predicted, and obtain the protein structure mapped with the mutation.

[0068] In one embodiment, the protein structure of the gene to be predicted is obtained by inputting the gene to be predicted into a protein structure generation model. A protein structure generation model refers to a model or software capable of generating the protein structure of a gene or polypeptide based on its amino acid sequence. Optionally, the protein structure generation model includes one or more of the following generation models: VAE, AlphaFold, AlphaFold2, and ProteoGAN; preferably, the protein structure generation model is AlphaFold2.

[0069] VAE (Variational Autoencoder) is a generative model based on joint probabilities.

[0070] AlphaFold accelerated and enabled large-scale discoveries, including deciphering the structure of the nuclear pore complex. AlphaFold predictions, with very high confidence, provided a nine-helix topology, offering clues about protein function that are crucial for unraveling the causes of rare genetic diseases.

[0071] AlphaFold2 discovered a novel protein-gated mechanism, glucose-6-phosphatase, an enzyme that identifies binding sites for inhibiting enzymes, and a transmembrane protein, Wolframin, located in the ER. It improved the accuracy of predictions to the atomic level, more quickly and accurately identifying enzyme active sites, with 350,000 predicted sites.

[0072] ProteoGAN, a conditional generative adversarial network, outperforms classic and recent deep learning baselines in protein sequence generation. It primarily expands protein screening with candidates that are farther from the known sequence space than previously possible, but is more likely to have functionality with relatively novel candidates than other methods.

[0073] In one embodiment, preferably, the protein structure of the gene to be predicted is obtained by inputting the wild-type gene into the AlphaFold2 protein structure generation model. The protein structure dataset in AlphaFold2 includes the AlphaFold Protein Structure Database (AlphaFoldDB, https: / / alphafold.ebi.ac.uk / ), containing approximately 365,000 structure predictions. AlphaFold2 is used to provide these proteins with 1400-amino acid fragments; proteins exceeding 2700 residues are excluded. All mutations are mapped to the predicted protein 3D structure; mutations that cannot be mapped to a structure due to isoform inconsistencies are excluded.

[0074] S103: Extract features from the protein structure mapped with mutations to obtain the structural features of the mutated protein.

[0075] In one embodiment, the structural features of a mutant protein include protein-level features, residue-level features, and mutation-level features.

[0076] In one specific embodiment, mutations are mapped onto wild-type protein structures generated based on AlphaFold2, and then the resulting mutated protein structures are used to extract mutant protein features. The extracted mutant protein features include 57 mutant protein structural features across three levels: protein-level features, residue-level features, and mutation-level features. These 57 mutant protein structural features include thermodynamic features, protein volume features, secondary structure-related features, relative solvent accessibility features of mutated residues, and empirical rule features that can determine whether a mutation has a significant impact on protein structure. Pathogenic mutations exhibit greater changes in protein stability than benign mutations.

[0077] Optionally, protein-level features include 16 thermodynamic features and 3 protein volume features; the thermodynamic features include 15 components such as total unfolding energy and its van der Waals collisions, hydrogen bond energy, and side chain entropy; the protein volume features include total volume, interstitial volume, and van der Waals volume.

[0078] Optionally, residue-level features include eight secondary structure-related features and the relative solvent accessibility (RSA) of the mutated residue, used to characterize the structural background of the mutated residue. For each mutated residue, its secondary structure is assigned using the DSSP (protein secondary structure dictionary) program, generating features corresponding to eight secondary structure types: DSSP(H), DSSP(E), DSSP(G), DSSP(I), DSSP(T), DSSP(S), and DSSP(C).

[0079] Optionally, mutation level features include 16 thermodynamic features describing changes in protein stability after mutation and 13 features derived from empirical rules for determining whether a mutation has a significant effect on protein structure; the 16 thermodynamic features include the free energy difference between the mutant protein and its 15 components; the 13 features derived from empirical rules for determining whether a mutation has a significant effect on protein structure are a combination of the structural background and physicochemical properties of the mutant / mutant residues, including glycine bending, proline in the α-helix, substitution of residues in the α-helix by proline / glycine, cysteine ​​residues, energy charge loss, energy hydrophobicity replacement, outlier substitution, disulfide bond breakage, buried salt bridge breakage, buried proline, energy replacement of glycine, protein unfolding free energy, and the unfolding free energy difference between mutant proteins.

[0080] Figure 3 This study demonstrates that the structural features of mutant proteins are associated with pathogenicity, particularly RSA and ΔΔG, which are significantly correlated with variant pathogenicity. From... Figure 3 As shown in Figure A, the RSA of pathogenic variants was significantly lower than that of benign variants (P<2.2e-16, Mann-Whitney U test), which is consistent with the fact that most proteins are less tolerant of hidden mutations than exposed mutations. Figure 3 As shown in B, ΔΔG measures the effect of a single amino acid substitution on protein stability; pathogenic variants showed greater changes in protein stability than benign variants (P < 2.2e-16, Mann-Whitney U test). Figure 3 As shown in Figure C, as expected, disulfide bond breaking has the highest positive predictive value among all structural features (odds ratio [OR] = 93.8, 95% confidence interval [CI] = 44.5–198, P < 2.2e-16, Pearson chi-square test). Almost all (98.72%) missense mutations that break disulfide bonds are pathogenic, supporting the important role of disulfide bonds in protein function. Figure 3 As shown in D, the association between the type of secondary structure in which the mutation occurs and its pathogenicity is consistent with prior knowledge. Mutations in the circular or irregular extensions of proteins (Dictionary of Protein Secondary Structures - DSSP-C) are often benign (OR = 0.32, 95% CI = 0.31–0.34, P < 2.2e-16, Pearson chi-square test). In contrast, mutations in regular secondary structures are often pathogenic, especially mutations in α-helices (DSSP-H; OR = 1.73, 95% CI = 1.66–1.79, P < 2.2e-16, Pearson chi-square test) or β-sheets (DSSP-E; OR = 1.97, 95% CI = 1.87–2.08, P < 2.2e-16, Pearson chi-square test).

[0081] S104: Input the mutation of the gene to be predicted into the gene family analyzer to obtain the gene family to which the mutation of the gene to be predicted belongs.

[0082] In one instance, gene family analysis methods can generally be divided into two types: one is based on sequence alignment to predict genes with similar sequences, thereby identifying a gene family; the other is based on structure for prediction. Optionally, the gene family analyzer includes any one or more of the following software: TBtools, HMM, Motif; preferably, the gene family analyzer is TBtools or HMM.

[0083] TBtools is an excellent bioinformatics software that can be used for gene family analysis. Its main tools include: processing GFF3 files, predicting MEME results, multigenomic gene structure analysis, collinearity analysis, and visualization.

[0084] Hidden Markov Models (HMMs) utilize information from conserved domains to locate corresponding genomic information, describing a Markov process with hidden, unknown parameters. The challenge lies in determining these hidden parameters from observable parameters and then using them for further analysis.

[0085] Furthermore, the gene families it belongs to mainly include collagen, ion channels, and other gene families.

[0086] S105: When the gene family to which the mutation of the gene to be predicted belongs is an ion channel, select the RSA in the structural features of the mutant protein.

[0087] Figure 4 This diagram illustrates the importance of RSA in predicting the pathogenicity of variants. Figure 4 A shows the ten most important features affecting the discriminative power of RSA. Among them, residue-level features contribute the most to the discriminative power of RSA, followed by two mutation-level features (ΔΔG and ΔVander). Seven of these ten most important features are protein-level features, demonstrating the important role of protein stability in predicting the pathogenicity of variants. Figure 4 B and Figure 4 C represents the gene set enrichment analysis (GSEA) plots for the ion channel gene family (B) and the collagen gene family (C), respectively. Notably, RSA alone can predict the pathogenicity of mutations in 76 out of 200 genes with an accuracy of >0.8 (with at least 10 pathogenic variants and 10 benign variants used for evaluation), especially showing significant predictive power for ion channel gene mutations.

[0088] In addition, from Figure 4The results showed that RSA had the highest accuracy in classifying ion channel genes (P = 8.67e-04); however, RSA alone was not satisfactory for 11% of the genes (21 out of 200 genes, with an accuracy of <0.5), which represented the collagen gene family (P = 1.99e-06). Figure 4 C). Where ΔG represents the difference in unfolding free energy between the mutant protein and the mutant protein; ΔG represents the unfolding free energy of the protein; and VdW represents van der Waals.

[0089] In a specific instance, such as Figure 2 As shown, when the gene family to which the mutation belongs is an ion channel, the RSA feature in the mutant protein structure feature is selected for prediction to obtain the classification result of whether the mutation is a pathogenic mutation or a benign mutation; when the gene family to which the mutation belongs is collagen, mutant protein structure features other than RSA are selected; when the gene family to which the mutation belongs is neither collagen nor an ion channel, mutant protein structure features are used for prediction to obtain the classification result.

[0090] S106: The classifier predicts whether the mutation is a pathogenic mutation or a benign mutation.

[0091] In one embodiment, the classifier employs one or more of the following: Boosting, XGBoost, and AdaBoost. Boosting is a mainstream representative technique in ensemble learning in machine learning. XGBoost, which has outperformed deep neural networks in many data analysis competitions, is an efficient implementation of the GradientBoost algorithm within the Boosting family.

[0092] Figure 5 This is a graph showing the effectiveness of RSA-based prediction of TSC2 protein mutations provided in an embodiment of the present invention: Figure 5 A predicted the SIGMA, EVE, and REVEL scores for all marker mutations in the TSC2 protein. Figure 5 B shows a heatmap of SIGMA scores for all possible amino acid substitutions in the TSC2 protein. As shown in 5A, the classification of mutations as pathogenic or benign using the SIGMA model is based on the RSA quantification of the exposure status of residues between mutated residues. The exposure status includes hidden residues (low RSA residues), intermediate residues (medium RSA residues), and exposed residues (high RSA residues).

[0093] Figure 5Both EVE and REVEL in the framework are predictors that predict the mutation effects of the acquired gene to be predicted, resulting in a mutation effect map of the gene to be predicted. Predictors include any one or more of the following software: Enformer, PROVEAN, EVE, DEOGEN2, MUTPRED, MutationAssessor, SIFT, ClinPred, PolyPhen2, and REVEL; the predictors are PROVEAN, EVE, DEOGEN2, and MUTPRED.

[0094] PROVEAN is a tool for predicting whether protein sequence variations affect protein function. Based on evolutionary conservation, neural network models, and the BLOSUM62 amino acid substitution scoring matrix, it identifies whether non-synonymous mutations or InDels affect protein biological function. Prediction scores range from -14 to 14, with a threshold of -2.5. Scores of -14 to 2.5 are predicted as Deleterious, and scores of -2.5 to 14 are predicted as Neutral. Lower scores indicate more harmful mutations, and vice versa.

[0095] EVE estimates the likelihood of each single amino acid variation being benign or pathogenic, understands the pathogenicity of human missense variations from the distribution of sequence variations across species, outperforms other computational prediction models in predicting clinical outcomes, and scores as high as or better than the current gold standard high-throughput experiments for testing the impact of mutations on biological function.

[0096] DEOGEN2 contains heterogeneous information about the molecular effects of mutations, the domains involved, gene relevance, and the interactions they participate in. This extensive contextual information is non-linearly mapped to a single harmfulness score for each mutation.

[0097] MUTPRED's function is to predict the pathogenicity and molecular mechanisms of amino acid substitutions, using Fasta-formatted amino acid sequences as the primary input. MUTPRED is a collection of machine learning tools that can predict the pathogenicity of protein-coding mutations to infer the molecular mechanisms of diseases.

[0098] REVEL is an ensemble method for predicting missense mutations. It combines the prediction results of multiple software programs (MutPred, FATHMM, VEST, PolyPhen, SIFT, PROVEAN, Mutation Assessor, Mutation Taster, LRT, GERP, SiPhy, phyloP, and phastCons) and uses a random forest algorithm, which shows good prediction results for rare missense mutations. The prediction score for a single mutation ranges from 0 to 1, and the threshold can be set according to the required sensitivity and specificity.

[0099] In a specific embodiment, the SIGMA (Structure-Informed Germline Missense Mutation Assessor) model is developed based on the close relationship between protein structure and function. It primarily uses machine learning algorithms such as gradient enhancement machines to evaluate the impact of missense variants within the context of protein structure. The construction of the SIGMA model mainly includes feature selection, feature fusion, feature computation, and model optimization. Specifically, the SIGMA model construction process is as follows: A mutation dataset containing label information is obtained; the mutation data is mapped to wild-type protein structures to obtain the mapped mutated protein structures; feature selection is performed on the mapped mutated protein structures to obtain mutated protein structural features; machine learning algorithms are used to process the mutated protein structural features to obtain predicted classification results; based on the predicted classification results and actual results, the machine learning algorithm is optimized to obtain the final mutation classification result of the gene to be predicted, outputting the SIGMA model.

[0100] Further, optionally, the mutation dataset includes one or more of the following datasets: gnomAD, HumVar, ExoVar, PredictSNP, VariBench, SwissVar, Humsavar, and ClinVar. The final labeled dataset containing 27,165 benign (negative) and 22,957 pathogenic (positive) missense mutations is used as the "gold standard" dataset. The dataset is divided into 80% for training and 20% for testing. The gnomAD database retains 27,928 mutations for model evaluation. Training the SIGMA model can be as follows: obtain labeled samples of different categories from the dataset (e.g., pathogenic / potentially pathogenic variants are labeled positive, while benign / potentially benign variants are labeled negative); repeat the protein structure generation and feature extraction processes described above to obtain a feature set; input the feature set into the SIGMA model to obtain the classification results of the samples; calculate the loss between the classification results and the true values ​​using a loss function; then perform backpropagation; update the parameters using an optimizer to obtain the trained SIGMA model.

[0101] Furthermore, the labeling information was defined as follows: pathogenic / potentially pathogenic variants were labeled positive, while benign / potentially benign variants were labeled negative. For gnomAD mutations, preprocessing was performed using SpliceAI to predict the gnomAD database (for deep intron regions: distance from the exon-intron boundary >50 nt). Due to the difficulty of deep intron prediction, only 56% deletions were observed when the threshold was set to 0.8. Therefore, the impact of mutations on splicing was predicted, and mutations with SpliceAI scores greater than 0.5 were excluded. Mutations introducing mystery splicing sites primarily affected mRNA splicing rather than protein structure, and selected common missense mutations (the highest allele frequency in all populations with gnomAD > 0.05) were labeled negative. In addition, the GOF / LOF database labeled 193 gain-of-function (GoF) and 921 loss-of-function (LoF) pathogenic missense variants.

[0102] Optionally, the preprocessing of the mutation dataset may also include using an algorithm to predict the impact of mutations on splicing and excluding mutations with scores exceeding a threshold; preferably, the SpliceAI algorithm is used to predict the impact of mutations on splicing and exclude mutations with scores greater than 0.5.

[0103] Optionally, pretreatment includes excluding mutations with more than a threshold of 2700 residues;

[0104] Optionally, preprocessing includes excluding mutations that cannot be mapped to the structure due to isomer inconsistencies;

[0105] Optionally, pretreatment may include excluding mutations with clinical significance that are “conflicting in their pathogenicity explanations”.

[0106] Furthermore, the machine learning algorithms used include one or more of the following: GBDT, GBM, SVM, RF, Adaboost, and Apriori. Additionally, model optimization includes one or more of the following methods: steepest descent, Newton's method, and quasi-Newton methods.

[0107] The above method is used to select a gene mutation pathogenicity analysis scheme. Based on the gene information of the mutation to be predicted, the classification result of whether the mutation is pathogenic can be predicted. Figure 6 This is a pathogenicity effect analysis diagram based on SIGMA+ predicted mutations provided in an embodiment of the present invention. Figure 6 A demonstrates the potential applications of combining SIGMA with other predictors. For example, combining SIGMA with four individual VEPs with high performance, namely DEOGEN2, EVE, PROVEAN, and MutPred, yields a more comprehensive predictor of variant pathogenicity. Various combinations of these five predictors enhance predictive performance (AUC range of 0.942–0.954). Figure 6 B specifically demonstrates the high predictive power of SIGMA+ for the pathogenicity of variants (AUC = 0.966). Here, SIGMA+ is the combination of all five predictors, namely the combination of SIGMA, DEOGEN2, EVE, PROVEAN and MutPred, which can achieve a high predictive power of AUC = 0.966, significantly distinguishing between benign and pathogenic variants. Figure 6 C demonstrates the correlation between VEP and deep mutation scanning (DMS) measurements. From Figure 6 The results show that SIGMA contributes the most to the combination, while DEOGEN2 contributes the least, indicating that SIGMA has strong discriminative ability, while SIGMA+ has a higher predictive ability for the pathogenicity of variants. Therefore, mining additional information sources may be more beneficial than iteratively building meta-predictors using more and more existing predictors.

[0108] This invention provides an RSA-based system for predicting mutation pathogenicity, comprising:

[0109] The acquisition module is used to acquire mutations in the gene to be predicted.

[0110] The structure generation module is used to map the obtained mutations of the gene to be predicted onto the protein structure of the gene to be predicted, so as to obtain the protein structure mapped with the mutations.

[0111] The feature processing module processes the mapped mutated protein structure to obtain the structural features of the mutated protein;

[0112] The gene family generation module inputs the obtained mutations of the gene to be predicted into the predictor to obtain the gene family to which the mutation of the gene to be predicted belongs.

[0113] The classification module selects RSA from the structural features of the mutant protein when the gene family to be predicted belongs to an ion channel, and obtains the classification result of whether the mutation is a pathogenic mutation or a benign mutation.

[0114] Figure 7 This invention provides an RSA-based predictive mutation pathogenicity analysis device, comprising: a memory and a processor; the device may further include: an input device and an output device.

[0115] Memory, processor, input devices, and output devices can be connected via a bus or other means, such as Figure 7 The example shown is a bus-based connection;

[0116] Memory is used to store program instructions;

[0117] The processor is used to call program instructions, which, when executed, are used to perform the aforementioned analysis method for predicting the pathogenicity of mutations based on RSA.

[0118] The present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described analysis method for predicting the pathogenicity of mutations based on RSA.

[0119] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and modules described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0120] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between apparatuses or modules, and may be electrical, mechanical, or other forms.

[0121] The modules described as separate components may or may not be physically separate. Similarly, the components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0122] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The aforementioned integrated modules can be implemented in hardware or as software functional modules.

[0123] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc.

[0124] Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware, and the program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0125] The computer device provided by the present invention has been described in detail above. For those skilled in the art, there will be changes in the specific implementation and application scope based on the ideas of the embodiments of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. An analysis method for predicting pathogenicity of a mutation based on RSA, characterized by, The method comprises: obtaining a mutation of a gene to be predicted, wherein the gene to be predicted comprises a wild-type gene; mapping the mutation of the gene to be predicted to a protein structure of the gene to be predicted to obtain a protein structure with the mutation mapped; extracting features of the protein structure with the mutation to obtain mutation protein structure features; inputting the mutation of the gene to be predicted into a gene family analyzer to obtain a gene family to which the mutation of the gene to be predicted belongs; when the gene family to which the mutation of the gene to be predicted belongs is an ion channel, selecting RSA in the mutation protein structure features for prediction to obtain a classification result that the mutation is a pathogenic mutation or a benign mutation.

2. The method of claim 1, wherein the method is based on the analysis of the pathogenicity of the predicted mutation of the RSA. The classification result that the mutation is a pathogenic mutation or a benign mutation is obtained based on RSA quantifying the exposure state between mutation residues.

3. The method of claim 1, wherein the method is based on the analysis of the pathogenicity of the predicted mutation of the RSA. When the gene family to which the mutation of the gene to be predicted belongs is a collagen gene family, selecting thermodynamic features, protein volume features, secondary structure related features and mutation level features in the mutation protein structure features for prediction to obtain a classification result; wherein the secondary structure related features include features assigned using the DSSP program: DSSP (H), DSSP (E), DSSP (G), DSSP (I), DSSP (T), DSSP (S), DSSP (C).

4. The method of claim 3, wherein the method is based on the analysis of the pathogenicity of the predicted mutation of the RSA. The mutation level features include features of empirical rules that can determine whether a mutation has a significant impact on a protein structure.

5. The method of claim 1, wherein the method is based on the analysis of the pathogenicity of the predicted mutation of the RSA. When the gene family to which the mutation of the gene to be predicted belongs is neither collagen nor ion channel, using the mutation protein structure features for prediction to obtain a classification result; wherein the mutation protein structure features include protein level features, residue level features and mutation level features.

6. The method of claim 5, wherein the method is based on the analysis of the pathogenicity of the predicted mutation of the RSA. The protein level features include thermodynamic features and protein volume features.

7. The method of claim 5, wherein the method is based on the analysis of the pathogenicity of the predicted mutation of the RSA. The residue level features include secondary structure related features and RSA, wherein the secondary structure related features include features assigned using the DSSP program: DSSP (H), DSSP (E), DSSP (G), DSSP (I), DSSP (T), DSSP (S), DSSP (C).

8. The method of claim 5, wherein the method is based on the analysis of the pathogenicity of the predicted mutation of the RSA. The mutation level features include thermodynamic features and features of empirical rules that can determine whether a mutation has a significant impact on a protein structure.

9. The method of claim 1, wherein the method is based on the analysis of the pathogenicity of the predicted mutation of the RSA. The mapping of the mutation of the gene to be predicted to the protein structure of the gene to be predicted further comprises mapping the mutation of the gene to be predicted to the protein structure of the gene to be predicted after preprocessing.

10. The method of claim 9, wherein the method is based on the analysis of the pathogenicity of the predicted mutation of the RSA. The preprocessing uses an algorithm to predict the impact of mutation splicing and excludes mutations with scores exceeding a threshold.

11. The method of claim 9, wherein the method is based on the analysis of the pathogenicity of the predicted mutation of the RSA. The preprocessing uses the SpliceAI algorithm to predict the impact of mutation splicing and excludes mutations with scores greater than 0.

5.

12. The method of claim 9, wherein the method is based on the analysis of the pathogenicity of the predicted mutation of the RSA. The preprocessing includes excluding mutations with a number of residues exceeding a threshold.

13. The method of claim 9, wherein the method is based on the analysis of the pathogenicity of the predicted mutation of the RSA. The preprocessing includes excluding mutations that cannot be mapped to a structure due to isomer inconsistency.

14. The method of claim 9, wherein the method is based on the analysis of the pathogenicity of the predicted mutation of the RSA. The preprocessing includes excluding mutations with "pathogenic interpretation contradictory" clinical significance.

15. The method of claim 1, wherein the method is based on the analysis of RSA prediction of pathogenicity of mutations. The protein structure of the gene to be predicted is obtained by inputting the gene to be predicted into a protein structure generation model.

16. The method of claim 15, wherein the method is based on the analysis of RSA prediction of pathogenicity of mutations. The protein structure generation model comprises one or more of the following models: AlphaFold, AlphaFold2, ProteoGAN.

17. The method of claim 15, wherein the method is based on the analysis of RSA prediction of pathogenicity of mutations. The protein structure generation model is AlphaFold2.

18. The method of claim 1, wherein the method is based on the analysis of RSA prediction of pathogenicity of mutations. The gene family analyzer analyzes the gene family to which the mutation of the gene to be predicted belongs by a conserved domain, and the gene family analyzer comprises any one or more of the following software: TBtools, HMM, Motif.

19. The method of claim 18, wherein the method is based on the analysis of RSA prediction of pathogenicity of mutations. The gene family analyzer is TBtools or HMM.

20. A device for predicting the pathogenicity of mutations based on RSA, characterized in that, The device comprises a memory and a processor; the memory is used to store program instructions; the processor is used to call the program instructions, and when the program instructions are executed, it is used to execute the method for analyzing the pathogenicity of the mutation based on RSA according to any one of claims 1-19.

21. A system for analyzing the pathogenicity of a mutation based on RSA prediction, comprising: The system comprises: An acquisition module is configured to acquire a mutation of a gene to be predicted. A structure generation module is configured to map the acquired mutation of the gene to be predicted onto a protein structure of the gene to be predicted to obtain a protein structure with the mutation mapped thereon. A feature processing module is configured to process the protein structure with the mutation mapped thereon to obtain a mutation protein structure feature. A gene family generation module is configured to input the mutation of the gene to be predicted into a predictor to obtain a gene family to which the mutation of the gene to be predicted belongs. A classification module is configured to select an RSA in the mutation protein structure feature for prediction when the gene family to which the mutation of the gene to be predicted belongs is an ion channel to obtain a classification result of the mutation being a pathogenic mutation or a benign mutation.

22. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the method for analyzing the pathogenicity of the mutation based on RSA according to any one of claims 1-19.

Citation Information

Patent Citations

  • Genetic variation pathogenicity prediction method and device, storage medium and computer equipment

    CN114300036A

  • Method for identifying deleterious genetic mutations

    WO2022198510A1