Method, system and device for predicting pathogenicity of mutations based on sigma

The SIGMA model integrates gene data and protein 3D structure information, and uses machine learning algorithms to analyze the pathogenicity of missense mutations. This solves the problem of insufficient utilization of protein structure information in existing technologies and achieves higher accuracy in predicting the pathogenicity of mutations and assisting in diagnosis.

CN115691666BActive Publication Date: 2026-02-10PEKING UNION MEDICAL COLLEGE HOSPITAL
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211290686.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-21
Publication Date
2026-02-10
Estimated Expiration
2042-10-21

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively utilize protein structure information to assess the pathogenicity of missense mutations, resulting in inaccurate analysis of the pathogenicity of gene data mutations.

Method used

The SIGMA model is used to integrate high-quality gene data and protein 3D structure information. Machine learning algorithms such as Gradient Boosting Machine (GBM) are used to predict the pathogenicity of mutations. Feature extraction and preprocessing techniques such as SpliceAI algorithm are combined to optimize the model and improve the accuracy of analysis.

Benefits of technology

It significantly improves the accuracy and depth of pathogenicity analysis of missense mutations, enabling more accurate assistance in the diagnosis and treatment selection of disease development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115691666B_ABST
    Figure CN115691666B_ABST
Patent Text Reader

Abstract

The present application relates to a SIGMA-based method, system and device for predicting the pathogenicity of mutations. The method comprises the following steps: obtaining mutations of a gene to be predicted; mapping the mutations of the gene to be predicted onto a protein structure of the gene to be predicted to obtain a protein structure with the mutations mapped; extracting features of the protein structure with the mutations to obtain mutation protein structure features; and inputting the mutation protein structure features into a SIGMA model to obtain a classification result of whether the mutations are pathogenic mutations or benign mutations. The mutation protein structure features include protein level features, residue level features and mutation level features. The method predicts the pathogenicity of mutations based on the SIGMA model and mutation protein structure features, explores the high prediction capability and potential application value of the pathogenicity of mutations, and has a beneficial promoting effect on the analysis and research of the pathogenicity of mutations of gene data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of gene data analysis, and more specifically, to an analytical method, apparatus, system, computer-readable storage medium, and application thereof for predicting the pathogenicity of mutations based on SIGMA. Background Technology

[0002] Interpreting human genetic variations relies heavily on predicting their effects on proteins. For example, protein truncation mutations often lead to loss of protein function and may be classified as pathogenic mutations. Most clinically relevant variations are missense variations, which have varying degrees of impact on protein structure and function.

[0003] Due to the vast number of possible missense mutations (approximately 76 million) in the human exome, characterizing every single one in vivo or in vitro is impractical. To computationally assess the pathogenicity of missense variants, numerous advanced variant effect prediction (VEP) algorithms, such as SIFT, REVEL, and EVE, have been developed. Features typically used to train VEPs include mutation frequency, evolutionary conservation, and the physicochemical properties of amino acids. Integrating results from multiple VEPs, in addition to using a single source of information, can significantly improve performance. Given the close relationship between protein structure and function, the structural context of a mutation represents promising information independent of mutation frequency and evolutionary conservation. However, the scarcity of high-resolution 3D protein structures hinders the development of VEPs that utilize protein structural information.

[0004] Recent advances in computer technology have enabled breakthroughs in protein structure modeling. Among the most successful achievements, the AlphaFold2 protein structure generation model predicts 3D protein structures with near-experimental accuracy, expanding the structural coverage of the human proteome from 17% to 98.5%. Extensive and accurate structural information holds unprecedented potential to help predict mutation effects. Therefore, using protein structural information to assess the pathogenicity of mutations will significantly advance research on the pathogenicity analysis of gene data mutations. Summary of the Invention

[0005] The method of this invention integrates mutation information of the gene to be predicted and 3D structure information of wild-type proteins based on a high-quality dataset, and is trained using a machine learning algorithm. It deeply mines the variation features in gene data, and then assesses the pathogenicity of mutations, solving related life science problems. It has strong innovation in the field of life sciences and will play an important role in promoting the analysis and research of the pathogenicity of mutations in gene data.

[0006] This application discloses a method for constructing a SIGMA model, including:

[0007] Obtain a dataset of mutations and genes, wherein the mutation data includes labeling information;

[0008] The gene is input into a protein structure generation model to obtain the protein structure;

[0009] The mutation is mapped onto the protein structure to obtain the protein structure mapped with the mutation;

[0010] Feature extraction is performed on the mapped mutated protein structure to obtain the mutated protein structure features;

[0011] The structural features of the mutant protein are processed using machine learning algorithms to obtain the predicted classification results;

[0012] Based on the predicted classification results and the actual results, the machine learning algorithm is optimized to obtain the SIGMA model;

[0013] The SIGMA (Structure-Informed Germline Missense mutation Assessor) model is developed based on the close relationship between protein structure and function. It uses a gradient enhancement machine algorithm to evaluate the impact of missense variants in the context of protein structure.

[0014] Optionally, the machine learning algorithm includes any one or more of GBM, SVM, LSTM, Transformer, Naive Bayes, and Bayesian Belief Network.

[0015] This application also discloses a pathogenicity analysis method based on SIGMA-predicted mutations, including:

[0016] Obtain the mutation of the gene to be predicted, wherein the gene to be predicted includes wild-type genes;

[0017] The gene to be predicted is input into the protein structure generation model to obtain the protein structure of the gene to be predicted.

[0018] The mutation of the gene to be predicted is mapped onto the protein structure of the gene to be predicted, and the protein structure mapped with the mutation is obtained.

[0019] Feature extraction is performed on the mapped mutated protein structure to obtain the mutated protein structure features;

[0020] The structural features of the mutant protein are input into the SIGMA model to obtain classification results of whether the mutation is pathogenic or benign.

[0021] Furthermore, the mutation or gene dataset includes one or more of the following datasets: dbSNP, dbVar, gnomAD, ExAC, RefSeq; the mutation dataset is preprocessed, which includes normalizing the mutation dataset to obtain mutation data;

[0022] Optionally, the preprocessing employs one or more of the following methods: Optionally, an algorithm is used to predict the impact of mutations on splicing and to exclude mutations with scores exceeding a threshold; Preferably, the SpliceAI algorithm is used to predict the impact of mutations on splicing and to exclude mutations with SpliceAI scores greater than 0.5.

[0023] Optionally, the preprocessing includes excluding mutations with a residue count exceeding a threshold;

[0024] Optionally, the preprocessing includes excluding mutations that cannot be mapped to a structure due to transcript inconsistency;

[0025] Optionally, the pretreatment includes excluding mutations with clinical significance that are “conflicting in their pathogenicity explanations”.

[0026] Furthermore, the protein structure generation model includes any one or more of the following generation models: AlphaFold, AlphaFold2, and ProteoGAN; preferably, the protein structure generation model is AlphaFold2.

[0027] Furthermore, the structural features of the mutant protein include protein-level features, residue-level features, and mutation-level features;

[0028] Optionally, the protein level characteristics include thermodynamic characteristics and protein volume characteristics;

[0029] Optionally, the residue-level features include secondary structure-related features and the relative solvent accessibility of the mutated residues, used to characterize the structural background of the mutated residues;

[0030] Optionally, the mutation level characteristics include thermodynamic characteristics and characteristics derived from empirical rules for determining whether a mutation has a significant effect on protein structure.

[0031] Optionally, the thermodynamic characteristics include van der Waals collisions, hydrogen bond energy, side chain entropy, van der Waals contribution, solvation pole, electrostatic interaction, energy ionization, and free energy.

[0032] Optionally, the protein volume characteristics include total volume, void volume, and van der Waals volume;

[0033] Optionally, the secondary structure-related features include eight features generated using the DSSP program: DSSP(H), DSSP(E), DSSP(G), DSSP(I), DSSP(T), DSSP(S), and DSSP(C).

[0034] Optionally, the features derived from the empirical rules for determining whether a mutation has a significant effect on protein structure include glycine bending, proline in the α-helix, substitution of residues in the α-helix by proline / glycine, energy charge loss, energy charge loss, energy hydrophobicity replacement, outlier substitution, disulfide bond breakage, buried salt bridge breakage, buried proline, energy replacement of glycine, the unfolding free energy of the protein, and the difference in unfolding free energy between the mutant protein.

[0035] A pathogenicity analysis device based on SIGMA-predicted mutations includes: a memory and a processor;

[0036] The memory is used to store program instructions;

[0037] The processor is used to call program instructions, which, when executed, are used to perform the pathogenicity analysis method based on SIGMA-predicted mutations.

[0038] A pathogenicity analysis system based on SIGMA-predicted mutations includes:

[0039] An acquisition module is used to acquire mutations in a gene to be predicted, the gene to be predicted including wild-type genes;

[0040] The structure generation module maps the mutation of the gene to be predicted to the protein structure of the gene to be predicted, thereby obtaining the protein structure mapped with the mutation.

[0041] The feature processing module extracts features from the mapped mutated protein structure to obtain the mutated protein structure features;

[0042] The classification module inputs the structural features of the mutant protein into the SIGMA model to obtain classification results of whether the mutation is a pathogenic mutation or a benign mutation.

[0043] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described pathogenicity analysis method based on SIGMA-predicted mutations.

[0044] Any of the following applications:

[0045] Application of the aforementioned equipment or systems in the selection of auxiliary mutation pathogenicity analysis schemes;

[0046] Application of the aforementioned equipment or systems in predicting the pathogenicity of mutations;

[0047] Application of the aforementioned equipment or systems in classifying gene mutation samples or predicting sample attributes;

[0048] The application of the above-mentioned equipment or systems in the diagnosis of the occurrence and development of pathogenic mutations.

[0049] This invention predicts the pathogenicity of mutations based on the SIGMA model and the structural characteristics of mutant proteins, aiming to explore its high predictive ability and potential application value for the pathogenicity of variants, and to promote the research on the pathogenicity analysis of gene data mutations.

[0050] Advantages of this application:

[0051] 1. This application innovatively discloses a new method for analyzing the pathogenicity of mutated genes. This method is based on the SIGMA model to deeply mine the life laws hidden behind gene data. Based on the characteristics of mutation mapping on protein structure, it comprehensively analyzes the mutation classification results, thereby improving the accuracy and depth of data analysis.

[0052] 2. This application innovatively preprocesses the mutation dataset, which includes normalizing the mutation data, using the SpliceAI algorithm to predict the impact of mutations on splicing and excluding mutations with scores greater than 0.5, and excluding mutations with residue counts exceeding a threshold, thus making full use of the mutation data;

[0053] 3. This application creatively discloses a pathogenicity analysis device and system based on SIGMA-predicted mutations. Through in-depth interpretation of sample gene data, this application can be more accurately applied to fields such as auxiliary diagnosis of disease occurrence and development related to gene data, and auxiliary selection of treatment plans. Attached Figure Description

[0054] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0055] Figure 1 This is a schematic flowchart of pathogenicity analysis based on SIGMA-predicted mutations provided in an embodiment of the present invention;

[0056] Figure 2 This is a protein structure feature map based on AlphaFold2 that maps to mutations, provided in an embodiment of the present invention.

[0057] Figure 3 This is a prediction effect evaluation chart based on SIGMA pathogenicity analysis provided in an embodiment of the present invention;

[0058] Figure 4 This is a Spearman correlation analysis statistical graph based on SIGMA-predicted mutations provided in an embodiment of the present invention;

[0059] Figure 5 This is a statistical chart of the importance features of predicted mutations based on SIGMA, provided in an embodiment of the present invention.

[0060] Figure 6 This is a statistical chart of classification results based on SIGMA-predicted mutations provided in an embodiment of the present invention;

[0061] Figure 7 This is a schematic diagram of a pathogenicity analysis device based on SIGMA-predicted mutations provided in an embodiment of the present invention;

[0062] Figure 8 This is a schematic diagram of a pathogenicity analysis system based on SIGMA prediction of mutations provided in an embodiment of the present invention. Detailed Implementation

[0063] To enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0064] In some of the processes described in the specification, claims, and accompanying drawings of this invention, multiple operations appearing in a specific order are included. However, it should be clearly understood that these operations may not be performed in the order they appear herein, or may be performed in parallel. The operation numbers, such as S101, S102, etc., are merely used to distinguish different operations and do not themselves represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be performed sequentially or in parallel.

[0065] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0066] Figure 1 This is a schematic flowchart of a pathogenicity analysis method based on SIGMA prediction of mutations provided by an embodiment of the present invention. Specifically, the method includes the following steps:

[0067] S101: Obtain the mutation of the gene to be predicted;

[0068] In one embodiment, obtaining mutations in the gene to be tested includes publicly available genetic data. Optionally, the mutation information of the gene to be tested also includes genetic information of the sample obtained by high-throughput sequencing. Optionally, the mutations in the gene to be predicted include one or more of the following datasets: dbSNP, dbVar, gnomAD, RefSeq, and ExAC, where dbSNP, dbVar, and RefSeq are all from the ClinVar database. A total of 27,165 benign and 22,957 pathogenic missense variants were retrieved from the gnomAD and ClinVar databases. For ClinVar mutations, single nucleotide missense mutations with an audit status of at least one star were retained (audited by an expert panel: providing criteria, multiple submitters without conflict; providing criteria, single submitter).

[0069] Furthermore, the preprocessing of mutation data includes one or more of the following methods: using an algorithm to predict the impact of mutations on splicing and excluding mutations with scores exceeding a threshold; preferably, using the SpliceAI algorithm to predict the impact of mutations on splicing and excluding mutations with SpliceAI scores greater than 0.5.

[0070] SpliceAI predicts splicing changes caused by single nucleotide variations, accurately predicting splicing sites (location and probability of abnormal splicing) from any mRNA precursor sequence, thus enabling the prediction of hidden splicing caused by mutations in non-coding RNA regions.

[0071] Optionally, the pretreatment also includes excluding mutations with more than a threshold number of residues; for example, proteins with more than 2,700 residues are excluded.

[0072] Optionally, pretreatment includes excluding mutations with clinical significance that presents conflicting explanations for their pathogenicity;

[0073] Optionally, preprocessing includes excluding mutations that cannot be mapped to the structure due to isomer inconsistencies;

[0074] In one embodiment, the mutated gene comprised 20,047 pathogenic (red) and 20,148 benign (blue) variants, and also included 27,928 mutations in six proteins characterized by the DMS experimental system, using 27 computational VEPs, such as EVE screening and DEOGEN2. The EVE score for each mutation was derived from predictions of other VEPs retrieved from the dbNSFP database (version 4.1a). These VEPs were divided into two categories: single predictors independent of other VEPs (n=16) and meta-predictors that integrated the results of other VEPs into input features (n=11).

[0075] S102: Map the mutation of the gene to be predicted to the protein structure of the gene to be predicted, and obtain the protein structure mapped with the mutation.

[0076] In one embodiment, the protein structure of the gene to be predicted is obtained by inputting the gene to be predicted into a protein structure generation model. Preferably, the protein structure of the gene to be predicted is the wild-type gene protein structure. Furthermore, the protein structure generation model refers to a model or software capable of generating the protein structure of a gene or polypeptide based on the amino acid sequence of the gene sequence or polypeptide. Optionally, the protein structure generation model includes one or more of the following generation models: AlphaFold, AlphaFold2, ProteoGAN, and VAE; preferably, the protein structure generation model is AlphaFold2.

[0077] AlphaFold accelerated and enabled large-scale discoveries, including deciphering the structure of the nuclear pore complex. AlphaFold predictions, with very high confidence, provided a nine-helix topology, offering clues about protein function that are crucial for unraveling the causes of rare genetic diseases.

[0078] AlphaFold2 discovered a novel protein-gated mechanism, glucose-6-phosphatase, identified enzymes that inhibit enzyme binding sites, and Wolframin, a transmembrane protein located in the ER. It improved prediction accuracy to the atomic level, enabling faster and more precise identification of enzyme active sites, with 350,000 predicted results. Currently, the AlphaFold database contains approximately 365,000 structure predictions.

[0079] ProteoGAN, a conditional generative adversarial network, outperforms classical and recent deep learning baselines in protein sequence generation. The main application of ProteoGAN is to expand protein screening with candidates that are farther from the known sequence space than previously possible, but whose relatively novel candidates are more likely to be functional than those of other methods.

[0080] VAE (Variational Autoencoder) is a generative model that is based on joint probabilities.

[0081] Specifically, the protein structure dataset for the protein structure generation model includes the AlphaFold Protein Structure Database (AlphaFoldDB). AlphaFold2 only provides fragments of 1400 amino acids for these proteins; proteins with more than 2700 residues are excluded, as are mutations that cannot be mapped to a structure due to isoform inconsistencies.

[0082] In a specific example, the wild-type gene is input into the AlphaFold2 model to obtain the wild-type gene protein structure. Then, the mutation of the gene to be predicted is mapped onto the wild-type gene protein structure to obtain the protein structure mapped to the mutation. The protein structure mapped to the mutation is a multi-dimensional structure.

[0083] S103: Extract features from the protein structure mapped with mutations to obtain the structural features of the mutated protein;

[0084] In one embodiment, 57 features were extracted from the protein structure mapped with mutations; the 57 extracted features included protein-level features, residue-level features, and mutation-level features, wherein pathogenic mutations exhibited greater changes in protein stability than benign mutations.

[0085] Optionally, protein-level features include 16 thermodynamic features and 3 protein volume features; among them, the thermodynamic features include 15 components such as total unfolding energy and its van der Waals collisions, hydrogen bond energy, and side chain entropy; the protein volume features include total volume, interstitial volume, and van der Waals volume.

[0086] Optionally, residue-level features include eight secondary structure-related features and the relative solvent accessibility (RSA) of the mutated residue, used to characterize the structural background of the mutated residue. For each mutated residue, its secondary structure is assigned using the DSSP (Dict. DSSP) program, generating eight secondary structure types: DSSP(H), DSSP(E), DSSP(G), DSSP(I), DSSP(T), DSSP(S), and DSSP(C). RSA quantifies the exposure status of the residue; mutations in hidden residues (i.e., low RSA residues) are more likely to be pathogenic than mutations in exposed residues (i.e., high RSA residues).

[0087] Optionally, mutation level features include 16 thermodynamic features describing changes in protein stability after mutation and 13 features derived from empirical rules for determining whether a mutation has a significant effect on protein structure; the 16 thermodynamic features include the free energy difference between the mutant protein and its 15 components; the other 13 features are a combination of the structural background and physicochemical properties of the mutant / mutant residues, including glycine bending, proline in the α-helix, substitution of residues in the α-helix by proline / glycine, cysteine ​​residues, energy charge loss, energy charge loss, energy hydrophobicity replacement, outlier substitution, disulfide bond breakage, buried salt bridge breakage, buried proline, energy displacement of glycine, protein unfolding free energy, and the unfolding free energy difference between the mutant protein and the mutant protein.

[0088] Figure 2This study demonstrates the wild-type protein structures generated using the AlphaFold2 generative model. Mapping mutations onto wild-type protein structures yields mutant protein structural features, which are correlated with pathogenicity, particularly the relative solvent accessibility (RSA) and free energy difference (ΔΔG) between mutant and wild-type proteins. Figure 2 As shown in A, the RSA of pathogenic variants was significantly lower than that of benign variants (P<2.2e-16, Mann-Whitney U test), consistent with the fact that most proteins are less tolerant of hidden mutations than exposed mutations. Figure 2 As shown in B, ΔΔG measures the effect of a single amino acid substitution on protein stability, with pathogenic variants exhibiting greater changes in protein stability than benign variants (P < 2.2e-16, Mann-Whitney U test). Figure 2 As shown in Figure C, as expected, disulfide bond breaking had the highest positive predictive value among all structural features (odds ratio [OR] = 93.8, 95% confidence interval [CI] = 44.5–198, P < 2.2e-16, Pearson chi-square test). Almost all (98.72%) missense mutations that broke disulfide bonds were pathogenic, further supporting the important role of disulfide bonds in protein function. Figure 2 As shown in D, the association between the type of secondary structure in which the mutation occurs and its pathogenicity is consistent with prior knowledge. Mutations in the circular or irregular extensions of proteins (Dictionary of Protein Secondary Structures [DSSP]-C) are often benign (OR=0.32, 95%CI=0.31). 0.34, P<2.2e 16. Pearson chi-square test). In contrast, mutations in regular secondary structures tend to be pathogenic, especially mutations in the α-helix (DSSP-H; OR=1.73, 95%CI=1.66-1.79, P<2.2e-16, Pearson chi-square test) or β-sheet (DSSPE; OR=1.97, 95%CI=1.87-2.08, P<2.2e-16, Pearson chi-square test).

[0089] S104: Input the structural features of the mutant protein into the SIGMA model to obtain the classification results.

[0090] In one embodiment, the machine learning algorithm employs one or more of GBM, SVM, LSTM, Transformer, Naive Bayes, and Bayesian Belief Network, and the classifier can be any of the following models: Random Forest, Decision Tree, Logistic Regression, Support Vector Machine, and Neural Network; no limitation is imposed here. Figure 3As shown, the constructed SIGMA model uses Gradient Boosting Machine (GBM) to train a classifier that distinguishes between pathogenic and benign variants, while using grid search for parameter tuning.

[0091] Specifically, such as Figure 3 The performance evaluation plot shown uses 40,195 mutations with structural features as the training dataset and employs a SIGMA model based on gradient boosting machine (GBM) to predict the pathogenicity of missense mutations. Figure 3 As shown in Figure A, the SIGMA score distribution for 20,047 pathogenic (red) and 20,148 benign (blue) variants in the training set, where the SIGMA scores are quantitative SIGMA scores (from 0 to 1) calculated using out-of-fold prediction, shows a bimodal distribution, indicating that SIGMA has strong discriminative power. For the training set, Figure 3 The area under the ROC curve (AUC) of receiver operating characteristics shown in B is 0.944 (95% CI = 0.942 - 0.946), indicating an incorrect prediction. Consistently, as... Figure 3 As shown in C, high prediction accuracy was achieved on the test set (AUC=0.933, 95%CI=0.928-0.938). Figure 3 As shown in D, when comparing the AUC of SIGMA and 16 individual variation effect predictors (VEPs) using the test set, SIGMA outperformed all individual VEPs, with AUCs ranging from 0.779 (FATHMM) to 0.929 (MutPred). Figure 3 E represents the correlation between VEP and Deep Mutation Scan (DMS) measurements, where Spearman's correlation is calculated between the functional score of the DMS experiment, the predicted score of SIGMA, and 16 individual VEPs. The closest competitor is DEOGEN2 (overall rho=0.387, Spearman correlation analysis; as shown) Figure 3 The EVE predictor, which contains extensive heterogeneous information, improved its performance ranking from sixth on the labeled dataset to third on the DMS dataset (e.g., EVE). Figure 3 E).

[0092] Furthermore, to limit data recurrence that could potentially overstate predictor performance, a deep mutation scanning DMS independent of the labeled dataset was used to assemble a test dataset for six proteins (BRCA1, P53, MSH2, PTEN, VKORC1, and HRAS). The study used high-throughput functional assays to quantify the functional impact of all possible mutations in the proteins / protein domains with continuous scoring, preserving a high-quality DMS study of 28,293 mutations in these six proteins. Figure 4As shown, SIGMA and DMS scores were performed on the six proteins, along with the correlation (Spearman correlation analysis) with high-quality DMS studies. The color of each hexagon represents the number of mutations. The results indicate that the SIGMA model based on the DMS dataset has the highest correlation among all individual predictor variables (overall rho=0.419, Spearman correlation analysis), especially for BRCA1 (rho=0.286), PTEN (rho=0.519), and HRAS (rho=0.404). Figure 3 E and Figure 4 ).

[0093] Figure 5 This demonstrates an importance feature analysis based on the SIGMA model, from... Figure 5 As can be seen from A, the ten most important features affecting the discriminative power of SIGMA are: residue-level feature RSA contributes the most to the discriminative power of SIGMA, followed by two mutation-level features (ΔΔG and ΔVander). Seven of these ten most important features are protein-level features, demonstrating the important role of protein stability in predicting the pathogenicity of variants. Figure 5 B and Figure 5 C represents the gene set enrichment analysis (GSEA) plots for ion channel genes (B) and collagen family genes (C), respectively. Notably, RSA alone can predict the pathogenicity of variants in 76 out of 200 genes with an accuracy >0.8 (with at least 10 pathogenic variants and 10 benign variants used for evaluation). Among these, RSA showed the highest classification accuracy for ion channel genes relative to solvent accessibility (P=8.67e-04); however, RSA alone performed unsatisfactorily for 11% of the genes (21 out of 200 genes, accuracy <0.5), which are rich in the collagen gene family (P=1.99e-06). Figure 5 (C). Additionally, ΔG, the difference in unfolding free energy between the mutant protein and the mutated protein; ΔG, the unfolding free energy of the protein; VdW, van der Waals.

[0094] In one embodiment, a final labeled dataset containing 27,165 benign (negative) and 22,957 pathogenic (positive) missense mutations is used as the "gold standard" dataset. The dataset is divided into 80% for training and 20% for testing. The gnomAD database retains 27,928 mutations for model evaluation. Training the SIGMA model can be performed as follows: obtaining labeled samples of different categories from the dataset (e.g., pathogenic / potentially pathogenic variants are labeled as positive, while benign / potentially benign variants are labeled as negative), repeating the protein structure generation and feature extraction processes described above to obtain a feature set, inputting the feature set into the SIGMA model to obtain the sample classification results, calculating the loss between the classification results and the true values ​​using a loss function, performing backpropagation, updating the parameters using an optimizer, and obtaining the trained SIGMA model.

[0095] Furthermore, the labeling information was defined as follows: pathogenic / potentially pathogenic variants were labeled positive, while benign / potentially benign variants were labeled negative. For gnomAD mutations, preprocessing was performed using SpliceAI to predict the gnomAD database (for deep intron regions: distance from the exon-intron boundary >50 nt). Due to the difficulty of deep intron prediction, only 56% deletions were observed when the threshold was set to 0.8. Therefore, the impact of mutations on splicing was predicted, and mutations with SpliceAI scores greater than 0.5 were excluded. Mutations introducing mystery splicing sites primarily affect mRNA splicing rather than protein structure, and selected common missense mutations (the highest allele frequency in all populations with gnomAD > 0.05) were labeled negative. In addition, the GOF / LOF database labeled 193 gain-of-function (GoF) and 921 loss-of-function (LoF) pathogenic missense variants.

[0096] In one embodiment, the SIGMA model is evaluated based on genomic data type to assess the association between the aforementioned structural features and variant pathogenicity. For dichotomous features, contingency tables are constructed, and the chi-square test is used to determine whether an association exists between the two dichotomous variables. For continuous features, logistic regression analysis is used to examine the association between the feature and variant pathogenicity. The strength of the association between the feature and variant pathogenicity is quantified using the odds ratio (OR) and 95% confidence interval (CI). All statistical analyses and data visualizations are performed using R software packages: h2o, caret, pROC, forestplot, ggpubr, ggsci, viridis, and cutpointr, where a p-value <0.05 is considered statistically significant.

[0097] In one embodiment, the above method is applied to the selection of an auxiliary gene mutation pathogenicity analysis scheme, and the classification result of whether the mutation is pathogenic is predicted based on the gene information of the sample. Figure 6 This is a classification result diagram of pathogenicity of mutations based on SIGMA prediction provided in an embodiment of the present invention. The classification result refers to the classification result of the predicted mutation as a pathogenic mutation or a benign mutation. Figure 6 The classification process illustrated uses a training dataset of 40,195 mutations with structural features, predicting the pathogenicity of missense mutations based on the Gradient Boosting Machine (GBM) in the SIGMA model. Based on the SIGMA predictions from the training set, we examine its performance using different thresholds, finding that the optimal threshold for determining variant pathogenicity is 0.498, achieving an accuracy of 85%, sensitivity of 88%, and specificity of 83%. Figure 6 As can be seen, the quantitative SIGMA score (from 0 to 1) for the potential pathogenicity of the variant shows a bimodal distribution, indicating that SIGMA has a strong discriminative ability.

[0098] Figure 7 This invention provides a pathogenicity analysis device based on SIGMA-predicted mutations. The device includes a memory and a processor; the device may also include an input device and an output device.

[0099] Processors, memory, input devices, and output devices can be connected via a bus or other means. Figure 7 Taking the bus connection method as an example;

[0100] Memory is used to store program instructions;

[0101] The processor is used to call program instructions, which, when executed, are used to perform the pathogenicity analysis method based on SIGMA-predicted mutations.

[0102] Figure 8 This invention provides a pathogenicity analysis system based on SIGMA-predicted mutations, comprising:

[0103] The acquisition module S801 is used to acquire the mutations of the gene to be predicted;

[0104] The structure generation module S802 maps the mutation of the gene to be predicted to the protein structure of the gene to be predicted, thereby obtaining the protein structure mapped with the mutation; wherein, the preferred gene to be predicted includes a wild-type gene.

[0105] The feature processing module S803 extracts features from the protein structure mapped with mutations to obtain the structural features of the mutated protein.

[0106] The classification module S804 inputs the structural features of the mutant protein into the SIGMA model to obtain the classification results of whether the mutation is a pathogenic mutation or a benign mutation.

[0107] The present invention provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the above-described pathogenicity analysis method based on SIGMA-predicted mutations.

[0108] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and modules described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0109] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between apparatuses or modules, and may be electrical, mechanical, or other forms.

[0110] The modules described as separate components may or may not be physically separate. Similarly, the components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0111] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.

[0112] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc.

[0113] Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware, and the program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0114] The computer device provided by the present invention has been described in detail above. For those skilled in the art, there will be changes in the specific implementation and application scope based on the ideas of the embodiments of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A pathogenicity analysis method based on SIGMA-predicted mutations, comprising: Obtain the mutations of the gene to be predicted; preprocess the mutations, the preprocessing including using the SpliceAI algorithm to predict the effect of mutation splicing and excluding mutations with a SpliceAI score greater than 0.5; The gene to be predicted is input into a protein structure generation model to obtain the protein structure of the gene to be predicted; the protein structure generation model is AlphaFold2. The mutation of the gene to be predicted is mapped onto the protein structure of the gene to be predicted, resulting in a protein structure mapped with the mutation; the features of the mutated protein structure include protein-level features, residue-level features, and mutation-level features; The protein-level characteristics include thermodynamic characteristics and protein volume characteristics; The residue-level features include secondary structure-related features and relative solvent accessibility of mutated residues, used to characterize the structural background of mutated residues; The mutation level characteristics include thermodynamic characteristics and characteristics derived from empirical rules for determining whether a mutation has a significant effect on protein structure. Feature extraction is performed on the mapped mutated protein structure to obtain the mutated protein structure features; The structural features of the mutant protein are input into the SIGMA model to obtain classification results of whether the mutation is pathogenic or benign. SIGMA model: Obtain a dataset of mutations and genes, where the mutation data includes labeling information; The gene is input into a protein structure generation model to obtain the protein structure; The mutation is mapped onto the protein structure to obtain the protein structure mapped with the mutation; Feature extraction is performed on the mapped mutated protein structure to obtain the mutated protein structure features; The structural features of the mutant protein are processed using machine learning algorithms to obtain the predicted classification results; Based on the predicted classification results and the actual results, the machine learning algorithm is optimized to obtain the SIGMA model; The machine learning algorithm is GBM.

2. The pathogenicity analysis method based on SIGMA-predicted mutations according to claim 1, characterized in that, The preprocessing also includes normalizing the mutations and inputting them into the protein structure generation model.

3. The pathogenicity analysis method based on SIGMA-predicted mutations according to claim 1, characterized in that, The preprocessing also includes excluding mutations with a residue count exceeding a threshold.

4. The pathogenicity analysis method based on SIGMA-predicted mutations according to claim 1, characterized in that, The preprocessing also includes excluding mutations that cannot be mapped to the structure due to isomer inconsistencies.

5. The pathogenicity analysis method based on SIGMA-predicted mutations according to claim 1, characterized in that, The pretreatment also includes excluding mutations with clinical significance that are "conflicting in their pathogenicity explanations".

6. The pathogenicity analysis method based on SIGMA-predicted mutations according to claim 1, characterized in that, The process involves inputting the gene to be predicted into a protein structure generation model to obtain the protein structure of the gene to be predicted. Furthermore, the gene to be predicted includes a wild-type gene.

7. The pathogenicity analysis method based on SIGMA-predicted mutations according to claim 1, characterized in that, The protein structure of the gene to be predicted is the wild-type protein structure.

8. The pathogenicity analysis method based on SIGMA-predicted mutations according to claim 1, characterized in that, The thermodynamic characteristics include van der Waals collisions, hydrogen bond energy, and side chain entropy.

9. The pathogenicity analysis method based on SIGMA-predicted mutations according to claim 1, characterized in that, The protein volume characteristics include total volume, interstitial volume, and van der Waals volume.

10. The pathogenicity analysis method based on SIGMA-predicted mutations according to claim 1, characterized in that, The secondary structure-related features include features generated using the DSSP program, including: DSSP-H, DSSP-E, DSSP-G, DSSP-I, DSSP-T, DSSP-S, and DSSP-C.

11. A pathogenicity analysis device based on SIGMA-predicted mutations, characterized in that, The device includes: A memory and a processor; the memory is used to store program instructions; the processor is used to invoke the program instructions, which, when executed, are used to perform the pathogenicity analysis method based on SIGMA prediction of mutations as described in any one of claims 1-10.

12. A pathogenicity analysis system based on SIGMA-predicted mutations, characterized in that, The system includes: An acquisition module is used to acquire mutations in the gene to be predicted; the mutations are preprocessed, the preprocessing including using the SpliceAI algorithm to predict the effect of mutation splicing and excluding mutations with a SpliceAI score greater than 0.5; The structure generation module maps the mutation of the gene to be predicted to the protein structure of the gene to be predicted, thereby obtaining the protein structure mapped with the mutation; the protein structure generation model is AlphaFold2. The feature processing module extracts features from the mapped mutated protein structure to obtain the mutated protein structure features; The structural features of the mutant protein include protein-level features, residue-level features, and mutation-level features; The protein-level characteristics include thermodynamic characteristics and protein volume characteristics; The residue-level features include secondary structure-related features and relative solvent accessibility of mutated residues, used to characterize the structural background of mutated residues; The mutation level features include thermodynamic features and features derived from empirical rules for determining whether a mutation has a significant impact on protein structure; the classification module inputs the mutated protein structure features into the SIGMA model to obtain classification results of whether the mutation is a pathogenic mutation or a benign mutation; the SIGMA model: acquires mutation and gene datasets, wherein the mutation data includes labeling information; The gene is input into a protein structure generation model to obtain the protein structure; The mutation is mapped onto the protein structure to obtain the protein structure mapped with the mutation; Feature extraction is performed on the mapped mutated protein structure to obtain the mutated protein structure features; The structural features of the mutant protein are processed using machine learning algorithms to obtain the predicted classification results; Based on the predicted classification results and the actual results, the machine learning algorithm is optimized to obtain the SIGMA model; the machine learning algorithm is GBM.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the pathogenicity analysis method based on SIGMA prediction of mutations as described in any one of claims 1-10.

Citation Information

Patent Citations

  • Prediction method and device for genetic diseases caused by gene mutation

    CN106960122A

  • Target identification method and device based on allosteric mechanism, and storage medium

    CN112397140A

  • Genetic variation pathogenicity prediction method and device, storage medium and computer equipment

    CN114300036A