Machine learning assisted integrated protein mutation phenotype prediction method and system
Through machine learning-assisted methods, the multi-dimensional characteristics of proteins are integrated to predict mutant phenotypes, which solves the shortcomings of existing methods in terms of accuracy, reliability and interpretability, and achieves a more comprehensive evaluation and prediction of protein mutation effects.
Patent Information
- Application Number
- CN202411891674.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-30
AI Technical Summary
The existing methods for predicting protein mutation phenotypes have limitations that rely on local sequence characteristics or structural information, lack of comprehensive evaluation of protein mutations' overall structure, function and network effects, and the inability to effectively integrate multi-dimensional features to predict. They have problems in accuracy and reliability and lack interpretability.
Using machine learning-assisted integrated protein mutant phenotype prediction method, we use homologous sequences to perform multi-sequence comparison, calculate Shannon entropy value and co-evolution coefficient, build an amino acid contact energy network and elastic network model, integrate sequence characteristics, structural characteristics and elastic network characteristics, and use machine learning methods to predict mutant phenotypes, and visualize and phenotype classification.
It improves the accuracy and reliability of protein mutation prediction, can study the effects of protein mutations from multiple angles, provide more comprehensive phenotypic prediction results, and provide biological mechanism mining information through interpretable machine learning methods.
Smart Images

Figure CN120072047A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of protein prediction, and specifically to a machine learning-assisted integrated protein mutation phenotype prediction method and system. Background Art
[0002] The relationship between protein mutations and phenotypes is crucial in biological and medical research. Especially in the research of disease occurrence, development, and treatment, the type, location of mutations, and their effects on protein function play a decisive role. However, the existing technologies mainly have the following problems:
[0003] Existing protein mutation analysis tools are generally used to predict relatively general problems such as pathogenicity, while our tool is used to predict the phenotypic effects of protein pathogenic mutations.
[0004] Traditional tools mainly use information such as sequence conservation. These tools generally only consider information at the amino acid sequence structure level. Our tool integrates network topology and kinetic parameters, etc. It not only considers the tertiary structure of proteins and protein dynamics, etc., but also improves the prediction accuracy and can study the effects of protein mutations from more perspectives.
[0005] Traditional tools analyze multiple proteins, while we analyze a single protein, which means obtaining more independent biological mechanism information for a single protein.
[0006] Many current tools use deep learning frameworks, with poor interpretability, and can only give a prediction result without mechanism mining. While the interpretable machine learning we use can perform global and individual interpretations, and can provide more information at the level of biological mechanism mining. Summary of the Invention
[0007] In view of the above existing problems, the present invention is proposed.
[0008] Therefore, the technical problem solved by the present invention is that existing protein mutation phenotype prediction methods have limitations in relying on local sequence features or structural information, lack a comprehensive assessment of the overall structure, function, and network effects of protein mutations, and are insufficient in effectively integrating multi-dimensional features for prediction. They have problems in terms of accuracy and reliability, and existing methods lack interpretability.
[0009] To solve the above technical problems, the present invention provides the following technical solution: A machine learning-assisted integrated protein mutation phenotype prediction method, including:
[0010] Obtain the homologous sequences of the target protein, perform multiple sequence alignment on the homologous sequences, and obtain the sequence information after alignment; based on the sequence information after alignment, perform conservation scoring on each amino acid site to evaluate the conservation degree of each site during evolution; based on the conservation scoring, calculate the Shannon entropy value and co-evolution coefficient of the target protein to characterize the mutation frequency and co-evolution relationship of amino acid sites; use the Shannon entropy value and co-evolution coefficient to calculate the free energy change of the target protein to evaluate the impact of mutations on protein stability; according to the free energy change, construct an amino acid contact energy network and calculate the topological features in the network; based on the amino acid contact energy network, calculate the relative accessible area of each site to evaluate the impact of mutations on the surface-exposed region of the protein; construct an elastic network model of the target protein to simulate the impact of mutations on protein dynamic stability; integrate the sequence features, structural features, and elastic network features of the protein to form a comprehensive feature set; based on the comprehensive feature set, use machine learning methods to predict the mutant phenotypes of the target protein; visualize the mutant phenotype prediction results and perform phenotype classification according to the prediction results.
[0011] As a preferred embodiment of the machine learning-assisted integrated protein mutant phenotype prediction method of the present invention, wherein: obtaining the sequence information after alignment includes obtaining the sequence information of the target protein and searching for homologous protein sequences in the database by standard methods; screening the search results, removing redundant sequences, and selecting an appropriate number of homologous sequences according to preset criteria; performing multiple sequence alignment on the screened homologous sequences to generate an alignment result.
[0012] As a preferred embodiment of the machine learning-assisted integrated protein mutant phenotype prediction method of the present invention, wherein: the Shannon entropy value is expressed as,
[0013]
[0014] where p j is the relative frequency of the jth amino acid at position i in the protein sequence; the co-evolution coefficient is expressed as,
[0015]
[0016] where P(xi,yj) is the joint probability of observing amino acid types x and y at sequence positions i and j respectively, and P(xi) is the marginal / singlet probability of x-type amino acids at the ith position.
[0017] As a preferred embodiment of the machine learning-assisted integrated protein mutation phenotype prediction method of the present invention, wherein: constructing the amino acid contact energy network includes constructing an amino acid contact energy network for each phenotype, importing the corresponding module to generate the contact energy network of the protein, and calculating the hydrophobicity and polarity of each amino acid site; through the differential analysis of the mutant type and the wild type, obtaining the global hydrophobicity and polarity changes; using graph theory methods to calculate the betweenness centrality, closeness centrality, and eigenvector centrality between each disease group and the wild type; saving each calculation result according to different disease groups to form a differential data file.
[0018] As a preferred embodiment of the machine learning-assisted integrated protein mutation phenotype prediction method of the present invention, wherein: the relative accessible area is expressed as
[0019]
[0020] wherein, A exposed represents the solvent-exposed surface area of the amino acid in the protein, and A total represents the solvent-exposed surface area of the amino acid in the isolated solution.
[0021] As a preferred embodiment of the machine learning-assisted integrated protein mutation phenotype prediction method of the present invention, wherein: constructing the elastic network model of the target protein includes constructing an elastic network of the protein according to the interaction of amino acid residues, and calculating the collective vibration mode, structural dynamics properties, and vibration correlation between nodes of the protein through the ANM and GNM models respectively; by analyzing the effectiveness, sensitivity, and stiffness of the protein, evaluating the impact of the mutation on the overall structure and function of the protein.
[0022] As a preferred embodiment of the machine learning-assisted integrated protein mutation phenotype prediction method of the present invention, wherein: using machine learning methods to predict the mutation phenotype of the target protein includes, for each prediction task, transforming it into multiple binary classification problems, and performing dimensionality reduction processing on the calculated feature parameters; adopting recursive feature elimination technology, and combining with the LightGBM algorithm for feature screening, optimizing the feature combination, evaluating its accuracy using cross-validation, and returning the best feature set; based on the selected feature set, using the LightGBM algorithm for training, and analyzing the positive and negative impacts of the features on the model results through the SHAP interpretation method; adopting a combination optimization of multiple machine learning models, obtaining the best prediction model by traversing the model combination; inputting the prediction results of different models into a logistic regression model for integration, calculating and outputting the final PhenoScore value, and selecting the best model combination according to the AUC value.
[0023] Another object of the present invention is to provide a machine learning-assisted integrated protein mutation phenotype prediction method system, which can solve the problem of one-sided evaluation of the impact of mutations in existing methods by constructing a machine learning-assisted integrated protein mutation phenotype prediction method system, integrating various information such as protein sequence features, structural features, amino acid contact energy network, and elastic network model, and optimizing the accuracy and reliability of mutation prediction.
[0024] To solve the above technical problems, the present invention provides the following technical solutions: A machine learning-assisted integrated protein mutation phenotype prediction method system, including: a homologous sequence module, used to obtain the homologous sequences of the target protein and perform multiple sequence alignment on the homologous sequences to obtain the aligned sequence information; a conservation assessment module, used to perform conservation scoring on each amino acid site based on the aligned sequence information to evaluate the conservation degree of each site during evolution; an entropy co-evolution module, used to calculate the Shannon entropy value and co-evolution coefficient of the target protein based on the conservation scoring to characterize the mutation frequency and co-evolution relationship of amino acid sites; a free energy assessment module, used to calculate the free energy change of the target protein using the Shannon entropy value and co-evolution coefficient to evaluate the impact of mutations on protein stability; a contact energy network module, used to construct an amino acid contact energy network based on the free energy change and calculate the topological features in the network; an accessible area module, used to calculate the relative accessible area of each site based on the amino acid contact energy network to evaluate the impact of mutations on the surface-exposed region of the protein; an elastic network module, used to construct an elastic network model of the target protein to simulate the impact of mutations on protein dynamic stability; a feature integration module, used to integrate the sequence features, structural features, and elastic network features of the protein to form a comprehensive feature set; a phenotype prediction module, used to perform mutation phenotype prediction of the target protein using machine learning methods based on the comprehensive feature set; a visualization module, used to visualize the mutation phenotype prediction results and perform phenotype classification according to the prediction results.
[0025] A computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, it implements the steps of the above-mentioned machine learning-assisted integrated protein mutation phenotype prediction method.
[0026] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the above-mentioned machine learning-assisted integrated protein mutation phenotype prediction method.
[0027] Advantages of the present invention: The machine learning-assisted integrated protein mutation phenotype prediction method provided by the present invention solves the problems of local analysis and lack of global consideration in the prior art by introducing machine learning techniques and multi-dimensional feature integration methods, comprehensively considering the sequence features, structural features, and elastic network features of proteins, and using multiple algorithms for feature extraction and modeling. The present invention comprehensively analyzes the effects of protein mutations by calculating multiple parameters (such as conservation, Shannon entropy, co-evolution coefficient, free energy change, amino acid contact energy network, relative accessible area, etc.), and combines the elastic network model and contact energy network to better simulate and predict the global effects of protein mutations on function. In addition, considering more complete features and having interpretability, it solves the phenotype prediction problem for individual proteins. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0029] Figure 1 FIG. is the overall flowchart of the machine learning-assisted integrated protein mutation phenotype prediction method provided by an embodiment of the present invention.
[0030] Figure 2 FIG. is the program running diagram of the machine learning-assisted integrated protein mutation phenotype prediction method provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the specific embodiments of the present invention in detail with reference to the drawings of the specification. Obviously, the described embodiments are some embodiments of the present invention, not all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0032] Many specific details are set forth in the following description to facilitate a thorough understanding of the present invention, but the present invention may be practiced in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the spirit of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0033] Embodiment 1
[0034] Refer to Figure 1 , for an embodiment of the present invention, a machine learning-assisted integrated protein mutation phenotype prediction method is provided, including:
[0035] Step S1: Obtain homologous sequences and perform multiple sequence alignment
[0036] Download the fasta file of Uniref50 in UniProt and place it in the server. Use the command sudo apt-get install ncbi-blast+ to download the blast program, and use the blastp program to build a local uniref50 protein sequence database (note that the database is relatively large, so make sure there is at least 60G of space). First, the user's file is transferred to the data directory, and Run.py starts running. First, read the protein sequence in the pdb file, import the module in APMA.py, and Feature_Cal / Blast_MSA.py in feature calculation is imported. Using the interface program for blastp in biopython, search for homologous sequences in the locally established Uniref50 protein sequence library using all species databases through Blastp, set evalue = 1e-50, and search for at most 3000 sequences, and save them into data / Blast_result.fasta. Because the conservation of sites needs to be statistically analyzed, it is not necessary to align all sequences (it takes too long and is inefficient). At the same time, sequences with only relatively high evalue cannot be considered otherwise it is inaccurate. So, first remove duplicates from all sequences and then perform random sampling. Randomly select 200 sequences from at most 3000 sequences with evalue greater than 1e-50 and rewrite them into the blast_results.fasta file. If the number of sequences found is less than 200, then align all sequences, place the wild-type sequence at the front >Input_seq, and then run Clustal Omega for multiple sequence alignment to generate a multiple sequence alignment file query_msa.fasta.
[0037] Step S2. Score the conservation of each site
[0038] The result of multiple sequence alignment can reflect the conservation of sites at the evolutionary level. To quantitatively evaluate the conservation of the site where the mutation occurs, score the conservation of the multiple sequence alignment result. Use the command sudo apt-get install rate4site in the server to install the mutation scoring tool rate4site, and use the rate4site tool to score the conservation of query_msa.fasta. Each site has a score, and the generated score.txt text is stored in data / score.txt. The lower the conservation score, the more conserved the site is at the sequence level.
[0039] Step S3. Calculate the entropy and co-evolution coefficient of the protein
[0040] The amino acid sites of amino acids can also use information entropy and mutual information to reflect evolutionary conservation. In the present invention, the Feature_Cal / sequence.py and Feature_Cal / sequence.py modules are imported to calculate S(i) and M(i), that is, the Shannon entropy and mutual information of the protein sequence (to represent the co-evolution between amino acid sites).
[0041] The Shannon information entropy of a protein
[0042] A concept based on information theory, which measures the uncertainty or information content of a system. In a protein sequence, Shannon entropy is used to describe the diversity or uncertainty of amino acids at a specific position (position i). To calculate Shannon entropy, the multiple sequence alignment result of the protein needs to be obtained. First, blastp is used to search for homologous sequences of the protein, and then clusteral omega is used for multiple sequence alignment (MSA). The obtained multiple sequence alignment result is used for subsequent calculations. The calculation formula of Shannon entropy S(i) is:
[0043]
[0044] where p j is the relative frequency of the jth amino acid at position i in the protein sequence. The value range of Shannon entropy is usually between 0 and 3. When there is only one amino acid at a position, the Shannon entropy is 0, indicating that the amino acid at this position is determined; while when all 20 amino acids appear with equal probability at this position, the Shannon entropy reaches the maximum value, indicating that the amino acid at this position is completely uncertain.
[0045] The co-evolution coefficient of a protein
[0046] The co-evolution coefficient M(i) of a protein is an index used to measure the degree of mutual correlation of proteins during evolution. The common method is based on the mutual information (MI) calculation formula, which can be expressed as
[0047]
[0048] P(xi,yj) is the joint probability of observing amino acid types x and y at sequence positions i and j respectively. P(xi) is the marginal / singlet probability of amino acid type x at the ith position. I(I,j) varies within the range of [0,Imax], corresponding to completely uncorrelated and most correlated residue pairs. The co-evolution of mutations is measured by the average MI value corresponding to each residue.
[0049] Step S4. Perform mutant construction and free energy change calculation
[0050] In addition to the conservation of amino acid sites, the free energy change before and after site mutation can reflect the degree of local structural change of the protein. In this invention, FoldX5 is used to evaluate the free energy change before and after amino acid mutation. Install the linux version of FoldX5, import the mutation / FoldX.py module to process files, start running FoldX, delete the pdb files starting with WT obtained from the run, rename the files in the form of pten.pdb according to the phenotype classification, rename them as ASD_1.pdb, ASD_2.pdb…Cancer_1.pdb…, and at the same time record the energy change in the.fdxout file obtained from the run. Each mutation corresponds to an energy change, and the ddG corresponding to each site is obtained.
[0051] Step S5. Construct the amino acid contact energy (AAN) network
[0052] The above mutation measurement indicators do not consider the three-dimensional structure of the protein. In order to measure the role of amino acid sites in the internal site regulation network of the protein and the topological structure of the network, in this embodiment, an amino acid contact energy network is constructed and the important topological centrality parameters of the nodes are calculated. Import Feature_Cal / AAWeb.py, construct an R file of AANetwork.R for each phenotype, use NACEN to construct a weighted contact energy network of protein sites to calculate the Node Hydrophobicity and Node polarity of each protein. Since the Hydrophobicity and Polarity also change due to amino acid changes, the global Hydrophobicity and Polarity differences are obtained by adding the mutant type and the wild type. At the same time, import the igraph package to calculate the global differences of Betweenness, Closeness, and Eigenvector between each disease group and the wild type, and save them in data / AAWeb according to different disease groups, in the form of ASD.txt…
[0053] Betweenness centrality
[0054] Betweenness centrality is related to the shortest path of communication between different nodes in the network. Betweenness centrality is a global parameter that can reflect the information of the entire network at a node to a certain extent. A node with a larger betweenness indicates that it occupies a higher position in network communication and will be an important bottleneck for signal transmission in the network.
[0055]
[0056] Among them, C B (v) represents the betweenness centrality of amino acid v, and σ st is the number of shortest paths from node s to node t, and σ st (v) is the number of shortest paths passing through node v.
[0057] Closeness centrality
[0058] Closeness centrality is also related to the shortest paths of communication between different nodes in the network. It is also a global parameter that measures the average communication fluency from a node to other nodes. The larger the Closeness, the smaller the sum of distances from the site to other sites, and the relatively easier the communication from it to other sites.
[0059]
[0060] Among them, C C (v) represents the closeness centrality of amino acid v, and d(v, u) is the length of the shortest path from node v to node u
[0061] Eigenvector centrality
[0062] Eigenvector centrality is an important concept in graph theory used to measure the importance of nodes in a graph. Different from degree centrality, eigenvector centrality not only focuses on the number of direct connections of a node but also considers the importance of the nodes it is connected to.
[0063]
[0064] Among them, C E (v) represents the eigenvector centrality of amino acid v, λ is the eigenvalue corresponding to the eigenvector, and A ij is the element at the i, j position in the adjacency matrix of the network, and x j is the centrality value of node j.
[0065] Step S6. Calculate the relative accessible surface area (RASA) of each site of the protein
[0066] The three-dimensional structure information of the protein can be studied not only through the amino acid network but also using the relative surface area of the sites. Install the dssp tool using the command sudo apt-get install dssp in the linux server, import the Feature_Cal / DSSP_RASA.py module, and calculate the rASA value of each site of the wild-type protein using the rolling ball algorithm by calling DSSP. The rASA can be calculated using the following formula:
[0067]
[0068] where A exposed represents the solvent-exposed accessible surface area of an amino acid in a protein, and A total represents the solvent-exposed accessible surface area of an amino acid in a separate solution. The area of the protein in contact with the solvent ranges from [0, 1]. The closer to 1, the more this amino acid is in the hydrophilic region, i.e., on the surface of the protein. The closer to 0, the more this amino acid is in the hydrophobic region (folded inside the protein structure). Generally, amino acids in the hydrophobic region have stronger functions.
[0069]
[0070] Step S7. Construct an elastic network model of the protein and integrate parameters
[0071] The effect of mutation is to disrupt the amino acid network of the original wild-type protein. Here, an elastic network model of the protein is introduced. The elastic network belongs to the category of network dynamics. When the protein network is subjected to external interference (i.e., amino acid site mutation), the elastic network will adaptively adjust and reorganize. This effect can be characterized by parameters such as Effectiveness, Sensitivity, and Stiffness. Sites with low effectiveness, high sensitivity, and high stiffness will cause significant changes in the structure of the entire network when mutated. The elastic network characteristics of the protein will be calculated by the ANM model and the GNM model.
[0072] ANM (Anisotropic Network Model): ANM is a force field model used to describe the collective motion and vibration of proteins. In ANM, the structure of the protein is modeled as an elastic network, where each amino acid residue is regarded as a node, and the interaction between them is represented as a spring connecting these nodes. ANM assumes that the dynamics of the protein structure can be described by the vibration of these springs, and these vibrations can be calculated through linear regularization. The ANM model is used to predict the collective mode vibration of proteins and related structural dynamic properties.
[0073] GNM (Gaussian Network Model): GNM is a statistical mechanics model used to describe the dynamic properties of proteins. In GNM, the structure of a protein is modeled as a Gaussian network, where each amino acid residue is regarded as a node, and the interactions between them are represented by a Gaussian potential function. GNM assumes that the dynamic properties of the protein structure are mainly influenced by the interactions between these nodes, and this influence can be described by calculating the correlations between the nodes. The GNM model is used to analyze the structural dynamic properties of proteins, such as vibrational correlations and dynamically communicating paths related to functions.
[0074] The total potential energy of the GNM and ANM models with N nodes can be calculated using the following formula:
[0075]
[0076] where γ represents the force constant of the spring between sites in the elastic network, represents the spatial displacement of the amino acid after perturbation, and Γ ij represents the (i, j)-element of the adjacency matrix of the amino acid network.
[0077] Step S8. Protein Mutation Integration Model
[0078] In this embodiment, a sequence model of proteins, an amino acid interaction network model, and a protein elastic network model have been introduced. Here, this embodiment proposes an integrated model to characterize the impact of a single-site protein. A protein can be simplified into a graph structure, where each site represents an amino acid, and the sequence features of the amino acid, such as Entropy, Coevolution, RASA, conservativeness, etc., are recorded at the site. At the same time, the amino acid sites interact with other sites, and the amino acid network retains the information of the mutual communication paths between the sites. Meanwhile, the elastic network model of the protein is used to calculate the dynamic stability of the protein sites in the network. In network dynamics, sites with high effectiveness, low sensitivity, and low stiffness tend to only locally affect the structure of the network and have little impact on the distal structure; while sites with low effectiveness, high sensitivity, and high stiffness tend to transmit the change signals generated by mutations to more distant positions when mutated, resulting in changes in the local network structure at the mutation position and affecting the distal structure at the same time, which to a certain extent reflects the allosteric effect of the protein. The occurrence of such changes will lead to changes in the communication paths from distal amino acids to proximal amino acids. Sites that could originally communicate are isolated or communicate through different paths, affecting other pathways of the protein. At the same time, the sequence-related information of the protein is also stored in the nodes, and classic energy and conservativeness-related information are also considered, thus establishing a relatively complete interpretable protein mutation integration model. Machine learning methods can combine diverse information. In this embodiment, the calculated features are divided into three categories: protein sequence-related information, protein contact energy network-related information, and protein elastic network-related information, and these three types of information are integrated.
[0079] Step S9. Machine learning: Dimensionality reduction - Interpretation - Search - Scoring
[0080] Dimensionality reduction: Import the machine learning module and start automatic modeling using the sklearn package. First, all classifications are divided into binary classifications one by one. First, perform RFE feature elimination on the 15 calculated parameters. The core algorithm used is LightGBM, and the cross-validation scores are returned through stratified 10-fold cross-validation. This score can be considered as the accuracy of the model. The RFE and cross-validation results are saved in model_construction.txt. The parameter combinations calculated by LightGBM are fine-tuned to obtain indicators for differentiating phenotypes.
[0081] Explanation: Train the LightGBM model with the obtained indicators parameters. For each classification task, use the shap package to interpret the model, draw the beeswarm plot and force plot, and explore the positive or negative influence degree of each parameter in the indicators on the final model judgment result.
[0082] Search & scoring: For each two - combination input, traverse all model combinations in XGBoost, LightGBM, and CatBoost. Input the predicted probability values into the logistic regression model for integration. Then calculate the AUC value of the integrated score PhenoScore for binary classification, return the model combination with the largest AUC value for classification, record the scoring value file of each locus in Outcome, and also record the best model combination and the best AUC value in Outcome's model_construction.txt.
[0083] Step S10. Visualization
[0084] Import the plotting module and use two Python packages, matplotlib and seaborn, for automatic plotting. Draw distribution images (violin plot + box plot + scatter plot) for each feature, perform principal component analysis PCA and plot for each phenotype, perform spearman correlation analysis between each feature, draw the correlation heat map between features, draw the feature ROC curve for each feature. Here, since it is a feature ROC curve, it does not represent the superiority or inferiority of the probability for 0 - 1 classification. So, flip the curves with AUC less than 0.5 to above 0.5, and draw the ROC_AUC value of the best score obtained. Draw the shap interpretation beeswarm plot and force plot for the machine learning model, draw the boxplot of the distribution of the integrated phenotype score PhenoScore, and also draw the ROC image of PhenoScore and calculate the AUC.
[0085] Example 2
[0086] Reference Figure 2 , which is an embodiment of the present invention, provides a machine - learning - assisted integrated protein mutation phenotype prediction method system, including:
[0087] As Figure 2 shown, Figure A shows the running route of deePheMut: first obtain the user's data -> comprehensively calculate the parameters of the protein data -> analyze the parameters and construct the model -> visualize the results. The specific details are as follows. Figure B shows the upload form of the web page, where protein structure information, mutation information, and email are uploaded.
[0088] The route details are as follows: A brief description of the overall steps: Obtaining homologous sequences and multiple sequence alignment: Search for homologous sequences of the protein through BLAST, extract up to 3000 sequences with an evalue less than 1e-50 from UniRef50, randomly select 200 for Clustal Omega multiple sequence alignment, and generate an alignment file for subsequent analysis. Feature calculation and scoring: Use Rate4site to calculate the site conservation score, combine FoldX to calculate the mutation energy change (ddG), and calculate the relative accessible surface area (rASA) through DSSP. Further analyze the sequence evolution characteristics using co-evolution and information entropy models. Based on the protein contact energy and elastic network model, construct an amino acid network, calculate the topological parameters of the nodes, and combine the elastic network model to describe the dynamic impact of mutations on the overall network. Use the gradient boosting algorithm to perform dimensionality reduction and classification task modeling on the features, optimize the parameter combination using RFE and cross-validation, explain the positive and negative impacts of key features on classification, and construct the final classification model by comprehensively scoring the predictions of multiple models. Use matplotlib and seaborn to generate feature distribution, correlation, and principal component analysis (PCA) plots to evaluate the model performance, draw the ROC curve and model interpretation plots, and display the final classification results and phenotype discrimination ability.
[0089] Homologous sequence module, used to obtain homologous sequences of the target protein and perform multiple sequence alignment on the homologous sequences to obtain the aligned sequence information; Conservation assessment module, used to perform conservation scoring on each amino acid site based on the aligned sequence information to evaluate the conservation degree of each site during evolution; Entropy co-evolution module, used to calculate the Shannon entropy value and co-evolution coefficient of the target protein based on the conservation score to characterize the mutation frequency and co-evolution relationship of amino acid sites; Free energy assessment module, used to calculate the free energy change of the target protein using the Shannon entropy value and co-evolution coefficient to evaluate the impact of mutations on protein stability; Contact energy network module, used to construct an amino acid contact energy network based on the free energy change and calculate the topological features in the network; Accessible area module, used to calculate the relative accessible surface area of each site based on the amino acid contact energy network to evaluate the impact of mutations on the surface-exposed region of the protein; Elastic network module, used to construct an elastic network model of the target protein to simulate the impact of mutations on protein dynamic stability; Feature integration module, used to integrate the sequence features, structural features, and elastic network features of the protein to form a comprehensive feature set; Phenotype prediction module, used to perform mutation phenotype prediction of the target protein using machine learning methods based on the comprehensive feature set; Visualization module, used to visualize the mutation phenotype prediction results and perform phenotype classification according to the prediction results.
[0090] Example 3
[0091] An embodiment of the present invention, which is different from the previous two embodiments, is as follows:
[0092] If the above-mentioned function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0093] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0094] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part with one or more wirings (electronic device), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber device, and portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or otherwise processing it as necessary, and then storing it in a computer memory.
[0095] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one of the following techniques well known in the art or a combination thereof can be used: discrete logic circuits having logic gate circuits for implementing logic functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.
[0096] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A machine learning-assisted integrated protein mutation phenotype prediction method, characterized in that: include: Obtaining homologous sequences of the target protein, and performing multiple sequence alignment on the homologous sequences to obtain aligned sequence information; Based on the compared sequence information, each amino acid site is scored for conservation to evaluate the degree of conservation of each site during the evolution process; Based on the conservation score, the Shannon entropy value and coevolution coefficient of the target protein are calculated to characterize the variation frequency and coevolution relationship of the amino acid sites; Using the Shannon entropy value and the coevolution coefficient, the free energy change of the target protein is calculated to evaluate the effect of the mutation on the protein stability; According to the free energy change, construct an amino acid contact energy network, and calculate the topological features in the network; Based on the amino acid contact energy network, the relative accessible area of each site is calculated to evaluate the effect of the mutation on the exposed area of the protein surface; Construct an elastic network model of the target protein to simulate the effect of mutations on the dynamic stability of the protein; Integrating the sequence features, structural features and elastic network features of the protein to form a comprehensive feature set; Based on the comprehensive feature set, a machine learning method is used to predict the mutant phenotype of the target protein; The mutant phenotype prediction results are visualized, and the phenotypes are classified according to the prediction results.
2. The method for predicting protein mutation phenotypes assisted by machine learning according to claim 1, characterized in that: The obtaining of the aligned sequence information includes obtaining the sequence information of the target protein and searching for homologous protein sequences in a database by a standard method; Screen the search results, remove redundant sequences, and select an appropriate number of homologous sequences according to preset criteria; Perform multiple sequence alignment on the screened homologous sequences to generate alignment results.
3. The method for predicting protein mutation phenotypes assisted by machine learning according to claim 2, characterized in that: The Shannon entropy value is expressed as, Among them, p j is the relative frequency of the jth amino acid at position i in the protein sequence; The coevolution coefficient is expressed as, where P(xi,yj) is the joint probability of amino acid types x and y being observed at sequence positions, i and j, respectively, and P(xi) is the marginal / singlet probability of amino acid type x at position i.
4. The machine learning-assisted integrated protein mutation phenotype prediction method according to claim 3, characterized in that: The construction of the amino acid contact energy network includes constructing an amino acid contact energy network for each phenotype, importing a corresponding module to generate a contact energy network of a protein, and calculating the hydrophobicity and polarity of each amino acid site; By analyzing the differences between mutants and wild types, global changes in hydrophobicity and polarity were obtained; Using graph theory methods, betweenness centrality, closeness centrality, and eigenvector centrality between each disease group and the wild type were calculated; Each calculation result is saved according to different disease groups to form a difference data file.
5. The machine learning-assisted integrated protein mutation phenotype prediction method according to claim 4, characterized in that: The relative accessible area is expressed as, Among them, A exposed A represents the exposed volume surface area of amino acids in protein. total Represents the exposed volumetric surface area of the amino acid in solution alone.
6. The machine learning-assisted integrated protein mutation phenotype prediction method according to claim 5, characterized in that: The elastic network model of the target protein is constructed by constructing an elastic network of the protein according to the interaction of amino acid residues, and calculating the collective vibration mode, structural dynamics properties and vibration correlation between nodes of the protein by ANM and GNM models respectively; Assess the effects of mutations on the overall structure and function of the protein by analyzing its effectiveness, sensitivity, and rigidity.
7. The machine learning-assisted integrated protein mutation phenotype prediction method according to claim 6, characterized in that: The method of using a machine learning method to predict the mutant phenotype of a target protein includes converting each prediction task into multiple binary classification problems and performing dimensionality reduction processing on the calculated characteristic parameters; Recursive feature elimination technology is used in combination with the LightGBM algorithm to screen features, optimize feature combinations, use cross-validation to evaluate their accuracy, and return the best feature set; Based on the selected feature set, the LightGBM algorithm is used for training, and the positive and negative effects of the features on the model results are analyzed through the SHAP interpretation method; Use multiple machine learning models for combined optimization and obtain the best prediction model by traversing the model combination; The prediction results of different models were input into the logistic regression model for integration, the final PhenoScore value was calculated and output, and the best model combination was selected based on the AUC value.
8. A system using the machine learning-assisted integrated protein mutation phenotype prediction method according to any one of claims 1 to 7, characterized in that: include: A homologous sequence module is used to obtain homologous sequences of target proteins and perform multiple sequence alignment on the homologous sequences to obtain aligned sequence information; A conservation evaluation module is used to score the conservation of each amino acid site based on the sequence information after the alignment, and evaluate the conservation degree of each site during the evolution process; An entropy coevolution module, for calculating the Shannon entropy value and coevolution coefficient of the target protein based on the conservation score, and characterizing the variation frequency and coevolution relationship of the amino acid sites; A free energy evaluation module, used to calculate the free energy change of the target protein using the Shannon entropy value and the coevolution coefficient, and evaluate the effect of the mutation on the protein stability; A contact energy network module, used to construct an amino acid contact energy network according to the free energy change, and calculate the topological features in the network; An accessible area module, used to calculate the relative accessible area of each site based on the amino acid contact energy network, and evaluate the effect of mutations on the exposed area of the protein surface; The elastic network module is used to construct an elastic network model of the target protein and simulate the effect of mutations on the dynamic stability of the protein; A feature integration module, used to integrate the sequence features, structural features and elastic network features of the protein to form a comprehensive feature set; A phenotype prediction module, used to predict the mutant phenotype of the target protein using a machine learning method based on the comprehensive feature set; The visualization module is used to visualize the mutant phenotype prediction results and perform phenotype classification according to the prediction results.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the machine learning-assisted integrated protein mutation phenotype prediction method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the machine learning-assisted integrated protein mutation phenotype prediction method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Protein phase separation characteristic prediction method based on artificial intelligence
CN121789772A
Monte Carlo optimized protein thermal stability screening method fusing evolutionary information
CN122050499A