Protein thermal stability point mutation prediction model construction method and prediction method

By constructing a protein thermal stability point mutation prediction model, using a three-dimensional convolutional neural network and a transfer learning neural network, combined with experimental data optimization model, the problem of limited prediction accuracy in the existing technology is solved, and more efficient prediction of protein thermal stability and functional activity is achieved.

CN119920310AActive Publication Date: 2025-05-02HUBEI UNIV

Patent Information

Application Number
CN202510001719.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-02
Estimated Expiration
2045-01-02

AI Technical Summary

Technical Problem

The prior art has experimental methods with high time and resource consumption, calculation methods with high computational costs and limited simulation accuracy, and machine learning models with insufficient comprehensive feature extraction, resulting in limited prediction accuracy.

Method used

A method for building a protein thermal stability point mutation prediction model is provided. By obtaining the three-dimensional structure of the protein, rasterizing the process to form multiple three-dimensional matrices, and combining them into a four-dimensional matrix with physical quantities as channels, training a three-dimensional convolutional neural network to build an initial model, and combining experimental data to optimize the final model through transfer learning neural networks.

Benefits of technology

This method can more comprehensively reflect the structural characteristics of proteins and improve prediction accuracy. Especially in the absence of comprehensive experimental data, it can still make reasonable inferences based on known experimental information and benchmark values, and optimize the prediction of protein thermal stability and functional activity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119920310A_ABST
    Figure CN119920310A_ABST
Patent Text Reader

Abstract

According to the protein thermal stability point mutation prediction model construction method disclosed by the invention, multi-dimensional physical quantity selection is performed within the preset distance range around the target residue, so that the structural characteristics of the protein can be reflected more comprehensively, and a richer information basis is provided for predicting the protein thermal stability point mutation. When the three-dimensional convolutional neural network is trained, interaction between atoms in a large range can be captured through a large-size convolution kernel, and a complex relation is revealed; local details can be focused through a small-size convolution kernel, structural features can be comprehensively described from different scales through combination of the two, and it is ensured that the model learns key information. The normalization method of the transfer learning neural network emphasizes information provided by experimental data, keeps the benchmark probability of non-mutated amino acid as a reference, can more accurately predict the influence of amino acid mutation on the thermal stability and functional activity of the protein, and especially, can predict the thermal stability and functional activity of the protein under the condition of lacking comprehensive experimental data. Reasonable inference can still be made depending on known experimental information and reference values.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of protein engineering technology, and more specifically, to a method for constructing a protein thermal stability point mutation prediction model and a prediction method. Background Art

[0002] Protein thermal stability is a key research area in protein engineering, which is of great significance for understanding protein functions, optimizing protein performance and developing new biopharmaceuticals. In the relevant technical fields, it involves the intersection of multiple disciplines such as bioinformatics, computational biology and protein chemistry, aiming to reveal the relationship between protein structure and function through a combination of theoretical calculation and experimental research, and design and modify proteins on this basis.

[0003] The existing technology mainly uses the following methods to study the thermal stability of proteins. Traditional experimental methods usually rely on site-directed mutation of proteins, and then measure the thermal stability changes of mutants through experimental means (such as differential scanning calorimetry, dynamic light scattering, etc.). Although this method is direct, it takes a lot of time and resources because each mutant requires tedious experimental preparation and testing processes. At the same time, due to the complexity of protein structure and function, it is inefficient to screen mutation sites one by one through experiments, and it is difficult to fully explore the effects of all possible mutation combinations on thermal stability.

[0004] Computational methods make up for the shortcomings of experimental methods to a certain extent. Some studies use molecular dynamics simulations to predict structural changes after protein mutations, and then evaluate their effects on thermal stability. This method can simulate the dynamic behavior of proteins in a virtual environment, but the computational cost is high, and the accuracy of the simulation is limited by factors such as force field parameters and simulation time. In addition, machine learning-based methods have also gradually been applied to the prediction of protein thermal stability, such as using known protein structures and thermal stability data to train classification or regression models to predict the effects of mutations on thermal stability. However, existing machine learning models are often not comprehensive enough in feature extraction and cannot fully capture key information related to thermal stability in protein structure, resulting in limited prediction accuracy. Summary of the invention

[0005] In response to at least one defect or improvement need in the prior art, the present application provides a method for constructing a protein thermal stability point mutation prediction model and a prediction method, which are used to overcome some of the defects mentioned in the background technology section.

[0006] To achieve the above objectives, in a first aspect, the present application provides a method for constructing a protein thermal stability point mutation prediction model, comprising:

[0007] Obtaining a three-dimensional structure of a preset type of protein, selecting a preset type of physical quantity within a preset distance range around the target residue for rasterization processing to form multiple three-dimensional matrices, and merging them into a four-dimensional matrix using the physical quantity type as a channel;

[0008] Using the four-dimensional matrix as training data, training a three-dimensional convolutional neural network to construct an initial protein thermal stability point mutation prediction model; the three-dimensional convolutional neural network includes a preset large-size convolution kernel for extracting global features and a preset small-size convolution kernel for extracting local features;

[0009] The output of the previous layer of the output layer of the initial protein thermal stability point mutation prediction model is used as input, and a normalization method is used to adjust the amino acid mutation probability in combination with enzyme activity experimental data and / or keep the baseline probability of non-mutated amino acids in the absence of enzyme activity experimental data as a reference, to train a transfer learning neural network and obtain an optimized final protein thermal stability point mutation prediction model.

[0010] Furthermore, the preset distance range is specifically:

[0011] The region between the spherical regions with a radius of 5 to 15 angstroms centered on the target residue.

[0012] Furthermore, the distance to the target residue is weighted by a distance weighting function w(r), and the specific formula includes:

[0013]

[0014] Where r represents the distance between the target residue and the surrounding atoms; α represents the adjustment parameter; and w(r) represents the weighted distance to the target residue.

[0015] Furthermore, the preset types of physical quantities include physical quantities that reflect the structural characteristics of proteins from the atomic level, the intermolecular force level, and the chemical group level;

[0016] The preset types of physical quantities specifically include one or more of the number of carbon atoms, the number of nitrogen atoms, the solvent accessible surface area, the average charge value, the number of hydrogen bond donors, the number of hydrogen bond acceptors, and the number of aromatic rings.

[0017] Furthermore, the sizes of the preset large-size convolution kernel and the preset small-size convolution kernel of the three-dimensional convolution layer of the three-dimensional convolutional neural network are 3×3×3 and 2×2×2, respectively, and the step size is set to 1.

[0018] Furthermore, the three-dimensional convolution layer performs a convolution operation on the input matrix through a sliding window operation. The convolution kernel performs element-level multiplication and addition operations with each local area of ​​the input to obtain the eigenvalue of the corresponding position. The specific formula includes:

[0019]

[0020] Among them, I x ,I y ,I z Respectively represent the size of the input matrix in three dimensions; f represents the size of the convolution kernel; s represents the step size; O x , O y , O z They represent the size of the convolution output matrix in the corresponding dimension.

[0021] Furthermore, the normalization method for adjusting the probability of amino acid mutation in combination with enzyme activity experimental data includes:

[0022] Assume that the initial mutation probability of all amino acids is 100%;

[0023] If the experimental data show that the mutation of a certain amino acid increases the activity of the enzyme, then the mutation probability corresponding to the amino acid will increase accordingly; otherwise, it will decrease accordingly;

[0024] The adjusted mutation probabilities are normalized so that the sum of all mutation probabilities is 100%; the specific formula includes:

[0025]

[0026] Among them, a i represents the relative activity percentage of the ith amino acid associated with the experimental data; N represents the total number of all amino acids; It represents the sum of the relative activity percentages of all amino acids; Represents the normalized amino acid mutation probability in the scenario with enzyme activity experimental data;

[0027] Normalization methods that keep the baseline probability of unmutated amino acids as a reference when there is no enzyme activity experimental data include:

[0028] The mutation probability of amino acids with no enzyme activity experimental data was set to 0%;

[0029] The mutation probability of unmutated amino acids was set to 100% as a baseline value, thus ensuring that the model's predictions could be referenced to its initial state in the absence of mutations;

[0030] The specific formula includes:

[0031]

[0032] Among them, p i represents the mutation probability after correction by experimental data; represents the sum of all amino acid mutation probabilities; Represents the normalized amino acid mutation probability in the scenario of no enzyme activity experimental data.

[0033] Furthermore, the fully connected layer of the transfer learning neural network uses ReLU as an activation function;

[0034] The loss function used when training the three-dimensional convolutional neural network includes:

[0035]

[0036] Among them, L(θ) represents the total loss function; represents the mutation probability predicted by the network; y i represents the actual mutation probability; θ j represents the weight parameter of the network; N represents the total number of training samples; λ represents the regularization term coefficient; m represents the number of parameters in the network.

[0037] In the second aspect, the present application provides a method for predicting protein thermal stability point mutations, wherein the structural data of the protein to be studied is input into the initial protein thermal stability point mutation prediction model or the final protein thermal stability point mutation prediction model in any of the prediction model construction methods described above for processing, and the mutation site and the target amino acid to be mutated are predicted.

[0038] Furthermore, the standard deviation and actual amino acid mutation probability of each candidate site were evaluated, and the eligible sites were ranked;

[0039] If there are multiple sites with standard deviations higher than the first preset value, these sites are sorted and recommended in descending order of standard deviation; the sorting criteria for standard deviations include:

[0040]

[0041] Among them, σ n represents the standard deviation corresponding to the candidate site n; Indicates the standard deviation ranking value;

[0042] If the actual amino acid mutation probability of multiple sites is not higher than the second preset value, the sites are ranked and recommended in ascending order of mutation probability; the ranking criteria for mutation probability include:

[0043]

[0044] Among them, p i represents the mutation probability of candidate site i; Represents the mutation probability ranking value;

[0045] Based on the weights of the standard deviation ranking value and the mutation probability ranking value, the final recommendation order is determined. The specific formula includes:

[0046]

[0047] Among them, w σ and w p They represent the weight of the standard deviation ranking value and the weight of the mutation probability ranking value, respectively, σ +w p =1;Rank i Indicates the comprehensive ranking value;

[0048] Based on the comprehensive ranking values, mutation sites and target amino acids to be mutated are recommended, and mutation sites that are ranked high based on the comprehensive ranking values ​​are preferentially selected for experimental research.

[0049] In general, the above technical solutions conceived by the present application can achieve the following beneficial effects compared with the prior art:

[0050] The method for constructing a protein thermal stability point mutation prediction model disclosed in the present application selects multidimensional physical quantities within a preset distance range around the target residue, which can more comprehensively reflect the structural characteristics of the protein and provide a richer information basis for predicting protein thermal stability point mutations. When training a three-dimensional convolutional neural network, a large-size convolution kernel can capture the interactions between atoms in a larger range and reveal complex relationships; a small-size convolution kernel can focus on local details. The combination of the two can comprehensively describe structural features from different scales to ensure that the model learns key information. Furthermore, the normalization method of the transfer learning neural network emphasizes the information provided by the experimental data, while keeping the baseline probability of the unmutated amino acid as a reference, which can more accurately predict the impact of amino acid mutations on protein thermal stability and functional activity, especially in the absence of comprehensive experimental data, and can still rely on known experimental information and benchmark values ​​to make reasonable inferences. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0052] Figure 1 A core flow chart of a method for constructing a protein thermal stability point mutation prediction model provided in an embodiment of the present application;

[0053] Figure 2A system architecture diagram of a protein thermal stability point mutation recommendation system based on transfer learning provided in an embodiment of the present application;

[0054] Figure 3 A flow chart of data processing for a feature engineering unit provided in an embodiment of the present application;

[0055] Figure 4 A three-dimensional convolutional neural network data processing flow chart provided in an embodiment of the present application;

[0056] Figure 5 A flowchart of a transfer learning neural network operation provided in an embodiment of the present application;

[0057] Figure 6 A logic diagram for sorting recommendation units provided in an embodiment of the present application;

[0058] Figure 7 A schematic block diagram of an electronic device suitable for implementing the protein thermal stability point mutation prediction model construction method and / or prediction method described above, provided in an embodiment of the present application. DETAILED DESCRIPTION

[0059] In order to make the purpose, technical solutions and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. In addition, the technical features involved in each embodiment of the present application described below can be combined with each other as long as they do not conflict with each other.

[0060] The terms "first", "second" or "nth" in the specification, claims or drawings of the present application are used to distinguish different objects rather than to describe a specific order. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units inherent to these processes, methods, products or devices.

[0061] This application selects multi-dimensional physical quantities through feature engineering units, and adopts reasonable distance ranges and weighting mechanisms, which can more comprehensively and accurately reflect the structural characteristics of proteins, and provide more abundant and accurate information for subsequent predictions. However, the prior art is relatively single in feature extraction, and it is impossible to fully consider the various key factors in protein structure, which affects the accuracy of prediction. In terms of neural network architecture, the convolutional neural network module and transfer learning neural network module of this application improve the learning ability and prediction performance of the model through optimized structural design (such as convolution layer parameter setting, full connection layer and pooling layer coordination, use of pre-training model and ReLU activation function, etc.). In contrast, the calculation model in the prior art may be unreasonable in structural design, with high calculation cost and unsatisfactory prediction effect. The normalization method of this application dynamically adjusts the probability of amino acid mutation in combination with experimental data, and reasonably handles special cases at the same time, so that the prediction is more in line with the actual situation, and the prior art is insufficient in optimizing prediction in combination with experimental data, and it is difficult to accurately reflect the comprehensive impact of mutation on protein function and thermal stability. The recommendation unit of this application optimizes the recommendation order based on multi-criteria sorting, which can provide more targeted and prioritized mutation sites for experiments. However, the existing technology lacks systematicity and accuracy in recommending mutation sites, making it difficult to efficiently guide experimental research. In summary, the existing technology has many limitations in the study of protein thermal stability point mutations. This application aims to overcome these shortcomings and provide a more effective prediction and recommendation technology.

[0062] One embodiment of the present application provides a recommendation scheme for protein thermal stability point mutations based on transfer learning. The recommendation system mainly includes four module units: feature engineering unit, convolutional neural network module, transfer learning neural network module and recommendation unit. The first three module units are essentially the construction process of protein thermal stability point mutation prediction model, which can be referred to Figure 1 ; The last module unit is the specific application of the constructed prediction model, which can be referred to Figure 2 The specific technical details of these four parts are introduced one by one below.

[0063] (1) Feature Engineering Unit

[0064] The three-dimensional structural data of various proteins were obtained from the RCSB protein structure database, including but not limited to enzymes, antibodies, structural proteins and other different types of proteins. Physical quantities within a specific distance range around the target residue were selected for rasterization to form multiple three-dimensional matrices, which were then merged into four-dimensional matrices using the types of physical quantities as channels as input data for the convolutional neural network module. The physical quantities include the number of carbon atoms, the number of nitrogen atoms, the number of other types of atoms, the solvent accessible surface area, the average charge value, the number of hydrogen bond donors, the number of hydrogen bond acceptors and the number of aromatic rings. These physical quantities reflect the structural characteristics of proteins from multiple dimensions such as the atomic level, the intermolecular force level and the chemical group level, providing comprehensive structural information input for subsequent neural network training.

[0065] The public protein structure database uses the RCSB protein structure database, which contains three-dimensional structural data of proteins from a variety of biological sources and with different functions. By using these data, the convolutional neural network module can capture the universal structural features between different proteins during training, and can also identify the structural differences of proteins with specific functions or sources. The structural data provided by this database not only includes the spatial arrangement of proteins and the interactions between atoms, but also covers multi-dimensional physical and chemical information such as solvent accessible surface area, charge distribution, and the number of hydrogen bond donors and acceptors. These data provide rich inputs for the convolutional neural network, ensuring that the network can learn about multiple levels of features such as protein conformational changes, interactions between amino acid residues, and protein stability.

[0066] In the feature engineering unit, when selecting physical quantities around the target residue, the specific distance range used is the space between a spherical region with a radius of 5 to 15 angstroms centered on the target residue. Figure 3 The figure shows the data processing flow chart of the feature engineering unit. The selection of this distance range is based on a deep understanding of the relationship between protein thermal stability and the surrounding environment. By selecting this range, it is possible to cover the key atoms and chemical groups that interact with the target residues, while avoiding the introduction of too much structural information irrelevant to thermal stability due to too large a selection range. Generally, a small distance range will lead to insufficient local information and fail to fully reflect the structural environment of the target residue, while a large range may contain too much background noise irrelevant to the target residue, affecting the training effect of the subsequent model. Therefore, by selecting physical quantities within the range of 5 to 15 angstroms, it is possible to effectively balance the integrity of information and the suppression of noise, thereby providing accurate and representative input features for the neural network.

[0067] In the specific implementation process, the feature engineering unit first determines the position of the target residue, and collects physical quantities in a spherical area with a radius of 5 to 15 angstroms with the residue as the center. These physical quantities include the number of carbon atoms, the number of nitrogen atoms, the charge distribution, the number of hydrogen bond donors and acceptors, the number of aromatic rings, etc. These physical quantities reflect the local chemical environment of the target residue and its interaction with surrounding atoms and groups to a certain extent. Subsequently, these physical quantities are converted into multiple three-dimensional matrices through rasterization processing, and finally merged into a four-dimensional matrix with the type of physical quantity as the channel, and sent to the neural network as input data for subsequent training.

[0068] In order to further ensure the rationality of the selection of this physical quantity, the distance to the target residue can be weighted by the following distance weighting function w(r), which indicates the contribution of atoms in different distance ranges to the thermal stability of the target residue. The weighting function can be expressed as:

[0069]

[0070] Among them, r is the distance between the target residue and the surrounding atoms; α is the adjustment parameter, which controls the influence of distance on the physical quantity. This weighting function shows that atoms closer to the target residue have a greater influence on thermal stability, so their weight is higher, while atoms farther away from the target residue gradually contribute less to thermal stability.

[0071] In practical applications, this weighting method can be used to adjust the forces of atoms and chemical groups in different distance ranges on thermal stability. By assigning reasonable weights to the selected physical quantities, it can be ensured that the input data more accurately reflects the local environmental characteristics of the target residue. This process not only helps the neural network learn the impact of local changes in protein structure on thermal stability, but also improves the generalization ability of the model, ensuring that its thermal stability can be accurately predicted when faced with different protein structures.

[0072] Therefore, the feature engineering unit effectively improves the representativeness of the input data through reasonable distance selection and weighting mechanism, providing accurate and comprehensive structural features for the subsequent convolutional neural network module and transfer learning neural network module, thereby enhancing the performance of the system in protein thermal stability prediction.

[0073] This application selects a variety of physical quantities such as the number of carbon atoms, the number of nitrogen atoms, the number of other types of atoms, the solvent accessible surface area, the average charge, the number of hydrogen bond donors, the number of hydrogen bond acceptors, the number of aromatic rings, etc. These physical quantities cover multiple dimensions such as the atomic level, the intermolecular force level, and the chemical group level. Compared with traditional methods that only rely on a single or a few characteristics, they can more comprehensively reflect the structural characteristics of proteins and provide a richer information basis for accurately predicting protein thermal stability point mutations.

[0074] This application sets the selection range of physical quantities around the target residue within a spherical region of 5 to 15 angstroms, and introduces a distance weighting function. This distance range is optimized to balance information integrity and noise suppression, covering key interacting atoms and groups while avoiding the introduction of too much irrelevant information. The weighting function assigns weights based on the distance between the atom and the target residue, highlighting the important influence of nearby atoms on thermal stability, allowing the input data to more accurately reflect local environmental characteristics, helping the neural network learn the impact of local changes on thermal stability, and improving the generalization ability of the model.

[0075] (2) Convolutional neural network module

[0076] It is composed of multiple groups of 3D convolutional layers and 3D pooling layers stacked in order, and then connected to multiple layers of fully connected layers, and finally outputs 20 values ​​corresponding to 20 standard amino acids. This module uses the full set or subset of data from a public protein structure database (such as the RCSB protein structure database) for training. The input four-dimensional matrix is ​​feature extracted through the 3D convolutional layer, the 3D pooling layer performs dimensionality reduction, and the fully connected layer integrates feature information to achieve the predicted output of amino acids.

[0077] By processing multi-dimensional data, convolutional neural networks can extract effective feature representations and use them to predict the impact of mutations in amino acid residues on thermal stability. In the convolutional neural network module, the three-dimensional convolution layer processes the input four-dimensional matrix data and can identify different structural patterns. In order to further improve the generalization ability of the network, normalization methods such as batch normalization and layer normalization are used during training to ensure that the input data is always balanced during training, thereby avoiding problems such as gradient vanishing or gradient exploding.

[0078] During the training of the convolutional neural network, the optimized loss function is as follows:

[0079]

[0080] Among them, L(θ) represents the total loss function, represents the mutation probability predicted by the network, y i represents the actual mutation probability, θ jRepresents the weight parameter of the network, N represents the total number of training samples, λ represents the regularization term coefficient, and m represents the number of parameters in the network. This loss function not only calculates the difference between the predicted value and the actual value through the mean square error (MSE), but also controls the complexity of the model through the regularization term to prevent overfitting.

[0081] In this way, the convolutional neural network is able to learn a wide range of protein structural features from the RCSB database, thereby achieving higher accuracy in mutation prediction for specific proteins.

[0082] In the convolutional neural network module, the filter (convolution kernel) size of the three-dimensional convolution layer is 3×3×3 or 2×2×2, and the step size is set to 1. Figure 4 The figure shows the structure of the convolutional neural network module. This setting effectively controls the computational complexity and ensures the training and prediction efficiency of the network while ensuring the effective extraction of local features in the three-dimensional structure of proteins. The filter size and step size of the three-dimensional convolution layer are crucial to extracting the structural features of proteins, especially when dealing with high-dimensional and complex structures. Appropriate filter design can help the network better learn local and global structural information.

[0083] Specifically, the 3×3×3 convolution kernel can capture the interaction information between atoms in the protein molecule over a larger spatial range. It can simultaneously take into account the spatial relationships and structural patterns in a larger range around the target residue, which helps to reveal more complex interactions between different atoms and amino acids. For example, this convolution kernel helps capture long-range interactions or large-scale environmental effects, which is crucial for understanding the thermal stability of proteins. In contrast, the 2×2×2 convolution kernel focuses on local detail features. It captures finer atomic arrangements and local chemical environments around the target residues by sliding over a smaller range, thereby providing the network with higher-resolution local information. The combined use of these two convolution kernels can comprehensively describe the three-dimensional structural features of proteins from a global and local perspective, respectively, ensuring that the model can accurately learn the key structural features that affect thermal stability and function.

[0084] In the implementation process, the convolution layer performs convolution operation on the input matrix through sliding window operation, and the convolution kernel performs element-level multiplication and addition operation with each local area of ​​the input to obtain the eigenvalue of the corresponding position. Set the convolution kernel size to f×f×f, the step size to s, and the size of the input matrix to I x ×I y ×I z , the size of the convolution output matrix is ​​O x ×O y ×O z It can be calculated by the following formula:

[0085]

[0086] Among them, I x ,I y ,I z Respectively represent the size of the input matrix in three dimensions, f is the size of the convolution kernel, s is the step size, O x ,O y ,O z is the size of the convolution output matrix in the corresponding dimension. By adjusting the size f and step size s of the convolution kernel, the receptive field of the convolution operation (i.e., the size of the area where the convolution kernel slides on the input matrix) can be controlled, thereby effectively capturing structural features of different scales.

[0087] This convolution kernel design can not only capture the microscopic features of local atomic interactions in proteins, but also learn the global relationships between atoms on a larger scale. By introducing multiple convolutional layers in the neural network and using convolution kernels of different sizes for hierarchical feature extraction, the network can gradually aggregate and refine structural information at different levels from local to global, thereby better reflecting the thermal stability and function of proteins in the final prediction.

[0088] The three-dimensional convolution layer not only effectively reduces the amount of calculation through reasonable convolution kernel size and step size settings, but also can efficiently extract the multi-scale features of the protein structure, providing strong data support for subsequent feature integration and thermal stability prediction.

[0089] The filter size of the three-dimensional convolution layer is 3×3×3 or 2×2×2, with a step size of 1. This setting can effectively extract the local features of the three-dimensional structure of the protein while controlling the computational complexity. The 3×3×3 convolution kernel can capture a wide range of atomic interactions and reveal complex relationships; the 2×2×2 convolution kernel focuses on local details. The combination of the two can comprehensively describe the structural features from different scales, ensuring that the model learns key information. The reasonable setting of the receptive field of the convolution operation can effectively reduce the amount of calculation and improve the efficiency of training and prediction.

[0090] Multiple groups of 3D convolutional layers and 3D pooling layers are stacked in an orderly manner and then connected to multiple layers of fully connected layers. The 3D pooling layer can reduce the data dimension, reduce the amount of calculation and retain important features. The fully connected layer further integrates the feature information to achieve accurate prediction output of amino acids. This structural design enables the network to gradually aggregate and refine information at different levels of structure, enhancing its ability to learn protein structural features.

[0091] (3) Transfer learning neural network module

[0092] The output of the previous layer of the convolutional neural network module output layer is used as the input of this module, which is composed of several layers of fully connected layers, one-dimensional convolutional layers-one-dimensional pooling layers or normalization layers-fully connected layers. The final output is 20 values ​​corresponding to 20 standard amino acids. The four-dimensional matrix obtained by feature engineering of the unmutated structure of the experimental protein is used as input, and the experimental observable physical quantities such as the protein melting temperature, dynamic light scattering particle size or enzymatic reaction Michaelis constant of the experimental mutant that characterize the thermal stability of the protein are normalized and output as the mutation probability of the corresponding amino acid. This module uses the characteristics of transfer learning and combines experimental data to further optimize the prediction of amino acid mutation probability, effectively improving the model's prediction accuracy for mutations related to the thermal stability of specific proteins.

[0093] The aforementioned convolutional neural network module can be trained to obtain the initial protein thermal stability point mutation prediction model. In order to further optimize the learning efficiency and accuracy of the initial protein thermal stability point mutation prediction model, the diversity of protein data in the RCSB database is used in the transfer learning neural network module to help the network quickly adapt to the characteristics of different proteins. In the process of transfer learning, the pre-trained model learns a large amount of standard protein data, so that the model can quickly capture potential structural patterns, and further refine the predictions based on experimental data. In this way, the transfer learning neural network module can improve the prediction accuracy of specific mutations, especially for some unseen protein structures, and can achieve better generalization.

[0094] In the transfer learning neural network module, the normalization method adjusts the probability of amino acid mutation by combining enzyme activity experimental data to ensure that the model's prediction results better reflect the actual functional activity of the protein.

[0095] Specifically, the normalization process is carried out through the following steps:

[0096] First, without any external correction, it is assumed that the initial mutation probability of all amino acids is 100%, that is, the potential for mutation of each amino acid is equal at the initial stage. This initial setting provides a unified starting point for subsequent corrections. Next, the model adjusts the mutation probability of each amino acid based on experimental data, especially the relative activity percentage of enzyme activity experiments. For example, assuming that the experimental data shows that the mutation of a certain amino acid significantly improves the activity of the enzyme, the mutation probability corresponding to the amino acid will increase accordingly. On the contrary, if a mutation leads to a decrease in enzyme activity, the mutation probability of the amino acid will be reduced. Finally, in order to ensure that the sum of all probabilities is 100%, all adjusted mutation probabilities will be standardized, that is, the final mutation probability of each amino acid is calculated by dividing its corrected probability value by the sum of all amino acid mutation probabilities.

[0097] This normalization method combines experimental data on enzyme activity, so that the model prediction not only focuses on the changes in protein thermal stability, but also considers the effects of mutations on enzyme activity and its function. Specifically, let p i is the initial mutation probability of the ith amino acid, a i is the relative activity percentage of the corresponding amino acid in the experimental data, the normalized amino acid mutation probability The calculation formula is:

[0098]

[0099] Among them, a i represents the relative activity percentage of the ith amino acid associated with the experimental data, N is the total number of all amino acids, is the sum of the relative activity percentages of all amino acids. represents the normalized amino acid mutation probability in the scenario with enzyme activity experimental data. Through this calculation, the normalized It ensures that the sum of the probabilities of all amino acids is 100%, and can reasonably distribute the mutation probabilities of different amino acids based on the enzyme activity data.

[0100] This normalization method enables the transfer learning neural network module to optimize the probability prediction of amino acid mutations based on enzyme activity experimental data, and more accurately reflect the comprehensive impact of different mutations on protein thermal stability and functional activity. In practical applications, using this normalization method, the recommendation system can not only optimize the thermal stability of proteins, but also maintain the enzyme activity and function after mutation, thereby providing more practical mutation recommendations in protein engineering.

[0101] In the normalization method in the transfer learning neural network module, in addition to adjusting the amino acid mutation probability in combination with the enzyme activity experimental data, amino acids without experimental data are also considered. For those amino acids without experimental data, their mutation probability is set to 0% to avoid unnecessary interference with the final prediction results. At the same time, for unmutated amino acids, their mutation probability is set to 100% as a baseline value to ensure that in the absence of mutations, the model's prediction can be referenced to its initial state. This processing method ensures that the normalized probability distribution can more reasonably reflect the mutation situation by highlighting the influence of amino acids for which experimental data has been obtained and combining the baseline probability of unmutated amino acids.

[0102] Specifically, assuming that the relative activity percentage corresponding to amino acid i in the experimental data is a i , where if there is no experimental value (i.e. a i =0), the mutation probability of the amino acid p iis set to 0%. For unmutated amino acids, the mutation probability is set to 100%. Then, the amino acid mutation probability of all experimental data will be normalized according to the following formula:

[0103]

[0104] in, represents the normalized amino acid mutation probability in the scenario of no enzyme activity experimental data, p i is the mutation probability after correction by experimental data (0% if there is no experimental data), is the sum of all amino acid mutation probabilities. The normalized probability value The probability of each amino acid was guaranteed to be reasonable, and in the absence of experimental data, its mutation probability was restricted to 0%, while the non-mutated amino acids were kept at 100%.

[0105] The purpose of this normalization method is to emphasize the information provided by the experimental data while keeping the baseline probability of non-mutated amino acids as a reference. In this way, the transfer learning neural network module can more accurately predict the impact of amino acid mutations on protein thermal stability and functional activity, especially in the absence of comprehensive experimental data, and can still make reasonable inferences based on known experimental information and baseline values. This not only enhances the practical application value of the model, but also makes the distribution of mutation probabilities more consistent with biological significance, especially in functional prediction and protein engineering, it can provide more targeted and practical results.

[0106] In the transfer learning neural network module, the fully connected layer uses ReLU (rectified linear unit) as the activation function. This choice plays an important role in the training and performance optimization of the model. The mathematical expression of ReLU is:

[0107] ReLU(x)=max(0,x);

[0108] Where x is the input signal. The main advantage of the ReLU function is that it is computationally simple and has nonlinear characteristics, which enables it to effectively process large-scale data while avoiding the problem of gradient vanishing. The gradient vanishing problem usually occurs when using traditional activation functions such as Sigmoid or Tanh. When the gradient of these functions is small, the training speed of the network will be significantly slowed down, and even the parameters cannot be effectively updated. In contrast, the gradient of the ReLU function is constant when the input is positive and does not depend on the size of the input, so it can accelerate the convergence of the network.

[0109] Specifically, the output of the ReLU function is non-negative, which makes it very suitable for probability-related tasks. In the transfer learning neural network module, the output probability of each amino acid mutation should be a non-negative value, and the sum should be 1 (or close to 1) to meet the basic requirements of probability. The non-negativity of the ReLU function ensures that the output of each node meets this requirement, which helps to ensure that the output results of the transfer learning neural network are reasonable and effective.

[0110] By introducing the ReLU activation function in the fully connected layer, the network can quickly adjust parameters during training, increasing its ability to learn mutations related to the thermal stability of specific proteins. The nonlinear characteristics of ReLU enable the network to capture more complex patterns in the input data, especially when dealing with high-dimensional data with multiple hierarchical structures. ReLU can effectively enhance the expressive power and complexity of the model, thereby improving the final prediction accuracy.

[0111] In order to further optimize the performance of the model, the ReLU activation function is also combined with other technologies such as batch normalization and regularization methods to ensure that the network can be stably trained when facing a large amount of data, avoid overfitting, and ensure that the model has good generalization ability. During the training process, the gradient of ReLU will be effectively propagated to the previous layer, so that the network can quickly and effectively learn the complex relationship between the three-dimensional structural characteristics and thermal stability of proteins.

[0112] The ReLU activation function not only has obvious advantages in reducing the amount of calculation and accelerating the convergence speed, but also its non-negative output characteristics and good gradient propagation characteristics make it play a key role in the protein thermal stability prediction model, especially in the transfer learning neural network module. Through the ReLU activation function, the network can maintain high efficiency and accuracy in the process of large-scale data processing and complex pattern learning, ensuring that the final output is in line with the rationality of biology and probability.

[0113] The transfer learning neural network module takes the output of the previous layer of the convolutional neural network module output layer as input, and uses structures such as fully connected layer stacking, one-dimensional convolution layer-one-dimensional pooling layer stacking, or normalization layer-fully connected layer stacking, and pre-training with public database data, and then fine-tuning with experimental protein data (such as protein melting temperature, dynamic light scattering particle size, or enzymatic reaction Michaelis constant, etc.), which effectively improves the prediction accuracy of mutations related to thermal stability of specific proteins. Especially when data is limited, the transfer learning characteristics can quickly adapt and optimize predictions. Figure 5 Shown is the operation flow chart of the transfer learning neural network module.

[0114] The fully connected layer of the transfer learning neural network module uses the ReLU activation function, which is simple to calculate and has strong nonlinear characteristics, avoiding the gradient vanishing problem and accelerating network convergence. At the same time, the output of the ReLU function is non-negative, which meets the mutation probability requirements of amino acids, helps to ensure the rationality of the output results, and can be used with other technologies to stabilize training, enhance the model's expression and generalization capabilities, and enable the network to better learn the relationship between protein structure and thermal stability.

[0115] In the transfer learning neural network module, the normalization method adjusts the probability of amino acid mutation by combining the enzyme activity experimental data. Initially, it is assumed that the probability of all amino acid mutations is 100%, and then adjusted according to the changes in enzyme activity in the experimental data. If the mutation increases the enzyme activity, the probability increases, and vice versa. Finally, the standardization ensures that the total probability is 100%. This method allows the prediction to take into account the impact of mutations on enzyme activity and function, optimizes the mutation probability prediction, and more accurately reflects the comprehensive impact.

[0116] For amino acids without experimental data, the mutation probability is set to 0% to avoid interference; the non-mutated amino acids are set to 100% as the benchmark value. By highlighting the experimental data information and combining the benchmark value, the normalized probability distribution is more reasonable. In the absence of comprehensive experimental data, reasonable inferences can still be made based on known information, thereby enhancing the value of the model in practical applications and providing more targeted and practical mutation suggestions.

[0117] (4) Recommended Unit

[0118] Based on the output results of the convolutional neural network module or the transfer learning neural network module, mutation sites and target amino acids to be mutated are recommended according to specific criteria. The criteria include sites where the standard deviation of the probabilities of several standard amino acids predicted by the model is higher than the preset value, i.e., sites with high uncertainty; and sites where the probability corresponding to the actual amino acids in the model-predicted structure is not higher than the set value, i.e., sites that do not match the experimental results. In this way, sites with high mutation potential and corresponding target amino acids are screened out, providing targeted guidance for the modification of protein thermal stability.

[0119] In the recommendation unit, in order to recommend mutation sites and target amino acids according to the criteria, the system first evaluates the standard deviation of each candidate site and the mutation probability of the actual amino acid, and then ranks the sites that meet the criteria. Figure 6The following is a logic diagram for sorting recommended units. Specifically, if there are multiple sites with standard deviations higher than the preset value, the system will sort and recommend these sites in descending order of standard deviation; if there are multiple sites with actual amino acid mutation probabilities not higher than the set value, they will be sorted and recommended in descending order of mutation probabilities. The core goal of this sorting method is to give priority to sites with higher mutation potential and sites whose experimental results are inconsistent with the predicted amino acid probability, so as to provide a more reasonable and prioritized mutation selection order for experimental research, thereby improving experimental efficiency and success rate.

[0120] To further quantify this ranking method, the standard deviation and mutation probability can be ranked and prioritized by the following formula. Assume that for the i-th candidate site, the standard deviation is σ i , the actual probability of amino acid mutation is p i , the system handles sorting based on the following two criteria:

[0121] Standard deviation sorting: For sites with larger standard deviations, it reflects that the uncertainty of the site is higher and may have greater potential in structure prediction. The sorting criteria are set as follows:

[0122]

[0123] Among them, σ n represents the standard deviation corresponding to the candidate site n, Sort values ​​by standard deviation. Sites with larger standard deviations are given higher priority.

[0124] Mutation probability ranking: For sites with a lower mutation probability, it means that the actual amino acid at the site deviates more from the model-predicted amino acid and may have a higher potential for optimization. The ranking criteria are set as follows:

[0125]

[0126] Among them, p i is the mutation probability of the ith site, The mutation probability ranking value. Sites with lower mutation probability will be given higher priority.

[0127] Through these two sorting methods, the final recommendation order can be obtained by combining the standard deviation sorting value and the mutation probability sorting value. The comprehensive sorting value is set as follows:

[0128]

[0129] Among them, w σ and w p are the weights of the standard deviation ranking value and the mutation probability ranking value, respectively, w σ+w p = 1. This weighted comprehensive ranking method can adjust the weights according to the actual situation to balance the impact of the two factors of standard deviation and mutation probability on the ranking results.

[0130] The recommended sites after sorting will be arranged according to the comprehensive sorting value, and sites with larger standard deviations and smaller mutation probabilities will be recommended first, so as to ensure that the recommended results give priority to sites with higher mutation potential and those that are inconsistent with experimental results. Through this sorting method, more priority and targeted mutation sites can be provided for experimental research, thereby improving the rationality of experimental design, optimizing the research process and increasing the success rate.

[0131] The recommendation unit recommends mutation sites and target amino acids based on the standard deviation of amino acid probability predicted by the model and the criteria set for the actual amino acid mutation probability. Sites with standard deviations higher than the preset value (high uncertainty) and actual amino acid probabilities not higher than the set value (inconsistent with experimental results) are ranked and recommended. Through a quantitative ranking formula, the sites are ranked according to the size of the standard deviation and the size of the mutation probability, and the weights of the two are set to obtain the final recommendation order. Sites with higher mutation potential and sites that are inconsistent with experimental results are given priority, thereby improving the accuracy and rationality of the recommendations, providing more priority and targeted mutation selections for experimental research, and helping to improve experimental efficiency and success rate.

[0132] The synergistic effect of the above-mentioned improvements, from feature engineering to provide precise input, neural network architecture optimization and training to enhance learning ability, normalization method to optimize probability prediction to precise recommendation of recommendation units, has comprehensively improved the system's prediction accuracy for protein thermal stability point mutations, and can more effectively guide thermal stability modification work in protein engineering, reduce experimental blindness, and save time and resources.

[0133] The use of public databases (such as the RCSB protein structure database) and the flexibility of the neural network architecture make this scheme not only applicable to various types of proteins, but also easily extendable to other related protein property research or engineering applications, with strong versatility and potential for application expansion.

[0134] The above is a separate introduction to each module unit. The following is a series introduction of each module unit in combination with two more specific embodiments to reveal the situation of the model in specific applications.

[0135] 1. A more specific embodiment of the present application

[0136] 1. Data preparation

[0137] The RCSB protein structure database is used to obtain the 3D structure data of various proteins, including but not limited to enzymes, antibodies, structural proteins and other different types of proteins.

[0138] For each protein, the target residue is determined, and physical quantities are selected within a spherical region with a radius of 5 to 15 angstroms centered on the target residue, such as the number of carbon atoms, the number of nitrogen atoms, the number of other types of atoms, the solvent accessible surface area, the average charge, the number of hydrogen bond donors, the number of hydrogen bond acceptors, and the number of aromatic rings.

[0139] The selected physical quantities are rasterized to form multiple three-dimensional matrices, which are then merged into four-dimensional matrices using the types of physical quantities as channels as input data for the convolutional neural network module.

[0140] 2. Training of Convolutional Neural Network Module

[0141] Construct a convolutional neural network module, set the filter size of the three-dimensional convolution layer to 3×3×3 and 2×2×2, with a step size of 1, and stack multiple groups of three-dimensional convolutional layers and three-dimensional pooling layers in order, followed by multiple layers of fully connected layers.

[0142] Use the entire set or subset of public protein structure database data for training. During the training process, batch normalization and layer normalization methods are used to optimize the loss function (including mean square error calculation and regularization terms to control model complexity) to ensure that the model can learn protein structural features, avoid overfitting, and improve generalization ability.

[0143] 3. Training and optimization of transfer learning neural network modules

[0144] The output of the previous layer of the output layer of the convolutional neural network module is used as the input of the transfer learning neural network module to construct a network structure consisting of several layers of fully connected layer stacking, one-dimensional convolution layer-one-dimensional pooling layer stacking, or normalization layer-fully connected layer stacking.

[0145] The four-dimensional matrix obtained by feature engineering of the unmutated structure of the experimental protein is used as input, and the experimental observable physical quantities characterizing the thermal stability of the protein, such as the protein melting temperature, dynamic light scattering particle size or Michaelis constant of the enzymatic reaction of the experimental mutant, are normalized and output as the mutation probability of the corresponding amino acid.

[0146] The fully connected layer uses the ReLU activation function, combined with batch normalization and regularization methods, to accelerate network convergence, enhance the expressiveness and generalization capabilities of the model, enable the model to better learn the relationship between protein structure and thermal stability, and improve the prediction accuracy of thermal stability-related mutations of specific proteins.

[0147] 4. Application of Recommendation Unit

[0148] The structural data of the protein to be studied is input into the trained system and processed through the convolutional neural network module and the transfer learning neural network module to obtain the probability of amino acid mutation at each site.

[0149] The recommendation unit sets criteria based on the standard deviation of the amino acid probability predicted by the model and the actual amino acid mutation probability, and screens out sites with a standard deviation higher than a preset value (such as 0.5) and sites with an actual amino acid probability not higher than a set value (such as 0.2).

[0150] The selected sites are sorted according to the quantitative sorting formula, sorted according to the standard deviation size and mutation probability size, and the weights are set comprehensively (such as the standard deviation sorting weight is 0.6, and the mutation probability sorting weight is 0.4) to obtain the final recommended order.

[0151] According to the final recommendation order, the top-ranked mutation sites are given priority for experimental research, such as conducting site-directed mutagenesis experiments on the top 5 sites to measure the thermal stability changes of the mutants and verify the accuracy and effectiveness of the recommendation system.

[0152] 2. Another more specific embodiment of the present application

[0153] 1. Research on specific proteins (taking a protein with a specific function as an example)

[0154] Select a protein with a specific function (such as a certain oxidoreductase) and obtain its three-dimensional structure data from the RCSB protein structure database.

[0155] The residues in the key functional region of the protein are taken as target residues. According to the method of feature engineering unit, physical quantities are selected and processed in a spherical region of 5 to 15 angstroms around the target residues to obtain four-dimensional matrix input data.

[0156] Using the existing experimental data related to the protein (such as enzyme activity experimental data), in the transfer learning neural network module, the amino acid mutation probability is adjusted through the normalization method combined with the experimental data. The mutation probability of amino acids without experimental data and non-mutated amino acids is processed according to the set rules to make the prediction more in line with the actual situation.

[0157] After being processed by the convolutional neural network module and the transfer learning neural network module, the recommendation unit recommends mutation sites and target amino acids based on multi-criteria sorting, and provides targeted mutation suggestions for the thermal stability modification of the specific protein, such as recommending that several specific sites be mutated to specific amino acids to improve its thermal stability while maintaining enzyme activity as much as possible.

[0158] 2. Model verification and optimization

[0159] Mutants were prepared experimentally according to the recommended sites, and the thermal stability changes of the mutants were determined by differential scanning calorimetry, dynamic light scattering and other experimental methods, and the changes in enzyme activity were determined by enzyme activity detection experiments.

[0160] Compare the experimental results with the prediction results of the recommendation system to evaluate the prediction accuracy. For example, if the thermal stability of a certain site is significantly improved after mutation, and the experimental verification results are consistent, it means that the model's prediction at this site is valid; if there is a deviation, analyze the cause, which may be that the physical quantities selected by the feature engineering are not comprehensive enough, the structure or parameters of the neural network model need to be adjusted, etc.

[0161] According to the verification results, the model is optimized, such as adjusting the selection weights of physical quantities in the feature engineering unit, optimizing the convolutional layer parameters of the convolutional neural network module, improving the combination structure of the fully connected layer and other layers in the transfer learning neural network module, etc., and then trained and verified again to continuously improve the model's predictive performance for protein thermal stability point mutations, so that it can more accurately guide experimental research and practical applications in protein engineering.

[0162] Figure 7 A block diagram of an electronic device suitable for implementing the protein thermal stability point mutation prediction model construction method and / or prediction method described above according to an embodiment of the present application is schematically shown. Figure 7 The electronic device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0163] like Figure 7 As shown, the electronic device 1000 described in this embodiment includes: a processor 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage part 1008 into a random access memory (RAM) 1003. Processor 1001 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), and the like. Processor 1001 may also include an onboard memory for caching purposes. Processor 1001 may include a single processing unit or multiple processing units for performing different actions of the protein thermal stability point mutation prediction model construction method and / or prediction method flow according to an embodiment of the present application.

[0164] In RAM 1003, various programs and data required for the operation of electronic device 1000 are stored. Processor 1001, ROM 1002 and RAM 1003 are connected to each other via bus 1004. Processor 1001 performs various operations of the protein thermal stability point mutation prediction model construction method and / or prediction method flow according to the embodiment of the present application by executing the program in ROM 1002 and / or RAM 1003. It should be noted that the program can also be stored in one or more memories other than ROM 1002 and RAM 1003. Processor 1001 can also perform various operations of the protein thermal stability point mutation prediction model construction method and / or prediction method flow according to the embodiment of the present application by executing the program stored in the one or more memories.

[0165] According to an embodiment of the present application, the electronic device 1000 may further include an input / output (I / O) interface 1005, which is also connected to the bus 1004. The electronic device 1000 may further include one or more of the following components connected to the I / O interface 1005: an input portion 1006 including a keyboard, a mouse, etc.; an output portion 1007 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 1008 including a hard disk, etc.; and a communication portion 1009 including a network interface card such as a LAN card, a modem, etc. The communication portion 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1010 as needed, so that a computer program read therefrom is installed into the storage portion 1008 as needed.

[0166] According to the protein thermal stability point mutation prediction model construction method and / or prediction method flow process of the embodiment of the present application can be implemented as a computer software program. For example, the embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes a program code for executing the protein thermal stability point mutation prediction model construction method and / or prediction method shown in the flow chart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 1009, and / or installed from the removable medium 1011. When the computer program is executed by the processor 1001, the above-mentioned functions defined in the system of the embodiment of the present application are executed. According to an embodiment of the present application, the system, equipment, device, module and / or unit described above, etc. can be realized by a computer program module.

[0167] The embodiments of the present application also provide a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the steps of the protein thermal stability point mutation prediction model construction method and / or prediction method according to the embodiments of the present application can be implemented.

[0168] According to an embodiment of the present application, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In an embodiment of the present application, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, an apparatus, or a device. For example, according to an embodiment of the present application, the computer-readable storage medium may include one or more memories other than the ROM 1002 and / or RAM 1003 described above.

[0169] It should be noted that the functional modules in the various embodiments of the present application can be integrated into a processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules. If the integrated modules are implemented in the form of software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product.

[0170] The flowchart and / or block diagram in the accompanying drawings illustrate the possible architecture, functions and operations of the system, method and computer program product according to various embodiments of the present application. In this regard, each box in the flowchart and / or block diagram can represent a part of a module, program segment or code, and a part of the above-mentioned module, program segment or code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. It should also be noted that each box in the block diagram or flowchart, and the combination of the boxes in the block diagram or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0171] Those skilled in the art will appreciate that the features described in the various embodiments and / or claims of the present application may be combined and / or combined in a variety of ways, even if such combinations and / or combinations are not explicitly described in the present application. In particular, without departing from the spirit and teachings of the present application, the technical features described in the various embodiments and / or claims of the present application may be combined and / or combined in a variety of ways, and all of these combinations and / or combinations fall within the scope of the present application.

[0172] Although the present application has been shown and described with reference to specific exemplary embodiments of the present application, it should be understood by those skilled in the art that various changes in form and details may be made to the present application without departing from the spirit and scope of the present application as defined by the appended claims and their equivalents. Therefore, the scope of the present application should not be limited to the above-mentioned embodiments, but should be determined not only by the appended claims, but also by the equivalents of the appended claims.

Claims

1. A method for constructing a protein thermal stability point mutation prediction model, characterized in that: include: Obtaining a three-dimensional structure of a preset type of protein, selecting a preset type of physical quantity within a preset distance range around the target residue for rasterization processing to form multiple three-dimensional matrices, and merging them into a four-dimensional matrix using the physical quantity type as a channel; Using the four-dimensional matrix as training data, training a three-dimensional convolutional neural network to construct an initial protein thermal stability point mutation prediction model; the three-dimensional convolutional neural network includes a preset large-size convolution kernel for extracting global features and a preset small-size convolution kernel for extracting local features; The output of the previous layer of the output layer of the initial protein thermal stability point mutation prediction model is used as input, and a normalization method is used to adjust the amino acid mutation probability in combination with enzyme activity experimental data and / or keep the baseline probability of non-mutated amino acids in the absence of enzyme activity experimental data as a reference, to train a transfer learning neural network and obtain an optimized final protein thermal stability point mutation prediction model.

2. The prediction model construction method according to claim 1, characterized in that: The preset distance range is specifically: The region between the spherical regions with a radius of 5 to 15 angstroms centered on the target residue.

3. The prediction model construction method according to claim 2, characterized in that: The distance to the target residue is weighted by the distance weighting function w(r), and the specific formula includes: Where r represents the distance between the target residue and the surrounding atoms; α represents the adjustment parameter; and w(r) represents the weighted distance to the target residue.

4. The prediction model construction method according to claim 1, characterized in that: The preset types of physical quantities include physical quantities that reflect the structural characteristics of proteins from the atomic level, the intermolecular force level, and the chemical group level; The preset types of physical quantities specifically include one or more of the number of carbon atoms, the number of nitrogen atoms, the solvent accessible surface area, the average charge value, the number of hydrogen bond donors, the number of hydrogen bond acceptors, and the number of aromatic rings.

5. The prediction model construction method according to claim 1, characterized in that: The sizes of the preset large-size convolution kernel and the preset small-size convolution kernel of the three-dimensional convolution layer of the three-dimensional convolutional neural network are 3×3×3 and 2×2×2 respectively, and the step size is set to 1.

6. The prediction model construction method according to claim 5, characterized in that: The three-dimensional convolution layer performs a convolution operation on the input matrix through a sliding window operation. The convolution kernel performs element-level multiplication and addition operations with each local area of ​​the input to obtain the eigenvalue of the corresponding position. The specific formula includes: Among them, I x ,I y ,I z Respectively represent the size of the input matrix in three dimensions; f represents the size of the convolution kernel; s represents the step size; O x , O y , O z They represent the size of the convolution output matrix in the corresponding dimension.

7. The prediction model construction method according to claim 1, characterized in that: Normalization methods for adjusting amino acid mutation probabilities based on enzyme activity experimental data include: Assume that the initial mutation probability of all amino acids is 100%; If the experimental data show that the mutation of a certain amino acid increases the activity of the enzyme, then the mutation probability corresponding to the amino acid will increase accordingly; otherwise, it will decrease accordingly; The adjusted mutation probabilities are normalized so that the sum of all mutation probabilities is 100%; the specific formula includes: Among them, a i represents the relative activity percentage of the ith amino acid associated with the experimental data; N represents the total number of all amino acids; It represents the sum of the relative activity percentages of all amino acids; Represents the normalized amino acid mutation probability in the scenario with enzyme activity experimental data; Normalization methods that keep the baseline probability of unmutated amino acids as a reference when there is no enzyme activity experimental data include: The mutation probability of amino acids with no enzyme activity experimental data was set to 0%; The mutation probability of unmutated amino acids was set to 100% as a baseline value, thus ensuring that the model's predictions could be referenced to its initial state in the absence of mutations; The specific formulas include: Among them, p i represents the mutation probability after correction by experimental data; represents the sum of all amino acid mutation probabilities; Represents the normalized amino acid mutation probability in the scenario of no enzyme activity experimental data.

8. The prediction model construction method according to claim 1, characterized in that: The fully connected layer of the transfer learning neural network uses ReLU as the activation function; The loss function used when training the three-dimensional convolutional neural network includes: Among them, L(θ) represents the total loss function; represents the mutation probability predicted by the network; y i represents the actual mutation probability; θ j represents the weight parameter of the network; N represents the total number of training samples; λ represents the regularization term coefficient; m represents the number of parameters in the network.

9. A method for predicting protein thermal stability point mutations, characterized in that: The structural data of the protein to be studied is input into the initial protein thermal stability point mutation prediction model or the final protein thermal stability point mutation prediction model in the prediction model construction method according to any one of claims 1 to 8 for processing to predict the mutation site and the target amino acid to be mutated.

10. The prediction method according to claim 9, characterized in that: Evaluate the standard deviation and actual amino acid mutation probability of each candidate site, and rank the sites that meet the criteria; If there are multiple sites with standard deviations higher than the first preset value, these sites are sorted and recommended in descending order of standard deviation; The sorting criteria for standard deviation include: Among them, σ n represents the standard deviation corresponding to the candidate site n; Indicates the standard deviation ranking value; If the actual amino acid mutation probability of multiple sites is not higher than the second preset value, the sites are ranked and recommended in ascending order of mutation probability; the ranking criteria for mutation probability include: Among them, p i represents the mutation probability of candidate site i; Represents the mutation probability ranking value; Based on the weights of the standard deviation ranking value and the mutation probability ranking value, the final recommendation order is determined. The specific formula includes: Among them, w σ and w p They represent the weight of the standard deviation ranking value and the weight of the mutation probability ranking value, respectively, σ +w p =1;Rank i Indicates the comprehensive ranking value; Based on the comprehensive ranking values, mutation sites and target amino acids to be mutated are recommended, and mutation sites that are ranked high based on the comprehensive ranking values ​​are preferentially selected for experimental research.

Citation Information

Patent Citations

  • Protein secondary structure prediction method based on multi-scale convolution attention neural network

    CN112767997A

  • GPCR thermal stability mutation prediction model construction method, prediction method and device

    CN114913914A

  • Protein stability prediction model construction method, prediction method and system

    CN117476104A

  • Thermal stability new enzyme design and transformation method based on deep learning model

    CN118782150A

  • Phi29 DNA polymerase mutant having increased thermal stability and use thereof

    US20200208126A1

Cited By

  • Protein phase separation prediction method and system based on physical information neural network, electronic equipment and storage medium

    CN122266477A