Methods, apparatus, and articles of manufacture to predict protein structure energy
Patent Information
- Application Number
- CN202410659393.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-27
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2044-05-27
AI Technical Summary
[0003]目前,现有技术的蛋白质结构能量预测方法至少存在在预测准确度较低的问题
[0011] As described above, the pre-trained energy prediction model in this embodiment can not only efficiently distinguish between steady-state and unfolded protein structures, but also assign lower energy values to protein structures with high accuracy, thus demonstrating good performance in protein design accuracy. Therefore, using the method for predicting protein structure energy provided in this embodiment, the protein structure energy of target protein structure data can be predicted relatively quickly and accurately, yielding effective prediction results.
Smart Images

Figure CN118447918B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of protein energy prediction technology, and in particular to a method, apparatus, electronic device, non-transitory computer-readable storage medium, and computer program product for predicting protein structural energy. Background Technology
[0002] In the development of antiviral drugs, protein engineering is required for drug design; among which, computer-aided protein design strategies are crucial; these computer-aided protein design strategies rely on accurate protein structure energy prediction methods (i.e., energy functions).
[0003] Currently, existing protein structure-energy prediction methods suffer from low prediction accuracy.
[0004] Therefore, how to provide a solution that can "predict the energy of protein structure relatively quickly and accurately" has become an urgent problem to be solved. Summary of the Invention
[0005] This disclosure provides a method, apparatus, electronic device, non-transitory computer-readable storage medium, and computer program product for predicting protein structure energy, in order to address the deficiencies in the prior art.
[0006] This disclosure provides a method for predicting protein structure energy, comprising: acquiring target protein structure data to be processed; wherein the target protein structure data includes main chain structure data and side chain structure data, and the protein structure types involved in the target protein structure data include steady state and unfolded state; inputting the target protein structure data into a pre-trained energy prediction model to obtain a protein structure energy matching the target protein structure data output by the pre-trained energy prediction model; wherein the pre-trained energy prediction model includes a cascaded pre-trained neural network and an output mapping unit, and the pre-trained neural network is trained using sample data constructed based on the protein structure data.
[0007] This disclosure also provides an apparatus for predicting protein structure energy, comprising: a data acquisition module configured to acquire target protein structure data to be processed; wherein the target protein structure data includes main chain structure data and side chain structure data, and the protein structure types involved in the target protein structure data include steady-state and unfolded states; and an energy prediction module configured to input the target protein structure data into a pre-trained energy prediction model to obtain a protein structure energy matching the target protein structure data output by the pre-trained energy prediction model; wherein the pre-trained energy prediction model includes a cascaded pre-trained neural network and an output mapping unit, and the pre-trained neural network is trained using sample data constructed based on the protein structure data.
[0008] This disclosure also provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method for predicting protein structure energy as described above.
[0009] This disclosure also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for predicting protein structure energy as described in any of the above.
[0010] This disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the method for predicting protein structure energy as described above.
[0011] As described above, the pre-trained energy prediction model in this embodiment can not only efficiently distinguish between steady-state and unfolded protein structures, but also assign lower energy values to protein structures with high accuracy, thus demonstrating good performance in protein design accuracy. Therefore, using the method for predicting protein structure energy provided in this embodiment, the protein structure energy of target protein structure data can be predicted relatively quickly and accurately, yielding effective prediction results. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a flowchart illustrating a method for predicting protein structure energy provided in an embodiment of this disclosure.
[0014] Figure 2 This is a schematic diagram of the structure of the pre-trained neural network provided in this embodiment of the present disclosure, in the case of "including three improved protein probability function GPAM-PPF sub-networks";
[0015] Figure 3 This is a schematic diagram of the structure of the GPAM-PPF subnetwork provided in the embodiments of this disclosure;
[0016] Figure 4 This is a schematic diagram of the structure of the pre-trained neural network provided in the embodiments of this disclosure, in the case of "including an improved protein probability function GPAM-PPF sub-network";
[0017] Figure 5 This is a schematic diagram of the device for predicting protein structure energy provided in this disclosure;
[0018] Figure 6 This is a schematic diagram of the structure of the electronic device provided in this disclosure. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this disclosure clearer, the technical solutions of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0020] Example:
[0021] The scheme for predicting protein structure energy disclosed herein will be described below with reference to the accompanying drawings.
[0022] Figure 1 This is a schematic flowchart of a method for predicting protein structure energy according to an exemplary embodiment of this disclosure. This embodiment can be applied to electronic devices (e.g., servers or cloud computing platforms), such as... Figure 1 As shown, the method for predicting protein structure energy includes the following steps:
[0023] S110. Obtain the structural data of the target protein to be processed.
[0024] The target protein structure data includes main chain structure data and side chain structure data, and the protein structure types involved in the target protein structure data include stable state and unfolded state.
[0025] The target protein structure data to be processed can be determined through relevant technical means in the field of protein engineering, and then stored in the form of a pre-defined type of file (e.g., a PDB file).
[0026] The execution entity of this embodiment (e.g., a server or cloud computing platform) can obtain the target protein structure data to be processed by communicating with the terminal storing the "target protein structure data to be processed" (e.g., a data storage server).
[0027] S120. Input the target protein structure data into the pre-trained energy prediction model to obtain the protein structure energy that matches the target protein structure data, output by the pre-trained energy prediction model.
[0028] The pre-trained energy prediction model includes a cascaded pre-trained neural network and an output mapping unit. The pre-trained neural network is trained using sample data constructed based on protein structure data.
[0029] The output mapping unit is mainly used to map the output of the pre-trained neural network into protein structure energy. The specific implementation details will be described later and will not be elaborated here.
[0030] It should be noted that the pre-trained neural network described in this embodiment is designed and determined based on the global knowledge distillation method or the marker point method. Specific implementation details will be described later and will not be repeated here.
[0031] As described above, the pre-trained energy prediction model in this embodiment can not only efficiently distinguish between steady-state and unfolded protein structures, but also assign lower energy values to protein structures with high accuracy, thus demonstrating good performance in protein design accuracy. Therefore, using the method for predicting protein structure energy provided in this embodiment, the protein structure energy of target protein structure data can be predicted relatively quickly and accurately, yielding effective prediction results.
[0032] exist Figure 1 Based on the embodiments, as an optional implementation, the input terminal of the pre-trained neural network serves as the input terminal of the pre-trained energy prediction model. The input terminal of the pre-trained neural network is connected to the input terminal of the output mapping unit, and the output terminal of the output mapping unit serves as the output terminal of the pre-trained neural network. The "cascaded" structure is formed from the pre-trained neural network and the output mapping unit.
[0033] exist Figure 1 Based on the embodiments, as an optional implementation, the pre-trained neural network can be implemented in the following way:
[0034] As an optional example, the pre-trained neural network is designed based on the global knowledge distillation method. The core idea of the global knowledge distillation method includes designing a teacher and student model approach: fixing the parameters of the teacher model, allowing both the teacher and student models to extract feature vectors from similar protein structure data, and using the distance between their feature vectors as the loss value to optimize the parameters of the student model. This allows the student model to learn the method used by the teacher model to extract feature vectors from this type of protein structure data, thus completing the statistical analysis of the preferences and constraints for this type of protein structure feature.
[0035] Specifically, refer to Figure 2 The design of a pre-trained neural network may include three improved protein probability function (GPAM-PPF) subnetworks, wherein one GPAM-PPF subnetwork is used as the teacher subnetwork (i.e., the teacher model in the figure), one GPAM-PPF subnetwork is used as the steady-state student subnetwork (i.e., ... Figure 2 (Medium steady-state student model), and a GPAM-PPF subnetwork used as the non-folded student subnetwork (i.e., Figure 2 (China-Africa Folded State Student Model).
[0036] The teacher subnetwork and the steady-state student subnetwork form a pair of training networks, and the teacher subnetwork and the non-folded student subnetwork form another pair of training networks.
[0037] During the training phase, the two pairs of training networks are trained on different categories of protein structure datasets to optimize the parameters of the student model. The training processes are independent of each other and do not interfere with each other. The two pairs of training networks are only connected through the teacher subnetwork with the same and fixed parameters.
[0038] As another optional example, the pre-trained neural network is designed based on the marker point method. The core idea of the marker point method includes: the constructed neural network part contains only the LMP energy function of a GPAM-PPF subnetwork, and is trained using a marker point dataset (data including steady-state and unfolded states) to complete the statistical analysis of protein structural feature preferences and constraints.
[0039] Specifically, the pre-trained neural network includes an improved protein probability function GPAM-PPF subnetwork.
[0040] During the training phase, the GPAM-PPF subnetwork needs to traverse protein structure data in two states (i.e., steady state and unfolded state). Therefore, in each round of training, all steady-state protein structure data can be traversed first, and then all unfolded protein structure data can be traversed.
[0041] Based on the above embodiments and implementation methods, as an optional implementation method, refer to... Figure 3The GPAM-PPF subnetwork includes a cascaded three-dimensional feature extraction module, a one-dimensional feature extraction module, four improved GPAM modules, and a single-layer convolutional layer.
[0042] The 3D feature extraction module includes 9 3D convolutional sub-modules, wherein each 3D convolutional sub-module includes a 3D data convolution Conv3D layer, a 3D data normalization BatchNormal3D layer, a random dropout layer, and a LeakyReLU activation function layer.
[0043] The input data for the three-dimensional feature extraction module is a protein main chain structure array, which can be 28-channel 20*20*20 dimensional data. After the convolutional dimensionality reduction operation of the convolutional layer, the final output is 256-channel 1*1*1 dimensional data, which is then processed and transformed into a 256-dimensional feature vector.
[0044] The one-dimensional feature extraction module includes two one-dimensional convolutional sub-modules, wherein each one-dimensional convolutional sub-module includes a one-dimensional data convolution Conv1 D layer, a one-dimensional data normalization BatchNormal1 D layer, a random dropout layer, and an activation function LeakyReLU layer.
[0045] The one-dimensional feature extraction module concatenates the feature vector output from the three-dimensional feature extraction module with the amino acid type data to form input data. There are 20 natural amino acid types, which are encoded to form a 20-dimensional feature vector. This 20-dimensional feature vector is then concatenated with the feature vector output from the three-dimensional feature extraction module to form 276-dimensional input data. The one-dimensional feature extraction module then inputs the concatenated 276-dimensional data, performs convolutional dimensionality reduction operations in the convolutional layer, and finally outputs 256-dimensional data.
[0046] The input data for the four improved GPAM modules is composed of the feature vector of the preceding protein structure data output by the preceding module and the new protein structure data.
[0047] Therefore, the four improved GPAM modules, arranged from closest to furthest from the main chain, address the dihedral angles of the side chains (e.g., Figure 3 The data of chi1 dihedral angle, chi2 dihedral angle, chi3 dihedral angle, and chi4 dihedral angle are incorporated into the feature vector, and finally the joint feature vector of protein main chain and side chain structure data is output.
[0048] The one-dimensional feature extraction module outputs 256-dimensional data, which is then concatenated with a sidechain dihedral data to form 280-dimensional data, which serves as the input data for the improved GPAM module. Finally, after the convolutional dimensionality reduction operation of the single-layer convolutional layer, the final output is 256-dimensional data.
[0049] Based on the above embodiments and implementation methods, as an optional implementation method, when the pre-trained neural network includes three improved protein probability function GPAM-PPF sub-networks, the training steps of the pre-trained neural network include:
[0050] First, it should be noted that steps 1) to 6) are used to train the training network consisting of the "teacher subnetwork and the steady-state student subnetwork to be trained"; steps 1) and 7) to 11) are used to train the training network consisting of the "teacher subnetwork and the non-folded student subnetwork to be trained". The two training networks can be trained simultaneously.
[0051] Step 1) Create a training dataset;
[0052] The training dataset includes a first training data subset and a second training data subset with the same amount of data; the first training data subset includes steady-state protein structure sample data, and the second training data subset includes unfolded protein structure sample data.
[0053] Optionally, the first and second training data subsets can be established based on the CATH dataset. The steady-state protein structure data in the first training data subset may contain 39,428 PDB files, from which 8,145,448 data points can be extracted; the amount of unfolded protein structure data in the first training data subset is the same as the amount of steady-state protein structure data.
[0054] Step 2) Based on the first training data subset, the same steady-state protein structure sample data are input into the teacher sub-network and the steady-state student sub-network to be trained, respectively, to obtain the first predicted feature vector output by the teacher sub-network and the second predicted feature vector output by the steady-state student sub-network to be trained; wherein, the network parameters of the teacher sub-network remain unchanged.
[0055] By setting the network parameters of the teacher subnetwork to remain constant, the consistency of the knowledge distillation objects of the steady-state student subnetwork to be trained in each round of training can be guaranteed.
[0056] Among them, reference Figure 2 The first predicted feature vector is Figure 2 The "steady-state feature vector extracted by the teacher model" and the second predicted feature vector are... Figure 2 The feature vector extracted from the steady-state model.
[0057] Step 3) Using the first preset loss function, based on the first predicted feature vector and the second predicted feature vector, obtain the function value of the first loss function.
[0058] As an alternative example, the first preset loss function can be represented by the following formula (1):
[0059]
[0060] in, This represents the function value of the first loss function (i.e., Figure 2 Loss shown p ); h1 represents the first predicted feature vector; h1 represents the second predicted feature vector; MSE represents the mean squared distance between the predicted feature vectors.
[0061] Step 4) Pass the first loss function value back to the steady-state student subnetwork to be trained.
[0062] Step 5) The steady-state student sub-network to be trained adjusts the network parameters of each network layer according to the function value of the first loss function.
[0063] As an optional example, the Adam optimizer can be used for subsequent gradient descent and learning rate optimization processes to adaptively and gradually adjust the network parameters of the steady-state student subnetwork to be trained. The initial learning rate of the Adam optimizer is 10. -4 betas are 0.5 and 0.999 respectively, and weight_decay is 5*10. -6 The Adam optimizer employs an exponential descent mode, with the descent exponent set to 0.9, reducing the learning rate once every 65,536 training iterations.
[0064] Step 6) Iteratively execute steps 3) to 5) until the first preset training completion condition is met, and obtain the steady-state student subnetwork from the steady-state student subnetwork to be trained;
[0065] The first preset training completion condition includes the number of iterations being greater than a first preset iteration number threshold, or the function value of the first preset loss function being less than a first preset convergence threshold.
[0066] In this embodiment of the disclosure, the first preset iteration number threshold and the first preset convergence threshold are not limited, and can be determined according to actual needs.
[0067] Step 7) Based on the second training data subset, the same unfolded protein structure sample data are input into the teacher sub-network and the unfolded student sub-network to be trained, respectively, to obtain the third predicted feature vector output by the teacher sub-network and the fourth predicted feature vector output by the unfolded student sub-network to be trained; wherein, the network parameters of the teacher sub-network remain unchanged.
[0068] By setting the network parameters of the teacher subnetwork to remain constant, the consistency of the knowledge distillation objects of the non-folded student subnetwork to be trained in each round of training can be guaranteed.
[0069] Among them, reference Figure 2 The third predicted feature vector is Figure 2 The "unfolded state feature vector extracted by the teacher model" and the fourth predicted feature vector are... Figure 2 The feature vector extracted from the non-folded student model.
[0070] Step 8) Using the second preset loss function, based on the third predicted feature vector and the fourth predicted feature vector, obtain the function value of the second loss function.
[0071] As an alternative example, the second preset loss function can be expressed using the following formula (2):
[0072]
[0073] in, This represents the function value of the second loss function (i.e., Figure 2 Loss shown n ); h1 represents the third predicted feature vector; h2 represents the fourth predicted feature vector; MSE represents the mean squared distance between the predicted feature vectors.
[0074] Step 9) Pass the second loss function value back to the non-folded student sub-network to be trained.
[0075] Step 10) The non-folded student sub-network to be trained adjusts the network parameters of each network layer according to the function value of the second loss function.
[0076] As an optional example, the Adam optimizer can be used for subsequent gradient descent and learning rate optimization processes to adaptively and progressively adjust the network parameters of the non-folded student subnetwork to be trained. The initial learning rate of the Adam optimizer is 10. -4 betas are 0.5 and 0.999 respectively, and weight_decay is 5*10. -6 The Adam optimizer employs an exponential descent mode, with the descent exponent set to 0.9, reducing the learning rate once every 65,536 training iterations.
[0077] Step 11) Iteratively execute steps 8) to 10) until the second preset training completion condition is met, and obtain the non-folded student sub-network from the non-folded student sub-network to be trained;
[0078] The second preset training completion condition includes either the number of iterations being greater than a second preset iteration threshold, or the function value of the second preset loss function being less than a second preset convergence threshold.
[0079] In this embodiment of the disclosure, the second preset iteration number threshold and the second preset convergence threshold are not limited, and can be determined according to actual needs.
[0080] Step 12) After synchronously iterating through steps 3) to 5) and steps 8) to 10), and obtaining the steady-state student subnetwork and the non-folded student subnetwork, the pre-trained neural network is composed of the teacher subnetwork, the steady-state student subnetwork, and the non-folded student subnetwork.
[0081] As described above, the training network consisting of the "teacher subnetwork and the steady-state student subnetwork to be trained" and the training network consisting of the "teacher subnetwork and the non-folded student subnetwork to be trained" are trained using different training data respectively. They calculate their respective loss values using formulas (1) and (2) respectively, and optimize the parameters of their respective student subnetworks. The training processes are independent of each other and do not interfere with each other. The two training networks are only connected through the teacher subnetwork with the same and fixed parameters.
[0082] Based on the above embodiments and implementation methods, as an optional implementation method, when the pre-trained neural network includes an improved protein probability function GPAM-PPF sub-network, the training steps of the pre-trained neural network include:
[0083] Step 1: Create a training dataset;
[0084] The training dataset includes a mixed sample data subset, a steady-state labeled data subset, and an unfolded data subset; the mixed sample data subset includes sample protein structure data that does not distinguish between structural types; the steady-state labeled data subset includes steady-state protein structure labeled data; and the unfolded data subset includes unfolded protein structure labeled data.
[0085] As an optional example, creating the training dataset may include the following steps:
[0086] Step I) Determine the target amount of data to be sampled from the sampled dataset using a preset sampling rule; wherein the sampled dataset includes a sampled steady-state protein structure dataset and a sampled unfolded protein structure dataset.
[0087] Optionally, the preset sampling rule can be expressed using the following calculation formula (3):
[0088]
[0089] Where Ds represents the marked point dataset, D represents the sampled dataset, |Ds| represents the target data volume (i.e., the data volume of the marked point dataset), |D| represents the data volume of the sampled dataset, lg represents the logarithmic function with base 10, and floor represents the floor operation.
[0090] Step II) Based on the target data volume, sample steady-state protein structure data from the sampled steady-state protein structure dataset as steady-state protein structure marker point data to obtain the steady-state marker data subset; Step III) Based on the target data volume, sample unfolded protein structure data from the sampled steady-state protein structure dataset as unfolded protein structure marker point data to obtain the unfolded data subset.
[0091] It should be noted that in protein structure datasets (e.g., the CATH dataset), sampling is typically performed on a per-PDB file basis. The augmented CATH dataset contains 39,428 PDB files representing steady-state protein structures and 39,428 PDB files representing unfolded protein structures. Therefore, the sampled subsets of steady-state and unfolded label data should each contain 16 PDB files. To ensure that the feature space of the label data subsets closely approximates the feature space of the sampled dataset, during the sampling process, steady-state and unfolded protein structures are further divided into four categories: predominantly α-helices, predominantly β-sheets, a mixture of α-helices and β-sheets, and a minority of secondary structures. A random function implemented using a Gaussian distribution is used to generate sampling numbers according to the proportions of the four categories, thus sampling the steady-state and unfolded label data subsets.
[0092] Step ②: Based on the mixed sample data subset, input the undifferentiated protein structure data of the samples into the GPAM-PPF sub-network to be trained to obtain the fifth predicted feature vector output by the GPAM-PPF sub-network to be trained; Step ③: Based on the steady-state labeled data subset, input the steady-state protein structure labeled data into the GPAM-PPF sub-network to be trained to obtain the sixth predicted feature vector output by the GPAM-PPF sub-network to be trained; Step ④: Based on the unfolded data subset, input the unfolded protein structure labeled data into the GPAM-PPF sub-network to be trained to obtain the seventh predicted feature vector output by the GPAM-PPF sub-network to be trained.
[0093] In steps ② to ④, refer to Figure 4 Training data can be simultaneously input into the GPAM-PPF subnetwork to be trained.
[0094] As an optional example, the GPAM-PPF subnetwork to be trained needs to traverse protein structure data in two states. Therefore, in each round of training, all steady-state protein structure data can be traversed first, and then all unfolded protein structure data can be traversed.
[0095] Among them, reference Figure 4 The fifth predicted feature vector is Figure 4 The "training data feature vector" and the sixth predicted feature vector are... Figure 4 The "steady-state marker data feature vector" and the seventh predicted feature vector are... Figure 4 The "feature vector of non-folded state marker point data" in the middle.
[0096] Step 5: Using the third preset loss function, based on the fifth, sixth, and seventh predicted feature vectors, obtain the function value of the third loss function.
[0097] As an optional example, the third preset loss function can be expressed using the following formulas (4) to (7):
[0098] L1 = cdist(h p h s (4)
[0099] L2 = cdist(h p h u (5)
[0100] L3 = cdist(h s h u (6)
[0101]
[0102] Among them, h p h represents the fifth predicted feature vector; s h represents the sixth predicted feature vector; u L1 represents the seventh predicted feature vector; L2 represents the Euclidean distance between the fifth and sixth predicted feature vectors; L3 represents the Euclidean distance between the sixth and seventh predicted feature vectors; cdist represents the function for calculating the Euclidean distance between feature vectors; L(h p h s h u) represents the third preset loss function; the condition "train data is stable" of the third preset loss function corresponds to the case where the training data is steady-state protein structure data, and the condition "train data is unfolded" corresponds to the case where the training data is unfolded protein structure data.
[0103] Step 6: Pass the third loss function value back to the GPAM-PPF subnetwork to be trained.
[0104] Step 7: The GPAM-PPF sub-network to be trained adjusts the network parameters of each network layer according to the function value of the third loss function.
[0105] As an optional example, the Adam optimizer can be used for subsequent gradient descent and learning rate optimization processes to adaptively and gradually adjust the network parameters of the GPAM-PPF sub-network being trained. The initial learning rate of the Adam optimizer is 10. -4 betas are 0.5 and 0.999 respectively, and weight_decay is 5*10. -6 The Adam optimizer employs an exponential descent mode, with the descent exponent set to 0.9, reducing the learning rate once every 65,536 training iterations.
[0106] Step 8: Iteratively execute steps 5 to 7 until the third preset training completion condition is met, and obtain the GPAM-PPF subnetwork from the GPAM-PPF subnetwork to be trained.
[0107] The third preset training completion condition includes either the number of iterations being greater than a third preset iteration number threshold, or the function value of the third preset loss function being less than a third preset convergence threshold.
[0108] In this embodiment of the disclosure, the third preset iteration number threshold and the third preset convergence threshold are not limited, and can be determined according to actual needs.
[0109] It should be noted that, since different proteins have different peptide chain lengths, in order to avoid the influence of peptide chain length on protein structural energy, the protein structural energy is set to be the average of the energies of amino acids at all positions. That is, the cdist function reduces dimensionality by calculating the average value and combines the results into a single value.
[0110] As described above, when training the GPAM-PPF subnetwork using steps ① to ⑧, the Euclidean distance between each pair of the "fifth predicted feature vector, sixth predicted feature vector, and seventh predicted feature vector" can be calculated using the above formula. Then, the calculation method of the loss value is adjusted according to the training data type, so that the "steady-state data" in the sample protein structure data that does not distinguish the structure type is close to the steady-state labeled data subset, and the "unfolded state data" in the sample protein structure data that does not distinguish the structure type is close to the unfolded state data subset.
[0111] Based on the above embodiments and implementation methods, as an optional implementation method, the creation of the training dataset includes: using preset data augmentation rules to perform data augmentation on the unfolded protein structure sample data and / or the unfolded protein structure marker data to obtain an expanded training dataset.
[0112] First and foremost, it needs to be explained that the unique physical and chemical properties of unfolded protein structures (e.g., high dynamism and structural diversity) make it difficult to achieve satisfactory training results by directly applying conventional data processing methods.
[0113] In view of this, in order to effectively process unfolded protein structure data and improve the training efficiency and prediction generalization ability of pre-trained neural networks, we adopted a data augmentation method (e.g., nearest neighbor counting). This method expands the unfolded protein structure dataset to enhance the pre-trained neural network's ability to identify and process such structures, thereby improving the overall performance of the pre-trained neural network.
[0114] The core idea of the nearest neighbor counting method is to enhance data based on the spatial proximity of each amino acid residue. Specifically, for each amino acid residue in the unfolded protein structure, its distance to other amino acid residues in three-dimensional space is calculated, and which residues are its nearest neighbors are determined according to a predetermined distance threshold. This process generates a proximity feature matrix, where each element represents the number of nearest neighbors for a specific amino acid residue.
[0115] Based on this, the original "sheetless protein structure sample data and / or sheetless protein structure marker data" are enhanced using this proximity feature matrix.
[0116] Specifically, as an alternative example, proximity features can be incorporated as additional input features into the model's training process. This not only increases the pre-trained neural network's ability to process unfolded protein structure data but also provides it with additional spatial structural information, helping it capture the complexity of dynamic protein structural changes.
[0117] Based on the above embodiments and implementation methods, as an optional implementation, in step S120, the pre-trained neural network can output a joint feature vector that matches the input target protein structure data. The joint feature vector represents the statistical representation of the pre-trained neural network's preference and constraints on the target protein structure. The output mapping unit can convert the joint feature vector into the protein structure energy based on a preset mapping rule.
[0118] As an optional example, in the case where the pre-trained neural network includes three improved protein probability function GPAM-PPF subnetworks, the "preset mapping rule" of the output mapping unit can be represented by the following calculation formula (8):
[0119] E(P, θ) t θ sp θ sn ) = MSE(θ t (P), θ sp (P))-MSE(θ t (P), θ sn (P)) (8)
[0120] Where, E(P, θ) t θ sp θ sn ) represents the energy of the protein structure; P represents the input target protein structure data; θ t Represents the network parameters of the teacher subnetwork; θ sp θ represents the network parameters of the steady-state student subnetwork. sn This represents the network parameters of the non-folded student subnetwork.
[0121] As another alternative example, in the case where the pre-trained neural network includes an improved protein probability function GPAM-PPF subnetwork, the "preset mapping rule" of the output mapping unit can be represented by the following formula (9):
[0122]
[0123] Wherein, E(P, P) s P u θ) represents the energy of the protein structure; P represents the input target protein structure data; θ represents the network parameters of the GPAM-PPF subnetwork; P s P represents the steady-state marker point; u This indicates a non-folded marker point.
[0124] It should be noted that, optionally, in the GPAM-PPF subnetwork trained based on the "marker point method", the representation "P" can be stored. s P represents the steady-state marker point; u The labeling parameter represents the "unfolded state label point". Therefore, during the application phase, the labeling parameter and the joint feature vector are sent to the output mapping unit, which can then use the above calculation formula (9) to calculate the energy of the protein structure.
[0125] In addition, during the joint training phase of the GPAM-PPF subnetwork and the output mapping unit, the output mapping unit can use the steady-state labeled data subset, which includes steady-state protein structure labeled data, and the unfolded data subset, which includes unfolded protein structure labeled data, combined with the joint feature vector output by the GPAM-PPF subnetwork, to calculate the protein structure energy.
[0126] As described above, the pre-trained energy prediction model in this embodiment can not only efficiently distinguish between steady-state and unfolded protein structures, but also assign lower energy values to protein structures with high accuracy, thus demonstrating good performance in protein design accuracy. Therefore, using the method for predicting protein structure energy provided in this embodiment, the protein structure energy of target protein structure data can be predicted relatively quickly and accurately, yielding effective prediction results.
[0127] The apparatus for predicting protein structure energy provided in this disclosure is described below. The apparatus for predicting protein structure energy described below can be referred to in conjunction with the method for predicting protein structure energy described above.
[0128] Figure 5 This is a schematic diagram of a device for predicting protein structure energy provided in an exemplary embodiment of this disclosure. Figure 5 As shown, a device for predicting the energy of protein structure includes:
[0129] The data acquisition module 510 is configured to: acquire target protein structure data to be processed; wherein the target protein structure data includes main chain structure data and side chain structure data, and the protein structure types involved in the target protein structure data include steady state and unfolded state; the energy prediction module 520 is configured to: input the target protein structure data into a pre-trained energy prediction model to obtain the protein structure energy matching the target protein structure data output by the pre-trained energy prediction model; wherein the pre-trained energy prediction model includes a cascaded pre-trained neural network and an output mapping unit, and the pre-trained neural network is trained using sample data constructed based on the protein structure data.
[0130] Figure 6 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 6 As shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630, and a communication bus 640, wherein the processor 610, the communication interface 620, and the memory 630 communicate with each other via the communication bus 640. The processor 610 can call logical instructions in the memory 630 to execute a method for predicting protein structure energy. This method includes: acquiring target protein structure data to be processed; wherein the target protein structure data includes main chain structure data and side chain structure data, and the protein structure types involved in the target protein structure data include steady-state and unfolded states; inputting the target protein structure data into a pre-trained energy prediction model to obtain a protein structure energy matching the target protein structure data output by the pre-trained energy prediction model; wherein the pre-trained energy prediction model includes a cascaded pre-trained neural network and an output mapping unit, and the pre-trained neural network is trained using sample data constructed based on the protein structure data.
[0131] Furthermore, the logical instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0132] On the other hand, this disclosure also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the methods provided above for predicting protein structure energy. The method includes: acquiring target protein structure data to be processed; wherein the target protein structure data includes main chain structure data and side chain structure data, and the protein structure types involved in the target protein structure data include steady-state and unfolded states; inputting the target protein structure data into a pre-trained energy prediction model to obtain a protein structure energy matching the target protein structure data output by the pre-trained energy prediction model; wherein the pre-trained energy prediction model includes a cascaded pre-trained neural network and an output mapping unit, and the pre-trained neural network is trained using sample data constructed based on protein structure data.
[0133] In another aspect, this disclosure also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, is implemented to perform the methods for predicting protein structure energy provided by the methods described above. The method includes: acquiring target protein structure data to be processed; wherein the target protein structure data includes main chain structure data and side chain structure data, and the protein structure types involved in the target protein structure data include steady-state and unfolded states; inputting the target protein structure data into a pre-trained energy prediction model to obtain a protein structure energy matching the target protein structure data output by the pre-trained energy prediction model; wherein the pre-trained energy prediction model includes a cascaded pre-trained neural network and an output mapping unit, and the pre-trained neural network is trained using sample data constructed based on the protein structure data.
[0134] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0135] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0136] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this disclosure, and are not intended to limit them. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure.
Claims
1. A method for predicting the structural energy of a protein, comprising: Obtain the target protein structure data to be processed; wherein, the target protein structure data includes main chain structure data and side chain structure data, and the protein structure types involved in the target protein structure data include steady state and unfolded state; The target protein structure data is input into a pre-trained energy prediction model to obtain the protein structure energy that matches the target protein structure data, output by the pre-trained energy prediction model. The pre-trained energy prediction model includes a cascaded pre-trained neural network and an output mapping unit. The pre-trained neural network is trained using sample data constructed based on protein structure data. The input terminal of the pre-trained neural network serves as the input terminal of the pre-trained energy prediction model. The input terminal of the pre-trained neural network is connected to the input terminal of the output mapping unit, and the output terminal of the output mapping unit serves as the output terminal of the pre-trained neural network. The pre-trained neural network includes three improved protein probability function GPAM-PPF subnetworks, wherein one GPAM-PPF subnetwork is used as a teacher subnetwork, one GPAM-PPF subnetwork is used as a steady-state student subnetwork, and one GPAM-PPF subnetwork is used as an unfolded student subnetwork. Alternatively, the pre-trained neural network may include an improved protein probability function GPAM-PPF subnetwork; In the case where the pre-trained neural network includes an improved protein probability function GPAM-PPF subnetwork, the training steps of the pre-trained neural network include: Step 1: Create a training dataset; The training dataset includes a mixed sample data subset, a steady-state labeled data subset, and an unfolded data subset; the mixed sample data subset includes protein structure data of samples without distinguishing structure type; the steady-state labeled data subset includes steady-state protein structure labeled data; and the unfolded data subset includes unfolded protein structure labeled data. Step ②: Based on the mixed sample data subset, input the undifferentiated sample protein structure data into the GPAM-PPF sub-network to be trained to obtain the fifth predicted feature vector output by the GPAM-PPF sub-network to be trained; Step 3: Based on the steady-state labeled data subset, input the steady-state protein structure labeled data into the GPAM-PPF sub-network to be trained to obtain the sixth predicted feature vector output by the GPAM-PPF sub-network to be trained; Step 4: Based on the unfolded state data subset, input the unfolded state protein structure marker point data into the GPAM-PPF sub-network to be trained to obtain the seventh predicted feature vector output by the GPAM-PPF sub-network to be trained; Step 5: Using the third preset loss function, based on the fifth, sixth, and seventh predicted feature vectors, obtain the function value of the third loss function; Step 6: Pass the third loss function value back to the GPAM-PPF subnetwork to be trained; Step 7: The GPAM-PPF sub-network to be trained adjusts the network parameters of each network layer according to the function value of the third loss function; Step ⑧: Iteratively execute steps ⑤ to ⑦ until the third preset training completion condition is met, and obtain the GPAM-PPF subnetwork from the GPAM-PPF subnetwork to be trained; The third preset training completion condition includes either the number of iterations being greater than a third preset iteration number threshold, or the function value of the third preset loss function being less than a third preset convergence threshold.
2. The method according to claim 1, characterized in that, The GPAM-PPF subnetwork includes a cascaded three-dimensional feature extraction module, a one-dimensional feature extraction module, four improved GPAM modules, and a single-layer convolutional layer. The three-dimensional feature extraction module includes nine three-dimensional convolutional sub-modules, wherein each three-dimensional convolutional sub-module includes a three-dimensional data convolution Conv3D layer, a three-dimensional data normalization BatchNormal3D layer, a random dropout layer, and an activation function LeakyReLU layer. The one-dimensional feature extraction module includes two one-dimensional convolutional sub-modules, wherein each one-dimensional convolutional sub-module includes a one-dimensional data convolution Conv1D layer, a one-dimensional data normalization BatchNormal1D layer, a random dropout layer, and an activation function LeakyReLU layer.
3. The method according to claim 1, characterized in that, In the case where the pre-trained neural network includes three improved protein probability function GPAM-PPF subnetworks, the training steps of the pre-trained neural network include: Step 1) Create a training dataset; The training dataset includes a first training data subset and a second training data subset with the same amount of data; the first training data subset includes steady-state protein structure sample data, and the second training data subset includes unfolded protein structure sample data. Step 2) Based on the first training data subset, input the same steady-state protein structure sample data into the teacher sub-network and the steady-state student sub-network to be trained, respectively, to obtain the first predicted feature vector output by the teacher sub-network and the second predicted feature vector output by the steady-state student sub-network to be trained; wherein, the network parameters of the teacher sub-network remain unchanged. Step 3) Using the first preset loss function, based on the first predicted feature vector and the second predicted feature vector, obtain the function value of the first loss function; Step 4) Pass the first loss function value back to the steady-state student subnetwork to be trained; Step 5) The steady-state student sub-network to be trained adjusts the network parameters of each network layer according to the function value of the first loss function; Step 6) Iteratively execute steps 3) to 5) until the first preset training completion condition is met, and obtain the steady-state student subnetwork from the steady-state student subnetwork to be trained; The first preset training completion condition includes the number of iterations being greater than the first preset iteration number threshold, or the function value of the first preset loss function being less than the first preset convergence threshold. Step 7) Based on the second training data subset, the same unfolded protein structure sample data are input into the teacher subnetwork and the unfolded student subnetwork to be trained, respectively, to obtain the third predicted feature vector output by the teacher subnetwork and the fourth predicted feature vector output by the unfolded student subnetwork to be trained; wherein, the network parameters of the teacher subnetwork remain unchanged. Step 8) Using the second preset loss function, based on the third predicted feature vector and the fourth predicted feature vector, obtain the function value of the second loss function; Step 9) Pass the second loss function value back to the unfolded student subnetwork to be trained; Step 10) The unfolded student sub-network to be trained adjusts the network parameters of each network layer according to the function value of the second loss function; Step 11) Iteratively execute steps 8) to 10) until the second preset training completion condition is met, and obtain the non-folded student sub-network from the non-folded student sub-network to be trained; The second preset training completion condition includes the number of iterations being greater than the second preset iteration number threshold, or the function value of the second preset loss function being less than the second preset convergence threshold. Step 12) After synchronously iterating through steps 3) to 5) and steps 8) to 10), and obtaining the steady-state student subnetwork and the non-folded student subnetwork, the pre-trained neural network is composed of the teacher subnetwork, the steady-state student subnetwork, and the non-folded student subnetwork.
4. The method according to claim 1 or 3, characterized in that, The creation of the training dataset includes: Using preset data augmentation rules, data augmentation is performed on the sample data of the unfolded protein structure and / or the marker data of the unfolded protein structure to obtain an expanded training dataset.
5. The method according to claim 1, characterized in that, The creation of the training dataset includes: Using preset sampling rules, the target amount of data to be sampled from the sampled dataset is determined; wherein, the sampled dataset includes a sampled steady-state protein structure dataset and a sampled unfolded protein structure dataset. Based on the target data volume, steady-state protein structure data is sampled from the sampled steady-state protein structure dataset as steady-state protein structure marker point data to obtain the steady-state marker data subset; Based on the target data volume, unfolded protein structure data is sampled from the sampled steady-state protein structure dataset as unfolded protein structure marker point data to obtain the unfolded data subset.
6. A device for predicting the energy of protein structure, comprising: The data acquisition module is configured to acquire target protein structure data to be processed; wherein the target protein structure data includes main chain structure data and side chain structure data, and the protein structure types involved in the target protein structure data include steady state and unfolded state; The energy prediction module is configured to: input the target protein structure data into a pre-trained energy prediction model to obtain the protein structure energy that matches the target protein structure data, output by the pre-trained energy prediction model. The pre-trained energy prediction model includes a cascaded pre-trained neural network and an output mapping unit. The pre-trained neural network is trained using sample data constructed based on protein structure data. The input terminal of the pre-trained neural network serves as the input terminal of the pre-trained energy prediction model. The input terminal of the pre-trained neural network is connected to the input terminal of the output mapping unit, and the output terminal of the output mapping unit serves as the output terminal of the pre-trained neural network. The pre-trained neural network includes three improved protein probability function GPAM-PPF subnetworks, wherein one GPAM-PPF subnetwork is used as a teacher subnetwork, one GPAM-PPF subnetwork is used as a steady-state student subnetwork, and one GPAM-PPF subnetwork is used as an unfolded student subnetwork. Alternatively, the pre-trained neural network may include an improved protein probability function GPAM-PPF subnetwork; In the case where the pre-trained neural network includes an improved protein probability function GPAM-PPF subnetwork, the training steps of the pre-trained neural network include: Step 1: Create a training dataset; The training dataset includes a mixed sample data subset, a steady-state labeled data subset, and an unfolded data subset; the mixed sample data subset includes protein structure data of samples without distinguishing structure type; the steady-state labeled data subset includes steady-state protein structure labeled data; and the unfolded data subset includes unfolded protein structure labeled data. Step ②: Based on the mixed sample data subset, input the undifferentiated sample protein structure data into the GPAM-PPF sub-network to be trained to obtain the fifth predicted feature vector output by the GPAM-PPF sub-network to be trained; Step 3: Based on the steady-state labeled data subset, input the steady-state protein structure labeled data into the GPAM-PPF sub-network to be trained to obtain the sixth predicted feature vector output by the GPAM-PPF sub-network to be trained; Step 4: Based on the unfolded state data subset, input the unfolded state protein structure marker point data into the GPAM-PPF sub-network to be trained to obtain the seventh predicted feature vector output by the GPAM-PPF sub-network to be trained; Step 5: Using the third preset loss function, based on the fifth, sixth, and seventh predicted feature vectors, obtain the function value of the third loss function; Step 6: Pass the third loss function value back to the GPAM-PPF subnetwork to be trained; Step 7: The GPAM-PPF sub-network to be trained adjusts the network parameters of each network layer according to the function value of the third loss function; Step ⑧: Iteratively execute steps ⑤ to ⑦ until the third preset training completion condition is met, and obtain the GPAM-PPF subnetwork from the GPAM-PPF subnetwork to be trained; The third preset training completion condition includes either the number of iterations being greater than a third preset iteration number threshold, or the function value of the third preset loss function being less than a third preset convergence threshold.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method for predicting protein structure energy as described in any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method for predicting protein structure energy as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Student model training method based on pre-training language model and text classification system
CN115526332A
Rapid prediction and correlation mechanism analysis method for molecular targets
CN115631808A