Protein Conformation Prediction Method, Device, Electronic Device and Storage Medium
By pre-sorting and fusion sorting rules for the candidate protein conformation of the target protein sequence, the problem of inaccurate prediction of protein conformation is solved and the accuracy of prediction is improved.
Patent Information
- Application Number
- CN202011438019.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-07
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2040-12-07
AI Technical Summary
In the prior art, there is a problem of inaccurate prediction of protein conformation prediction methods corresponding to protein sequences.
By obtaining the M candidate protein conformations corresponding to the target protein sequence, pre-sorting is performed based on the N protein sorting rules, N pre-sorting results are obtained, and then processing these results based on the fusion sorting rules to obtain the target sorting results, and finally determining the predicted protein conformation.
Through the processing of two-level sorting and fusion sorting rules, the accuracy of protein conformation prediction is improved, and the problem of poor sorting stability in traditional tools is avoided.
Smart Images

Figure CN114613429B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and more specifically, to a method, apparatus, electronic device, and storage medium for predicting protein conformation. Background Art
[0002] Proteins are the basic organic substances that make up human cells and play a very important role in the human body. Understanding the spatial structure of proteins (i.e., protein conformation) is of great significance. However, in related technologies, the method for predicting the protein conformation corresponding to a protein sequence has the problem of inaccurate prediction. Summary of the Invention
[0003] In view of the above problems, embodiments of this application propose a method, apparatus, electronic device, and storage medium for predicting protein conformation to improve the above problems.
[0004] In a first aspect, an embodiment of this application provides a method for predicting protein conformation. The method includes: obtaining M candidate protein conformations corresponding to a target protein sequence, where M is a positive integer greater than 1; pre-ranking the M candidate protein conformations based on N protein ranking rules to obtain N pre-ranking results corresponding to the M candidate protein conformations, where N is a positive integer greater than 0; processing the N pre-ranking results based on a fusion ranking rule to obtain a target ranking result corresponding to the M candidate protein conformations; and determining a predicted protein conformation corresponding to the target protein sequence according to the target ranking result.
[0005] In a second aspect, an embodiment of this application provides a device for predicting protein conformation. The device includes: a candidate protein conformation obtaining module, a pre-ranking module, a target ranking result obtaining module, and a predicted protein conformation determining module. The candidate protein conformation obtaining module is configured to obtain M candidate protein conformations corresponding to a target protein sequence, where M is a positive integer greater than 1. The pre-ranking module is configured to pre-rank the M candidate protein conformations based on N protein ranking rules to obtain N pre-ranking results corresponding to the M candidate protein conformations, where N is a positive integer greater than 0. The target ranking result obtaining module is configured to process the N pre-ranking results based on a fusion ranking rule to obtain a target ranking result corresponding to the M candidate protein conformations. The predicted protein conformation determining module is configured to determine a predicted protein conformation corresponding to the target protein sequence according to the target ranking result.
[0006] In a third aspect, an embodiment of this application provides an electronic device, including a processor and a memory; one or more programs are stored in the memory and configured to be executed by the processor to implement the above method.
[0007] Fourthly, an embodiment of the present application provides a computer-readable storage medium, in which program code is stored. When the program code is run by a processor, the above-mentioned method is executed.
[0008] Fifthly, an embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned method.
[0009] A protein conformation prediction method, device, electronic device and storage medium provided by an embodiment of the present application obtain M candidate protein conformations corresponding to a target protein sequence, then pre-rank the M candidate protein conformations based on N protein ranking rules to obtain N pre-ranking results corresponding to the M candidate protein conformations, and then process the N pre-ranking results based on a fusion ranking rule to obtain a target ranking result corresponding to the M candidate protein conformations. Finally, according to the target ranking result, the predicted protein conformation corresponding to the target protein sequence is determined. Thus, through the foregoing manner, on the basis of pre-ranking the M candidate proteins corresponding to the target protein sequence according to N protein ranking rules, the N pre-ranking results are processed based on the fusion ranking rule. Since the fusion ranking rule is processed based on the N ranking results obtained by pre-ranking, it is equivalent to performing two-level ranking on the protein conformation, integrating different ranking methods of various protein ranking rules, avoiding the problem of poor sorting stability of traditional tools, enabling better prediction accuracy on various protein structures, and thus improving the overall accuracy of the predicted protein conformation. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0011] Figure 1 FIG. shows a schematic diagram of an application environment proposed by an embodiment of the present application;
[0012] Figure 2 FIG. shows a schematic diagram of another application environment proposed by an embodiment of the present application;
[0013] Figure 3 FIG. shows a flowchart of a protein conformation prediction method proposed by an embodiment of the present application;
[0014] Figure 4 The flowchart of another protein conformation prediction method proposed by the embodiments of the present application is shown;
[0015] Figure 5 It shows Figure 4 The flowchart of an implementation manner of S210 in a protein conformation prediction method proposed by the shown embodiment;
[0016] Figure 6 It shows Figure 4 The flowchart of an implementation manner of S230 in a protein conformation prediction method proposed by the shown embodiment;
[0017] Figure 7 The schematic diagram of a training process integrating sorting rules proposed by the embodiments of the present application is shown;
[0018] Figure 8 The flowchart of another protein conformation prediction method proposed by the embodiments of the present application is shown;
[0019] Figure 9 The schematic diagram of the specific application process of a protein conformation prediction method proposed by the embodiments of the present application is shown;
[0020] Figure 10 The block diagram of a protein conformation prediction device proposed by the embodiments of the present application is shown;
[0021] Figure 11 The structural block diagram of another electronic device for executing the protein conformation prediction method according to the embodiments of the present application is shown;
[0022] Figure 12 The storage unit for storing or carrying the program code for implementing the protein conformation prediction method according to the embodiments of the present application is shown. Detailed implementation manners
[0023] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0024] Artificial Intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including the theory, methods, technologies, and application systems for perceiving the environment, acquiring knowledge, and using knowledge to achieve the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.
[0025] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields involved, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0026] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0027] Among them, with the development of machine learning technology, machine learning technology has been widely studied and applied in multiple fields. The technical solution provided by the embodiments of this application involves the application of machine learning technology in the biomedical field. Specifically, it involves a method for predicting protein conformation. Among them, protein conformation refers to the spatial structure of proteins.
[0028] Protein is the material basis of life, an organic macromolecule, the basic organic matter that constitutes cells, the main bearer of life activities, and the target of pharmaceuticals. When the human body gets sick, essentially, there are problems in protein production. Currently, small molecule drugs and large molecule drugs in pharmaceuticals interact with the proteins in the human body and ultimately maintain the original signal pathways of human cells. Therefore, understanding protein conformation is of great significance.
[0029] Protein prediction algorithms, which predict the three-dimensional structure of proteins based on protein structure features, are one of the important means of analyzing protein structures. However, existing protein prediction algorithms generate a large number of protein conformations. For example, the Rosetta tool, which is a comprehensive software for simulating the structures of macromolecules and uses the simulated annealing algorithm to predict the three-dimensional structure of proteins, generally generates 50,000 protein candidate conformations. Or, using the AlphaFold system also requires generating a large number of protein candidate conformations. However, only one of the large number of protein conformations is the true three-dimensional structure of the protein that is closest to the protein sequence. Therefore, protein conformation selection is becoming increasingly important.
[0030] In some ways, the most correct protein conformation can be predicted by sorting a large number of protein conformations. However, in related technologies, various methods for sorting protein conformations, for example, methods that use the distances between the backbone and side-chain atoms of proteins, or the planar angles between the backbone and side-chain atoms, or the dihedral angles between the backbone and side-chain atoms, etc. as comparison data for sorting. These protein sorting methods have relatively large performance differences on different protein structures, and have poor stability in protein conformation sorting, which in turn leads to low overall accuracy of the predicted protein conformations.
[0031] Therefore, the inventors proposed the protein conformation sorting method, device, electronic device, and storage medium provided in this application. In this application, after obtaining M candidate protein conformations corresponding to the target protein sequence, first, based on N protein sorting rules, pre-sort the M candidate protein conformations to obtain N pre-sorting results corresponding to the M candidate protein conformations. Then, based on the fusion sorting rule, process the N pre-sorting results to obtain the target sorting result corresponding to the M candidate protein conformations. Finally, according to the target sorting result, determine the predicted protein conformation corresponding to the target protein sequence.
[0032] Thus, through the foregoing method, on the basis of pre-sorting the M candidate proteins corresponding to the target protein sequence according to N protein sorting rules, and then processing the N pre-sorting results based on the fusion sorting rule. Since the fusion sorting rule processes based on the N sorting results obtained from the pre-sorting, it is equivalent to performing two-level sorting on the protein conformations, integrating different sorting methods of multiple protein sorting rules, avoiding the problem of poor sorting stability of traditional tools, enabling better prediction accuracy on various protein structures, and thus improving the overall accuracy of the predicted protein conformations.
[0033] Before further elaborating on the embodiments of this application, an application environment involved in the embodiments of this application is introduced.
[0034] As Figure 1As shown Figure 1 As shown is a schematic diagram of the application environment involved in the embodiments of the present application. Among them, there are a client 110, a server 120, a protein conformation prediction module 130, a protein conformation pre-sorting module 140, and a fusion sorting module 150. Among them, the client 110 is used to provide the protein sequence for which protein conformation prediction is required. The client 110 can be installed on the local terminal of a pharmaceutical company or a biological research institution. The client 110 sends the protein sequence for which protein conformation prediction is required as the target protein sequence to the server 120. After receiving the target protein sequence, the server 120 further sends the target protein sequence to the protein conformation prediction module 130. The protein conformation prediction module 130 predicts multiple candidate protein conformations based on the target protein sequence, and then sends the multiple candidate protein conformations to the protein conformation pre-sorting module 140, so that the protein conformation pre-sorting module 140 pre-sorts the multiple candidate protein conformations based on multiple protein sorting rules to obtain multiple pre-sorting results. The protein conformation pre-sorting module 140 then sends the multiple pre-sorting results to the fusion sorting module 150, so that the fusion sorting module 150 processes them according to the fusion sorting rules to obtain the target sorting result, and finally sends the target sorting result to the server, so that the server determines the predicted protein conformation corresponding to the target protein sequence according to the target sorting result and returns the predicted protein conformation to the client.
[0035] It should be noted that Figure 1 is an exemplary application environment, and the method provided by the embodiments of the present application can also run in other application environments.
[0036] Optionally, in addition to being able to run independently of the server 120 on different hardware devices as shown Figure 1 the protein conformation prediction module 130, the protein conformation pre-sorting module 140, and the fusion sorting module 150 can also run in the server 120 as shown Figure 2 In Figure 2In the environment shown, a server module responsible for communicating with the client 110 can run in the server 120. After receiving the target protein sequence, the server module can transfer the target protein sequence to the locally running protein conformation prediction module 130 based on inter-process communication. Correspondingly, after predicting multiple candidate protein conformations based on the target protein sequence, the protein conformation prediction module 130 can also send the multiple candidate protein conformations to the protein conformation pre-sorting module 140 based on inter-process communication, so that the protein conformation pre-sorting module 140 can pre-sort the multiple candidate protein conformations based on multiple protein sorting rules to obtain multiple pre-sorting results. The protein conformation pre-sorting module 140 can also send the multiple pre-sorting results to the fusion sorting module 150 based on inter-process communication, so that the fusion sorting module 150 can process them according to the fusion sorting rules to obtain the target sorting result. The fusion sorting module 150 can also send the target sorting result to the server module based on inter-process communication, so that the server module can determine the predicted protein conformation corresponding to the target protein sequence according to the target sorting result and return the predicted protein conformation to the client 110.
[0037] Optionally, the functions performed by the protein conformation prediction module 130, the protein conformation pre-sorting module 140, and the fusion sorting module 150 can also be all performed by the client 110.
[0038] It should be noted that the server 120 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The electronic device where the client 110 is located can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto.
[0039] Next, each embodiment of the present application will be specifically described with reference to the accompanying drawings.
[0040] Please refer to Figure 3 , Figure 3 The flowchart of a protein conformation prediction method proposed in an embodiment of the present application is shown. In this embodiment, the method can be applied to a server, and the method includes:
[0041] S110, obtaining M candidate protein conformations corresponding to the target protein sequence, where M is a positive integer greater than 1.
[0042] It is understandable that the target protein sequence is the protein sequence for subsequent protein conformation prediction. In one application scenario, the server can receive a protein conformation prediction request uploaded by the client. The request may carry a protein sequence for predicting the protein conformation, and this protein sequence is the target protein sequence. After receiving the target protein sequence, the server can first obtain M candidate protein conformations corresponding to the target protein sequence. Here, M is a positive integer greater than 1.
[0043] Among them, the candidate protein conformations can be considered as the M protein conformations obtained after predicting the target protein sequence using a protein prediction algorithm. These protein conformations are collectively referred to as candidate protein conformations. Only one of the candidate protein conformations is the closest to the true spatial structure of the protein corresponding to the target protein sequence, and the true protein conformation corresponding to the target protein sequence needs to be screened from the M candidate protein conformations.
[0044] S120, based on N protein sorting rules, pre-sort the M candidate protein conformations to obtain N pre-sorting results corresponding to the M candidate protein conformations, where N is a positive integer greater than 0.
[0045] Among them, after the server obtains the M candidate protein conformations corresponding to the target protein sequence, it can use N protein sorting rules to perform a pre-sorting on the M candidate protein conformations respectively to obtain the pre-sorting results. The number of types of protein sorting rules, that is, N, can be set as needed. For example, one, two, three or more different protein sorting rules can be selected to perform a pre-sorting on the M candidate protein conformations respectively. When one protein sorting rule is selected to perform a pre-sorting on multiple candidate protein conformations, one pre-sorting result can be obtained. When three different protein sorting rules are selected to perform a pre-sorting on multiple candidate protein conformations respectively, three pre-sorting results can be obtained. Here, N is a positive integer greater than 0.
[0046] Among them, there can be multiple categories of protein sorting rules. For example, the protein sorting rule can be a sorting rule based on the local information of the candidate protein. Specifically, the protein sorting rule can perform a local scoring on the local information of the candidate protein conformation and sort according to the score. It can also be a sorting rule based on the global statistical information of the candidate protein. Specifically, the protein sorting rule can score the global statistical information of the candidate protein conformation and sort according to the score.
[0047] It should be noted that the local information of proteins can include various types of local information, and each type of local information can correspond to a specific protein sorting rule. For example, for the local information of the distance between the main chain and side chain atoms of a protein, a protein sorting rule can be corresponding; for the planar angle between the main chain and side chain atoms of a protein, another protein sorting rule can be corresponding; for the dihedral angle between the main chain and side chain atoms of a protein, another protein sorting rule can be corresponding.
[0048] In addition, the global statistical information of proteins can also include various types of global statistical information. For each type of global statistical information respectively, a protein sorting rule can also be corresponding.
[0049] Based on this, the N protein sorting rules in the above-mentioned pre-sorting of M candidate protein conformations based on N protein sorting rules can be either a combination of multiple different sorting rules under the same type of rules or a combination of sorting rules under different types of rules. That is, the multiple sorting rules can be a combination of one or more protein sorting rules under the sorting rules based only on the local information of the candidate protein, or a combination of one or more protein sorting rules under the sorting rules based on the global statistical information of the candidate protein, or can also include a combination composed of one or more protein sorting rules under the sorting rules based on the local information of the candidate protein and one or more protein sorting rules under the sorting rules based on the global statistical information of the candidate protein.
[0050] S130. Based on the fusion sorting rule, process the N pre-sorting results to obtain the target sorting results corresponding to the M candidate protein conformations.
[0051] Among them, the target sorting results can be considered as the sorting results obtained by the server after successively pre-sorting and processing through the fusion sorting rule for the M candidate protein objects corresponding to the target protein sequence.
[0052] Considering that the protein spatial structure is very complex, whether it is the sorting rule based on the local information of the candidate protein or the sorting rule based on the global statistical information of the candidate protein, it is very difficult to show good sorting performance on all protein structures. And because the data used by various sorting rules are different, the performance differences of various sorting rules on different protein spatial structures are large. For a target protein sequence whose protein spatial structure is unknown, it is not known which sorting rule will achieve better sorting effect after sorting. Therefore, after obtaining the N pre-sorting results corresponding to the M candidate protein conformations, the multiple pre-sorting results can be further processed based on the fusion sorting rule to obtain the target sorting results corresponding to the M candidate protein conformations.
[0053] The fusion sorting rule can integrate different statistical methods of N protein sorting rules, avoid the stability problems existing in using a single protein sorting rule, and enable good sorting performance on different protein structures.
[0054] As a way, a confidence level can be set in advance for the sorting position of the protein conformation in each pre-sorting result within the fusion sorting rule. For example, the confidence level of the candidate protein conformation ranked first in each pre-sorting result is X1, the confidence level of the candidate protein conformation ranked second in each pre-sorting result is X2, and the confidence level of the candidate protein conformation ranked third in each pre-sorting result is X3. Among them, X1 is greater than X2, and X2 is greater than X3, and so on. A confidence level can be set for each candidate protein conformation in each pre-sorting result. After obtaining multiple pre-sorting results, the confidence levels corresponding to the same candidate protein conformation in the N pre-sorting results can be added together to obtain the total confidence level corresponding to each candidate protein conformation. Then, according to the magnitude relationship of the total confidence levels of each candidate protein conformation, each candidate protein conformation is sorted to obtain the target sorting result corresponding to the M candidate protein conformations of the target protein sequence.
[0055] Among them, considering that the number of candidate protein conformations corresponding to the target protein sequence is large, and the higher the score of the candidate protein conformation ranked earlier in each pre-sorting result, the closer it is to the real conformation. Therefore, in order to simplify the calculation process, it can be considered to set confidence levels only for the preset number of protein conformations ranked close to the front in each pre-sorting result, and set the confidence levels of the remaining protein conformations to 0, or directly not set confidence levels. In this way, the number of candidate protein conformations used for calculation can be simplified, and the calculation accuracy will not be reduced.
[0056] Among them, weight coefficients corresponding to various protein sorting rules can also be pre-stored in the fusion sorting rule according to experience or experimental results. For example, in previous experiments, it is known that the accuracy of the protein sorting rule based on the planar angle between the main chain and side chain atoms of the protein is greater than that of other protein sorting rules. Then, a relatively large weight coefficient can be assigned to the protein sorting rule based on the planar angle between the main chain and side chain atoms of the protein, and a relatively small weight coefficient can be assigned to other sorting rules. Thus, when determining the total confidence level of each candidate protein conformation, the weight coefficients corresponding to each protein sorting rule can be multiplied in advance to obtain the confidence level corresponding to a certain candidate protein conformation under this protein sorting rule. Furthermore, the confidence levels corresponding to a certain candidate protein conformation under all protein sorting rules are added together to obtain the total confidence level of the M candidate protein conformations.
[0057] As another method, the N pre-ranking results can be processed by a trained neural network to obtain the target ranking results corresponding to the M candidate protein conformations. In this case, the fusion ranking rule can be trained based on the N sample pre-ranking results and the standard ranking results.
[0058] S140, determining a predicted protein conformation corresponding to the target protein sequence according to the target sorting result.
[0059] It can be understood that the target ranking result includes the ranking results obtained after each protein candidate conformation is processed by pre-ranking and fusion ranking rules in sequence, and has high ranking accuracy for various protein structures. Then, after obtaining the target ranking results corresponding to M candidate protein conformations, the predicted protein conformation corresponding to the target protein sequence can be determined according to the target ranking results.
[0060] Among them, considering that there is only one true protein conformation corresponding to the target protein sequence, as a method, the candidate protein conformation ranked first in the target sorting result can be determined as the predicted protein conformation corresponding to the target protein sequence.
[0061] In addition, considering that the candidate proteins ranked high in the target ranking results are likely to be the real protein spatial structures corresponding to the target protein sequence, in order to further improve the prediction accuracy, as another way, a preset number of candidate protein conformations ranked high in the target ranking results can also be used as the predicted protein conformations corresponding to the target protein sequence. In this case, determining the predicted protein conformation corresponding to the target protein sequence according to the target ranking results includes: obtaining a preset number of candidate protein conformations ranked high in the target ranking results as the predicted protein conformations corresponding to the target protein sequence.
[0062] The number corresponding to the preset quantity can be set as needed, for example, it can be set to three, five or ten.
[0063] In a specific application scenario, after obtaining a preset number of candidate protein conformations that rank high in the target sorting results as predicted protein conformations corresponding to the target protein sequence, the server can send the predicted protein conformations to the client. The client's experts can determine the true protein conformation from the predicted protein conformations based on manual experience.
[0064] Among them, various protein conformations in the embodiments of the present application can be directly transmitted or processed using pdb (protein data bank, protein three-dimensional structure data) files.
[0065] A protein conformation prediction method provided by an embodiment of the present application obtains M candidate protein conformations corresponding to a target protein sequence, then pre-orders the M candidate protein conformations based on N protein sorting rules to obtain N pre-sorting results corresponding to the M candidate protein conformations, and then processes the N pre-sorting results based on a fusion sorting rule to obtain a target sorting result corresponding to the M candidate protein conformations. Finally, according to the target sorting result, the predicted protein conformation corresponding to the target protein sequence is determined. Thus, through the foregoing method, on the basis of pre-ordering the M candidate proteins corresponding to the target protein sequence according to N protein sorting rules, and then processing the N pre-sorting results based on the fusion sorting rule. Since the fusion sorting rule is processed based on the N sorting results obtained by pre-sorting, it is equivalent to performing two-level sorting on the protein conformation, integrating different sorting methods of various protein sorting rules, avoiding the problem of poor sorting stability of traditional tools, enabling better prediction accuracy on various protein structures, and thus improving the overall accuracy of the predicted protein conformation.
[0066] As a way, the fusion sorting rule can be trained based on multiple sample pre-sorting results and a standard sorting result. In this case, please refer to Figure 4 , Figure 4 shown in the flowchart of a protein conformation prediction method proposed in an embodiment of the present application. In this embodiment, the method can be applied to a server, and the method includes:
[0067] S210, obtain m sample protein conformations and the corresponding standard sorting result, where the m sample protein conformations are obtained based on the sample protein sequence.
[0068] In this embodiment, the fusion sorting rule can be pre-stored in the server. Before storing the fusion sorting rule, it is necessary to establish the fusion sorting rule. As a way, the fusion sorting rule can be directly established in the server.
[0069] Specifically, the server pre-obtains a training sample set, where the training sample set includes multiple groups of training samples. Each group of training samples includes a sample protein sequence, m sample protein conformations corresponding to the sample protein sequence, and a standard sorting result corresponding to the sample protein sequence. Among them, the method for obtaining m sample protein conformations based on the sample protein sequence can refer to the content of obtaining M candidate protein conformations corresponding to the target protein sequence in the foregoing content.
[0070] As a way, as Figure 5 shown, obtaining m sample protein conformations and the corresponding standard sorting result includes:
[0071] S211. Obtain the m sample protein conformations corresponding to the sample protein sequence and the standard protein conformation corresponding to the sample protein sequence.
[0072] Among them, the method for obtaining the m sample protein conformations corresponding to the sample protein sequence can refer to the method for obtaining multiple candidate protein conformations corresponding to the target protein sequence in the foregoing content.
[0073] Among them, the standard protein conformation corresponding to the sample protein sequence refers to the true protein conformation corresponding to the sample protein sequence. The standard protein conformation corresponding to the sample protein sequence can be obtained from experiments. For example, a microscope can be used to take pictures to obtain the standard protein conformation corresponding to the sample protein sequence.
[0074] S212. Obtain the structural differences between each of the m sample protein conformations and the standard protein conformation.
[0075] After obtaining the m sample protein conformations corresponding to the sample protein sequence and the corresponding standard protein conformation, each of the sample protein conformations corresponding to the same sample protein sequence can be compared with the standard protein conformation corresponding to the sample protein sequence to obtain the structural differences between each of the m sample protein conformations and the standard protein conformation.
[0076] S213. Sort the m sample protein conformations based on the magnitudes of the structural differences to obtain a standard sorting result.
[0077] After obtaining the structural differences between each of the m sample protein conformations and the standard protein conformation, each of the sample protein conformations can be sorted based on the structural differences between each sample protein conformation and the standard protein conformation. Those with smaller differences are sorted earlier, and those with larger differences are sorted later.
[0078] As a way, the above S212 and S213 can use the scoring tool LDDT (The Local Distance Difference Test score measures) for sorting. LDDT can score each of the multiple predicted sample protein conformations according to the standard protein conformation. The scores given by LDDT represent the differences between the sample protein conformations and the standard protein conformation. Finally, sorting can be performed according to the scores to obtain a standard sorting result.
[0079] S220. Based on N protein sorting rules, pre-sort the m sample protein conformations to obtain N sample pre-sorting results corresponding to the m sample protein conformations.
[0080] Among them, the method for obtaining multiple sample pre-sorting results corresponding to multiple sample protein conformations can refer to the method for obtaining N pre-sorting results corresponding to multiple candidate protein conformations in the foregoing content.
[0081] S230. Based on the N sample pre-sorting results and the standard sorting results, train the initial model to obtain the fusion sorting rule.
[0082] After the server obtains the N sample pre-sorting results and the standard sorting results, it can use the N sample pre-sorting results and the standard sorting results as a set of training samples and input them into the initial model for training. After training with multiple sets of training samples, the fusion sorting rule can be obtained, which can also be called the fusion sorting model.
[0083] Among them, the initial model can be an SVM (Support Vector Machine), a neural network with one or more fully connected layers, etc.
[0084] As a way, as Figure 6 shown, based on the N sample pre-sorting results and the standard sorting results, train the initial model to obtain the fusion sorting rule, including:
[0085] S231. Input the N sample pre-sorting results into the initial model to obtain the predicted sorting result of the initial model.
[0086] S232. Construct a loss function based on the predicted sorting result and the standard sorting result.
[0087] S233. Update the parameters of the initial model based on the loss function.
[0088] In this embodiment, the training process of the fusion sorting rule includes: inputting the training data into the initial model, the initial model processes the N sample pre-sorting results, outputs a predicted sorting result, the initial model constructs a loss function based on the predicted sorting result and the standard sorting result, and updates the parameters in the initial model through the backpropagation algorithm. The above process is equivalent to completing one training of the initial model. By using multiple training samples for training and making the model repeat the above process, the fusion sorting rule can be obtained. Optionally, the loss function can adopt a square loss function, a cross-entropy loss function, or other loss functions.
[0089] Next, in combination with Figure 7 , a specific example is used to elaborate in detail on the training process of the fusion sorting rule in the embodiment of the present application.
[0090] As Figure 7As shown, a neural network with a single-layer simple fully connected layer is used as the initial model. The three sample sorting results obtained by the three protein sorting rules are input into the fully connected layer of the initial model to obtain the predicted sorting result output by the fully connected layer. Then, the predicted sorting result and the standard sorting result are jointly used to construct a loss function. Here, the loss function is selected as the squared loss function Loss, and Loss = (predicted sorting result - standard sorting result) * (predicted sorting result - standard sorting result). After obtaining the loss function, the parameters in the initial model are updated using the backpropagation algorithm, thereby completing one training of the initial model.
[0091] Among them, as Figure 7 shown, the standard sorting result is obtained by scoring the m sample protein conformations corresponding to the sample protein sequence and the standard protein conformation corresponding to the sample protein sequence using the scoring tool LDDT according to the structural difference and sorting according to the scoring results.
[0092] S240, obtain M candidate protein conformations corresponding to the target protein sequence, where M is a positive integer greater than 1.
[0093] S250, based on N protein sorting rules, pre-sort the M candidate protein conformations to obtain N pre-sorting results corresponding to the M candidate protein conformations, where N is a positive integer greater than 0.
[0094] S260, based on the fusion sorting rule, process the N pre-sorting results to obtain the target sorting result corresponding to the M candidate protein conformations.
[0095] S270, determine the predicted protein conformation corresponding to the target protein sequence according to the target sorting result.
[0096] A protein conformation prediction method provided by an embodiment of the present application obtains M candidate protein conformations corresponding to a target protein sequence, then pre-ranks the M candidate protein conformations based on N protein ranking rules to obtain N pre-ranking results corresponding to the M candidate protein conformations, and then processes the N pre-ranking results based on a fusion ranking rule trained based on N sample pre-ranking results and a standard ranking result to obtain a target ranking result corresponding to the M candidate protein conformations. Finally, according to the target ranking result, the predicted protein conformation corresponding to the target protein sequence is determined. Thus, through the foregoing method, on the basis of pre-ranking the M candidate proteins corresponding to the target protein sequence according to N protein ranking rules, and then processing the N pre-ranking results based on a fusion ranking rule trained based on N sample pre-ranking results and a standard ranking result. Since the fusion ranking rule processes based on the N ranking results obtained by pre-ranking, it is equivalent to performing two-level ranking on protein conformations, integrating different ranking methods of multiple protein ranking rules, avoiding the problem of poor ranking stability of traditional tools, enabling better prediction accuracy on various protein structures, and thus improving the overall accuracy of the predicted protein conformation.
[0097] As a method, after obtaining m sample protein conformations and their corresponding standard ranking results, the m sample protein conformations and their corresponding standard ranking results can also be directly used as training samples to train an initial model, thereby obtaining another ranking model. In this way, the function of ranking the M candidate protein conformations can also be realized, and thus the ranking result can be obtained. However, using this method, the advantages of multiple protein ranking rules cannot be considered, and the problem of overfitting is likely to occur.
[0098] Please refer to Figure 8 , Figure 8 The flowchart of a protein conformation prediction method proposed by an embodiment of the present application is shown. In this embodiment, the method can be applied to a server, and the method includes:
[0099] S310, search for homologous protein sequences corresponding to the target protein sequence in the protein database.
[0100] The protein database stores a variety of known protein structures, and each of these protein structures corresponds to a protein sequence. The corresponding relationship between the protein sequences and protein structures in the protein database was collected and saved in previous experiments.
[0101] Through the protein database, homologous protein sequences corresponding to the target protein sequence can be found. Among them, the homologous protein sequence has obvious similarity with the amino acid sequence of the target protein sequence and performs the same or similar functions in different organisms or within the same organism.
[0102] S320. Obtain the structural features corresponding to the target protein sequence based on the structural features corresponding to the homologous protein sequences.
[0103] After finding the homologous protein sequences corresponding to the target protein sequence, it is possible to continue to search for the protein structures corresponding to the homologous protein sequences in the protein data, so as to obtain the structural features corresponding to the homologous protein sequences, and then the structural features corresponding to the target protein sequence can be obtained based on the structural features corresponding to the homologous protein sequences.
[0104] As one way, the structural features corresponding to the homologous protein sequences can be directly used as the structural features corresponding to the target protein sequence.
[0105] As another way, the structural features corresponding to the homologous protein sequences can be input into the trained structure prediction neural network to obtain the predicted structural features output by the structure prediction network, and the predicted structural features are used as the structural features corresponding to the target protein sequence. Among them, there can be multiple predicted structural features.
[0106] S330. Obtain M candidate protein conformations corresponding to the target protein sequence based on the structural features corresponding to the target protein sequence.
[0107] After obtaining the structural features corresponding to the target protein sequence, a protein prediction algorithm can be used to predict multiple candidate protein conformations for each structural feature.
[0108] Among them, the protein prediction algorithm can be the Rosetta tool that uses the simulated annealing algorithm to predict protein structures, or the AphaFold tool.
[0109] S340. Pre-rank the M candidate protein conformations based on N protein ranking rules to obtain N pre-ranking results corresponding to the M candidate protein conformations, where N is a positive integer greater than 0.
[0110] As one way, the protein ranking rules can be based on the local information of the candidate proteins for ranking. Optionally, the protein ranking rules include at least one of the korpe energy ranking rule, the goap (generalized orientation-dependent all-atom potential) ranking rule, or the dope (Discrete Optimized Protein Energy) ranking rule.
[0111] S350. Process the N pre-ranking results based on the fusion ranking rules to obtain the target ranking results corresponding to the M candidate protein conformations.
[0112] S360. Determine the predicted protein conformation corresponding to the target protein sequence according to the target sorting result.
[0113] A protein conformation prediction method provided by an embodiment of the present application can, on the basis of pre-sorting M candidate proteins corresponding to the target protein sequence according to N protein sorting rules, further process the N pre-sorting results based on the fusion sorting rule. Since the fusion sorting rule processes based on the N sorting results obtained from the pre-sorting, it is equivalent to performing two-level sorting on the protein conformation, integrating different sorting methods of multiple protein sorting rules, avoiding the problem of poor sorting stability of traditional tools, enabling better prediction accuracy on various protein structures, and thus improving the overall accuracy of the predicted protein conformation.
[0114] In addition, since the amino acid sequences of homologous protein sequences and the target protein sequence have obvious similarities and perform the same or similar functions in different organisms or within the same organism, therefore, obtaining the predicted structural features of the target protein sequence based on the structural features of the homologous protein sequence corresponding to the target protein sequence, and then obtaining the M candidate protein conformations corresponding to the target protein sequence based on the predicted structural features can improve the accuracy of obtaining the candidate protein conformations.
[0115] Next, Figure 9 introduce a specific application process of the protein conformation prediction method provided by an embodiment of the present application.
[0116] As Figure 9 shown, first, the server receives the target protein sequence 410 sent by the client, searches for the homologous protein sequence corresponding to the target protein sequence 410 and the structural features corresponding to the homologous protein sequence in the protein database 420, then inputs the structural features corresponding to the homologous protein sequence into the trained structure prediction module 430 (a structure prediction neural network can be preset in the structure prediction module), obtains multiple predicted structural features output by the structure prediction module 430, and finally uses the protein conformation prediction module 440 (a protein prediction algorithm can be preset in the protein conformation prediction module) to predict multiple candidate protein conformations for each structural feature, so as to obtain the M candidate protein conformations corresponding to the target protein sequence.
[0117] Then, input the M candidate protein conformations into the protein conformation pre-sorting module 450. The protein conformation pre-sorting module 450 is preset with the GOAP sorting rule, the DOPE sorting rule, and the Korpe energy sorting rule. Use the GOAP sorting rule, the DOPE sorting rule, and the Korpe energy sorting rule respectively within the protein conformation pre-sorting module 450 to sort the M candidate protein conformations, and obtain the pre-sorting results corresponding to the GOAP sorting rule, the pre-sorting results corresponding to the DOPE sorting rule, and the pre-sorting results corresponding to the Korpe energy sorting rule respectively.
[0118] Next, use the three sorting results as features and input them into the fusion sorting module 460 (the fusion sorting rule is preset inside the fusion sorting module). After being processed by the fusion sorting module 460, the target protein sorting result can be obtained.
[0119] Finally, obtain the preset number of candidate protein conformations with the top rankings from the target sorting result, as the predicted protein conformations corresponding to the target protein sequence, and return them to the client.
[0120] Among them, the above-mentioned server can be a server deployed on Tencent Cloud.
[0121] It should be noted that this application provides some specific implementable examples above. On the premise of non-conflict, the examples of each embodiment can be combined arbitrarily to form a new protein conformation prediction method. It should be understood that for any new protein conformation prediction method formed by the combination of arbitrary examples, it should fall within the protection scope of this application.
[0122] Please refer to Figure 10 , Figure 10 which shows a block diagram of a protein conformation prediction device 500 proposed in an embodiment of this application. The device 500 includes: a candidate protein conformation acquisition module 510, a pre-sorting module 520, a target sorting result acquisition module 530, and a predicted protein conformation determination module 540.
[0123] The candidate protein conformation acquisition module 510 is used to acquire M candidate protein conformations corresponding to the target protein sequence, where M is a positive integer greater than 1.
[0124] As a way, the candidate protein conformation acquisition module 510 includes:
[0125] The homologous protein sequence search sub-module is used to search for the homologous protein sequence corresponding to the target protein sequence from the protein database;
[0126] The structure feature acquisition sub-module is used to obtain the structure feature corresponding to the target protein sequence based on the structure feature corresponding to the homologous protein sequence;
[0127] A candidate protein conformation obtaining sub-module, configured to obtain M candidate protein conformations corresponding to the target protein sequence based on the structural features corresponding to the target protein sequence.
[0128] A pre-sorting module 520, configured to pre-sort the M candidate protein conformations based on N protein sorting rules to obtain N pre-sorting results corresponding to the M candidate protein conformations, where N is a positive integer greater than 0.
[0129] As a way, the protein sorting rules are sorted based on the local information of the candidate protein. Optionally, the protein sorting rules include at least one of an energy sorting rule, a generalized orientation-related all-atom statistical potential rule, or a discrete optimization protein energy rule.
[0130] A target sorting result obtaining module 530, configured to process the N pre-sorting results based on a fusion sorting rule to obtain a target sorting result corresponding to the M candidate protein conformations.
[0131] As a way, the fusion sorting rule is trained based on N sample pre-sorting results and a standard sorting result.
[0132] A predicted protein conformation determining module 540, configured to determine a predicted protein conformation corresponding to the target protein sequence according to the target sorting result.
[0133] As a way, the predicted protein conformation determining module 540 includes: a predicted protein conformation obtaining sub-module.
[0134] The predicted protein conformation obtaining sub-module is configured to obtain a preset number of candidate protein conformations ranked at the top in the target sorting result as the predicted protein conformation corresponding to the target protein sequence.
[0135] As a way, the apparatus 500 further includes:
[0136] A sample obtaining module, configured to obtain m sample protein conformations and corresponding standard sorting results, where the m sample protein conformations are obtained based on sample protein sequences.
[0137] A sample pre-sorting module, configured to pre-sort the m sample protein conformations based on N protein sorting rules to obtain N sample pre-sorting results corresponding to the m sample protein conformations.
[0138] A training module, configured to train an initial model based on the N sample pre-sorting results and the standard sorting result to obtain a fusion sorting rule.
[0139] As a way, the sample obtaining module includes:
[0140] A sample acquisition sub-module, configured to acquire m sample protein conformations corresponding to a sample protein sequence and a standard protein conformation corresponding to the sample protein sequence.
[0141] A structural difference acquisition sub-module, configured to acquire the structural differences between each of the m sample protein conformations and the standard protein conformation.
[0142] A standard sorting result acquisition sub-module, configured to sort the m sample protein conformations based on the magnitudes of the structural differences to obtain a standard sorting result.
[0143] As a manner, the training module includes:
[0144] A predicted sorting result acquisition sub-module, configured to input N sample pre-sorting results into an initial model to obtain a predicted sorting result of the initial model.
[0145] A loss function construction sub-module, configured to construct a loss function based on the predicted sorting result and the standard sorting result.
[0146] A parameter update sub-module, configured to update the parameters of the initial model based on the loss function.
[0147] A protein conformation prediction device provided by an embodiment of the present application can, on the basis of pre-sorting M candidate proteins corresponding to a target protein sequence according to N protein sorting rules, further process the N pre-sorting results based on a fusion sorting rule. Since the fusion sorting rule processes the N sorting results obtained from the pre-sorting, it is equivalent to performing two-level sorting on the protein conformation, integrating different sorting methods of multiple protein sorting rules, avoiding the problem of poor sorting stability of traditional tools, enabling better prediction accuracy for various protein structures, and thus improving the overall accuracy of the predicted protein conformation.
[0148] It should be noted that the device embodiment in the present application corresponds to the foregoing method embodiment. The specific principle in the device embodiment can refer to the content in the foregoing method embodiment, and will not be elaborated herein.
[0149] Next, a description will be given in conjunction with Figure 11 an electronic device provided by the present application.
[0150] Please refer to Figure 11, based on the above protein conformation prediction method, another electronic device 200 including a processor 104 that can execute the aforementioned protein conformation prediction method is also provided in an embodiment of the present application. The electronic device 200 can be a device such as a smartphone, a tablet computer, a computer, or a portable computer. The electronic device 200 further includes a memory 104, a network module 106, and a screen 108. Among them, a program that can execute the content in the foregoing embodiments is stored in the memory 104, and the processor 102 can execute the program stored in the memory 104.
[0151] Among them, the processor 102 may include one or more cores for processing data and a message matrix unit. The processor 102 connects various parts within the entire electronic device 200 through various interfaces and lines, and executes various functions of the electronic device 200 and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 104, and by calling data stored in the memory 104. Optionally, the processor 102 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 102 may integrate a combination of one or several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing display content; the modem is used to process wireless communication. It can be understood that the above modem may not be integrated into the processor 102 and may be implemented separately by a communication chip.
[0152] The memory 104 may include random access memory (RAM) and may also include read-only memory. The memory 104 can be used to store instructions, programs, codes, code sets, or instruction sets. The memory 104 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for implementing at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the following various method embodiments, etc. The data storage area may also store data created during the use of the terminal 100 (such as phone book, audio and video data, etc.).
[0153] The network module 106 is used to receive and transmit electromagnetic waves, realize the mutual conversion between electromagnetic waves and electrical signals, so as to communicate with a communication network or other devices, such as communicating with an audio playback device. The network module 106 may include various existing circuit elements for performing these functions, such as, antennas, radio frequency transceivers, digital signal processors, encryption / decryption chips, subscriber identity module (SIM) cards, memories, and so on. The network module 106 can communicate with various networks such as the Internet, an enterprise intranet, a wireless network or communicate with other devices through a wireless network. The above-mentioned wireless network may include a cellular phone network, a wireless local area network or a metropolitan area network. For example, the network module 106 can interact with a base station.
[0154] The screen 108 can display interface content and can also be used to respond to touch gestures.
[0155] It should be noted that, in order to implement more functions, the electronic device 200 may also protect more components. For example, it may also protect a structured light sensor for collecting face information or may also protect a camera for collecting irises, etc.
[0156] Please refer to Figure 12 , which shows a structural block diagram of a computer-readable storage medium provided by an embodiment of the present application. Program code is stored in the computer-readable medium 1100, and the program code can be called by a processor to execute the method described in the above method embodiments.
[0157] The computer-readable storage medium 1100 may be an electronic memory such as a flash memory, an EEPROM (electrically erasable programmable read-only memory), an EPROM, a hard disk or a ROM. Optionally, the computer-readable storage medium 1100 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 1100 has a storage space for the program code 1110 that executes any method step in the above method. These program codes can be read from one or more computer program products or written into one or more computer program products. The program code 1110 may be compressed in a suitable form, for example.
[0158] Based on the above protein conformation prediction method, according to one aspect of the embodiments of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the above various optional implementation manners.
[0159] In summary, a method, apparatus, electronic device, storage medium, computer program product or computer program for predicting protein conformation provided by the embodiments of the present application can, on the basis of pre-ranking M candidate proteins corresponding to a target protein sequence according to N protein ranking rules, further process the N pre-ranking results based on a fusion ranking rule. Since the fusion ranking rule processes the N ranking results obtained from the pre-ranking, it is equivalent to performing two-level ranking on the protein conformation, integrating different ranking methods of multiple protein ranking rules, avoiding the problem of poor ranking stability of traditional tools, enabling better prediction accuracy for various protein structures, and thus improving the overall accuracy of the predicted protein conformation.
[0160] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for predicting protein conformation, characterized in that, Including: Obtaining M candidate protein conformations corresponding to a target protein sequence, where M is a positive integer greater than 1; Based on N protein sorting rules, pre-sorting the M candidate protein conformations to obtain N pre-sorting results corresponding to the M candidate protein conformations, where N is a positive integer greater than 0; Processing the N pre-sorting results through a fusion sorting model to obtain a target sorting result corresponding to the M candidate protein conformations; the fusion sorting model is obtained by training an initial model with N sample pre-sorting results and corresponding standard sorting results corresponding to m sample protein conformations; the N sample pre-sorting results are obtained by sorting m sample protein conformations corresponding to a sample protein sequence according to the N protein sorting rules; determining a predicted protein conformation corresponding to the target protein sequence according to the target sorting result; Wherein, the obtaining M candidate protein conformations corresponding to the target protein sequence includes: Searching in a protein database for a homologous protein sequence corresponding to the target protein sequence; Taking the structural features corresponding to the homologous protein sequence as the structural features corresponding to the target protein sequence; Through a protein prediction algorithm, predicting according to the structural features corresponding to the target protein sequence to obtain M candidate protein conformations corresponding to the target protein sequence.
2. The method according to claim 1, characterized in that, Before the processing the N pre-sorting results through the fusion sorting model to obtain a target sorting result corresponding to the M candidate protein conformations, the method further includes: Obtaining m sample protein conformations and corresponding standard sorting results, the m sample protein conformations are obtained based on a sample protein sequence; Based on the N protein sorting rules, pre-sorting the m sample protein conformations to obtain N sample pre-sorting results corresponding to the m sample protein conformations; Based on the N sample pre-sorting results and the standard sorting results, training the initial model to obtain the fusion sorting model.
3. The method according to claim 2, characterized in that, The obtaining m sample protein conformations and corresponding standard sorting results includes: Obtaining m sample protein conformations corresponding to a sample protein sequence and a standard protein conformation corresponding to the sample protein sequence; Obtaining the structural differences between the m sample protein conformations and the standard protein conformation respectively; Based on the magnitudes of the structural differences, sorting the m sample protein conformations to obtain the standard sorting result.
4. The method according to claim 2, characterized in that, The training the initial model based on the N sample pre-sorting results and the standard sorting results to obtain the fusion sorting model includes: Inputting the N sample pre-sorting results into the initial model to obtain a predicted sorting result of the initial model; Constructing a loss function based on the predicted sorting result and the standard sorting result; Based on the loss function, updating the parameters of the initial model, and the trained initial model is used as the fusion sorting model.
5. The method according to claim 1, characterized in that, The N protein ranking rules include a rule for ranking based on at least one local information of the candidate protein conformation, and a rule for ranking based on one local information is a protein ranking rule.
6. The method according to claim 5, characterized in that, The protein ranking rule comprises at least one of an energy ranking rule, a generalized orientation-dependent all-atom statistical potential rule, or a discrete optimized protein energy rule.
7. The method according to any one of claims 1-6, characterized in that, Determining the predicted protein conformation corresponding to the target protein sequence according to the target sorting result includes: A preset number of the candidate protein conformations that are ranked high in the target ranking result are obtained as predicted protein conformations corresponding to the target protein sequence.
8. A device for predicting protein conformation, characterized in that, include: A candidate protein conformation acquisition module is used to acquire M candidate protein conformations corresponding to the target protein sequence, where M is a positive integer greater than 1; A pre-sorting module, used for pre-sorting the M candidate protein conformations based on N protein sorting rules, to obtain N pre-sorting results corresponding to the M candidate protein conformations, wherein N is a positive integer greater than 0; The target ranking result obtaining module is used to process the N pre-ranking results through a fusion ranking model to obtain the target ranking results corresponding to the M candidate protein conformations; the fusion ranking model is obtained by training the initial model through the N sample pre-ranking results corresponding to the m sample protein conformations and the corresponding standard ranking results; the N sample pre-ranking results are obtained by respectively ranking the m sample protein conformations corresponding to the sample protein sequences according to the N protein ranking rules; A predicted protein conformation determination module, used to determine the predicted protein conformation corresponding to the target protein sequence according to the target sorting result; Wherein, the candidate protein conformation acquisition module comprises: A homologous protein sequence search submodule is used to search for homologous protein sequences corresponding to the target protein sequence from a protein database; A structural feature acquisition submodule, used to use the structural features corresponding to the homologous protein sequence as the structural features corresponding to the target protein sequence; The candidate protein conformation obtaining submodule is used to predict M candidate protein conformations corresponding to the target protein sequence based on the structural features corresponding to the target protein sequence through a protein prediction algorithm.
9. The device according to claim 8, wherein, The protein conformation prediction device also includes: A sample acquisition module, used to acquire m sample protein conformations and corresponding standard sorting results, wherein the m sample protein conformations are obtained based on sample protein sequences; A sample pre-sorting module, used to pre-sort the m sample protein conformations based on the N protein sorting rules, and obtain N sample pre-sorting results corresponding to the m sample protein conformations; The training module is used to train the initial model based on the N types of sample pre-sorting results and the standard sorting results to obtain the fusion sorting model.
10. The device according to claim 9, wherein, The sample acquisition module comprises: A sample acquisition submodule, used to acquire m sample protein conformations corresponding to a sample protein sequence and a standard protein conformation corresponding to the sample protein sequence; A structural difference acquisition submodule, used for acquiring the structural differences between the m sample protein conformations and the standard protein conformation; The standard sorting result acquisition submodule is used to sort the m sample protein conformations based on the size of the structural differences to obtain the standard sorting result.
11. The device according to claim 9, wherein, The training module comprises: A prediction ranking result acquisition submodule is used to input the N kinds of sample pre-ranking results into the initial model to obtain the prediction ranking results of the initial model; A loss function construction submodule, used to construct a loss function based on the predicted sorting result and the standard sorting result; The parameter updating submodule is used to update the parameters of the initial model based on the loss function.
12. The device according to claim 8, wherein, The N protein ranking rules include a rule for ranking based on at least one local information of the candidate protein conformation, and a rule for ranking based on one local information is a protein ranking rule.
13. The device according to claim 12, wherein, The protein ranking rule comprises at least one of an energy ranking rule, a generalized orientation-dependent all-atom statistical potential rule, or a discrete optimized protein energy rule.
14. The device according to any one of claims 8 - 13, wherein, The predicted protein conformation determination module includes a predicted protein conformation acquisition submodule, which is used to obtain a preset number of candidate protein conformations that are ranked high in the target sorting results as predicted protein conformations corresponding to the target protein sequence.
15. An electronic device, wherein, include: one or more processors; Memory; One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the method according to any one of claims 1 to 7.
16. A computer-readable storage medium, wherein, The computer-readable storage medium stores program codes, which can be called by a processor to execute the method according to any one of claims 1 to 7.
17. A computer program product, wherein, The method comprises computer instructions, which, when executed by a processor, implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Protein remote homology detecting method and device
CN104636636A
Protein sequence design implementation method based on multi-objective optimization
CN111554346A
Protein structure analysis method, protein structure analyzing instrument, program and recording medium
US20090319234A1