Prediction method, model training method, device and electronic equipment for complex structure
By acquiring the sequence and structural information of receptors and ligands, and combining the orientation information with a multi-round iteratively corrected structural prediction model, the problem of insufficient accuracy in predicting complex structures in existing technologies has been solved, and accurate prediction of complex structures with the same sequence but different structures has been achieved.
Patent Information
- Application Number
- CN202411303237.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-18
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2044-09-18
AI Technical Summary
Existing technologies struggle to accurately predict the structure of receptor-ligand binding complexes, especially when sequences are identical but structures differ, leading to poor prediction accuracy.
By obtaining the amino acid sequence of the receptor and the sequence information of the ligand, as well as the protein structure of the receptor, the orientation information of the ligand relative to the protein structure is determined. Combined with the structure prediction model, multiple rounds of iterative correction are performed to improve the prediction accuracy.
It enables accurate differentiation of complex structures with the same sequence but different structures, thus improving the accuracy of predicting complex structures.
Smart Images

Figure CN119360947B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, in particular to the technical field of deep learning and biological computing, and more particularly to a complex structure prediction method, a model training method and device, and an electronic device. BACKGROUND
[0002] Proteins, as macromolecular compounds supporting almost all cellular operations, play an important role in different biological fields, such as enzyme reactions, cell signal transduction, metabolic regulation, and gene expression. Among them, the technology of predicting the three-dimensional structure (tertiary structure) of proteins in space according to the amino acid category of protein chains (primary structure) has a very high research value in the field of life sciences. For example, in the new drug discovery scenario, researchers need to find ligands (polypeptides, small molecules, etc.) for receptors, and the binding mode / interfacial structure of the receptor and the ligand can provide useful auxiliary information for the research of the receptor-related drugs.
[0003] Therefore, how to accurately predict the complex structure of the combination of the receptor and the ligand is very important. SUMMARY
[0004] The present application provides a complex structure prediction method, a model training method, a device and an electronic device.
[0005] According to an aspect of the present application, a complex structure prediction method is provided, comprising:
[0006] obtaining the amino acid sequence of a receptor and the sequence information of a ligand to be combined, and the protein structure of the receptor;
[0007] determining the orientation information of the ligand relative to the protein structure based on the protein structure of the receptor;
[0008] determining the target complex structure of the combination of the receptor and the ligand according to the orientation information of the ligand relative to the protein structure, the amino acid sequence of the receptor and the sequence information of the ligand.
[0009] According to another aspect of the present application, a model training method is provided, comprising:
[0010] obtaining a training sample, wherein the training sample comprises the amino acid sequence of a receptor and the sequence information of a ligand to be combined, and the orientation information of the ligand relative to the protein structure of the receptor;
[0011] inputting the training sample into a structure prediction model to obtain the predicted complex structure of the receptor and the ligand;
[0012] determine a loss function of the structure prediction model according to a difference between the predicted complex structure and a target complex structure annotated by the training sample;
[0013] train the structure prediction model according to the loss function, to obtain the trained structure prediction model.
[0014] According to an aspect of the present application, a complex structure prediction device is provided, comprising:
[0015] an obtaining module configured to obtain sequence information of a receptor and a ligand to be combined, and a protein structure of the receptor;
[0016] a determining module configured to determine orientation information of the ligand relative to the protein structure based on the protein structure of the receptor;
[0017] a predicting module configured to determine a target complex structure of the receptor and the ligand combined according to the orientation information of the ligand relative to the protein structure, the sequence information of the receptor and the sequence information of the ligand.
[0018] According to another aspect of the present application, a model training device is provided, comprising:
[0019] an obtaining module configured to obtain a training sample, wherein the training sample comprises sequence information of a receptor and a ligand to be combined, and orientation information of the ligand relative to a protein structure of the receptor;
[0020] a predicting module configured to input the training sample into a structure prediction model, to obtain a predicted complex structure of the receptor and the ligand;
[0021] a determining module configured to determine a loss function of the structure prediction model according to a difference between the predicted complex structure and a target complex structure annotated by the training sample;
[0022] a training module configured to train the structure prediction model according to the loss function, to obtain a trained structure prediction model.
[0023] According to another aspect of the present application, an electronic device is provided, comprising:
[0024] at least one processor; and
[0025] a memory in communication connection with the at least one processor; wherein
[0026] The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method described in the above embodiments.
[0027] According to another aspect of the present application, a non-transitory computer readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to perform the method described in the above embodiments.
[0028] According to another aspect of the present application, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the steps of the method described in the above embodiments.
[0029] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0030] The accompanying drawings are used to better understand the present application, and do not limit the present application. Among them:
[0031] Figure 1 A flowchart of a prediction method of a complex structure provided by an embodiment of the present application is shown;
[0032] Figure 2 A flowchart of another prediction method of a complex structure provided by an embodiment of the present application is shown;
[0033] Figure 3 A flowchart of another prediction method of a complex structure provided by an embodiment of the present application is shown;
[0034] Figure 4 A flowchart of another prediction method of a complex structure provided by an embodiment of the present application is shown;
[0035] Figure 5 A schematic diagram of a ligand position determination provided by an embodiment of the present application is shown;
[0036] Figure 6 A flowchart of another prediction method of a complex structure provided by an embodiment of the present application is shown;
[0037] Figure 7 A structural schematic diagram of a structure prediction model provided by an embodiment of the present application is shown;
[0038] Figure 8 A flowchart of a model training method provided by an embodiment of the present application is shown;
[0039] Figure 9A structural schematic diagram of a prediction device for a complex structure provided by an embodiment of the present application is shown in the figure.
[0040] Figure 10 A structural schematic diagram of a model training device provided by an embodiment of the present application is shown in the figure.
[0041] Figure 11 A structural schematic diagram of an electronic device provided by an embodiment of the present application is shown in the figure. DETAILED DESCRIPTION
[0042] Exemplary embodiments of the present application are described below with reference to the accompanying drawings, which include various details of the embodiments of the present application to assist in understanding, and should be considered as merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Also, in order to be clear and concise, descriptions of well-known functions and structures are omitted in the following description.
[0043] In related art, protein structure prediction methods based on artificial intelligence AI models, such as the artificial intelligence program AlphaFold, are gradually approaching experimental observation results. The model directly predicts the complex structure of the receptor and the ligand binding from the sequence information of the receptor and the ligand, and the complex structure is also called the binding conformation. However, the same sequence of protein complex in nature can have different structures, as an example, as shown in the left figure in Figure 1 Chain A of protein 7eid can bind to two ligands C or D, and the two ligands are peptide chains, wherein the amino acid sequences of the peptide chain C and the peptide chain D are completely identical, both of which are KGGKGLGKGG. The left figure shows the real complex structure of chain A binding to the peptide chain C and the peptide chain D, i.e. the real binding conformation of ACD. If the input of the model only considers sequence information, the model is difficult to distinguish the binding mode of 7eid sample A-C chain and A-D chain. Taking the sequence information of A chain and C chain as an example, as shown in the right figure in Figure 1 The predicted complex structure of AC binding is quite different from the real result in the left figure, resulting in poor prediction accuracy.
[0044] To solve the above problems, the present application provides a complex structure prediction method, a model training method, a device and an electronic device.
[0045] The complex structure prediction method, the model training method, the device and the electronic device of the embodiments of the present application are described below with reference to the accompanying drawings.
[0046] It should be noted that in the technical solutions of the present application, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in the technical solutions of the present application are all carried out on the premise of obtaining the consent of the user, and all comply with the relevant legal regulations and do not violate public order and good customs.
[0047] Figure 2 A flowchart of a prediction method of a complex structure provided by an embodiment of the present application is shown.
[0048] The prediction method of the complex structure of the embodiment of the present application can be executed by a prediction device of the complex structure of the embodiment of the present application, and the device can be configured in an electronic device to realize the prediction of the complex structure.
[0049] The electronic device can be any device with computing capability, such as a personal computer, a mobile terminal, a server, etc. The mobile terminal can be a vehicle-mounted device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, etc. with various operating systems, touch screens and / or display screens.
[0050] As shown in the method, the method comprises the following steps: Figure 2
[0051] In step 201, the amino acid sequence of the receptor and the sequence information of the ligand to be combined, and the protein structure of the receptor are obtained.
[0052] The receptor is a macromolecular compound, usually a protein. The receptor includes small molecules and polypeptides, etc. The amino acid sequence of the receptor is usually a long chain composed of multiple amino acids, which is the primary structure. The protein structure of the receptor is the tertiary structure. The ligand to be combined can be combined with the receptor to obtain a complex three-dimensional structure of the receptor and the ligand, also known as the binding conformation.
[0053] The amino acid sequence can be obtained from an existing test set, or the amino acid sequence can be collected online, such as through web crawler technology, or the amino acid sequence can be provided by a user, which is not limited in the embodiment of the present application.
[0054] In the case of a small molecule ligand, the sequence information of the ligand includes small molecule information, and the small molecule information includes a series of sequence data describing the structure, properties and functions of the small molecule. In the case of a peptide chain, the sequence information of the ligand includes the amino acid sequence of the peptide chain.
[0055] As an implementation, the protein structure of the receptor can be predicted by using an artificial intelligence program AlphaFold in related technologies, wherein AlphaFold has high accuracy in predicting the three-dimensional structure of a single substance, so that the amino acid sequence of the receptor is input into the model of AlphaFold trained to obtain the protein structure of the receptor. As another implementation, if the real protein structure of the receptor is a known structure, it can be directly obtained from a stored database.
[0056] Step 202, based on the protein structure of the receptor, determine the orientation information of the ligand relative to the protein structure.
[0057] In an implementation manner of the embodiment of the present application, the prediction according to the sequence data has poor accuracy, and therefore, the orientation information of the ligand relative to the protein structure when the ligand and the receptor are combined is determined on the protein structure of the receptor, wherein the orientation information includes the position region of the ligand and the direction information of the position region, can indicate the spatial position and direction of the ligand when the ligand and the receptor are combined, and realize the positioning of the ligand position.
[0058] Step 203, according to the orientation information of the ligand relative to the protein structure, the amino acid sequence of the receptor and the sequence information of the ligand, determine the target complex structure of the combination of the receptor and the ligand.
[0059] In the embodiment of the present application, the amino acid sequence of the receptor and the sequence information of the ligand provide the sequence information of the receptor and the ligand, and the possible combination position and combination structure information of the receptor and the ligand can be determined according to the sequence information, but due to the case that the same sequence corresponds to different structures, in order to improve the accuracy, the orientation information of the ligand when the ligand and the receptor are combined is provided in the present application, the initial spatial position and structure information of the combination of the ligand and the receptor are provided, and then the structure is continuously corrected based on the above sequence information and orientation information, to obtain the final target complex structure predicted, and the prediction accuracy is improved.
[0060] In the complex structure prediction method of the embodiment of the present application, the amino acid sequence of the receptor and the sequence information of the ligand to be combined, and the protein structure of the receptor are obtained, the orientation information of the ligand relative to the protein structure is determined based on the protein structure of the receptor, and the target complex structure of the combination of the receptor and the ligand is determined according to the orientation information of the ligand relative to the protein structure, the amino acid sequence of the receptor and the sequence information of the ligand. By providing the orientation information of the ligand relative to the receptor when the ligand and the receptor are combined, the different structures of the complex with the same sequence in nature can be better distinguished, and the complex structure prediction effect is improved.
[0061] Based on the above embodiment, Figure 3 Another flowchart of the complex structure prediction method provided by the embodiment of the present application is shown as follows, Figure 3 The method comprises the following steps:
[0062] Step 301, obtain the amino acid sequence of the receptor and the sequence information of the ligand to be combined, and the protein structure of the receptor.
[0063] Step 302, based on the protein structure of the receptor, determine the orientation information of the ligand relative to the protein structure.
[0064] Wherein, the step 301 and step 302 can refer to the explanation and description in the foregoing embodiments, the principle is the same, and here is not described again.
[0065] Step 303, according to the orientation information of the ligand relative to the protein structure, determine the initial complex structure of the receptor and the ligand binding.
[0066] In the embodiments of the application, the orientation information of the ligand includes the spatial relative position relationship when the ligand and the receptor are combined, and under the condition of determining the protein structure of the receptor and the orientation information of the ligand when the ligand and the receptor are combined, the initial complex structure of the receptor and the ligand binding can be determined, which indicates the initial conformation of the ligand and the receptor binding.
[0067] Step 304, according to the initial complex structure, the amino acid sequence of the receptor and the sequence information of the ligand, determine the target complex structure of the receptor and the ligand binding.
[0068] In the embodiments of the application, the initial complex structure of the ligand and the receptor binding is provided, the initial spatial position and structure information of the ligand and the receptor binding is provided, and then the sequence information and the initial structure information are continuously corrected to obtain the final predicted target complex structure, which realizes the distinction of the complex structures with the same sequence but different structures in nature based on the initial complex structure, and improves the accuracy of the complex prediction.
[0069] In the complex structure prediction method of the embodiments of the application, the amino acid sequence of the receptor and the sequence information of the ligand to be combined, and the protein structure of the receptor are obtained, the orientation information of the ligand relative to the protein structure is determined based on the protein structure of the receptor, and the target complex structure of the receptor and the ligand binding is determined according to the orientation information of the ligand relative to the protein structure, the amino acid sequence of the receptor and the sequence information of the ligand. By providing the orientation information of the ligand relative to the receptor when the ligand and the receptor are combined, the initial spatial position relationship of the ligand and the receptor binding is indicated, and the prediction is based on the orientation information and the sequence information, which can better distinguish the ligands with the same sequence but different structures and the complex structures with the same sequence but different structures in nature, thereby improving the accuracy of the final predicted target complex structure.
[0070] Based on the above embodiments, Figure 4 Another flowchart of the complex structure prediction method provided in the embodiments of the application is shown as follows, Figure 4 The method comprises the following steps:
[0071] Step 401, obtaining the amino acid sequence of the receptor and the sequence information of the ligand to be combined, and the protein structure of the receptor.
[0072] In step 401, the protein structure is displayed on the interactive interface of the electronic device.
[0073] In step 402, the protein structure is displayed on the interactive interface of the electronic device.
[0074] In the embodiments of the present application, the protein structure is displayed on the interactive interface of the electronic device to determine the orientation information of the ligand of the protein structure set through interaction.
[0075] In step 403, the binding position is determined on the protein structure according to the detected first instruction.
[0076] The first instruction can be generated during the interaction between the interaction device and the electronic device, for example, the interaction device is a mouse or a touch panel, or the first instruction is realized by operation in the interactive interface of the electronic device, for example, a click operation. The first instruction is used to determine the binding position of the ligand and the receptor when the ligand and the receptor are combined, also known as the binding site.
[0077] In step 404, the position region of the ligand in space and the direction information of the position region are determined according to the detected second instruction and the binding position.
[0078] Similarly, the second instruction can refer to the description of the first instruction, and will not be described in detail in the embodiments. The second instruction is used to draw from the binding position as the starting point to determine the position region of the ligand and the direction information of the position region in the space around the protein structure.
[0079] Through interaction, prior knowledge such as expert experience of drug researchers can be obtained to determine the approximate orientation information of the ligand based on the protein structure. By providing the relative position information of the receptor and the ligand, the position and direction of the ligand are provided when the complex structure is predicted based on sequence information, and the prediction accuracy is improved.
[0080] As an example, Figure 5 A schematic diagram of ligand position determination provided by the embodiments of the present application is shown in FIG. 1. Figure 5 As shown in FIG. 1, the receptor is chain A of protein 7eid, and the ligand can be peptide chain C or peptide chain D, wherein the amino acid sequences of the peptide chain C and the peptide chain D are completely identical. Figure 5 As shown in FIG. 1, the binding point position of the peptide chain C combined with the chain A, the position region of the peptide chain C and the direction information of the position region, and the binding point position of the peptide chain D combined with the chain A, and the position region of the peptide chain D and the direction information of the position region are displayed. Figure 5As can be seen, the amino acid sequences of the peptide chain C and the peptide chain D are consistent, but the orientation information, i.e., the relative position information, is different, that is, the position and direction of the ligand are different, so as to identify the sequence information of different ligands or complexes in nature according to the orientation information of the added ligand, and improve the accuracy of model prediction.
[0081] In step 405, the initial three-dimensional positions of the atoms of the ligand in the position region are determined according to the position region and the direction information of the position region.
[0082] In step 406, the initial complex structure is obtained according to the protein structure of the receptor and the initial three-dimensional positions of the atoms of the ligand.
[0083] In the embodiment of the application, according to the position region and the direction information of the position region, the three-dimensional coordinate interval range of the corresponding position region when the ligand and the receptor are combined can be determined, and then the initial three-dimensional positions of the atoms included in the ligand are initialized in the position region where the ligand is located according to the random algorithm, so as to further determine the positions of the atoms of the ligand from the region, to accurately obtain the initial complex structure of the ligand and the receptor set, and to provide accurate spatial structure information.
[0084] In step 407, the initial complex structure is locally transformed to obtain the first complex structure.
[0085] In the embodiment of the application, the first complex structure is obtained by locally transforming the initial complex structure, which realizes the local noise addition to the complex structure, and the local loading makes the first complex structure carry more original structure information, which can improve the accuracy of subsequent model prediction. Specifically, the following methods can be used to achieve the above purpose.
[0086] In one implementation manner of the embodiment of the application, a target atom in the initial complex structure is randomly determined, which can be a target atom in the atoms of the ligand, i.e., the target atom is part of all atoms of the ligand, or the target atom is a target atom including the atoms in the ligand and the receptor, i.e., the target atom is part of all atoms included in the ligand and the receptor. Then, the position of the target atom is offset, so as to change the initial complex structure to obtain the first complex structure.
[0087] In another implementation manner of the embodiment of the application, the initial complex structure is twisted, so as to change the initial complex structure to obtain the first complex structure.
[0088] In still another implementation manner of the embodiment of the application, the initial complex structure is rotated, so as to change the initial complex structure to obtain the first complex structure.
[0089] At step 408, a structure prediction model is used to determine a target complex structure of the receptor and the ligand based on the first complex structure, the amino acid sequence of the receptor, and the sequence information of the ligand.
[0090] The structure prediction model may be, for example, a model of an artificial intelligence program AlphaFold, and is not limited in the present embodiment.
[0091] In the present embodiment, the structure prediction model is used to determine the target complex structure of the receptor and the ligand by continuously correcting the three-dimensional positions of the atoms in the complex structure based on the first complex structure, the amino acid sequence of the receptor, and the sequence information of the ligand, thereby improving accuracy.
[0092] The method for predicting a complex structure according to the present embodiment can determine the three-dimensional coordinate range of the position region of the ligand when the ligand and the receptor are combined based on the position region of the ligand and the direction information of the position region determined based on prior experience, and then initialize the initial three-dimensional positions of the atoms included in the ligand in the position region of the ligand based on a random algorithm, so as to further determine the positions of the atoms of the ligand from the region, to accurately obtain the initial complex structure of the ligand and the receptor, and to provide accurate spatial structure information. Furthermore, the first complex structure is obtained by performing local transformation on the initial complex structure, which realizes local noise addition to the complex structure, and the local addition makes the first complex structure carry more original structure information, which can improve the accuracy of subsequent model prediction.
[0093] Based on the above embodiments, Figure 6 Another flowchart of a method for predicting a complex structure according to the present embodiment is shown in FIG. 6. Figure 6 The method includes the following steps:
[0094] At step 601, the amino acid sequence of the receptor and the sequence information of the ligand to be combined, and the protein structure of the receptor are obtained.
[0095] At step 602, the orientation information of the ligand relative to the protein structure is determined based on the protein structure of the receptor.
[0096] At step 603, the initial complex structure of the receptor and the ligand is determined based on the orientation information of the ligand relative to the protein structure.
[0097] At step 604, the initial complex structure is locally transformed to obtain the first complex structure.
[0098] The steps 601 to 604 can refer to the related explanations and descriptions in the foregoing embodiments, and the principles are the same, which will not be described here again.
[0099] Step 605: The coding network in the structural prediction model is used to encode the amino acid sequence of the receptor and the sequence information of the ligand to be bound, so as to obtain the fusion coding feature.
[0100] Among them, the coding network refers to a network that can encode the sequence information of amino acids and ligands, such as the coding network of AlphaFold.
[0101] As an example, such as Figure 7 As shown, the ligand is a small molecule. To increase the amount of input information, the input data includes the amino acid sequence of the receptor and the sequence information of the small molecule. Covalent information can also be included. First, the amino acid sequence of the receptor protein and the structural template are searched in gene databases and structural databases (Genetic search) and (Template search). Reference conformation information (Conformer generation) is obtained based on the small molecule ligand search. Then, the amino acid sequence, structural template, reference conformation, covalent information, etc., are input into the input embedder, template module, multiple sequence alignment (MSA) module, etc. in the coding network in the form of sequences. Finally, the fused coding features are obtained through the 48-layer pairforme module.
[0102] The fusion coding features include original sequence features, first coding features, and second coding features. The first coding features carry distance information and relative position coding information between atoms, while the second coding features are obtained by concatenating multiple input sequences and carry topological features, geometric features, interaction features, combination features, and functional features between atoms.
[0103] Step 606: Using the decoding network in the structural prediction model, perform a set number of iterations based on the first composite structure and fusion coding features to update the structure of the receptor-ligand binding complex.
[0104] In this embodiment, the decoding network is also called a diffusion module. During the iteration process of a set number of rounds, the inputs of different rounds of iteration may be different, as explained below:
[0105] In this embodiment of the application, for the first iteration in an iteration with a set number of rounds, the iteration process includes:
[0106] The first complex structure and the fusion encoding feature are input into the decoding network for decoding to obtain the complex structure of the receptor and the ligand binding updated in the first round of iteration. Specifically, the first complex structure and the fusion encoding feature are input into the decoding network for decoding, and the decoding network predicts the predicted positions of the atoms included in the first complex structure based on the input feature, that is, the first complex structure is denoised by predicting the positions of the atoms to correct the first complex structure to obtain the updated complex result.
[0107] For any iteration process after the first round of iteration, which is a non-first round of iteration, the method comprises:
[0108] The complex structure of the receptor and the ligand binding updated in the previous round of iteration is obtained, and the complex structure of the receptor and the ligand binding updated in the previous round of iteration and the fusion encoding feature are input into the decoding network for decoding to obtain the complex structure of the receptor and the ligand binding updated in the current round of iteration. Through the non-first round of iteration, further noise removal is performed to correct the first complex structure to obtain the updated complex result, so that the prediction result is closer to the true result.
[0109] In the embodiment, the number of rounds is not limited.
[0110] It should be noted that, Figure 7 The confidence module in the step 601 is configured to score the plurality of complex structures obtained in each round of prediction to determine the complex structure with the highest score as the prediction result of each round.
[0111] In step 607, the complex structure of the receptor and the ligand binding updated in the last round of iteration is taken as the target complex structure of the receptor and the ligand binding.
[0112] In the embodiment, the structure prediction model is used for multiple rounds of iteration to gradually denoise and continuously correct the complex structure to complete the final prediction of the complex structure. The initial complex structure of the complex is determined by the orientation information of the ligand, which not only carries the binding position information of the ligand, but also carries more detailed information such as the shape, extension direction and area of the ligand, so that an accurate complex structure can be predicted in the prediction process.
[0113] The prediction method of the complex structure of the embodiment of the application realizes step-by-step noise reduction by multiple rounds of iteration of the structure prediction model, continuously corrects the complex structure, and finally predicts the complex structure. The initial complex structure of the complex is determined by providing the orientation information of the ligand, which not only carries the binding position information of the ligand, but also carries more detailed information such as the shape, extension direction and area of the ligand, so that an accurate complex structure can be predicted in the prediction process.
[0114] Based on the above embodiment, Figure 8 A flowchart of a model training method provided by the embodiment of the application is shown in Figure 8 The method comprises the following steps:
[0115] Step 801, obtaining a training sample, wherein the training sample comprises the amino acid sequence of the receptor and the sequence information of the ligand to be combined, and the orientation information of the ligand relative to the protein structure of the receptor.
[0116] As an implementation manner, the orientation information of the ligand relative to the protein structure of the receptor can be based on interactive addition, and specific implementation can refer to the related explanations in the foregoing embodiments, and the principle is the same and will not be repeated here.
[0117] As another implementation manner, the orientation information of the ligand relative to the protein structure of the receptor in the training sample can also be obtained by locally transforming the structure of the ligand in the complex structure of the ligand and the receptor that already exists, for example, by position offset determination.
[0118] In the training process, the application can use multiple samples for training, and the training samples can include the two samples described above to increase the form of the sample and improve the efficiency of sample generation.
[0119] Step 802, inputting the training sample into a structure prediction model to obtain a predicted complex structure of the receptor and the ligand.
[0120] The principle of the model structure and prediction can refer to the related explanations in the foregoing embodiments, and the principle is the same and will not be repeated here.
[0121] Step 803, determining a loss function of the structure prediction model according to the difference between the predicted complex structure and the target complex structure labeled by the training sample.
[0122] In an implementation manner of the embodiment of the application, the difference between the predicted complex structure and the target complex structure can be determined according to the difference between the predicted position of each atom in the predicted complex structure and the real position of each atom in the target complex structure.
[0123] At step 804, the structure prediction model is trained according to the loss function, to obtain a trained structure prediction model.
[0124] In an embodiment of the present application, the parameters of the structure prediction model are adjusted according to the loss function, to obtain the trained structure prediction model.
[0125] It should be noted that the foregoing steps 801 to 804 need to be repeatedly executed multiple times, and different training samples can be used each time, so that the training is stopped when the loss function is less than a threshold, or the training is stopped when the number of repeated executions is greater than a threshold. The structure prediction model obtained after the last model parameter adjustment is used as the trained structure prediction model.
[0126] In the model training method of the embodiment of the present application, compared with the method in the related art which only considers sequence information, the present application can better distinguish different structures of the same sequence complex in nature by providing the orientation information of the ligand relative to the receptor when the ligand and the receptor are combined, and through model training, the trained model can distinguish substances with the same sequence in nature, thereby improving the prediction effect and accuracy of the complex structure.
[0127] To implement the above-mentioned embodiments, an embodiment of the present application further provides a complex structure prediction device. Figure 9 A structural schematic diagram of a complex structure prediction device provided by an embodiment of the present application is shown.
[0128] As shown in the figure, the device comprises: Figure 9 The acquisition module 91 is configured to acquire the amino acid sequence of the receptor, the sequence information of the ligand to be combined, and the protein structure of the receptor.
[0129] The determination module 92 is configured to determine the orientation information of the ligand relative to the protein structure based on the protein structure of the receptor.
[0130] The prediction module 93 is configured to determine the target complex structure of the receptor and the ligand combined according to the orientation information of the ligand relative to the protein structure, the amino acid sequence of the receptor, and the sequence information of the ligand.
[0131] Further, in an implementation manner of the embodiment of the present application, the prediction module 93 is further configured to:
[0132] determine the initial complex structure of the receptor and the ligand combined according to the orientation information of the ligand relative to the protein structure;
[0133]
[0134] According to the initial complex structure, the amino acid sequence of the receptor, and sequence information of the ligand, a target complex structure in which the receptor and the ligand are combined is determined.
[0135] In an implementation form of the determining module 92, the determining module 92 is further configured to:
[0136] display the protein structure on an interactive interface of the electronic device;
[0137] determine a binding position on the protein structure according to the detected first instruction;
[0138] determine a position region of the ligand in space and direction information of the position region according to the detected second instruction and the binding position.
[0139] In an implementation form of the predicting module 93, the predicting module 93 is further configured to:
[0140] determine initial three-dimensional positions of atoms of the ligand in the position region according to the position region and the direction information of the position region;
[0141] obtain the initial complex structure according to the protein structure of the receptor and the initial three-dimensional positions of the atoms of the ligand.
[0142] In an implementation form of the predicting module 93, the predicting module 93 is further configured to:
[0143] perform local transformation on the initial complex structure to obtain a first complex structure;
[0144] determine a target complex structure in which the receptor and the ligand are combined according to the first complex structure, the amino acid sequence of the receptor, and sequence information of the ligand by using a structure prediction model.
[0145] In an implementation form of the predicting module 93, the predicting module 93 is further configured to:
[0146] encode the amino acid sequence of the receptor and sequence information of the ligand to be combined by using an encoding network in the structure prediction model to obtain fused encoding features;
[0147] perform iteration for a set number of rounds according to the first complex structure and the fused encoding features by using a decoding network in the structure prediction model to update a complex structure in which the receptor and the ligand are combined;
[0148] use the complex structure in which the receptor and the ligand are combined obtained by the last round of iteration as a target complex structure in which the receptor and the ligand are combined.
[0149] In an implementation form of the method according to the application, the first iteration of the iterations of the set number of iterations comprises:
[0150] inputting the first complex structure and the fusion encoding feature into the decoding network for decoding to obtain the complex structure of the receptor and the ligand binding updated in the first iteration.
[0151] In an implementation form of the method according to the application, the non-first iteration of the iterations of the set number of iterations comprises:
[0152] obtaining the complex structure of the receptor and the ligand binding updated in the previous iteration of the current iteration;
[0153] inputting the complex structure of the receptor and the ligand binding updated in the previous iteration of the current iteration and the fusion encoding feature into the decoding network for decoding to obtain the complex structure of the receptor and the ligand binding updated in the current iteration.
[0154] It should be noted that the explanation of the foregoing method embodiments is also applicable to the device of this embodiment, and therefore will not be described here again.
[0155] In the complex structure prediction device according to the application, the amino acid sequence of the receptor and the sequence information of the ligand to be combined, and the protein structure of the receptor are obtained, the orientation information of the ligand relative to the protein structure is determined based on the protein structure of the receptor, and the target complex structure of the receptor and the ligand binding is determined according to the orientation information of the ligand relative to the protein structure, the amino acid sequence of the receptor and the sequence information of the ligand. By providing the orientation information of the ligand relative to the receptor when the ligand and the receptor are combined, the initial spatial positional relationship of the ligand and the receptor binding is indicated, and based on the orientation information and the sequence information, the target complex structure with different structures can be better distinguished in nature with the same sequence, and the accuracy of the target complex structure finally predicted is improved.
[0156] In order to implement the foregoing embodiments, the model training device according to the application is also provided.
[0157] Figure 10 A structural schematic diagram of the model training device according to the application is provided.
[0158] As shown in Figure 10 The device comprises:
[0159] The obtaining module 11 is configured to obtain a training sample, wherein the training sample comprises an amino acid sequence of a receptor and sequence information of a ligand to be combined, and orientation information of the ligand relative to a protein structure of the receptor.
[0160] A prediction module 12 is configured to input the training sample into a structure prediction model to obtain a predicted complex structure of the receptor and the ligand.
[0161] A determination module 13 is configured to determine a loss function of the structure prediction model according to a difference between the predicted complex structure and a target complex structure labeled by the training sample.
[0162] A training module 14 is configured to train the structure prediction model according to the loss function to obtain a trained structure prediction model.
[0163] It should be noted that the above-mentioned method embodiments are also applicable to the device embodiments, and thus will not be described here again.
[0164] In the model training device provided in the embodiments of the present application, the orientation information of the ligand relative to the receptor when the ligand and the receptor are combined can be provided, so that different structures of the same sequence complex in nature can be better distinguished, and through model training, the trained model can distinguish substances with the same sequence in nature, thereby improving the prediction effect and accuracy of the complex structure.
[0165] According to the embodiments of the present application, the present application also provides an electronic device, a readable storage medium and a computer program product.
[0166] Figure 11 A structural schematic diagram of an electronic device provided in the embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are merely examples, and are not intended to limit the implementations described and / or claimed in this document.
[0167] As Figure 11As shown, the device 700 includes a computing unit 701 that can perform various appropriate actions and processes in accordance with a computer program stored in a ROM (Read-Only Memory) 702 or a computer program loaded into a RAM (Random Access Memory) 703 from the storage unit 708. Various programs and data required for the operation of the device 700 can also be stored in the RAM 703. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. An I / O (Input / Output) interface 705 is also connected to the bus 704.
[0168] A plurality of components in the device 700 are connected to the I / O interface 705, including an input unit 706 such as a keyboard, a mouse, and the like; an output unit 707 such as various types of displays, speakers, and the like; a storage unit 708 such as a magnetic disk, an optical disk, and the like; and a communication unit 709 such as a network card, a modem, a wireless communication transceiver, and the like. The communication unit 709 allows the device 700 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0169] The computing unit 701 can be various general and / or special-purpose processing components having processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), various special-purpose AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, a DSP (Digital Signal Processor), and any appropriate processor, controller, microcontroller, and the like. The computing unit 701 performs various methods and processes described above. For example, in some embodiments, the above-described methods can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the methods described above can be performed. Alternatively, in other embodiments, the computing unit 701 can be configured to perform the above-described methods by any other appropriate means (e.g., by means of firmware).
[0170] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a Field Programmable Gate Array (FPGA), an Application-Specific Integrated Circuit (ASIC), an Application Specific Standard Product (ASSP), a System on a Chip (SOC), a Complex Programmable Logic Device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0171] Program code for carrying out methods of the present application can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine and partially on a remote machine or entirely on a remote machine or server.
[0172] In the context of this application, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include a linearly-programmed electronic storage, a portable computer diskette, a hard disk, a RAM, a ROM, an EPROM (Electrically Programmable Read-Only-Memory), or a flash memory, an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0173] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0174] The systems and techniques described here can be implemented in a computing system that includes a back-end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front-end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a LAN (Local Area Network), a WAN (Wide Area Network), the Internet, and a blockchain network.
[0175] The computer system can include clients and servers. This relationship can be between a client and a server that are typically remote from each other and typically interact through a communication network. The relationship between client and server exists by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service (Virtual Private Server, virtual private server).
[0176] According to the embodiments of the present application, the present application also provides a computer program product, when the instruction processor in the computer program product executes, executes the method proposed in the above embodiments of the present application.
[0177] It should be understood that the steps shown above can be reordered, added or deleted using various forms of flow. For example, the steps described in the present application can be executed in parallel, sequentially or in different order, as long as the desired results of the technical solutions disclosed in the present application can be achieved, which is not limited herein.
[0178] The above detailed description does not constitute a limitation on the scope of protection of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A method of predicting a composite structure, characterized by, The method comprises: obtaining sequence information of a receptor and a ligand to be combined, and a protein structure of the receptor; determining orientation information of the ligand relative to the protein structure based on the protein structure of the receptor; determining a target complex structure of the receptor and the ligand combined according to the orientation information of the ligand relative to the protein structure, the sequence information of the receptor and the sequence information of the ligand; the method comprises: displaying the protein structure on an interactive interface of an electronic device; determining a binding position on the protein structure according to a detected first instruction, wherein the first instruction is generated through an interactive process; determining a position region of the ligand in space and direction information of the position region according to a detected second instruction and the binding position, wherein the second instruction is used to start from the binding position to draw and determine the position region of the ligand and the direction information of the position region in the space around the protein structure; the method comprises: determining initial three-dimensional positions of each atom of the ligand in the position region according to the position region and the direction information of the position region; obtaining an initial complex structure according to the protein structure of the receptor and the initial three-dimensional positions of each atom of the ligand; determining a target complex structure of the receptor and the ligand combined according to the initial complex structure, the sequence information of the receptor and the sequence information of the ligand, wherein the method comprises: performing local transformation on the initial complex structure to obtain a first complex structure, and the local transformation comprises: performing position offset on a target atom, or twisting or rotating the initial complex structure; using a structure prediction model to determine a target complex structure of the receptor and the ligand combined according to the first complex structure, the sequence information of the receptor and the sequence information of the ligand.
2. The method of claim 1, wherein, the method comprises: using an encoding network in the structure prediction model to encode the sequence information of the receptor and the ligand to be combined to obtain fused encoding features; using a decoding network in the structure prediction model to perform a set number of iterations according to the first complex structure and the fused encoding features to update the complex structure of the receptor and the ligand combined; using the complex structure of the receptor and the ligand combined obtained through the last iteration as the target complex structure of the receptor and the ligand combined.
3. The method of claim 2, wherein, the first iteration in the set number of iterations comprises: inputting the first complex structure and the fusion encoding feature into the decoding network for decoding to obtain the complex structure of the receptor and the ligand binding updated in the first round of iteration.
4. The method of claim 2, wherein, The non-first iteration in the set number of iterations includes: obtaining the complex structure of the receptor and the ligand binding updated in the previous iteration of the current iteration; inputting the complex structure of the receptor and the ligand binding updated in the previous iteration of the current iteration and the fusion encoding feature into the decoding network for decoding to obtain the complex structure of the receptor and the ligand binding updated in the current iteration.
5. A model training method, comprising: comprise: obtaining a training sample, wherein the training sample comprises amino acid sequence of a receptor and sequence information of a ligand to be combined, and orientation information of the ligand relative to a protein structure of the receptor, wherein it comprises: displaying the protein structure on an interactive interface of an electronic device; determining a binding position on the protein structure according to a detected first instruction, wherein the first instruction is generated through an interactive process; determining a position region of the ligand in space and direction information of the position region according to a detected second instruction and the binding position, wherein the second instruction is used to start from the binding position to draw and determine the position region of the ligand and the direction information of the position region in the space around the protein structure; inputting the training sample into a structure prediction model to obtain a predicted complex structure of the receptor and the ligand, wherein it comprises: determining initial three-dimensional positions of atoms of the ligand in the position region according to the position region and the direction information of the position region; obtaining an initial complex structure according to the protein structure of the receptor and the initial three-dimensional positions of the atoms of the ligand, and performing local transformation on the initial complex structure to obtain a first complex structure, wherein the local transformation comprises: position offset of target atoms, twisting or rotating of the initial complex structure; inputting the first complex structure, the amino acid sequence of the receptor and the sequence information of the ligand into the structure prediction model; determining a loss function of the structure prediction model according to the difference between the predicted complex structure and a target complex structure labeled by the training sample; training the structure prediction model according to the loss function to obtain a trained structure prediction model.
6. A device for predicting a composite structure, characterized by comprise: an obtaining module, configured to obtain amino acid sequence of a receptor and sequence information of a ligand to be combined, and a protein structure of the receptor; a determining module, configured to determine orientation information of the ligand relative to the protein structure based on the protein structure of the receptor; a prediction module, configured to determine a target complex structure of the receptor and the ligand binding according to the orientation information of the ligand relative to the protein structure, the amino acid sequence of the receptor and the sequence information of the ligand; the determining module is further configured to: display the protein structure on an interactive interface of an electronic device; According to the detected first instruction, a binding position is determined on the protein structure, wherein the first instruction is generated through an interaction process; According to the detected second instruction and the binding position, a position region of the ligand in space and direction information of the position region are determined, wherein the second instruction is used to draw from the binding position as a starting point to determine the position region of the ligand and the direction information of the position region in the space around the protein structure; The prediction module is further configured to: According to the position region and the direction information of the position region, initial three-dimensional positions of atoms of the ligand in the position region are determined; According to the protein structure of the receptor and the initial three-dimensional positions of the atoms of the ligand, an initial complex structure is obtained; A first complex structure is obtained by performing a local transformation on the initial complex structure, wherein the local transformation includes performing a position offset on a target atom, or twisting or rotating the initial complex structure; A target complex structure of the receptor and the ligand is determined by using a structure prediction model according to the first complex structure, an amino acid sequence of the receptor, and sequence information of the ligand.
7. The apparatus of claim 6, wherein, The prediction module is further configured to: A fusion encoding feature is obtained by using an encoding network in the structure prediction model to encode the amino acid sequence of the receptor and the sequence information of the ligand to be combined; A complex structure of the receptor and the ligand is updated by performing a preset number of iterations according to the first complex structure and the fusion encoding feature by using a decoding network in the structure prediction model; The complex structure of the receptor and the ligand obtained by the last iteration is taken as a target complex structure of the receptor and the ligand.
8. The apparatus of claim 7, wherein, The first iteration in the preset number of iterations includes: The first complex structure and the fusion encoding feature are input into the decoding network for decoding to obtain the complex structure of the receptor and the ligand updated by the first iteration.
9. The apparatus of claim 7, wherein, The non-first iteration in the preset number of iterations includes: The complex structure of the receptor and the ligand updated by the previous iteration is obtained; The complex structure of the receptor and the ligand updated by the current iteration is obtained by inputting the complex structure of the receptor and the ligand updated by the previous iteration and the fusion encoding feature into the decoding network for decoding.
10. A model training apparatus, comprising: The prediction module includes: The obtaining module is configured to obtain a training sample, wherein the training sample includes an amino acid sequence of a receptor, sequence information of a ligand to be combined, and orientation information of the ligand relative to a protein structure of the receptor, and the prediction module includes: The protein structure is displayed on an interaction interface of an electronic device; A binding position is determined on the protein structure according to a detected first instruction; A position region of the ligand in space and direction information of the position region are determined according to a detected second instruction and the binding position; The prediction module is further configured to: According to the position region and the direction information of the position region, initial three-dimensional positions of atoms of the ligand in the position region are determined; According to the protein structure of the receptor and the initial three-dimensional positions of the atoms of the ligand, an initial complex structure is obtained; A first complex structure is obtained by performing a local transformation on the initial complex structure, wherein the local transformation includes performing a position offset on a target atom, or twisting or rotating the initial complex structure; A target complex structure of the receptor and the ligand is determined by using a structure prediction model according to the first complex structure, an amino acid sequence of the receptor, and sequence information of the ligand. The prediction module is further configured to: A fusion encoding feature is obtained by using an encoding network in the structure prediction model to encode the amino acid sequence of the receptor and the sequence information of the ligand to be combined; A complex structure of the receptor and the ligand is updated by performing a preset number of iterations according to the first complex structure and the fusion encoding feature by using a decoding network in the structure prediction model; The complex structure of the receptor and the ligand obtained by the last iteration is taken as a target complex structure of the receptor and the ligand. The first iteration in the preset number of iterations includes: The first complex structure and the fusion encoding feature are input into the decoding network for decoding to obtain the complex structure of the receptor and the ligand updated by the first iteration. The non-first iteration in the preset number of iterations includes: The complex structure of the receptor and the ligand updated by the previous iteration is obtained; The complex structure of the receptor and the ligand updated by the current iteration is obtained by inputting the complex structure of the receptor and the ligand updated by the previous iteration and the fusion encoding feature into the decoding network for decoding. The prediction module includes: The obtaining module is configured to obtain a training sample, wherein the training sample includes an amino acid sequence of a receptor, sequence information of a ligand to be combined, and orientation information of the ligand relative to a protein structure of the receptor, and the prediction module includes: The protein structure is displayed on an interaction interface of an electronic device; A binding position is determined on the protein structure according to a detected first instruction; A position region of the ligand in space and direction information of the position region are determined according to a detected second instruction and the binding position; The prediction module is configured to input the training sample into a structure prediction model to obtain a predicted complex structure of the receptor and the ligand, and includes: determining initial three-dimensional positions of atoms of the ligand in the position region according to the position region and direction information of the position region; obtaining an initial complex structure according to the protein structure of the receptor and the initial three-dimensional positions of the atoms of the ligand, and performing local transformation on the initial complex structure to obtain a first complex structure, the local transformation including: performing position offset on a target atom, or twisting or rotating the initial complex structure; and inputting the first complex structure, the amino acid sequence of the receptor, and sequence information of the ligand into the structure prediction model. The determination module is configured to determine a loss function of the structure prediction model according to a difference between the predicted complex structure and a target complex structure labeled by the training sample. The training module is configured to train the structure prediction model according to the loss function to obtain the trained structure prediction model.
11. An electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-4, or perform the method of claim 5.
12. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to perform the method of any one of claims 1-4, or perform the method of claim 5.
13. A computer program product comprising a computer program which, when executed by a processor, implements the steps of the method of any one of claims 1-4, or implements the steps of the method of claim 5.
Citation Information
Patent Citations
Composite structure acquisition method and device, equipment, storage medium and program product
CN116959558A
Generation method and device of composite molecular structure, electronic equipment and storage medium
CN116978448A
Protein-ligand interaction prediction method and optimization method
CN117275610A