Method, device and computer equipment for evaluating ligand binding sites of a complex
By using deep learning to evaluate the binding sites of ligands in a complex and using a neural network model to calculate the binding strength, the problem of insufficient accuracy of traditional scoring functions is solved, and higher precision ligand binding site evaluation and drug design are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI ZELIXIR BIOTECH CO LTD
- Filing Date
- 2022-01-26
- Publication Date
- 2026-04-21
AI Technical Summary
Traditional scoring functions are not very accurate in evaluating ligand docking sites in complexes, resulting in inaccurate evaluations.
By employing deep learning methods, the distance information between protein and ligand atoms in the candidate complex conformation is obtained by binding the target ligand to multiple candidate sites of the target protein. The binding strength is then evaluated using a neural network model to determine the optimal binding site.
It improves the accuracy of ligand binding site assessment in complexes, reduces the false positive rate in molecular docking, and enhances the precision of drug design.
Smart Images

Figure CN115910234B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of biotechnology, specifically to a method, apparatus, and computer device for evaluating the ligand binding position of a complex. Background Technology
[0002] With the continuous development and popularization of computer technology, using computer simulation to assist drug development has become a very effective means in the drug development process, and technology-assisted drug development can save a lot of time and money.
[0003] In drug development, it is often necessary to evaluate the docking sites between ligands and proteins in a complex to screen for high-quality complexes. Currently, when evaluating suitable docking sites for ligands in proteins, traditional scoring functions are often used to score the complexes obtained after the ligands dock at different sites. The appropriate docking site for the ligand in the protein is then determined based on the scoring results obtained from the scoring function.
[0004] The traditional scoring function is not very accurate, which leads to inaccurate assessment of ligand docking positions in the complex. Summary of the Invention
[0005] This application provides a method, apparatus, and computer device for evaluating the ligand binding position of a complex. This method can effectively improve the accuracy of evaluating the ligand binding position of a complex.
[0006] The first aspect of this application provides a method for evaluating the ligand binding site of a complex, the method comprising:
[0007] The target ligands are bound to multiple candidate binding sites in the target protein to obtain the candidate complex conformation corresponding to each candidate binding site.
[0008] Obtain distance information between each atom of the target protein and each atom of the target ligand in each candidate complex conformation;
[0009] Based on the distance information, determine the binding strength information of the target protein and the target ligand in each candidate complex conformation;
[0010] The target bonding location is determined based on the bonding strength information.
[0011] Accordingly, a second aspect of this application provides a device for evaluating the ligand binding site of a complex, the device comprising:
[0012] The binding unit is used to bind the target ligand to multiple candidate binding sites in the target protein, thereby obtaining the candidate complex conformation corresponding to each candidate binding site.
[0013] The acquisition unit is used to acquire distance information between each atom of the target protein and each atom of the target ligand in each candidate complex conformation;
[0014] The first determining unit is configured to determine the binding strength information of the target protein and the target ligand in each candidate complex conformation based on the distance information.
[0015] The second determining unit is used to determine the target bonding position based on the bonding strength information.
[0016] A third aspect of this application also provides a computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to perform steps in the method for evaluating the ligand binding position of a complex provided in the first aspect of this application.
[0017] A fourth aspect of this application provides a computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps in the method for evaluating the ligand binding position of a complex provided in the first aspect of this application.
[0018] The fifth aspect of this application provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps in the ligand binding site evaluation method for a complex provided in the first aspect.
[0019] The ligand binding site evaluation method for complexes provided in this application involves binding the target ligand to multiple candidate binding sites in the target protein to obtain a candidate complex conformation corresponding to each candidate binding site; obtaining distance information between each atom of the target protein and each atom of the target ligand in each candidate complex conformation; determining the binding strength information between the target protein and the target ligand in each candidate complex conformation based on the distance information; and determining the target binding site based on the binding strength information.
[0020] Therefore, the ligand binding position evaluation method for complexes provided in this application can calculate the binding strength between the protein and ligand in the complex by analyzing the distance information between protein atoms and ligand atoms in the complex. Then, it determines the optimal binding position of the ligand based on the different binding strengths corresponding to different binding positions of the ligand. In other words, this method can accurately calculate the binding strength corresponding to different complex conformations obtained after the ligand binds to different binding positions of the protein, and then determine the optimal binding position based on the binding strength, thereby improving the accuracy of ligand binding position evaluation in complexes. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a schematic diagram of a scenario for evaluating the ligand binding site of the complex in this application;
[0023] Figure 2 This is a flowchart illustrating the method for evaluating the ligand binding site of the complex provided in this application;
[0024] Figure 3 This is a graph showing the test results of the ligand binding strength assessment method in this application in a specific database;
[0025] Figure 4 The graph shows the test results of the conformation optimization method provided in this application in a specific database;
[0026] Figure 5 This is another test result diagram of the conformation optimization method provided in this application in a specific database;
[0027] Figure 6 This is a schematic diagram of the structure of the ligand binding site evaluation device for the complex provided in this application;
[0028] Figure 7 This is a schematic diagram of the structure of the computer device provided in this application. Detailed Implementation
[0029] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0030] This invention provides a method, apparatus, and computer device for evaluating the ligand binding site of a complex. The method for evaluating the ligand binding site of a complex can be used in the apparatus for evaluating the ligand binding site of a complex. The apparatus can be integrated into a computer device, which can be a terminal or a server. The terminal can be a mobile phone, tablet computer, laptop computer, smart TV, wearable smart device, personal computer (PC), or vehicle terminal, etc. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The server can also be a node in a blockchain.
[0031] Please see Figure 1 This is a schematic diagram of a scenario illustrating the ligand binding site evaluation method for the complex provided in this application. As shown, server A receives an evaluation request from terminal B. The evaluation request includes the target protein and target ligand contained in the complex to be evaluated, as well as an evaluation instruction to evaluate the ligand binding site in the complex. Then, server A binds the target ligand to multiple candidate binding sites in the target protein, obtaining a candidate complex conformation corresponding to each candidate binding site; obtains the distance information between each atom of the target protein and each atom of the target ligand in each candidate complex conformation; determines the binding strength information between the target protein and the target ligand in each candidate complex conformation based on the distance information; and determines the target binding site based on the binding strength information. Then, the target binding site is returned to terminal B as the evaluation result.
[0032] It should be noted that, Figure 1 The schematic diagram illustrating the ligand binding site evaluation scenario of the complex is merely an example. The ligand binding site evaluation scenario of the complex described in the embodiments of this application is intended to more clearly illustrate the technical solution of this application and does not constitute a limitation on the technical solution provided in this application. Those skilled in the art will understand that as the ligand binding site evaluation scenario of the complex evolves and new business scenarios emerge, the technical solution provided in this application is equally applicable to similar technical problems.
[0033] The implementation scenarios described above will be explained in detail below.
[0034] In related technologies, when evaluating ligand binding sites in complexes, scoring functions are generally used to score the complex conformations obtained by ligands binding to different binding sites. However, traditional scoring functions are mostly based on experience, resulting in low accuracy and limited generalization ability. Therefore, this application provides a method for evaluating ligand binding sites in complexes, which can improve the accuracy of ligand binding site evaluation to a certain extent.
[0035] This application will describe the embodiments from the perspective of a ligand binding site evaluation device for a complex, which can be integrated into a computer device. The computer device can be a terminal or a server. The terminal can be a mobile phone, tablet computer, laptop computer, smart TV, wearable smart device, personal computer (PC), or in-vehicle terminal, etc. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery network (CDN) acceleration services, and big data and artificial intelligence platforms. Figure 2 The diagram shown is a flowchart illustrating the method for evaluating the ligand binding site of the complex provided in this application. The method includes:
[0036] Step 101: The target ligand is bound to multiple candidate binding sites in the target protein to obtain the candidate complex conformation corresponding to each candidate binding site.
[0037] Here, the target ligand and target protein refer to the ligand and protein that need to bind to form a complex. The binding of the ligand to the protein can also be referred to as attaching the ligand to the protein. Specifically, the target protein and target ligand can be the protein and ligand included in the binding site evaluation command sent from other terminals.
[0038] In virtual drug screening, it is often necessary to ligand-bind to proteins to determine the drug's efficacy against viruses. One method to determine whether a ligand (small molecule or short peptide) can effectively bind to the pathogenic protein is to calculate the binding strength between the protein and the ligand; a stronger binding strength indicates a tighter bond between the two. However, the binding strength is significantly affected by the ligand's binding site; the same ligand will result in different binding strengths at different binding sites. Therefore, the evaluation and optimization of binding sites are crucial in drug design.
[0039] The docking of ligands into the active pocket of proteins mainly involves two steps: First, sampling, which searches for multiple binding sites within the active pocket. Second, scoring, which uses a scoring function to predict the binding energy between the ligand and protein at different binding sites. Here, binding energy refers to the binding potential energy between the protein and ligand, i.e., the binding strength. Most current molecular docking software uses traditional scoring functions; however, these functions lack precision, resulting in poor accuracy in assessing ligand binding sites during docking and consequently a high false positive rate. Traditional scoring functions are generally classified into three categories: physics-based, experience-based, and knowledge-based. Traditional methods have low accuracy in predicting protein-ligand interactions and limited generalization ability. In recent years, applying machine learning methods to predict protein-ligand interactions has become a trend, especially with the introduction of deep learning. Algorithms such as convolutional neural networks and graph neural networks can significantly improve the efficiency and accuracy of predicting protein-ligand interactions. This application proposes a method for evaluating the ligand binding position of a complex based on deep learning, which aims to improve the accuracy of evaluating the ligand binding position in the complex.
[0040] Once the target protein and ligand for docking are identified, the first step is to determine multiple possible binding sites within the target protein. Specifically, molecular docking software can be used to identify several possible binding sites of the ligand within the protein's active pocket; these sites can be referred to as candidate binding sites. AutoDock Vina can be used as the molecular docking software; it is an open-source molecular simulation software primarily used for performing ligand-protein docking.
[0041] After identifying multiple possible candidate binding sites for the ligand in the protein's active pocket, the target ligand can be individually paired with these candidate binding sites to obtain the candidate complex conformation corresponding to each site. This pairing of the target ligand with these candidate sites can be achieved through computer-simulated docking, yielding the candidate complex conformation for each site.
[0042] Step 102: Obtain the distance information between each atom of the target protein and each atom of the target ligand in each candidate complex conformation.
[0043] After the candidate complex conformation corresponding to each candidate binding site is determined by simulation, the distance information between the atoms of the target protein and the atoms of the target ligand in each candidate complex conformation can be further calculated.
[0044] Proteins contain amino acids and heavy atoms, with the heavy atoms being non-hydrogen atoms. Specifically, the amino acid types include twenty basic amino acids and one additional group. The twenty basic amino acids are: glycine, alanine, valine, leucine, isoleucine, methionine, proline, tryptophan, serine, tyrosine, cysteine, phenylalanine, asparagine, glutamine, threonine, aspartic acid, glutamic acid, lysine, arginine, and histidine. The additional group of amino acids can be defined as other amino acids (OTH). Amino acids contain carbon, nitrogen, hydrogen, and oxygen atoms. The heavy atom types in proteins include carbon (C), nitrogen (N), oxygen (O), sulfur (S), and the additional atom type (DU).
[0045] The atoms in the target ligand can be classified into seven types according to their element type: carbon, nitrogen, oxygen, sulfur, phosphorus, halogens, and others. Specifically, halogens can be any one of fluorine, chlorine, bromine, and iodine.
[0046] Calculating the distance information between the atoms of the target protein and the atoms of the target ligand in each candidate complex conformation can be done by calculating the distance information between each atom of the target protein and each atom of the target ligand, or by calculating the distance information between the atoms of the amino acid and heavy atom combination type in the above protein and the ligand atoms.
[0047] To calculate the distance information between atoms, each candidate complex conformation can be arranged in a predefined three-dimensional coordinate system, and the coordinate information of each atom can be obtained. Then, the distance information between atoms can be calculated based on the coordinate information.
[0048] Step 103: Determine the binding strength information of the target protein and the target ligand in each candidate complex conformation based on the distance information.
[0049] In this process, after calculating the distance information between each atom of the target protein and each atom of the target ligand in each candidate complex conformation, the binding strength information between the target protein and the target ligand in each candidate complex conformation can be evaluated based on the distance information.
[0050] In some embodiments, determining the binding strength information of the target protein and the target ligand in each candidate complex conformation based on distance information includes:
[0051] 1. Generate ligand binding strength assessment features for each candidate complex conformation based on distance information;
[0052] 2. Input the ligand binding strength evaluation features corresponding to each candidate complex conformation into the preset neural network model to obtain the binding strength information of the target protein and the target ligand in each candidate complex conformation.
[0053] In this embodiment, a deep learning method can be used to evaluate the binding strength of the target protein and the target ligand in each candidate complex conformation through a neural network model. This neural network model can be a pre-defined neural network model used to output corresponding ligand binding strength information based on the input ligand binding strength evaluation features.
[0054] Specifically, the ligand binding strength features for each candidate complex can be generated first based on the distance information between protein atoms and ligand atoms in each candidate complex. Then, the ligand binding strength features for each candidate complex are input into the aforementioned trained preset neural network model to evaluate the ligand binding strength, thereby obtaining the ligand binding strength information for each candidate complex.
[0055] In some embodiments, generating ligand binding strength evaluation features for each candidate complex conformation based on distance information includes:
[0056] 1.1 Identify a first number of first target atoms in the target protein based on different combinations of amino acids and heavy atoms;
[0057] 1.2. Identify a second number of second target atoms in the target ligand;
[0058] 1.3 Determine the target distance information between each first target atom and each second target atom from the distance information;
[0059] 1.4. Generate ligand binding strength assessment features for each candidate complex conformation based on the target distance information.
[0060] In this embodiment, the ligand binding strength characteristics of each candidate complex can be determined based on the distance information between different types of protein heavy atoms and ligand atoms. Specifically, as mentioned earlier, there are 21 types of amino acids and 5 types of heavy atoms in proteins. Therefore, the types of atoms in the protein can be determined by different combinations of amino acids and heavy atoms, resulting in 21*5=105 types of protein atoms, which can be referred to as the first target atoms. When the same heavy atom binds to different amino acids, it can be identified as a different protein atom. Furthermore, as mentioned earlier, there are 7 types of ligand atoms, which can be referred to as the second target atoms. Therefore, there are 105*7=735 ways for the atomic pairs between the protein and the ligands to bind. For each atomic binding method, the distance between the corresponding atomic pairs can be determined based on the distance information between all protein atoms and all ligand atoms calculated above. Then, a feature value is calculated based on the distance between the atomic pairs. These feature values form a 735-dimensional feature vector, which can be used as the ligand binding strength evaluation feature in this application.
[0061] In some embodiments, generating ligand binding strength evaluation features for each candidate complex conformation based on target distance information includes:
[0062] 1.4.1 Calculate the first feature dimension of the target distance information to obtain the first sub-ligand binding strength evaluation feature corresponding to each candidate complex conformation;
[0063] 1.4.2 Calculate the second feature dimension of the target distance information to obtain the second ligand binding strength evaluation feature corresponding to each candidate complex conformation;
[0064] 1.4.3. For each candidate complex conformation, the first sub-ligand binding strength assessment feature and the second sub-ligand binding strength assessment feature are spliced together to obtain the ligand binding strength assessment feature corresponding to each candidate complex conformation.
[0065] In this embodiment, the characteristic value corresponding to each type of atom pair is calculated, specifically the interaction force between the two atoms in the atom pair. The interaction force between atoms in an atom pair can include van der Waals interactions and electrostatic interactions. In this embodiment, the negative exponent of the atomic distance can be used to characterize the interaction between atom pairs. Since the negative exponents of the distances corresponding to van der Waals interactions and electrostatic interactions are different, the van der Waals potential energy and electrostatic potential energy between the two atoms can be calculated based on the distance information between the atoms for each type of atom pair. These can be referred to as the first sub-ligand set strength evaluation feature and the second sub-ligand binding strength evaluation feature, respectively. Using the two potential energy values corresponding to each atom pair as the parameters of the feature, a 735*2 = 1470-dimensional feature vector corresponding to each candidate complex conformation can be obtained. This feature vector is the ligand binding strength evaluation feature corresponding to the candidate complex conformation.
[0066] The method for calculating the corresponding eigenvalues based on the distance information between atomic pairs in this application can be performed using the following formula (1):
[0067]
[0068] Where RA represents the type of atom in the protein, one of the aforementioned 105 types of protein atoms; L represents the type of atom in the ligand, one of the aforementioned 7 types of ligand atoms; i is the exponent, specifically 1 or 6; and d is the distance between atoms in the atom pair.
[0069] Thus, for each candidate complex conformation, 1470 E values can be calculated, and then the ligand binding strength evaluation features corresponding to the candidate complex conformation can be constructed based on these 1470 values.
[0070] In some embodiments, since excessively small distances between atoms in an atom pair may lead to an explosion of eigenvalues, this application embodiment allows for certain judgments and restrictions on the distances between atoms in an atom pair. The distance d between atoms in an atom pair is restricted to satisfy the following inequality:
[0071] 0.3nm≤d≤2.0nm
[0072] Here, nm represents the length unit nanometer. When the distance between atoms in an atom pair is less than 0.3 nanometers, d is set to 0.3 nanometers; when the distance between atoms in an atom pair is greater than 2 nanometers, the corresponding characteristic value E is set to 0.
[0073] In some embodiments, before inputting the ligand binding strength assessment features corresponding to each candidate complex conformation into a preset neural network model to obtain the binding strength information of the target protein and the target ligand in each candidate complex conformation, the method further includes:
[0074] A. Obtain training samples, which include ligand binding strength evaluation features corresponding to multiple sample complex conformations and binding strength information corresponding to each sample complex conformation.
[0075] B. The pre-set neural network model is trained by taking the ligand binding strength evaluation features corresponding to the conformations of multiple sample complexes as input and the binding strength information corresponding to the conformations of each sample complex as output.
[0076] In this embodiment, when processing the ligand binding strength evaluation features corresponding to each candidate complex conformation using a preset neural network model to obtain the ligand binding strength features, the preset neural network model needs to be a trained neural network model. Therefore, the preset neural network model needs to be trained before being used for ligand binding strength evaluation. The preset neural network model can be a convolutional neural network model or other neural network models; no limitation is imposed here.
[0077] Specifically, this application can employ a supervised training method to train the pre-defined neural network model. This requires first obtaining training samples for training the pre-defined neural network model. The training samples include multiple input features and a label value corresponding to each input feature. Here, the input features can be ligand binding strength evaluation features corresponding to the sample complex conformation, and their corresponding label values can be binding strength information corresponding to the sample complex conformation.
[0078] Thus, when training a preset neural network model using training samples, the ligand binding strength evaluation features corresponding to the sample complex are input into the preset neural network model to obtain its output binding strength information. Then, the model parameters of the preset neural network model are adjusted based on the difference between the output binding strength information and the label value until the preset neural network model converges, resulting in the trained preset neural network model.
[0079] In some embodiments, training samples are obtained, which include ligand binding strength evaluation features corresponding to multiple sample complex conformations and binding strength information corresponding to each sample complex conformation, including:
[0080] A1. Obtain the first number of sample complexes;
[0081] A2. Re-dock the ligands in each sample complex to obtain the second number of sample complex conformations;
[0082] A3. Calculate the binding strength information corresponding to the conformation of each sample complex;
[0083] A4. Obtain the sample distance information between protein atoms and ligand atoms in the conformation of each sample complex;
[0084] A5. Generate ligand binding strength assessment features corresponding to the complex conformation of each sample based on the sample distance information. The ligand binding strength assessment features corresponding to the complex conformation of each sample and the binding strength information corresponding to the complex conformation of each sample constitute training samples.
[0085] In this embodiment, the training samples can be obtained by re-docking and calculating the standard conformation of the complex. Specifically, a first number of sample complexes can be downloaded from the database, wherein the complex conformation corresponding to the sample complex is the standard conformation, that is, the binding strength between the protein and the ligand is the greatest in this conformation.
[0086] Then, molecular docking software is used to re-dock each sample complex, resulting in multiple sample complex conformations for each complex. At this point, the root mean square error (RMSD) between the multiple sample complex conformations and the standard conformation of the sample complex can be calculated, and the RMSD corresponding to each sample complex conformation is used as its corresponding tag value, that is, its binding strength information. It can be understood that the smaller the RMSD value, the smaller the difference between the complex conformation and the standard conformation, indicating a stronger binding strength between the ligand and the protein in the complex conformation, and a more stable state. Then, the RMSD value can be calculated for each sample complex conformation, serving as the tag value for that sample complex conformation.
[0087] Then, for the multiple sample complex conformations obtained from the re-docking, their corresponding binding strength assessment features can be generated. The method for generating the binding strength assessment features of the sample complex conformations here is consistent with the method for generating the binding strength assessment features of the candidate complex conformations described above, and will not be repeated here.
[0088] After generating the binding strength evaluation features and RMSD corresponding to each sample complex conformation, the binding strength evaluation features and RMSD corresponding to each sample complex conformation can be used as training samples to train the preset neural network model.
[0089] Step 104: Determine the target bonding location based on the bonding strength information.
[0090] As mentioned earlier, the input parameters for training the preset neural network model are the binding strength evaluation features of each sample complex conformation, and the label value is the RMSD corresponding to each sample complex conformation. Therefore, after inputting the binding strength features of each candidate complex conformation into the trained preset neural network model, the RMSD value corresponding to each candidate complex conformation will be output.
[0091] Understandably, a smaller RMSD value indicates a smaller difference between the candidate complex conformation and the standard conformation, which in turn indicates a stronger binding strength between the protein and ligand in the candidate complex conformation. Once the RMSD value corresponding to each candidate complex conformation is determined, the optimal binding site in the protein can be identified.
[0092] Among them, determining the target binding location based on binding strength information includes:
[0093] 1. Determine the target complex conformation from the candidate complex conformations based on the binding strength information;
[0094] 2. The binding site corresponding to the conformation of the target complex is determined as the target binding site.
[0095] In the embodiments of this application, after determining the binding strength information corresponding to each candidate complex conformation, that is, after determining the RMSD value corresponding to each candidate complex conformation, the candidate complex conformation with the smallest RMSD value can be determined as the most stable complex conformation, and it is determined as the target complex conformation.
[0096] Furthermore, the target binding site for attaching the target ligand to the target protein can be determined based on the conformation of the target complex; this target binding site is the optimal binding site.
[0097] In some embodiments, the method for evaluating the ligand binding site of the complex provided in this application may further include:
[0098] A. Obtain the spatial characteristics of the target ligand in the conformation of the target complex. The spatial characteristics include variable parameters that characterize the spatial position of each atom in the target ligand.
[0099] B. Spatial features are adjusted based on the binding strength information between proteins and ligands in the conformation of the target complex to obtain the target spatial features;
[0100] C. Determine the spatial position of each atom of the target ligand in the conformation of the target complex based on the spatial characteristics of the target.
[0101] In complexes, the RMSD value corresponding to the complex conformation is influenced not only by the binding position of the ligand on the protein but also by the spatial position information of the atoms within the ligand. However, compared to the influence of the ligand's binding position on the protein, the spatial position information of the atoms within the ligand has a smaller impact on the RMSD value. Therefore, when evaluating the ligand binding position of the complex, the influence of changes in the spatial position of the atoms within the ligand is not considered, and the spatial positions of the atoms within the ligand are assumed to be fixed.
[0102] However, after determining the optimal binding site of the ligand in the protein, the optimal spatial positions of the atoms in the ligand can be further determined, thereby obtaining a smaller RMSD value and a more stable complex conformation. Therefore, this application also provides a conformation optimization method, which can be used to further optimize the conformation of the complex after the ligand binding site is determined.
[0103] Therefore, after determining the optimal binding site for the target ligand to attach to the target protein, the spatial features of the target ligand within the conformation of the target complex can be further obtained. These spatial features include variable parameters characterizing the spatial position of each atom in the target ligand. Specifically, the bond relationships between the atoms of the target ligand in the conformation of the target complex can be extracted, and the control features of the target ligand can be characterized using a 6+k dimension vector. Here, 6 refers to the three-dimensional coordinates of the first atom of the target ligand attached to the target protein, as well as the rotation angles around the X, Y, and Z axes. The three-dimensional coordinates are (x, y, z), and the rotation angles are (α, β, γ). k represents the torsion angles (θ1, θ2, θ3, ..., θ) of the k rotatable bonds within the target ligand. k ).
[0104] In this context, the spatial features of the target ligand include 6+k eigenvalues, each a variable corresponding to a different complex conformation. The three-dimensional coordinates of each atom in the target ligand can then be calculated using these 6+k eigenvalue variables. Further, the ligand binding strength feature corresponding to the target complex conformation can be calculated based on the three-dimensional coordinates of the target ligand atoms and the target protein. This ligand binding strength feature is then processed using a pre-defined neural network to obtain the RMSD value corresponding to the target conformation. It is understood that this RMSD value is related to the variable parameters in the spatial features of the target ligand.
[0105] In this way, by adjusting the variable parameters, the RMSD can be minimized, thereby obtaining the target control feature corresponding to the minimum RMSD value.
[0106] After determining the target spatial features, the spatial position of each atom in the target ligand can be calculated based on the specific values of the variable parameters in the target spatial features, thereby obtaining the complex conformation with the smallest RMSD value.
[0107] In some embodiments, spatial features are modulated based on information about the binding strength between proteins and ligands in the conformation of the target complex to obtain target spatial features, including:
[0108] B1. Based on spatial features, recalculate the binding strength information between the protein and ligand in the conformation of the target complex to obtain a binding strength representation function containing variable parameters;
[0109] B2. The gradient descent method is used to process the binding strength representation function to obtain the value of each variable parameter;
[0110] B3. Determine the target space characteristics based on the value of each variable parameter.
[0111] In this embodiment, the spatial features are adjusted based on the binding strength information between the protein and ligand in the conformation of the target complex. Specifically, this can involve first calculating the binding strength information represented by variable parameters in the spatial features, i.e., obtaining the RMSD represented by the variable parameters in the spatial features. Specifically, this can involve calculating the representation function of the RMSD represented by the variable parameters in the spatial features.
[0112] Since the representation function is fully differentiable, it can be iteratively solved using gradient descent, essentially optimizing the spatial characteristics of the target ligand repeatedly. When the number of iterations reaches a preset number or the decrease in the RMSD value is less than a preset range, the representation function is considered converged, yielding the final value of each variable parameter. The final values of the variable parameters determine the target spatial characteristics of the target ligand.
[0113] The following examples illustrate the accuracy of the ligand binding strength assessment method included in the ligand binding site assessment method for complexes provided in this application. For instance... Figure 3 The figure shown is a graph of the test results of the ligand binding strength assessment method in this application in a specific database. Figure 3 The results of the ligand binding strength assessment method (also known as the scoring function) in this application were tested in CASF-2016. CASF-2016 provides a standard for evaluating scoring functions. CASF-2016 contains 285 natural protein-ligand complexes. Each protein-ligand was re-docked using molecular docking software, generating approximately 100 complex conformations for testing the docking performance of the scoring function provided in this application. Figure 3As shown, the scoring function provided in this application exhibits stronger predictive performance as the RMSD range increases. When the RMSD is between 0 and 5 angstroms, the Spearman correlation is close to 0.7.
[0114] like Figure 4 The figure shows the test results of the conformation optimization method provided in this application in a specific database. As shown, the figure illustrates the probability distribution of successful conformation optimization for each protein-ligand in the aforementioned CASF-2016 test set. The conformation optimization method provided in this application has a success rate of over 50% in most protein-ligand conformation optimizations.
[0115] like Figure 5 The figure shows another test result of the conformation optimization method provided in this application on a specific database. As shown in the figure, the probability of successful conformation optimization is displayed in different RMSD intervals. The conformation optimization method provided in this application has an optimization success rate of nearly 70% in the RMSD interval of 1 to 3.
[0116] As described above, the ligand binding site evaluation method for complexes provided in this application involves binding the target ligand to multiple candidate binding sites in the target protein to obtain a candidate complex conformation corresponding to each candidate binding site; obtaining distance information between each atom of the target protein and each atom of the target ligand in each candidate complex conformation; determining the binding strength information between the target protein and the target ligand in each candidate complex conformation based on the distance information; and determining the target binding site based on the binding strength information.
[0117] Therefore, the ligand binding position evaluation method for complexes provided in this application can calculate the binding strength between the protein and ligand in the complex by analyzing the distance information between protein atoms and ligand atoms in the complex. Then, it determines the optimal binding position of the ligand based on the different binding strengths corresponding to different binding positions of the ligand. In other words, this method can accurately calculate the binding strength corresponding to different complex conformations obtained after the ligand binds to different binding positions of the protein, and then determine the optimal binding position based on the binding strength, thereby improving the accuracy of ligand binding position evaluation in complexes.
[0118] To better implement the above methods, embodiments of this application also provide a device for evaluating the ligand binding position of a complex, which can be integrated into a terminal or server.
[0119] For example, such as Figure 6 The diagram shown is a schematic representation of the structure of a ligand binding site evaluation device for a complex provided in an embodiment of this application. The device may include a binding unit 301, an acquisition unit 302, a first determination unit 303, and a second determination unit 304, as follows:
[0120] Binding unit 301 is used to bind the target ligand to multiple candidate binding sites in the target protein to obtain the candidate complex conformation corresponding to each candidate binding site;
[0121] Acquisition unit 302 is used to acquire distance information between each atom of the target protein and each atom of the target ligand in each candidate complex conformation;
[0122] The first determining unit 303 is used to determine the binding strength information of the target protein and the target ligand in each candidate complex conformation based on the distance information.
[0123] The second determining unit 304 is used to determine the target bonding position based on the bonding strength information.
[0124] In some embodiments, the first determining unit includes:
[0125] A subunit is generated to generate ligand binding strength evaluation features corresponding to each candidate complex conformation based on the distance information.
[0126] The evaluation subunit is used to input the ligand binding strength evaluation features corresponding to each candidate complex conformation into a preset neural network model to obtain the binding strength information of the target protein and the target ligand in each candidate complex conformation.
[0127] In some embodiments, the generating subunit includes:
[0128] A first determining module is used to determine a first number of first target atoms in the target protein based on different combinations of amino acids and heavy atoms;
[0129] The second determining module is used to determine a second number of second target atoms in the target ligand;
[0130] The third determining module is used to determine the target distance information between each first target atom and each second target atom from the distance information;
[0131] The first generation module is used to generate ligand binding strength evaluation features corresponding to each candidate complex conformation based on the target distance information.
[0132] In some embodiments, the first generation module includes:
[0133] The first calculation submodule is used to calculate the first feature dimension of the target distance information to obtain the first sub-ligand binding strength evaluation feature corresponding to each candidate complex conformation.
[0134] The second calculation submodule is used to calculate the second feature dimension of the target distance information to obtain the second sub-ligand binding strength evaluation feature corresponding to each candidate complex conformation.
[0135] The splicing submodule is used to splice the first subligand binding strength evaluation feature and the second subligand binding strength evaluation feature corresponding to each candidate complex conformation to obtain the ligand binding strength evaluation feature corresponding to each candidate complex conformation.
[0136] In some embodiments, the ligand binding site evaluation device for the complex provided in this application further includes:
[0137] The first acquisition subunit is used to acquire training samples, which include ligand binding strength evaluation features corresponding to multiple sample complex conformations and binding strength information corresponding to each sample complex conformation.
[0138] The training subunit is used to train a preset neural network model by taking the ligand binding strength evaluation features corresponding to the conformations of the multiple sample complexes as input and the binding strength information corresponding to the conformations of each sample complex as output.
[0139] In some embodiments, the first acquisition subunit includes:
[0140] The first acquisition module is used to acquire a first number of sample complexes;
[0141] The docking module is used to re-dock the ligands in each sample complex to obtain a second number of sample complex conformations;
[0142] The calculation module is used to calculate the binding strength information corresponding to the conformation of each sample complex.
[0143] The second acquisition module is used to acquire sample distance information between protein atoms and ligand atoms in the conformation of each sample complex.
[0144] The second generation module is used to generate ligand binding strength evaluation features corresponding to the complex conformation of each sample based on the sample distance information. The ligand binding strength evaluation features corresponding to the complex conformation of each sample and the binding strength information corresponding to the complex conformation of each sample constitute training samples.
[0145] In some embodiments, the second determining unit includes:
[0146] The first determining subunit is used to determine the target complex conformation among the candidate complex conformations based on the binding strength information;
[0147] The second determining subunit is used to determine the binding site corresponding to the conformation of the target complex as the target binding site.
[0148] In some embodiments, the ligand binding site evaluation device for the complex provided in this application further includes:
[0149] The second acquisition subunit is used to acquire the spatial features of the target ligand in the conformation of the target complex, the spatial features including variable parameters characterizing the spatial position of each atom in the target ligand;
[0150] The regulation subunit is used to regulate the spatial features based on the binding strength information between the protein and the ligand in the conformation of the target complex to obtain the target spatial features;
[0151] The third determining subunit is used to determine the spatial position of each atom of the target ligand in the conformation of the target complex based on the target spatial characteristics.
[0152] In some embodiments, the adjustment subunit includes:
[0153] The calculation module is used to recalculate the binding strength information of the protein and ligand in the conformation of the target complex based on the spatial features, and obtain a binding strength representation function containing the variable parameters;
[0154] The processing module is used to process the binding strength representation function using the gradient descent method to obtain the value of each variable parameter;
[0155] The target spatial features are determined based on the value of each variable parameter.
[0156] In practice, each of the above units can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units, please refer to the previous method embodiments, which will not be repeated here.
[0157] As described above, the ligand binding site evaluation device for the complex provided in this application embodiment binds the target ligand to multiple candidate binding sites in the target protein through the binding unit 301, thereby obtaining a candidate complex conformation corresponding to each candidate binding site; the acquisition unit 302 acquires the distance information between each atom of the target protein and each atom of the target ligand in each candidate complex conformation; the first determination unit 303 determines the binding strength information between the target protein and the target ligand in each candidate complex conformation based on the distance information; and the second determination unit 304 determines the target binding site based on the binding strength information.
[0158] Therefore, the ligand binding position evaluation method for complexes provided in this application can calculate the binding strength between the protein and ligand in the complex by analyzing the distance information between protein atoms and ligand atoms in the complex. Then, it determines the optimal binding position of the ligand based on the different binding strengths corresponding to different binding positions of the ligand. In other words, this method can accurately calculate the binding strength corresponding to different complex conformations obtained after the ligand binds to different binding positions of the protein, and then determine the optimal binding position based on the binding strength, thereby improving the accuracy of ligand binding position evaluation in complexes.
[0159] Furthermore, this application also provides a conformation optimization method that can optimize the spatial position information of atoms in the ligand to obtain a more stable complex conformation. This improves the efficiency and accuracy of conformation optimization.
[0160] This application also provides a computer device, which can be a terminal or a server, such as... Figure 7 The diagram shown is a structural schematic of the computer device provided in this application. Specifically:
[0161] The computer device may include components such as a processing unit 401 with one or more processing cores, a storage unit 402 with one or more storage media, a power module 403, and an input module 404. Those skilled in the art will understand that... Figure 7 The computer device structure shown does not constitute a limitation on the computer device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:
[0162] The processing unit 401 is the control center of the computer device. It connects various parts of the computer device via various interfaces and lines, and performs various functions and processes data by running or executing software programs and / or modules stored in the storage unit 402, and by calling data stored in the storage unit 402. Optionally, the processing unit 401 may include one or more processing cores; preferably, the processing unit 401 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processing unit 401.
[0163] Storage unit 402 can be used to store software programs and modules. Processing unit 401 executes various functional applications and data processing by running the software programs and modules stored in storage unit 402. Storage unit 402 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, and web page access); the data storage area may store data created based on the use of the computer device. In addition, storage unit 402 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, storage unit 402 may also include a memory controller to provide processing unit 401 with access to storage unit 402.
[0164] The computer equipment also includes a power supply module 403 that supplies power to various components. Preferably, the power supply module 403 can be logically connected to the processing unit 401 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply module 403 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0165] The computer device may also include an input module 404, which can be used to receive input numeric or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0166] Although not shown, the computer device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processing unit 401 in the computer device loads the executable files corresponding to the processes of one or more applications into the storage unit 402 according to the following instructions, and the processing unit 401 runs the applications stored in the storage unit 402 to realize various functions, as follows:
[0167] The target ligand is bound to multiple candidate binding sites in the target protein to obtain the candidate complex conformation corresponding to each candidate binding site; the distance information between each atom of the target protein and each atom of the target ligand in each candidate complex conformation is obtained; the binding strength information between the target protein and the target ligand in each candidate complex conformation is determined based on the distance information; and the target binding site is determined based on the binding strength information.
[0168] It should be noted that the computer device provided in this application embodiment and the method in the above embodiment belong to the same concept. The specific implementation of each of the above operations can be found in the previous embodiments, and will not be repeated here.
[0169] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0170] To this end, embodiments of the present invention provide a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute steps in any of the methods provided in the embodiments of the present invention. For example, the instructions can execute the following steps:
[0171] The target ligand is bound to multiple candidate binding sites in the target protein to obtain the candidate complex conformation corresponding to each candidate binding site; the distance information between each atom of the target protein and each atom of the target ligand in each candidate complex conformation is obtained; the binding strength information between the target protein and the target ligand in each candidate complex conformation is determined based on the distance information; and the target binding site is determined based on the binding strength information.
[0172] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0173] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0174] Since the instructions stored in the computer-readable storage medium can execute the steps of any of the methods provided in the embodiments of the present invention, the beneficial effects that any of the methods provided in the embodiments of the present invention can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.
[0175] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a storage medium. A processor of a computer device reads the computer instructions from the storage medium and executes the computer instructions, causing the computer device to perform the methods provided in various alternative implementations of the aforementioned method for evaluating the ligand binding position of the complex.
[0176] The foregoing has provided a detailed description of the method, apparatus, and computer equipment for evaluating the ligand binding position of complexes provided in the embodiments of the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, those skilled in the art will recognize that there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A method for evaluating the ligand binding site of a complex, characterized in that, The method includes: The target ligands are bound to multiple candidate binding sites in the target protein to obtain the candidate complex conformation corresponding to each candidate binding site. Obtain distance information between each atom of the target protein and each atom of the target ligand in each candidate complex conformation; Based on the distance information, determine the binding strength information of the target protein and the target ligand in each candidate complex conformation; The target bonding location is determined based on the bonding strength information; The step of determining the binding strength information of the target protein and the target ligand in each candidate complex conformation based on the distance information includes: Based on the distance information, ligand binding strength evaluation features are generated for each candidate complex conformation. The ligand binding strength evaluation features corresponding to each candidate complex conformation are input into a preset neural network model to obtain the binding strength information of the target protein and the target ligand in each candidate complex conformation. The step of generating ligand binding strength evaluation features for each candidate complex conformation based on the distance information includes: A first number of first target atoms are determined in the target protein based on different combinations of amino acids and heavy atoms; A second number of second target atoms are identified in the target ligand; The target distance information between each first target atom and each second target atom is determined from the distance information; Based on the target distance information, generate ligand binding strength evaluation features corresponding to each candidate complex conformation; The step of generating ligand binding strength evaluation features for each candidate complex conformation based on the distance information includes: The target distance information is calculated using the first feature dimension to obtain the first sub-ligand binding strength evaluation feature corresponding to each candidate complex conformation; The target distance information is calculated using the second feature dimension to obtain the second sub-ligand binding strength evaluation feature corresponding to each candidate complex conformation; The binding strength assessment features of the first and second sub-ligands corresponding to each candidate complex conformation are spliced together to obtain the ligand binding strength assessment features corresponding to each candidate complex conformation.
2. The method according to claim 1, characterized in that, Before inputting the ligand binding strength assessment features corresponding to each candidate complex conformation into a preset neural network model to obtain the binding strength information of the target protein and the target ligand in each candidate complex conformation, the method further includes: Acquire training samples, which include ligand binding strength evaluation features corresponding to multiple sample complex conformations and binding strength information corresponding to each sample complex conformation. The preset neural network model is trained by taking the ligand binding strength evaluation features corresponding to the conformations of the multiple sample complexes as input and the binding strength information corresponding to the conformations of each sample complex as output.
3. The method according to claim 2, characterized in that, The acquisition of training samples includes ligand binding strength evaluation features corresponding to multiple sample complex conformations and binding strength information corresponding to each sample complex conformation, including: Obtain the first number of sample complexes; The ligands in each sample complex were re-coupled to obtain a second number of sample complex conformations; Calculate the binding strength information corresponding to the conformation of each sample complex; Obtain sample distance information between protein atoms and ligand atoms in each sample complex conformation; Based on the sample distance information, a ligand binding strength evaluation feature corresponding to the complex conformation of each sample is generated. The ligand binding strength evaluation feature corresponding to the complex conformation of each sample and the binding strength information corresponding to the complex conformation of each sample constitute a training sample.
4. The method according to claim 1, characterized in that, Determining the target binding location based on the binding strength information includes: The target complex conformation is determined from the candidate complex conformations based on the binding strength information. The binding site corresponding to the conformation of the target complex is determined as the target binding site.
5. The method according to claim 4, characterized in that, The method further includes: The spatial characteristics of the target ligand in the conformation of the target complex are obtained, and the spatial characteristics include variable parameters characterizing the spatial position of each atom in the target ligand; The spatial features are adjusted based on the binding strength information between the protein and ligand in the conformation of the target complex to obtain the target spatial features; The spatial position of each atom of the target ligand in the conformation of the target complex is determined based on the target spatial characteristics.
6. The method according to claim 5, characterized in that, The spatial features are adjusted based on the binding strength information between the protein and ligand in the conformation of the target complex to obtain the target spatial features, including: Based on the spatial features, the binding strength information between the protein and ligand in the conformation of the target complex is recalculated to obtain a binding strength representation function containing the variable parameters; The gradient descent method is used to process the binding strength representation function to obtain the value of each variable parameter; The target spatial features are determined based on the value of each variable parameter.
7. A device for evaluating the ligand binding site of a complex, characterized in that, The device includes: The binding unit is used to bind the target ligand to multiple candidate binding sites in the target protein, thereby obtaining the candidate complex conformation corresponding to each candidate binding site. The acquisition unit is used to acquire distance information between each atom of the target protein and each atom of the target ligand in each candidate complex conformation; The first determining unit is configured to determine the binding strength information of the target protein and the target ligand in each candidate complex conformation based on the distance information. The second determining unit is used to determine the target bonding location based on the bonding strength information; The first determining unit includes: A subunit is generated to generate ligand binding strength evaluation features corresponding to each candidate complex conformation based on the distance information. The evaluation subunit is used to input the ligand binding strength evaluation features corresponding to each candidate complex conformation into a preset neural network model to obtain the binding strength information of the target protein and the target ligand in each candidate complex conformation. The generating subunit includes: A first determining module is used to determine a first number of first target atoms in the target protein based on different combinations of amino acids and heavy atoms; The second determining module is used to determine a second number of second target atoms in the target ligand; The third determining module is used to determine the target distance information between each first target atom and each second target atom from the distance information; The generation module is used to generate ligand binding strength evaluation features corresponding to each candidate complex conformation based on the target distance information. The generation module includes: The first calculation submodule is used to calculate the first feature dimension of the target distance information to obtain the first sub-ligand binding strength evaluation feature corresponding to each candidate complex conformation. The second calculation submodule is used to calculate the second feature dimension of the target distance information to obtain the second sub-ligand binding strength evaluation feature corresponding to each candidate complex conformation. The splicing submodule is used to splice the first subligand binding strength evaluation feature and the second subligand binding strength evaluation feature corresponding to each candidate complex conformation to obtain the ligand binding strength evaluation feature corresponding to each candidate complex conformation.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to perform the steps in the method for evaluating the ligand binding position of the complex according to any one of claims 1 to 6.
9. A computer device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the steps in the method for evaluating the ligand binding position of the complex according to any one of claims 1 to 6.
Citation Information
Patent Citations
Methods and systems for identifying ligand-protein binding sites
CN107111691A
Method, device and equipment for predicting affinity between drug and target spot
CN113517038A
Drug-target protein affinity prediction method and system
CN113823352A