Statistical calculation method of absolute binding free energy
Through a twin network model based on data sampling and pairing, the relative free energy of protein binding to ligand molecules is predicted, and the statistical calculation is converted into absolute binding free energy, which solves the problem of large prediction errors in the existing technology, and significantly improves the accuracy and stability of the prediction.
Patent Information
- Application Number
- CN202510041399.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-09
AI Technical Summary
The existing free energy prediction methods for protein binding to ligand molecules have large errors, and the prediction accuracy and stability are low.
A twin network model based on data sampling pairing is adopted, by constructing a reference data set and a target molecular data set, proteins and ligand molecules with high similarity are excluded, and the twin network model is trained to predict relative binding free energy, and converted into absolute binding free energy through statistical calculations.
The accuracy and stability of combined free energy predictions are significantly improved, the error in numerical calculations is reduced, and the instability of a single prediction is reduced through the pairing of multiple reference molecules.
Smart Images

Figure CN119964653A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for calculating absolute binding free energy, and in particular to a method for calculating absolute binding free energy of protein and drug small molecule based on data sampling pairing and applying twin network. Background Art
[0002] In the field of biochemistry, proteins can be represented as amino acid sequences represented by 20 characters. This sequence information is unique to each protein. Small molecule compounds that can bind to proteins can be represented as SMILES strings. SMILES strings played an important role in the early field of bioinformatics. For example, they can be used to quickly search molecular databases by searching strings. It is a commonly used chemical language, and deep learning can be used to predict the properties of drug molecules from SMILES strings. The amino acid sequence of proteins and the SMILES of ligand molecules already contain the required information. Therefore, the molecular structure of the complex composed of proteins and small molecules can be represented as a sequence in the form of a string and a special chemical language, SMILES, respectively, and this type of chemical language can convey the corresponding biochemical information.
[0003] Calculating the binding free energy between proteins and ligand small molecules through theoretical methods is a key technology for achieving rational drug screening and design. At present, the calculation methods include molecular docking scoring, calculation based on physical models, prediction based on artificial intelligence methods, and inference based on data knowledge, which have become a common research paradigm in drug design. The prediction of binding free energy is crucial and has important applications in early drug screening and lead compound optimization. Among them, absolute binding free energy is an indicator to measure the strength of binding between targets and small molecules, and its calculation accuracy directly affects the ability to predict ligand ranking. Relative free energy refers to the difference in binding free energy between new ligands and lead compounds. Relative binding free energy has a wider range of application scenarios. For example, in the process of drug discovery, relative binding free energy can help evaluate the binding ability of different compounds with target proteins, thereby screening out the most promising candidate drugs. By calculating the effects of different chemical modifications on binding free energy, the structure of drug molecules can be optimized and their activity and selectivity can be improved. By calculating the changes in binding free energy, the effects of mutations on protein function can be predicted, guiding protein design and engineering.
[0004] By calculating the change in binding free energy between drugs and targets, it can help study the molecular mechanism behind drug resistance and provide information for the development of new drugs. Absolute free energy calculation is suitable for scenarios where the binding ability of specific compounds to targets needs to be accurately evaluated, and is especially important in drug screening and target specificity research. Relative binding free energy is more suitable for comparing the binding affinity between different compounds or mutants, and is used for drug optimization and mutation impact assessment. These two methods have their own advantages and are usually used in combination at different stages of drug design and for different research goals. Among them, the change in binding energy before and after mutation (ΔΔG) is the key data for measuring drug resistance. If the energy change caused by the mutation exceeds a certain value, it indicates that the drug molecule cannot bind to the mutated protein. Research on drug resistance caused by protein mutations has also attracted research interest from academia and pharmaceutical companies. Imatinib is a tyrosine kinase inhibitor mainly used to treat chronic myeloid leukemia. Patients will develop drug resistance after taking it. It has been reported that 80% of patients who received imatinib treatment for 5 years developed drug resistance due to point mutations in the BCR-ABL1 kinase region, which caused imatinib to be unable to bind to proteins. Schrodinger used the free energy perturbation (FEP) method to calculate the free energy changes of kinase inhibitors caused by 144 ABL kinase mutations in the clinic. The FEP calculation gave a calculation error of the change in binding energy (ΔΔG) caused by the mutation within 1.1 kal / mol, and the accuracy of the calculation method reached 88%. However, FEP consumes a lot of computing resources and is difficult to carry out on a large scale. Tencent Quantum Lab has collected and compiled a database MdrDB for changes in binding energy caused by protein mutations. The database provides relevant data on changes in protein-ligand affinity caused by mutations in the three-dimensional structure of proteins. Zheng Mingyue's research group at the Shanghai Institute of Materia Medica developed a deep learning model PBCNet (pairwise binding comparison network) to quickly and accurately predict the relative binding free energy of ligands. PBCNet first extracts protein structural features from protein binding pockets through a graph convolutional neural network, then fuses the features, and finally uses the extracted features to predict the relative binding free energy between ligands.
[0005] In the process of calculating relative binding free energy, there is a problem of directly subtracting two relatively close numbers. From the numerical analysis, it can be seen that the calculation method of calculating relative free energy by absolute free energy is prone to introduce calculation errors. Therefore, the calculation of binding free energy requires high-precision calculation, and new methods are developed to accurately predict binding free energy. Summary of the invention
[0006] Purpose of the invention: The purpose of the present invention is to provide a statistical calculation method for absolute binding free energy to solve the problems of large errors, low prediction accuracy and stability in existing protein and ligand molecule binding free energy prediction methods.
[0007] Technical solution: A statistical calculation method of absolute binding free energy according to the present invention comprises the following steps:
[0008] The reference data set and target molecule data set are constructed based on the binding constants between proteins and ligand molecules in the database, the protein amino acid sequences and the SMILES of the ligand molecules;
[0009] The protein molecules with similar protein amino acid sequences and ligand molecules with similar molecular fingerprints in the reference data set are excluded to obtain a reference molecule set;
[0010] The target molecule data set is divided into a training molecule set and a test molecule set according to the protein mutation type, and the training molecule set is paired with the reference molecule set to obtain a paired sampling data set;
[0011] The twin network model is trained using the paired sampling data set to obtain a prediction model;
[0012] Based on the prediction model, a test molecule set is input and a predicted relative binding free energy is output;
[0013] Convert the predicted relative binding free energies into absolute binding free energy values.
[0014] The present invention transforms the calculation of absolute free energy into the calculation of relative free energy, thereby avoiding the error of numerical calculation; in the present invention, each target molecule can be paired with multiple reference molecules, so that more known data can be used as reference values; the calculated results in the present invention can be averaged by statistical calculation, thereby avoiding the instability of a single prediction.
[0015] Preferably, the construction of the reference data set and the target molecule data set comprises:
[0016] The binding constant between the protein and the ligand molecule is directly selected or calculated from the database, and then converted into ΔG. The calculation formula is ΔG = -0.5961 × logKi, where logKi is the logarithm of the binding constant. The reference data set and the target molecule data set are constructed using ΔG, protein amino acid sequence and SMILES of the ligand molecule.
[0017] Preferably, the exclusion of protein molecules with similar protein amino acid sequences and ligand molecules with similar molecular fingerprints in the reference data set includes:
[0018] The similarity between each protein amino acid sequence in the reference data set and all sequences in the database was calculated by BLAST, and the protein amino acid sequences with similarity calculation results greater than 90 in the reference data set were excluded;
[0019] The Tanimoto similarity of the ligand molecules was calculated, and the ligand molecules with similarity settlement results greater than 0.9 were excluded.
[0020] Preferably, pairing the training molecule set with the reference molecule set comprises:
[0021] An appropriate number of reference molecules are selected from the reference molecule set, and the training molecules in the training molecule set are paired with an appropriate number of reference molecules. When pairing, the small molecules and the target molecules are sorted according to their similarity, and the top N molecules are selected as reference molecules. The selected reference molecules form molecular pairs with the molecules in the training molecule set in a one-to-many form to obtain a paired sampling data set.
[0022] Preferably, the method of training the twin network model using the paired sampling data set includes: encoding the protein amino acid sequence and the ligand molecule in the paired sampling data set respectively, and inputting the encoded data into the twin network model for model training.
[0023] In order to better distinguish the difference between two input data, the present invention adopts a twin network, inputs the target molecule and the reference molecule into a network with shared weights, and uses the linear difference as the output of the network model. Compared with a single sample network, the twin network can learn two samples in the form of shared weights, and can better distinguish the slight differences in data. The present invention applies data sampling pairing and the twin network to the prediction of the binding free energy of proteins and ligand molecules, realizing a more accurate calculation method than the single sample data learning neural network model.
[0024] Preferably, the encoding of the protein amino acid sequence and the ligand molecule in the paired sampling data set respectively includes: converting the protein amino acid sequence in the paired sampling data set into a .pt file using the ESM model; converting the SMILES of the ligand molecule into an ECFP fingerprint using a module in RDKit; and inputting the .pt file and the ECFP fingerprint into the twin network model for training.
[0025] Preferably, the inputting of the test molecule set comprises: pairing the test molecule set with a reference molecule set to obtain a test set, inputting the test set into a prediction model, and outputting different relative binding free energy prediction values.
[0026] Preferably, said converting the predicted relative binding free energy into an absolute binding free energy value comprises:
[0027] The following formula is used to convert different relative binding free energy predictions into a statistically averaged absolute binding free energy value:
[0028]
[0029] Among them, ΔG i is the final predicted value after statistical averaging, ΔΔG i,j is the relative binding free energy change caused by the i-th and j-th molecules, ΔG ref,j is the absolute binding free energy of the jth molecule.
[0030] The present invention further discloses a computer-readable storage medium storing a computer program, wherein the computer program implements the steps of the above method when executed by a processor.
[0031] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages:
[0032] 1. The goal in the neural network model is changed from absolute free energy calculation to relative free energy calculation, which avoids the errors caused by outliers in numerical calculations. The distribution of free energy can be obtained by adjusting the number of reference molecules, so that the average binding free energy can be calculated by statistical methods, avoiding accidental errors caused by single calculations. Compared with traditional calculation methods that do not use pairing, the present invention significantly improves the accuracy of predictions. In 21 tests, the average value of the Pearson correlation coefficient increased from the original (corresponding to two cases: NoPair_BioLip_Seq90 and NoPair) 0.18 and 0.88 to 0.91, indicating that the linear relationship between the test results has become closer. This change reflects the improvement of this calculation method in capturing the relationship between variables in the data, showing a significant improvement in its performance;
[0033] 2. By pairing with known reference data, the training molecule set can be expanded to provide more training data for the neural network model in the artificial intelligence method. Each target molecule can be paired with multiple reference molecules, so that more known data can be used as reference values. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 Schematic diagram of database construction and paired sampling;
[0035] Figure 2 Schematic diagram of the twin network model and the feedforward neural network model;
[0036] Figure 3 Pearson correlation coefficient diagram of the predicted value and experimental value of relative binding free energy;
[0037] Figure 4Pearson correlation coefficient plot of the predicted and experimental values of absolute binding free energy;
[0038] Figure 5 The mean and variance of Pearson correlation coefficient when using and not using paired sampling method;
[0039] Figure 6 Scatter plot of predicted values and true values when using different numbers of reference samples;
[0040] Figure 7 Pearson correlation coefficient diagram for different sample sizes. DETAILED DESCRIPTION
[0041] The technical solution of the present invention is further described below in conjunction with the accompanying drawings.
[0042] Example 1: A statistical calculation method for absolute binding free energy is as follows:
[0043] S1. Construct reference dataset and target molecule dataset and perform data normalization.
[0044] The reference dataset and target molecule dataset include items such as the amino acid sequence of the protein, the SMILES of the ligand molecule, and the binding constant between the protein and the ligand molecule. Existing databases such as BioLip and PDBbind can be used to obtain data under these corresponding items. Since the data in the existing databases are scattered in different files, the reference dataset and target molecule dataset need to be reconstructed.
[0045] Take the BioLip database as an example, which contains 457,346 entries. First, find the rows in columns 14-17 that are not empty from the BioLiP_nonredu.txt file. Columns 14-17 are the binding affinity in the original literature, the binding affinity in BindingMOAD, the binding affinity in the PDBbind-CN database, and the binding affinity in BindingDB. There may be multiple values in columns 14-17. Try to select the value of the binding constant -logKd / Ki. If there is no -logKd / Ki, calculate it from Ki. Use the math package to calculate -logKd / Ki. Then, convert the value of -logKd / Ki into ΔG through the formula ΔG = -0.5961×logKi.
[0046] In order to obtain the SMILES of a small molecule, it is necessary to perform string matching in the ligand.tsv file according to the characters in the "Ligand ID" in the file to obtain the SMILES of the corresponding ligand molecule.
[0047] S2. Screening of reference datasets based on protein amino acid sequence alignment results and ligand molecule alignment results
[0048] The software BLAST was used to calculate the similarity between the amino acid sequence of each protein in the reference data set and the amino acid sequences of all proteins in the database, and the amino acid sequences of proteins in the reference data set with similarity calculation results greater than 90 were excluded;
[0049] Then the Tanimoto similarity of the ligand molecules was calculated, and the ligand molecules with similarity settlement results greater than 0.9 in the reference data set were excluded.
[0050] The purpose of this step is to avoid information leakage. All protein molecules with similar amino acid sequences and all ligand molecules with similar molecular fingerprints are excluded from the reference data set. After screening in this step, the number of proteins with amino acid sequence similarity between 70-90, 50-70, and 0-50 were 3646, 11638, and 81745, respectively.
[0051] Three data sets with sequence similarities of 70-90, 50-70, and 0-50 were used as reference molecule sets and named Seq90, Seq70, and Seq50, respectively.
[0052] S3. Pairing target molecules with reference molecules
[0053] In order to make full use of the existing data set, the data in the target molecule data set is first divided into a test molecule set and a training molecule set according to the protein mutation type. After the division, the ΔG in the test molecule set is assumed to be unknown and needs to be predicted. This assumption is for cross-validation. The mutation types contained in the target molecule data set are {M244V, L248R, L248V, G250E, Y253F, E255K, E255V, V299L, T315A, T315I, F317C, F317I, F317L, F317V, M351T, E355A, F359C, F359I, F359V, H396R, E459K}. For example, when the selected mutation type is L248R, the data corresponding to L248R is used as the test molecule set, and the rest are used as the training molecule set. According to this principle, all mutations are calculated in turn.
[0054] An appropriate number of reference molecules are selected from the reference molecule set, and the training molecules in the training molecule set are paired with an appropriate number of reference molecules, such as Figure 1As shown, when pairing, the small molecules and the target molecules are ranked according to their similarity, and the top N molecules are selected as reference molecules. After the reference molecules are selected, the training molecules in the training molecule set are paired with an appropriate number (the present invention uses N molecules as a tuning parameter) of reference molecules in a one-to-many form to obtain a paired sampling data set. The number of reference molecules N is set to 1, 10, 50, 100, 200, and 300. By setting different numbers of reference molecules, it is verified how many reference molecules are more appropriate for screening and selection.
[0055] S4. Encode protein sequences and ligand molecules in paired sampling datasets separately
[0056] The protein sequence is encoded using the latest ESM (Evolutionary Scale Modeling) model, which is a method that uses deep learning technology to predict protein structure and function. The protein sequence itself can be represented by a string. The core idea of ESM is to regard the protein amino acid sequence as a language, each amino acid as a character, and then use an autoregressive neural network to learn the statistical laws of this language.
[0057] The specific encoding method is as follows:
[0058] The present invention uses the weight model of esm2_t33_650M_UR50D. The output of ESM is a feature representation, which is a high-dimensional vector containing the structural and functional information of the protein. The protein amino acid sequences of the target molecule and the reference molecule are converted into .pt files through ESM.
[0059] Then, the SMILES of the small molecule was read using the RDKit software, and the module in the RDKit was used to convert the small molecule SMILES into the commonly used ECFP fingerprint.
[0060] After encoding the protein and ligand molecules separately, the .pt file and ECFP fingerprint can be input into the neural network model for training and prediction.
[0061] S5. Obtaining the prediction model and its predicted relative binding free energy
[0062] During the training of the twin network model, the twin network model is trained using a paired sampling data set obtained by pairing the training molecule with the reference molecule. After the twin network model training is completed, a prediction model is obtained, and then the prediction model is used to predict the target molecule in the test molecule set.
[0063] In order to obtain the distribution of prediction results for the test molecule set, the test molecule set is also paired with the reference molecule set to obtain a test set, which is then input into the prediction model to output different relative binding free energy prediction values.
[0064] S6. The predicted relative binding free energy is classified into groups according to the SMILES of the molecules, and the statistical average is calculated as the predicted value of the model, thereby converting the relative binding free energy into an absolute binding free energy value
[0065] The calculation of relative binding free energy requires a relatively higher calculation accuracy. If the absolute free energy is calculated first, errors may be introduced in the calculation of relative binding free energy. The calculation formula is:
[0066] ΔΔG=(ΔG i +ε i )-(ΔG j +ε j ) (1)
[0067] where ΔG i and ΔG j is the absolute free energy of the two states, ε i and ε j is the calculation error.
[0068] If the relative free energy is calculated first and then passed through a known reference stationary point, the relative free energy can be converted into absolute free energy. Since the relative binding free energy is usually a small number, the calculation error will not be amplified by adding the relative binding free energy to the reference stationary point. The calculation formula is as follows:
[0069] ΔG i =(ΔΔG i,j +ε i )+ΔG ref,j (2)
[0070] Among them, ΔΔG i,j is the relative binding free energy, ΔG ref,j is the reference stagnation point, usually precisely measured experimentally, ΔG i is the calculated absolute binding free energy.
[0071] Through the above analysis, the present invention proposes a binding free energy calculation method based on data sampling, which pairs the target molecule to be calculated with the reference molecule in the database, first calculates the relative free energy, and then calculates the absolute free energy through the reference value.
[0072] Finally, formula (3) is used to convert the predicted relative binding free energy result into a statistically average absolute binding free energy value.
[0073]
[0074] in, is the final absolute binding free energy calculated by statistical average, ΔΔG i,j is the relative binding free energy change caused by the i-th and j-th molecules, ΔG ref,j is the absolute binding free energy of the jth molecule, i.e., the reference stationary point, and N is the number of reference molecules.
[0075] S7. Comparison between the twin network model and the single-sample learning neural network model
[0076] The present invention compares the training and prediction performance of the twin network model and the ordinary feedforward neural network model on the reference data set and the target molecule data set.
[0077] The model diagram of the twin network is as follows Figure 2 As shown, the twin network model and the ordinary feedforward neural network model used in the present invention are both composed of three-layer neural networks, and the number of neurons contained in each layer is 2048, 8 and 1 respectively. A Dropout layer is used, and its value is 0.5, and the activation function used is relu. The neural network model is trained for 100 epochs using the training molecule set constructed by the present invention. The difference is that the twin network pairs the training molecules and the reference molecules and inputs them into the same network model, so that the network model can share parameters. The single-sample neural network model (i.e., the ordinary feedforward neural network model) only inputs the reference molecule into the model.
[0078] S8. Comparison of results using paired sampling and not using paired sampling
[0079] The paired sampling dataset contains training molecules and reference molecules, and uses the twin network model for training and prediction. Figure 3 In the data set without paired sampling, the single-sample neural network model is used for training and prediction. In order to compare with the present invention, in the common feedforward neural network model, the training molecules (marked as NoPair) and all molecules with amino acid sequence similarity between 70-90 (marked as NoPair_BioLip_Seq90) are used as training sets.
[0080] The calculation was performed on 21 sets of mutation data, and the results are as follows Figure 3As shown in Figure 2, the correlation coefficient between the directly predicted relative binding free energy and the experimental relative binding free energy is low. This is because the relative binding free energy itself is relatively small, and a small perturbation in the prediction can cause a large change in the data. Therefore, formula (3) is used to convert the relative binding free energy into absolute binding free energy. The calculation results are shown in Figure 2. Figure 4 As shown in Figure 2, it is found that the results of the data using paired sampling are significantly better than those without paired sampling. In order to facilitate analysis, the mean and variance of Pearson's test for the 21 groups of data were further calculated. The results are shown in Figure 2. Figure 5 As shown, seven combinations were calculated: using only training molecules (NoPair), using all molecules with amino acid sequence similarity between 70-90 (NoPair_BioLip_Seq90), pairing the target molecule with the wild-type molecule (Pair_Wild), pairing the target molecule with the mutant molecule (Pair_mut), pairing the target molecule with molecules with amino acid sequence similarity between 70-90 (Pair_Seq90), pairing the target molecule with molecules with amino acid sequence similarity between 50-70 (Pair_Seq70), and pairing the target molecule with molecules with amino acid sequence similarity between 0-50 (Pair_Seq50). It can also be seen from the analysis of the mean and root mean square error that when the sequence similarity is 70-90, the mean of the Pearson coefficient corresponding to Pair_Seq90 is 0.91 and the variance is 0.13. The mean of the Pearson coefficient obtained when the paired sampling data is not used (NoPair and NoPair_BioLip_Seq90) is 0.88 and 0.18, and the variance is 0.17 and 0.27. In addition, the test will also be conducted on the internal self-paired data set. When only one wild type (Pair_Wild) is used, the Pearson coefficient is -0.23, and when the mutant data (Pair_mut) set is used, the Pearson coefficient is 0.86. The calculated results are all lower than the prediction results of the Pair_Seq90 sampling pairing proposed in the present invention. This experiment proves that the use of the paired sampling method can improve the accuracy of the prediction results.
[0081] S9. Comparison of results using different numbers of reference molecules
[0082] Using different numbers of reference molecules can obtain different reference values of ΔΔG i,j , and then by adding it to the reference value, ΔΔG i,j Transformed into ΔG iWhen multiple reference molecules are used, a predicted distribution for each molecule can be obtained. The present invention uses the average value as the final result. Of course, in practical applications, other methods can also be used to obtain an expected absolute free energy value from the distribution. The randomness of the results when using a single reference molecule can be avoided by statistical averaging.
[0083] In order to reduce the amount of computation of the twin network model and considering that there are not too many reference molecules in practical applications, the number of reference molecules N is set to 1, 10, 50, 100, 200, and 300. The optimal number of reference molecules is determined through testing.
[0084] from Figure 6 As can be seen from the results of the scatter plot, as the number of reference molecules increases, the slope of the fitted line gradually increases, and reaches a maximum value when the number of molecules is 50. When the number of molecules exceeds 100, the slope of the fitted line and the correlation coefficient decrease again. The reason is that too many molecules are introduced, which will bring in some abnormal values. In order to evaluate the applicability and stability of the present invention in different systems, the present invention calculates the average value and average variance of the Pearson coefficient in 21 tasks. The calculation results are shown in Figure 2. Figure 7 As shown, the naming is performed according to the sequence similarity used and the number of selected molecules. When the molecules with amino acid sequence similarity between 50-70 are selected, they are named Seq70_N1, Seq70_N10, Seq70_N50, Seq70_N100, Seq70_N200, Seq70_N300, respectively. The average Pearson coefficients of Seq70_N50 and Seq90_N50 are 0.89 and 0.92, respectively. The calculation results show that when the number of reference molecules in the calculation is set to 50, the calculated Pearson coefficient is higher than that of other numbers. This result shows that the number of reference molecules needs to strike a balance between diversity and chaos.
[0085] The present invention compares the paired sampling method with or without sampling, and proves that the predicted results using the paired sampling method are significantly better than the neural network model of single sample learning. In addition, the influence of the number of reference molecules on the prediction results was further tested, and it was found that when the number of reference molecules was set to 50, the Pearson coefficient was the highest, the variance was the smallest, and the prediction performance was optimal.
[0086] The present invention provides a method for calculating absolute binding free energy based on data sampling. The results show that the paired sampling method can be used to predict the distribution of binding free energy, and more reliable prediction data can be obtained through statistical averaging, which improves the performance of the neural network model in predicting binding free energy. The present invention provides a numerical method for evaluating the binding strength between proteins and ligand molecules, which can be used for early drug screening and judging drug resistance in the case of mutations, and provides a more accurate calculation method for artificial intelligence-assisted drug screening.
Claims
1. A statistical calculation method for absolute binding free energy, characterized in that: The steps include: The reference data set and target molecule data set are constructed based on the binding constants between proteins and ligand molecules in the database, the protein amino acid sequences and the SMILES of the ligand molecules; The protein molecules with similar protein amino acid sequences and ligand molecules with similar molecular fingerprints in the reference data set are excluded to obtain a reference molecule set; The target molecule data set is divided into a training molecule set and a test molecule set according to the protein mutation type, and the training molecule set is paired with the reference molecule set to obtain a paired sampling data set; The twin network model is trained using the paired sampling data set to obtain a prediction model; Based on the prediction model, a test molecule set is input and a predicted relative binding free energy is output; Convert the predicted relative binding free energies into absolute binding free energy values.
2. The statistical calculation method of absolute binding free energy according to claim 1, characterized in that: The construction of the reference data set and the target molecule data set comprises: The binding constant between the protein and the ligand molecule is directly selected or calculated from the database, and then converted into ΔG. The calculation formula is ΔG = -0.5961 × lgKi, where lgKi is the logarithm of the binding constant. The reference data set and the target molecule data set are constructed using ΔG, protein amino acid sequence and SMILES of the ligand molecule.
3. The statistical calculation method of absolute binding free energy according to claim 1, characterized in that: The protein molecules with similar protein amino acid sequences and ligand molecules with similar molecular fingerprints in the excluded reference data set include: The similarity between each protein amino acid sequence in the reference data set and all sequences in the database was calculated by BLAST, and the protein amino acid sequences with similarity calculation results greater than 90 in the reference data set were excluded; The Tanimoto similarity of the ligand molecules was calculated, and the ligand molecules with similarity settlement results greater than 0.9 were excluded.
4. The statistical calculation method of absolute binding free energy according to claim 1, characterized in that: The pairing of the training molecule set with the reference molecule set comprises: An appropriate number of reference molecules are selected from the reference molecule set, and the training molecules in the training molecule set are paired with an appropriate number of reference molecules. When pairing, the small molecules and the target molecules are sorted according to their similarity, and the top N molecules are selected as reference molecules. The selected reference molecules form molecular pairs with the molecules in the training molecule set in a one-to-many form to obtain a paired sampling data set.
5. The statistical calculation method of absolute binding free energy according to claim 1, characterized in that: The method of training the twin network model using the paired sampling data set includes: encoding the protein amino acid sequence and the ligand molecule in the paired sampling data set respectively, and inputting the encoded data into the twin network model for model training.
6. The statistical calculation method of absolute binding free energy according to claim 5, characterized in that: The encoding of the protein amino acid sequence and the ligand molecule in the paired sampling data set includes: using the Evolutionary ScaleModeling model to convert the protein amino acid sequence in the paired sampling data set into a .pt file; using the module in the RDKit to convert the SMILES of the ligand molecule into an ECFP fingerprint; and inputting the .pt file and the ECFP fingerprint into the twin network model for training.
7. The statistical calculation method of absolute binding free energy according to claim 1, characterized in that: The inputting of the test molecule set comprises: pairing the test molecule set with the reference molecule set to obtain a test set, inputting the test set into the prediction model, and outputting different relative binding free energy prediction values.
8. The statistical calculation method of absolute binding free energy according to claim 7, characterized in that: The conversion of the predicted relative binding free energy into an absolute binding free energy value comprises: The following formula is used to convert different relative binding free energy predictions into a statistically averaged absolute binding free energy value: Among them, ΔG i is the final predicted value after statistical averaging, ΔΔG i,j is the relative binding free energy change caused by the i-th and j-th molecules, ΔG ref,j is the absolute binding free energy of the jth molecule.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.