A tyk2 inhibitor and screening method and application thereof based on fusion of deep learning scoring and traditional molecular docking scoring
By combining deep learning scoring and traditional molecular docking scoring methods, the problem of high false positive rate in traditional molecular docking methods was solved, and efficient screening of TYK2 inhibitors was achieved, and active compounds were screened out.
Patent Information
- Application Number
- CN202411760870.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-03
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-12-03
AI Technical Summary
Traditional molecular docking methods have a high proportion of false positive results when screening TYK2 inhibitors, resulting in low screening efficiency and inability to effectively identify the correct molecular conformation.
A method integrating deep learning scoring and traditional molecular docking scoring was used to perform data fusion by combining the Watvina scoring function and the CNNscore scoring function to screen potential TYK2 inhibitors.
The false positive rate in virtual screening was significantly reduced, the screening efficiency, accuracy and screening efficiency were improved, and compounds with TYK2 inhibitory activity could be quickly identified.
Smart Images

Figure CN119724407B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biomolecular structure prediction, and specifically to a TYK2 inhibitor and a screening method and application thereof based on the fusion of deep learning scoring and traditional molecular docking scoring. Background Art
[0002] Janus kinases (JAKs) are members of a family of non-receptor tyrosine kinases that includes JAK1, JAK2, JAK3, and TYK2. These kinases play key roles in traditional Type I and Type II immune cytokine receptor signaling pathways, making them potential targets for the treatment of a variety of inflammatory and autoimmune diseases. JAK2 plays an important role in multiple physiological processes, including tumor growth, while TYK2 has become a research hotspot due to its ability to interfere with the signaling pathways of IL-23 and IL-12. These two cytokines are closely associated with the expansion and survival of TH17 cells and, in turn, are implicated in the pathogenesis of major autoimmune diseases such as psoriasis and inflammatory bowel disease.
[0003] In recent years, with advances in drug design and screening technologies, significant progress has been made in the research and development of inhibitors targeting JAK2 and TYK2. For example, deucravacitinib (an N-trideuterated methylpyridazine derivative) was approved by the U.S. Food and Drug Administration (FDA) as the first TYK2 JH2 inhibitor. It effectively targets and stabilizes the TYK2 JH2 domain, thereby blocking the activity of TYK2 JH1. Furthermore, deucravacitinib exhibits high selectivity and low toxicity, offering significant advantages over other FDA-approved JAK inhibitors.
[0004] Molecular docking, a structure-based virtual screening method, plays an important role in identifying promising bioactive molecules from molecular databases. However, a major challenge in molecular docking is the high rate of false positive results, where the top-ranked molecular conformations are often incorrect. Previous studies have shown that when docking σ2, more than 98% of the top 100 molecules screened from 1.3 billion molecules exist in incorrect tautomeric forms. The root cause of this problem lies in the limitations of traditional molecular docking methods.
[0005] Traditional molecular docking methods typically adopt a search and scoring framework, in which search algorithms are used to explore potential ligand conformations, while scoring functions (SFs) are used to select the optimal ligand conformation and estimate the protein-ligand (PL) binding strength. However, relying on simplified force fields (FFs) and empirical energy terms to estimate binding energy may compromise accuracy. In addition, oversimplification of search algorithms and scoring functions may lead to inaccurate predictions of binding conformations and affinities, ultimately reducing the enrichment power of virtual screening (VS). These problems directly lead to a high proportion of false positive results, that is, the top-ranked molecular conformations are often incorrect. Therefore, a strategy that can reduce the false positive rate in virtual screening and improve screening efficiency is particularly important. Summary of the Invention
[0006] To address the problems existing in the prior art, the present invention proposes a TYK2 inhibitor and a screening method and application thereof based on the fusion of deep learning scoring and traditional molecular docking scoring. The screening method is highly efficient, the TYK2 inhibitor obtained by screening has a novel structure, and can effectively inhibit the enzymatic activity of TYK2 in vitro.
[0007] In order to achieve the above-mentioned purpose, the present invention adopts the following technical solutions:
[0008] The present invention provides a method for screening TYK2 inhibitors based on the fusion of deep learning scoring and traditional molecular docking scoring, which comprises the following steps:
[0009] S1: Obtain the TYK2 protein structure and small molecules in the compound library and pre-process them;
[0010] S2: performing molecular docking on the small molecule pre-treated in step S1 to obtain the docking conformation and its corresponding Watvina docking score;
[0011] S3: Calculating the relative binding free energy of the molecular docking conformations that meet the Watvina docking score requirements in step S2;
[0012] S4: re-scoring the molecular docking conformations whose relative binding free energy meets the requirements in step S3 based on deep learning, and outputting the deep learning scores;
[0013] S5: performing fusion scoring on the molecules that have been re-scored and meet the conditions in step S4, and selecting candidate compounds based on the fusion score ranking;
[0014] S6: performing in vitro activity tests on the candidate compounds in step S5 to screen out TYK2 inhibitors with inhibitory activity.
[0015] Furthermore, in step S1, the pre-processing of the protein structure includes deleting redundant chains, deleting solvent molecules, and adding hydrogen atoms and charges.
[0016] Furthermore, when performing molecular docking in step S2, the molecular docking is expanded in three coordinate directions according to the position information of the crystallized ligand of PDB ID 6nzp. Get the 3D information of the docking box.
[0017] Furthermore, in step S3, the small molecules that meet the Watvina docking score criteria are small molecules with a score no higher than -7.
[0018] Furthermore, in step S4, the molecules that meet the relative binding free energy requirement are small molecules whose relative binding energy is less than -60 Kcal / mol.
[0019] Furthermore, in step S5, the small molecules that meet the re-scoring conditions are those with a re-scored CNNscore of not less than 0.95.
[0020] Furthermore, the specific steps of relative binding free energy in step S3 are:
[0021] S31: Use PDBFixer to process the protein structure file pre-processed in step S1, including repairing missing residues, replacing non-standard residues, adding missing atoms and hydrogen atoms; use OpenMM's Modeller module to add an explicit water model; use openbable to convert the small molecule docking conformation file obtained in step S2 into sdf format; merge the processed protein structure and small molecule, and remove any residues with a distance less than water molecules, avoiding the presence of water molecules that collide with small molecules;
[0022] S32: using LangevinIntegrator of OpenMM to perform energy minimization on the complex system of step S32 to achieve a stable energy state;
[0023] S33: Obtain the potential energy of the complex system of step S32 through OpenMM; construct systems containing only the protein and the ligand respectively, and then calculate the potential energy of each; use the total energy of the complex system minus the sum of the total energies of the protein and the ligand when they are independent to obtain the relative binding energy.
[0024] Furthermore, in step S2, the molecular docking score is performed by Watvina; and in step S4, the re-scoring based on deep learning is CNNscore.
[0025] Furthermore, in step S5, the fusion score is obtained by multiplying the Watvina docking score of step S2 by the deep learning score of step S4.
[0026] The present invention also provides a TYK2 inhibitor, which is obtained by screening through the method described above.
[0027] Furthermore, the chemical structure of the TYK2 inhibitor is specifically:
[0028]
[0029] The present invention also provides use of the TYK2 inhibitor in preparing medicines for treating TYK2-mediated diseases.
[0030] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0031] To reduce the false positive rate in virtual screening, this paper employs a novel data fusion strategy, combining a deep learning scoring function (CNNscore) with an empirical scoring function (Watvina). The CNNscore scoring function is capable of learning complex features and evaluating the quality of small molecule and protein binding conformations. Previous studies have shown that CNNscore outperforms the Autodock Vina scoring function in predicting conformations and performing virtual screening.
[0032] Watvina is a molecular docking software used in the present invention. Based on the molecular docking engine of AutodockVina, Watvina has been optimized in terms of scoring function and conformational search algorithm. Watvina's scoring mechanism is based on empirical factors such as van der Waals forces, hydrogen bonds, polar-polar repulsion, and hydrophobic attraction. Unlike Autodock Vina, Watvina takes into account the contribution of all hydrogen atoms. The conformational search adopts a simplified genetic algorithm-BFGS combination strategy; in addition, the torsion penalty of conjugated rotatable single bonds is also calculated.
[0033] This study utilizes data fusion technology to multiply the deep learning docking score (CNNscore) with the empirical docking score (Watvina) to reduce the false positive rate in virtual screening and improve screening efficiency. This fusion strategy significantly improves the accuracy of screening for JAK2 and TYK2 inhibitors, particularly for TYK2 JH2 inhibitors, and helps quickly identify potentially effective compounds from a large library of compounds.
[0034] 2. The present invention uses layered virtual screening, combined with manual visual analysis, and then conducts in vitro activity testing on selected compounds. The results show that two of the 40 screened compounds have good TYK2 inhibitory activity. This demonstrates that the virtual screening method constructed in this invention is accurate and reliable, and can quickly and efficiently identify potential inhibitors from complex compound libraries, improving screening efficiency and saving experimental costs and time. Subsequent in vitro activity experiments further validate the biological activity of the virtual screening results, improving the accuracy of the screening results and reducing the false positive rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 The results are based on the DUDE dataset test to compare the performance of data fusion technology and Watvina screening alone;
[0036] Figure 2 This is the methodological validation result based on the TYK2 target pair data fusion technology;
[0037] Figure 3 It is to screen the chemical structures of the top 40 small molecules;
[0038] Figure 4 These are the IC50 curves of compounds STL250599 and STK595738. Specific implementation methods
[0039] The technical solution of the present invention is further described in detail with reference to the following specific examples.
[0040] In the following examples, unless otherwise specified, the experimental methods used are conventional methods, and the materials and reagents used can be purchased from biological or chemical reagent companies.
[0041] Based on testing on the DUDE dataset, this study explored methods for optimally integrating deep learning scoring with traditional molecular docking scoring. The screening performance of this data fusion method on the DUDE dataset was also investigated. Furthermore, the effectiveness of this method in virtual screening was tested on PCBA and CASF datasets. Finally, the effectiveness of this data fusion method for screening TYK2 inhibitors was verified for the TYK2 target. A large library of 12 million compounds was virtually screened using Watvina scoring, CNNscore scoring, and MMGBSA calculations. Data fusion technology was incorporated into the screening process to improve the hit rate of virtual screening. Finally, compounds were selected for in vitro activity verification, identifying lead compounds with TYK2 inhibitory activity.
[0042] Example 1: Exploration and verification of data fusion method
[0043] 1. Research on Data Fusion Methods on DUD-E Dataset
[0044] (1) Preparation of the DUD-E dataset: 102 protein targets and their corresponding active and decoy molecules were obtained from the DUD-E website (https: / / dude.docking.org / ). Proteins and small molecules were optimized and converted to the same format using the rdkit2pdbqt.py script.
[0045] (2) Molecular docking using Watvina software: Molecular docking of the optimized molecules was performed using Watvina software (https: / / github.com / biocheming / watvina) to obtain the docked conformation and docking score.
[0046] (3) Rescoring using CNNscore: CNNscore (https: / / github.com / gnina / gnina) was used to rescore the conformation of Watvina after docking.
[0047] (4) Explore the best data fusion method:
[0048] Method 1: Select the conformation with the lowest Watvina docking score and multiply its score by the corresponding CNNscore.
[0049] Method 2: Select the conformation with the highest CNNscore and multiply its score by the corresponding Watvina docking score.
[0050] Method 3: Calculate the multiplication of the Watvina score and the CNNscore for each conformation and select the conformation with the lowest multiplication score.
[0051] The advantages and disadvantages of the three methods were evaluated by enrichment factor. The results are shown in Table 1. Finally, method 1 was found to be the best data fusion method.
[0052] Table 1 Research on data fusion methods based on DUDE dataset
[0053] Data fusion method EF0.5% EF1% EF5% Method 1 13.40 10.99 5.12 Method 2 10.86 8.87 4.42 Method 3 9.87 8.66 4.75
[0054] 2. Comparison of the screening performance of the data fusion method and Watvina scoring alone on the DUD-E dataset
[0055] Based on the minimum Watvina docking score obtained above and the fusion score obtained from the best data fusion method (Method 1), the top 1%, 5%, and 10% enrichment factors were calculated. Statistical analysis was performed using a two-tailed Mann-Whitney U test to assess the difference in screening performance between data fusion scoring and Watvina scoring alone.
[0056] The results are as follows Figure 1 As shown, the fusion method (Method 1) significantly outperforms Watvina scoring alone on the DUD-E dataset. Specifically, on the top 1% EF, the average performance of the fusion method improved by X1, with a p-value below 0.001, demonstrating statistically significant screening performance. Furthermore, on the 5% and 10% EF, the average performance of the fusion method improved by X2 and X3, respectively, further validating the effectiveness and superiority of the fusion method.
[0057] 3. Comparison of the screening performance of the data fusion method and Watvina scoring alone on the LIT-PCBA dataset
[0058] (1) Preparation of the LIT-PCBA database: 15 protein targets, 7761 active molecules, and 382674 inactive molecules were obtained from the LIT-PCBA official website (http: / / drugdesign.unistra.fr / LIT-PCBA). Proteins and small molecules were optimized and converted to the same format using the rdkit2pdbqt.py script.
[0059] (2) Molecular docking using Watvina: Use Watvina to perform molecular docking on the optimized molecules to obtain the docked conformation and docking score.
[0060] (3) Re-scoring using CNNscore: Use CNNscore to re-score the conformation after Watvina docking.
[0061] (4) Fusion of Watvina scoring and CNNscore: Use the above method 1 for data processing.
[0062] (5) Calculate the enrichment factor to evaluate the screening performance of the data fusion method: Calculate the enrichment factors of the top 0.5%, top 1%, and top 5% using only Watvina scoring and using data fusion scoring.
[0063] The results are shown in Table 2. In the top 0.5%, top 1%, and top 5% enrichment factors, the data fusion method is superior to using only Watvina scoring, and the data fusion method improves the screening performance.
[0064] Table 2 Comparison of the performance of data fusion technology and Watvina screening alone based on PCBA dataset test
[0065]
[0066] 4. Validation of the data fusion method for TYK2 target
[0067] (1) Obtain active compounds from the ChEMBL database: 50 TYK2 JH2 active compounds were obtained from the ChEMBL website (https: / / www.ebi.ac.uk / chembl / ). After optimizing these compounds, the format was converted using the rdkit2pdbqt.py script.
[0068] (2) Generate decoy molecules using the DUD-E database: Submit active molecules on the DUD-E website (https: / / dude.docking.org / ). Each active molecule generates 50 decoy molecules, for a total of 2,500 decoy molecules. After optimizing all decoy molecules, convert the format using the rdkit2pdbqt.py script.
[0069] (3) Use Watvina to perform molecular docking on the active molecules and the bait molecules: Use Watvina to perform molecular docking on the optimized molecules to obtain the docked conformation and docking score.
[0070] (4) Re-scoring using CNNscore: Use CNNscore to re-score the conformation after Watvina docking.
[0071] (5) Fusion of Watvina scoring and CNNscore scoring: Use the above method 1 for data processing.
[0072] (6) Comparison of the screening performance of the data fusion method with that of using only CNNscore or Watvina scoring: Calculate the BEDROC (α = 321.9, 80.5, 20.0) and the top 1%, top 5% and top 10% enrichment factors of using only CNNscore or Watvina scoring and the data fusion method.
[0073] The results are as follows Figure 2 The results showed that for the TYK2 target, the data fusion method had better screening performance than relying solely on CNNscore or Watvina scoring.
[0074] Example 2: Method for screening TYK2 inhibitors based on fusion of deep learning scoring and traditional molecular docking scoring
[0075] 1. Obtaining and processing of TYK2 JH2 protein structure and small molecule library
[0076] (1) Protein preparation: Obtain TYK2 JH2 protein structure (PDB ID: 6NZP) from RCSB Protein Data Bank (PDB). Remove extra chains, delete solvent molecules, add hydrogen atoms and charges. Format the pre-processed protein through the rdkit2pdbqt.py script.
[0077] (2) Obtaining and processing of small molecules: 12 million small molecules are obtained from the TopScience database, and these small molecules are formatted through the rdkit2pdbqt.py script.
[0078] 2. Molecular docking using Watvina
[0079] According to the position information of the crystallization ligand of pdbid 6nzp, expand in three coordinate directions Obtain the three-dimensional information of the docking box. Use Watvina to perform molecular docking of small molecules and JH2 protein structure to obtain docking conformation and docking score. Remove small molecules with a score higher than -7.
[0080] 3. Calculation of relative binding energy
[0081] (1) Assemble the protein and small molecules and perform solvation treatment
[0082] Use PDBFixer (https: / / github.com / openmm / pdbfixer) to preprocess the protein file obtained in step 1(1), including repairing missing residues, replacing non-standard residues, adding missing atoms and hydrogen atoms, etc. Use the Modeller module of OpenMM (https: / / github.com / openmm / openmm) to add an explicit water model. Convert the pdbqt file of the small molecule conformation obtained in step 2 to sdf format using openbable. Merge the processed protein and small molecules, and remove any water molecules with a distance less than from the ligand to avoid the presence of water molecules colliding with small molecules.
[0083] (2) Energy minimization
[0084] Use the LangevinIntegrator of OpenMM to perform energy minimization on the complex system to achieve a stable energy state.
[0085] (3) Energy calculation
[0086] The potential energy of the complex system was obtained using OpenMM, employing the GBSA implicit solvation model. Systems containing only the protein and ligand were constructed separately, and the potential energies of each were calculated. Finally, the relative binding energy was calculated by subtracting the sum of the independent energies of the protein and ligand from the total energy of the complex system.
[0087] 4. Use CNNscore for rescoring and screening
[0088] CNNscore was used to re-score the conformations obtained in step 3 that satisfied the relative binding energy of less than -60 Kcal / mol, and small molecules with a CNNscore lower than 0.95 were eliminated.
[0089] 5. Fusion scoring and ranking
[0090] The docking score in step 2 is multiplied by the CNNscore in step 4 to obtain the fusion score. The fusion score is sorted as shown in Table 3, and the top 40 small molecules are retained. The chemical structures of these small molecules are as follows Figure 3 shown.
[0091] Table 3 Screening data of the top 40 small molecules
[0092]
[0093]
[0094] 6. In vitro activity test
[0095] The top 40 small molecules obtained in step 5 were tested for in vitro activity using the TYK2 JH2 Pseudokinase Domain Inhibitor Screening Assay Kit.
[0096] Experimental conditions: A 384-well plate was used, and the system volume was set to 20 μL. A preliminary screening of 40 compounds was performed, with the small molecule concentration set to 50 μM. 100 nM TYK2 JH2 protein and inhibitor were incubated at room temperature for 10 minutes, followed by the addition of 30 nM JH2 Probe 1 to initiate the binding reaction. After incubation at room temperature for 1 hour, the plate was read using a SpectraMax i3 plate reader (Molecular Devices) at an emission wavelength of 485 nm and an excitation wavelength of 535 nm. The results are shown in Table 3. Among the 40 compounds, STL250599 and STK595738 exhibited significant TYK2 inhibitory activity.
[0097] Further evaluation: Two of the screened compounds, STL250599 and STK595738, were further evaluated. These two compounds were dissolved in DMSO and prepared into 11 half-log dilution series, with the highest concentration being 100 μM. The final compound concentration range was 10 μM to 169.2 pM, and the final DMSO concentration was 1%. The biological evaluation process was the same as the initial evaluation. By plotting the IC50 curves of STL250599 and STK595738, as shown in the following figure: Figure 4 As shown, the half-maximal inhibitory concentrations were 13.756 μM and 9.985 μM, respectively.
[0098] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, it is still possible for a person skilled in the art to modify the technical solutions described in the aforementioned embodiments, or to replace some of the technical features therein with equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions claimed to be protected by the present invention.
Claims
1. A method for screening TYK2 inhibitors based on the fusion of deep learning scoring and traditional molecular docking scoring, characterized in that: It includes the following steps: S1: Obtain the TYK2 protein structure and small molecules in the compound library and pre-process them; S2: performing molecular docking on the small molecule pre-treated in step S1 to obtain the docking conformation and its corresponding Watvina docking score; S3: Calculating the relative binding free energy of the molecular docking conformations that meet the Watvina docking score requirements in step S2; S4: re-scoring the molecular docking conformations whose relative binding free energy meets the requirements in step S3 based on deep learning, and outputting the deep learning scores; S5: performing fusion scoring on the molecules that have been re-scored and meet the criteria in step S4, and selecting candidate compounds based on the fusion score ranking; the fusion score is obtained by multiplying the Watvina docking score in step S2 by the deep learning score in step S4; S6: performing in vitro activity tests on the candidate compounds in step S5 to screen out TYK2 inhibitors with inhibitory activity.
2. The method according to claim 1, characterized in that In step S3, the small molecules that meet the Watvina docking score criteria are small molecules with a score no higher than -7.
3. The method according to claim 1, characterized in that In step S4, the small molecules that meet the relative binding free energy requirement are those with a relative binding energy less than -60 Kcal / mol.
4. The method according to claim 1, wherein In step S5, the small molecules that meet the re-scoring conditions are those with a re-scored CNNscore of not less than 0.
95.
5. The method according to claim 1, wherein The specific steps of the relative binding free energy in step S3 are: S31: Use PDBFixer to process the protein structure file pre-processed in step S1, including repairing missing residues, replacing non-standard residues, and adding missing atoms and hydrogen atoms; use the Modeller module of OpenMM to add an explicit water model; use openbable to convert the small molecule docking conformation file obtained in step S2 into sdf format; merge the processed protein structure and small molecule, and remove any water molecules with a distance less than 1.5 Å from the ligand to avoid water molecules colliding with the small molecule; S32: using LangevinIntegrator of OpenMM to perform energy minimization on the complex system of step S32 to achieve a stable energy state; S33: Obtain the potential energy of the complex system of step S32 through OpenMM; construct systems containing only the protein and the ligand respectively, and then calculate the potential energy of each; use the total energy of the complex system minus the sum of the total energies of the protein and the ligand when they are independent to obtain the relative binding energy.
6. The method according to claim 1, characterized in that In step S2, the molecular docking score is performed by Watvina; in step S4, the re-scoring based on deep learning is performed by CNNscore.
7. A TYK2 inhibitor, characterized in that The TYK2 inhibitor is obtained by screening according to any one of claims 1 to 6.
8. The TYK2 inhibitor according to claim 7, characterized in that The chemical structure of the TYK2 inhibitor is specifically: 。 9. Use of the TYK2 inhibitor according to claim 8 in the preparation of a medicament for treating TYK2-mediated diseases.
Citation Information
Patent Citations
Screening method of hURAT1 inhibitor, screened compound and application thereof
CN116189808A
Application of small molecule compounds in preparation of novel coronavirus and ACE2 receptor binding inhibitor drugs
CN116230114A