A CDK12 inhibitor, its screening method and application
By combining computer-aided drug design with artificial intelligence, the screening system has solved the problems of insufficient selectivity and high toxicity of existing CDK12 inhibitors, and screened out novel CDK12 inhibitors with high selectivity and stability, realizing the therapeutic potential of liver cancer and renal tubular epithelial cell carcinoma.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGSU PROVINCIAL GOVERNMENT HOSPITAL
- Filing Date
- 2026-05-08
- Publication Date
- 2026-07-31
AI Technical Summary
Existing CDK12 inhibitors suffer from insufficient selectivity, significant toxic side effects, poor pharmacokinetic performance, and limited clinical evidence. Furthermore, the high structural homology between CDK12 and CDK13 makes precise targeted therapy difficult.
A screening system combining computer-aided drug design and artificial intelligence was adopted. Through molecular docking, deep learning model prediction and molecular dynamics simulation, novel CDK12 inhibitors with high selectivity and stability were screened out. The binding affinity of compounds in the allosteric binding pocket was predicted by graph neural network model, and the binding free energy was calculated by MM/GBSA method to screen out optimized compounds.
Several novel CDK12 inhibitors with higher selectivity and stability were successfully obtained, providing a foundation for the development of subsequent targeted drugs. The effectiveness of the inhibitors was verified through in vitro experiments.
Smart Images

Figure CN122493968A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of targeted drug development, specifically relating to a CDK12 inhibitor and its screening method and application. Background Technology
[0002] Cyclin-dependent kinase 12 (CDK12) is a highly conserved serine / threonine protein kinase belonging to the cyclin-dependent kinase (CDK) family and playing a crucial role in transcriptional regulation. The CDK12 gene is located on the long arm of human chromosome 17, region 1, band 2 (17q12), encoding a protein of approximately 164 kDa composed of 1490 amino acids. This protein structure includes two proline-rich motifs (PRM), an arginine / serine domain (RSdomain), and a C-terminal kinase domain (KD). The PRM mediates protein-protein interactions, the RS domain participates in precursor mRNA splicing, and the kinase domain exhibits a typical bilobal conformation.
[0003] CDK12 primarily functions by forming a complex with its specific cyclin, such as Cyclin K. This complex directly participates in and regulates gene transcription by phosphorylating a key site in the C-terminal domain (CTD) of RNA polymerase II (RNAP II). Furthermore, CDK12 plays crucial roles in several important cell biological processes, including DNA damage repair (DDR), precursor mRNA splicing, intronic polyadenylation (IPA), and maintaining genome stability.
[0004] CDK12 plays a crucial biological role in various solid tumors and other diseases. Its dysfunction can promote tumorigenesis and development by affecting DDR (radical retinal occlusion), regulating IPA (intra-invasive prostate), promoting R-loop formation, and increasing genomic instability. This has been widely reported in diseases such as breast cancer, ovarian cancer, esophageal cancer, gastric cancer, and prostate cancer. Although existing CDK12 / 13 inhibitors (such as THZ531) have shown some antitumor activity in in vitro and in vivo models, they still suffer from insufficient selectivity, significant toxic side effects, poor pharmacokinetic performance, and limited clinical evidence. Furthermore, the high structural homology between CDK12 and CDK13 makes achieving "precise targeting" even more difficult. Therefore, developing novel CDK12 inhibitors with higher selectivity, lower toxicity, and better pharmacokinetic properties has become an urgent need to further translate CDK12-targeted therapy into practical applications. Summary of the Invention
[0005] To address the shortcomings of existing technologies, the present invention aims to provide a CDK12 inhibitor, its screening method, and its application, thereby solving the problems in the prior art.
[0006] The objective of this invention can be achieved through the following technical solutions: A method for screening CDK12 inhibitors includes the following steps: S1, The obtained three-dimensional structure of CDK12 protein is standardized and preprocessed to determine the dual candidate target sites for virtual screening. The dual candidate target sites include an ATP ortho-binding pocket and an allosteric binding pocket. S2, obtain the original compounds, and construct a compound screening library after three-dimensional conformation generation and charge distribution processing; S3, the first docking scoring of the compound screening library and the dual candidate target sites is performed using a molecular docking procedure, and the binding affinity of the compounds in the allosteric binding pocket is predicted simultaneously using a graph neural network model, and a set of candidate compounds that meet the preset affinity threshold is extracted. S4, extract the molecular fingerprint of each molecule in the candidate compound set, construct a matrix based on the similarity coefficient between molecules and perform cluster analysis, and extract the central compound of each cluster to construct a representative molecule set; S5, Molecular docking and pharmacokinetics evaluation are performed on the representative molecular set, and priority candidate compounds are screened from the dual candidate target sites respectively; S6. Construct the complex system of the preferred candidate compound and CDK12 protein, perform molecular dynamics simulation and calculate the binding free energy based on the MM / GBSA method, and select the compound with the lowest binding free energy as the CDK12 inhibitor.
[0007] Furthermore, the allosteric binding pocket is located on the A chain of the CDK12 protein and belongs to a non-ATP competitive binding region.
[0008] Furthermore, the graph neural network model is the PLANET model; The process of cross-predicting the binding affinity of compounds within the allosteric binding pocket using a graph neural network model includes: using the three-dimensional structure of the allosteric binding pocket and the two-dimensional structure of the compound ligand as the joint input of the PLANET model to predict the pKd value of the compound, and setting the affinity extraction threshold to pKd greater than 7.0.
[0009] Furthermore, the extraction of molecular fingerprints includes: generating a Morgan molecular fingerprint with a radius of 2 and a fingerprint length of 2048 bits.
[0010] Furthermore, the similarity coefficient is the Tanimoto similarity coefficient; the clustering analysis uses the K-Means algorithm, and the number of extracted clusters is set to 100.
[0011] Furthermore, the parameters for the molecular dynamics simulation are set as follows: an explicit water model is used and the simulation is performed under a time length of 20 ns; the binding free energy is calculated by extracting uniform conformation frames within the last 2 ns of the trajectory. S6 also includes: extracting the hydrogen bond interaction network of the complex, limiting the distance between donor and acceptor atoms to ≤3.5 Å and the angle to ≥120°, and statistically analyzing the hydrophobic contact characteristics of residues with a distance <5 Å.
[0012] A CDK12 inhibitor was obtained by screening using the above-mentioned screening method.
[0013] Furthermore, the CDK12 inhibitor is CNP0123054, and its structural formula is: .
[0014] The above-mentioned CDK12 inhibitors are used in the preparation of drugs for the treatment of liver cancer and / or renal tubular epithelial cell carcinoma.
[0015] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the CDK12 inhibitor screening method described above.
[0016] The beneficial effects of this invention are: This invention constructs a comprehensive screening system based on computer-aided drug design and artificial intelligence. By combining molecular docking (AutoDock Vina), affinity prediction using deep learning models, and molecular dynamics simulation, it enables efficient virtual screening of CDK12 kinase pockets. Simultaneously, it utilizes databases such as DrugBank to conduct drug repositioning, forming a three-level progressive strategy of "structure screening - AI optimization - repositioning verification". This has successfully yielded several novel CDK12 inhibitor candidate compounds with higher selectivity and stability, laying the foundation for subsequent targeted drug development. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a schematic diagram of the virtual screening process for the CDK12 inhibitor of the present invention; Figure 2 This is a detailed diagram of the virtual screening of CDK12 inhibitors according to the present invention.
[0019] Figure 3 This is a schematic diagram illustrating the inhibition of CDK12 expression in HepG2 cells by CNP0123054 of the present invention; Figure 4 This is a schematic diagram illustrating the inhibition of CDK12 expression in HK2 cells by CNP0123054 of the present invention. Detailed Implementation
[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. In the present invention, unless otherwise expressly stated, all percentage concentrations (%) involving solutions or reagents refer to volume percentages (v / v).
[0021] Example 1 like Figure 1 and Figure 2 As shown, a method for screening CDK12 inhibitors includes the following steps: S1. By performing pretreatment, structure optimization and energy minimization operations on the acquired or predicted CDK12 protein, the binding pockets of the CDK12 protein (ATP binding pocket and allosteric binding pocket) are determined. The 3D structure of the target protein was obtained from authoritative literature or the Protein Structure Database (RCSB PDB). If the experimental structure was unavailable, the AlphaFold2 deep learning-based model was used for high-precision structure prediction. The obtained initial structure underwent normalization preprocessing, including removal of water molecules, cofactors, and irrelevant strands, and supplementation of missing residues or atoms. PDB2PQR was used for hydrogenation and charge assignment, followed by atom type standardization using AutoDockTools, and finally conversion to PDBQT format to complete the structural preparation before molecular docking. The binding site (pocket) was determined using the following methods: the ortho-constitutional site was determined based on the known ligand structure, and potential allosteric sites were predicted using the CavityPlus tool (developed by Professor Lu Hua Lai's group at Peking University). The pocket region with the best energy and geometric features was selected for subsequent docking. This site is located in protein chain A, has a large volume, surface area, and deep geometric features. After verifying its spatial accessibility through PyMOL visualization, it was identified as a candidate pocket for virtual screening.
[0022] The ATP-binding site (pocket 1) of CDK12 is a crucial structural region for its catalytic activity, responsible for recognizing and binding ATP to complete substrate phosphorylation. Therefore, it has become a primary design target for small molecule inhibitors (e.g., ...). Figure 2 On the other hand, computational tools also predicted and identified a potential allosteric binding region (pocket 2), located on chain A of the protein. This site has a large volume and depth and can regulate protein activity through conformational changes, making it a potential binding site for allosteric regulators (such as...). Figure 2 (B and C in the example). Therefore, in this embodiment, the ATP-binding pocket and the predicted allosteric binding pocket are selected as two candidate target sites for virtual screening.
[0023] S2 collects original compounds from multi-source databases and performs deduplication, conformation generation, protonation prediction and charge assignment to finally build a screening library with structural diversity and meeting docking requirements. Original compounds were collected from the commercial database ChemDiv (1,648,197 compounds), the drug database DrugBank (9,469 approved drugs), and the natural product database COCONUT (407,271 natural products). The original compounds were standardized using OpenBabel or RDKit for duplicate removal, structural normalization, and stereoisomer enumeration. Gypsum-DL was used for 3D conformation generation, protonation state prediction (considering physiological pH conditions), hydrogenation, and charge partitioning (using AM1-BCC or Gasteiger methods), with the output in a docking-compatible format (e.g., PDBQT). Finally, a multi-dimensional screening library was constructed, incorporating structural diversity, drug-likeness, and synthetic feasibility.
[0024] S3. Using molecular docking software and deep learning models, parallel calculations were performed to score the binding pockets of the screening library molecules and CDK12 protein, and a set of candidate compounds was initially screened out in order of predicted binding energy from low to high. High-throughput parallel docking was achieved using a high-performance computing cluster (LSF scheduling, over 10,000 physical cores, GPU acceleration). AutoDock Vina was used for dual-pocket mesh docking (20×20×20 Å mesh, exhaustiveness=32), retaining the 5,000 compounds with the best predicted binding energies from each pocket. For allosteric pockets, the PLANET graph neural network model (using the 3D pocket structure and 2D ligand structure as input) was further used to predict the pKd values of the compounds, selecting 5,000 high-affinity compounds with pKd>7.0, and cross-validating the two sets of results. Simultaneously, the KarmaDock model was used for supplementary docking and ranking. Finally, the candidate compound set was selected by ranking the compounds from lowest to highest predicted binding energy.
[0025] S4. For the top-ranked cross-validation compounds, a candidate molecule set is selected by calculating molecular fingerprints and applying clustering algorithms for chemical spatial classification and representative molecule extraction. After cross-validation, compounds were selected, and Morgan molecular fingerprints (radius 2, fingerprint length 2,048 bits) were generated using RDKit. Based on these fingerprints, the Tanimoto similarity coefficients between compounds were calculated using the following formula: Where a and b represent the number of fingerprint features of molecules A and B, respectively, and c represents the number of features shared by both. Using the Scikit-Learn toolkit in Python, combined with the Tanimoto similarity matrix obtained by the above method, cluster analysis was performed on the compounds. The K-Means algorithm was used for clustering, with a set number of clusters of 100. The cluster center compound of each cluster was extracted based on the clustering results, thereby constructing a diverse subset of compounds to ensure the diversity of the screening results in chemical space and reduce redundancy in subsequent verification.
[0026] S5 utilizes high-precision molecular docking and ADMET comprehensive evaluation to screen priority candidate compounds with two binding pockets from known drugs in DrugBank; Thirty-eight known drugs from DrugBank were screened from clustered compounds, and these compounds underwent refined molecular docking and ADMET property evaluation. After generating multiconformations in the Maestro-LigPrep module, high-precision molecular docking was performed using the Schrödinger Glide module (OPLS4 force field, Extra Precision mode), and docking scores were calculated. Simultaneously, the ADMET prediction platform system was used to evaluate pharmacokinetic and toxicity-related indicators, including hERG inhibition risk, hepatotoxicity, Caco-2 permeability, and oral bioavailability. Considering both docking scores and ADMET properties, four priority candidate drugs were ultimately selected from each of the two binding pockets.
[0027] S6 uses molecular dynamics simulations and binding free energy calculations to evaluate the stability of the complexes between candidate molecules and the target, and screens out several final candidate molecules with the lowest binding free energy. Molecular dynamics simulations of complexes formed by representative small molecules selected through cross-validation and cluster analysis and target proteins were performed using GROMACS 2023.4 software. The atomic partial charges of the small molecule ligands were calculated using the Multiwfn program to obtain the RESP2(0.5) charge, and Sobtop was used to generate topology and coordinate files suitable for GROMACS. GAFF force fields were used for parameterization. Protein components were described using the Amberff14SB force field. The TIP3P explicit water model was selected as the solvent environment, and appropriate amounts of Na⁺ and Cl⁻ ions were added to maintain the system's electroneutrality.
[0028] The simulated system was subjected to periodic boundary conditions in all three dimensions. First, the system's energy was minimized using the steepest descent algorithm, followed by pre-equilibrium phases under NVT and NPT ensembles. At 300 K, NVT equilibrium was performed for 100 ps using the Velocity-rescale method (coupling time constant 0.1 ps); subsequently, NPT equilibrium was performed for 100 ps using Berendsen pressure coupling (coupling time constant 2.0 ps) with the pressure controlled at 1 bar. Long-range electrostatic interactions were calculated using the SPME algorithm, with a real-world cutoff distance of 12 Å. All hydrogen bond lengths were constrained using the LINCS algorithm. The simulation integration step size was set to 2 fs, and the system conformation was saved every 10 ps.
[0029] A 20 ns production phase simulation was performed for each complex system. 500 conformations were uniformly extracted from the last 2 ns of each trajectory, and the MM / GBSA binding free energy was calculated using the gmx_MMPBSA tool. The binding free energy is expressed as follows: Further, the binding modes of the screened candidate compounds were systematically compared with those of reported CDK12 inhibitors (including the ATP-competitive inhibitor SR-4835 and the allosteric inhibitor BSJ-01-175). Hydrogen bond interaction networks (donor-acceptor atomic distance ≤ 3.5 Å, angle ≥ 120°) were analyzed using HBPLUS software, and hydrophobic contacts (residue distance < 5 Å) were statistically analyzed using the VMD Timeline plugin. Heatmaps were generated based on MM / GBSA residue energy decomposition results, allowing for a multi-dimensional comparison of the binding characteristics between the candidate compounds and the reference inhibitors. This analysis provides structural evidence for the inhibitory mechanisms of the candidate molecules and offers important guidance for subsequent point mutation verification experiments and molecular structure optimization.
[0030] In this embodiment, the docking score of AutoDock Vina, the binding affinity prediction of the PLANET model, and the binding free energy calculated by MM / GBSA were used to evaluate compounds at ATP binding sites (Pocket 1) and allosteric sites (Pocket 2) in a multidimensional way. Finally, four representative potential CDK12 inhibitor candidates were screened from each of the two types of pockets (Table 1).
[0031] Table 1 Docking scores and binding free energy fractions of candidate compounds The structural formulas of the eight compounds shown in Table 1 are as follows: CNP0123054: CNP0250284: G435-0657: Y043-6272: J030-0840: L286-0483: L759-0739: SA04-0089: As shown in Table 1, the compounds screened at the ATP binding site (Pocket 1) generally exhibited good docking scores (ranging from -11.691 to -11.504 kcal / mol), indicating strong initial binding ability to the target protein; the corresponding MM / GBSA binding free energies ranged from -41.01 to -38.66 kcal / mol, indicating good binding stability. In contrast, the compounds at the allosteric site (Pocket 2) showed weaker initial docking affinity than those in Pocket 1 (docking scores ranging from -9.84 to -9.299 kcal / mol), but their MM / GBSA binding free energies were in the range of -46.78 to -36.81 kcal / mol, suggesting strong binding stability in dynamic simulations. Further analysis revealed that all eight candidate compounds screened met or largely met the drug-likeness evaluation criteria of Lipinski, Pfizer, and Golden Triangle, demonstrating good drug-likeness characteristics.
[0032] Multidimensional computational evaluation can effectively screen CDK12 inhibitor candidates that possess both high binding affinity and good pharmacokinetic potential. These representative compounds from different binding sites provide diverse and high-potential starting points for subsequent in vitro biochemical validation, binding mode comparison analysis, and structural optimization.
[0033] The three-dimensional docking conformation of CNP0123054 in the CDK12 protein binding pocket is as follows: Figure 2 As shown in D, the schematic diagram illustrates the interaction of key amino acid residues of CNP0123054 in the CDK12 protein binding pocket. Figure 2 As shown in E, CNP0123054 can stably bind to the active pocket of the CDK12 protein and form synergistic effects with multiple key amino acid residues, including hydrogen bonds and hydrophobic interactions, thereby significantly enhancing the binding stability of the ligand to the target protein, suggesting its potential as a CDK12 inhibitor.
[0034] S7, through wet experiments, validated the final candidate molecule and confirmed the CDK12 inhibitor.
[0035] The wet experiment process includes: S71, Cell Culture The HepG2 human hepatocellular carcinoma cell line was cultured in DMEM / HighGlucose complete medium containing 10% fetal bovine serum and 1% penicillin-streptomycin solution under standard conditions at 37°C and 5% CO2 saturated humidity. The human renal tubular epithelial cell line HK-2 was cultured under the following conditions: 10% FBS, 1% P / S, DMEM / F12, 37°C, 5% CO2.
[0036] S72, Cell Intervention Cells were seeded at an appropriate density in six-well plates. After sufficient adhesion and growth to approximately 70–80% confluence, the medium was replaced with culture medium containing different concentrations (0, 10, 100, 500 nmol / L and 1, 10 μmol / L) of candidate CDK12 inhibitors for drug treatment. Cells were collected and proteins extracted 24 h after intervention. Changes in the expression of CDK12 and its downstream pathway-related molecules were detected to evaluate the efficacy of the compounds.
[0037] S73, Western blot (1) Tissue sample processing and protein extraction Mouse liver tissue was obtained, briefly rinsed in pre-chilled PBS, blotted dry with filter paper, and weighed. The tissue was placed in a pre-chilled glass homogenizer, and 200 μL of pre-chilled RIPA lysis buffer (containing 1× protease inhibitor and phosphatase inhibitor) was added per 20 mg of tissue. The homogenate was thoroughly homogenized on ice. The homogenate was transferred to centrifuge tubes and centrifuged at 12,000×g for 15 minutes at 4°C. The supernatant was collected as the total protein sample. The sample was aliquoted and stored at -80°C, avoiding repeated freeze-thaw cycles.
[0038] (2) Protein concentration quantification The BCA protein quantification kit was used, and the standard curve working solution was prepared according to the manufacturer's instructions. Protein samples were appropriately diluted with lysis buffer and then mixed with the working solution in 96-well plates, incubated at 37°C for 30 minutes. The absorbance at 562 nm was measured using a microplate reader, and the protein concentration of each sample was calculated based on the standard curve.
[0039] (3) SDS-PAGE gel electrophoresis Mix the protein sample with 5× loading buffer and heat in a metal bath at 99°C for 10 minutes to denature the protein. Load 20–30 μg of total protein into each well, along with pre-stained protein molecular weight standards (markers). Electrophoresis is performed in 1× Tris-Glycine SDS-PAGE buffer at a constant voltage of 120 V for approximately 60 minutes.
[0040] (4) Transfer membrane After electrophoresis, the gel and the PVDF membrane pre-activated in methanol for 1 minute were placed in a transfer clamp and transferred for 60 minutes at a constant current of 400 mA under ice bath conditions.
[0041] (5) Closed and primary antibody incubation The PVDF membrane was placed in a rapid blocking solution and blocked by shaking at room temperature for 15 min. The membrane was then incubated overnight at 4°C with anti-CDK12 / Ser2 / Ser7 primary antibody.
[0042] (6) Secondary antibody incubation Recover the primary antibody and wash the membrane three times with TBST for 10 minutes each time. Add HRP-labeled goat anti-rabbit secondary antibody and incubate at room temperature with shaking for 1 hour. Wash again three times with TBST for 10 minutes each time.
[0043] (7) Chemiluminescence imaging Mix equal volumes of ECL chemiluminescence working solution A and solution B, and uniformly drop the mixture onto the membrane. Incubate at room temperature for 1–2 minutes. Acquire the signal using a chemiluminescence imaging system, adjust the exposure time according to the signal intensity, and save the image as a digital file. Grayscale analysis can be performed using ImageJ or similar software, and the protein can be standardized using an internal reference protein (such as β-actin) to compare the relative expression levels of the target protein.
[0044] Example 2 In this embodiment, the inhibitory effect of the candidate compound CNP0123054 on CDK12 protein expression was evaluated by in vitro experiments. For the HepG2 cell experiment, the experimental materials, cell culture conditions, drug treatment methods, and Western blot detection steps used in this example are the same as those in S7 of Example 1.
[0045] For the HK2 cell experiment, the operation is basically the same as the HepG2 cell experiment, except that: (1) HepG2 cells are replaced with human renal tubular epithelial HK2 cells and cultured in DMEM / F12 medium containing 10% fetal bovine serum; (2) In the Western blotting assay, in addition to using anti-CDK12 antibody, Ser2 and Ser7 antibodies, which are related to the phosphorylation sites of RNA polymerase II C terminal domain (CTD), are used to detect the protein expression levels of CDK12, Ser2 and Ser7 in HK2 cells at the same time.
[0046] Experimental results are as follows Figure 3 and Figure 4 As shown; Figure 3 In the figure, A reflects the protein expression level of CDK12 in HepG2 cells as detected by Western blotting. Figure 3 B in the text reflects the... Figure 3 Quantitative analysis of protein expression levels in A in the study; Figure 4 In the figure, A reflects the protein expression levels of CDK12, Ser2, and Ser7 in HK2 cells as detected by Western blotting. Figure 4 B in the text reflects the... Figure 4 Quantitative analysis of protein expression levels in A in the study.
[0047] from Figure 3 and Figure 4 The results show that CNP0123054 can significantly downregulate CDK12 protein levels in various cell types, including HepG2 and HK2, and this effect is dose-dependent. In HK2 cells, the changes in CDK12 protein levels are consistent with the changes in Ser2 and Ser7 phosphorylation levels mediated by CNP0123054 on RNA polymerase II CTD, further supporting the inhibitory effect of CNP0123054 on CDK12.
[0048] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0049] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention.
Claims
1. A method of screening for a CDK12 inhibitor, characterized by, Includes the following steps: S1, The obtained three-dimensional structure of CDK12 protein is standardized and preprocessed to determine the dual candidate target sites for virtual screening. The dual candidate target sites include an ATP ortho-binding pocket and an allosteric binding pocket. S2, obtain the original compounds, and construct a compound screening library after three-dimensional conformation generation and charge distribution processing; S3, the first docking scoring of the compound screening library and the dual candidate target sites is performed using a molecular docking procedure, and the binding affinity of the compounds in the allosteric binding pocket is predicted simultaneously using a graph neural network model, and a set of candidate compounds that meet the preset affinity threshold is extracted. S4, extract the molecular fingerprint of each molecule in the candidate compound set, construct a matrix based on the similarity coefficient between molecules and perform cluster analysis, and extract the central compound of each cluster to construct a representative molecule set; S5, Molecular docking and pharmacokinetics evaluation are performed on the representative molecular set, and priority candidate compounds are screened from the dual candidate target sites respectively; S6. Construct the complex system of the preferred candidate compound and CDK12 protein, perform molecular dynamics simulation and calculate the binding free energy based on the MM / GBSA method, and select the compound with the lowest binding free energy as the CDK12 inhibitor.
2. The method for screening CDK12 inhibitors according to claim 1, characterized in that, The allosteric binding pocket is located on the A chain of the CDK12 protein and belongs to a non-ATP-competitive binding region.
3. The method for screening CDK12 inhibitors according to claim 1, characterized in that, The graph neural network model is the PLANET model; The process of cross-predicting the binding affinity of compounds within the allosteric binding pocket using a graph neural network model includes: using the three-dimensional structure of the allosteric binding pocket and the two-dimensional structure of the compound ligand as the joint input of the PLANET model to predict the pKd value of the compound, and setting the affinity extraction threshold to pKd greater than 7.
0.
4. The method for screening CDK12 inhibitors according to claim 1, characterized in that, The extraction of molecular fingerprints includes generating a Morgan molecular fingerprint with a radius of 2 and a fingerprint length of 2048 bits.
5. The method for screening CDK12 inhibitors according to claim 1, characterized in that, The similarity coefficient is the Tanimoto similarity coefficient; the clustering analysis uses the K-Means algorithm, and the number of extracted clusters is set to 100.
6. The method for screening CDK12 inhibitors according to claim 1, characterized in that, The parameters for the molecular dynamics simulation are set as follows: an explicit water model is used and the simulation is performed under a time length of 20 ns. The binding free energy is calculated by extracting uniform conformation frames within the last 2 ns of the trajectory. S6 also includes: extracting the hydrogen bond interaction network of the complex, limiting the distance between donor and acceptor atoms to ≤3.5 Å and the angle to ≥120°, and statistically analyzing the hydrophobic contact characteristics of residues with a distance <5 Å.
7. A CDK12 inhibitor, characterized in that, The sample was obtained by screening using the screening method described in any one of claims 1-6.
8. The CDK12 inhibitor according to claim 7, characterized in that, The CDK12 inhibitor is CNP0123054, and its structural formula is as follows: .
9. The use of the CDK12 inhibitor according to claim 7 or 8 in the preparation of a medicament for treating liver cancer and / or renal tubular epithelial cell carcinoma.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps in the screening method for CDK12 inhibitors as described in any one of claims 1-6.