Virtual Screening Method, Device, Equipment and Storage Medium of LCK Inhibitor Based on Deep Learning

Through a virtual screening method based on deep learning, the efficient LCK inhibitor 1232030-35-1 was screened, solving the problem of lack of specific LCK inhibitors in the prior art and achieving effective treatment of T-ALL.

CN118571312BActive Publication Date: 2025-07-01HENAN CANCER HOSPITAL +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410664841.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-27
Publication Date
2025-07-01
Estimated Expiration
2044-05-27

AI Technical Summary

Technical Problem

The lack of efficient and specific LCK inhibitors in the prior art makes it difficult to effectively treat T-cell acute lymphoblastic leukemia (T-ALL), an invasive hematologic malignant tumor.

Method used

Using a deep learning-based method, data cleaning and clustering are performed by obtaining molecular data of LCK targets, and virtual screening is performed using PLANET deep learning algorithm and Glide SP molecular docking program to screen out potential LCK inhibitors, and in vitro inhibitory activity detection and biological effect evaluation are performed.

Benefits of technology

The LCK inhibitor 1232030-35-1 with extremely low IC50 value was successfully screened, which significantly inhibited LCK signaling, showed strong anti-leukemia effects in vitro and in vivo, and extended the survival of T-ALL xenograft mice.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118571312B_ABST
    Figure CN118571312B_ABST
Patent Text Reader

Abstract

The present invention relates to a virtual screening method for LCK inhibitors based on deep learning, comprising: obtaining molecular data of the LCK target, cleaning the data, and obtaining a dedicated evaluation dataset of the LCK target; calculating the Morgan fingerprints of each molecule in the dedicated dataset and establishing a similarity matrix, and clustering each molecule to form multiple clusters; obtaining an active molecule set in each of the clusters, and generating a decoy molecule set according to an active molecule set; respectively obtaining ligand small molecules and receptor proteins, and evaluating and determining a deep learning algorithm and a molecular docking program according to the active molecule set and the decoy molecule set; respectively obtaining a molecular database and a protein, and performing virtual screening according to the determined deep learning algorithm and molecular docking program to obtain quasi-target molecules; detecting the in vitro inhibitory activity of the quasi-target molecules against LCK to obtain final target molecules; and evaluating the biological effects of the final target molecules. The present invention also relates to a device, equipment, and a storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer-aided drug design, and more particularly to a virtual screening method, device, equipment and storage medium of LCK inhibitors based on deep learning. Background Art

[0002] Lymphocyte-specific protein tyrosine kinase (LCK) plays a crucial role in the development and activation of T cells. Dysregulation of LCK signaling has been shown to drive the tumorigenesis of T cell acute lymphoblastic leukemia (T-ALL), thus providing a therapeutic target for leukemia treatment. LCK is a member of the SRC family kinases (SFKs), specifically expressed in T cells, and plays a central role in the activation of T cells mediated by the T cell receptor (TCR) signal. Similar to mature T cells, T lymphoid progenitor cells in the thymus also require LCK to transmit the strong proliferation signal triggered by the pre-TCR complex during early T cell development. In this context, dysregulation of pre-TCR-LCK activity has been shown to lead to uncontrolled cell expansion and is prone to induce T-ALL, which is an aggressive hematological malignancy.

[0003] More than 40% of clinical T-ALL cases exhibit constitutive activation of pre-TCR-LCK signals. Unbiased analysis of the global phosphoproteome of human T-ALL consistently shows widespread activation of LCK in T-ALL cell lines and primary samples. These findings indicate that blocking LCK in T-ALL has therapeutic potential against different cell lines. Despite efforts to discover effective LCK inhibitors, the number of candidate drugs with high activity and specificity in relevant disease models remains limited. Summary of the Invention

[0004] To address the above deficiencies, according to the first aspect of the present invention, a virtual screening method of LCK inhibitors based on deep learning is provided, including: obtaining molecular data of the LCK target, cleaning the data, and obtaining a dedicated evaluation dataset of the LCK target; calculating the Morgan fingerprints of each molecule in the dedicated evaluation dataset of the LCK target and establishing a similarity matrix, and clustering each molecule to form multiple clusters; obtaining an active molecule set in each cluster, and generating a decoy molecule set according to the active molecule set; respectively obtaining ligand small molecules and receptor proteins, and evaluating and determining a deep learning algorithm and a molecular docking program according to the active molecule set and the decoy molecule set; respectively obtaining a molecular database and a protein, and performing virtual screening according to the determined deep learning algorithm and molecular docking program to obtain quasi-target molecules through the virtual screening; detecting the in vitro inhibitory activity of the quasi-target molecules against LCK to obtain final target molecules; and evaluating the biological effects of the final target molecules.

[0005] Optionally, it further includes: simulating and verifying the LCK inhibitory effect of the final target molecule through molecular dynamics.

[0006] Optionally, the evaluated deep learning algorithm includes: the PLANET deep learning algorithm, and the evaluated molecular docking programs include four molecular docking programs: the AutoDock-GPU molecular docking program, the Autodock Vina molecular docking program, the LeDock molecular docking program, and the Glide SP molecular docking program; the determined molecular docking program includes at least one of the four evaluated molecular docking programs.

[0007] Optionally, the determined deep learning algorithm is the PLANET deep learning algorithm, and the determined molecular docking program is the Glide SP molecular docking program.

[0008] Optionally, the virtual screening includes at least two steps. Among them, the first virtual screening step is to screen through the PLANET deep learning algorithm, and the second virtual screening step is to screen again through the Glide SP molecular docking program.

[0009] Optionally, the score passing the first virtual screening step is set to be higher than 7, and the score passing the second virtual step is set to be lower than -7.

[0010] Optionally, the quasi-target molecules are respectively the first quasi-target molecule, the second quasi-target molecule, the third quasi-target molecule, and the fourth quasi-target molecule.

[0011] Optionally, detecting the in vitro inhibitory activity of the quasi-target molecule against LCK includes that the IC50 value of the third quasi-target molecule is 0.43 nM.

[0012] Optionally, the third quasi-target molecule is the final target molecule, and its chemical formula is:

[0013]

[0014] According to a second aspect of the present invention, there is provided a virtual screening device for LCK inhibitors based on deep learning, comprising: an acquisition and clustering module: used to acquire molecular data of the LCK target, clean the data, and obtain a dedicated evaluation dataset for the LCK target, calculate the Morgan fingerprints of each molecule in the dedicated evaluation dataset for the LCK target and establish a similarity matrix, and cluster each of the molecules to form a plurality of clusters; an acquisition and generation module: used to randomly select an active molecule set from each of the clusters, and generate a decoy molecule set according to the active molecule set; an evaluation and determination module: used to respectively acquire ligand small molecules and receptor proteins, and evaluate and determine a deep learning algorithm and a molecular docking program according to the active molecule set and the decoy molecule set; a virtual screening module: respectively acquire a molecular database and a protein, and perform virtual screening according to the determined deep learning algorithm and molecular docking program, and obtain quasi-target molecules through the virtual screening; a detection and evaluation module: used to detect the in vitro inhibitory activity of the quasi-target molecules against LCK, obtain final target molecules; and perform a biological effect evaluation on the final target molecules.

[0015] According to a third aspect of the present invention, there is provided an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above method.

[0016] According to a fourth aspect of the present invention, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the above method.

[0017] According to a fifth aspect of the present invention, there is provided a computer program product, comprising a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0018] Embodiments of the present invention identify potential LCK inhibitors and compare the efficacy of the deep learning algorithm PLANET based on a graph neural network and traditional docking tools. PLANET not only outperforms several mainstream screening tools in terms of efficiency, but also exhibits performance comparable to that of the screening suite. The present invention integrates PLANET with GlideSP to reconstruct and generate conformations. This combined method successfully identified four effective LCK inhibitors from the MCE drug-like molecule dataset. The novelty of these molecules was confirmed by skeletal novelty assessment. The similarity range of these four LCK inhibitors is between 0.2 and 0.4, indicating that they have significant novelty. Among them, the most prominent molecule is 1232030-35-1, whose IC50 The value is extremely low, only 0.43 nM. Through rapid screening by deep learning scoring methods and conformational selection using Glide, the powerful screening results not only highlight the potential of 1232030-35-1 as an LCK inhibitor but also represent the first experimental verification of PLANET's ability to selectively screen active molecules against specific targets. Brief Description of the Drawings

[0019] Hereinafter, the drawings of the exemplary embodiments of the present invention are shown by way of example, and the same or similar reference numerals are used in the respective drawings to represent the same or similar elements. In the drawings:

[0020] Figure 1 Shows the flowchart of the virtual screening method for LCK inhibitors based on deep learning of the exemplary embodiments of the present invention.

[0021] Figure 2A Shows the structural profiles of the active compounds and decoy compounds of the exemplary embodiments of the present invention.

[0022] Figure 2B Shows the molecular weight distributions of the active compounds and decoy compounds of the exemplary embodiments of the present invention.

[0023] Figure 2C Shows the LogP values of the active compounds and decoy compounds of the exemplary embodiments of the present invention.

[0024] Figure 3A Shows the ROC curves of the docking software performance of each model of the exemplary embodiments of the present invention.

[0025] Figure 3B Shows the schematic diagrams of the software molecular docking scores of the active compounds and decoy compounds of each model of the exemplary embodiments of the present invention.

[0026] Figure 4A Shows the flowchart of the virtual screening targeting LCK of the exemplary embodiments of the present invention.

[0027] Figure 4B Shows the structural diagrams of four small molecules obtained by virtual screening of the exemplary embodiments of the present invention.

[0028] Figure 5A Shows the inhibitory effects of four quasi-target molecules of the exemplary embodiments of the present invention.

[0029] Figure 5B Shows the inhibition rates of four quasi-target molecules of the exemplary embodiments of the present invention.

[0030] Figure 6AShows the comparison graph of the kinase inhibitory activity between the final target molecule of the exemplary embodiment of the present invention and 302962-49-8.

[0031] Figure 6B Shows the comparison graph of the cell viability between the final target molecule of the exemplary embodiment of the present invention and 302962-49-8.

[0032] Figure 6C Shows the schematic diagram of evaluating the leukemia burden of T-ALL xenograft mice after drug administration by bioimaging in the embodiment of the present invention.

[0033] Figure 6D Shows the quantification of luciferase signal in NPG mice inoculated with Jurkat-Luc cells in the embodiment of the present invention.

[0034] Figure 6E Shows the Kaplan-Meier survival curve of T-ALL xenograft in the embodiment of the present invention.

[0035] Figure 7A Shows the schematic diagram of the RMSD value between the final target molecule of the exemplary embodiment of the present invention and LCK.

[0036] Figure 7B Shows the time schematic diagram of the interaction within the final target molecule 1232030-35-1-LCK complex of the exemplary embodiment of the present invention.

[0037] Figure 7C Shows the schematic diagram in the interaction of the final target molecule 1232030-35-1-LCK complex of the exemplary embodiment of the present invention.

[0038] Figure 7D Shows the schematic diagram of the overall binding mode between the final target molecule 1232030-35-1 of the exemplary embodiment of the present invention and LCK.

[0039] Figure 7E Shows the three-dimensional visualization schematic diagram of the protein-ligand interaction of the exemplary embodiment of the present invention.

[0040] Figure 7F Shows the two-dimensional schematic diagram of the protein-ligand interaction of the exemplary embodiment of the present invention.

[0041] Figure 8 Shows the structural schematic diagram of the virtual screening device for LCK inhibitors based on deep learning in the exemplary embodiment of the present invention.

[0042] Figure 9 Shows the structural schematic diagram of the electronic device of the exemplary embodiment of the present invention. Detailed implementation manners

[0043] To better explain the present invention for easier understanding, the present invention will be described in detail below in conjunction with the accompanying drawings through specific implementation manners.

[0044] In the present invention, the term "and / or" is intended to cover all possible combinations and sub - combinations of the listed elements, including any one of the separately listed elements, any sub - combination or all elements, without necessarily excluding other elements. Unless otherwise specified, the terms "first", "second", etc. are used to describe various elements without intending to limit the positional relationship, temporal relationship or importance relationship of these elements. Such terms are only used to distinguish one element from another. Unless otherwise specified, the orientation or positional relationship indicated by terms such as "front, rear, upper, lower, left, right" is usually based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of description and simplification of the description, and cannot be understood as a limitation on the protection scope of the present invention.

[0045] The screening methods based on artificial intelligence have developed rapidly and exceeded traditional docking screening tools in terms of speed. However, most of these methods are evaluated using datasets such as DUD - E (Database of Useful Decoys: Enhanced, accessible at https: / / dude.docking.org) and LIT - PCBA, lacking verification in real - world screening scenarios. The present invention has established an LCK validation dataset for evaluating the capabilities of different scoring methods in LCK inhibitor screening work. Among them, the deep - learning algorithm (scoring method) PLANET (Protein - Ligand Affinity Prediction Network) has shown good inhibitory performance. Through virtual screening, the present invention has identified four candidate compounds with LCK inhibition potential. And the present invention has conducted a comprehensive biological effect evaluation on the best - performing molecule 1232030 - 35 - 1, confirming its inhibitory effect on intracellular LCK signal transduction and its anti - leukemia efficacy in vitro and in vivo. To deepen the understanding and lay a foundation for the further development of this molecule in the future, the present invention further uses molecular dynamics simulations to clarify the intricate binding kinetics within the 1232030 - 35 - 1 - LCK complex.

[0046] Figure 1 The flowchart of the virtual screening method for LCK inhibitors based on deep learning according to an exemplary embodiment of the present invention is shown.

[0047] S102: Obtain the molecular data of the LCK target, clean the data, and obtain a dedicated evaluation dataset for the LCK target. Calculate the Morgan fingerprints of each molecule in the dedicated evaluation dataset for the LCK target and establish a similarity matrix, and cluster the molecules to form multiple clusters.

[0048] In some embodiments, first, compound data related to the LCK target, IC 50 values and SMILES format structures were extracted from the ChEMBL database. These structures were cleaned and standardized using RDKit, incomplete records were removed, and duplicate records were deleted based on InChI keys. This refinement process resulted in an edited set containing 1,688 unique molecules.

[0049] Second, the Morgan fingerprints of each molecule were calculated and converted to a vector format for ease of calculation. To evaluate molecular similarity, a similarity matrix was created using the Tanimoto coefficient and the cdist method. Finally, we divided the molecules into 50 different clusters.

[0050] Among them, the main dataset came from the ChEMBL database (access date: 2023-10-11I). After processing, duplicate entries and those lacking labels or SMILES information were removed. In addition, this dataset also included pIC 50 values converted from the original IC 50 measurements, thus providing a more detailed understanding of molecular activity.

[0051] -lnIC 50 = pIC 50 ;

[0052] Then, Morgan fingerprints were generated using RDkit (version 2019.09.1), and then a similarity matrix was created. These techniques provide a quantitative measure of the similarity between different molecules based on fingerprints. Finally, clustering operations were performed using the sklearn clustering tool to divide the molecules into 50 different clusters.

[0053] S104: Obtain a set of active molecules in each of the clusters and generate a set of decoy molecules based on the set of active molecules.

[0054] DUD-E contains a protein-ligand compound benchmark. It contains various experimentally confirmed active compounds and their affinities for different targets, as well as the corresponding decoy molecules that cannot bind to the targets. These decoys have similar physicochemical properties to the active compounds, and their two-dimensional topological structures are different from the original compounds. DUD-E is a commonly used benchmark for promoting computational docking methods.

[0055] In previous clustering studies, 50 clusters were identified, and one active compound was randomly selected from each cluster, for a total of 50 active compounds. Then, 2,900 decoy compounds were made using DUD-E (database) based on these active compounds. Figure 2AThe structural distributions of the active compounds and decoy compounds are shown and differentiated by a legend. To better compare the physicochemical properties of the active compounds and decoy compounds, we used the open-source cheminformatics toolkit RDKit to evaluate their molecular aspects. As Figure 2B and Figure 2C shown, there are no significant differences in the molecular weight distributions ( Figure 2B ) and LogP values ( Figure 2C ) of the active compounds and decoy compounds.

[0056] In some embodiments, the molecular weight of each compound was obtained using the Descriptors.MolWt function in RDKit, and the partition coefficient or LogP was calculated using the Descriptors.MolLogP function. In terms of data visualization, embodiments of the present invention used the matplotlib and seaborn libraries. In some embodiments, the "Set2" color palette was used to draw the charts. As Figure 2B and Figure 2C shown, the generated Violin Plots, which combine the characteristics of box plots and kernel density plots, provide a comprehensive view of the data distribution. In addition, we also showed the quartile values in the violin plots to understand the data distribution in detail. By adopting these techniques, we aim to comprehensively analyze the molecular properties and their distributions of the active compounds and decoy compounds.

[0057] S106: Obtain the ligand small molecules and receptor proteins respectively, and evaluate and determine the deep learning algorithm and molecular docking program according to the active molecule set and the decoy molecule set.

[0058] To determine the most accurate molecular docking method, the present invention evaluated five different algorithms and programs: the PLANET deep learning algorithm, AutoDock-GPU, Autodock Vina, LeDock, and Glide SP, four molecular docking programs. In the test, 50 LCK active compounds and 2,900 LCK decoy compounds were used. In some embodiments, we used the T-test and AUC-ROC to visualize the results to measure the accuracy of each software. As Figure 3A and Figure 3B shown, the performances of the PLANET deep learning algorithm and Glide SP are comparable. However, one of the limitations of the standalone PLANET is that it cannot generate the three-dimensional conformation of the compounds. Therefore, we first used the PLANET deep learning algorithm as the first virtual screening to screen the MCE drug database for preparation for the subsequent steps. Subsequently, embodiments of the present invention utilized GlideSP to screen the compounds.

[0059] PLANET Deep Learning Algorithm

[0060] PLANET uses the illustrated three-dimensional structure of the target protein binding pocket and the two-dimensional chemical structure or SMILES of the ligand molecule as its main input information. The model is exhaustively trained using a three-objective strategy, focusing on three intertwined tasks: measuring the binding affinity between the protein and the ligand, mapping the contact interface between the protein and the ligand, and formulating the ligand distance matrix. In some embodiments of the present invention, PLANET was obtained from Github via the following link: (https: / / github.com / ComputArtCMCG / PLANET / ). For PLANET, the SMILES representation of the ligand was used as the input. The protein with the PDB ID: 3byo and its associated crystal ligand were obtained from the RCSB database. In addition, we also used the Protein Preparation Wizard in the 2021-4 suite. Subsequently, docking was performed using the PLANET_run.py script. In addition, embodiments of the present invention also obtained prediction scores.

[0061] AutoDock-GPU Molecular Docking Program

[0062] To prepare for the docking simulation, the present invention used MGLTools (version 1.5.7). In the first step, the prepare_ligand4.py script was used, which helps to ensure the accurate representation of hydrogen atoms in the ligand, thus generating an output file in pdbqt format. For the preparation of the receptor, we obtained the PDB code 3byo from the RCSB database and used the prepare_receptor4.py script. At this stage, we thoroughly verified the hydrogen atoms and removed any detected water molecules. In the second step, the prepare_gpf4.py script was used to design the binding site, setting the box_range parameter to 50,50,50 and centering the grid box at the coordinates 25.78, 37.61, 85.30. To improve work efficiency, the ligand files were divided into eight different folders, with the aim of facilitating the parallel execution of the script. During the actual molecular docking process, Autodock-gpu (version 1.5.3) was selected. To ensure accuracy, we limited the output structures and poses of each ligand to a maximum of five.

[0063] AutoDock Vina Molecular Docking Program

[0064] In AutoDock Vina (version 1.2.3), prepare_ligand4.py was used to prepare the ligand, and prepare_pdb_split_alt_confs.py was used to generate the configure.txt file. This configuration file includes the following parameters: center_x = 25.78, center_y = 37.61, center_z = 85.30; size_x = 50, size_y = 50, size_z = 50; energy_range = 3; exhaustiveness = 48; num_modes = 20. Then, prepare_receptor4.py was used to prepare the receptor, which included removing water and verifying the presence of hydrogen. Finally, we used Vina to perform the molecular docking process.

[0065] LeDock molecular docking program

[0066] LeDock is used for fast and accurate flexible docking of small molecules with proteins, with a pose prediction accuracy of over 90% on the Astex diversity set and a running time of about 3 seconds for each drug-like molecule. It uses the SYBYL Mol2 format for small molecule input. In the present invention, Lepro was used for protein preparation, and the ligand file was divided into five folders to facilitate parallel script execution. Conformational sampling used the default settings, combining simulated annealing and evolutionary optimization. The docking score was calculated using the default scoring function.

[0067] Glide SP molecular docking program

[0068] In the protein preparation stage, we used the Protein Preparation Wizard in 2021 - 4 to add missing hydrogen atoms, determine partial charges, define side chains / loops, and finally, optimize hydrogen bonds and limit energy minimization. In terms of ligand processing, we used the LigPrep module under default settings and the Epik tool (version 5.6) under the OPLS4 force field to perform operations such as hydrogenation and isomer generation within the pH range of 7.0 ± 2.0, ensuring that each ligand has up to 32 stereoisomers. Before docking, we created the necessary files centered on the protein pocket coordinates, set the original ligand coordinates as the docking focus, and used In 2021 - 4, a grid file after removing the original ligand was generated. The Glide module was used for molecular docking, which is crucial for predicting three - dimensional coordination and understanding molecular interactions that affect drug specificity and efficacy. When initiating the molecular docking simulation, we used the SP precision mode, introduced the pre - arranged LCK target receptor and ligand grid files, as well as the OPLS4 force field, limited the maximum output structures / coordinations of each ligand to 10 to obtain the best results, and then ran the program.

[0069] S108: Obtain the molecular database and protein respectively. According to the determined deep - learning algorithm and molecular docking program, perform virtual screening, and obtain quasi - target molecules through the virtual screening.

[0070] The molecular database is crucial for the virtual screening process. In the present invention, molecules were carefully screened from the MCE - type drug database. To ensure the maximum integrity and reliability of the screening task, the embodiments of the present invention implemented a strict pre - processing strategy. The RDKit was used to systematically identify and delete duplicate molecular structures in the database. Finally, the MCE - type drug database consisted of 8,617 molecules. Then, we stored these refined molecular structures in the SMILES format, ready for the first - stage screening by PLANET.

[0071] During the protein preparation process, the co - crystal structure with PDB ID: 3byo was obtained from the RCSB Protein Data Bank. Subsequently, we used the Protein Preparation Wizard in the 2021 - 4 suite to guide our program. This included adding missing hydrogens, determining partial charges, defining side chains / loops, optimizing hydrogen - bond assignments, and performing constrained energy minimization. Then, the crystal ligand was extracted from the co - crystal protein structure and the ligand was saved as a separate SDF file. This file is called the crystal ligand file in PLANET, used to determine the binding pocket and overwrite the provided coordinates when specified.

[0072] In the validation experiments of some embodiments, PLANET and Glide SP (standard precision) mode performed best. Although PLANET performed well in some aspects, it lacked the ability to generate three - dimensional conformations of compounds, which limited the intuitive observation of the interaction between compounds and LCK. To overcome this limitation, the embodiments of the present invention provided a comprehensive virtual screening workflow that effectively utilized PLANET and Glide SP. These two algorithms / programs complement each other. As Figure 4A shown, first use PLANET to initiate the screening process, focusing on the MCE - type drug database, and screen out the top 1,262 compounds for further study. Then use Glide SP was used to improve the accuracy of compound conformations. The scope was narrowed down to molecules with a PLANET score higher than 7 and a Glide SP score lower than -7 kcal / mol, and a total of 474 compounds were finally screened out. Glide SP score lower than -7 kcal / mol, and a total of 474 compounds were finally screened out.

[0073] Through cooperation and consultation with medicinal chemistry experts, we identified four molecules as shown below: the first quasi-target molecule, the second quasi-target molecule, the third quasi-target molecule, and the fourth quasi-target molecule. Among them, Figure 4B as shown below: the first quasi-target molecule, the second quasi-target molecule, the third quasi-target molecule, and the fourth quasi-target molecule. Among them,

[0074] The structural formula of the first quasi-target molecule 1197958-12-5 is:

[0075]

[0076] The structural formula of the second quasi-target molecule 781613-23-8 is:

[0077]

[0078] The structural formula of the third quasi-target molecule 1232030-35-1 is:

[0079]

[0080] The structural formula of the fourth quasi-target molecule 1234356-69-4 is:

[0081]

[0082] Further analysis was carried out. To evaluate the structural uniqueness of these molecules, the Morgan molecular fingerprint and Tanimoto similarity calculation methods were used. Compared with existing LCK inhibitors, the Skeleton novelty of these compounds was between 0.2 and 0.4, thus confirming the novelty of the selected compounds. After completing this meticulous process, bioassay tests were conducted, as detailed later.

[0083] In some embodiments, the prepared MCE drug database was used, where the input ligands were in SMI format, and the docking was performed using the PLANET_run.py script. In the protein preparation stage, The protein preparation wizard in the 2021-4 suite was used to add missing hydrogen atoms, determine partial charges, define side chains / loops, optimize hydrogen bond assignments, and perform constrained energy minimization. In ligand processing, the LigPrep module and Epik tool (version 5.6) of the same suite were used to handle operations such as hydrogenation, desalting, and isomer generation within the pH range of 7.0 ± 2.0, generating up to 32 stereoisomers for each ligand. When preparing for docking, necessary files were created centered on the coordinates of the protein pocket, and the coordinates of the original ligands were set as the focus of docking. Subsequently, these ligands were removed, and the 2021-4 suite was used to generate the grid file. At the start of the simulation, a pre-arranged LCK target receptor and ligand grid file were used, the force field was set to OPLS4, and the maximum number of output structures / coordinations for each ligand was 10 to ensure the best results.

[0084] More specifically, in the data screening and processing step, experimental data from a predefined molecular scoring database was used, and the data was stored in CSV file format. Two screening thresholds were set: to visualize the screening results, we used the Matplotlib library to create a scatter plot of molecular scores with a gradient background from red to green, enhancing the visual effect. This gradient was generated using a uniform grid in the Numpy library and applying a custom linear segmented colormap. On the scatter plot, the positions of all molecules were marked in blue, and the threshold was represented by a red dashed line at a Glide_SP value of -7. According to the results of the last round of scoring, a total of 474 molecules with a PLANET score greater than 7 and a Glide_SP score less than -7 were screened out.

[0085] In the similarity calculation step, the RDKit library was used to convert SMILES strings into chemical structures and generate Morgan molecular fingerprints (radius 2), generate molecular fingerprints and calculate similarities. Then, Tanimoto similarity calculations were performed between the fingerprint of each query compound and the fingerprints of all compounds in the database. The compound with the highest similarity in each query result was found, and its SMILES string and similarity value were recorded. The results were visualized using the Matplotlib library to draw a bar chart.

[0086] S110: Detect the in vitro inhibitory activity of the quasi-target molecule against LCK to obtain the final target molecule; evaluate the biological effects of the final target molecule.

[0087] First, the in vitro inhibitory activity of these compounds against LCK at a concentration of 25 μM was evaluated using the homogeneous time-resolved fluorescence (HTRF) assay. Table 1 and Figure 5AThe inhibitory results of each molecule are shown, and the inhibition rates of four quasi-target molecules are close to 100% at the tested concentrations. Further study on their half maximal inhibitory concentration (IC 50 ) values. Combining Figure 5B with Table 1 shows that the compound with the identifier 1232030-35-1 has the strongest inhibitory effect. Its IC 50 value is 0.43 nM, comparable to the efficacy of the multi-target kinase inhibitor TG 100572 used as a positive control in the current system. This finding highlights the great inhibitory potential of 1232030-35-1 against LCK.

[0088] Table 1. Scores of different compounds at each stage and the inhibition rates of small molecules among them

[0089]

[0090] Note: a All quantitative data (expressed as mean ± SEM) are from at least two independent experiments.

[0091] The present invention further analyzed the inhibitory effect of 1232030-35-1 on LCK signal transduction in T-ALL cell line Jurkat cells by immunoblotting. The results showed that after exposure to 1232030-35-1 at a concentration of 100 nM for 6 hours, the phosphorylation of the Y394 site of LCK was significantly reduced, which is a key modification for the full activation of LCK kinase activity. In addition, the phosphorylation of ZAP70 kinase and its adaptor LAT (two downstream molecules of LCK) also continuously decreased after treatment with 1232030-35-1. This indicates that the LCK signaling pathway in leukemia cells is indeed inhibited. Notably, as Figure 6A shown, immunoblots of total and phosphorylated LCK, total and phosphorylated ZAP70, and total and phosphorylated LAT in Jurkat cells treated with 302962-49-8 or 1232030-35-1 at the specified concentrations. Among them, HSP90 was used as a loading control. 1232030-35-1 showed similar kinase inhibitory activity to dasatinib (302962-49-8), a widely studied LCK inhibitor in T-ALL.

[0092] According to the results of the combined virtual screening and activity testing, the anti-leukemia effect of 1232030-35-1 was further evaluated. First, a cell proliferation assay was performed using the Cell Counting Kit-8 (CCK-8). As Figure 5B shown, 1232030-35-1 has a higher inhibitory efficiency on Jurkat cell viability, with an IC 50 value of 61.25 nM, while the IC 50The value is 1288 nM. As Figures 6C to 6E shown, Figure 6C The leukemia burden of T-ALL xenograft mice after drug administration was evaluated by bioluminescence imaging, and the time points after transplantation are shown above. Figure 6D At day 21, the luciferase signal was quantified in NPG mice inoculated with Jurkat-Luc cells. Figure 6E The Kaplan-Meier survival curves for T-ALL xenografts (n = 4 per group) are shown, where *p < 0.05; **p < 0.01; ***p < 0.001. Data showed that continuous daily intraperitoneal injection of 1232030-35-1 (25 mg / kg) more effectively inhibited the in vivo expansion of leukemia cells compared with injection of 302962-49-8 (25 mg / kg) or the vehicle control (as Figure 6C and Figure 6D shown). Finally, in vivo administration of 1232030-35-1 significantly prolonged the survival of mice in the leukemia xenograft model ( Figure 6E ). Specifically as follows:

[0093] IC 50 assay

[0094] The compound was initially dissolved in dimethyl sulfoxide to make a 10 mM stock solution, and then diluted 50-fold to the desired test concentration. A concentration range from 0.1 nM to 10,000 nM was formed by 10-fold serial dilution. The assay buffer included 5x buffer, 5 mM MgCl2, 1 mM DTT, and ddH2O, and was used to prepare two solutions: 2x ATP and substrate mixture, and 2x kinase and metal mixture. 25 nL of the compound solution was transferred to a 384-well plate using an Echo 655, and then 2.5 μL of the 2x kinase and metal mixture was added, and incubated at 25 °C for 10 minutes. Then, 2.5 μL of the 2x substrate and ATP solution was added, and incubated at 25 °C for 50 minutes. 2x XL665 and antibody mixture were prepared separately. After incubation, 5 μL of the kinase detection reagent was added and incubated at 25 °C for 60 minutes. Fluorescence emission was measured at 620 nm and 665 nm wavelengths using a microplate reader. Data were processed using GraphPad 7.0 software for dose-response analysis, and the IC 50 value was calculated using the specified equation:

[0095] Y = Bottom+(Top - Bottom) / (1 + 10^((LogIC 50 - X)*hillslope)).

[0096] Immunoblotting

[0097] Cells were lysed with RIPA buffer (NCM Biotech) containing protease inhibitor cocktail (ABclonal), and protein concentration was determined using a BCA assay kit (Thermo Fisher Scientific). 50 μg of total cellular protein was separated by SDS-PAGE and transferred to a polyvinylidene difluoride (PVDF) membrane (Millipore). After blocking with NcmBlot blocking buffer (NCM Biotech) for 40 minutes, the primary antibodies were added and the PVDF membrane was incubated overnight at 4 °C. Before detection with Clarity Western ECL Substrate (Bio-Rad), the PVDF membrane was incubated with an appropriate amount of horseradish peroxidase-conjugated secondary antibody for 1 hour at room temperature. Protein bands were visualized using a BLT GelView 6000Plus system (BLT). The antibodies used were as follows: LCK (CY7108, Abways Technology), p-LCK Y394 (CY6323, Abways Technology), ZAP70 (CY6937, Abways Technology), p-ZAP70 Y319 (#2701, Cell Signaling Technology), LAT (CY10662, Abways Technology), p-LAT Y191 (AY0148, Abways Technology), and HSP90 (sc-13119, Santa Cruz Biotechnology).

[0098] Cell culture and CCK8 assay

[0099] The human T-ALL cell line Jurkat was purchased from the American Type Culture Collection (ATCC) and cultured in RPMI-1640 (Gibco) supplemented with 10% fetal bovine serum (FBS, Gemini), 1% penicillin / streptomycin (Hyclone), 1% non-essential amino acids (Gibco), 2 mM L-glutamine (Sigma), 1 mM sodium pyruvate (Sigma), and 55 μM β-mercaptoethanol (Sigma). Cells were cultured in complete medium at 37 °C with 5% CO2.

[0100] Cell growth inhibition was detected using a Cell Counting Kit-8 (CCK-8, Dojindo Laboratories). First, Jurkat cells were counted and then 10,000 cells per well were seeded into a 96-well cell culture plate. Then, after culturing at 37 °C and 5% CO2 for 24 hours, a series of concentrations of 302962-49-8 (MedChemExpress, HPLC purity ≥99%) or 1232030-35-1 (MedChemExpress, HPLC purity ≥98%) were diluted with the corresponding culture medium and the culture medium was replaced. After co-culturing for 48 hours, 10 μl of CCK-8 reagent was added to each well, and after incubating at 37 °C for 1 hour, the optical density (OD) value at a wavelength of 450 nm was measured using a Sunrise multimode microplate reader (Tecan Life Sciences). The percentage of each concentration relative to the control group was used as the cell survival rate. The IC 50 value was calculated using Graphpad Prism 7.0.

[0101] Human T-ALL xenograft

[0102] NOD.Cg-Prkdcscid Il2rgtm1Vst / Vst (NPG) mice were purchased from Beijing Vitalstar Biotechnology Co., Ltd. Animal experiments were conducted in accordance with animal ethics regulations, and the research protocol was approved by the Animal Ethics Committee of Zhengzhou University. For the construction of T-ALL xenografts, Jurkat cells were first infected with a lentivirus expressing green fluorescent protein (GFP) and luciferase (pWPXLd-Luciferase-GFP). Then, one million sorted GFP-positive cells were intravenously injected into 6-week-old NPG mice that had received 1.5 Gy of irradiation. Treatment started on the 7th day after transplantation. 302962-49-8 or 1232030-35-1 (25 mg / kg) was intraperitoneally injected into the mice daily. D-Luciferin (150 mg / kg) was intraperitoneally injected and disease progression was evaluated by in vivo bioluminescence imaging (AmiHTX, Spectral Instruments Imaging).

[0103] In some embodiments, it further includes: performing simulation analysis and verification on the LCK inhibitory effect of the final target molecule by molecular dynamics.

[0104] Calculating the root mean square deviation (RMSD) of the protein-ligand complex can provide insights into the degree of conformational change relative to the initial conformation during the simulation, thereby observing the stability of the complex. In our study, we used The 2021-4 kit conducted in-depth simulations on LCK and 1232030-35-1. We extended the simulation time to 500 ns to ensure a comprehensive understanding of the system dynamics over a longer time scale. As we monitored the simulation, the RMSD values of the protein and ligand showed signs of stability in the last 10 ns (as Figure 7A shown). This indicates that our system has reached an equilibrium state, and any subsequent dynamics mainly represent fluctuations around this equilibrium state.

[0105] As Figure 7B and 7C shown, the data indicate that Met319 has a strong interaction with 1232030-35-1 at almost all observed time points. To further investigate this observation, we selected the final conformation for more detailed study. We plotted graphs to highlight the interactions between ligand atoms and specific protein residues ( Figures 7D to 7F ). Notably, our analysis shows that the 8-ethyl-2-(methylamino)pyrido[2,3-d]pyrimidin-7(8H)-one molecule of 1232030-35-1 forms a hydrogen bond with Met319.

[0106] To prevent recurrence, overcome drug resistance, and minimize the associated toxicity of intensive chemotherapy in T-ALL patients, there is an urgent need to develop new targeted drugs. In-depth understanding of T-ALL genetics, especially the widespread application of next-generation sequencing technologies, has opened the way for new treatment methods. Nevertheless, genomics-guided therapies have not been successfully translated into clinical practice for T-ALL to date. Recent studies have emphasized that in the molecular pathological evolution of T-ALL, in addition to driver gene mutations, complex RNA and protein signaling networks that maintain malignant proliferation and survival are also of great significance. Comprehensive proteomic analysis shows that the highly active SFK family (including LCK, SRC, LYN, and FYN, etc.) and serine / threonine kinases (CDK1 / 2, AKT, and PAK1 / 2) are specific dependence factors for the survival of T-ALL cells. Among them, LCK is the most active kinase in most primary T-ALL samples. Consistent with this, genome-scale CRISPR-Cas9 screening has also confirmed the dependence of T-ALL cells on LCK rather than other SFKs. Therefore, identifying potent LCK inhibitors has important clinical significance for future T-ALL treatment.

[0107] In the embodiments of the present invention, a comprehensive screening was conducted to identify potential LCK inhibitors. The present invention compared the efficacy of the deep learning algorithm PLANET based on graph neural networks and traditional docking tools. PLANET not only outperformed several mainstream screening tools in terms of efficiency but also showed performance comparable to Screen for comparable performance of the kit. The method of the present invention integrates PLANET with Glide SP of to generate conformations. This combined method successfully identified four effective LCK inhibitors from the MCE library of drugs. The novelty of these molecules was confirmed by the assessment of skeletal novelty. Compared with the existing LCK activity inhibitors, the similarity range of these four LCK inhibitors is between 0.2 and 0.4, indicating their significant novelty. The most prominent molecule is 1232030-35-1, whose IC 50 value is extremely low, only 0.43 nM. This successful identification demonstrates the effectiveness of our workflow, namely rapid screening by the deep learning scoring method and conformational selection using Glide. The strong screening results not only highlight the potential of 1232030-35-1 as an LCK inhibitor, but also represent the first experimental verification of the ability of PLANET to selectively screen for active molecules against specific targets. These results highlight the ability and practicality of integrating cutting-edge computational methods in the field of drug discovery.

[0108] Increasing evidence indicates that phosphorylation of LCK at the Y394 site is crucial for the full activation of its enzymatic activity. After TCR stimulation, LCK phosphorylates the immunoreceptor tyrosine-based activation motif (ITAM) in TCR, leading to sequential phosphorylation of ZAP70 at the Y319 and Y493 sites. After complete activation of ZAP70, it further phosphorylates the Y191 site of its key adaptor LAT to coordinate downstream signal cascades. Our protein phosphorylation analysis based on immunoblotting showed that after treatment with low concentrations of 1232030-35-1, the activation phosphorylation of LCK, ZAP70, and LAT in T-ALL cells was continuously inhibited, indicating that the pre-TCR-LCK signal transduction was globally inactivated. This was not only manifested as a strong anti-proliferative effect on T-ALL cells at the in vitro level but also showed good in vivo therapeutic effects on T-ALL mice. Encouragingly, we found that the anti-T-ALL efficacy of 1232030-35-1 in vitro and in vivo was significantly stronger than that of dasatinib (302962-49-8), which was previously proposed as an emerging potential inhibitor of LCK for T-ALL treatment. Notably, although the two compounds had similar inhibitory efficiencies on the phosphorylation of LCK signal transduction proteins, continuous administration of 1232030-35-1 could more significantly reduce the systemic leukemia burden and prolong the survival of experimental mice. Therefore, 1232030-35-1 is expected to show efficacy in patients with poor response to dasatinib monotherapy. Even more promising is that the combined use of the two may produce a more persistent and potent inhibition of LCK signals and coordinate the antagonistic leukemia effect. These data confirm the effectiveness of our screening workflow in isolating potent LCK inhibitors.

[0109] In addition, molecular dynamics simulations were performed in the embodiments of the present invention to precisely elucidate the binding kinetics of 1232030-35-1 to LCK, so as to clearly understand its interaction mechanism. The results showed that the 8-ethyl-2-(methylamino)pyrido[2,3-d]pyrimidin-7(8H)-one molecule of 1232030-35-1 formed a hydrogen bond with Met319, which was consistent with the binding behavior observed in known LCK inhibitors. This consistency in the interaction with known LCK inhibitors not only verified the inhibitory potential of 1232030-35-1 but also provided a structural basis for its mechanism of action. This information helps to guide future modification and optimization of this compound to improve its efficacy as an LCK inhibitor.

[0110] According to the embodiments of the present invention, Figure 8 is a schematic framework diagram of a virtual screening device for LCK inhibitors based on deep learning according to an exemplary embodiment of the present invention, as Figure 8As shown in the figure, an embodiment of the present invention further provides a virtual screening device for LCK inhibitors based on deep learning. The virtual screening device 800 for LCK inhibitors based on deep learning includes:

[0111] An acquisition and clustering module 802: used to acquire molecular data of the LCK target, clean the data, obtain a dedicated evaluation dataset for the LCK target, calculate the Morgan fingerprints of each molecule in the dedicated evaluation dataset for the LCK target and establish a similarity matrix, and cluster each molecule to form multiple clusters;

[0112] An acquisition and generation module 804: used to randomly select an active molecule set from each of the clusters and generate a decoy molecule set according to the active molecule set;

[0113] An evaluation and determination module 806: used to respectively acquire ligand small molecules and receptor proteins, and evaluate and determine a deep learning algorithm and a molecular docking program according to the active molecule set and the decoy molecule set;

[0114] A virtual screening module 808: respectively acquire a molecular database and a protein, perform virtual screening according to the determined deep learning algorithm and molecular docking program, and generate quasi-target molecules through the virtual screening;

[0115] A detection and evaluation module 810: used to detect the in vitro inhibitory activity of the quasi-target molecules against LCK, generate final target molecules; and evaluate the biological effects of the final target molecules.

[0116] It should be noted that the device embodiments in the embodiments of the present invention may refer to the above method embodiments and will not be elaborated here.

[0117] According to an embodiment of the present invention, the present invention further provides an electronic device, a readable storage medium, and a computer program product.

[0118] According to an embodiment of the present invention, the present invention provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method according to any one of the above.

[0119] According to an embodiment of the present invention, the present invention provides a computer program product, which includes: a computer program. The computer program is stored in a readable storage medium. At least one processor of the electronic device can read the computer program from the readable storage medium, and at least one processor executes the computer program to cause the electronic device to execute the solution provided in any one of the above embodiments.

[0120] According to an embodiment of the present invention, the present invention further provides an electronic device. FIG. 4 shows a schematic block diagram of an exemplary electronic device 900 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0121] As Figure 9 shown, the device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in the ROM 902 or a computer program loaded from a storage unit into the RAM 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. The I / O interface 905 is also connected to the bus 904.

[0122] A plurality of components in the device 400 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0123] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 executes the various methods and processes described above, such as the problem processing method. For example, in some embodiments, the problem processing method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the problem processing method described above can be executed. Alternatively, in other embodiments, the computing unit 901 can be configured to execute the problem processing method in any other suitable manner (e.g., by means of firmware).

[0124] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-a-chip systems (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0125] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0126] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0127] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube)) or LCD for displaying information to the user; and a keyboard and a pointing device by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback; and input from the user can be received in any form.

[0128] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components, or a computing system including frontend components, or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0129] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact via a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services. The server can also be a server of a distributed system, or a server combined with a blockchain.

[0130] Embodiments of the present invention identify potential LCK inhibitors and compare the efficacy of the deep learning algorithm PLANET based on graph neural networks with that of traditional docking tools. PLANET not only outperforms several mainstream screening tools in terms of efficiency but also exhibits performance comparable to that of the screening kit. The present invention integrates PLANET with Glide SP to reconstruct and generate conformations. This combined method successfully identifies four effective LCK inhibitors from the MCE class of drug molecule datasets. The novelty of these molecules is confirmed by skeletal novelty assessment. The similarity range of these four LCK inhibitors is between 0.2 - 0.4, indicating their significant novelty. Among them, the most prominent molecule is 1232030 - 35 - 1, whose IC 50 value is extremely low, only 0.43 nM. Through rapid screening by the deep learning scoring method and conformational selection using Glide, the strong screening results not only highlight the potential of 1232030 - 35 - 1 as an LCK inhibitor but also represent the first experimental verification of PLANET's ability to selectively screen active molecules against specific targets.

[0131] It is understood that the present invention is described by means of some embodiments. Those skilled in the art will know that, without departing from the spirit and scope of the present invention, various changes or equivalent substitutions can be made to these features and embodiments. Additionally, under the teachings of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the scope protected by the present invention.

Claims

1. A virtual screening method for LCK inhibitors based on deep learning, characterized in that: include: Acquiring molecular data of the LCK target, cleaning the data, and obtaining a dedicated data set for evaluation of the LCK target, calculating the Morgan fingerprint of each molecule in the dedicated data set for evaluation of the LCK target and establishing a similarity matrix, and clustering the molecules to form a plurality of clusters; Obtaining an active molecule set in each of the clusters, and generating a decoy molecule set based on the active molecule set; The ligand small molecules and the receptor proteins are obtained respectively, and the deep learning algorithm and the molecular docking program are evaluated and determined according to the active molecule set and the decoy molecule set, wherein: The deep learning algorithms evaluated include: PLANET deep learning algorithm, and the molecular docking programs evaluated include four molecular docking programs: AutoDock-GPU molecular docking program, Autodock Vina molecular docking program, LeDock molecular docking program, and The determined molecular docking program is Glide SP; the determined molecular docking program includes at least one of the four molecular docking programs evaluated; the determined deep learning algorithm is PLANET deep learning algorithm, and the determined molecular docking program is Glide SP molecular docking program; A molecular database and a protein are obtained respectively, and virtual screening is performed according to a determined deep learning algorithm and a molecular docking program, and a quasi-target molecule is obtained through the virtual screening, wherein: The virtual screening includes at least two steps, wherein the first virtual screening step is screening by using the PLANET deep learning algorithm, and the second virtual screening step is screening by using The Glide SP molecular docking program was used for another screening; The in vitro inhibitory activity of the quasi-target molecule on LCK is detected to obtain the final target molecule; and the biological effect of the final target molecule is evaluated.

2. The LCK inhibitor virtual screening method based on deep learning according to claim 1, characterized in that: The score for passing the first virtual screening step is set to be higher than 7, and the score for passing the second virtual screening step is set to be lower than -7.

3. The LCK inhibitor virtual screening method based on deep learning according to claim 2, characterized in that: The quasi-target molecules are respectively a first quasi-target molecule, a second quasi-target molecule, a third quasi-target molecule, and a fourth quasi-target molecule.

4. The LCK inhibitor virtual screening method based on deep learning according to claim 3, characterized in that: The third quasi-target molecule is the final target molecule, and its chemical formula is:

Citation Information

Patent Citations

  • Method for improving virtual screening capability of docking software based on machine learning algorithm

    CN111402967A

  • Compound library construction method and device based on artificial intelligence, equipment and storage medium

    CN113436686A