Screening Method, Device, Computer Equipment, Storage Medium and Program Product

Through the screening method of multi-level prediction model, the problem of low accuracy of virtual drug screening in the prior art is solved, the screening efficiency and accuracy are improved, and the drug discovery process is promoted.

CN116994672BActive Publication Date: 2025-06-24TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310912286.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-24
Publication Date
2025-06-24
Estimated Expiration
2043-07-24

AI Technical Summary

Technical Problem

In the prior art, the accuracy of the virtual screening method of drugs is low, resulting in low efficiency in the screening task.

Method used

Multi-level screening is performed by inputting the structural information of the molecule into the cell activity prediction model, the main protease activity prediction model and the ADMET prediction model to improve the accuracy and efficiency of the screening.

Benefits of technology

It improves the accuracy and efficiency of drug screening, ensures that the selected target molecules have high biological activity and pharmacokinetic properties, and thus accelerates the drug discovery process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116994672B_ABST
    Figure CN116994672B_ABST
Patent Text Reader

Abstract

The present application relates to a screening method, apparatus, computer device, storage medium, and computer program product. The method includes: obtaining a plurality of molecules from a molecular library and obtaining the structural information of the plurality of molecules; inputting the structural information of the plurality of molecules into a cell activity prediction model, and performing a first screening process on the plurality of molecules based on the output of the cell activity prediction model to obtain a first molecule; inputting the structural information of the first molecule into a main protease activity prediction model, and performing a second screening process on the first molecule based on the output of the main protease activity prediction model to obtain a second molecule; inputting the structural information of the second molecule into an ADMET prediction model, and performing a third screening process on the second molecule based on the output of the ADMET prediction model to obtain a third molecule; determining a target molecule based on the third molecule. Using this method can effectively improve the working efficiency of the screening task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of virtual drug screening, and particularly to a screening method, device, computer device, storage medium, and computer program product. Background Art

[0002] Drug screening is an important way to discover drug lead compounds. The complexity of drug screening lies in the huge parameter space, that is, researchers need to screen out small molecule lead compounds against target proteins from a huge molecular library.

[0003] In the prior art, the above screening process is usually implemented based on virtual screening technology, that is, by computer simulating the interaction between the target and small molecules in the molecular library, calculating relevant data, and screening small molecules based on the relevant data.

[0004] However, the screening method based on traditional virtual screening technology has low accuracy, which in turn leads to low working efficiency of the screening task. Summary of the Invention

[0005] Based on this, it is necessary to provide a screening method, device, computer device, computer-readable storage medium, and computer program product with relatively high working efficiency for the above technical problems.

[0006] In a first aspect, the present application provides a screening method. The method includes:

[0007] Obtain a plurality of molecules from a molecular library, and obtain the structural information of the plurality of molecules; input the structural information of the plurality of molecules into a cell activity prediction model, and perform a first screening process on the plurality of molecules based on the output of the cell activity prediction model to obtain a first molecule; input the structural information of the first molecule into a main protease activity prediction model, and perform a second screening process on the first molecule based on the output of the main protease activity prediction model to obtain a second molecule; input the structural information of the second molecule into an ADMET prediction model, and perform a third screening process on the second molecule based on the output of the ADMET prediction model to obtain a third molecule; determine a target molecule based on the third molecule.

[0008] In one embodiment, before inputting the structural information of the plurality of molecules into the cell activity prediction model, the method further includes: obtaining the structural information of molecules with known activity and the inhibition rate of molecules with known activity; determining similarity information according to the structural information of molecules with known activity and the structural information of the plurality of molecules; determining the SA score of the plurality of molecules based on the similarity information and the inhibition rate of molecules with known activity; performing an initial screening process on the plurality of molecules according to the SA score of the plurality of molecules and a preset score threshold.

[0009] In one embodiment, inputting the structural information of the plurality of molecules into a cell activity prediction model, and performing a first screening process on the plurality of molecules based on the output of the cell activity prediction model to obtain a first molecule, includes: inputting the structural information of the plurality of molecules into the cell activity prediction model to obtain the cell activity information corresponding to the plurality of molecules; and performing the first screening process on the plurality of molecules based on the cell activity information and a preset cell activity threshold to obtain the first molecule.

[0010] In one embodiment, inputting the structural information of the first molecule into a main protease activity prediction model, and performing a second screening process on the first molecule based on the output of the main protease activity prediction model to obtain a second molecule, includes: inputting the structural information of the first molecule into the main protease activity prediction model to obtain the main protease activity information corresponding to the first molecule; determining the affinity information between the first molecule and a target based on molecular docking technology and the main protease activity information; and performing the second screening process on the first molecule based on the main protease activity information, the affinity information, a preset main protease activity threshold, and a preset affinity threshold to obtain the second molecule.

[0011] In one embodiment, determining a target molecule based on the third molecule includes: determining the covalent binding ability information of the third molecule based on quantum chemistry methods; performing a fourth screening process on the third molecule according to the covalent binding ability information and a preset covalent binding ability threshold to obtain a fourth molecule; and determining the target molecule based on the fourth molecule.

[0012] In one embodiment, determining the target molecule based on the fourth molecule includes: performing a clustering process on the fourth molecule to obtain a plurality of molecule groups; for each molecule group, determining one molecule as a representative molecule according to a preset rule, and determining a plurality of representative molecules based on the plurality of molecule groups; and determining the plurality of representative molecules as the target molecules.

[0013] In a second aspect, the present application also provides a screening device. The device includes:

[0014] An acquisition module, configured to acquire a plurality of molecules from a molecular library and acquire the structural information of the plurality of molecules;

[0015] A first execution module, configured to input the structural information of the plurality of molecules into a cell activity prediction model, and perform a first screening process on the plurality of molecules based on the output of the cell activity prediction model to obtain a first molecule;

[0016] A second execution module, configured to input the structural information of the first molecule into a main protease activity prediction model, and perform a second screening process on the first molecule based on the output of the main protease activity prediction model to obtain a second molecule;

[0017] A third execution module, configured to input the structural information of the second molecule into an ADMET prediction model, and perform a third screening process on the second molecule based on the output of the ADMET prediction model to obtain a third molecule;

[0018] A determination module, configured to determine a target molecule based on the third molecule.

[0019] In one embodiment, the screening device further includes a fourth execution module, and the fourth execution module is configured to: obtain the structural information of a molecule with known activity and the inhibition rate of the molecule with known activity; determine similarity information according to the structural information of the molecule with known activity and the structural information of the multiple molecules; determine the SA score of the multiple molecules based on the similarity information and the inhibition rate of the molecule with known activity; perform an initial screening process on the multiple molecules according to the SA scores of the multiple molecules and a preset score threshold.

[0020] In one embodiment, the first execution module is specifically configured to: input the structural information of the multiple molecules into the cell activity prediction model to obtain the cell activity information corresponding to the multiple molecules; perform the first screening process on the multiple molecules based on the cell activity information and a preset cell activity threshold to obtain the first molecule.

[0021] In one embodiment, the second execution module is specifically configured to: input the structural information of the first molecule into the main protease activity prediction model to obtain the main protease activity information corresponding to the first molecule; determine the affinity information between the first molecule and the target based on molecular docking technology and the main protease activity information; perform the second screening process on the first molecule based on the main protease activity information, the affinity information, a preset main protease activity threshold, and a preset affinity threshold to obtain the second molecule.

[0022] In one embodiment, the third execution module is specifically configured to: determine the covalent binding ability information of the third molecule based on quantum chemistry; perform a fourth screening process on the third molecule according to the covalent binding ability information and a preset covalent binding ability threshold to obtain a fourth molecule; determine the target molecule based on the fourth molecule.

[0023] In one embodiment, the third execution module is specifically configured to: perform clustering processing on the fourth molecule to obtain multiple molecule groups; for each molecule group, determine one molecule as a representative molecule according to a preset rule, and determine multiple representative molecules based on the multiple molecule groups; determine the multiple representative molecules as the target molecules.

[0024] In a third aspect, the present application also provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps described in any one of the above first aspects are implemented.

[0025] In a fourth aspect, the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, the steps described in any one of the above first aspects are implemented.

[0026] In a fifth aspect, the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps described in any one of the above first aspects are implemented.

[0027] The above screening method, device, computer device, storage medium, and computer program product obtain a plurality of molecules from a molecular library and obtain the structural information of the plurality of molecules; input the structural information of the plurality of molecules into a cell activity prediction model, and perform a first screening process on the plurality of molecules based on the output of the cell activity prediction model to obtain a first molecule; input the structural information of the first molecule into a main protease activity prediction model, and perform a second screening process on the first molecule based on the output of the main protease activity prediction model to obtain a second molecule; input the structural information of the second molecule into an ADMET prediction model, and perform a third screening process on the second molecule based on the output of the ADMET prediction model to obtain a third molecule; determine a target molecule based on the third molecule. The screening method provided by the present application first performs a first screening process on a plurality of molecules based on the output of the cell activity prediction model to obtain a first molecule, then performs a second screening process on the first molecule based on the output of the main protease activity prediction model to obtain a second molecule, then performs a third screening process on the second molecule based on the output of the ADMET prediction model to obtain a third molecule, and finally determines a target molecule based on the third molecule. The screening method provided by the present application screens molecules based on the outputs of multiple models to obtain a target molecule. Since different models can predict different properties of molecules, that is, screen molecules based on multiple dimensions, therefore, the target molecule obtained by using the screening method provided by the present application has a high accuracy, thereby effectively improving the working efficiency of the screening task. Description of the Drawings

[0028] Figure 1 It is a schematic flowchart of the screening method in an embodiment;

[0029] Figure 2 It is a schematic flowchart of the method for initially screening a plurality of molecules in an embodiment;

[0030] Figure 3Schematic flowchart of a first screening process for the plurality of molecules based on the output of the cell activity prediction model in one embodiment to obtain the first molecule

[0031] Figure 4 Schematic flowchart of a second screening process for the first molecule based on the output of the main protease activity prediction model in one embodiment to obtain the second molecule

[0032] Figure 5 Schematic flowchart of a method for determining the target molecule based on the third molecule in one embodiment

[0033] Figure 6 Schematic flowchart of a method for determining the target molecule based on the fourth molecule in one embodiment

[0034] Figure 7 Schematic flowchart of another screening method in one embodiment

[0035] Figure 8 Structure block diagram of a screening device in one embodiment

[0036] Figure 9 Structure block diagram of another screening device in one embodiment

[0037] Figure 10 Internal structure diagram of a computer device in one embodiment

[0038] Figure 11 Schematic diagram of the ROC curve of the cell activity prediction model in one embodiment

[0039] Figure 12 Schematic diagram of the model integration effect of the cell activity prediction model in one embodiment

[0040] Figure 13 Schematic diagram of the ROC curve of the main protease activity prediction model in one embodiment

[0041] Figure 14 Schematic diagram of the model integration effect of the main protease activity prediction model in one embodiment

[0042] Figure 15 Schematic diagram of the proportion of inactive molecules under different preset score thresholds when k = 15 in one embodiment Detailed implementation manners

[0043] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0044] Drug screening is an important way to discover drug lead compounds. The complexity of drug screening lies in the huge parameter space, that is, researchers need to screen out small molecule lead compounds targeting the target protein from a huge molecular library.

[0045] In the prior art, the above screening process is usually realized based on virtual screening technology, that is, by computer simulating the interaction between the target and small molecules in the molecular library, calculating relevant data, and screening small molecules based on the relevant data.

[0046] However, the screening method based on traditional virtual screening technology has low accuracy and slow calculation speed, which leads to low work efficiency of the screening task.

[0047] In view of this, the embodiments of the present application provide a screening method, which can effectively improve the work efficiency of the screening task.

[0048] The screening method provided by the embodiments of the present application may be executed by a computer device, and the computer device may be a server.

[0049] In one embodiment, as Figure 1 shown, a screening method is provided, and the method includes the following steps:

[0050] Step 101, obtain a plurality of molecules from the molecular library, and obtain the structural information of the plurality of molecules.

[0051] Optionally, the molecular library refers to a collection composed of entity compounds with specific structures or functions and their related information under a certain specific standard, and is usually used as a tool for new drug discovery or cell induction.

[0052] In one possible implementation, the molecular library may be a bioactive molecular library.

[0053] In another possible implementation, the molecular library may also be a natural product molecular library.

[0054] In another possible implementation, the molecular library may also be a drug-like diversity molecular library.

[0055] Optionally, the structural information refers to the description of the three-dimensional arrangement of atoms in the molecule. The molecular structure largely affects the reactivity, polarity, phase state, color, magnetism and biological activity of chemical substances.

[0056] In one possible implementation, the structural information of the plurality of molecules may be obtained based on a structural information database.

[0057] Step 102: Input the structural information of the multiple molecules into the cell activity prediction model, and perform a first screening process on the multiple molecules based on the output of the cell activity prediction model to obtain the first molecule.

[0058] Optionally, the cell activity prediction model refers to a model that can predict the cell activity of molecules. The cell activity prediction model includes the integrated language representation model BERT, the geometric conformation enhanced AI algorithm model GEM, and the three-dimensional molecular pre-training model Uni-Mol.

[0059] In an optional embodiment of the present application, the initial BERT model, GEM model, and Uni-Mol model are trained respectively based on the labeled cell activity dataset and the method of random undersampling, and the trained models are integrated by the Bagging method to obtain the cell activity prediction model. Since the cell activity prediction model is an integrated model, the variance of the model is reduced, thereby effectively improving the stability. As Figure 11 and Figure 12 shown, Figure 11 shows the ROC curve and AUC value of the cell activity prediction model. Figure 11 The solid line in Figure 11 is the ROC curve, Figure 12 and the dotted line in Figure 12 is the random guess line, which shows the model integration effect of the cell activity prediction model. The dotted line in

[0060] In a possible implementation manner, the structural information of the multiple molecules can be input into the cell activity prediction model. The cell activity prediction model will output the cell activity information corresponding to the multiple molecules based on the structural information of the multiple molecules. The multiple molecules are screened according to the cell activity information, and the molecules with poor cell activity in the multiple molecules are excluded to obtain the first molecule.

[0061] Step 103: Input the structural information of the first molecule into the main protease activity prediction model, and perform a second screening process on the first molecule based on the output of the main protease activity prediction model to obtain the second molecule.

[0062] Optionally, the main protease activity prediction model refers to a model that can predict the main protease activity of molecules. The main protease prediction model includes the integrated language representation model BERT, the geometric conformation enhanced AI algorithm model GEM, and the three-dimensional molecular pre-training model Uni-Mol.

[0063] In an alternative embodiment of the present application, the initial BERT model, GEM model, and Uni-Mol model are respectively trained based on the labeled main protease activity dataset and the method of random undersampling, and the trained models are integrated by the Bagging method to obtain the main protease activity prediction model. Since the main protease activity prediction model is an integrated model, the variance of the model is reduced, thereby effectively improving the stability. As Figure 13 and Figure 14 shown, Figure 13 the ROC curve and AUC value of the main protease activity prediction model are shown. Figure 13 The solid line in Figure 13 is the ROC curve, Figure 14 and the dashed line in Figure 14 is the random guess line, which shows the model integration effect of the main protease activity prediction model. The dashed line in

[0064] is the best score (the mean of AP and AUC) among all models. Compared with the individual models, the main protease activity prediction model integrated by the hard voting-based Baggging method has better performance on the validation set.

[0065] Step 104: Input the structural information of the second molecule into the ADMET prediction model, and perform a third screening process on the second molecule based on the output of the ADMET prediction model to obtain a third molecule.

[0066] Optionally, the ADMET prediction model can be ADMETlab2.0, which is an online prediction platform based on comprehensive pharmacokinetics and toxicity.

[0067] In a possible implementation manner, the structural information of the second molecule can be input into the ADMET prediction model to obtain various ADMET property information of the second molecule, such as Caco-2, IGC50, LogS, TPSA, LC50, mw, LogP, LC50DM, QED, and nRig. Based on the above properties, the toxicity, metabolic ability, and permeability of the second molecule are determined. Based on the preset toxicity threshold, preset metabolic ability threshold, and preset permeability threshold, the second molecule is screened, and the molecules with strong toxicity, poor metabolic ability, and poor permeability in the second molecule are removed to obtain a third molecule.

[0068] Step 105: Determine the target molecule based on the third molecule.

[0069] In a possible implementation manner, the third molecule can be directly determined as the target molecule.

[0070] In another possible implementation manner, the covalent binding ability information of the third molecule can also be obtained, and the third molecule is screened based on the covalent binding ability information, and the molecule obtained after the screening process is determined as the target molecule.

[0071] In another possible implementation manner, the third molecule can also be clustered to obtain multiple molecule groups. For each molecule group, one molecule is determined as the representative molecule according to a preset rule, multiple representative molecules are determined based on the multiple molecule groups, and the multiple representative molecules are determined as the target molecules.

[0072] In another possible implementation manner, the covalent binding ability information of the third molecule can also be obtained, the third molecule is screened based on the covalent binding ability information, and the screened third molecule is clustered to obtain multiple molecule groups. For each molecule group, one molecule is determined as the representative molecule according to a preset rule, multiple representative molecules are determined based on the multiple molecule groups, and the multiple representative molecules are determined as the target molecules.

[0073] The above screening method obtains multiple molecules from a molecular library and obtains the structural information of the multiple molecules; inputs the structural information of the multiple molecules into a cell activity prediction model, and performs a first screening process on the multiple molecules based on the output of the cell activity prediction model to obtain a first molecule; inputs the structural information of the first molecule into a main protease activity prediction model, and performs a second screening process on the first molecule based on the output of the main protease activity prediction model to obtain a second molecule; inputs the structural information of the second molecule into an ADMET prediction model, and performs a third screening process on the second molecule based on the output of the ADMET prediction model to obtain a third molecule; determines the target molecule based on the third molecule. The screening method provided in this application first performs a first screening process on multiple molecules based on the output of the cell activity prediction model to obtain a first molecule, then performs a second screening process on the first molecule based on the output of the main protease activity prediction model to obtain a second molecule, then performs a third screening process on the second molecule based on the output of the ADMET prediction model to obtain a third molecule, and finally determines the target molecule based on the third molecule. The screening method provided in this application screens molecules based on the outputs of multiple models to obtain the target molecule. Since different models can predict different properties of molecules, that is, screening molecules based on multiple dimensions, therefore, the target molecule obtained by using the screening method provided in this application has a high accuracy, thereby effectively improving the working efficiency of the screening task.

[0074] In one embodiment, as Figure 2 shown, before inputting the structural information of the plurality of molecules into the cell activity prediction model, the method further includes the following steps:

[0075] Step 201, obtain the structural information of the molecules with known activity and the inhibition rate of the molecules with known activity.

[0076] Optionally, the molecules with known activity refer to the molecules with known biological activity.

[0077] In a possible implementation manner, the molecules with known activity can be obtained based on an active compound library, the structural information of the molecules with known activity can be obtained based on a structural information database, and the inhibition rate of the molecules with known activity can be determined according to the molecules with known activity.

[0078] Step 202, determine similarity information according to the structural information of the molecules with known activity and the structural information of the plurality of molecules.

[0079] In a possible implementation manner, for each molecule in the plurality of molecules, based on the Tanimoto coefficient, the similarity information between each molecule in the molecules with known activity and the structural information of this molecule is determined, and the above process is performed for each molecule in the plurality of molecules to obtain the similarity information of each molecule in the plurality of molecules.

[0080] Step 203, determine the SA score of the plurality of molecules based on the similarity information and the inhibition rate of the molecules with known activity.

[0081] In a possible implementation manner, the SA score calculation formula is:

[0082]

[0083] where n is the total number of molecules with known activity, k is a hyperparameter, which can be set according to different data sets. The larger the k value, the higher the reference weight of the highly similar molecules. Similarity is the similarity information of each molecule in the plurality of molecules, and Inhibition is the inhibition rate of the molecules with known activity.

[0084] As described above, k is a hyperparameter, so it is necessary to determine the optimal value of k. In a possible implementation manner, the present application constructs a scoring function, and determines the optimal value of k according to the scoring function. The scoring function is:

[0085]

[0086] where, Mean(SA active) refers to the average SA score of known active molecules, Mean(SA inactive ) refers to the average SA score of inactive molecules, SA max refers to the maximum SA score among all molecules, SA min refers to the minimum SA score among all molecules.

[0087] This scoring function is used to calculate the difference between the average SA value of active molecules and the average SA value of inactive molecules among known active molecules. The higher the score, the better the current value of k.

[0088] Step 204: Perform an initial screening process on the multiple molecules according to the SA scores of the multiple molecules and a preset score threshold.

[0089] Optionally, the preset score threshold can be set in advance by a technician according to actual needs.

[0090] In a possible implementation manner, the preset score threshold can be determined by calculating the SA value of each known active molecule, and the multiple molecules are screened according to the preset score threshold to eliminate inactive molecules. As Figure 15 shown, it is the proportion of inactive molecules under different preset score thresholds when the value of k is 15. Figure 15 The X-axis represents the SA score, and the Y-axis represents the proportion of inactive molecules under different preset score thresholds. It can be seen from Figure 15 that when the SA score of a molecule among the multiple molecules is lower than -35, the structure of this molecule is extremely similar to that of inactive molecules, and there is a 97% probability that this molecule is also inactive. Therefore, this type of molecule can be eliminated in advance, thereby effectively reducing the workload of subsequent screening tasks and effectively improving work efficiency.

[0091] In one embodiment, as Figure 3 shown, inputting the structural information of the multiple molecules into a cell activity prediction model, and performing a first screening process on the multiple molecules based on the output of the cell activity prediction model to obtain a first molecule, including the following steps:

[0092] Step 301: Input the structural information of the multiple molecules into the cell activity prediction model to obtain the cell activity information corresponding to the multiple molecules.

[0093] Step 302: Perform the first screening process on the multiple molecules based on the cell activity information and a preset cell activity threshold to obtain the first molecule.

[0094] Optionally, the preset cell activity threshold can be set by a technician according to actual needs.

[0095] In a possible implementation manner, for each of the multiple molecules, the cellular activity of the molecule is determined according to the cellular activity information of the molecule. If the cellular activity of the molecule is greater than the preset cellular activity threshold, the molecule is retained; if the cellular activity of the molecule is less than or equal to the preset cellular activity threshold, the molecule is excluded. The above process is performed on each of the multiple molecules, and finally the molecules that are retained are determined as the first molecules.

[0096] As described above, the structural information of the multiple molecules is input into the cellular activity prediction model to obtain the cellular activity information corresponding to the multiple molecules, and then based on the cellular activity information and the preset cellular activity threshold, the first screening process is performed on the multiple molecules to obtain the first molecules. The cellular activity prediction model has been described above. Since the cellular activity prediction model has good stability, the cellular activity information obtained based on the cellular activity prediction model is highly reliable, so that the result accuracy of the first screening process is higher, and thus the accuracy of the screening task is effectively improved.

[0097] In one embodiment, as Figure 4 shown, the structural information of the first molecule is input into the main protease activity prediction model, and based on the output of the main protease activity prediction model, the second screening process is performed on the first molecule to obtain the second molecule, including the following steps:

[0098] Step 401: Input the structural information of the first molecule into the main protease activity prediction model to obtain the main protease activity information corresponding to the first molecule.

[0099] Step 402: Based on the molecular docking technology and the main protease activity information, determine the affinity information between the first molecule and the target.

[0100] Optionally, the molecular docking technology refers to the process of mutually recognizing two or more molecules through geometric matching and energy matching to find the best matching mode.

[0101] In a possible implementation manner, the molecular docking technology is implemented based on software. The affinity between the first molecule and the target is scored based on the molecular docking technology, and the affinity information is determined based on the score.

[0102] Step 403: Based on the main protease activity information, the affinity information, the preset main protease activity threshold, and the preset affinity threshold, perform the second screening process on the first molecule to obtain the second molecule.

[0103] Optionally, the preset main protease activity threshold and the preset affinity threshold can be preset by those skilled in the art according to actual needs.

[0104] In a possible implementation, for each molecule in the first molecule, the main protease activity and affinity of the molecule are determined according to the main protease activity information and affinity information of the molecule. If the main protease activity of the molecule is greater than the preset main protease activity threshold and the affinity of the molecule is greater than the preset affinity threshold, the molecule is retained. If the main protease activity of the molecule is less than or equal to the preset main protease activity threshold, or the affinity of the molecule is less than or equal to the preset affinity threshold, the molecule is excluded. The above process is performed on each molecule in the first molecule, and finally the molecules that are retained are determined as the second molecule.

[0105] As described above, the structural information of the first molecule is input into the main protease activity prediction model to obtain the main protease activity information corresponding to the first molecule. The affinity information between the first molecule and the target is determined based on the molecular docking technology and the main protease activity information. Based on the main protease activity information, the affinity information, the preset main protease activity threshold, and the preset affinity threshold, the first molecule is subjected to the second screening process to obtain the second molecule. The main protease activity prediction model has been described above. Since the main protease activity prediction model has good stability, the main protease activity information obtained based on the main protease activity prediction model is highly reliable, so that the result of the second screening process is more accurate, thereby effectively improving the accuracy of the screening task.

[0106] In one embodiment, as Figure 5 shown, determining the target molecule based on the third molecule includes the following steps:

[0107] Step 501: Determine the covalent binding ability information of the third molecule based on quantum chemistry.

[0108] Optionally, the covalent binding ability information is used to characterize the covalent binding ability of each molecule in the third molecule to the main protease.

[0109] In a possible implementation, first calculate the proton affinity ability information between the cyano group in each molecule of the third molecule and Cys145 of the main protease based on quantum chemistry, and then determine the covalent binding ability of each molecule in the third molecule to the main protease according to the proton affinity ability information.

[0110] Step 502: Perform a fourth screening process on the three molecules according to the covalent binding ability information and the preset covalent binding ability threshold to obtain the fourth molecule.

[0111] Optionally, the preset covalent binding ability threshold can be set in advance by a technician according to actual needs.

[0112] In a possible implementation, for each molecule in the third molecules, the covalent binding ability of the molecule is determined according to the covalent binding ability information of the molecule. If the covalent binding ability of the molecule is greater than the preset covalent binding ability threshold, the molecule is retained. If the covalent binding ability of the molecule is less than or equal to the preset covalent binding ability threshold, the molecule is excluded. The above process is performed on each molecule in the third molecules, and finally the retained molecules are determined as the fourth molecules.

[0113] Step 503: Determine the target molecule based on the fourth molecules.

[0114] In one embodiment, as Figure 6 shown, the determining the target molecule based on the fourth molecules includes the following steps:

[0115] Step 601: Perform clustering processing on the fourth molecules to obtain multiple molecule groups.

[0116] Optionally, the clustering processing refers to a method of dividing a data set into different classes or clusters according to a specific criterion, so that the similarity of data objects within the same class or cluster is as large as possible, and the difference of data objects not in the same class or cluster is also as large as possible.

[0117] In a possible implementation, the Butina algorithm in rdkit can be used to perform clustering processing on the fourth molecules to obtain multiple molecule groups of different classes.

[0118] Step 602: For each molecule group, determine one molecule as the representative molecule according to the preset rule, and determine multiple representative molecules based on the multiple molecule groups.

[0119] Step 603: Determine the multiple representative molecules as the target molecules.

[0120] In a possible implementation, for each molecule in each molecule group, obtain the SA score, cell activity information, main protease activity information, covalent binding ability information, and ADMET property information of the molecule. Determine the comprehensive score of the molecule based on the SA score, cell activity information, main protease activity information, covalent binding ability information, and ADMET property information of the molecule. Perform the above process on each molecule in the molecule group, and finally determine the representative molecule in the molecule group according to the comprehensive score of each molecule. Perform the above process on each molecule group to determine the representative molecule of each molecule group. After determining multiple representative molecules based on multiple molecule groups, determine the multiple representative molecules as the target molecules.

[0121] The above-mentioned clustering process for the fourth molecule is performed to obtain multiple molecule groups. For each molecule group, one molecule is determined as a representative molecule according to a preset rule. Multiple representative molecules are determined based on the multiple molecule groups, and the multiple representative molecules are determined as the target molecules. The method of selecting representative molecules from the molecule groups of each class and using the representative molecules as the target molecules, and the target molecules will be used to test the developed drugs. Therefore, selecting representative molecules from the molecule groups of each class for testing can avoid the molecular structures used for testing being too similar, effectively improving the working efficiency of the screening task.

[0122] In one embodiment, as Figure 7 shown, another screening method is provided. The method includes the following steps:

[0123] Step 701: Obtain multiple molecules from the molecular library and obtain the structural information of the multiple molecules; obtain the structural information of the molecules with known activity and the inhibition rate of the molecules with known activity; determine the similarity information based on the structural information of the molecules with known activity and the structural information of the multiple molecules; determine the SA scores of the multiple molecules based on the similarity information and the inhibition rate of the molecules with known activity; perform an initial screening process on the multiple molecules according to the SA scores of the multiple molecules and a preset score threshold.

[0124] Step 702: Input the structural information of the multiple molecules into the cell activity prediction model to obtain the cell activity information corresponding to the multiple molecules; perform the first screening process on the multiple molecules based on the cell activity information and a preset cell activity threshold to obtain the first molecule.

[0125] Step 703: Input the structural information of the first molecule into the main protease activity prediction model to obtain the main protease activity information corresponding to the first molecule; determine the affinity information between the first molecule and the target based on molecular docking technology and the main protease activity information; perform the second screening process on the first molecule based on the main protease activity information, the affinity information, a preset main protease activity threshold, and a preset affinity threshold to obtain the second molecule.

[0126] Step 704: Input the structural information of the second molecule into the ADMET prediction model and perform a third screening process on the second molecule based on the output of the ADMET prediction model to obtain the third molecule.

[0127] Step 705: Determine the covalent binding ability information of the third molecule based on quantum chemistry method; perform a fourth screening process on the three molecules according to the covalent binding ability information and a preset covalent binding ability threshold to obtain a fourth molecule; perform a clustering process on the fourth molecule to obtain multiple molecule groups; for each molecule group, determine one molecule as a representative molecule according to a preset rule, and determine multiple representative molecules based on the multiple molecule groups; determine the multiple representative molecules as target molecules.

[0128] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0129] Based on the same inventive concept, the embodiments of the present application further provide a screening device for implementing the above-mentioned screening method. The solution provided by this device for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more of the following screening device embodiments can refer to the limitations on the screening method in the above text, and will not be repeated here.

[0130] In one embodiment, as Figure 8 shown, a screening device 800 is provided, including: an acquisition module 801, a first execution module 802, a second execution module 803, a third execution module 804, and a determination module 805, where:

[0131] The acquisition module 801 is configured to acquire multiple molecules from a molecular library and acquire the structural information of the multiple molecules;

[0132] The first execution module 802 is configured to input the structural information of the multiple molecules into a cell activity prediction model, and perform a first screening process on the multiple molecules based on the output of the cell activity prediction model to obtain a first molecule;

[0133] The second execution module 803 is configured to input the structural information of the first molecule into a main protease activity prediction model, and perform a second screening process on the first molecule based on the output of the main protease activity prediction model to obtain a second molecule;

[0134] The third execution module 804 is configured to input the structural information of the second molecule into the ADMET prediction model, and perform a third screening process on the second molecule based on the output of the ADMET prediction model to obtain a third molecule;

[0135] The determination module 805 is configured to determine a target molecule based on the third molecule.

[0136] In one embodiment, as Figure 9 shown, another screening device 900 is further provided. In addition to including each module included in the screening device 800, the screening device 900 further includes a fourth execution module 806. The fourth execution module 806 is configured to: obtain the structural information of the molecule with known activity and the inhibition rate of the molecule with known activity; determine similarity information according to the structural information of the molecule with known activity and the structural information of the multiple molecules; determine the SA score of the multiple molecules based on the similarity information and the inhibition rate of the molecule with known activity; perform an initial screening process on the multiple molecules according to the SA scores of the multiple molecules and a preset score threshold.

[0137] In one embodiment, the first execution module 802 is specifically configured to: input the structural information of the multiple molecules into the cell activity prediction model to obtain the cell activity information corresponding to the multiple molecules; perform the first screening process on the multiple molecules based on the cell activity information and a preset cell activity threshold to obtain the first molecule.

[0138] In one embodiment, the second execution module 803 is specifically configured to: input the structural information of the first molecule into the main protease activity prediction model to obtain the main protease activity information corresponding to the first molecule; determine the affinity information between the first molecule and the target based on molecular docking technology and the main protease activity information; perform the second screening process on the first molecule based on the main protease activity information, the affinity information, a preset main protease activity threshold, and a preset affinity threshold to obtain the second molecule.

[0139] In one embodiment, the third execution module 804 is specifically configured to: determine the covalent binding ability information of the third molecule based on quantum chemistry; perform a fourth screening process on the three molecules according to the covalent binding ability information and a preset covalent binding ability threshold to obtain a fourth molecule; determine the target molecule based on the fourth molecule.

[0140] In one embodiment, the third execution module 804 is specifically configured to: perform a clustering process on the fourth molecule to obtain multiple molecule groups; for each molecule group, determine one molecule as a representative molecule according to a preset rule, and determine multiple representative molecules based on the multiple molecule groups; determine the multiple representative molecules as the target molecules.

[0141] Each module in the above screening device can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in or independent of a processor in a computer device in the form of hardware, or stored in a memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.

[0142] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structural diagram can be as Figure 10 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a screening method.

[0143] Those skilled in the art can understand that Figure 10 the structure shown in

[0144] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0145] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, any step in the above embodiment is implemented.

[0146] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, any step in the above embodiment is implemented.

[0147] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories.

[0148] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0149] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A screening method, characterized in that, The method includes: Obtaining a plurality of molecules from a molecular library and obtaining the structural information of the plurality of molecules; Inputting the structural information of the plurality of molecules into a cell activity prediction model, and performing a first screening process on the plurality of molecules based on the output of the cell activity prediction model to obtain a first molecule; Inputting the structural information of the first molecule into a main protease activity prediction model, and performing a second screening process on the first molecule based on the output of the main protease activity prediction model to obtain a second molecule; Inputting the structural information of the second molecule into an ADMET prediction model, and performing a third screening process on the second molecule based on the output of the ADMET prediction model to obtain a third molecule; Determining the covalent binding ability information of the third molecule based on quantum chemistry methods; performing a fourth screening process on the third molecule according to the covalent binding ability information and a preset covalent binding ability threshold to obtain a fourth molecule; Performing clustering processing on the fourth molecule to obtain a plurality of molecule groups; for each molecule group, determining one molecule as a representative molecule according to a preset rule, and determining a plurality of representative molecules based on the plurality of molecule groups; determining the plurality of representative molecules as target molecules.

2. The method according to claim 1, wherein Before inputting the structural information of the plurality of molecules into the cell activity prediction model, the method further includes: Obtaining the structural information of molecules with known activity and the inhibition rate of molecules with known activity; Determining similarity information according to the structural information of the molecules with known activity and the structural information of the plurality of molecules; Determining the SA score of the plurality of molecules based on the similarity information and the inhibition rate of the molecules with known activity; Performing an initial screening process on the plurality of molecules according to the SA score of the plurality of molecules and a preset score threshold.

3. The method according to claim 1, characterized in that, The step of inputting the structural information of the plurality of molecules into the cell activity prediction model and performing a first screening process on the plurality of molecules based on the output of the cell activity prediction model to obtain a first molecule includes: Inputting the structural information of the plurality of molecules into the cell activity prediction model to obtain the cell activity information corresponding to the plurality of molecules; Performing the first screening process on the plurality of molecules based on the cell activity information and a preset cell activity threshold to obtain the first molecule.

4. The method according to claim 1, wherein The step of inputting the structural information of the first molecule into the main protease activity prediction model and performing a second screening process on the first molecule based on the output of the main protease activity prediction model to obtain a second molecule includes: Inputting the structural information of the first molecule into the main protease activity prediction model to obtain the main protease activity information corresponding to the first molecule; Determining the affinity information between the first molecule and the target based on molecular docking technology and the main protease activity information; Performing the second screening process on the first molecule based on the main protease activity information, the affinity information, a preset main protease activity threshold, and a preset affinity threshold to obtain the second molecule.

5. The method according to claim 1, wherein Inputting the structural information of the second molecule into an ADMET prediction model, and performing a third screening process on the second molecule based on the output of the ADMET prediction model to obtain a third molecule, includes: Inputting the structural information of the second molecule into an ADMET prediction model to obtain various ADMET property information of the second molecule; Determining the toxicity information, metabolic capacity information, and permeability information of the second molecule based on the various ADMET property information; Performing the third screening process on the second molecule based on the toxicity information, metabolic capacity information, permeability information, a preset toxicity threshold, a preset metabolic capacity threshold, and a preset permeability threshold to obtain the third molecule.

6. The method according to claim 1, wherein The covalent binding ability information is used to characterize the covalent binding ability of each molecule in the third molecule to the main protease.

7. A screening device, characterized in that, The device includes: An acquisition module, configured to acquire a plurality of molecules from a molecular library and acquire the structural information of the plurality of molecules; A first execution module, configured to input the structural information of the plurality of molecules into a cell activity prediction model, and perform a first screening process on the plurality of molecules based on the output of the cell activity prediction model to obtain a first molecule; A second execution module, configured to input the structural information of the first molecule into a main protease activity prediction model, and perform a second screening process on the first molecule based on the output of the main protease activity prediction model to obtain a second molecule; A third execution module, configured to input the structural information of the second molecule into an ADMET prediction model, and perform a third screening process on the second molecule based on the output of the ADMET prediction model to obtain a third molecule; A determination module, configured to determine the covalent binding ability information of the third molecule based on quantum chemistry; perform a fourth screening process on the three molecules based on the covalent binding ability information and a preset covalent binding ability threshold to obtain a fourth molecule; perform a clustering process on the fourth molecule to obtain a plurality of molecule groups; for each molecule group, determine one molecule as a representative molecule according to a preset rule, and determine a plurality of representative molecules based on the plurality of molecule groups; determine the plurality of representative molecules as target molecules.

8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Drug screening method

    CN103294933A

  • Method for screening small-molecule inhibitors by taking cathepsin D as target point

    CN107346379A