Drug compound discovery method and system based on artificial intelligence
Through an artificial intelligence-based drug compound discovery method, combining screening, structural optimization, protonation prediction, protein repair, hydrogen network optimization, docking conformation prediction and comprehensive scoring, the problem of difficult to quickly and accurately identify drug active compounds in the existing technology is solved, efficient and accurate drug screening is achieved, R&D costs and time are reduced, and personalized medical and environmental protection is promoted.
Patent Information
- Application Number
- CN202510487493.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to perform mixed scoring based on multiple models, and it is difficult to quickly and accurately identify compounds with potential drug activity, affecting screening efficiency and accuracy.
A method of drug compound discovery based on artificial intelligence is adopted, including screening of predicted molecular populations, structural optimization, protonation prediction, protein repair, hydrogen network optimization, docking conformation prediction and comprehensive scoring to obtain comprehensive scoring scores and perform molecular screening.
It significantly improves the efficiency and accuracy of drug screening, especially in the field of immunomodulators, reduces R&D costs and time, improves the success rate of new drugs on marketing, and promotes the development of personalized medical care, while reducing the generation of chemical waste, which is a friendly environment.
Smart Images

Figure CN120015166A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of compound discovery, and in particular to a drug compound discovery method and system based on artificial intelligence. Background Art
[0002] In the field of drug research and development, drug discovery is a long and complex process, and the discovery and optimization of lead compounds is a key link in this process. The traditional research and development approach mainly relies on the experience of medicinal chemists, involving a large amount of experiments and resource investment, resulting in a long research and development cycle and high costs. With the rapid development of artificial intelligence technology, especially the widespread application of machine learning and deep learning algorithms, AI has gradually shown great potential in drug research and development, especially in molecular virtual screening.
[0003] At present, the Chinese invention patent with application number CN202310234744.9 discloses a virtual screening method and application of quorum sensing lead compounds. The main process includes: the input molecular compound structure is preprocessed to construct a molecular adjacency matrix, and sent to the GNN1 network to generate compound features; the input protein sequence, extracts its protein amino acid composition and dipeptide frequency to form a preliminary protein feature vector, and sends it to the cross network to generate cross-fusion features; at the same time, the protein sequence generates a corresponding contact map, which is then sent to the GNN2 network to generate protein sequence features; finally, the three feature combinations are sent to the fully connected layer to predict the affinity value. This invention can be used to discover new compounds with quorum sensing activity, providing new ideas and means for the control and prevention of bacteria such as Ralstonia solanacearum; at the same time, this method can efficiently screen out compounds that bind to PhcA and PhcR proteins, thereby discovering compounds with quorum sensing activity.
[0004] The above-mentioned technologies are difficult to perform mixed scoring based on multiple models, making it difficult to quickly and accurately identify compounds with potential drug activity, thus affecting screening efficiency and accuracy. Summary of the invention
[0005] The technical problem solved by the present invention is that it is difficult to perform mixed scoring based on multiple models in the prior art, and it is difficult to quickly and accurately identify compounds with potential drug activity, which affects the screening efficiency and accuracy.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: A drug compound discovery method based on artificial intelligence comprises the following steps: Step S1, screening the molecular population to be predicted to obtain a rough screening molecular population; Step S2, performing structural optimization and protonation prediction on the coarse screened molecular population to obtain a predicted molecular population; Step S3, performing protein repair and hydrogen network optimization on the predicted molecular population to obtain the pre-processed molecular population; Step S4, performing docking conformation prediction on the pretreated molecular group to obtain a docked molecular group, performing docking conformation prediction on the docked molecular group to obtain a comprehensive score; Step S5, performing molecular screening based on the comprehensive scoring value and the docked molecular population to obtain the screened molecular population.
[0007] Preferably, the step S1 includes the following sub-steps: Step S101, inputting a group of molecules to be predicted into a drug rough screening algorithm based on drug chemical rules to obtain a rough screening score, wherein the group of molecules to be predicted includes the molecules to be predicted; Step S102, obtaining a coarse screening molecular group based on the coarse screening score; If the coarse screening score is equal to 1, it means that the molecule to be predicted is accepted and output as a coarse screening molecule, and the coarse screening molecules are merged and output as a coarse screening molecule group; If the coarse screening score is equal to 0, it means that the molecule to be predicted is abandoned.
[0008] Preferably, the drug rough screening algorithm based on drug chemical rules in step S101 is: The i-th molecule to be predicted in the group of molecules to be predicted is roughly screened and scored. The mathematical expression of the rough screening score is: ; in, Score the coarse screen. is the number of all molecules to be predicted in the group of molecules to be predicted, is the Dirac function, is the reference molecule, is the number of molecules to be predicted in the group of molecules to be predicted, is a natural number greater than 0, For the Molecules to be predicted, is the similarity between the molecule to be predicted and the reference molecule, is the maximum similarity threshold of the reference molecule, is the minimum similarity threshold of the reference molecule, To satisfy The five principles of drug-like The five principles of drugs are molecular weight < 500D, ≤5, hydrogen bond receptors <10, PSA <140, for Calculate the function, for The calculation function accepts the minimum threshold, for Calculate the function, For receiving Calculate the maximum threshold of the function, For receiving Evaluates the minimum threshold for a function.
[0009] Preferably, step S2 includes the following sub-steps: Step S201, structural optimization of the coarse screen molecular group, the logic of the structural optimization is: Input the coarse-screened molecular group, check and delete the metal atoms; If the coarse screened molecular group contains metal atoms, delete the metal atoms, and after deleting the metal atoms, check and delete the metal complexes; If the coarse screened molecular population contains a metal complex, delete the entire metal complex, generate a 3D structure after deleting the metal complex, and check the number of atoms; If the coarse-screened molecule group contains coarse-screened molecules with more than 56 atoms, the process is terminated; If the coarse sieving molecule group does not contain a coarse sieving molecule with an atomic number exceeding 56, then determine whether the coarse sieving molecule in the coarse sieving molecule group has tautomerism; If tautomerism exists, all tautomers are generated, and other isomers with the highest score less than or equal to 1 are screened to form an isomer list; Step S202, performing protonation prediction on the coarse screening molecular group after structural optimization, obtaining a number of predicted molecules, and merging the predicted molecules to output as a predicted molecular group.
[0010] Preferably, the logic of the protonation prediction is: According to the isomer list, it is judged whether the coarse-screened molecules in the coarse-screened molecule group need to be protonated. According to the preset judgment conditions, the integrity of the atomic structure of the coarse-screened molecules is maintained and the charge is removed. The protonation state of the coarse-screened molecules is predicted based on the protonated fragment library using a machine learning method. After predicting the protonation state of the coarse-screened molecules, it is judged whether a 3D structure is generated. If not, the existing results are retained. If generated, the 3D structure is optimized based on quantum mechanics and molecular mechanics methods, and the optimal structure is screened out based on energy and it is judged whether the coarse-screened molecules have stereoisomerism. If stereoisomerism exists, the original stereoisomer center is skipped and retained to generate all stereoisomers. According to the 3D structure of the stereoisomer, the generated 3D structure is optimized and sorted to generate a predicted molecular group.
[0011] Preferably, step S3 includes the following sub-steps: Step S301, protein repair is performed on the predicted molecular population, and the logic of the protein repair is: Accepting the predicted molecular group, checking whether there are heterologous molecules in the predicted molecular group, wherein the heterologous molecules include ligands and cofactors, correcting errors in the heterologous molecules, verifying that the naming and numbering of the protein chains in the predicted molecular group after the errors in the heterologous molecules are correct, and if the naming and numbering of the protein chains are incorrect, repairing the protein chains and missing information, deleting polypeptide chains whose length is less than a preset chain length in the predicted molecular group after the protein chains and missing information are repaired, checking DNA, RNA and non-standard atoms, and deleting virtual atoms; Step S302, performing hydrogen network optimization on the predicted molecular group after protein repair, obtaining preprocessed molecules, and merging and outputting them as the preprocessed molecular group.
[0012] Preferably, the logic of hydrogen network optimization is: Display the missing atoms in the predicted molecular group after protein repair, repair the atomic information of the predicted molecules with missing atoms, check whether the three-dimensional coordinates in the atomic information are complete, and if the three-dimensional coordinates in the atomic information are complete, display the missing amino acid residues of the protein in the predicted molecular group after repairing the atomic information of the missing atoms, repair the atoms in the missing amino acid residues, detect non-standard atomic types, delete atoms that do not meet the preset specifications, screen specific compounds in the composite system, detect and repair disulfide bonds in the proteins in the predicted molecular group, generate a predicted molecular group with correct bridging structure, optimize the three-dimensional structure of the protein in the predicted molecular group, and obtain the preprocessed molecular group.
[0013] Preferably, step S4 includes the following sub-steps: Step S401, docking the pre-treated molecular population to obtain the docked molecular population, the docking logic is: Set the docking pocket size and pocket position, convert the protein pockets and small molecules in the .pdb file format specified in the pre-processed molecular group into the .ipdqt file format, and perform docking operations on the converted pre-processed molecular group; Step S402, performing docking conformation prediction on the pre-processed molecular group after docking to obtain a comprehensive score. The logic of the docking conformation prediction is: The pre-processed molecular population after docking scoring is input into the hybrid scoring and sampling algorithm, the logic of which is: The data is compressed after being scored using CrownVina. The mathematical expression for compressing the data after being scored using CrownVina is: ; in, This is the compressed score after using CrownVina scoring. The score after scoring CrownVina, The maximum value after scoring CrownVina, The minimum score after scoring CrownVina; The data is compressed after scoring using PKingScore. The mathematical expression for compressing the data after scoring using PKingScore is: ; in, It is the compressed score after scoring using PKingScore. The score after PKingScore scoring, The maximum score after PKingScore scoring. The minimum score after PKingScore scoring; Perform comprehensive scoring, the mathematical expression of the comprehensive scoring is: ; in, For the comprehensive scoring value, For preset The corresponding weights, For preset The corresponding weight.
[0014] Preferably, step S5 includes the following sub-steps: Step S501, input the comprehensive scoring value and the docked molecular population into a screening algorithm, the mathematical expression of the screening algorithm is: ; in, is the screening algorithm function, is the number of reference molecules, For the The comprehensive scoring results of each molecule are For the The molecular scoring results are is the maximum value of the reference molecule scoring threshold, is the minimum value of the reference molecule scoring threshold, To describe the Dirac type function that describes whether the pocket molecule to be screened needs to suppress the minimum value of the reference molecule scoring threshold, is the number of pockets that do not need to be suppressed, If condition 1 is met, If the condition is not met, it is 0. It is a Dirac type function that describes whether the molecule to be screened does not need to inhibit the pocket and meets the maximum value of the reference molecule scoring threshold. If condition 1 is met, If the condition is not met, it is 0; Step S502, judging whether the molecules in the docked molecule group satisfy the condition that the product of the two parts is 1 according to the screening algorithm, and if the product of the two parts is 1, the screened molecule group is obtained by screening.
[0015] A drug compound discovery system based on artificial intelligence, which is applied to the drug compound discovery method based on artificial intelligence, comprising a primary screening module, an optimization prediction module, a repair optimization module, a conformation prediction module and a secondary screening module; The primary screening module is used to screen the molecular population once to obtain a coarse screening molecular population; The optimization prediction module is used to perform structural optimization and protonation prediction on the coarse screening molecular population to obtain the predicted molecular population; The repair optimization module is used to perform protein repair and hydrogen network optimization on the predicted molecular population to obtain the pre-processed molecular population; The conformation prediction module is used to perform docking conformation prediction on the pre-treated molecular group, obtain the docked molecular group, perform docking conformation prediction on the docked molecular group, and obtain a comprehensive scoring value; The secondary screening module is used to perform molecular screening based on the comprehensive scoring value and the docked molecular population to obtain the screened molecular population.
[0016] Beneficial effects of the present invention: The present invention combines small molecule structure optimization with protonation prediction and mixed scoring, significantly improving the efficiency and accuracy of drug screening, especially in the field of immunomodulators. This technology reduces R&D costs and time, increases the success rate of new drugs on the market, and meets the market demand for highly effective immunomodulators. At the same time, it promotes the development of personalized medicine, and can reduce chemical waste by optimizing experimental processes, which is environmentally friendly. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 A flowchart of a method for discovering drug compounds based on artificial intelligence according to an embodiment of the present invention; Figure 2 A protein repair and optimization flow chart of a drug compound discovery method based on artificial intelligence provided by one embodiment of the present invention; Figure 3 A schematic diagram of the system structure of an artificial intelligence-based drug compound discovery system provided for one embodiment of the present invention. DETAILED DESCRIPTION
[0018] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, but not all of the embodiments.
[0019] Example 1, reference Figure 1 , provides an artificial intelligence-based drug compound discovery method, comprising the following steps: Step S1, screening the molecular population to be predicted to obtain a rough screening molecular population; Step S2, performing structural optimization and protonation prediction on the coarse screened molecular population to obtain a predicted molecular population; Step S3, performing protein repair and hydrogen network optimization on the predicted molecular population to obtain the pre-processed molecular population; Step S4, performing docking conformation prediction on the pretreated molecular group to obtain a docked molecular group, performing docking conformation prediction on the docked molecular group to obtain a comprehensive score; Step S5, performing molecular screening based on the comprehensive scoring value and the docked molecular population to obtain the screened molecular population.
[0020] This system aims to screen all molecules by combining small molecule structure optimization, protonation prediction, rough screening based on physical rules, protein repair, hybrid docking of classical physical algorithms and AI models, and a hybrid scoring system based on multiple models. By integrating these advanced technologies, it can quickly and accurately identify compounds with potential drug activity, especially in the field of immunomodulators, to meet the growing demand for efficient and safe immunomodulators in the autoimmune disease treatment market, and focus on using AI technology to optimize and improve the discovery process of immunomodulator compounds.
[0021] Small molecule structure optimization is a key step in drug design, involving adjustments to the geometric structure of the compound to minimize energy and find the most stable conformation. Protonation prediction is crucial for drug design, and Epik software uses machine learning to predict the pKa value and protonation state distribution of complex drug-like molecules. Rough screening based on physical rules can quickly exclude compounds that do not meet basic physical and chemical standards. Protein repair technology is also a key link in drug design, involving the identification and repair of anomalies in protein structure. Hybrid docking methods combine classical physical algorithms and artificial intelligence models to improve the accuracy and efficiency of docking. Hybrid scoring systems based on multiple models combine different scoring functions to provide a comprehensive score to more accurately predict the activity and selectivity of compounds. Together, these technologies form a one-stop, high-throughput, high-precision, and high-efficiency compound screening platform that aims to accelerate the development of new drugs and improve the success rate by integrating the latest AI technologies and computational methods.
[0022] Although AI technology in this field has made significant progress, there are still some technical and application deficiencies that need further development and improvement. For example, the accuracy and generalization ability of small molecule structure optimization still have room for improvement, especially when dealing with protonation predictions in complex biological environments. Rough screening based on physical rules may lack comprehensive consideration of the multi-faceted properties of compounds, and more biological properties and pharmacodynamic data need to be further integrated. The automation and accuracy of protein repair technology also need to be further improved to reduce manual intervention and improve efficiency. The integration efficiency of algorithms and the convergence speed of models in hybrid docking methods are technical challenges, especially when dealing with large-scale data sets. The comprehensive evaluation capabilities of hybrid scoring systems, especially in terms of data dependence and quality issues, also need to be further improved. These challenges are the key directions for future research and development. By solving these problems, the success rate of AI technology in the compound discovery stage can be further improved, and more efficient and accurate immunomodulator compound discovery can be achieved.
[0023] In order to screen out potential candidate drug molecules more efficiently and accurately, we first screened a large number of molecules based on drug rules and reference molecule similarity. The drug rules here include Lipinski's five rules, QED calculation function and logP. The purpose is to screen out molecules that meet the general drug characteristics and have some similarity with reference molecules. Then, the molecules after the rough screening are structurally optimized and the protonation state is predicted to improve the accuracy of the subsequent docking. At the same time, in order to smoothly proceed with the subsequent docking, we developed a protein repair and optimization module to repair the protein and optimize the hydrogen network. After completing the previous pretreatment, the molecules that meet the rough screening rules are docked and docking conformation is predicted. Since the two docking algorithms are completely based on different rules to predict the docking affinity, a comprehensive scoring algorithm is developed to combine the two scoring algorithms to give a more robust and accurate scoring algorithm to sort the molecules after docking.
[0024] Step S1 includes the following sub-steps: Step S101, inputting a group of molecules to be predicted into a drug rough screening algorithm based on drug chemical rules to obtain a rough screening score, wherein the group of molecules to be predicted includes the molecules to be predicted.
[0025] Step S102, obtaining a coarse screening molecular group based on the coarse screening score.
[0026] If the coarse screening score is equal to 1, it means that the molecule to be predicted is accepted and output as a coarse screening molecule, and the coarse screening molecules are merged and output as a coarse screening molecule group.
[0027] If the coarse screening score is equal to 0, it means that the molecule to be predicted is abandoned.
[0028] The present invention proposes a drug rough screening algorithm based on medicinal chemistry rules. By integrating the basic principles and rules of medicinal chemistry, the algorithm can quickly screen out compounds that meet specific physicochemical properties. This method not only improves the screening efficiency, but also reduces the workload of subsequent experiments, providing an efficient pre-screening tool for the early stages of drug discovery.
[0029] The drug rough screening algorithm based on drug chemical rules in step S101 is: The i-th molecule to be predicted in the group of molecules to be predicted is roughly screened and scored. The mathematical expression of the rough screening score is: ; in, Score the coarse screen. is the number of reference molecules, is the Dirac function, is the reference molecule, is the number of molecules to be predicted in the group of molecules to be predicted, is a natural number greater than 0, For the Molecules to be predicted, is the similarity between the predicted molecule and the reference molecule calculated based on the Dice algorithm, is the maximum similarity threshold of the reference molecule, is the minimum similarity threshold of the reference molecule, To satisfy Five principles of drug-like The five principles of drugs are molecular weight < 500D, ≤5, hydrogen bond receptors <10, PSA <140, for Calculate the function, for The calculation function accepts the minimum threshold, for Calculate the function, For receiving Calculate the maximum threshold of the function, For receiving Evaluates the minimum threshold for a function.
[0030] The reference molecules mentioned here are some molecules that have been recognized to have certain activity and selectivity. They are used as references in order to screen out molecules with better activity and selectivity than them.
[0031] Lipinski's five rules are a set of empirical rules used in medicinal chemistry to evaluate whether a compound has drug-likeness. They were proposed by Christopher A. Lipinski in 1997. This set of rules is mainly used to predict the oral bioavailability of small molecule compounds. The five criteria of the rules are as follows: 1. Molecular Weight: the molecular weight of the compound should be less than 500Da; 2. Lipid-water partition coefficient (LogP): The lipid-water partition coefficient (LogP, usually expressed as octanol-water partition coefficient) of the compound should be less than 5; 3. The number of hydrogen bond donors (HBD): the number of hydrogen bond donors (such as -OH, -NH, etc.) in the compound should be less than 5; 4. The number of hydrogen bond acceptors (HBA): the number of hydrogen bond acceptors (such as -O, -N, etc.) in the compound should be less than 10; 5. Number of rotatable bonds: the number of rotatable bonds in the compound should be less than 10.
[0032] The core idea of Lipinski's five rules is that if a compound violates any two or more of the above rules, then it may not have good oral bioavailability. These rules are widely used to screen potential drug candidates in the early stages of drug development.
[0033] The present invention proposes a small molecule structure optimization method, which combines the advantages of quantum mechanics and molecular mechanics to predict the most stable conformation of small molecules. At the same time, a protonation prediction method is also proposed, which uses machine learning technology to accurately predict the protonation state of small molecules under specific pH conditions, which is crucial for the bioavailability and efficacy of drugs.
[0034] The main algorithms for small molecule structure optimization, protonation state prediction, tautomerism and stereoisomerism state prediction include: Molecular conformation optimization based on classical force fields and dynamics.
[0035] Predict the protonation state of a molecule based on a protonation state fragment library retrieval algorithm.
[0036] It is combined with the Rdkit library to predict molecular tautomerism and stereoisomerism; and finally the optimal structure is screened based on molecular potential energy.
[0037] Step S2 includes the following sub-steps: Step S201, structural optimization of the coarse screen molecular group, the logic of structural optimization is: Enter a coarse screen of molecular groups, check and delete metal atoms.
[0038] If the coarse screened molecular group contains metal atoms, delete the metal atoms, and after deleting the metal atoms, check and delete the metal complexes.
[0039] If the coarse screened molecular population contains a metal complex, delete the entire metal complex, generate a 3D structure after deleting the metal complex, and check the number of atoms.
[0040] If the coarse screening molecule group contains a coarse screening molecule with an atomic number exceeding 56, the process is terminated.
[0041] If the coarse screening molecule group does not contain coarse screening molecules with an atomic number exceeding 56, it is determined whether the coarse screening molecules in the coarse screening molecule group have tautomerism.
[0042] If tautomerism exists, all tautomers are generated, and other isomers with the highest score less than or equal to 1 are screened to form an isomer list.
[0043] Step S202, performing protonation prediction on the coarse screening molecular group after structural optimization, obtaining a number of predicted molecules, and merging the predicted molecules to output as a predicted molecular group.
[0044] The logic of protonation prediction is: According to the isomer list, it is judged whether the coarse-screened molecules in the coarse-screened molecule group need to be protonated. According to the preset judgment conditions, the integrity of the atomic structure of the coarse-screened molecules is maintained and the charge is removed. The protonation state of the coarse-screened molecules is predicted based on the protonated fragment library using a machine learning method. After predicting the protonation state of the coarse-screened molecules, it is judged whether a 3D structure is generated. If not, the existing results are retained. If generated, the 3D structure is optimized based on quantum mechanics and molecular mechanics methods, and the Rdkit library is combined to screen out the optimal structure and judge whether the coarse-screened molecules have stereoisomerism. If stereoisomerism exists, the original stereoisomer center is skipped and retained to generate all stereoisomers. According to the 3D structure of the stereoisomer, the generated 3D structure is optimized and sorted to generate a predicted molecular group.
[0045] Step S3 includes the following sub-steps: Step S301, performing protein repair on the predicted molecular population.
[0046] In protein structure research, the present invention proposes a protein repair and optimization method that can identify and correct anomalies in protein models and improve the accuracy of protein models. This is of great significance for subsequent molecular docking and drug design because it directly affects the prediction results of drug-target binding.
[0047] like Figure 2 As shown in the protein repair and optimization flow chart, the logic of protein repair is: Accept the predicted molecular group, check whether there are heterologous molecules in the predicted molecular group, heterologous molecules include ligands and cofactors, correct errors in heterologous molecules, verify that the naming and numbering of protein chains in the predicted molecular group after correcting the errors of heterologous molecules are correct, if the naming and numbering of protein chains are incorrect, repair the protein chains and missing information, delete the polypeptide chains whose length is less than the preset chain length in the predicted molecular group after repairing the protein chains and missing information, check DNA, RNA and non-standard atoms, and delete virtual atoms.
[0048] Step S302, performing hydrogen network optimization on the predicted molecular group after protein repair, obtaining preprocessed molecules, and merging and outputting them as the preprocessed molecular group.
[0049] The logic of hydrogen network optimization is: Display the missing atoms in the predicted molecular group after protein repair, repair the atomic information of the predicted molecules with missing atoms, check whether the three-dimensional coordinates in the atomic information are complete, and if the three-dimensional coordinates in the atomic information are complete, display the missing amino acid residues of the protein in the predicted molecular group after repairing the atomic information of the missing atoms, repair the atoms in the missing amino acid residues, detect non-standard atomic types, delete atoms that do not meet the preset specifications, screen specific compounds in the composite system, detect and repair disulfide bonds in the proteins in the predicted molecular group, generate a predicted molecular group with correct bridging structure, optimize the three-dimensional structure of the protein in the predicted molecular group, and obtain the preprocessed molecular group.
[0050] The present invention integrates multiple docking methods, including classical docking and AI docking technology. This integrated method can evaluate the interaction between small molecules and proteins from different angles, improving the accuracy and reliability of docking results. By combining the physical basis of classical methods and the predictive ability of AI methods, the present invention can more comprehensively understand and predict the binding characteristics of drugs and targets.
[0051] Step S4 includes the following sub-steps: Step S401, docking the pre-treated molecular population to obtain the docked molecular population. The docking logic is: The size and position of the docking pocket are set, the protein pockets and small molecules in the .pdb file format specified in the pre-processed molecular group are converted into .ipdqt file format, and the docking operation is performed on the converted pre-processed molecular group.
[0052] Step S402, performing docking conformation prediction on the pre-processed molecular group after docking to obtain a comprehensive score. The logic of the docking conformation prediction is: The pre-processed molecular population after docking scoring is input into the hybrid scoring and sampling algorithm, the logic of which is: The data is compressed after being scored using CrownVina. The mathematical expression for compressing the data after being scored using CrownVina is: ; in, This is the compressed score after using CrownVina scoring. The score after scoring CrownVina, The maximum value after scoring CrownVina, The minimum score after scoring CrownVina.
[0053] The data is compressed after being scored using PKingScore. The mathematical expression for compressing the data after being scored using PKingScore is: ; in, It is the compressed score after scoring using PKingScore. The score after PKingScore scoring, The maximum score after PKingScore scoring. It is the minimum score after PKingScore scoring.
[0054] Comprehensive scoring is performed, and the mathematical expression of comprehensive scoring is: ; in, For the comprehensive scoring value, For preset The corresponding weights, For preset The corresponding weight.
[0055] The present invention integrates multiple docking methods, including classical docking and AI docking technology. This integrated method can evaluate the interaction between small molecules and proteins from different angles, improving the accuracy and reliability of docking results. By combining the physical basis of classical methods and the predictive ability of AI methods, the present invention can more comprehensively understand and predict the binding characteristics of drugs and targets.
[0056] Although the two docking scoring algorithms can complete the virtual screening scoring of molecules separately, we usually use the two algorithms in combination, which is also one of the characteristics of the present invention. After the two algorithms participate in the calculation at the same time, the subsequent comprehensive scoring algorithm gives a new score value based on the two scoring algorithms as the final indicator for evaluating the activity of the molecule.
[0057] As for why CrownVina and PKingScore are used at the same time and the final comprehensive score is given, it is because CrownVina is better at predicting molecular conformations and the scoring accuracy is relatively stable for molecules with medium, high and low activities, while PKingScore is mainly good at scoring molecules with higher activities, but has poor predictive ability for molecules with medium and low activities, so comprehensive scoring can effectively make up for their shortcomings.
[0058] Finally, the present invention proposes a candidate molecule screening algorithm based on comprehensive scoring and reference ligands. This algorithm not only considers the activity and selectivity of the compound, but also refers to the structure and properties of known active molecules, thereby increasing the probability of screening out compounds with potential drug activity. This method is particularly suitable for the discovery of immunomodulators because it can identify compounds with high selectivity for specific targets.
[0059] Step S5 includes the following sub-steps: Step S501, input the comprehensive scoring value and the docked molecular population into the screening algorithm. The mathematical expression of the screening algorithm is: ; in, is the screening algorithm function, is the number of reference molecules, For the The comprehensive scoring results of each molecule are For the The molecular scoring results are is the maximum value of the reference molecule scoring threshold, is the minimum value of the reference molecule scoring threshold, To describe the Dirac type function that describes whether the pocket molecule to be screened needs to suppress the minimum value of the reference molecule scoring threshold, The condition is met as 1, If the condition is not met, it is 0. It is a Dirac type function that describes whether the molecule to be screened that does not require inhibition of the pocket meets the maximum value of the reference molecule scoring threshold. If condition 1 is met, If the condition is not met, the value is 0.
[0060] Step S502, judging whether the molecules in the docked molecule group satisfy the condition that the product of the two parts is 1 according to the screening algorithm, and if the product of the two parts is 1, the screened molecule group is obtained by screening.
[0061] Through the integration and innovative application of these technologies, the present invention aims to provide a more efficient and precise immunomodulator compound discovery platform to accelerate the development of new drugs and improve the success rate.
[0062] The technical solution adopted by the present invention to solve the technical problem is: 1. In order to deal with various problems encountered in high-throughput processing of small molecules, including removing impurity ions and repairing protonation states, we developed a molecular state optimization algorithm. This algorithm can automatically identify and remove impurity ions in molecules to ensure the purity of molecules, while repairing the protonation state of molecules to adapt to different physiological environments. The implementation of this algorithm greatly improves the automation level and accuracy of small molecule processing, and provides high-quality molecular input for subsequent molecular docking and screening.
[0063] 2. In response to the possible missing atoms and residues in protein structures and the need to optimize the hydrogen network, we have developed a protein repair method. This method can identify incomplete parts in protein structures and intelligently supplement missing atoms and residues based on the physicochemical properties of proteins and bioinformatics data. In addition, it can optimize the hydrogen network of proteins, ensure the stability and accuracy of protein structures, and provide a more reliable protein model for molecular docking.
[0064] 3. In order to improve the accuracy and efficiency of molecular docking, we developed a molecular docking conformation and scoring function optimization method based on equivariant graph neural network and Gaussian mixture neural network, namely PKingScore docking scoring model. This method uses equivariant graph neural network to capture the geometric and topological characteristics of molecules, and uses Gaussian mixture neural network to optimize the conformation of molecular docking and improve the scoring function. This method, which combines deep learning and graph theory, can more accurately predict the binding conformation of molecules and provide more reasonable scores, thereby improving the reliability of docking results.
[0065] 4. In order to comprehensively evaluate the activity and selectivity of molecules, we developed a multi-model hybrid scoring algorithm. This algorithm integrates multiple different scoring models, including physics-based, knowledge-based, and machine learning-based models to provide a comprehensive evaluation. Through this hybrid scoring method, we can evaluate the activity and selectivity of molecules from multiple perspectives, thereby more accurately predicting the efficacy and safety of molecules. This approach is particularly suitable for the discovery of immunomodulators because it can identify compounds with high selectivity for specific targets while reducing potential side effects.
[0066] 5. The entire virtual screening system takes into account both accuracy and automation to the greatest extent, and the virtual screening process requires almost no additional intervention.
[0067] Compared with the prior art, the present invention has the following beneficial effects: Improve screening efficiency and accuracy: By integrating small molecule structure optimization, protonation prediction, physical rule-based coarse screening, protein repair, hybrid docking of classical physical algorithms and AI models, and a hybrid scoring system based on multiple models, the present invention can quickly and accurately identify compounds with potential drug activity, especially in the field of immunomodulators, significantly improving screening efficiency and accuracy.
[0068] Reduce R&D costs and time: The one-stop, high-throughput, high-precision, and high-efficiency compound screening platform of the present invention can reduce the experimental verification steps in the drug development process, thereby reducing R&D costs and shortening drug market launch time.
[0069] Improving the success rate of new drug launches: By improving the accuracy and reliability of compound screening, the present invention helps to improve the success rate of new drug development, especially in the field of immunomodulators, which is crucial to meeting the growing demand for efficient and safe immunomodulators in the autoimmune disease treatment market.
[0070] Promoting the development of personalized medicine: The present invention can identify compounds with high selectivity for specific targets, which will help develop more targeted and personalized drugs and promote the development of personalized medicine.
[0071] Environmentally friendly: The present invention helps to reduce the generation of chemical waste by reducing experimental steps and optimizing experimental processes, and has a positive impact on environmental protection.
[0072] Example 2, reference Figure 3 , provides an artificial intelligence-based drug compound discovery system, including a primary screening module, an optimization prediction module, a repair optimization module, a conformation prediction module and a secondary screening module.
[0073] The primary screening module is used to screen the molecular population once to obtain a coarse screening molecular population.
[0074] The optimization prediction module is used to perform structural optimization and protonation prediction on the coarse screening molecular population to obtain the predicted molecular population.
[0075] The repair optimization module is used to perform protein repair and hydrogen network optimization on the predicted molecular population to obtain the preprocessed molecular population.
[0076] The conformation prediction module is used to perform docking conformation prediction on the pre-treated molecular group, obtain the docked molecular group, perform docking conformation prediction on the docked molecular group, and obtain a comprehensive scoring value.
[0077] The secondary screening module is used to perform molecular screening based on the comprehensive scoring value and the molecular population after docking to obtain the molecular population after screening.
[0078] It should be understood by those skilled in the art that the embodiments of the present invention may be provided as methods, systems or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program codes. Among them, the storage medium may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, referred to as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, referred to as EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, referred to as EPROM), programmable read-only memory (Programmable Read-Only Memory, referred to as PROM), read-only memory (Read-Only Memory, referred to as ROM), magnetic memory, flash memory, magnetic disk or optical disk. These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0079] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A drug compound discovery method based on artificial intelligence, characterized in that: The steps include: Step S1, screening the molecular population to be predicted to obtain a rough screening molecular population; Step S2, performing structural optimization and protonation prediction on the coarse screened molecular population to obtain a predicted molecular population; Step S3, performing protein repair and hydrogen network optimization on the predicted molecular population to obtain the pre-processed molecular population; Step S4, performing docking conformation prediction on the pretreated molecular group to obtain a docked molecular group, performing docking conformation prediction on the docked molecular group to obtain a comprehensive score; Step S5, performing molecular screening based on the comprehensive scoring value and the docked molecular population to obtain the screened molecular population.
2. The method for drug compound discovery based on artificial intelligence according to claim 1, characterized in that: The step S1 includes the following sub-steps: Step S101, inputting a group of molecules to be predicted into a drug rough screening algorithm based on drug chemical rules to obtain a rough screening score, wherein the group of molecules to be predicted includes the molecules to be predicted; Step S102, obtaining a coarse screening molecular group based on the coarse screening score; If the coarse screening score is equal to 1, it means that the molecule to be predicted is accepted and output as a coarse screening molecule, and the coarse screening molecules are merged and output as a coarse screening molecule group; If the coarse screening score is equal to 0, it means that the molecule to be predicted is abandoned.
3. The method for drug compound discovery based on artificial intelligence according to claim 2, characterized in that: The drug rough screening algorithm based on the drug chemical rules in step S101 is: The i-th molecule to be predicted in the group of molecules to be predicted is roughly screened and scored. The mathematical expression of the rough screening score is: ; in, Score the coarse screen. is the number of all molecules to be predicted in the group of molecules to be predicted, is the Dirac function, As the reference molecule, is the number of molecules to be predicted in the group of molecules to be predicted, is a natural number greater than 0, For the Molecules to be predicted, is the similarity between the molecule to be predicted and the reference molecule, is the maximum similarity threshold of the reference molecule, is the minimum similarity threshold of the reference molecule, To satisfy The five principles of drug-like The five principles of drugs are molecular weight < 500D, ≤5, hydrogen bond receptors <10, PSA <140, for Calculate the function, for The calculation function accepts the minimum threshold, for Calculate the function, For receiving Calculate the maximum threshold of the function, For receiving Evaluates the minimum threshold for a function.
4. The method for drug compound discovery based on artificial intelligence according to claim 3, characterized in that: The step S2 includes the following sub-steps: Step S201, structural optimization of the coarse screen molecular group, the logic of the structural optimization is: Input the coarse-screened molecular group, check and delete the metal atoms; If the coarse screened molecular group contains metal atoms, delete the metal atoms, and after deleting the metal atoms, check and delete the metal complexes; If the coarse screened molecular population contains a metal complex, delete the entire metal complex, generate a 3D structure after deleting the metal complex, and check the number of atoms; If the coarse-screened molecule group contains coarse-screened molecules with more than 56 atoms, the process is terminated; If the coarse sieving molecule group does not contain a coarse sieving molecule with an atomic number exceeding 56, then determine whether the coarse sieving molecule in the coarse sieving molecule group has tautomerism; If tautomerism exists, all tautomers are generated, and other isomers with the highest score less than or equal to 1 are screened to form an isomer list; Step S202, performing protonation prediction on the coarse screening molecular group after structural optimization, obtaining a number of predicted molecules, and merging the predicted molecules to output as a predicted molecular group.
5. The method for drug compound discovery based on artificial intelligence according to claim 4, characterized in that: The logic of the protonation prediction is: According to the isomer list, it is judged whether the coarse-screened molecules in the coarse-screened molecule group need to be protonated. According to the preset judgment conditions, the integrity of the atomic structure of the coarse-screened molecules is maintained and the charge is removed. The protonation state of the coarse-screened molecules is predicted based on the protonated fragment library using a machine learning method. After predicting the protonation state of the coarse-screened molecules, it is judged whether a 3D structure is generated. If not, the existing results are retained. If generated, the 3D structure is optimized based on quantum mechanics and molecular mechanics methods, and the optimal structure is screened out based on energy and it is judged whether the coarse-screened molecules have stereoisomerism. If stereoisomerism exists, the original stereoisomer center is skipped and retained to generate all stereoisomers. According to the 3D structure of the stereoisomer, the generated 3D structure is optimized and sorted to generate a predicted molecular group.
6. The method for drug compound discovery based on artificial intelligence according to claim 5, characterized in that: The step S3 includes the following sub-steps: Step S301, protein repair is performed on the predicted molecular population, and the logic of the protein repair is: Accepting the predicted molecular group, checking whether there are heterologous molecules in the predicted molecular group, wherein the heterologous molecules include ligands and cofactors, correcting errors in the heterologous molecules, verifying that the naming and numbering of the protein chains in the predicted molecular group after the errors in the heterologous molecules are correct, and if the naming and numbering of the protein chains are incorrect, repairing the protein chains and missing information, deleting polypeptide chains whose length is less than a preset chain length in the predicted molecular group after the protein chains and missing information are repaired, checking DNA, RNA and non-standard atoms, and deleting virtual atoms; Step S302, performing hydrogen network optimization on the predicted molecular group after protein repair, obtaining preprocessed molecules, and merging and outputting them as the preprocessed molecular group.
7. The method for drug compound discovery based on artificial intelligence according to claim 6, characterized in that: The logic of hydrogen network optimization is: Display the missing atoms in the predicted molecular group after protein repair, repair the atomic information of the predicted molecules with missing atoms, check whether the three-dimensional coordinates in the atomic information are complete, and if the three-dimensional coordinates in the atomic information are complete, display the missing amino acid residues of the protein in the predicted molecular group after repairing the atomic information of the missing atoms, repair the atoms in the missing amino acid residues, detect non-standard atomic types, delete atoms that do not meet the preset specifications, screen specific compounds in the composite system, detect and repair disulfide bonds in the proteins in the predicted molecular group, generate a predicted molecular group with correct bridging structure, optimize the three-dimensional structure of the protein in the predicted molecular group, and obtain the preprocessed molecular group.
8. The method for drug compound discovery based on artificial intelligence according to claim 7, characterized in that: Step S4 includes the following sub-steps: Step S401, docking the pre-treated molecular population to obtain the docked molecular population, the docking logic is: Set the docking pocket size and pocket position, convert the protein pockets and small molecules in the .pdb file format specified in the pre-processed molecular group into the .ipdqt file format, and perform docking operations on the converted pre-processed molecular group; Step S402, performing docking conformation prediction on the pre-processed molecular group after docking to obtain a comprehensive score. The logic of the docking conformation prediction is: The pre-processed molecular population after docking scoring is input into the hybrid scoring and sampling algorithm, the logic of which is: The data is compressed after being scored using CrownVina. The mathematical expression for compressing the data after being scored using CrownVina is: ; in, This is the compressed score after using CrownVina scoring. The score after scoring CrownVina, The maximum value after scoring CrownVina, The minimum score after scoring CrownVina; The data is compressed after scoring using PKingScore. The mathematical expression for compressing the data after scoring using PKingScore is: ; in, It is the compressed score after scoring using PKingScore. The score after PKingScore scoring, The maximum score after PKingScore scoring. The minimum score after PKingScore scoring; A comprehensive score is given, and the mathematical expression of the comprehensive score is: ; in, For the comprehensive scoring value, For preset The corresponding weights, For preset The corresponding weight.
9. The method for drug compound discovery based on artificial intelligence according to claim 8, characterized in that: The step S5 includes the following sub-steps: Step S501, input the comprehensive scoring value and the docked molecular population into a screening algorithm, the mathematical expression of the screening algorithm is: ; in, is the screening algorithm function, is the number of reference molecules, For the The comprehensive scoring results of each molecule are For the The molecular scoring results are is the maximum value of the reference molecule scoring threshold, is the minimum value of the reference molecule scoring threshold, To describe the Dirac type function that describes whether the pocket molecule to be screened needs to suppress the minimum value of the reference molecule scoring threshold, is the number of pockets that do not need to be suppressed, If condition 1 is met, If the condition is not met, it is 0. It is a Dirac type function that describes whether the molecule to be screened does not need to inhibit the pocket to meet the maximum value of the reference molecule scoring threshold. If condition 1 is met, If the condition is not met, it is 0; Step S502, judging whether the molecules in the docked molecule group satisfy the condition that the product of the two parts is 1 according to the screening algorithm, and if the product of the two parts is 1, the screened molecule group is obtained by screening.
10. A drug compound discovery system based on artificial intelligence, which is applied to a drug compound discovery method based on artificial intelligence as claimed in any one of claims 1 to 9, characterized in that: It includes a primary screening module, an optimization prediction module, a repair optimization module, a conformation prediction module and a secondary screening module; The primary screening module is used to screen the molecular population once to obtain a coarse screening molecular population; The optimization prediction module is used to perform structural optimization and protonation prediction on the coarse screening molecular population to obtain the predicted molecular population; The repair optimization module is used to perform protein repair and hydrogen network optimization on the predicted molecular population to obtain the pre-processed molecular population; The conformation prediction module is used to perform docking conformation prediction on the pre-treated molecular group, obtain the docked molecular group, perform docking conformation prediction on the docked molecular group, and obtain a comprehensive scoring value; The secondary screening module is used to perform molecular screening based on the comprehensive scoring value and the docked molecular population to obtain the screened molecular population.
Citation Information
Patent Citations
Virtual screening method and application of quorum sensing lead compound
CN116189759A