Generative chemistry space virtual screening method based on large language model
By combining large language models and Bayesian optimization algorithms, and utilizing graph neural network surrogate models and molecular generators, the limitations and blind spots in virtual screening technology are solved, enabling efficient and flexible compound screening and improving the efficiency of drug discovery.
Patent Information
- Application Number
- CN202511180535.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-11-28
AI Technical Summary
Existing virtual screening technologies suffer from limitations in screening scope and blindness in the screening process, resulting in high computational and time costs. Furthermore, they cannot effectively utilize feedback for guidance, making it difficult to efficiently discover novel drug molecules.
By combining a large language model and a Bayesian optimization algorithm, and using a graph neural network surrogate model and a molecular generator, compounds in the chemical space are generated using cue engineering, and the screening process is updated and optimized in real time, with feedback used to guide the screening.
This enables the efficient and flexible discovery of molecules with medicinal potential in generative chemical space, reducing computational resources and time overhead, expanding the screening scope, and improving the efficiency of drug discovery.
Smart Images

Figure CN121034464A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of molecular generation, in particular to a method for generating a chemical space virtual screening based on a large language model. BACKGROUND
[0002] The process of drug research and development has the characteristics of long cycle, high investment, high risk and multi-disciplinary intersection, and usually needs to go through multiple stages such as target discovery, drug design, preclinical research, clinical trials, regulatory approval and post-marketing monitoring, which takes 10-15 years and costs more than 1 billion US dollars. There is a high elimination rate in the process of drug research and development, and successful cases often rely on the combination of technological innovation and accidental discovery, and need to be continuously optimized to balance efficacy, safety and commercial feasibility. Virtual screening is a key technology for accelerating the early process of drug research and development. It quickly and cost-effectively screens potential active molecules from a vast library of compounds through computer simulation technology, significantly improving the efficiency of drug discovery. Specifically, virtual screening uses molecular docking, pharmacophore modeling, machine learning and other methods to quickly predict the binding ability of compounds and target proteins, and preferentially locks candidate molecules to reduce the blindness of experimental screening.
[0003] Compared with traditional wet experiments, virtual screening technology can significantly shorten the research and development cycle, reduce the cost of consumption, and help discover structurally novel molecules that are difficult to identify by traditional methods. In addition, with the rise of artificial intelligence technology and big data technology, virtual screening also has the potential to optimize drug design, providing accurate direction for subsequent synthesis and experimental verification, and accelerating the process from target validation to lead compound optimization. However, the current virtual screening technology still has some shortcomings. On the one hand, the target library for screening still needs to be prepared or carefully customized, making the screening range highly limited; on the other hand, the screening process cannot be guided by real-time feedback, making the screening target more blind, and brute-force screening of a large candidate molecule library still causes huge computational resources and time consumption. Recently, molecular generation methods based on deep learning technology have provided rich possibilities for exploring vast chemical space. These emerging methods not only provide novel compound libraries for traditional virtual screening technology, but also can directly customize lead compounds for different target proteins, but they do not solve the fundamental drawbacks of traditional virtual screening technology, i.e. the limitation of screening range and the blindness of screening process, and these methods are not sufficient in technology to meet the prerequisites for solving these limitations, and large language models based on prompt engineering provide a good opportunity to break through the technical bottleneck of traditional virtual screening.
[0004] With a large number of parameters and a large amount of training data, a large language model has obtained deep domain knowledge and has accurate instruction following ability. Its friendly interaction with humans makes it gradually evolve into a tool for solving scientific problems in addition to solving daily question and answer problems. How to use prompt engineering to release the emergent ability of a large language model and guide the large language model to solve practical problems has gradually become the focus of researchers. In the field of drug discovery, in addition to general large models, researchers have even developed a variety of special large models to solve a variety of sub-problems in the field of life sciences, which also provides a good practical basis for large language model-based molecular design. Although the large language model can be used as a flexible control molecular generation tool, it does not solve the low efficiency problem in traditional virtual screening, and the Bayesian optimization algorithm as an efficient intelligent optimization algorithm provides crucial support for quickly and accurately screening lead compounds with drug development potential in the generated molecular library. The present application is based on these technologies and organically combines a variety of advanced algorithms to achieve efficient virtual screening of the generated chemical space. SUMMARY
[0005] In view of the defects in the prior art, the purpose of the embodiments of the present application is to provide a large language model-based generated chemical space virtual screening method to solve the problems in the above background art.
[0006] To achieve the above purpose, the present application provides the following technical solutions:
[0007] A large language model-based generated chemical space virtual screening method, comprising the following steps:
[0008] Step 1: initialization phase;
[0009] S11: determine the target protein of the virtual screening, and prepare a predefined external compound library required for starting the virtual screening of the generated chemical space;
[0010] S12: use molecular docking software to simulate the interaction mode of the molecules in the predefined external compound library with the target protein, and obtain the docking scores of these molecules, i.e. the binding affinities;
[0011] S13: train a graph neural network proxy model using the molecules and their corresponding binding affinities as a tool for subsequent rapid assessment of the affinity of generated molecules;
[0012] S14: use a collection function to select a batch of compounds with exploration value from the current compound library as a reference for a large language model-based molecular generator to guide the large language model to generate similar molecules;
[0013] S15: Based on the large language model-based molecule generator, a batch of new molecules similar in structure are generated by prompting engineering with each selected compound as a reference, and then handed over to the agent model for evaluation, and the process of generating a chemical space cycle screening is started from it;
[0014] Step two: cycle phase;
[0015] S21: The graph neural network agent model predicts the binding affinity of the newly generated molecules;
[0016] S22: The evaluation function evaluates the predicted results and selects the next step to explore from them;
[0017] S23: The selected molecules are simulated using molecular docking software to obtain the binding affinity, which is regarded as the true value, and added to the training set of the graph neural network agent model;
[0018] S24: The selected molecules are handed over to the large language model-based molecule generator to generate a batch of new molecules as new references to form a new compound library to be screened;
[0019] S25: The graph neural network agent model is retrained every certain round to adapt to the screening process in real time;
[0020] S26: Repeat the above steps until the specified round and find new molecules that meet the requirements.
[0021] As a further scheme of the present application, the target protein in step S11 and the determination of the external compound library include that the target protein is any protein of interest, and the structure thereof is obtained from the RCSB database by downloading the PDB format file, and the external compound library is any compound library, which can also be a special user-customized compound set of interest.
[0022] As a further scheme of the present application, the true value calculation of the molecular binding affinity in step S12 includes using the SMINA software to perform molecular docking experiments of molecules and target proteins, and the score returned by the software is the binding affinity of the molecules.
[0023] As a further scheme of the present application, the agent model for predicting the binding affinity of molecules in step S13 includes that the graph neural network model takes the molecular graph as input, and the binding affinity output by the SMINA software as the true label, and the training data comes from the pre-prepared external compound library, that is, after the molecules in the library are evaluated by the molecular docking software SMINA, all the affinity data are retained for the graph neural network to learn the relationship between the two;
[0024] In the subsequent continuous iterations, a part of the newly generated molecules is selected, and after being evaluated by the docking software, the "molecule-affinity tag" data pair is added to the previous training set, and the agent model is retrained to follow the screening process in real time.
[0025] As a further scheme of the present application, the collection function of the molecules to be explored in step S14 includes that in the initialization stage, the target of the collection function is a predefined compound library, and in the subsequent iteration process, the target of the collection function is a set of newly generated molecules in each iteration.
[0026] As a further scheme of the present application, the exploration of the generation chemical space in step S15 includes using a large language model as a molecule generator, and realizing this function in a prompt engineering manner, that is, by giving a "reference molecule" SMILES string, while applying a purposeful constraint, to make the large language model generate a compound similar to the reference molecule but structurally inconsistent.
[0027] As a further scheme of the present application, the selection of the large language model uses a large language model GPT4o as a molecule generator.
[0028] In summary, the embodiments of the present application have the following beneficial effects compared with the prior art:
[0029] Mainly composed of two parts of a large language model and a Bayesian optimization method. The large language model serves as a medium for exploring the generation of chemical space, and modifies some potential reference materials in a prompt engineering manner, provides an updated generation of compound library, and the principle thereof is to the manual design process of a biochemist, that is, to design around a potential skeleton structure, and explore various possible derivative compounds. The purpose of the Bayesian optimization as a screening module is to quickly select a batch of compounds with exploration value for the generator to explore. Benefited from the high efficiency of the Bayesian optimization algorithm;
[0030] The present application can avoid large-scale evaluation and efficiently find molecules with drug development potential. By combining the large language model with the Bayesian optimization algorithm, the present application can be guided by feedback in real time, and the screening process in the generation of chemical space can be promoted, so that the narrow virtual screening is extended to the general, and the efficiency, flexibility and expandability are combined, which can be applied to various customized scenes.
[0031] In order to more clearly illustrate the structural features and effects of the present application, the present application will be described in detail below in combination with the drawings and specific embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0032] Figure 1 is the flowchart of the initialization stage of the present application.
[0033] Figure 2 is the iterative flow chart of the screening phase of the present application.
[0034] Figure 3 is the model structure diagram of the present application.
[0035] Figure 4 is the result diagram of the present application when a batch of compounds are randomly given as the starting point.
[0036] Figure 5 is the result diagram of the present application when a known drug compound is given as the starting point. DETAILED DESCRIPTION
[0037] In order to make the purpose, technical scheme and advantages of the present application more clear, the present application will be further described in detail below in combination with the drawings and examples. It should be understood that the specific examples described herein are only used to explain the present application and do not limit the present application.
[0038] The specific implementation of the present application will be described in detail below in combination with specific examples.
[0039] In one embodiment, a large language model-based generated chemical space virtual screening method, see Figures 1-5 , comprising the following steps:
[0040] Step one: initialization phase;
[0041] S11: determine the target protein of virtual screening, and prepare the predefined external compound library required for starting virtual screening of the generated chemical space;
[0042] S12: simulate the interaction mode of molecules in the predefined external compound library with the target protein using molecular docking software, and obtain the docking scores of these molecules, i.e. binding affinity;
[0043] S13: train a graph neural network proxy model using molecules and their corresponding binding affinities as tools for subsequent rapid assessment of generated molecule affinity;
[0044] S14: use a collection function to select a batch of compounds with exploration value from the current compound library as a reference for the large language model-based molecule generator, guiding the large language model to generate similar molecules;
[0045] S15: the large language model-based molecule generator generates a batch of structurally similar new molecules by prompting engineering with each selected compound as a reference, and then hands them over to the proxy model for evaluation, and from this starts the process of cyclic screening of the generated chemical space;
[0046] Step two: cyclic phase;
[0047] S21: The graph neural network agent model predicts the binding affinity of the newly generated molecule;
[0048] S22: The collection function evaluates the predicted results and selects the object for the next exploration from them;
[0049] S23: The selected molecule is simulated using molecular docking software to obtain the binding affinity, which is regarded as the true value, and is added to the training set of the graph neural network agent model;
[0050] S24: The selected molecule is handed over to the large language model-based molecule generator to generate a new batch of molecules as new references, forming a new compound library to be screened;
[0051] S25: The graph neural network agent model is retrained every certain round to adapt to the screening process in real time;
[0052] S26: Repeat the above steps until the specified round and find a new molecule that meets the requirements.
[0053] Further, referring to Figures 1-5 , the determination of the target protein and the external compound library in step S11 includes that the target protein is any protein of interest, and its structure is obtained from the RCSB database by downloading the PDB format file, and the external compound library is any compound library, which can also be a special user-customized compound set of interest.
[0054] Further, referring to Figures 1-5 , the true value calculation of the molecular binding affinity in step S12 includes using the SMINA software to perform molecular docking experiments of the molecule and the target protein, and the score returned by the software is the binding affinity of the molecule.
[0055] Further, referring to Figures 1-5 , the agent model for predicting the binding affinity of the molecule in step S13 includes that the graph neural network model takes the molecular graph as input, and the binding affinity output by the SMINA software as the true label, and the training data comes from the pre-prepared external compound library, that is, after the molecules in the library are evaluated by the molecular docking software SMINA, all the affinity data are retained for the graph neural network to learn the relationship between the two;
[0056] In subsequent continuous iterations, a part of the newly generated molecules is selected, and after being evaluated by the docking software, the "molecule-affinity label" data pair is added to the previous training set, and the agent model is retrained to follow the screening process in real time.
[0057] Further, referring to Figures 1-5The selection of the acquisition function of the molecule to be explored in step S14 includes that, in the initialization stage, the target of the acquisition function is a predefined compound library, and in the subsequent iteration process, the target of the acquisition function is a set of newly generated molecules in each iteration.
[0058] Further, referring to Figures 1-5 The exploration of the generated chemical space in step S15 includes using a large language model as a molecule generator, and implementing this function in a prompt engineering manner, that is, by giving a SMILES string of a "reference molecule", while applying a purposeful constraint, to make the large language model generate a compound similar to the reference molecule but structurally inconsistent.
[0059] Further, referring to Figures 1-5 The selection of the large language model uses a large language model GPT4o as a molecule generator.
[0060] In this embodiment, the traditional virtual screening method often faces a huge compound library, which can be a general compound library such as PubChem, ZINC, or a customized compound library. The virtual screening process will screen compounds on this library for a specific target protein, which usually requires huge time and computing resource overhead. At the same time, this compound library also limits the scope of virtual screening, that is, only a limited set can be screened, and cannot be expanded to the real and huge chemical space. This feature greatly reduces the possibility of discovering drug molecules, because these compound libraries have been used many times, and most of the drug molecules have been developed. In addition, another disadvantage of traditional virtual screening is that when a drug molecule is known to have potential drug development potential, the screening algorithm cannot use this valuable feedback for further exploration, making the screening process overall brute force, purposeless, and leading to inefficient screening. In view of these disadvantages, the present application proposes a virtual screening method for generated chemical space based on a large language model and a Bayesian optimization algorithm, which aims to break through the limitation of the compound library and perform online and efficient virtual screening in the generated chemical space through feedback-guided manner, in order to discover new drug molecules in the huge generated chemical space.
[0061] The virtual screening method for generated chemical space based on a large language model and a Bayesian optimization algorithm proposed by the present application will be described and explained in detail through a specific example below.
[0062] Referring to Figure 1The present application initiates the virtual screening of the generated chemical space through an initialization process. Here, the aldo-keto reductase family 1 member B10 (AKR1B10, PDBID: 4GQ0) is first determined as the target protein for virtual screening, and ZINC100K is provided as an external compound library to provide some initial candidate targets. Subsequently, the present application randomly selects a batch of molecules from the ZINC100K library, obtains their binding affinity scores with AKR1B10 through molecular docking software, and trains a graph neural network proxy model with these as training data, in order to quickly predict the binding affinity values of the molecules. Next, the present application uses a selection function to select a certain number of molecules from the batch of randomly selected molecules from ZINC100K to hand over to the generative model for exploration. Finally, the generative model will generate a series of structural analogs with the help of the prompt engineering, taking the molecules selected by the selection function as the reference, and start the iterative exploration process of the generated chemical space from the batch of generated molecules. In particular, for the prompt engineering part, the present application not only needs to give the SMILES string of the reference molecule, but also needs to specify the number of generated molecules, and specify the actions that can be performed when modifying the reference (such as adding, deleting, replacing), the targets that can be modified (such as individual atoms, aromatic rings, or others), etc. In addition, in order to ensure the diversity and effectiveness of the generation, the present application also enhances the SMILES string before generation, and corrects the generated results using regular expressions after generation.
[0063] In the screening phase of the generated chemical space, the present application continuously generates and selects molecules through an iterative process, which is shown in Figure 2 , and the model graph is shown in Figure 3 . Specifically, after obtaining the generated molecules in the initialization phase, the present application uses a graph neural network proxy model to predict the binding affinity of the batch of molecules, and uses a selection function to select the molecular structure for further exploration. For the selected results, the present application evaluates their real binding affinity scores through molecular docking software, and on the other hand, hands them over to the large language model to obtain a new generated set, and starts the next round of iterative process. In each round of iterative process, the molecules evaluated by the molecular docking software will be retained and added to the training set of the graph neural network proxy model, and the graph neural network proxy model will be retrained every certain number of rounds to keep up with the screening in real time. Overall, the virtual screening process of the generated chemical space continuously advances in the generated chemical space through the mode of feedback guidance and prompt generation, while the Bayesian optimization module is responsible for making quick and reasonable choices in the process of advancement, to provide accurate guidance for the exploration of the generated space.
[0064] The present application can have two modes of application. If starting from randomly selected molecules, the present application is a virtual screening framework for generating chemical space, an example of which is screening against AKR1B10, the results of which are shown in Figure 4 If the starting molecules are replaced by known drug molecules, the present application simultaneously has the function of lead optimization, but it can also be summarized in the large framework of virtual screening for generating chemical space, an example of which is screening against papain-like protease (PLpro, PDB ID: 8UOB), as shown in Figure 5 .
[0065] The above description is merely preferred embodiments of the present application, but not to limit the present application, any modification, equivalent replacement and improvement made within the spirit and principle of the present application shall be included in the scope of protection of the present application.
Claims
1. A method for generating a chemical space virtual screening based on a large language model, characterized in that, The method comprises the following steps: Step one: initialization stage; S11: determine the target protein of virtual screening, and prepare the predefined external compound library required for starting virtual screening of the generated chemical space; S12: simulate the interaction mode of molecules in the predefined external compound library with the target protein using molecular docking software, and obtain the docking scores of these molecules, i.e. the binding affinity; S13: train a graph neural network proxy model using molecules and their corresponding binding affinities as a tool for subsequent rapid assessment of the affinity of generated molecules; S14: use a collection function to select a batch of compounds with exploration value from the current compound library as a reference for a large language model-based molecule generator to guide the large language model to generate similar molecules; S15: the large language model-based molecule generator generates a batch of new molecules similar to each selected compound as a reference, and then submits them to the proxy model for evaluation, and starts the process of cyclic screening of the generated chemical space; Step two: cyclic stage; S21: the graph neural network proxy model predicts the binding affinity of the newly generated molecules; S22: the collection function evaluates the predicted results and selects the next exploration object therefrom; S23: use molecular docking software to simulate the selected molecules to obtain the binding affinity, which is regarded as the true value, and add it to the training set of the graph neural network proxy model; S24: submit the selected molecules to the large language model-based molecule generator to generate a batch of new molecules as new references to form a new compound library to be screened; S25: retrain the graph neural network proxy model every certain number of rounds to adapt to the screening process in real time; S26: repeat the above steps until a specified number of rounds and a new molecule with satisfactory performance is found.
2. The large language model based method for generating a chemical space virtual screen according to claim 1, wherein, The determination of the target protein and the external compound library in S11 includes that the target protein is any protein of interest, and its structure is obtained from the RCSB database by downloading a PDB file, and the external compound library is any compound library, which is a special user-customized set of compounds of interest.
3. The method of claim 2, wherein the method is performed by a computer system comprising one or more processors and memory storing instructions that, when executed by the one or more processors, cause the computer system to perform the method. The true value calculation of the molecular binding affinity in S12 includes performing a molecular docking experiment of the molecule and the target protein using the SMINA software, and the score returned by the software is the binding affinity of the molecule.
4. The method of claim 3, wherein the method is performed by a computer system. The proxy model for predicting the binding affinity of molecules in S13 includes that the graph neural network model takes a molecular graph as input, and the binding affinity output by the SMINA software as the true label, and the training data comes from the pre-prepared external compound library, i.e. after evaluating the molecules in the library by the molecular docking software SMINA, all the affinity data is retained for the graph neural network to learn the relationship between them; In subsequent continuous iterations, a part of the newly generated molecules is selected, and after being evaluated by the docking software, the "molecule-affinity label" data pair is added to the previous training set, and the proxy model is retrained to follow the screening process in real time.
5. The method of claim 4, wherein the method is performed by a computer system comprising one or more processors and memory storing instructions that, when executed by the one or more processors, cause the computer system to perform the method. The collection function of the S14 selected to explore the molecule includes that in the initialization stage, the target of the collection function is a predefined compound library, and in the subsequent iteration process, the target of the collection function is a newly generated molecule set of each iteration.
6. The method of claim 5, wherein the method is performed by a computer system comprising one or more processors and memory storing instructions that, when executed by the one or more processors, cause the computer system to perform the method. The exploration of the S15 to generate a chemical space includes using a large language model as a molecule generator, and realizing this function in a prompt engineering manner, that is, by giving a "reference molecule" SMILES string, while applying a purposeful constraint, to make the large language model generate a compound similar to the reference molecule but structurally inconsistent.
7. The large language model based method for generating a chemical space virtual screen according to claim 6, wherein, The selection of the large language model uses a large language model GPT4o as a molecule generator.