Lead compound discovery method based on large language model
By combining fragment-based drug design methods with large language models and using large language models to generate new molecules, the limitations of virtual screening methods in chemical space exploration are solved, and efficient discovery of leading compounds is achieved, and design effects at the level of human experts are achieved.
Patent Information
- Application Number
- CN202510196427.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-13
AI Technical Summary
Existing virtual screening methods have limitations when exploring chemical spaces, making it difficult to cover the entire chemical space, making it difficult to discover new valuable molecules.
Combining fragment-based drug design methods and large language models, new molecules with potential high affinity are generated through large language models, and fragments are used as cues to stimulate the intrinsic potential of large language models to achieve effective exploration of the entire chemical space.
The same level of pilot compound discovery as human experts have been achieved, and better results and higher diversity have been achieved, significantly improving design efficiency and quality.
Smart Images

Figure CN120148609A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of drug research and development, and particularly relates to a method for discovering lead compounds based on large language models. Background Art
[0002] Drug research and development is a process with a long cycle and high cost, usually taking 10 to 20 years and costing up to $2.6 billion. In this process, the discovery of lead compounds is the most time-consuming and costly link, that is, first discover some lead compounds with certain drug properties, and then carry out subsequent optimization, clinical trials, etc. Currently, virtual screening is the most commonly used method for discovering lead compounds, but it has obvious limitations. The chemical space is huge. Research results show that the size of the entire chemical space is on the order of 10 60 ~10 100 orders of magnitude, while the chemical space that can be covered by current virtual screening is about 10 10 or so. It can be seen that the virtual screening method can only cover a very small part of the entire chemical space, and this part of the chemical space has almost been explored by researchers, and it is almost impossible to discover new valuable molecules. Therefore, exploring a larger chemical space becomes particularly important.
[0003] To overcome the limitations of virtual screening, fragment-based drug design methods have emerged. Research results show that the size of the fragment-based chemical space is about 10 7 or so, and the fragment space that can be explored by fragment-based drug design methods is about 10 5 or so, which can basically cover the entire chemical space and has great advantages compared with virtual screening. However, conventional fragment-based drug design methods highly rely on the participation of domain experts. From the selection of fragments to the processing and optimization of fragments, experts need to guide with their rich domain knowledge and experience, which makes the whole process highly subjective, difficult to fully explicitly model, and not conducive to reducing the design cost. With the rapid development of artificial intelligence, large language models have become a research hotspot. Utilizing the scale effect of large models and researching methods such as prompt to guide large models to release their inherent capabilities is one of the important directions of AI4Science. Large language models have a huge number of parameters, learn a large amount of domain knowledge during the training stage, and with their powerful emergence ability, they have powerful engineering capabilities. Research results show that in tasks involving chemical space structure information, large language models can already reach the level of human experts. How to guide large language models to release their inherent capabilities to solve practical problems is the focus of people's research. For this reason, the present invention proposes a method for discovering lead compounds based on large language models. Summary of the Invention
[0004] The object of the present invention is to provide a method for discovering lead compounds based on large language models, aiming to solve the problems raised in the above background art.
[0005] The object of the present invention is achieved through the following technical solutions:
[0006] A method for discovering lead compounds based on large language models, comprising the following steps:
[0007] Step 1: Determine the target and construct an initial compound library;
[0008] Step 2: Fragment generation and formation of an initial fragment library;
[0009] Step 3: Update the fragment library;
[0010] Calculate the frequency of fragment occurrence and the contribution degree of the fragment to the affinity of the target protein, and update the fragment library;
[0011] Step 4: Calculate the fragment sampling weight and select candidate fragments;
[0012] Integrate the frequency of fragment occurrence and the contribution degree of the fragment to the affinity of the target protein, calculate the fragment sampling weight, and sample on the fragment library according to the fragment sampling weight to obtain the most promising candidate fragments;
[0013] Step 5: The large language model generates molecules based on the candidate fragments;
[0014] Step 6: Evaluate the quality of the generated molecules and update the molecule library;
[0015] Step 7: Repeat steps 2 to 6 in a loop until a preset number of iterations is reached.
[0016] Furthermore, the specific process of step 1 is as follows:
[0017] Select a protein or RNA related to the disease as the drug target to determine the goal of drug research and development; search for its structural information in the PDB protein library according to the target number, and select small molecule drugs that have been marketed or clinically tested in the small molecule drug library of drug-like compounds to construct an initial compound library.
[0018] Furthermore, step 2 specifically includes the following steps:
[0019] Step 21: Break the molecules in the initial compound library into multiple fragments:
[0020]
[0021] Among them, represents the set of fragments split from molecule s i ; represents the molecule si The j-th fragment split out;
[0022] Step 22: Aggregate the fragment sets generated after cleavage to form an initial fragment library.
[0023] Further, the specific steps of step 3 are as follows:
[0024] Step 31: Calculate the degree of affinity contribution of each fragment to the target protein:
[0025]
[0026] Among them, Affinity() is an affinity calculation function; Indicates the j-th fragment split out from molecule s i The j-th fragment split out; F is the fragment split from molecule s i The fragment split out; Is the degree of affinity contribution of the fragment to the target protein;
[0027] Step 32: Calculate the frequency of occurrence of each fragment in the fragment library as
[0028] Step 33: Update the fragment library:
[0029]
[0030] Among them, FDB k Indicates the fragment library after the k-th iteration; FDB k+1 Indicates the fragment library after the (k + 1)-th iteration.
[0031] Further, the specific steps of step 4 are as follows:
[0032] Step 41: Synthesize the frequency of fragment occurrence and its degree of affinity contribution to the target protein, and calculate the fragment sampling weight:
[0033]
[0034] Among them, Fq k Is the fragment frequency distribution function of the k-th fragment library, Is the fragment sampling weight; Is the degree of affinity contribution of the fragment to the target protein;
[0035] Step 42: Sample on the fragment library according to the fragment sampling weight to obtain the most promising candidate fragments:
[0036]
[0037] Among them, F Sel Is the candidate fragment; is a maximum function, indicating to obtain the maximum
[0038] Furthermore, the specific process of step 5 is as follows:
[0039] Take the sampled candidate fragments as input, introduce the generative pre-trained Transformer model, and use the role definition method in prompt engineering to dynamically adjust the prompt words input to the model in each round, guiding the model to explore different chemical space regions. By fully stimulating the inherent capabilities of the model, the model generates new molecules with potentially high affinity according to the input candidate fragments.
[0040] Furthermore, the specific process of step 6 is as follows:
[0041] Use the smina docking software to evaluate the quality of the generated new molecules and obtain their affinity values; merge the newly evaluated data with the old data to construct and update the compound library;
[0042] Before docking, consult the literature based on the structural information of the target protein to find the position of the protein pocket, and define the position and size of the protein pocket according to the residue position;
[0043] When docking new molecules, use the example small molecule as a benchmark to locate the size and position of the protein pocket;
[0044] In the smina docking software, set addbox to 1, exhaustiveness to 16, and set the random seed.
[0045] Compared with the prior art, the beneficial effects of the present invention are:
[0046] The present invention combines the fragment-based drug design method with an advanced large language model, utilizes the rich domain knowledge and inherent potential of the large language model, and replaces the limitations of traditional reliance on human expert experience and intuition. In this process, the fragments act as prompts to stimulate the large language model to generate new molecules with high affinity and diversity, thereby achieving effective exploration of the entire chemical space. Experimental results show that the method of the present invention can achieve the same level of results as human experts in the lead compound discovery task for a given target protein, and compared with similar methods, it not only obtains better results but also exhibits higher diversity. This method provides a brand-new technical means for drug research and development, significantly improving the design efficiency and quality of lead compounds. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 is a flowchart of the method of the present invention.
[0048] Figure 2 It is the docking score distribution diagram of the generated results.
[0049] Figure 3 It is the interaction between the generated result and the target protein. Specific implementation mode
[0050] For a clearer understanding of the technical features, objectives, and beneficial effects of the present invention, the technical solutions of the present invention are described in detail below, but it should not be construed as a limitation on the scope of implementation of the present invention.
[0051] The following describes the specific implementation of the present invention in detail in combination with specific embodiments.
[0052] The present invention provides a method for discovering lead compounds based on a large language model, and its flowchart is as Figure 1 shown, and the discovery method includes the following steps:
[0053] Step 1: Determine the target and construct an initial compound library.
[0054] Select a protein or RNA related to the disease as the drug target to determine the goal of drug research and development. Search for its structural information according to the target number in the PDB protein library, and select small molecule drugs that have been marketed or clinically tested in common small molecule compound libraries for drugs (such as pubchem, GDB, etc.) to construct an initial compound library.
[0055] Target properties and input features of small molecules:
[0056] The target property of the small molecule is the affinity score with the target protein or other properties (such as toxicity, hydrophilicity, synthesizability, etc.). The input of the small molecule is its structural features, including its molecular formula, SMILE expression, relative atomic coordinates and other structural information.
[0057] Step 2: Fragment generation and formation of an initial fragment library. Specifically, it includes:
[0058] Step 21: Break the molecules in the initial compound library into multiple fragments:
[0059]
[0060] Among them, represents the set of fragments split from molecule s i ; represents the jth fragment split from molecule s i ;
[0061] The specific process of the step of breaking the molecules in the initial compound library into multiple fragments is as follows:
[0062] For each molecule in the initial compound library, considering the synthetic accessibility, rationality, and structure of the fragments comprehensively, select the synthetic bonds, rotatable bonds, or bonds formed by common chemical reactions of small molecules to break, generating reasonable complete fragments for subsequent processes.
[0063] Step 22: Aggregate the fragment sets generated after breaking to form an initial fragment library.
[0064] Step 3: Update the fragment library. Specifically include:
[0065] Step 31: Calculate the contribution degree of affinity of each fragment to the target protein:
[0066]
[0067] Among them, Affinity() is the affinity calculation function; F is the fragment obtained by splitting molecule s i ; is the contribution degree of affinity of the fragment to the target protein;
[0068] Step 32: Calculate the frequency of occurrence of each fragment in the fragment library as to reflect the distribution of the entire fragment chemical space;
[0069] Step 33: Update the fragment library:
[0070]
[0071] Among them, FDB k represents the fragment library after the completion of the k-th iteration; FDB k+1 represents the fragment library after the completion of the (k + 1)-th iteration.
[0072] As the algorithm iterates and explores the chemical space, the fragment library is updated dynamically. The distribution of the fragment library converges from the distribution of the initial compound library to the distribution of the entire chemical space, and then can guide the subsequent exploration of the entire chemical space.
[0073] Step 4: Calculate the fragment sampling weight and select candidate fragments. According to the fitting of the fragment library to the actual chemical space fragment distribution, sample in the chemical fragment space, and comprehensively consider the contribution degree of affinity of the fragment to the target protein and its frequency of occurrence to obtain the most potential region for subsequent exploration. Specifically include:
[0074] Step 41: Comprehensively consider the frequency of occurrence of the fragment and its contribution degree of affinity to the target protein, and calculate the fragment sampling weight:
[0075]
[0076] Among them, Fq kis the fragment frequency distribution function of the k-th round fragment library, is the fragment sampling weight.
[0077] Step 42: Sample on the fragment library according to the fragment sampling weight to obtain the candidate fragment most worthy of exploration:
[0078]
[0079] where F Sel is the candidate fragment; is the maximum function, indicating to obtain the maximum
[0080] Step 5: The large language model generates molecules based on the candidate fragments.
[0081] Input the candidate fragments obtained by sampling into the large language model, and use techniques such as prompt engineering to enable the large language model to be induced by the input molecular fragments and stimulate its inherent ability to generate new molecules. With its huge learnable parameters and rich domain knowledge contained, the large language model demonstrates significant natural language interaction advantages during the interaction process.
[0082] Specifically, the large language model adopts the Generative Pretrained Transformer model (GPT), so the specific process of this step can be obtained as follows:
[0083] Take the candidate fragments obtained by sampling as the input and introduce the Generative Pretrained Transformer model (GPT). Use the role definition method in prompt engineering to dynamically adjust the prompt words input to the GPT model in each round, guide the model to explore different chemical space regions, and by fully stimulating the inherent ability of the GPT model, enable it to generate new molecules with potential high affinity according to the input candidate fragments. This process aims to gradually explore and discover ideal lead compounds according to the affinity preference through continuous iteration and optimization.
[0084] Step 6: Evaluate the quality of the generated molecules and update the molecule library;
[0085] Use the smina docking software to evaluate the quality of the generated new molecules and obtain their affinity values. Merge the newly evaluated data with the old data to construct and update the compound library to include more potential high-affinity molecules.
[0086] Before docking, consult relevant literature according to the structural information of the target protein to find the position of the protein pocket, and define the position and size of the protein pocket according to the residue position.
[0087] When docking new molecules, use the example small molecule as a benchmark to locate the size and position of the protein pocket to ensure the accuracy of docking.
[0088] In the smina docking software, set addbox to 1 to control the range of the docking pocket and prevent the molecule from docking to other positions; set exhaustiveness to 16 to control the accuracy of its docking iteration, and set a random seed to ensure the reproducibility of the results.
[0089] Step 7: Repeat steps 2 to 6 in a loop until the preset number of iterations is reached.
[0090] Effect verification experiment:
[0091] The actual effect of this method is as Figure 2 and Figure 3 shown. Figure 2 shows the statistical laws of the generated molecules. It can be seen from the figure that the lead compound discovery method based on the large language model proposed by us has a similar affinity distribution to the molecules carefully designed by human experts when starting from a random library, basically reaching the level of human experts. Figure 3 shows the interaction between the generated molecules and the target protein. It can be seen from the figure that the molecules generated based on the lead compound discovery method based on the large language model proposed by us can form strong hydrogen bonds and π bonds in the interaction with the target protein, showing strong binding ability.
[0092] In summary, the present invention combines the fragment-based drug design method with the large language model, replaces human expert guidance in drug design with the rich domain knowledge of the large language model, and uses fragments as prompts to stimulate the internal potential of the large language model. The two complement each other's advantages to complete the exploration of the entire chemical space, and design lead compounds reaching the level of human experts in actual use, providing a new technical means for drug research and development.
[0093] The above is only the preferred embodiment of the present invention. It should be noted that for those skilled in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, which should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicality of the patent.
Claims
1. A lead compound discovery method based on a large language model, characterized in that: The following steps are involved: Step 1: Identify the target and construct the initial compound library; Step 2: Fragment generation and initial fragment library formation; Step 3: Update the fragment library; Calculate the frequency of occurrence of the fragments and the degree of affinity contribution of the fragments to the target protein and update the fragment library; Step 4: Calculation of fragment sampling weights and selection of candidate fragments; The frequency of occurrence of the fragments and the degree of contribution of the fragments to the affinity of the target protein are comprehensively considered to calculate the fragment sampling weights, and the fragment library is sampled according to the fragment sampling weights to obtain the candidate fragments most worth exploring; Step 5: The large language model generates molecules based on the candidate fragments; Step 6: Evaluate the quality of generated molecules and update the molecule library; Step 7: Repeat steps 2 to 6 until the preset number of iterations is reached.
2. The method for discovering lead compounds based on a large language model according to claim 1, characterized in that: The specific process of step 1 is as follows: Select disease-related proteins or RNA as drug targets to determine the goals of drug development; search for their structural information in the PDB protein library according to the target number, select small molecule drugs that have been marketed or clinically tested in the drug-like small molecule compound library, and build an initial compound library.
3. The method for discovering lead compounds based on a large language model according to claim 1, characterized in that: The step 2 specifically includes the following steps: Step 21: Fragment the molecules in the initial compound library into multiple fragments: in, Represented by the molecule s i The collection of split fragments; Represented by the molecule s i The jth segment split out; Step 22: Summarize the fragment collections generated after fragmentation to form an initial fragment library.
4. The method for discovering lead compounds based on a large language model according to claim 1, characterized in that: The step 3 specifically comprises the following steps: Step 31: Calculate the affinity contribution of each fragment to the target protein: Among them, Affinity() is the affinity calculation function; Represented by the molecule s i The jth fragment separated; F is the fragment of molecule s i The fragments obtained by splitting; is the degree of contribution of the fragment to the affinity of the target protein; Step 32: Calculate the frequency of each fragment in the fragment library: Step 33: Update the snippet library: Among them, FDB k Represents the fragment library after the kth iteration; FDB k+1 Represents the fragment library after the k+1th round of iteration.
5. The method for discovering lead compounds based on a large language model according to claim 1, characterized in that: The step 4 specifically comprises the following steps: Step 41: Calculate the fragment sampling weight based on the frequency of occurrence of the fragment and its affinity contribution to the target protein: where Fq k is the fragment frequency distribution function of the k-th round fragment library, is the fragment sampling weight; is the degree of contribution of the fragment to the affinity of the target protein; Step 42: According to the fragment sampling weight, sample the fragment library to obtain the candidate fragments most worth exploring: Among them, F Sel is a candidate fragment; To find the maximum value function, we need to The largest 6. The method for discovering lead compounds based on a large language model according to claim 1, characterized in that: The specific process of step 5 is as follows: The sampled candidate fragments are used as input, and a generative pre-trained Transformer model is introduced. The role definition method in prompt engineering is used to dynamically adjust the prompt words input to the model in each round to guide the model to explore different chemical space regions. By fully stimulating the model's inherent capabilities, the model can generate new molecules with potentially high affinity based on the input candidate fragments.
7. The method for discovering lead compounds based on a large language model according to claim 1, characterized in that: The specific process of step 6 is as follows: Use SMINA docking software to evaluate the quality of the generated new molecules and obtain their affinity values; merge the newly evaluated data with the old data to build and update the compound library; Before docking, the literature is consulted based on the structural information of the target protein to find the location of the protein pocket, and the location and size of the protein pocket are defined based on the residue positions; When docking a new molecule, use the example small molecule as a benchmark to locate the size and position of the protein pocket; in the smina docking software, set addbox to 1, exhaustiveness to 16, and set the random seed.
Citation Information
Patent Citations
Automatic drug design method and system, computing equipment and computer readable storage medium
CN112116963A
Method for generating and screening pilot active molecules based on artificial intelligence fragmentation technology
CN118800363A
Artificial intelligence assisted computational fragment-based drug design
WO2024182496A2
Cited By
Cooperative interaction strategy optimization method based on multiple agents
CN120354878A
Collaborative interaction strategy optimization method based on multi-agent
CN120354878B
3D molecule sampling optimization method and device, electronic equipment and storage medium
CN121963844A
A 3D molecule sampling optimization method and device, electronic equipment and storage medium
CN121963844B