Method and system for generating compound structure, and computer program product
By providing a library of compound structures and physicochemical parameters to a large language model (LLM) and generating compound structures under constraints, the problem of insufficient efficiency and accuracy in compound generation in existing technologies is solved, achieving efficient and accurate compound structure generation.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- HANGZHOU DEEP PRINCIPLE TECHNOLOGY CO LTD
- Filing Date
- 2025-10-15
- Publication Date
- 2026-04-23
AI Technical Summary
Existing technologies struggle to efficiently and accurately utilize large language models to generate compound structures, especially under constraints regarding the charge and physicochemical parameters of transition metal complexes.
By providing a library of compound structures, background compounds with known physicochemical parameters, constraints, and expected physicochemical parameters to a Large Language Model (LLM), candidate compounds that meet the conditions are generated using LLM. These candidate compounds are then combined with background compounds, and the expected compound structures are generated iteratively through optimization.
It enables efficient and accurate compound structure generation in the formation of transition metal complexes and other compounds, improving the efficiency and precision of compound generation.
Smart Images

Figure CN2025127789_23042026_PF_FP_ABST
Abstract
Description
Methods, systems, and computer program products for generating compound structures Technical Field
[0001] This application relates to the field of computational chemistry, and more specifically, to a method, system, and computer program product for generating compound structures using large language models. Background Technology
[0002] In computational chemistry, compound screening can be viewed as a search problem in a confined space, and the importance of artificial intelligence in this field is increasingly recognized. With the development of Large Language Models (LLMs), whether LLMs can be used for efficient and accurate compound generation has become a focus of attention. Summary of the Invention
[0003] This application proposes a method for generating compound structures using an LLM (Liquid Matrix Modeling), comprising: providing the LLM with a compound structure material library and background compounds with known physicochemical parameters; providing the LLM with constraint conditions and expected physicochemical parameters; using the LLM to generate candidate compounds that conform to the provided physicochemical parameters based on the compound structure material library and the physicochemical parameters under the constraint conditions; combining the candidate compounds with the background compounds to form a new combination of background compounds; and providing the LLM with the compound structure material library and the new combination of background compounds to iteratively generate the structure of the expected compound.
[0004] According to embodiments of this application, when providing the LLM with a background compound having known physicochemical parameters, the LLM is also provided with the physicochemical parameters of the background compound; when providing the LLM with a combination of the compound structure material library and the new background compound, the LLM is also provided with the physicochemical parameters of the new background compound.
[0005] According to an embodiment of this application, the limiting condition includes the charge of the transition metal complex (TMC).
[0006] According to embodiments of this application, the anticipated physicochemical parameters include at least two different physicochemical parameters.
[0007] According to embodiments of this application, the compounds include organic compounds, inorganic compounds, organometallic complexes, polymeric compounds, small molecule compounds, and biological macromolecules.
[0008] According to an embodiment of this application, combining the candidate compound with the background compound to form a new background compound combination includes: combining the candidate compound and the top K compounds with the best physicochemical parameters among the background compounds to form a new background compound combination.
[0009] According to an embodiment of this application, generating candidate compounds that conform to the provided physicochemical parameters using the LLM based on the compound structure material library and the physicochemical parameters under the limiting conditions includes: generating the candidate compounds based on new compound materials outside the compound structure material library.
[0010] According to an embodiment of this application, after generating the candidate compound based on a new compound material outside the compound structure material library, the method further includes integrating the new compound material into the compound structure material library.
[0011] According to the embodiments of this application, the compound is TMC, the compound structure material library is a ligand library, and the physicochemical parameters include the HOMO-LUMO band gap and polarizability.
[0012] This application also provides a system for generating compound structures using an LLM (Liquid Matrix Modeling), comprising a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to perform the following steps: providing the LLM with a compound structure material library and background compounds with known physicochemical parameters; providing the LLM with limiting conditions and expected physicochemical parameters; using the LLM to generate candidate compounds that conform to the provided physicochemical parameters based on the compound structure material library and the physicochemical parameters under the limiting conditions; combining the candidate compounds with the background compounds to form a new combination of background compounds; and providing the LLM with the compound structure material library and the new combination of background compounds to iteratively generate the expected compound structure.
[0013] This application also provides a computer program product, including computer program instructions that, when executed by a processor, perform the following steps: providing the LLM with a compound structure material library and background compounds with known physicochemical parameters; providing the LLM with limiting conditions and expected physicochemical parameters; using the LLM to generate candidate compounds that conform to the provided physicochemical parameters based on the compound structure material library and the physicochemical parameters under the limiting conditions; combining the candidate compounds with the background compounds to form a new combination of background compounds; and providing the LLM with the compound structure material library and the new combination of background compounds to iteratively generate the structure of the expected compound.
[0014] The methods, systems, and computer programs for generating compound structures provided in this application can utilize knowledge within an LLM and external chemical data, bypassing the need to design complex mathematical formulas, and provide efficient and accurate compound structure generation. Attached Figure Description
[0015] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0016] Figure 1 is a flowchart of a method for generating compound structures using LLM according to an embodiment of this application;
[0017] Figure 2 is a schematic diagram of a method for generating compound structures using LLM according to an embodiment of this application;
[0018] Figure 3 is a schematic diagram of a system for generating compound structures using LLM according to an embodiment of this application;
[0019] Figure 4 is a structural diagram of 50 ligands and their complexes in the embodiments of this application;
[0020] Figure 5 shows the structures of 10 new ligands generated using LLM;
[0021] Figure 6 compares the effectiveness of the method for generating compound structures using LLM with other methods;
[0022] Figure 7 is a schematic diagram of the analysis of the top 20 TMCs out of 200 generated candidate TMCs. Detailed Implementation
[0023] To better understand this application, the technical solutions of this application will be described in more detail with reference to the accompanying drawings. It should be understood that these detailed descriptions are merely illustrative of exemplary embodiments of this application and are not intended to limit the scope of this application in any way. Throughout the specification, the same reference numerals refer to the same elements. The expression "and / or" includes any combination or all combinations of one or more of the associated listed items.
[0024] It should be noted that in this specification, terms such as "first," "second," and "third" are used only to distinguish one feature from another and do not indicate any limitation on the features. In the accompanying drawings, the dimensions, scale, and shape of the illustrations have been slightly adjusted for ease of explanation. The drawings are for illustrative purposes only and are not strictly to scale. As used herein, the terms "approximately," "about," and similar terms are used as expressions of approximation, not as expressions of degree, and are intended to illustrate inherent deviations in measurements or calculations that will be recognized by one of ordinary skill in the art.
[0025] It should also be understood that expressions such as "comprising," "including," "having," "containing," and / or "comprising" are open-ended rather than closed-ended expressions in this specification, indicating the presence of the stated features, elements, and / or components, but not excluding the presence of one or more other features, elements, components, and / or combinations thereof. Furthermore, when expressions such as "at least one of..." appear after a list of listed features, they modify the entire list of features, not just a single feature in the list. Additionally, when describing embodiments of this application, the word "may" is used to mean "one or more embodiments of this application." And the term "exemplary" is intended to refer to examples or illustrations.
[0026] Unless otherwise specified, all terms used herein (including engineering and technical terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. It should also be understood that, unless expressly stated herein, terms defined in common dictionaries shall be interpreted as having the meaning consistent with their meaning in the context of the relevant art, and not as having an idealized or overly formalized meaning.
[0027] It should be noted that, unless otherwise specified, the embodiments and features in the examples of this application can be combined with each other. Furthermore, unless explicitly limited or contradicted by the context, the specific steps included in the methods described in this application are not limited to the order in which they are described, but can be performed in any order or in parallel. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0028] Figure 1 is a flowchart of a method for generating compound structures using LLM according to an embodiment of this application. Figure 2 is a schematic diagram of a method for generating compound structures using LLM according to an embodiment of this application.
[0029] In step S1010, a compound structure material library 2200 and background compounds 2300 with known physicochemical parameters are provided to the LLM 2100. The LLM is a deep learning model trained and optimized using a large amount of text data, capable of understanding, processing, and / or generating natural language text. The LLM can handle various natural language tasks, such as text classification, question answering, and dialogue. The LLM in this application can be O1-Preview or Claude-3.5-Sonnet. However, the implementation of this application is not limited to these. This application is applicable to any LLM with inherent chemical knowledge. The compound structure material library 2200 can contain building blocks for compounds, such as atoms, ions, ligands, etc. The background compounds 2300 with known physicochemical parameters can be compounds of the same class as the expected compounds.
[0030] Molecular orbital theory posits that electrons in a molecule no longer belong solely to any single atom but move throughout the entire molecule. Atomic orbitals in a molecule form molecular orbitals through linear combinations. Molecular orbitals are classified into two types: bonding orbitals and antibonding orbitals. Bonding orbitals allow electrons to form chemical bonds between atoms, while antibonding orbitals push electrons away from the regions between atoms. According to molecular orbital theory, electrons fill molecular orbitals in order of increasing energy. The highest-energy occupied orbital is called the most occupied molecular orbital (HOMO), and the lowest-energy unoccupied orbital is called the least occupied molecular orbital (LUMO).
[0031] The HOMO-LUMO energy gap (i.e., the band gap) refers to the energy difference between the highest occupied molecular orbital (HOMO) and the lowest unoccupied molecular orbital (LUMO). This energy gap has a significant impact on the chemical properties and reactivity of molecules. Generally speaking, the smaller the HOMO-LUMO energy gap, the less stable the chemical bonds in the molecule. This is because a small energy gap means that electrons can easily jump from the HOMO to the LUMO, leading to the breaking of chemical bonds. Conversely, a large energy gap indicates relatively stable chemical bonds. Secondly, the HOMO-LUMO energy gap is an important parameter for measuring the reactivity of molecules. The smaller the energy gap, the higher the reactivity of the molecule. This is because electrons can easily jump between the HOMO and LUMO, thus participating in chemical reactions. For example, in photochemical reactions, molecules with smaller HOMO-LUMO energy gaps are more likely to absorb photons and initiate reactions. In addition, the HOMO-LUMO energy gap is also related to many properties of molecules, such as redox properties and spectral properties. For example, the higher the HOMO energy level, the easier it is for the molecule to lose electrons; the lower the LUMO energy level, the easier it is for the molecule to gain electrons. Furthermore, the HOMO-LUMO energy gap also affects the absorption and emission spectra of molecules. Therefore, studying the physicochemical properties of organometallic complexes, such as the HOMO-LUMO energy gap, has practical significance in compound applications such as catalysts, optoelectronic materials, and drug design.
[0032] Physicochemical parameters can be characterizing parameters of the compound. For example, in the case of a TMC, such characterizing parameters can be the HOMO-LUMO gap, polarizability, etc., and the background compound 2300 can be a TMC constructed using ligands from the compound structure material library 2200.
[0033] In step S1020, constraint conditions 2400 and expected physicochemical parameters 2500 are provided to the LLM. It should be noted that there is no order restriction between steps S1010 and S1020. Furthermore, the order in which the compound structure library 2200, background compounds 2300, constraint conditions 2400, and physicochemical parameters 2500 are input into the LLM 2100 is not restricted. Constraint conditions 2400 can be set according to the purpose of compound generation. For example, in the case of a TMC, these constraints can be set such that the total charge of the TMC is -1, 0, or 1. As another example, these constraints can also be set such that the ligands in the generated candidate compounds should be ligands present in the compound structure library 2200. The expected physicochemical parameters 2500 can be an optimization task set according to the generation purpose. Depending on the optimization task, one physicochemical parameter 2500 can be set for single-objective optimization, or at least two different physicochemical parameters 2500 can be set for multi-objective optimization. For example, in the case of a compound called TMC, the expected physicochemical parameter 2500 could be "a larger bandgap than the HOMU-LUMO bandgap of a given TMC" and / or "a larger polarizability than the polarizability of a given TMC".
[0034] In step S1030, candidate compounds 2600 that meet the provided physicochemical parameters are generated using LLM 2100 based on the compound structure material library 2200 and physicochemical parameters 2300 under restrictive conditions. Specifically, the number of candidate compounds 2600 to be generated can be set when generating using LLM 2100.
[0035] In step S1040, candidate compound 2600 is combined with background compound 2300 to form a new combination of background compounds. There are several ways to combine them. For example, assuming background compound 2300 includes 20 TMCs, and LLM 2100 generates 10 TMCs each time, the top 20 compounds with the best physicochemical parameters from these 30 TMCs can be combined to form a new combination of background compounds. Alternatively, mimicking the scenario of researchers exploring new compounds, all 30 TMCs can be provided to LLM 2100 as new background compounds, along with the physicochemical parameters of the 10 generated TMCs. Correspondingly, the initial background compound 2300 is provided to LLM 2100 along with its corresponding physicochemical parameters. In this case, LLM 2100 can not only learn positively from the generated TMCs with excellent physicochemical parameters but also learn negatively from the generated TMCs with poor physicochemical parameters, thereby optimizing the next round of generation.
[0036] In step S1050, a combination of the compound structure library and new background compounds is provided to the LLM 2100 to iteratively generate the desired compound structure. The iteration termination condition can be set according to the actual situation, such as a predetermined number of iterations. For single-objective optimization, a smaller number of iterations can be set, such as 20; for multi-objective optimization, a larger number of iterations can be set, such as 40.
[0037] The compounds involved in this application include, but are not limited to, organic compounds, inorganic compounds, organometallic complexes, polymers, small molecules, and biomacromolecules, including but not limited to TMC.
[0038] In the task of generating TMCs, the compound structure material library 2200 can contain 50 ligands. All 50 ligands are monodentate, primarily composed of elements such as carbon, hydrogen, nitrogen, oxygen, phosphorus, and sulfur; some molecules also contain halogens (such as fluorine, chlorine, bromine, and iodine). Molecules range from lightweight monatomic molecules to medium-molecular-weight compounds. The charge states are primarily 0 (25) or -1 (25).
[0039] Furthermore, the ligands in the complexes from the compound structure library 2200, the new background compounds 2300, and the candidate compounds 2600 can be monodentate ligands, polydentate ligands, etc., and the coordinating atoms include, but are not limited to, electron-donating atoms such as O, N, and S. The ligand charge states include, but are not limited to, 0, negative charge, and multiple negative charges. The types of ligands can include monoligands and multiple different ligands. In addition, the charge states of the metal atoms can be 0, positive charge, and multiple positive charges, and the types of ligand atoms include, but are not limited to, various metal atoms including Pd. Furthermore, the complex morphology can be mononuclear, binuclear, or polynuclear complexes. In the case of polynuclear complexes, the complex can exhibit a one-dimensional chain structure, a two-dimensional planar structure, and a three-dimensional structure. In complexes with multiple metal central atoms, the metal central atoms can be the same or different.
[0040] Furthermore, the compound structure library 2200, new background compounds 2300, and candidate compounds 2600 are not limited to organometallic complexes or transition metal-organic complexes, but also include, but are not limited to, organic compounds, inorganic compounds, organometallic complexes, polymers, small molecules, and biomacromolecules. For example, during generation, a certain number of new compounds are first generated using LLM2100, and then new compound combinations are constructed to form a new set of background compounds 2300. In this case, the newly generated compounds can be incorporated into the compound structure library 2200 to enhance the diversity of the generated candidate compounds 2600.
[0041] According to another embodiment of this application, when generating new compounds, it is permissible to generate candidate compounds 2600 using new compound materials outside the compound structure material library 2200. These new materials can come from the knowledge base included with the LLM 2100, or they can be generated directly by the LLM 2100.
[0042] For example, in the task of generating TMCs, the compound structure material library 2200 can fix 50 ligands and their complexes. Correspondingly, the background compound 2300 can be several TMCs (e.g., 20 TMCs) from the 1.37M Pd(II) square planar structures [i.e., Pd(II) square planar] formed by these 50 ligands and the central atom. During generation, 10 new ligands are first generated using LLM 2100, for example, 5 electrically neutral ligands and 5 ionic ligands. Then, new possible combinations of TMCs are constructed using these 10 new ligands, thereby forming a new set of background compounds 2300. In this case, the newly generated ligands can be incorporated into the compound structure material library 2200 to increase the diversity of the generated candidate compounds 2600.
[0043] Figure 4 shows an example from the compound structure library 2200, containing, but not limited to, the specific chemical structures of 50 ligands and their complexes. In the example compound in Figure 4, L... x The ligands are designated L1-L4, and for the same complex, L1-L4 can be independently identical or different from each other. For example, the 50 ligands in Figure 4 are all monodentate ligands, mainly composed of elements such as carbon, hydrogen, nitrogen, oxygen, phosphorus, and sulfur; some molecules also contain halogens (such as fluorine, chlorine, bromine, and iodine). Molecular weights range from lightweight monatomic molecules to medium-molecular-weight compounds. The charge states are 0 (25 ligands) or -1 (25 ligands). The central metal atom in Figure 4 can be Pd or other metal atoms. Figure 3 is a schematic diagram of a system for generating compound structures using LLM according to another embodiment of this application. Figure 5 is a structural diagram of 10 new ligands generated using LLM.
[0044] As shown in Figure 3, the computer system includes one or more processors, a communication unit, etc. The one or more processors include, for example, one or more central processing units (CPUs) 301, and / or one or more graphics processing units (GPUs) 313, etc. The processors can perform various appropriate actions and processes according to executable instructions stored in read-only memory (ROM) 302 or executable instructions loaded from storage unit 308 into random access memory (RAM) 303. The communication unit 312 may include, but is not limited to, a network interface card (NIC), which may include, but is not limited to, an InfiniBand (IB) NIC.
[0045] The processor can communicate with read-only memory 302 and / or random access memory 303 to execute the executable instructions, and is connected to communication unit 312 via bus 304 and communicates with other target devices via communication unit 312, thereby completing the operation corresponding to any of the methods proposed in the embodiments of this application, such as: providing the LLM with a compound structure material library and background compounds with known physicochemical parameters; providing the LLM with limiting conditions and expected physicochemical parameters; using the LLM to generate candidate compounds that meet the provided physicochemical parameters based on the compound structure material library and the physicochemical parameters under the limiting conditions; combining the candidate compounds with the background compounds to form a new combination of background compounds; and providing the LLM with the compound structure material library and the new combination of background compounds to iteratively generate the structure of the expected compound.
[0046] Furthermore, RAM 303 can also store various programs and data required for device operation. CPU 301, ROM 302, and RAM 303 are interconnected via bus 304. With RAM 303 present, ROM 302 is an optional module. RAM 303 stores executable instructions, or executable instructions are written to ROM 302 during runtime. These executable instructions cause processor 301 to perform operations corresponding to the aforementioned communication method. Input / output interface (I / O interface) 305 is also connected to bus 304. Communication unit 312 can be integrated or configured with multiple sub-modules (e.g., multiple IB network cards) linked on the bus.
[0047] The following components are connected to I / O interface 305: input section 306 including keyboard, mouse, etc.; output section 307 including cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; storage section 308 including hard disk, etc.; and communication section 309 including network interface card, such as LAN card, modem, etc. Communication section 309 performs communication processing via a network such as the Internet. Drive 310 is also connected to I / O interface 305 as needed. Removable media 311, such as disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 310 as needed.
[0048] It should be noted that the architecture shown in Figure 3 is only one optional implementation. In practice, the number and type of components in Figure 3 can be selected, reduced, added, or replaced according to actual needs. Different functional components can also be implemented using separate or integrated configurations. For example, the GPU and CPU can be separate or the GPU can be integrated onto the CPU; the communication unit 312 can be separate or integrated onto the CPU or GPU, and so on. All these alternative implementations fall within the scope of protection disclosed in this invention.
[0049] Specifically, according to this application, the process described with reference to flowchart 1 can be implemented as a computer program product. For example, this application proposes a computer program product including computer-readable instructions that, when executed by a processor, perform the following operations: providing the LLM with a compound structure material library and background compounds with known physicochemical parameters; providing the LLM with limiting conditions and expected physicochemical parameters; using the LLM to generate candidate compounds that conform to the provided physicochemical parameters based on the compound structure material library and the physicochemical parameters under the limiting conditions; combining the candidate compounds with the background compounds to form a new combination of background compounds; and providing the LLM with the compound structure material library and the new combination of background compounds to iteratively generate the structure of the expected compound.
[0050] In such an implementation, the computer program product can be downloaded and installed from a network via the communication unit 309, and / or read and installed from the removable medium 311. When the computer program product is executed by the central processing unit (CPU) 301, it performs the functions defined in the method of this application.
[0051] The technical solutions of this application may be implemented in many ways. For example, they may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The order of steps used to illustrate the method is provided only for the purpose of more clearly illustrating the technical solutions. Unless specifically limited, the method steps of this application are not limited to the order specifically described above. Furthermore, in some embodiments, this application may also be implemented as a storage medium for storing computer program products.
[0052] The following examples further illustrate the implementation methods of this application. In these examples, the LLM used is o1-preview. Given that this model is currently more thoroughly trained and performs better with English prompts, examples of prompts in this application are provided in English, along with their meanings. It is important to note that the numbers given in the examples are merely for the purpose of more specifically explaining the methods of this application's implementation methods, and not for limiting purposes.
[0053] Example 1 - Using LLM for bi-objective iterative optimization of TMC based on a given compound structure library
[0054] The user prompt "I have a pool of 50 ligands in a CSV file format below." is used to retrieve a library of compound structures consisting of 50 ligands.
[0055] Next, input the ligand data file in CSV format. Each line of data contains the string "SMILES", the ID, the charge, the connecting atom, and the connecting atom index, as follows:
[0056] c1ccccn1,RUCBEY-subgraph-1,0,N,1
[0057] CP(C)C,WECJIA-subgraph-3,0,P,1
[0058] N#CC,KEYRUB-subgraph-1,0,N,1
[0059] [C-]#[N+]c1c(C)cccc1C,NURKEQ-subgraph-2,0,C,1
[0060] O,MEBXUN-subgraph-1,0,O,1
[0061] n1c(cccc1C)C,BIFMOV-subgraph-1,0,N,1
[0062] CP(C)c1ccccc1,CUJYEL-subgraph-2,0,P,1
[0063] n1ccc(cc1)C,EZEXEM-subgraph-1,0,N,1
[0064] n1cccc(c1)Cl,FOMVUB-subgraph-2,0,N,1
[0065] [C-]#[N+]C(C)(C)C,EFIHEJ-subgraph-3,0,C,1
[0066] CN1[C]N(C)C=C1,LETTEL-subgraph-1,0,C,2
[0067] n1ccn(c1)C,KAKKIR-subgraph-3,0,N,1
[0068] [C-]#[N+]C1CCCCC1,BICRIQ-subgraph-3,0,C,1
[0069] S(=O)(C)C,UPEGAZ-subgraph-2,0,S,1
[0070] O1CCNCC1,CEVJAP-subgraph-2,0,N,1
[0071] n1[nH]c(cc1C)C,BABTUT-subgraph-3,0,N,1
[0072] C(C)NCC,ZEJJEF-subgraph-3,0,N,1
[0073] n1ccc(cc1)N(C)C,KULGAZ-subgraph-2,0,N,1
[0074] [C-]#[O+],CIGDAA-subgraph-1,0,C,1
[0075] c1ccnc(c1)N,HOVMIP-subgraph-3,0,N,1
[0076] N,ULUSIE-subgraph-1,0,N,1
[0077] S(C)C,IBEKUV-subgraph-1,0,S,1
[0078] CS(=O)C,REQSUD-subgraph-2,0,O,1
[0079] NCC,BOSJIF-subgraph-1,0,N,1
[0080] n1c(cc(cc1C)C)C,GUVMEP-subgraph-0,0,N,1
[0081] [Cl-],MAZJIJ-subgraph-0,-1,Cl,1
[0082] [Br-],OBONEA-subgraph-1,-1,Br,1
[0083] [I-],CORTOU-subgraph-2,-1,I,1
[0084] [CH3-],LEVGUO-subgraph-2,-1,C,1
[0085] [C-]1=C(F)C(=C(C(=C1F)F)F)F,REBWEB-subgraph-2,-1,C,1
[0086] C1=C[C-]=CC=C1,DOGPAS-subgraph-1,-1,C,3
[0087] O=N(=O)[O-],IJIMIX-subgraph-1,-1,O,1
[0088] [N-]=[N+]=[N-],PEJGAN-subgraph-1,-1,N,1
[0089] S1(=O)(=O)[N-]C(=O)c2c1cccc2,BIFZEX-subgraph-0,-1,N,1
[0090] [C-]#N,IRIXUC-subgraph-3,-1,C,1
[0091] [S-]C#N,SAYGOO-subgraph-0,-1,S,1
[0092] [F-],UROGIS-subgraph-1,-1,F,1
[0093] [C-](F)(F)F,MAQKEX-subgraph-1,-1,C,1
[0094] O=N[O-],LUQWUQ-subgraph-1,-1,N,1
[0095] C[C-]=O,QAYDID-subgraph-2,-1,C,2
[0096] C1CC(=O)[N-]C1=O,MOYDOV-subgraph-3,-1,N,1
[0097] [C-]1=CC=C(C=C1)F,NIZQUK-subgraph-1,-1,C,1
[0098] [S-]C#N,SAYHIJ-subgraph-1,-1,N,1
[0099] O=[C-]OC,CIQGOY-subgraph-0,-1,C,1
[0100] [S-]c1ccccc1,VUFZUT-subgraph-1,-1,S,1
[0101] [O-]c1ccccc1,ZOQFIU-subgraph-0,-1,O,1
[0102] [S-]c1c(c(cc(c1F)F)F)F,GUQBUQ-subgraph-0,-1,S,1
[0103] [C-]1=CC=C(C=C1)C,LEZYUM-subgraph-2,-1,C,1
[0104] [N-]1C(=O)c2c(C1=O)cccc2,RAJXUX-subgraph-2,-1,N,1
[0105] c1c(C#[C-])cccc1,QEWZOH-subgraph-3,-1,C,4
[0106] Subsequently, the input task, “I am interested in making a Pd(II)square planer transition metal complex (TMC)”, is given, which instructs the LLM to generate a Pd(II)square planar TMC.
[0107] Along with the input to the task, input the constraints and expected physicochemical parameters.
[0108] The following are the suggested keywords for the expected physicochemical parameters (or optimization targets):
[0109] “Design objective: 1. Find new TMCs with both a larger HOMO-LUMO gap and a larger polarizability (these two values will be normalized and evaluated together) than the given TMCs presented later on.” The above optimization objective aims to perform bi-objective optimization, simultaneously pursuing a larger HOMO-LUMO band gap and a higher polarizability. Specifically, it requires that the generated TMCs have a higher HOMO-LUMO band gap and polarizability than the subsequently given TMCs.
[0110] The prompts for the limiting conditions are as follows:
[0111] Constraints:
[0112] The total charge of the TMC should be-1,0,or 1.
[0113] All ligands in the TMC need to be those present in this csv file provided above.”
[0114] The above constraints limit the total charge of the generated TMC to -1, 0, or 1. Furthermore, the above constraints also require that all ligands in the TMC be those already included in the CSV file provided to the LLM.
[0115] Then, a background compound with known physicochemical parameters (in this example, TMC) is provided. The prompt words are as follows:
[0116] “Here is some TMCs their measures of total charge,polarisability,and HOMO-LUMO gap.They are provided in a format of{$TMC,${total charge},${polarisability},${HOMO-LUMO gap}}.The$TMC should be in a format of Pd_$L1_$L2_$L3_$L4,where Pd is the metal center,and$L1,$L2,$L3,and$L4 are the id of the ligands(listed in the csv file)and follow a clockwise ordering.Note that the$TMC has cyclic symmetry for the ligands,so that Pd_$L1_$L2_$L3_$L4,Pd_$L2_$L3_$L4_$L1,Pd_$L3_$L4_$L1_$L2, and Pd_$L4_$L1_$L2_$L3 are the same TMC.Below are the TMCs and their ground-truth total charge, polarizability, and HOMO-LUMO gap.”
[0117] The above prompt explains the data format of the upcoming TMC. The TMC data will then be provided:
[0118] {Pd_KEYRUB-subgraph-1_MEBXUN-subgraph-1_BIFMOV-subgraph-1_RAJXUX-subgraph-2,1,250.599,2.465}
[0119] {Pd_MOYDOV-subgraph-3_ULUSIE-subgraph-1_NURKEQ-subgraph-2_NURKEQ-subgraph-2,1,307.955,2.784}
[0120] {Pd_MEBXUN-subgraph-1_BABTUT-subgraph-3_BOSJIF-subgraph-1_REBWEB-subgraph-2,1,217.567,2.94}
[0121] {Pd_RUCBEY-subgraph-1_ZEJJEF-subgraph-3_EFIHEJ-subgraph-3_DOGPAS-subgraph-1,1,282.377,2.245}
[0122] {Pd_KAKKIR-subgraph-3_QAYDID-subgraph-2_IRIXUC-subgraph-3_GUQBUQ-subgraph-0,-1,229.135,1.674}
[0123] {Pd_SAYHIJ-subgraph-1_LUQWUQ-subgraph-1_SAYGOO-subgraph-0_REQSUD-subgraph-2,-1,178.701,2.251}
[0124] {Pd_REQSUD-subgraph-2_MAQKEX-subgraph-1_OBONEA-subgraph-1_LUQWUQ-subgraph-1,-1,146.398,2.21}
[0125] {Pd_CIQGOY-subgraph-0_BOSJIF-subgraph-1_UROGIS-subgraph-1_NURKEQ-subgraph-2,0,209.993,2.11}
[0126] {Pd_MAQKEX-subgraph-1_DOGPAS-subgraph-1_KULGAZ-subgraph-2_ZEJJEF-subgraph-3,0,272.928,2.306}
[0127] {Pd_ZOQFIU-subgraph-0_HOVMIP-subgraph-3_NURKEQ-subgraph-2_MEBXUN-subgraph-1,1,285.01,1.786}
[0128] Finally, additional information can be provided to the LLM, such as the number of TMCs to be generated, suggested generation methods, and restrictions on the output data format. You can request the LLM to generate 5 new TMCs, and suggest using crossover and / or mutation in genetic algorithms to generate them. Furthermore, you can request that the LLM avoid generating duplicate TMCs. The following are some prompts:
[0129] "Please propose 5new TMCs that achieve my design objective. You can make ligand crossover(ie,swap ligands in the TMCs)or ligand mutations(ie,substitute ligands in TMCs with some other ligands in the pool)based on your chemistry knowledge to achieve the design objective.You can also directly propose new TMCs by taking advantage of your chemistry knowledge to propose Pd TMCs with maximize HOMO-LUMO gap and maximize polarizability.Be sure to use the id of the provided pool of 50ligands and control the total charge of the TMC to be-1,0,or 1.In addition,please DO NOT propose duplicated TMCs that I have listed above.
[0130] Your output should follow the format:{<< <explanation>>>:$EXPLANATION,<< <tmc>>>:[$TMC],<<<TOTAL_CHARGE>>>:${total charge},<< <polarisability>>>:$POLARISABILITY,<<<HOMO-LUMO gap> >>:${HOMO-LUMO gap}}.Here are the requirements that you should fulfill:
[0131] 1.$EXPLANATION should be your analysis of why the new TMC would achieve my design objective.
[0132] 2.The$TMC should be in a format of[Pd_$L1_$L2_$L3_$L4],where Pd is the center metal.$L1,$L2,$L3,and$L4 should be the id of the ligands(listed in the csv file)and follow a clockwise ordering.
[0133] 3 $POLARISABILITY and${HOMO-LUMO gap}are your predictions of the polarisability and HOMO-LUMO gap for$TMC based on your chemistry knowledge and provided data.”
[0134] Next, LLM will generate candidate compounds as required. The following shows the forms of the candidate compounds generated by LLM:
[0135] {<< <explanation>>>:This TMC combines the large,aromatic BIFMOV-subgraph-1 ligand with the strongly electron-withdrawing EFIHEJ-subgraph-3 ligand.The combination should increase polarizability due to the extended π-system,while the electron-withdrawing groups should widen the HOMO-LUMO gap.The neutral MEBXUN-subgraph-1(water)ligand balances the charge.,<< <tmc>>>:[Pd_BIFMOV-subgraph-1_EFIHEJ-subgraph-3_BIFMOV-subgraph-1_MEBXUN-subgraph-1],<<<TOTAL_CHARGE>>>:1,<< <polarisability>>>:310.5,<<<HOMO-LUMO gap> >>:2.95}
[0136] Then, the TMC generated above can be combined with the background compound initially provided to the LLM and iteratively generated according to the above process.
[0137] Table 1 shows examples 1-10 of complexes generated by the LLM model based on the dual optimization objectives of polarity and HOMO-LUMO bandgap. The newly generated complexes 1-10 are the top 10 complexes with the largest product of bandgap and polarity among the complex structures generated through multiple iterations. That is, the complex 1-10 has both large polarity and a relatively large HOMO-LUMO bandgap.
[0138] Example 2 - Iterative optimization of single-objective TMC with novel ligands using LLM
[0139] The user prompt "I have a pool of 50 ligands in a CSV file format below." leads to a compound structure library 2200 composed of 50 ligands. It's important to note that newly generated ligands will be added to the compound structure library 2200 in subsequent iterations. For example, if m new ligands are generated in the first round, a chemical structure library of 50+m ligands will be provided to the LLM in subsequent iterations.
[0140] Next, input the ligand data file in CSV format. Each line of data contains the string "SMILES", ID, charge, connecting atom, and connecting atom number. The prompts are similar to the format given in Example 1, and will not be repeated here.
[0141] Subsequently, input the generation task, "I am interested in making a Pd(II) square planer transition metal complex (TMC)," thereby instructing LLM to generate a Pd(II) square planar TMC.
[0142] Along with the input for the task generation, input the constraints and expected physicochemical parameters. The prompts are as follows:
[0143] "My design objective is to maximize its polarisability while making the total charge of the TMC be -1, 0 or 1." The above optimization objective aims to perform single-objective optimization, maximizing only the polarisability; the constraint is to ensure that the total charge of the generated TMC is -1, 0 or 1.
[0144] Then, a background compound with known physicochemical parameters (in this example, TMC) is provided. The prompt words are as follows:
[0145] "I have made 100 TMCs and measured their total charge and polarisability.They are provided in a format of{$TMC,${total charge},${polarisability},{error message(if any)}}.
[0146] The$TMC should be in a format of Pd_$L1_$L2_$L3_$L4,where Pd is the metal center,$L1,$L2,$L3,and$L4 are the id of the ligands(listed in the csv file)and follow a clockwise ordering.Note that the$TMC has cyclic symmetry for the ligands so that Pd_$L1_$L2_$L3_$L4,Pd_$L2_$L3_$L4_$L1,Pd_$L3_$L4_$L1_$L2, and Pd_$L4_$L1_$L2_$L3 are the same TMC. Below are the TMCs and their ground-truth total charge and polarisability.”
[0147] The above prompts explain the data format of the TMCs that will be provided. The TMC data will then be provided with prompts similar to those in Example 1, and will not be repeated here. Similarly, if n new TMCs are generated in the first round of generation, these n newly generated TMCs will be added to the background compounds in subsequent iterations of optimization. At this point, there will be 100+n TMCs.
[0148] Subsequently, the LLM can be instructed using natural language to construct the TMC using new ligands beyond the initially provided 50 ligands, with the following prompts:
[0149] "Grounded on your chemistry knowledge, look at the pattern of the provided data and think about what makes a ligand combination give extremely large polarisability for a TMC. Then based on your chemistry knowledge and100example TMCs provided, can you give me TEN new TMCs that would help me further maximize the polarisability? I am expecting that the new proposed TMC has a larger polarisability than the largest in given TMCs. For the TEN proposed TMCs, you must design new ligands to achieve the objective.”
[0150] Finally, additional information can be provided to the LLM, depending on the purpose of the generation. This includes the number of TMCs to be generated, suggested generation methods, and restrictions on the output data format. For example, the LLM can be asked to generate 5 or 10 new TMCs. Alternatively, it can be suggested that the LLM utilize crossover and / or mutation in genetic algorithms to generate TMCs. Furthermore, the LLM can be asked not to generate duplicate TMCs.
[0151] Then, new prompt words can be constructed using the chemical structure material library containing new ligands generated above and the TMC containing new ligands, and the process can be iteratively generated according to the above procedure.
[0152] Table 2 shows examples 1-10 of the complexes generated by optimization based on the polarity optimization objective using the LLM model. All of the newly generated complexes 1-10 exhibit excellent polarity.
[0153] Example 3 - Excellent performance of compounds generated using LLM
[0154] LLM-based compound generation exhibits significant advantages over other algorithms. To demonstrate the superiority of LLM-based compound generation over other algorithms, Figure 6 compares the performance of single-step compound generation with other methods.
[0155] In the process of generating compounds using LLM, a non-iterative approach (i.e., a single-step method) can also be employed. For example, a similar approach to that given in Example 1 can be used, except that after each generation of candidate compounds by LLM, they are not re-fed to LLM for iterative optimization. LLM can be used to generate 10 compounds at a time, repeated 20 times, and then the top 20 compounds out of the 200 generated compounds can be analyzed and compared. Figure 6 shows that even using this single-step method, generating compounds using LLM demonstrates superiority over other algorithms.
[0156] Figure 6a shows the HOMO-LUMO bandgap distribution of the top 20 TMCs out of 200 generated by each method. Blue represents the random algorithm, red represents the genetic algorithm (GA), orange represents Claude-3.5-Sonnet, green represents O1-Preview, purple represents O1-Mini, and sky blue represents GPT-4O. These results demonstrate the superior performance of LLM in TMC generation tasks with limited sample sizes. Figure 6b shows the cumulative HOMO-LUMO bandgap probability and its overall distribution for all TMCs proposed by various methods. Figure 6c shows the proportion of valid (within the 1.37M TMC space) and unique (i.e., non-repeating) TMCs among the 200 generated TMCs. Figure 6, subplot d, illustrates the HOMO-LUMO bandgap distribution of the top 20 TMCs out of 200 generated TMCs when different numbers of known TMCs are provided in the prompt. Only the results from claude-3.5-sonnet (orange) and o1-preview (green) are shown in the figure. Additionally, the TMCs with the largest HOMO-LUMO bandgap derived from claude-3.5-sonnet and o1-preview are also displayed. The average HOMO-LUMO bandgap of the top 20 TMCs derived from the random algorithm is represented by the blue dashed line. Notably, o1-preview demonstrates superior performance compared to the random algorithm using only one initial TMC sample. Increasing the number of initial TMCs generally improves the HOMO-LUMO bandgap of the model-generated TMCs (although there is a marginal effect), indicating that LLM can achieve satisfactory results with limited initial data, providing researchers with a balance between data requirements and performance. The colors of different atoms in the d-sub-diagram of Figure 6 are represented as follows: palladium (Pd) is sky blue, carbon (C) is gray, nitrogen (N) is blue, oxygen (O) is red, phosphorus (P) is purple, fluorine (F) is green, and hydrogen (H) is white. The purple dashed lines in TMC represent the coordination bonds between Pd(II) and the ligands.
[0157] As another example, Figure 7 illustrates the analysis of the top 20 TMCs out of 200 generated candidate TMCs using an iterative optimization method. o1-preview identifies better TMCs faster than the genetic algorithm. In Figure 7, the random algorithm is represented in blue, the genetic algorithm in orange, and the o1-preview method in green and gray. Green data indicates the case where, in each iteration, only the 20 TMCs with the largest HOMO-LUMO band gap from the previous round of candidate TMCs and known TMCs are used in the next round of iterations. Gray data indicates the case where, in each iteration, all historical information (including physicochemical parameters) of the generated TMCs is retained, and this historical information is used in the next iteration. Although the two o1-preview methods show little difference in results in the early iterations (5 iterations), the performance gap becomes apparent as optimization progresses. Retaining all historical data enables the model to generate TMCs with higher HOMO-LUMO band gaps. The above description is merely an implementation method of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of protection involved in this application is not limited to technical solutions formed by specific combinations of the above-mentioned technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-mentioned technical features or their equivalent features without departing from the described technical concept. For example, technical solutions formed by substituting the above-mentioned features with technical features disclosed in this application (but not limited to) that have similar functions.< / polarisability> < / tmc> < / explanation> < / polarisability> < / tmc> < / explanation>
Claims
1. A method for generating compound structures using a large language model, comprising: The large language model is provided with a library of compound structures and background compounds with known physicochemical parameters. The large language model is provided with constraints and expected physicochemical parameters; Using the large language model, candidate compounds that meet the provided physicochemical parameters are generated based on the compound structure library and the physicochemical parameters under the specified constraints; The candidate compounds are combined with the background compounds to form a new combination of background compounds; The combination of the compound structure material library and the new background compounds is provided to the large language model to iteratively generate the desired compound structure.
2. The method according to claim 1, wherein: When providing background compounds with known physicochemical parameters to the large language model, the physicochemical parameters of the background compounds are also provided to the large language model. When the combination of the compound structure material library and the new background compound is provided to the large language model, the physicochemical parameters of the new background compound are also provided to the large language model.
3. The method of claim 1, wherein, The expected physicochemical parameters include at least two different physicochemical parameters.
4. The method of claim 1, wherein, The compounds include organic compounds, inorganic compounds, organometallic complexes, polymers, small molecules, and biological macromolecules.
5. The method of claim 1, wherein, Combining the candidate compound with the background compound to form a new background compound combination includes: The candidate compounds and the top K compounds with the best physicochemical parameters among the background compounds are combined to form a new combination of background compounds.
6. The method of claim 1, wherein, Using the large language model, based on the compound structure library and the physicochemical parameters, candidate compounds that meet the provided physicochemical parameters are generated under the specified constraints, including: The candidate compounds are generated based on new compound materials outside the aforementioned compound structure material library.
7. The method of claim 6, wherein, After generating the candidate compound based on new compound materials outside the compound structure material library, the method further includes integrating the new compound materials into the compound structure material library.
8. The method of claim 1, wherein, The compound is a transition metal complex, the compound structure library is a ligand library, and the physicochemical parameters include the HOMO-LUMO band gap and polarizability.
9. The method of claim 8, wherein, The limiting conditions include the charge of transition metal complexes, metal atoms and / or ligands.
10. A system for generating a compound structure using a large language model, comprising a memory, a processor, and a computer program stored on the memory, wherein the computer program comprises the following steps: The processor executes the computer program to perform the following steps: The large language model is provided with a library of compound structures and background compounds with known physicochemical parameters. The large language model is provided with constraints and expected physicochemical parameters; Using the large language model, candidate compounds that meet the provided physicochemical parameters are generated based on the compound structure library and the physicochemical parameters under the specified constraints; The candidate compounds are combined with the background compounds to form a new combination of background compounds; The combination of the compound structure material library and the new background compounds is provided to the large language model to iteratively generate the desired compound structure.
11. A computer program product comprising computer program instructions that, when executed by a processor, perform the following steps: The large language model is provided with a library of compound structures and background compounds with known physicochemical parameters. The large language model is provided with constraints and expected physicochemical parameters; Using the large language model, candidate compounds that meet the provided physicochemical parameters are generated based on the compound structure library and the physicochemical parameters under the specified constraints; The candidate compounds are combined with the background compounds to form a new combination of background compounds; The combination of the compound structure material library and the new background compounds is provided to the large language model to iteratively generate the desired compound structure.