Inverse Synthesis Route Planning Method and System Based on Multimodal Large Model
Through the cross-modal alignment and reaction mechanism inference of the molecular structure of the multimodal large language model, synthesizers are constructed and reactants are generated, which solves the problems of low efficiency of existing inverse synthesis methods and lack of explanatory information fusion, and realizes efficient and accurate reverse synthesis route planning.
Patent Information
- Application Number
- CN202510574155.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-05-06
AI Technical Summary
The existing inverse synthesis methods have low efficiency in path planning, complex processes and over-reliance on expert experience, and the multimodal information fusion lacks explanatory nature, making it difficult to achieve efficient and accurate chemical reaction path prediction.
A multimodal large language model is adopted to construct a multimodal inference task chain through cross-modal alignment of molecular structures, reaction mechanism inference, synthesis sub-construction and reactant generation, and an inverse synthesis route is generated, and the best solution is screened through autoregression generation.
It realizes intelligent reverse derivation from target molecules to reactants, improves the accuracy, generalization ability and interpretability of reverse synthesis prediction, and provides efficient and accurate chemical synthesis and drug design solutions.
Smart Images

Figure CN120089250B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the cross - field of chemoinformatics and artificial intelligence, and specifically relates to an inverse synthesis path planning method based on a multimodal large language model (MLLM), aiming to achieve efficient and accurate prediction of chemical reaction paths and reactants, and providing an intelligent solution for fields such as chemical synthesis and drug design. Background Art
[0002] Inverse synthesis analysis is a key task in the field of chemical synthesis, aiming to deduce feasible reaction paths and starting materials from target molecules in a reverse manner. Traditional inverse synthesis methods mainly rely on the experience of chemical experts, with low efficiency and difficulty in dealing with complex molecular structures. In recent years, with the development of machine learning and artificial intelligence technologies, data - driven inverse synthesis methods have gradually emerged, but still face many challenges.
[0003] Existing inverse synthesis methods are mainly divided into template - based methods, template - free methods, and semi - template - based methods. Template - based methods reduce the search space by matching predefined reaction templates, but their generalization ability is limited and it is difficult to handle novel reactions. Template - free methods transform the inverse synthesis problem into a translation problem and use deep learning models (such as Transformer) to generate reaction paths, but usually require complex test integration strategies and have a large computational burden. Semi - template - based methods decompose the product into synthons and gradually complete the reactant prediction, but usually require multi - stage models, with a cumbersome process and lack of knowledge transfer.
[0004] In recent years, general large language models (LLMs) have achieved remarkable success in natural language processing tasks, but there is still a large gap in their application in chemical tasks. Although there have been studies attempting to apply large language models to the chemical field, their performance in inverse synthesis tasks is still far from that of dedicated models. In addition, multimodal large language models (MLLMs) show powerful reasoning abilities by integrating multimodal information such as text and images, but their application in the chemical field is still in its infancy. Existing methods usually fuse multimodal information through simple concatenation or attention - based aggregation, ignoring the domain gap between molecular and reaction knowledge and lacking interpretability.
[0005] Therefore, it is necessary to develop an inverse synthesis model that can effectively integrate various modalities of molecular information, combine chemical knowledge with data - driven methods to enhance the accuracy, generalization ability, and interpretability of inverse synthesis prediction. Summary of the Invention
[0006] The purpose of the present invention is to solve the above problems in the prior art and provide an inverse synthesis route planning method and system based on a multimodal large model.
[0007] The specific technical solutions adopted by the present invention are as follows:
[0008] In a first aspect, the present invention provides an inverse synthesis route planning method based on a multimodal large model, which includes:
[0009] S1. Extract the structural features of the three-dimensional molecular structure of the target reaction product through a molecular encoder, and then project the structural features into the representation space of the large language model through a linear layer, so as to obtain the projected structural features and input them into the large language model, and the large language model predicts the SMILES sequence of the target reaction product;
[0010] S2. Drive the large language model through prompt words to sequentially execute a multimodal reasoning task chain composed of a reaction type prediction subtask, a synthon construction subtask, and a reactant generation subtask according to the predicted SMILES sequence, and obtain an inverse synthesis route for synthesizing the target reaction product from the reactants;
[0011] In the reaction type prediction subtask, the large language model needs to reason about the reaction mechanism according to the predicted SMILES sequence of the target reaction product and predict the reaction type for synthesizing the target reaction product;
[0012] In the synthon construction subtask, the large language model needs to assign each atom in the target reaction product to the corresponding synthon, generate an atom-level synthon assignment mask through an autoregressive prediction method, so as to decompose the target reaction product into different synthons, and then predict the activity of each atom on each synthon, determine whether each atom belongs to the reaction center atom, and generate an atom-level activity mask for each synthon;
[0013] In the reactant generation subtask, the large language model needs to combine the information already inferred in the current multimodal reasoning task chain and finally generate the SMILES sequence of the reactant corresponding to each synthon.
[0014] As a preference of the above first aspect, the three-dimensional molecular structure of the target reaction product is generated by the SMILES sequence of the target reaction product through the RDKit tool.
[0015] As a preference of the above first aspect, in the process of generating the atom-level synthon assignment mask, the mask tokens used include a first start token, a first end token, and classification tokens for different synthons, and all the mask tokens used are pre-added to the vocabulary of the large language model.
[0016] As a preference of the above first aspect, in the process of generating the atom-level activity mask for each synthon, the mask tokens used include a second start token, a second end token, and a binary classification token indicating whether the atom belongs to the reaction center atom, and all the mask tokens used are pre-added to the vocabulary of the large language model.
[0017] Preferably, in the first aspect above, the molecular encoder adopts a pre-trained Uni-Mol model, and the large language model adopts a pre-trained Qwen2-0.5B model or LLAMA model.
[0018] Preferably, in the first aspect above, the model framework composed of the molecular encoder, the linear layer, and the large language model needs to be pre-trained in two stages. In the first stage, the molecular encoder and the large language model need to be frozen, and the linear layer is optimized by the error loss between the target reaction product SMILES sequence predicted by the large language model and its true label. In the second stage, the molecular encoder, the linear layer, and the large language model need to be unfrozen, and the molecular encoder, the linear layer, and the large language model are jointly optimized by the error loss between the target reaction product SMILES sequence and the reactant SMILES sequence predicted by the large language model and their respective true labels.
[0019] Preferably, in the first aspect above, the large language model needs to execute the multi-modal reasoning task chain multiple times to obtain different retrosynthetic routes, and then screen the best retrosynthetic route from the generated set of retrosynthetic routes according to the currently specified screening principle.
[0020] In a second aspect, the present invention provides a retrosynthetic route planning system based on a multi-modal large model, which includes:
[0021] A data input module for a user to input the three-dimensional molecular structure of the target reaction product;
[0022] A retrosynthetic route generation module for generating a retrosynthetic route according to the three-dimensional molecular structure of the target reaction product input by the user, according to the retrosynthetic route planning method based on a multi-modal large model described in any one of the above first aspects;
[0023] A result output module for outputting or visually displaying the result data generated in the retrosynthetic route generation module.
[0024] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the retrosynthetic route planning method based on a multi-modal large model described in any one of the above first aspects is implemented.
[0025] In a fourth aspect, the present invention provides a computer electronic device, which includes a memory and a processor;
[0026] The memory is used to store a computer program;
[0027] The processor is used to implement the inverse synthesis route planning method based on the multimodal large model according to any of the solutions in the first aspect as described above when executing the computer program.
[0028] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0029] 1) Through steps such as cross-modal alignment of molecular structures, reaction mechanism reasoning, synthon construction, reactant generation, and inverse synthesis path screening, the present invention can achieve intelligent reverse derivation from the target molecule to the reactants, providing an efficient and accurate solution for fields such as chemical synthesis and drug design.
[0030] 2) By constructing a bridging task, the present invention forces the large language model to first predict the SMILES representation of the target reaction product based on structural features at the beginning of the reasoning chain, thereby aligning different modal representations of the molecule in the representation space of the large language model and providing more accurate multimodal information support for subsequent inverse synthesis reasoning.
[0031] 3) The present invention constructs a multimodal reasoning task chain composed of a reaction type prediction sub-task, a synthon construction sub-task, and a reactant generation sub-task. Through reaction mechanism reasoning and synthon construction, the system can provide clear guidance on chemical reaction mechanisms, enhancing the interpretability of the reasoning process.
[0032] 4) The present invention can obtain multiple possible reaction paths using the autoregressive generation method and select the optimal solution through a screening mechanism, further improving the experimental feasibility of the inverse synthesis path. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is a schematic diagram of the steps of the inverse synthesis route planning method based on the multimodal large model;
[0034] Figure 2 It is the basic flowchart of the inverse synthesis route planning method of the present invention;
[0035] Figure 3 It is the overall framework diagram of the inverse synthesis route planning of the present invention;
[0036] Figure 4 It is a schematic diagram of the model data flow during the reasoning process of the present invention for performing inverse synthesis route planning;
[0037] Figure 5 It is a schematic diagram of the module composition of the inverse synthesis route planning system based on the multimodal large model;
[0038] Figure 6 It is a schematic diagram of the structure of the computer electronic device;
[0039] Figure 7Schematic diagram of the retrosynthetic route planning process with phenylacetic acid as the target reaction product in the embodiments of the present invention. Detailed implementation manners
[0040] To make the above objects, features and advantages of the present invention more obvious and understandable, the following will describe the detailed implementation manners of the present invention with reference to the accompanying drawings. Many specific details are set forth in the following description to facilitate a thorough understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below. The technical features in the various embodiments of the present invention can be combined correspondingly without conflict.
[0041] In the description of the present invention, it should be understood that the terms "first" and "second" are only used for the purpose of distinguishing descriptions, and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features.
[0042] The present invention provides a retrosynthetic route planning method based on a multimodal large model, aiming to overcome the limitations of existing retrosynthetic methods in low path planning efficiency, complex processes, and over-reliance on expert experience. Through steps such as cross-modal alignment of molecular structures, reaction mechanism reasoning, synthon construction, reactant generation, and retrosynthetic path screening, it realizes intelligent reverse derivation from the target molecule to the reactant, providing an efficient and accurate solution for fields such as chemical synthesis and drug design.
[0043] To facilitate an accurate understanding of the essence of the technical solution of the present invention, before discussing the implementation of the specific technical solution, relevant technical concepts involved in the present invention are defined and introduced in advance.
[0044] Definition 1: Large language model
[0045] A large language model (LLM) is a natural language processing (NLP) model based on deep learning. It usually adopts a large-scale neural network architecture (such as Transformer) and is pre-trained with large-scale text data. The large language model can understand, generate, and reason about natural language text, and performs well in tasks such as question answering, translation, code generation, and scientific research.
[0046] Definition 2: Multimodal large language model
[0047] A Multimodal Large Language Model (MLLM) is a large language model that can process multiple data modalities (such as text, images, videos, point clouds, molecular structures, etc.). By integrating data from different modalities, MLLMs enhance the model's understanding and reasoning capabilities for complex tasks (such as chemical reaction prediction, drug design, material discovery, etc.). Different from traditional large language models, MLLMs can simultaneously process and correlate cross-modal information to achieve more accurate scientific reasoning.
[0048] Definition 3: Retrosynthesis
[0049] Retrosynthesis is a key problem in chemical synthesis, which involves deducing suitable precursor molecules (reactants) and their synthetic routes from a target molecule (product). Retrosynthesis usually combines chemical knowledge, empirical rules, database information, and artificial intelligence models for analysis, aiming to optimize the synthetic route, thereby improving synthesis efficiency and reducing costs. This technique is widely used in the fields of drug design, organic synthesis, and materials science to help chemists efficiently plan synthetic routes.
[0050] Definition 4: Reaction Center
[0051] The Reaction Center refers to the atoms or atomic groups that directly participate in the breaking or formation of chemical bonds in a chemical reaction. It is the core region where structural changes occur during the reaction and determines the type and path of the reaction. In organic chemistry, the identification of the reaction center is the basis for understanding and designing reaction mechanisms, as well as the key to optimizing reaction conditions and improving product selectivity.
[0052] Definition 5: Simplified Molecular Input Line Entry System
[0053] The Simplified Molecular Input Line Entry System (SMILES) is a text format used to describe the structure of chemical molecules. Through a series of character encodings, SMILES represents the atoms, chemical bonds, and their connection methods in a molecule, converting complex chemical molecular structures into computer-readable text. This format facilitates the recording, transmission, and processing of molecules and is widely used in chemical databases and molecular modeling tools as the standard input format.
[0054] Definition 6: Molecular Graph
[0055] A molecular graph is a molecular representation method based on graph theory, where atoms are regarded as vertices of the graph and chemical bonds are regarded as edges of the graph. Molecular graphs can visually describe the topological structure of molecules and provide structured data input for machine learning models such as graph neural networks. Molecular graphs are widely used in tasks such as molecular property prediction, reaction prediction, and drug discovery, and can effectively capture the complex relationships between molecules, improving the accuracy of molecular modeling.
[0056] Definition 7: Synthon
[0057] A synthon is a key concept in retrosynthetic analysis, referring to the intermediate molecular fragments generated during the decomposition of the target molecule. These fragments represent the key bond-breaking positions of the target molecule in the retrosynthetic pathway and can be further transformed into practically usable reactants through chemical reactions. The identification and segmentation of synthons are the core steps of the retrosynthetic task, directly affecting the feasibility and efficiency of the reaction pathway. By rationally designing synthons, chemists can more efficiently plan synthetic routes, thus accelerating the progress in fields such as new drug development and material design.
[0058] As Figure 1 shown, in a preferred embodiment of the present invention, a retrosynthetic route planning method based on a multimodal large model is provided, which includes two steps S1 and S2. The specific implementation of the two steps is introduced below.
[0059] S1. Extract the structural features of the three-dimensional molecular structure of the target reaction product through a molecular encoder, and then project the structural features into the representation space of a large language model (LLM) through a linear layer, so as to obtain the projected structural features and input them into the large language model, and the large language model predicts the SMILES sequence of the target reaction product.
[0060] It should be noted that the target reaction product in the present invention refers to the reaction product targeted by the retrosynthetic route planning. The three-dimensional molecular structure of the target reaction product input into the molecular encoder can be directly input with a known three-dimensional molecular structure. However, due to the incomplete three-dimensional molecular structure data of the reaction product, for a reaction product without a three-dimensional molecular structure, its SMILES sequence can be obtained first, and then a three-dimensional molecular structure can be generated based on this SMILES sequence through the RDKit tool.
[0061] In the above step S1, a molecule-reaction bridging task is actually constructed. In this task, the three-dimensional molecular structure of the target product is first input, and its structural features are extracted by a dedicated molecular encoder. Subsequently, these features are projected into the representation space of the large language model through a linear layer, that is, the SMILES representation of the reaction product in the one-dimensional text space is obtained to achieve the initial alignment between the molecular structure and the text modality. To further enhance the alignment effect, the large language model is forced to predict the SMILES representation of the target reaction product based on the extracted projected structural features. This process not only activates the pre-trained knowledge of the molecular encoder but also ensures the fine-grained alignment between the molecular topology information and the SMILES text modality, laying a solid foundation for subsequent retrosynthetic reasoning.
[0062] It should be noted that this molecule-reaction bridging task does not simply generate the SMILES representation of the target reaction product and input it into the large language model. Instead, it requires the large language model to predict the SMILES representation of the target reaction product based on the structural features at the beginning of the inference chain. In traditional retrosynthetic route planning methods, the spatial gap between the molecular-level structure encoder and the reaction-level SMILES space is ignored in the initial input information. Therefore, when using simple input-level fusion, the ambiguity of the subtle association between the product and the reactant is caused. However, as a delicate task, these precise correlations are crucial for retrosynthesis. Therefore, the goal of the present invention is to first bridge this gap by establishing a delicate combination at the beginning of the inference chain. The present invention provides precise complementary information by binding the structural features and the molecular formula atom by atom through the bridging task of forcing the large language model to predict the SMILES molecular formula of the product based on the molecular structure features. Moreover, since the structural features are the only possible source of information to solve this bridging task, the pre-trained knowledge in the molecular encoder will be fully activated, thus providing fine-grained aligned multimodal information for subsequent inference tasks.
[0063] S2. Drive the large language model through prompts (Prompts) to sequentially execute a multimodal inference task chain composed of a reaction type prediction subtask, a synthon construction subtask, and a reactant generation subtask based on the predicted SMILES sequence, and obtain a retrosynthetic route for synthesizing the target reaction product from the reactants.
[0064] As Figure 2As shown in the figure, it is a flowchart showing the reverse synthesis route planning method of the present invention. First, through the bridging task in step S1 above, the SMILES sequence of the target reaction product is predicted based on the three-dimensional molecular structure diagram (this SMILES sequence can be used to align with the true label of the SMILES sequence during the training stage to calculate the loss function), and then by constructing prompts, each subtask in the multi-modal inference task chain is sequentially executed. Each execution of the multi-modal inference task chain can obtain a reverse synthesis route. In practical applications, the large language model needs to execute the multi-modal inference task chain multiple times to obtain different reverse synthesis routes, and then screen the best reverse synthesis route from the set of generated reverse synthesis routes according to the currently specified screening principle. Based on the above process of the reverse synthesis route planning method, the corresponding overall reverse synthesis route planning framework is as Figure 3 shown, which shows an exemplary execution process of the bridging task and the multi-modal inference task chain. Among them, the cross-modal alignment of the molecular structure in step 1 corresponds to the bridging task in step S1, and steps 2-4 correspond to the multi-modal inference task chain in step S2.
[0065] In the multi-modal inference task chain of the present invention, it includes a reaction type prediction subtask, a synthon construction subtask, and a reactant generation subtask. The specific implementation methods of each subtask will be introduced below.
[0066] In the reaction type prediction subtask of the present invention, the large language model needs to perform reasoning on the reaction mechanism based on the predicted SMILES sequence of the target reaction product, and predict the reaction type for synthesizing the target reaction product.
[0067] The function of the above reaction type prediction subtask is to perform reasoning on the reaction mechanism on the basis of cross-modal alignment of the molecular structure to accurately guide reverse synthesis analysis. The reaction type is crucial in reverse synthesis reasoning, directly affecting the determination of reaction sites, reaction conditions, and possible intermediates and precursor molecules, thus determining the rationality of the reverse synthesis path. Therefore, in this step, through the extensive chemical knowledge contained in the large language model, combined with the molecular structure characteristics and the aligned SMILES information, the reaction mechanism and reaction type that the target molecule may involve are intelligently inferred. This step not only effectively reduces the reverse synthesis search space, but also provides clear guidance on the chemical reaction mechanism for the subsequent synthon construction, making the reasoning process more reasonable and interpretable, thereby improving the reliability and practicality in complex molecular synthesis design. After completing the reaction mechanism reasoning, it enters the synthon construction stage, aiming to efficiently split the molecular structure of the reaction product into multiple synthon units to provide clear guidance for the subsequent reactant generation.
[0068] In the synthetic building block task of the present invention, the large language model needs to assign each atom in the target reaction product to the corresponding synthetic building block, generate an atom-level synthetic building block assignment mask through autoregressive prediction, so as to decompose the target reaction product into different synthetic building blocks, and then predict the activity of each atom on each synthetic building block, determine whether each atom belongs to the reaction center atom, and generate an atom-level activity mask for each synthetic building block.
[0069] It should be noted that the autoregressive prediction method in the present invention means generating a mask label of the synthetic building block to which each atom in the target reaction product belongs one by one. When generating a mask label for each atom, the autoregressive method needs to be used, that is, considering the mask labels of all the atoms that have been generated to generate the mask label of the current atom.
[0070] It should be noted that in the process of generating the atom-level synthetic building block assignment mask, the mask labels used include the first start label, the first end label, and the classification labels of different synthetic building blocks, and all the mask labels used are pre-added to the vocabulary of the large language model. Therefore, assuming that there are a total of n atoms in the target reaction product, the final atom-level synthetic building block assignment mask is a sequence containing n + 2 labels. The first one is the first start label, the last one is the first end label, and the middle n ones respectively correspond to the synthetic building blocks to which the n atoms belong. However, it should be noted that the synthetic building block classification label here is only used to distinguish different synthetic building blocks, and different synthetic building block classification labels can be distinguished as long as they have different indexes. For example <syn1> , <syn2> , <syn3>It can represent the first synthon, the second synthon, and the third synthon. The forms of the first start marker and the first end marker are not limited, and <sosyn>and <eosyn>They are used as the first start tag and the first end tag respectively. Generally, 5 to 10 tags are sufficient for n.
[0071] Similarly, assuming a synthon contains m atoms, the atom-level activity mask generated for this synthon is a sequence containing m + 2 tags. The first tag is the second start tag, the last tag is the second end tag, and the middle m tags respectively record whether the m atoms belong to the reaction center atoms. In the process of generating the atom-level activity mask for each synthon, the mask tags used include the second start tag, the second end tag, and the binary classification tag indicating whether the atom belongs to the reaction center atom.
[0072] The binary classification tag of the reaction center atoms here is only used to distinguish whether different atoms belong to the reaction center atoms. The forms of the two types of tags are not limited, and similarly, the forms of the second start tag and the second end tag are not limited. In the embodiments of the present invention, <sorc>Indicates the second start marker, <eorc>Indicates the second end marker, <nrc>Represents a non-reactive center atom, <rc>Represents the reaction center atom.
[0073] It should be noted that since the general large language model's vocabulary often does not have the above special markers, all the mask markers used in the process of generating the atomic-level synthon assignment mask and the mask markers used in the process of generating the atomic-level activity mask for each synthon need to be added to the large language model's vocabulary in advance to expand the vocabulary and participate in the model's training process.
[0074] In the embodiments of the present invention, referring to Figure 3 As shown, after completing the reaction mechanism reasoning, the synthon construction stage can be entered, so that the molecular structure of the reaction product can be efficiently split into multiple synthon units through the synthon construction subtask in the multimodal reasoning task chain, providing clear guidance for the subsequent generation of reactants. The specific vocabulary expansion, model training, and reasoning processes can be implemented as follows:
[0075] 1) Structured Marking and Vocabulary Expansion. To achieve precise segmentation of the target molecular structure, first expand the large language model's vocabulary by introducing the above special mask markers to represent different synthons. These markers can automatically identify and label which synthon each atom belongs to according to the structural characteristics of the molecule, that is, clarify which synthon unit each atom belongs to in the retrosynthesis process. This attribution relationship directly reflects the splitting strategy of the molecule in the retrosynthetic path. In this way, the model can directly output the atomic-level assignment mask in the text generation task.
[0076] 2) Data-Driven Atomic Assignment Learning. During the training process, use molecular data with atomic assignment labels to guide the LLM to learn how to identify the appropriate atomic attribution from the molecular structure. Through the autoregressive prediction method, the model gradually generates the synthon sequence to ensure that each atom is correctly assigned to the corresponding synthon.
[0077] 3) Reaction Center Identification and Activity Prediction. To further improve the chemical rationality of the atomic assignment strategy, a reaction center identification and activity prediction mechanism is introduced. This mechanism provides key chemical constraints for atomic assignment by analyzing which atoms in the synthon may act as reaction centers in the retrosynthesis process. Specifically, first, by deeply fusing the structural features extracted by the molecular encoder with the synthon features, the ability to capture key reaction sites in the molecule is enhanced, and then by predicting which atoms may belong to the reaction center atoms, the atomic assignment results in the synthon conform to the laws of real chemical reactions.
[0078] Thus, after completing the synthon splitting through the synthon construction subtask, the reactant generation stage can be entered.
[0079] In the reactant generation subtask of the present invention, the large language model needs to combine the information inferred in the current multi-modal reasoning task chain and finally generate the reactant SMILES sequences corresponding to each synthon. Specifically, using the autoregressive generation method, the possible reactant SMILES sequences are gradually generated based on the determined synthons. In this process, it is necessary to refer to the information inferred in the current multi-modal reasoning task chain, including the reaction type and the activity information of the synthons, to ensure that the retrosynthetic path conforms to the basic chemical reaction logic.
[0080] In addition, in practical applications, the above large language model needs to execute the multi-modal reasoning task chain multiple times to obtain different retrosynthetic routes, and then screen the best retrosynthetic route from the generated set of retrosynthetic routes according to the currently specified screening principle. After generating multiple possible retrosynthetic routes, each retrosynthetic route represents a different retrosynthetic scheme. The user can specify the screening principle to the large model to enable it to screen the best retrosynthetic route, such as maximizing the yield, minimizing the cost, or minimizing the reaction condition requirements as the screening conditions. Of course, the model can also screen the best retrosynthetic route based on the default screening principle, that is, comprehensively considering factors such as the availability of raw materials, reaction conditions, and reaction yields to screen the best retrosynthetic scheme from the generated set of routes.
[0081] It should be noted that the molecular encoder used in the embodiments of the present invention adopts the pre-trained Uni-Mol model, and the large language model adopts the pre-trained Qwen2-0.5B model, but the present invention is not limited thereto, and other molecular encoders and large language models (such as the LLAMA model) can also be used.
[0082] In addition, it should be noted that as Figure 4 shown, the model framework composed of the above-mentioned molecular encoder Uni-Mol, linear layer, and large language model Qwen2-0.5B is shown, which constitutes a multi-modal large language model. The model data flow in the process of performing retrosynthetic route planning reasoning of this multi-modal large language model needs to be trained before being used for actual reasoning tasks. In the embodiments of the present invention, a two-stage training method can be adopted. In the first stage, the molecular encoder and the large language model need to be frozen, and the linear layer is optimized by the error loss between the target reaction product SMILES sequence predicted by the large language model and its true label. In the second stage, the molecular encoder, linear layer, and large language model need to be unfrozen, and the molecular encoder, linear layer, and large language model are jointly optimized by the error losses between the target reaction product SMILES sequence and the reactant SMILES sequence predicted by the large language model and their respective true labels. The fine-tuning of the large language model can be achieved by adding a low-rank adaptation LoRA module to the Qwen2-0.5B model.
[0083] In summary, through steps such as cross-modal alignment of molecular structures, reaction mechanism reasoning, synthon construction, reactant generation, and retrosynthetic route screening, the present invention can achieve intelligent reverse derivation from the target molecule to the reactants, providing an efficient and accurate solution for fields such as chemical synthesis and drug design.
[0084] It should be noted that the method steps shown in S1~S2 can essentially be implemented in the form of a computer program and integrated into a software system in the form of functional modules for invocation.
[0085] Thus, based on the same inventive concept, as Figure 5 shown, the present invention also provides a retrosynthetic route planning system based on a multimodal large model, which includes:
[0086] A data input module for the user to input the three-dimensional molecular structure of the target reaction product;
[0087] A retrosynthetic route generation module for generating a retrosynthetic route according to the three-dimensional molecular structure of the target reaction product input by the user, in accordance with the retrosynthetic route planning method based on the multimodal large model described above;
[0088] A result output module for outputting or visually displaying the result data generated in the retrosynthetic route generation module.
[0089] In addition, based on the same inventive concept, as Figure 6 shown, the present invention also provides a computer electronic device corresponding to the retrosynthetic route planning method based on a multimodal large model provided in the above embodiment, which includes a memory and a processor;
[0090] The memory is used to store a computer program;
[0091] The processor is used to implement the retrosynthetic route planning method based on the multimodal large model described above when executing the computer program;
[0092] In addition, when the logical instructions in the above memory are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0093] Accordingly, based on the same inventive concept, the present invention provides a computer-readable storage medium corresponding to an inverse synthesis route planning method based on a multimodal large model. A computer program is stored on the storage medium, and when the computer program is executed by a processor, the inverse synthesis route planning method as described above can be implemented.
[0094] Accordingly, based on the same inventive concept, the present invention provides a computer program product including a computer program / instructions. When the computer program / instructions are executed by a processor, the inverse synthesis route planning method as described above can be implemented.
[0095] Specifically, in the computer-readable storage media of the above three embodiments, the stored computer program is executed by a processor, and the foregoing steps S1 to S2 can be executed.
[0096] It can be understood that the above storage medium may include a Random Access Memory (RAM), and may also include a Non-Volatile Memory (NVM), such as at least one disk memory. At the same time, the storage medium may also be various media such as a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc that can store program codes.
[0097] It can be understood that the above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0098] In addition, it should be noted that those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the above-described system can refer to the corresponding process in the foregoing method embodiments, and will not be elaborated herein. In the embodiments provided in the present application, the division of steps or modules in the system and method is only a logical function division, and there may be other division methods in actual implementation. For example, multiple modules or steps may be combined or integrated together, and a module or step may also be split.
[0099] To demonstrate the advantages of the inverse synthesis route planning method based on the multi-modal large model shown in S1 - S2 of the present invention, in Example 1 below, Figure 3 the shown model framework is applied to an exemplary target product to demonstrate its technical effects. Meanwhile, in Example 2, the performance of this method is tested through specific data sets.
[0100] Example 1
[0101] In this example, taking phenylacetic acid as the target reaction product, a method for inverse synthesis route planning based on a multi-modal large model is elaborated in detail. Its model framework and inference process are still as Figure 3 and Figure 4 shown. Figure 7 The basic process of the method in this example is given, and the specific steps are described as follows:
[0102] Step 1: Cross-modal alignment of molecular structures
[0103] Given the SMILES representation C1=CC=C(C=C1)CC(=O)O of the target product phenylacetic acid, the corresponding three-dimensional molecular structure is generated using the RDKit tool. Then, the structural features are extracted through the pre-trained molecular encoder Uni-Mol, including the spatial positions of atoms, the connection modes of bonds, and the overall conformation of the molecule. Subsequently, these features are projected into the feature space of the large language model Qwen2-0.5B through a linear mapping layer to achieve the initial alignment of the molecular structure and the text modality. To further enhance the alignment effect, the fine-tuned Qwen2-0.5B model is used to predict the SMILES representation of the product based on the extracted molecular features. This process not only activates the pre-trained knowledge of the molecular encoder but also ensures the fine-grained alignment of the molecular topological information and the SMILES text modality, laying a solid foundation for subsequent inverse synthesis reasoning.
[0104] Step 2: Reaction mechanism reasoning
[0105] Construct reaction mechanism reasoning prompts, such as "Predict the reaction types for synthesizing phenylacetic acid C1=CC=C(C=C1)CC(=O)O". Using the Qwen2-0.5B model to reason, the possible reaction types are found to be "nucleophilic substitution reaction and hydrolysis reaction", effectively narrowing the inverse synthesis search space and providing clear guidance for subsequent synthon construction, making the reasoning process more reasonable and interpretable.
[0106] Step 3: Synthon construction
[0107] The synthon construction stage aims to efficiently split the molecular structure of the reaction product into multiple synthon units to provide guidance for subsequent reactant generation. This process includes the following sub-steps:
[0108] Step 3.1: Structured Marking and Vocabulary Expansion. To achieve precise segmentation of the target molecular structure, first expand the vocabulary of the Qwen2-0.5B model before model training, introducing a series of special structured markings <sosyn> , <syn1> , <syn2> , …, <eosyn>are used to represent different synthons. These markers are used to label which synthon unit each atom belongs to during the retrosynthesis process.
[0109] Step 3.2: Data-driven atom assignment learning. During the previous model training process, molecular data with atom assignment labels is used to guide the Qwen2-0.5B model to learn how to identify appropriate atom assignments from the molecular structure. Construct atom assignment prompts, such as "Assign each atom in the reaction product C1=CC=C(C=C1)CC(=O)O to the corresponding synthon, <sosyn>Indicates the start, <eosyn>Indicates the end, <syni>"indicating which synthon an atom belongs to". Through the autoregressive prediction method, the Qwen2-0.5B model gradually generates the synthon sequence to ensure that each atom is correctly assigned to the corresponding synthon. See Figure 3 and Figure 7 As shown, in this embodiment, C1=CC=C(C=C1)C belongs to synthon 1, i.e., <SYN1>, while C(=O)O belongs to synthon 2, i.e., <SYN2>. Therefore, the sequence of synthons is <sosyn> , <syn1> , <syn1> , …, <syn1> , <syn2> , <syn2> , <syn2> , <eosyn>. For the sake of simplicity of representation, several are omitted here <syn1>。
[0110] Step 3.3: Reaction center identification and activity prediction. For each synthon subunit, construct the prompt "Predict the synthon <syni>The activity of each atom in it is judged to determine whether it is located at the reaction center. <sorc>Indicates the start, <eorc>Indicates the end, <nrc>Represents a non-reactive center atom, <rc>"Indicates the reaction center atom". Use the Qwen2-0.5B model to predict whether each atom in each synthon unit is a reaction center or connected to a leaving group. For the synthon <syn1>, prediction result <sorc> , <nrc> , <nrc> , …, <rc> , <eorc>, for synthon <syn2>, prediction result <sorc> , <rc> , <nrc> ,…, <nrc> , <eorc>, it can be determined that the reaction center should be located at the carbon-carbon bond outside the benzene ring.
[0111] Step 4: Generation of Reactants and Screening of Retrosynthetic Routes
[0112] After the synthon splitting is completed, it enters the final stage of reactant generation and retrosynthetic route screening. The goal of this stage is to use the information obtained in the previous steps to predict reactant molecules, reaction routes, and screen the best retrosynthetic scheme. The specific approach is as follows:
[0113] Step 4.1: Using the Qwen2-0.5B model, the reactant SMILES sequence gradually generated based on the synthon C1=CC=C(C=C1)C is C1=CC=C(C=C1)CBr, and the reactant it represents is benzyl bromide. The reactant SMILES sequence gradually generated based on the synthon C(=O)O is [C-]#N.[Na+], and the reactant it represents is sodium cyanide. Thus, a route for synthesizing phenylacetic acid using benzyl bromide and sodium cyanide based on the hydrolysis method of phenylacetonitrile is obtained.
[0114] Step 4.2: Repeat Steps 1 - Step 4.1 to generate multiple possible reaction routes. For example, Route 2 is the reductive carboxylation of benzaldehyde O=CC1=CC=CC=C1 and carbon dioxide O=C=O under alkaline conditions, and Route 3 is the Friedel-Crafts alkylation reaction of benzene C1=CC=CC=C1 and chloroacetic acid O=C(O)CCl catalyzed by a Lewis acid. Using the chemical knowledge reasoning in the large model, it can be known that the reaction conditions of Route 2 and Route 3 are relatively harsh and the applicable range is limited. Therefore, the best retrosynthetic scheme screened from the generated route set is Route 1.
[0115] Example 2
[0116] In this example, the inverse synthesis route planning method based on the multi-modal large model shown in S1~S2 above (denoted as the method of the present invention) was used to conduct experiments on the publicly available chemical reaction dataset Mol-Instructions and USPTO-50K to verify the performance of the method. Among them, Mol-Instructions contains a total of 143,536 chemical reactions, and USPTO-50K contains approximately 50,000 chemical reactions.
[0117] In this example, the model selection and parameter settings are as follows: Uni-Mol is used as the molecular encoder, and Qwen2-0.5B is used as the LLM. The rank parameter of LoRA is set to 8, and the scaling parameter is set to 16. The number of synthon structural labels is set to the empirical value of 10. During the training process, 10 rounds of training were conducted on the Mol-Instructions dataset, and 2 rounds of training were conducted on the USPTO-50K dataset.
[0118] In this embodiment, the experimental evaluation metrics include:
[0119] (1) Top-k: The Top-k accuracy measures whether the correct result is included in the top k predictions of the model. The higher the value, the more accurate the prediction.
[0120] (2) Levenshtein: The Levenshtein distance measures the edit distance between the predicted product and the true product at the character level. The smaller the value, the more accurate the prediction.
[0121] (3) RDK FTS: RDK fingerprint similarity. This metric generates molecular fingerprints based on the RDKit tool and evaluates the structural similarity between the predicted molecule and the true molecule in terms of their RDK fingerprints. The higher the value, the more accurate the prediction.
[0122] (4) MACCS FTS: MACCS fingerprint similarity. This metric generates molecular fingerprints based on MACCS Keys and evaluates the structural similarity between the predicted molecule and the true molecule in terms of their molecular fingerprints. The higher the value, the more accurate the prediction.
[0123] (5) Morgan FTS: Morgan fingerprint similarity. This metric generates molecular fingerprints based on the Morgan algorithm and evaluates the structural similarity between the predicted molecule and the true molecule in terms of their Morgan fingerprints. The higher the value, the more accurate the prediction.
[0124] In this embodiment, the method of the present invention is compared with several methods. The retrosynthesis prediction methods used as controls are divided into two groups. The first group is the large model-based methods, including: (1) VICUNA: A general large language model that performs retrosynthesis reasoning through instruction tuning. (2) MOL-INS: A large language model based on molecular instruction learning that directly generates reaction paths. (3) INSTRUCTMOL: A large language model method that combines instruction tuning and template-free prediction, suitable for general retrosynthesis tasks. (4) UNIMOT: A large model method that converts molecular graphs into natural language inputs, emphasizing molecular semantic consistency.
[0125] The second group is the method based on professional small models, including: (1) RETROSIM: A template-based method that relies on a fixed template library for prediction. (2) LOCALRETRO: A template-based method that improves template adaptability through context pruning. (3) RETROXPERT: A semi-template-based method that first predicts the reaction center and then generates reactants. (4) G2RETRO: A semi-template-based method that predicts reaction paths based on graph editing operations. (5) GTA: A template-free method that simulates atomic-level operations in chemical reactions through graph-to-action. (6) GRAPH2EDITS: A template-free method that, based on graph neural networks, gradually reconstructs reactants by learning sequences of editing operations on molecular structures.
[0126] The experimental results of the method of the present invention and each control method based on large models on the Mol-Instructions dataset are shown in Table 1. The experimental results of the method of the present invention and each control method based on professional small models on the USPTO-50K dataset are shown in Table 2.
[0127] Table 1 Comparison results of the present method and methods based on large models on the Mol-Instructions dataset
[0128]
[0129] Table 2 Comparison results of the present method and other methods based on professional small models on the USPTO-50K dataset
[0130]
[0131] As shown in Table 1, on the Mol-Instructions dataset, the method of this embodiment is significantly superior to all control methods based on large models in various metrics. As shown in Table 2, on the USPTO-50K dataset, the method of this embodiment is superior to other professional small models based on templates, semi-templates, and template-free.
[0132] In addition, in this embodiment, an ablation study of the method of the present invention is also carried out on the USPTO-50K dataset to verify the effectiveness of the key structures in the model. The naming of each ablation variant experiment is as follows: (1) w / o Bridging: It means removing the cross-modal alignment task of molecular structures. This task was originally used to align molecular structures with text information to promote modal fusion. (2) w / oType: It means removing the reaction mechanism reasoning task, that is, not using reaction categories for additional guidance. (3) w / o Syn. Seg.& Act.: It means removing both atomic assignment and reaction center identification in the synthon construction task. The combined removal can evaluate the overall role of the structure reasoning module. (4) w / o Syn. Act.: It means removing only the reaction center identification in the synthon construction task. (5) w / o All: It means removing all tasks and only retaining the original SMILES sequence as the input, which is equivalent to a pure language modeling method. In this embodiment, experiments are carried out on different ablation variants, and the performance on the USPTO-50K dataset is shown in Table 3 as follows:
[0133] Table 3 Effects of Ablation Variants on the USPTO-50K Dataset
[0134]
[0135] As shown in Table 3, by comparing different ablation variant experiments, it can be seen from w / o Bridging that the cross-modal alignment task of molecular structures plays a key role in cross-modal information integration and is the core component for improving the model's understanding ability. From w / oType, it can be known that reaction types, as prior knowledge, can enhance the model's structure perception ability and assist in reaction path judgment. From w / o Syn. Act., it can be seen that removing the reaction center identification module in the synthon construction task leads to a decrease in model performance, indicating that this module is important for accurately identifying reaction centers and improving prediction accuracy. Further, from the results of w / o Syn. Seg. &Act., it can be seen that when both the atomic assignment and reaction center identification tasks are removed simultaneously, the model performance drops more significantly, indicating that the cooperation between the two helps to enhance the model's reasoning ability. Through the above analysis, the segmented task design helps the model better understand chemical reactions. Each task performs its own functions and cooperates with each other in the model, jointly improving the model performance.
[0136] The above are the preferred embodiments of the present invention. The present invention should not be limited to the content disclosed in this embodiment and the accompanying drawings. Any equivalent or modified implementation completed without departing from the spirit disclosed by the present invention falls within the protection scope of the present invention. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.< / eorc> < / nrc> < / nrc> < / rc> < / sorc> < / eorc> < / rc> < / nrc> < / nrc> < / sorc> < / rc> < / nrc> < / eorc> < / sorc> < / syni> < / eosyn> < / syn2> < / syn2> < / syn2> < / syn1> < / syn1> < / syn1> < / sosyn> < / syni> < / eosyn> < / sosyn> < / eosyn> < / syn2> < / syn1> < / sosyn> < / rc> < / nrc> < / eorc> < / sorc> < / eosyn> < / sosyn> can be adopted in the embodiments of the present invention. <sosyn>and <eosyn>They are used as the first start tag and the first end tag respectively. Generally, 5 to 10 tags are sufficient for n.
[0071] Similarly, assuming a synthon contains m atoms, the atom-level activity mask generated for this synthon is a sequence containing m + 2 tags. The first tag is the second start tag, the last tag is the second end tag, and the middle m tags respectively record whether the m atoms belong to the reaction center atoms. In the process of generating the atom-level activity mask for each synthon, the mask tags used include the second start tag, the second end tag, and the binary classification tag indicating whether the atom belongs to the reaction center atom.
[0072] The binary classification tag of the reaction center atoms here is only used to distinguish whether different atoms belong to the reaction center atoms. The forms of the two types of tags are not limited, and similarly, the forms of the second start tag and the second end tag are not limited. In the embodiments of the present invention, <sorc>Indicates the second start marker, <eorc>Indicates the second end marker, <nrc>Represents a non-reactive center atom, <rc>Represents the reaction center atom.
[0073] It should be noted that since the general large language model's vocabulary often does not have the above special markers, all the mask markers used in the process of generating the atomic-level synthon assignment mask and the mask markers used in the process of generating the atomic-level activity mask for each synthon need to be added to the large language model's vocabulary in advance to expand the vocabulary and participate in the model's training process.
[0074] In the embodiments of the present invention, referring to Figure 3 As shown, after completing the reaction mechanism reasoning, the synthon construction stage can be entered, so that the molecular structure of the reaction product can be efficiently split into multiple synthon units through the synthon construction subtask in the multimodal reasoning task chain, providing clear guidance for the subsequent generation of reactants. The specific vocabulary expansion, model training, and reasoning processes can be implemented as follows:
[0075] 1) Structured Marking and Vocabulary Expansion. To achieve precise segmentation of the target molecular structure, first expand the large language model's vocabulary by introducing the above special mask markers to represent different synthons. These markers can automatically identify and label which synthon each atom belongs to according to the structural characteristics of the molecule, that is, clarify which synthon unit each atom belongs to in the retrosynthesis process. This attribution relationship directly reflects the splitting strategy of the molecule in the retrosynthetic path. In this way, the model can directly output the atomic-level assignment mask in the text generation task.
[0076] 2) Data-Driven Atomic Assignment Learning. During the training process, use molecular data with atomic assignment labels to guide the LLM to learn how to identify the appropriate atomic attribution from the molecular structure. Through the autoregressive prediction method, the model gradually generates the synthon sequence to ensure that each atom is correctly assigned to the corresponding synthon.
[0077] 3) Reaction Center Identification and Activity Prediction. To further improve the chemical rationality of the atomic assignment strategy, a reaction center identification and activity prediction mechanism is introduced. This mechanism provides key chemical constraints for atomic assignment by analyzing which atoms in the synthon may act as reaction centers in the retrosynthesis process. Specifically, first, by deeply fusing the structural features extracted by the molecular encoder with the synthon features, the ability to capture key reaction sites in the molecule is enhanced, and then by predicting which atoms may belong to the reaction center atoms, the atomic assignment results in the synthon conform to the laws of real chemical reactions.
[0078] Thus, after completing the synthon splitting through the synthon construction subtask, the reactant generation stage can be entered.
[0079] In the reactant generation subtask of the present invention, the large language model needs to combine the information inferred in the current multi-modal reasoning task chain and finally generate the reactant SMILES sequences corresponding to each synthon. Specifically, using the autoregressive generation method, the possible reactant SMILES sequences are gradually generated based on the determined synthons. In this process, it is necessary to refer to the information inferred in the current multi-modal reasoning task chain, including the reaction type and the activity information of the synthons, to ensure that the retrosynthetic path conforms to the basic chemical reaction logic.
[0080] In addition, in practical applications, the above large language model needs to execute the multi-modal reasoning task chain multiple times to obtain different retrosynthetic routes, and then screen the best retrosynthetic route from the generated set of retrosynthetic routes according to the currently specified screening principle. After generating multiple possible retrosynthetic routes, each retrosynthetic route represents a different retrosynthetic scheme. The user can specify the screening principle to the large model to enable it to screen the best retrosynthetic route, such as maximizing the yield, minimizing the cost, or minimizing the reaction condition requirements as the screening conditions. Of course, the model can also screen the best retrosynthetic route based on the default screening principle, that is, comprehensively considering factors such as the availability of raw materials, reaction conditions, and reaction yields to screen the best retrosynthetic scheme from the generated set of routes.
[0081] It should be noted that the molecular encoder used in the embodiments of the present invention adopts the pre-trained Uni-Mol model, and the large language model adopts the pre-trained Qwen2-0.5B model, but the present invention is not limited thereto, and other molecular encoders and large language models (such as the LLAMA model) can also be used.
[0082] In addition, it should be noted that as Figure 4 shown, the model framework composed of the above-mentioned molecular encoder Uni-Mol, linear layer, and large language model Qwen2-0.5B is shown, which constitutes a multi-modal large language model. The model data flow in the process of performing retrosynthetic route planning reasoning of this multi-modal large language model needs to be trained before being used for actual reasoning tasks. In the embodiments of the present invention, a two-stage training method can be adopted. In the first stage, the molecular encoder and the large language model need to be frozen, and the linear layer is optimized by the error loss between the target reaction product SMILES sequence predicted by the large language model and its true label. In the second stage, the molecular encoder, linear layer, and large language model need to be unfrozen, and the molecular encoder, linear layer, and large language model are jointly optimized by the error losses between the target reaction product SMILES sequence and the reactant SMILES sequence predicted by the large language model and their respective true labels. The fine-tuning of the large language model can be achieved by adding a low-rank adaptation LoRA module to the Qwen2-0.5B model.
[0083] In summary, through steps such as cross-modal alignment of molecular structures, reaction mechanism reasoning, synthon construction, reactant generation, and retrosynthetic route screening, the present invention can achieve intelligent reverse derivation from the target molecule to the reactants, providing an efficient and accurate solution for fields such as chemical synthesis and drug design.
[0084] It should be noted that the method steps shown in S1~S2 can essentially be implemented in the form of a computer program and integrated into a software system in the form of functional modules for invocation.
[0085] Thus, based on the same inventive concept, as Figure 5 shown, the present invention also provides a retrosynthetic route planning system based on a multimodal large model, which includes:
[0086] A data input module for the user to input the three-dimensional molecular structure of the target reaction product;
[0087] A retrosynthetic route generation module for generating a retrosynthetic route according to the three-dimensional molecular structure of the target reaction product input by the user, in accordance with the retrosynthetic route planning method based on the multimodal large model described above;
[0088] A result output module for outputting or visually displaying the result data generated in the retrosynthetic route generation module.
[0089] In addition, based on the same inventive concept, as Figure 6 shown, the present invention also provides a computer electronic device corresponding to the retrosynthetic route planning method based on a multimodal large model provided in the above embodiment, which includes a memory and a processor;
[0090] The memory is used to store a computer program;
[0091] The processor is used to implement the retrosynthetic route planning method based on the multimodal large model described above when executing the computer program;
[0092] In addition, when the logical instructions in the above memory are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0093] Accordingly, based on the same inventive concept, the present invention provides a computer-readable storage medium corresponding to an inverse synthesis route planning method based on a multimodal large model. A computer program is stored on the storage medium, and when the computer program is executed by a processor, the inverse synthesis route planning method as described above can be implemented.
[0094] Accordingly, based on the same inventive concept, the present invention provides a computer program product including a computer program / instructions. When the computer program / instructions are executed by a processor, the inverse synthesis route planning method as described above can be implemented.
[0095] Specifically, in the computer-readable storage media of the above three embodiments, the stored computer program is executed by a processor, and the foregoing steps S1 to S2 can be executed.
[0096] It can be understood that the above storage medium may include a Random Access Memory (RAM), and may also include a Non-Volatile Memory (NVM), such as at least one disk memory. At the same time, the storage medium may also be various media such as a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc that can store program codes.
[0097] It can be understood that the above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0098] In addition, it should be noted that those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the above-described system can refer to the corresponding process in the foregoing method embodiments, and will not be elaborated herein. In the embodiments provided in the present application, the division of steps or modules in the system and method is only a logical function division, and there may be other division methods in actual implementation. For example, multiple modules or steps may be combined or integrated together, and a module or step may also be split.
[0099] To demonstrate the advantages of the inverse synthesis route planning method based on the multi-modal large model shown in S1 - S2 of the present invention, in Example 1 below, Figure 3 the shown model framework is applied to an exemplary target product to demonstrate its technical effects. Meanwhile, in Example 2, the performance of this method is tested through specific data sets.
[0100] Example 1
[0101] In this example, taking phenylacetic acid as the target reaction product, a method for inverse synthesis route planning based on a multi-modal large model is elaborated in detail. Its model framework and inference process are still as Figure 3 and Figure 4 shown. Figure 7 The basic process of the method in this example is given, and the specific steps are described as follows:
[0102] Step 1: Cross-modal alignment of molecular structures
[0103] Given the SMILES representation C1=CC=C(C=C1)CC(=O)O of the target product phenylacetic acid, the corresponding three-dimensional molecular structure is generated using the RDKit tool. Then, the structural features are extracted through the pre-trained molecular encoder Uni-Mol, including the spatial positions of atoms, the connection modes of bonds, and the overall conformation of the molecule. Subsequently, these features are projected into the feature space of the large language model Qwen2-0.5B through a linear mapping layer to achieve the initial alignment of the molecular structure and the text modality. To further enhance the alignment effect, the fine-tuned Qwen2-0.5B model is used to predict the SMILES representation of the product based on the extracted molecular features. This process not only activates the pre-trained knowledge of the molecular encoder but also ensures the fine-grained alignment of the molecular topological information and the SMILES text modality, laying a solid foundation for subsequent inverse synthesis reasoning.
[0104] Step 2: Reaction mechanism reasoning
[0105] Construct reaction mechanism reasoning prompts, such as "Predict the reaction types for synthesizing phenylacetic acid C1=CC=C(C=C1)CC(=O)O". Using the Qwen2-0.5B model to reason, the possible reaction types are found to be "nucleophilic substitution reaction and hydrolysis reaction", effectively narrowing the inverse synthesis search space and providing clear guidance for subsequent synthon construction, making the reasoning process more reasonable and interpretable.
[0106] Step 3: Synthon construction
[0107] The synthon construction stage aims to efficiently split the molecular structure of the reaction product into multiple synthon units to provide guidance for subsequent reactant generation. This process includes the following sub-steps:
[0108] Step 3.1: Structured Marking and Vocabulary Expansion. To achieve precise segmentation of the target molecular structure, first expand the vocabulary of the Qwen2-0.5B model before model training, introducing a series of special structured markings <sosyn> , <syn1> , <syn2> , …, <eosyn>are used to represent different synthons. These markers are used to label which synthon unit each atom belongs to during the retrosynthesis process.
[0109] Step 3.2: Data-driven atom assignment learning. During the previous model training process, molecular data with atom assignment labels is used to guide the Qwen2-0.5B model to learn how to identify appropriate atom assignments from the molecular structure. Construct atom assignment prompts, such as "Assign each atom in the reaction product C1=CC=C(C=C1)CC(=O)O to the corresponding synthon, <sosyn>Indicates the start, <eosyn>Indicates the end, <syni>"indicating which synthon an atom belongs to". Through the autoregressive prediction method, the Qwen2-0.5B model gradually generates the synthon sequence to ensure that each atom is correctly assigned to the corresponding synthon. See Figure 3 and Figure 7 As shown, in this embodiment, C1=CC=C(C=C1)C belongs to synthon 1, i.e., <SYN1>, while C(=O)O belongs to synthon 2, i.e., <SYN2>. Therefore, the sequence of synthons is <sosyn> , <syn1> , <syn1> , …, <syn1> , <syn2> , <syn2> , <syn2> , <eosyn>. For the sake of simplicity of representation, several are omitted here <syn1>。
[0110] Step 3.3: Reaction center identification and activity prediction. For each synthon subunit, construct the prompt "Predict the synthon <syni>The activity of each atom in it is judged to determine whether it is located at the reaction center. <sorc>Indicates the start, <eorc>Indicates the end, <nrc>Represents a non-reactive center atom, <rc>"Indicates the reaction center atom". Use the Qwen2-0.5B model to predict whether each atom in each synthon unit is a reaction center or connected to a leaving group. For the synthon <syn1>, prediction result <sorc> , <nrc> , <nrc> , …, <rc> , <eorc>, for synthon <syn2>, prediction result <sorc> , <rc> , <nrc> ,…, <nrc> , <eorc>, it can be determined that the reaction center should be located at the carbon-carbon bond outside the benzene ring.
[0111] Step 4: Generation of Reactants and Screening of Retrosynthetic Routes
[0112] After the synthon splitting is completed, it enters the final stage of reactant generation and retrosynthetic route screening. The goal of this stage is to use the information obtained in the previous steps to predict reactant molecules, reaction routes, and screen the best retrosynthetic scheme. The specific approach is as follows:
[0113] Step 4.1: Using the Qwen2-0.5B model, the reactant SMILES sequence gradually generated based on the synthon C1=CC=C(C=C1)C is C1=CC=C(C=C1)CBr, and the reactant it represents is benzyl bromide. The reactant SMILES sequence gradually generated based on the synthon C(=O)O is [C-]#N.[Na+], and the reactant it represents is sodium cyanide. Thus, a route for synthesizing phenylacetic acid using benzyl bromide and sodium cyanide based on the hydrolysis method of phenylacetonitrile is obtained.
[0114] Step 4.2: Repeat Steps 1 - Step 4.1 to generate multiple possible reaction routes. For example, Route 2 is the reductive carboxylation of benzaldehyde O=CC1=CC=CC=C1 and carbon dioxide O=C=O under alkaline conditions, and Route 3 is the Friedel-Crafts alkylation reaction of benzene C1=CC=CC=C1 and chloroacetic acid O=C(O)CCl catalyzed by a Lewis acid. Using the chemical knowledge reasoning in the large model, it can be known that the reaction conditions of Route 2 and Route 3 are relatively harsh and the applicable range is limited. Therefore, the best retrosynthetic scheme screened from the generated route set is Route 1.
[0115] Example 2
[0116] In this example, the inverse synthesis route planning method based on the multi-modal large model shown in S1~S2 above (denoted as the method of the present invention) was used to conduct experiments on the publicly available chemical reaction dataset Mol-Instructions and USPTO-50K to verify the performance of the method. Among them, Mol-Instructions contains a total of 143,536 chemical reactions, and USPTO-50K contains approximately 50,000 chemical reactions.
[0117] In this example, the model selection and parameter settings are as follows: Uni-Mol is used as the molecular encoder, and Qwen2-0.5B is used as the LLM. The rank parameter of LoRA is set to 8, and the scaling parameter is set to 16. The number of synthon structural labels is set to the empirical value of 10. During the training process, 10 rounds of training were conducted on the Mol-Instructions dataset, and 2 rounds of training were conducted on the USPTO-50K dataset.
[0118] In this embodiment, the experimental evaluation metrics include:
[0119] (1) Top-k: The Top-k accuracy measures whether the correct result is included in the top k predictions of the model. The higher the value, the more accurate the prediction.
[0120] (2) Levenshtein: The Levenshtein distance measures the edit distance between the predicted product and the true product at the character level. The smaller the value, the more accurate the prediction.
[0121] (3) RDK FTS: RDK fingerprint similarity. This metric generates molecular fingerprints based on the RDKit tool and evaluates the structural similarity between the predicted molecule and the true molecule in terms of their RDK fingerprints. The higher the value, the more accurate the prediction.
[0122] (4) MACCS FTS: MACCS fingerprint similarity. This metric generates molecular fingerprints based on MACCS Keys and evaluates the structural similarity between the predicted molecule and the true molecule in terms of their molecular fingerprints. The higher the value, the more accurate the prediction.
[0123] (5) Morgan FTS: Morgan fingerprint similarity. This metric generates molecular fingerprints based on the Morgan algorithm and evaluates the structural similarity between the predicted molecule and the true molecule in terms of their Morgan fingerprints. The higher the value, the more accurate the prediction.
[0124] In this embodiment, the method of the present invention is compared with several methods. The retrosynthesis prediction methods used as controls are divided into two groups. The first group is the large model-based methods, including: (1) VICUNA: A general large language model that performs retrosynthesis reasoning through instruction tuning. (2) MOL-INS: A large language model based on molecular instruction learning that directly generates reaction paths. (3) INSTRUCTMOL: A large language model method that combines instruction tuning and template-free prediction, suitable for general retrosynthesis tasks. (4) UNIMOT: A large model method that converts molecular graphs into natural language inputs, emphasizing molecular semantic consistency.
[0125] The second group is the method based on professional small models, including: (1) RETROSIM: A template-based method that relies on a fixed template library for prediction. (2) LOCALRETRO: A template-based method that improves template adaptability through context pruning. (3) RETROXPERT: A semi-template-based method that first predicts the reaction center and then generates reactants. (4) G2RETRO: A semi-template-based method that predicts reaction paths based on graph editing operations. (5) GTA: A template-free method that simulates atomic-level operations in chemical reactions through graph-to-action. (6) GRAPH2EDITS: A template-free method that, based on graph neural networks, gradually reconstructs reactants by learning sequences of editing operations on molecular structures.
[0126] The experimental results of the method of the present invention and each control method based on large models on the Mol-Instructions dataset are shown in Table 1. The experimental results of the method of the present invention and each control method based on professional small models on the USPTO-50K dataset are shown in Table 2.
[0127] Table 1 Comparison results of the present method and methods based on large models on the Mol-Instructions dataset
[0128]
[0129] Table 2 Comparison results of the present method and other methods based on professional small models on the USPTO-50K dataset
[0130]
[0131] As shown in Table 1, on the Mol-Instructions dataset, the method of this embodiment is significantly superior to all control methods based on large models in various metrics. As shown in Table 2, on the USPTO-50K dataset, the method of this embodiment is superior to other professional small models based on templates, semi-templates, and template-free.
[0132] In addition, in this embodiment, an ablation study of the method of the present invention is also carried out on the USPTO-50K dataset to verify the effectiveness of the key structures in the model. The naming of each ablation variant experiment is as follows: (1) w / o Bridging: It means removing the cross-modal alignment task of molecular structures. This task was originally used to align molecular structures with text information to promote modal fusion. (2) w / oType: It means removing the reaction mechanism reasoning task, that is, not using reaction categories for additional guidance. (3) w / o Syn. Seg.& Act.: It means removing both atomic assignment and reaction center identification in the synthon construction task. The combined removal can evaluate the overall role of the structure reasoning module. (4) w / o Syn. Act.: It means removing only the reaction center identification in the synthon construction task. (5) w / o All: It means removing all tasks and only retaining the original SMILES sequence as the input, which is equivalent to a pure language modeling method. In this embodiment, experiments are carried out on different ablation variants, and the performance on the USPTO-50K dataset is shown in Table 3 as follows:
[0133] Table 3 Effects of Ablation Variants on the USPTO-50K Dataset
[0134]
[0135] As shown in Table 3, by comparing different ablation variant experiments, it can be seen from w / o Bridging that the cross-modal alignment task of molecular structures plays a key role in cross-modal information integration and is the core component for improving the model's understanding ability. From w / oType, it can be known that reaction types, as prior knowledge, can enhance the model's structure perception ability and assist in reaction path judgment. From w / o Syn. Act., it can be seen that removing the reaction center identification module in the synthon construction task leads to a decrease in model performance, indicating that this module is important for accurately identifying reaction centers and improving prediction accuracy. Further, from the results of w / o Syn. Seg. &Act., it can be seen that when both the atomic assignment and reaction center identification tasks are removed simultaneously, the model performance drops more significantly, indicating that the cooperation between the two helps to enhance the model's reasoning ability. Through the above analysis, the segmented task design helps the model better understand chemical reactions. Each task performs its own functions and cooperates with each other in the model, jointly improving the model performance.
[0136] The above are the preferred embodiments of the present invention. The present invention should not be limited to the content disclosed in this embodiment and the accompanying drawings. Any equivalent or modified implementation completed without departing from the spirit disclosed by the present invention falls within the protection scope of the present invention. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.< / eorc> < / nrc> < / nrc> < / rc> < / sorc> < / eorc> < / rc> < / nrc> < / nrc> < / sorc> < / rc> < / nrc> < / eorc> < / sorc> < / syni> < / eosyn> < / syn2> < / syn2> < / syn2> < / syn1> < / syn1> < / syn1> < / sosyn> < / syni> < / eosyn> < / sosyn> < / eosyn> < / syn2> < / syn1> < / sosyn> < / rc> < / nrc> < / eorc> < / sorc> < / eosyn> < / sosyn> < / syn2> < / syn1>
Claims
1. A reverse synthesis route planning method based on a multimodal large model, characterized in that, Including: S1. Extract the structural features of the three-dimensional molecular structure of the target reaction product through a molecular encoder, and then project the structural features into the representation space of the large language model through a linear layer, so as to obtain the projected structural features and input them into the large language model, and the large language model predicts the SMILES sequence of the target reaction product; S2. Drive the large language model through prompt words to sequentially execute the multimodal inference task chain composed of the reaction type prediction subtask, synthon construction subtask, and reactant generation subtask according to the predicted SMILES sequence, and obtain the retro-synthesis route for synthesizing the target reaction product from the reactants; In the reaction type prediction subtask, the large language model needs to reason about the reaction mechanism based on the predicted SMILES sequence of the target reaction product and predict the reaction type for synthesizing the target reaction product; In the synthon construction subtask, the large language model needs to assign each atom in the target reaction product to the corresponding synthon, generate an atom-level synthon assignment mask through autoregressive prediction, so as to decompose the target reaction product into different synthons, and then predict the activity of each atom on each synthon, judge whether each atom belongs to the reaction center atom, and generate an atom-level activity mask for each synthon; In the reactant generation subtask, the large language model needs to combine the information already inferred in the current multimodal inference task chain and finally generate the SMILES sequence of the reactant corresponding to each synthon.
2. The inverse synthesis route planning method based on a multimodal large model according to claim 1, wherein The three-dimensional molecular structure of the target reaction product is generated by the SMILES sequence of the target reaction product through the RDKit tool.
3. The inverse synthesis route planning method based on a multimodal large model according to claim 1, wherein In the process of generating the atom-level synthon assignment mask, the mask tokens used include the first start token, the first end token, and the classification tokens of different synthons, and all the mask tokens used are pre-added to the vocabulary of the large language model.
4. The inverse synthesis route planning method based on a multimodal large model according to claim 1, characterized in that In the process of generating the atom-level activity mask for each synthon, the mask tokens used include the second start token, the second end token, and the binary classification token indicating whether the atom belongs to the reaction center atom, and all the mask tokens used are pre-added to the vocabulary of the large language model.
5. The inverse synthesis route planning method based on a multi-modal large model according to claim 1, wherein The molecular encoder uses the pre-trained Uni-Mol model, and the large language model uses the pre-trained Qwen2-0.5B model or LLAMA model.
6. The inverse synthesis route planning method based on a multi-modal large model according to claim 1, wherein, The model framework composed of the molecular encoder, linear layer, and large language model needs to be pre-trained in two stages. In the first stage, the molecular encoder and the large language model need to be frozen, and the linear layer is optimized through the error loss between the SMILES sequence of the target reaction product predicted by the large language model and its true label. In the second stage, the molecular encoder, linear layer, and large language model need to be unfrozen, and the molecular encoder, linear layer, and large language model are jointly optimized through the error losses between the SMILES sequence of the target reaction product and the SMILES sequence of the reactants predicted by the large language model and their respective true labels.
7. The inverse synthesis route planning method based on a multimodal large model according to claim 1, wherein The large language model needs to execute the multi-modal reasoning task chain multiple times to obtain different retrosynthetic routes, and then screen the best retrosynthetic route from the generated set of retrosynthetic routes according to the currently specified screening principle.
8. An inverse synthesis route planning system based on a multimodal large model, characterized in that, It includes: A data input module for the user to input the three-dimensional molecular structure of the target reaction product; A retrosynthetic route generation module for generating a retrosynthetic route according to the three-dimensional molecular structure of the target reaction product input by the user, in accordance with the retrosynthetic route planning method based on a multi-modal large model as described in any one of claims 1 to 7; A result output module for outputting or visually displaying the result data generated in the retrosynthetic route generation module.
9. A computer-readable storage medium, characterized in that, A computer program is stored on the storage medium, and when the computer program is executed by a processor, the retrosynthetic route planning method based on a multi-modal large model as described in any one of claims 1 to 7 is implemented.
10. A computer electronic device, characterized in that, It includes a memory and a processor; The memory is used for storing a computer program; The processor is used for implementing the retrosynthetic route planning method based on a multi-modal large model as described in any one of claims 1 to 7 when executing the computer program.
Citation Information
Patent Citations
Artificial intelligence-based retrosynthesis prediction method and device, equipment and storage medium
CN111524557A
Method and device for processing synthesis and inverse synthesis molecular graph prediction model
CN116705197A