Template-free molecule inverse synthesis method based on three-dimensional conformation enhancement

By designing the Retro3D model, combining 1D SMILES sequence and 3D conformation information, the problem of inaccurate prediction results in the prior art when processing complex molecules is solved, and higher reactant prediction accuracy and calculation efficiency are achieved.

CN120015140AActive Publication Date: 2025-05-16EAST CHINA NORMAL UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510154551.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-05-16
Estimated Expiration
2045-02-12

AI Technical Summary

Technical Problem

The existing molecular inverse synthesis methods are less accurate and feasible when dealing with molecules with complex three-dimensional structures and multiple reaction centers, and it is difficult to effectively fuse 1D SMILES representations with 3D spatial position information.

Method used

A template-free molecular inverse synthesis method based on three-dimensional conformation enhancement is proposed. By designing a Retro3D model, this model combines the Transformer architecture to process 1D SMILES sequence and 3D conformation information, uses the Atom-align Fusion module to align 1D and 3D embeddings, and applies the Distance-weighted Attention mechanism and SMILES alignment module to enhance the model's understanding of spatial relationships and the accuracy of reactant prediction.

Benefits of technology

It improves the accuracy of reactant prediction and the rationality of the generation results, especially when dealing with molecules with complex three-dimensional structures and multiple reaction centers, it performs better, significantly improves the accuracy of Top-k prediction, and optimizes the computational efficiency and processing capabilities of complex molecules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120015140A_ABST
    Figure CN120015140A_ABST
Patent Text Reader

Abstract

The invention discloses a template-free molecule inverse synthesis method based on three-dimensional conformation enhancement, and aims to improve the accuracy and generalization ability of inverse synthesis prediction. According to the method, the 1D SMILES sequence and the 3D conformation information are fused, and the challenges of 1D-3D representation alignment and effective utilization of spatial information are solved. According to the method, an Atom-align Fusion module is adopted to keep alignment of atomic labeling and 3D representation, a Distance-weighted Attention mechanism is introduced, and attention distribution is optimized through a molecular space structure. In addition, in combination with SMILES alignment, data enhancement (random atom arrangement and root atom alignment) and attention guidance loss, the prediction ability of the model is further improved. The method does not need to depend on a template library or a molecular editing tool, remarkably improves the Top-k accuracy on a USPTO-50k data set, exceeds an existing template-free method, reaches an advanced level, and particularly shows better performance in inverse synthesis prediction of a complex molecular structure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of molecular retrosynthesis and the field of generative models. Specifically, it relates to a template-free molecular retrosynthesis method based on three-dimensional conformation enhancement, which proposes a molecular retrosynthesis generative model Retro3D based on 3D information fusion, and the model can effectively combine the 1D SMILES representation of the molecular structure with the 3D spatial position information to generate a reactant SMILES sequence. Background Art

[0002] Currently, the field of molecular retrosynthesis is gradually becoming an important research direction in chemical informatics and drug discovery. With the continuous development of artificial intelligence and machine learning technologies, especially driven by deep learning models, data-driven molecular design methods have made significant progress. The molecular retrosynthesis task aims to predict reactants based on given target products and is one of the key technologies in the fields of drug design, materials science, etc.

[0003] At present, the existing molecular retrosynthesis methods can be roughly divided into template-based methods and template-free methods. Template methods rely on existing reaction template libraries to predict reactants by comparing and matching reaction centers, while template-free methods focus more on learning the laws of molecular reactions from big data without relying on traditional chemical knowledge bases. However, these methods often ignore the spatial structural information of molecules, resulting in low accuracy and feasibility of prediction results when dealing with molecules with complex stereostructures and multiple reaction centers.

[0004] In recent years, more and more studies have begun to try to introduce molecular 3D structural information into retrosynthesis prediction tasks to overcome the limitations of traditional methods in dealing with complex molecules. However, how to effectively integrate 1D SMILES representation with 3D spatial position information and utilize it in the model remains a challenge. Therefore, molecular retrosynthesis generation models that combine 1D and 3D information have become one of the current research hotspots. Summary of the invention

[0005] The technical problem to be solved by the present invention is to provide a template-free molecular retrosynthesis method based on three-dimensional conformation enhancement. The method is based on a molecular retrosynthesis generation model based on 3D information fusion, which can improve the accuracy of reactant prediction and the rationality of generation results, especially when dealing with molecules with complex stereostructures and multiple reaction centers, and has better performance.

[0006] The specific technical solution for achieving the purpose of the present invention is:

[0007] A template-free molecular retrosynthesis method based on three-dimensional conformation enhancement, the method comprising the following steps:

[0008] Step A, building a Retro3D model; obtaining the molecular SMILES representation and 3D coordinate data, and designing a Retro3D model based on the Transformer architecture, processing both 1D SMILES sequence information and 3D conformation information;

[0009] Step B, generate a 1D embedding representation of the SMILES sequence; encode the molecular SMILES sequence by word vector, convert each SMILES symbol into a 1D embedding representation, and characterize each atom and its connection relationship;

[0010] Step C, extracting 3D conformation information and generating 3D position embedding; extracting the spatial position information of each atom from the 3D conformation of the molecule, performing message passing through the ComENet method, generating the neighborhood information of each atom, and forming a 3D position embedding;

[0011] Step D, Atom-align Fusion module aligns 1D and 3D embeddings; aligns 1D SMILES embedding with 3D position embedding to generate fused embedding F 3D , ensure that SMILES characters are aligned with their corresponding 3D spatial positions;

[0012] Step E, calculate and construct a 3D distance matrix; generate a preliminary 3D distance matrix by calculating the Euclidean distance between each pair of atoms; use the Gaussian basis function to convert the three-dimensional distance into a high-dimensional representation to enhance the model's understanding of spatial relationships;

[0013] Step F, apply the distance-weighted attention mechanism; introduce 3D distance weights into the self-attention mechanism, redistribute the attention weights according to the spatial relationship between atoms, and the model can focus on the atom pairs in the chemical reaction area;

[0014] Step G: SMILES alignment enhances the correspondence between products and reactants; using the atomic mapping relationship between reactants and products, a SMILES alignment mapping is constructed, and the correspondence between atoms is guided in model training through the cross-attention mechanism;

[0015] Step H: Data enhancement and model training; data enhancement is performed by randomly rearranging the atomic order and aligning the root atoms of products and reactants to improve the generalization ability of the model; during the training process, the loss function is optimized and the model parameters are adjusted;

[0016] Step I: Reasoning and reactant generation; In the reasoning stage, a beam search strategy is used to select the Top-k results from the generated candidate reactants and output the optimal reactant SMILES sequence.

[0017] Further, the step A specifically comprises:

[0018] S11, obtain the molecular SMILES representation to encode the structural information of the molecule, including:

[0019] Use SMILES to represent the 1D structure of a molecule. The SMILES string represents the structure of a molecule through atomic symbols, chemical bonds, and ring structure symbols.

[0020] Use 3D coordinate data to represent the spatial position information of each atom in the molecule to capture the three-dimensional structure of the molecule;

[0021] S12, extract the 3D conformation information of the molecule; extract the 3D conformation data of the molecule through the molecular simulation tool RDKit; the specific operations are: 1) use RDKit or other molecular modeling tools to generate the 3D conformation of the molecule; 2) extract the three-dimensional coordinates (x i , y i , z i ), forming a 3D coordinate matrix;

[0022] S13, design the architecture of Retro3D model; including: encoder: responsible for receiving the SMILES sequence and 3D conformation information at the input end and encoding it into latent representation; decoder: performs reverse reasoning based on the representation generated by the encoder to predict the SMILES sequence of the reactant;

[0023] S14, selecting an input representation for embedding processing; combining the extracted 3D conformation information with the SMILES sequence and converting them into a unified vector representation through an embedding layer; including:

[0024] Segment the SMILES string, and then use the word embedding layer to convert each SMILES symbol into a corresponding vector representation;

[0025] The extracted 3D coordinate data is normalized and converted into an embedded representation consistent with the length of the SMILES sequence;

[0026] Combine SMILES embedding and 3D position embedding to form a joint representation;

[0027] S15, select the model architecture hyperparameters; these include:

[0028] Model layers: Select the appropriate number of encoder and decoder layers, set to 6 layers to ensure the expressiveness of the model;

[0029] Number of attention heads: Set the number of heads in the multi-head attention layer to 8 to enhance the model's ability to pay attention to different structural information;

[0030] Model dimension: Set the dimension of the input vector to 512 dimensions to ensure that the model can effectively process large-scale data;

[0031] S16, initialize parameters and build training process; including:

[0032] Initialization parameters: Initialize model parameters using pre-training methods or random initialization;

[0033] Define loss function: Use cross entropy loss function to optimize prediction accuracy during training;

[0034] Optimizer selection: Select Adam optimizer and set the learning rate and weight decay parameters;

[0035] S17, model training setup and data preprocessing; prepare for model training, standardize, denoise and clean SMILES data to ensure the quality of input data; divide the training set, validation set and test set into a ratio of 8:1:1.

[0036] Further, the step B specifically includes:

[0037] S21, word segmentation of SMILES sequence; first, use the word segmentation algorithm to decompose the SMILES sequence into atomic symbols, chemical bond symbols and ring structure symbols; for example, the molecule CCO will be decomposed into three symbols: C, C, O;

[0038] S22, symbol mapping: for each symbol obtained by word segmentation, use a predefined symbol set to map each symbol to a unique index;

[0039] S23, construct a word embedding layer; use word embedding technology to construct an embedding layer for SMILES symbols; learn the vector representation of each symbol through training on a large amount of molecular data, so that similar chemical bonds or atoms have similar vector representations;

[0040] S24, generate the initial symbol embedding vector; use the word embedding layer to map each SMILES symbol to a vector space of fixed dimension; the vector representation of the symbol is input into the model;

[0041] S25, add position information; use position encoding technology to convert the position information of the symbol into a vector and add it to the symbol's embedding vector; ensure that the model can understand the relative position of the symbol in the sequence.

[0042] Further, the step C specifically comprises:

[0043] S31, using molecular simulation tools to generate 3D conformations; first, using the molecular simulation tool RDKit to generate the 3D conformations of the molecules; this process includes optimizing the three-dimensional structure of the molecules to ensure that it complies with chemical rules;

[0044] S32, extract the 3D coordinates of each atom; extract the 3D coordinates of each atom; the position of each atom is expressed as (x i , y i , z i ), where x i , y i , z i is the coordinate of the ith atom in three-dimensional space;

[0045] S33, calculating the local spatial relationship between atoms; calculating the distance, angle and local spatial relationship between each pair of atoms;

[0046] S34, generating 3D position embedding of atoms; generating a corresponding 3D position embedding vector for each atom according to the 3D coordinates and local spatial relationship of each atom;

[0047] S35, normalize the 3D position information; normalize the coordinates of each atom to zero mean and unit variance to ensure that the 3D position representations between molecules are trained on the same scale;

[0048] S36, generating a 3D position embedding matrix; organizing the 3D position embedding vectors of atoms into a matrix for representing the spatial position information of all atoms in the molecule; the dimension of this matrix is ​​N*D, where N is the number of atoms in the molecule and D is the embedding dimension of each atomic position;

[0049] S37, combines 3D position embedding with 1D SMILES embedding; through the Atom-align Fusion module, the 1D embedding representation of the SMILES sequence and the 3D position embedding representation are combined to generate the final joint embedding; the model can process both 1D SMILES sequence information and 3D spatial position information at the same time, thereby enhancing the understanding of molecular structure;

[0050] S38, optimize 3D position embedding; during the model training process, the representation of 3D position embedding is optimized through the feedback of the loss function so that the embedding can better capture the spatial relationship between molecules.

[0051] Further, the step D specifically includes:

[0052] S41, extract 1D and 3D embedding representations; the 1D embedding is a vector sequence representing the SMILES symbol, while the 3D embedding is the position vector of each atom in three-dimensional space;

[0053] S42, determining an alignment method of the 1D and 3D embeddings; aligning the 1D embeddings and the 3D embeddings to the same atomic index, so that the 1D embedding and the 3D position embedding of each atom can correspond in the same space;

[0054] S43, pad the 3D position embedding to match the length of the 1D embedding; use zero vectors to fill non-atomic positions and pad the 3D position embedding to a dimension equal to the length of the 1D sequence;

[0055] S44, calculate the weight coefficient of fusion embedding; define the weight coefficient and , which are used to control the proportion of 1D sequence embedding and 3D position embedding in the final fused embedding respectively; this weight coefficient is dynamically adjusted during the training process to ensure that the model can properly handle 1D and 3D information;

[0056] S45, performing weighted fusion of 1D and 3D embedding; according to the calculated weight coefficient and , weighted fusion of 1D SMILES embedding and 3D position embedding; the specific formula is: ;in, is the 3D position embedding, is SMILES embedding The final fused embedding combines the 1D molecular structure information with the 3D spatial structure information.

[0057] Further, the step E specifically includes:

[0058] S51, extracting the 3D coordinate information of atoms; extracting the 3D coordinates of each atom from the 3D conformation of the molecule; ensuring that the coordinate information of each atom is correct and complete;

[0059] S52, calculating the Euclidean distance between each pair of atoms; based on the 3D coordinate data extracted from step S51, calculating the Euclidean distance between each pair of atoms in the molecule ;formula: , this distance metric reflects the relative position relationship between two atoms in three-dimensional space;

[0060] S53, generate a preliminary 3D distance matrix; calculate the Euclidean distance , construct a symmetric 3D distance matrix with a size of , where N is the number of atoms in the molecule, and each element in the matrix represents the distance between two corresponding atoms;

[0061] S54, normalizing the distance matrix; in order to ensure that the scales of different molecules are consistent, the 3D distance matrix needs to be normalized; each element in the distance matrix is ​​divided by the mean, and the distance values ​​are standardized so that the values ​​of the distance matrix have a uniform scale;

[0062] S55, apply Gaussian function to transform the distance matrix; the distance between each pair of atoms is converted into Mapping to a high-dimensional space, obtaining a new representation between each pair of atoms ;formula: ,in, It is The mean of the Gaussian basis functions, is the standard deviation, is the distance between atoms i and j;

[0063] S56, generating a weighted 3D distance matrix; converting the Gaussian basis function to The information is integrated into the matrix to generate a weighted 3D distance matrix ; Each element in the matrix represents the weighted distance between the atom pair (i, j);

[0064] S57, perform nonlinear transformation of the matrix; map the weighted 3D distance matrix through a multi-layer perceptron , generate the final 3D distance matrix;

[0065] S58, constructing the final representation of the 3D distance matrix; transforming the transformed 3D distance matrix Input into the model for training and used to adjust the self-attention weight of the model.

[0066] Further, the step F specifically includes:

[0067] S61, combines spatial distance with the standard attention mechanism; when calculating the attention score between atoms, in addition to the original similarity calculation based on vector representation, As an additional term, the standard attention score is adjusted so that spatially closer atom pairs can receive larger attention weights;

[0068] S62, construct the final attention score; combine the standard attention score and weighted spatial information to construct the final attention score ; Calculation method: ,in, and are the query vector and bond vector of atom i and atom j respectively, dim is the dimension of the vector, It is a hyperparameter used to adjust the influence of 3D spatial weight on attention;

[0069] S63, use the Softmax function to normalize the attention score; apply the Softmax function to Normalization is performed to ensure that the sum of all attention weights is 1; through the Softmax function, the attention weights between each atom pair will be assigned according to their relative distance and vector similarity;

[0070] S64, calculate the weighted attention output; use the normalized attention weights , perform weighted summation on the values ​​of each atom to obtain the final attention output ; Calculation formula: , in, is the value vector of atom j, N(i) is the set of atoms adjacent to atom i, is the attention weight of atom i to atom j; the final Will be passed as input to the next layer of the model.

[0071] Further, the step G specifically includes:

[0072] S71, performing atomic mapping to generate a SMILES alignment graph; atomic mapping refers to the pairing relationship between corresponding atoms in reactants and products; the SMILES alignment graph is generated by analyzing the atomic mapping relationship between the products and reactants;

[0073] S72, design an attention-guided loss function; this loss function guides the model to allocate attention to the correct atomic correspondence during training by comparing the difference between the model's predicted attention distribution and the actual SMILES alignment map.

[0074] Further, the step H specifically includes:

[0075] S81, apply Root-align SMILES enhancement; select the root atom in the product and find the corresponding atom in the reactant; generate a reactant sequence with high similarity through the alignment of the root atoms;

[0076] S82, use early stopping strategy to prevent overfitting; monitor the loss value on the validation set during training, and stop training early if it is found that the performance of the model on the validation set has not improved further; prevent the model from over-learning on the training set, resulting in a decrease in generalization ability.

[0077] Further, the step I specifically includes:

[0078] S91, beam search initialization; in beam search, the beam size is defined as 10, which means that in each generation step, the model retains the best top 10 candidate results; initially, the beam search generates the reactant sequence starting from the first symbol of the product sequence, and each step selects the most likely next symbol according to the model's prediction until the complete reactant sequence is generated.

[0079] Compared with existing template methods and traditional generation models, Retro3D of the present invention achieves better performance in retrosynthesis tasks, especially when dealing with molecules with complex stereostructures, and can generate more accurate and reasonable reactant prediction results.

[0080] The beneficial effects of the present invention are mainly reflected in the following four aspects: (1) Improving the prediction accuracy of molecular retrosynthesis. By combining 1D SMILES sequence and 3D molecular conformation information, Retro3D can capture the spatial structural characteristics of molecules, optimize reaction center identification, and generate reactants that conform to chemical rules. Experimental results show that compared with the traditional template-free method, the present invention has achieved a significant improvement in the Top-k prediction accuracy, especially in molecules with stereochemical characteristics. (2) Enhancing the scalability and adaptability of the model. Since the present invention does not rely on a fixed reaction template, but directly learns the reaction rules from the data through a deep learning method, it can be applied to data sets of different sizes, such as USPTO-50K, and can be extended to other chemical reaction databases. This method is particularly suitable for new or rare reaction types and has strong versatility and application potential. (3) Optimizing computational efficiency and improving reasoning speed. The beam search is used to optimize the reactant generation process, and the generation search space is reduced through the SMILES alignment module, so that the model can improve computational efficiency while ensuring prediction accuracy. Experiments show that with the same computing resources, the present invention can generate high-quality reactants at a faster reasoning speed, improving the practicality of the model on large-scale data sets. (4) Enhanced processing capabilities for complex molecules. Through the distance-weighted attention mechanism, the present invention can use 3D molecular information to guide the model to focus on spatially related atomic pairs, thereby generating more reasonable reactants on molecules with complex stereostructures (such as multi-chiral centers, heterocyclic systems, bridged ring structures, etc.), improving the chemical feasibility of the prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0081] Figure 1 is a flow chart of the present invention;

[0082] Figure 2 A retrosynthetic framework diagram of a molecular retrosynthetic method provided by an embodiment of the present invention;

[0083] Figure 3This is a flowchart of the multi-head self-attention mechanism;

[0084] Figure 4 Flowchart of the multi-head cross attention mechanism. DETAILED DESCRIPTION

[0085] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0086] See also Figure 1 The present invention provides a template-free molecular retrosynthesis method based on three-dimensional conformation enhancement, comprising:

[0087] Step A: Build the Retro3D model framework. This step includes selecting appropriate molecular representation methods (such as SMILES) and 3D coordinate data, and designing a Transformer-based Retro3D model that can process both 1D SMILES sequences and 3D conformation information. Through the encoder-decoder architecture, the input layer is constructed and the corresponding embedding representation is provided to the model.

[0088] Step B: Generate 1D embedding representation of SMILES sequence. By encoding the molecular SMILES sequence with word vectors, each SMILES symbol is converted into a 1D embedding representation to ensure that each atom and its connection relationship are effectively represented.

[0089] Step C, extract 3D conformation information and generate 3D position embedding. In this step, the spatial position information of each atom is extracted from the 3D conformation of the molecule, and the neighborhood information of each atom is generated through message passing through the ComENet method to form a 3D position embedding.

[0090] Step D, Atom-align Fusion module aligns 1D and 3D embeddings. In this step, the 1D SMILES embedding is aligned with the 3D position embedding to generate a fused embedding $F_{\text{3D}}$, ensuring that the SMILES atoms are aligned with their corresponding 3D spatial positions for subsequent processing.

[0091] Step E, calculate and construct a 3D distance matrix. A preliminary 3D distance matrix is ​​generated by calculating the Euclidean distance between each pair of atoms. These distances are then converted into a high-dimensional representation using a Gaussian basis function to enhance the model's understanding of spatial relationships.

[0092] Step F, apply the distance-weighted attention mechanism. Introduce 3D distance weights in the self-attention mechanism to redistribute attention weights according to the spatial relationship between atoms. In this way, the model can focus on atom pairs that have an important impact on chemical reactions.

[0093] Step G: SMILES alignment module enhances the correspondence between products and reactants. The atomic mapping relationship between reactants and products is used to construct a SMILES alignment map (SAM), and the correct atomic correspondence is guided in model training through the cross-attention mechanism, thereby improving the accuracy of prediction.

[0094] Step H: Data augmentation and model training. Data augmentation is performed by randomly rearranging the atomic order and aligning the root atoms of products and reactants to improve the generalization ability of the model. During the training process, the loss function is optimized and the model parameters are adjusted.

[0095] Step I: Inference and reactant generation: In the inference phase, a beam search strategy is used to select the top-k results from the generated candidate reactants, and finally the optimal reactant SMILES sequence is output.

[0096] Step A: Build the Retro3D model framework. The specific implementation steps are as follows:

[0097] S11, select an appropriate molecular representation method. In this step, select an appropriate molecular representation method to encode the structural information of the molecule, mainly including: using SMILES (Simplified Molecular Input Line Entry System) to represent the 1D structure of the molecule. The SMILES string characterizes the structure of the molecule through atomic symbols, chemical bonds, and ring structure symbols; using 3D coordinate data to represent the spatial position information of each atom in the molecule to capture the three-dimensional structure of the molecule.

[0098] S12, extracting 3D conformation information of the molecule. This step extracts the 3D conformation data of the molecule through a molecular simulation tool (such as RDKit). The specific operations are as follows: 1) Generate the 3D conformation of the molecule using RDKit or other molecular modeling tools. 2) Extract the position of each atom (such as the coordinate x i , y i , z i ), forming a 3D coordinate matrix as input for subsequent processing.

[0099] S13, design the architecture of the Retro3D model. In this step, the core architecture of the Retro3D model is designed. Retro3D is based on the Transformer architecture, which includes: 1) Encoder: responsible for receiving input data (SMILES sequence and 3D conformation information) and encoding it into hidden representations. 2) Decoder: performs reverse reasoning based on the representation generated by the encoder to predict the SMILES sequence of the reactant.

[0100] S14, select input representation and perform embedding processing. Combine the extracted 3D conformation information with the SMILES sequence and convert them into a unified vector representation through the embedding layer. Specifically include:

[0101] Tokenize the SMILES string, and then use the word embedding layer to convert each SMILES symbol into a corresponding vector representation.

[0102] The extracted 3D coordinate data are normalized and converted into an embedding representation that matches the length of the SMILES sequence.

[0103] The SMILES embedding and the 3D position embedding are combined to form a joint representation that is input into the Transformer model.

[0104] S15, select appropriate Transformer architecture hyperparameters. In this step, set the hyperparameters of the Transformer model to ensure that it can handle the complex structural information of molecules. It mainly includes:

[0105] Number of model layers: Choose an appropriate number of encoder and decoder layers, generally 6 to 12 layers, to ensure the expressiveness of the model.

[0106] Number of attention heads: Set the number of attention heads for each layer of Transformer, usually 8 to 16 heads, to enhance the model's attention to different structural information.

[0107] Model Dimension: Set an appropriate dimension for each input vector, such as 512 dimensions or higher, to ensure that the model can effectively process large-scale data.

[0108] S16, initialize parameters and build training process. In this step, initialize the parameters of the model and build the training process:

[0109] Initialization parameters: Initialize model parameters using pre-training methods or random initialization.

[0110] Define loss function: Select a suitable loss function, such as cross-entropy loss, to optimize prediction accuracy during training.

[0111] Optimizer selection: Select an appropriate optimizer (such as the Adam optimizer) and set hyperparameters such as learning rate and weight decay.

[0112] S17, model training setup and data preprocessing. In this step, the model training preparation is performed, and the SMILES data is standardized, denoised, and cleaned to ensure the quality of the input data. The data is divided according to the ratio of training set, validation set, and test set (e.g., 70% training, 15% validation, and 15% test) to ensure that the model can be effectively trained on different data sets.

[0113] Step B, generating a 1D embedding representation of the SMILES sequence, specifically includes:

[0114] S21, SMILES sequence segmentation. First, obtain the SMILES representation of the molecule, and use the segmentation algorithm to decompose the SMILES sequence into basic units (tokens) such as atomic symbols, chemical bond symbols, and ring structure symbols. For example, the molecule CCO will be decomposed into three symbols: C, C, and O.

[0115] S22, symbol mapping. For each symbol obtained by word segmentation, use a predefined symbol set (such as a standard symbol set provided by RDKit or other tools) to map it, and map each symbol to a unique index to ensure that each symbol has a corresponding unique identifier.

[0116] S23, build word embedding layer. Use word embedding technology (such as Word2Vec, GloVe, etc.) to build the embedding layer of SMILES symbols. Through training on a large amount of molecular data, learn the vector representation of each symbol so that similar chemical bonds or atoms have similar vector representations, ensuring that the molecular structure information can be effectively transmitted.

[0117] S24, generate initial symbol embedding vectors. Use the word embedding layer to map each SMILES symbol to a vector space of fixed dimension. Usually, the dimension of these vectors is set to 512 dimensions or higher to ensure sufficient expressiveness. The vector representation corresponding to each symbol will be used as the input of the symbol in the model.

[0118] S25, adding position information. In order to ensure that the order information in the SMILES sequence can be passed to the model, the position information is added to the embedding vector of each symbol. Usually, the positional encoding technology is used to convert the position information of each symbol into a vector and add it to the embedding vector of the symbol. This ensures that the model can understand the relative position of the symbol in the SMILES sequence.

[0119] Step C, extracting 3D conformation information and generating 3D position embedding, specifically includes:

[0120] S31, Generate 3D conformation using molecular simulation tools. First, select a suitable molecular simulation tool (such as RDKit or Gaussian, etc.) to generate the 3D conformation of the molecule. By using these tools, molecular dynamics simulation or energy minimization is performed based on the 2D structural information of the molecule to generate a stable 3D conformation. This process usually includes optimizing the three-dimensional structure of the molecule to ensure that it complies with chemical rules, such as the rationality of interatomic distances and angles.

[0121] S32, extract the 3D coordinates of each atom. After obtaining a stable 3D conformation through molecular simulation tools, extract the spatial position information (i.e., 3D coordinates) of each atom. The position of each atom can be expressed as a triple (x i , y i , z i ), where x i , y i , z i is the coordinate of the ith atom in three-dimensional space.

[0122] S33, calculate the local spatial relationships between atoms. After extracting the 3D coordinates of each atom, calculate the distance, angle and other spatial relationships between each pair of atoms. These local spatial relationships are used to enhance the model's understanding of atomic proximity and structure. For example, calculate the Euclidean distance between two atoms. , so as to facilitate the subsequent spatial feature fusion.

[0123] S34, generating 3D position embeddings of atoms. According to the 3D coordinates and local spatial relationships of each atom, a corresponding 3D position embedding vector is generated for each atom. Typically, the positions of atoms are encoded using position encoding techniques to generate fixed-dimensional embedding representations (such as 512 dimensions or higher), which contain the spatial position information of atoms.

[0124] S35, standardize the 3D position information. In order to avoid the scale differences between different molecules affecting the learning of the model, the 3D position embedding of atoms is usually standardized. By standardizing the coordinates of each atom to zero mean and unit variance, it is ensured that the 3D position representations between molecules are trained on the same scale, thereby improving the robustness of the model.

[0125] S36, generate a 3D position embedding matrix. Organize the 3D position embedding vector of each atom into a matrix, which is used to represent the spatial position information of all atoms in the molecule. The dimension of this matrix is ​​usually N*D, where N is the number of atoms in the molecule and D is the embedding dimension of each atomic position (usually 512 or higher). This matrix will be passed as input to the subsequent steps of the Transformer model.

[0126] S37, combines 3D position embedding with 1D SMILES embedding. Through the Atom-align Fusion module, the 1D embedding representation of the SMILES sequence and the 3D position embedding representation are combined to generate the final joint embedding. In this way, the model can process both 1D SMILES sequence information and 3D spatial position information at the same time, thereby enhancing the understanding of molecular structure.

[0127] S38, optimize 3D position embedding. During the model training process, the representation of 3D position embedding is optimized through the feedback of the loss function, so that the embedding can better capture the spatial relationship between molecules, thereby improving the prediction accuracy. This step ensures that the 3D position embedding can effectively help the model generate synthetic routes and predict reactants during the prediction process.

[0128] Step D, Atom-align Fusion module aligns 1D and 3D embeddings, specifically including:

[0129] S41, extract 1D and 3D embedding representations. First, obtain the 1D embedding representation from the SMILES sequence and the 3D position embedding representation from the 3D conformation of the molecule. Prior to this, the SMILES sequence has been converted into a vector representation through word segmentation and word embedding techniques, and the 3D position embedding is a spatial position vector generated by extracting the 3D coordinates of each atom. At this point, the 1D embedding is a vector sequence representing the SMILES symbol, and the 3D embedding is the position vector of each atom in three-dimensional space.

[0130] S42, determine the alignment of 1D and 3D embeddings. Since 1D SMILES embedding and 3D position embedding represent different aspects of molecules (1D represents the linear relationship of molecular structure, 3D represents the geometric relationship of molecular spatial structure), it is necessary to ensure the alignment of these two embeddings in the model. The specific method is to align the 1D embedding and 3D embedding to the same atomic index so that the 1D embedding and 3D position embedding of each atom can correspond in the same space.

[0131] S43, pad the 3D position embedding to match the length of the 1D embedding. Since the length of the SMILES sequence (i.e., the number of atom symbols) may be different from the number of atoms at the 3D position, the 3D position embedding needs to be padded. Specifically, if the length of the 3D position embedding is less than the length of the SMILES sequence, the 3D position embedding is padded to a dimension equal to the length of the SMILES sequence, usually using zero padding or special symbols to fill non-atom positions.

[0132] S44, calculate the weight coefficient of the fused embedding. In order to effectively fuse 1D and 3D embedding, define the weight coefficient and , which are used to control the contribution ratio of 1D SMILES embedding and 3D position embedding in the final fused embedding. The weight coefficient is dynamically adjusted during the training process to ensure that the model can properly handle 1D and 3D information.

[0133] S45, performing weighted fusion of 1D and 3D embedding. According to the calculated weight coefficient and , weighted fusion of 1D SMILES embedding and 3D position embedding. The specific formula is: .in, is the 3D position embedding, is SMILES embedding is the final fusion embedding. This fusion embedding combines the 1D molecular structure information with the 3D spatial structure information and provides it to the subsequent model for processing.

[0134] Step E, calculate and construct a 3D distance matrix, specifically including:

[0135] S51, extracting the 3D coordinate information of atoms. First, extract the 3D coordinates (x i , y i , z i ). These coordinate data can be obtained through molecular modeling tools (such as RDKi). In this step, ensure that the coordinate information of each atom is correct and complete so that the spatial relationship between atoms can be calculated later.

[0136] S52, calculating the Euclidean distance between each pair of atoms. Based on the 3D coordinate data extracted from step S51, the Euclidean distance between each pair of atoms in the molecule is calculated. The formula is as follows: This distance metric reflects the relative position relationship between two atoms in three-dimensional space and is the core calculation for constructing a 3D distance matrix.

[0137] S53, generating a preliminary 3D distance matrix. The Euclidean distance obtained by calculating , construct a symmetric 3D distance matrix, where N is the number of atoms in the molecule, and each element in the matrix represents the distance between two corresponding atoms

[0138] S54, normalize the distance matrix. In order to ensure that the scales of different molecules are consistent, the 3D distance matrix needs to be normalized. By dividing each element in the distance matrix by the maximum distance or mean, the distance values ​​are standardized so that the values ​​of the distance matrix have a uniform scale. This helps to improve the robustness of the model to spatial relationships and reduce inconsistencies caused by different molecular sizes.

[0139] S55, apply Gaussian function to transform the distance matrix. In order to increase the expressive power of the matrix, the distance matrix can be transformed into a higher dimensional representation using Gaussian basis function. The distance between each pair of atoms is transformed into Mapping to a high-dimensional space, obtaining a new representation between each pair of atoms The formula is as follows: ,in, It is The mean of the Gaussian basis functions, is the standard deviation, is the distance between atoms i and j.

[0140] S56, generating a weighted 3D distance matrix. The information is integrated into the matrix to generate a weighted 3D distance matrix . Each element in this matrix It represents the weighted distance between the atomic pair (i, j), which can reflect the complexity of the spatial relationship between atoms.

[0141] S57, perform nonlinear transformation of the matrix. In order to enhance the model's ability to express spatial structure, nonlinear transformation is used to further process the 3D distance matrix. The weighted 3D distance matrix is ​​processed by a multi-layer perceptron (MLP) or other nonlinear layers. The final 3D distance matrix is ​​generated by mapping. This transformation helps capture the nonlinear spatial relationship between atoms and enables the model to handle more complex spatial dependencies.

[0142] S58, construct the final representation of the 3D distance matrix. The 3D distance matrix after normalization, Gaussian transformation and nonlinear transformation As the final 3D spatial relationship matrix, it is passed as input to the subsequent model for training. This matrix will be used to adjust the model's attention mechanism to ensure that the model can focus on important spatial relationships in molecules.

[0143] Step F, apply the Distance-weighted Attention mechanism, which includes:

[0144] S61, combining spatial distance with standard attention mechanism. In order to enable the model to strike a balance between spatial structure and sequence information, it is necessary to combine the attention weight based on 3D distance with the traditional self-attention mechanism. Specifically, when calculating the attention score between atoms, in addition to the original similarity calculation based on vector representation (such as dot product calculation), it is also necessary to As an additional term, the standard attention score is adjusted so that spatially closer atom pairs can receive larger attention weights, thereby enhancing the model's perception of spatial relationships.

[0145] S62: Construct the final attention score. Combine the standard attention score and the weighted spatial information to construct the final attention score. The calculation method is as follows: ,in, and are the query and key vectors of atom i and atom j respectively, d is the dimension of the vector, is a hyperparameter used to adjust the influence of 3D spatial weight on attention.

[0146] S63, use the Softmax function to normalize the attention score. In order to obtain the final attention weight, the Softmax function is applied to Normalization is performed to ensure that the sum of all attention weights is 1. Through the Softmax function, the attention weights between each atom pair will be assigned according to their relative distance and vector similarity.

[0147] S64, calculate the weighted attention output. Use the normalized attention weights , perform weighted summation on the values ​​of each atom to obtain the final attention output The calculation formula is as follows: , in, is the value vector of atom j, N(i) is the set of atoms adjacent to atom i, is the attention weight of atom i to atom j. The final Will be passed as input to the next layer of the model.

[0148] Step G, the SMILES alignment module enhances the correspondence between products and reactants, specifically including:

[0149] S71, perform atomic mapping to generate a SMILES alignment map (SAM). Generate a SMILES Alignment Map (SAM) by analyzing the atomic mapping relationship between products and reactants. Atomic mapping refers to the pairing relationship between corresponding atoms in reactants and products. For example, in a specific reaction, an atom in the reactant may become an atom in the product. Through these mapping relationships, a SMILES alignment map is constructed, which contains the corresponding relationship between atoms in reactants and products, helping the model understand the structural connection between the two.

[0150] S72, design an attention guidance loss function. In order to enhance the model's learning of SMILES alignment, an attention guidance loss function is designed. This loss function encourages the model to learn the correct atomic correspondence through attention during training by comparing the difference between the model's predicted attention distribution and the actual SMILES alignment map. Specifically, the loss function measures the alignment error by calculating the cross-entropy loss, thereby guiding the model to pay more attention to the structural mapping between products and reactants when generating reactants.

[0151] Step H, data enhancement and model training, specifically includes:

[0152] S81, apply Root-align SMILES enhancement. Select a root atom in the product and make sure to find the corresponding atom in the reactant as the root atom. By aligning the root atoms, a reactant SMILES sequence with high similarity is generated. This method can reduce the structural differences between reactants and products, thereby enhancing the stability of model learning and helping the model focus more on learning the key changes of the reaction.

[0153] S82, use early stopping strategy to prevent overfitting. To avoid overfitting of the model during training, use early stopping strategy. During training, monitor the loss value on the validation set. If the performance of the model on the validation set is not further improved, stop training early. This helps prevent the model from over-learning on the training set, resulting in a decrease in generalization ability.

[0154] Step I: The SMILES alignment module enhances the correspondence between products and reactants, specifically including:

[0155] S91, beam search initialization. In order to generate the optimal reactant SMILES sequence, a beam search algorithm is used. In beam search, a beam size is defined, for example, set to 10, which means that in each generation step, the model retains the top 10 optimal candidate results. Initially, the beam search generates the reactant sequence starting from the first symbol of the product SMILES sequence, and each step selects the most likely next symbol according to the model's prediction until the complete reactant SMILES sequence is generated.

[0156] S92, use early stopping strategy to prevent overfitting. To avoid overfitting of the model during training, use early stopping strategy. During training, monitor the loss value on the validation set. If the performance of the model on the validation set is not further improved, stop training in advance. This helps prevent the model from over-learning on the training set, resulting in a decrease in generalization ability.

[0157] The method of the present invention first extracts the SMILES representation of the target molecule through a molecular simulation tool (such as RDKit) and generates its corresponding 3D conformation. The generated 3D conformation includes the spatial coordinates of each atom. Then, the embedded representation of the 1D SMILES sequence is aligned with the embedded representation of the 3D spatial position information through the Atom-alignFusion module. This module combines the 1D SMILES representation and the 3D position embedding to ensure that when input to the model, the linear information and spatial information of the molecular structure can be integrated in the self-attention stage (see Figure 3 ) effective integration.

[0158] Next, the distance-weighted attention mechanism is used to calculate the attention of the combination of 1D and 3D embeddings. By introducing a Gaussian weighting function, the attention weights are adjusted according to the 3D distance between atoms in the molecule, so that the model can pay more attention to pairs of atoms with similar spatial structures. This mechanism calculates the Euclidean distance between each pair of atoms and adjusts the model's attention allocation based on this distance information, ensuring that the model can generate reactants in the cross-attention stage (see Figure 4 ) focuses on key spatial relationships in molecules.

[0159] During the model training process, data augmentation strategies are used for optimization. First, the SMILES sequence is randomized (such as randomly disrupting the atomic order) to generate enhanced training data. This process can increase the diversity of training data, allowing the model to adapt to more diverse reactant predictions. Then, the SMILES sequences of products and reactants are aligned through the SMILES alignment module to generate a SMILES Alignment Map (SAM), and the model is guided to learn the atomic mapping relationship between products and reactants by introducing attention-guided loss.

[0160] In the inference phase, the reactant SMILES sequence is generated through a beam search algorithm. In beam search, the model starts with the product SMILES sequence and gradually generates the reactant sequence through the decoder. During the generation process, the model selects the optimal path based on the probability of each candidate symbol until the complete reactant SMILES sequence is generated. The generated reactant candidate sequence is verified to ensure that it complies with chemical rules and is valid.

[0161] Finally, through the above steps, the molecular retrosynthesis generation model combining 1D and 3D information can effectively predict the reactant SMILES sequence, especially when facing complex molecules, and can generate more accurate and reasonable reactant prediction results.

[0162] Example 1

[0163] See also Figure 2 In this embodiment, a molecular retrosynthesis generation model Retro3D based on 3D information fusion is constructed. The model can effectively combine the 1D SMILES representation of the molecular structure with the 3D spatial position information to generate the reactant SMILES sequence. Next, the model is applied to the molecular retrosynthesis task in the USPTO-50k dataset to verify its effect. The following describes the working process in detail:

[0164] First, we extract reaction data from the USPTO-50k dataset, which contains 50,016 atom-annotated reactions, of which 40,008 reactions are used as training sets, 5,001 reactions are used as validation sets, and 5,007 reactions are used as test sets. Each reaction in the dataset includes a product SMILES sequence and a corresponding reactant SMILES sequence.

[0165] Then, molecular simulation tools (such as RDKit) are used to generate corresponding 3D conformations for each reaction and extract the 3D position embeddings of each reactant and product molecule. These 3D position information are input into the Retro3D model together with the 1D SMILES representations of the products and reactants, and the 1D and 3D embeddings are aligned through the Atom-align Fusion module, and the spatial relationship is introduced into the attention calculation through the Distance-weighted Attention mechanism.

[0166] In order to enhance the training data, random SMILES enhancement and root-align SMILES enhancement strategies are used to enhance the original SMILES data. By randomizing the SMILES sequence and aligning the root atoms to generate different reactant samples, the diversity of the data set is increased and the generalization ability of the model is improved.

[0167] Next, the dataset was divided into training set, validation set, and test set in a ratio of 8:1:1. The Retro3D model was trained on the training set, the model parameters were adjusted on the validation set, and the performance of the model was evaluated on the test set. In the experiment, the Retro3D model was compared with existing template methods, semi-template methods, and other template-free methods, including RetroSim, GLN, Retroformer, etc., to evaluate its Top-k accuracy in generating reactant SMILES sequences.

[0168] Table 1

[0169] The experimental results in Table 1 show that Retro3D surpasses the existing template-free methods in terms of accuracy and effectiveness in generating reactant SMILES sequences, regardless of whether there is a reaction type or not, and also performs quite well in comparison with template methods and semi-template methods. This shows that Retro3D can effectively combine 1D and 3D information and achieve good prediction results in complex molecular retrosynthesis tasks.

Claims

1. A template-free molecular retrosynthesis method based on three-dimensional conformation enhancement, characterized in that: The method comprises the following steps: Step A, building a Retro3D model; obtaining the molecular SMILES representation and 3D coordinate data, and designing a Retro3D model based on the Transformer architecture, processing both 1D SMILES sequence information and 3D conformation information; Step B, generate a 1D embedding representation of the SMILES sequence; encode the molecular SMILES sequence by word vector, convert each SMILES symbol into a 1D embedding representation, and characterize each atom and its connection relationship; Step C, extracting 3D conformation information and generating 3D position embedding; extracting the spatial position information of each atom from the 3D conformation of the molecule, performing message passing through the ComENet method, generating the neighborhood information of each atom, and forming a 3D position embedding; Step D, Atom-align Fusion module aligns 1D and 3D embeddings; aligns 1D SMILES embedding with 3D position embedding to generate fused embedding F 3D , ensure that the SMILES characters are aligned with their corresponding 3D spatial positions; Step E, calculate and construct a 3D distance matrix; generate a preliminary 3D distance matrix by calculating the Euclidean distance between each pair of atoms; use the Gaussian basis function to convert the three-dimensional distance into a high-dimensional representation to enhance the model's understanding of spatial relationships; Step F, apply the distance-weighted attention mechanism; introduce 3D distance weights into the self-attention mechanism, redistribute the attention weights according to the spatial relationship between atoms, and the model can focus on the atom pairs in the chemical reaction area; Step G: SMILES alignment enhances the correspondence between products and reactants; using the atomic mapping relationship between reactants and products, a SMILES alignment mapping is constructed, and the correspondence between atoms is guided in model training through the cross-attention mechanism; Step H: Data enhancement and model training; data enhancement is performed by randomly rearranging the atomic order and aligning the root atoms of products and reactants to improve the generalization ability of the model; during the training process, the loss function is optimized and the model parameters are adjusted; Step I: Reasoning and reactant generation; In the reasoning stage, a beam search strategy is used to select the Top-k results from the generated candidate reactants and output the optimal reactant SMILES sequence.

2. The template-free molecular retrosynthesis method according to claim 1, characterized in that: The step A specifically comprises: S11, obtain the molecular SMILES representation to encode the structural information of the molecule, including: Use SMILES to represent the 1D structure of a molecule. The SMILES string represents the structure of a molecule through atomic symbols, chemical bonds, and ring structure symbols. Use 3D coordinate data to represent the spatial position information of each atom in the molecule to capture the three-dimensional structure of the molecule; S12, extract the 3D conformation information of the molecule; extract the 3D conformation data of the molecule through the molecular simulation tool RDKit; the specific operations are: 1) use RDKit or other molecular modeling tools to generate the 3D conformation of the molecule; 2) extract the three-dimensional coordinates (x i , y i , z i ), forming a 3D coordinate matrix; S13, design the architecture of Retro3D model; including: encoder: responsible for receiving the SMILES sequence and 3D conformation information at the input end and encoding it into latent representation; decoder: performs reverse reasoning based on the representation generated by the encoder to predict the SMILES sequence of the reactant; S14, selecting an input representation for embedding processing; combining the extracted 3D conformation information with the SMILES sequence and converting them into a unified vector representation through an embedding layer; including: Segment the SMILES string, and then use the word embedding layer to convert each SMILES symbol into a corresponding vector representation; The extracted 3D coordinate data is normalized and converted into an embedded representation consistent with the length of the SMILES sequence; Combine SMILES embedding and 3D position embedding to form a joint representation; S15, select model architecture hyperparameters; including: Model layers: Select the number of encoder and decoder layers, set to 6 layers to ensure the expressiveness of the model; Number of attention heads: The number of heads in the multi-head attention layer is set to 8 to enhance the model's ability to pay attention to different structural information; Model dimension: The dimension of the input vector is set to 512 to ensure that the model can effectively process large-scale data; S16, initialize parameters and build training process; including: Initialization parameters: Initialize model parameters using pre-training methods or random initialization; Define loss function: Use cross entropy loss function to optimize prediction accuracy during training; Optimizer selection: Select Adam optimizer and set the learning rate and weight decay parameters; S17, model training setup and data preprocessing; prepare for model training, standardize, denoise and clean SMILES data to ensure the quality of input data; divide the training set, validation set and test set into a ratio of 8:1:

1.

3. The template-free molecular retrosynthesis method according to claim 1, characterized in that: The step B specifically comprises: S21, word segmentation of SMILES sequence; first, the word segmentation algorithm is used to decompose the SMILES sequence into atomic symbols, chemical bond symbols and ring structure symbols; S22, symbol mapping: for each symbol obtained by word segmentation, use a predefined symbol set to map each symbol to a unique index; S23, construct a word embedding layer; use word embedding technology to construct an embedding layer for SMILES symbols; learn the vector representation of each symbol through training on a large amount of molecular data, so that similar chemical bonds or atoms have similar vector representations; S24, generate the initial symbol embedding vector; use the word embedding layer to map each SMILES symbol to a vector space of fixed dimension; the vector representation of the symbol is input into the model; S25, add position information; use position encoding technology to convert the position information of the symbol into a vector and add it to the symbol's embedding vector; ensure that the model can understand the relative position of the symbol in the sequence.

4. The template-free molecular retrosynthesis method according to claim 1, characterized in that: The step C specifically comprises: S31, using molecular simulation tools to generate 3D conformations; first, using the molecular simulation tool RDKit to generate the 3D conformations of the molecules; this process includes optimizing the three-dimensional structure of the molecules to ensure that it complies with chemical rules; S32, extract the 3D coordinates of each atom; extract the 3D coordinates of each atom; the position of each atom is expressed as (x i , y i ,z i ), where x i , y i , z i is the coordinate of the ith atom in three-dimensional space; S33, calculating the local spatial relationship between atoms; calculating the distance, angle and local spatial relationship between each pair of atoms; S34, generating 3D position embedding of atoms; generating a corresponding 3D position embedding vector for each atom according to the 3D coordinates and local spatial relationship of each atom; S35, normalize the 3D position information; normalize the coordinates of each atom to zero mean and unit variance to ensure that the 3D position representations between molecules are trained on the same scale; S36, generating a 3D position embedding matrix; organizing the 3D position embedding vectors of atoms into a matrix for representing the spatial position information of all atoms in the molecule; the dimension of this matrix is ​​N*D, where N is the number of atoms in the molecule and D is the embedding dimension of each atomic position; S37, combines 3D position embedding with 1D SMILES embedding; through the Atom-align Fusion module, the 1D embedding representation of the SMILES sequence and the 3D position embedding representation are combined to generate the final joint embedding; the model can simultaneously process 1DSMILES sequence information and 3D spatial position information, thereby enhancing the understanding of molecular structure; S38, optimize 3D position embedding; during the model training process, the representation of 3D position embedding is optimized through the feedback of the loss function so that the embedding can better capture the spatial relationship between molecules.

5. The template-free molecular retrosynthesis method according to claim 1, characterized in that: The step D specifically comprises: S41, extract 1D and 3D embedding representations; the 1D embedding is a vector sequence representing the SMILES symbol, while the 3D embedding is the position vector of each atom in three-dimensional space; S42, determining an alignment method of the 1D and 3D embeddings; aligning the 1D embeddings and the 3D embeddings to the same atomic index, so that the 1D embedding and the 3D position embedding of each atom can correspond in the same space; S43, pad the 3D position embedding to match the length of the 1D embedding; use zero vectors to fill non-atomic positions and pad the 3D position embedding to a dimension equal to the length of the 1D sequence; S44, calculate the weight coefficient of fusion embedding; define the weight coefficient and , which are used to control the proportion of 1D sequence embedding and 3D position embedding in the final fused embedding respectively; this weight coefficient is dynamically adjusted during the training process to ensure that the model can process 1D and 3D information; S45, performing weighted fusion of 1D and 3D embedding; according to the calculated weight coefficient and , weighted fusion of 1D SMILES embedding and 3D position embedding; the specific formula is: ;in, is the 3D position embedding, is SMILES embedding The final fused embedding combines the 1D molecular structure information with the 3D spatial structure information.

6. The template-free molecular retrosynthesis method according to claim 1, characterized in that: The step E specifically comprises: S51, extracting the 3D coordinate information of atoms; extracting the 3D coordinates of each atom from the 3D conformation of the molecule; ensuring that the coordinate information of each atom is correct and complete; S52, calculating the Euclidean distance between each pair of atoms; based on the 3D coordinate data extracted from step S51, calculating the Euclidean distance between each pair of atoms in the molecule ;formula: , this distance metric reflects the relative position relationship between two atoms in three-dimensional space; S53, generating a preliminary 3D distance matrix; by calculating the Euclidean distance , construct a symmetric 3D distance matrix with a size of , where N is the number of atoms in the molecule, and each element in the matrix represents the distance between two corresponding atoms; S54, normalizing the distance matrix; dividing each element in the distance matrix by the mean, and standardizing the distance values ​​so that the values ​​of the distance matrix have a uniform scale; S55, apply Gaussian function to transform the distance matrix; the distance between each pair of atoms is converted into Mapping to a high-dimensional space, obtaining a new representation between each pair of atoms ;formula: ,in, It is The mean of the Gaussian basis functions, is the standard deviation, is the distance between atoms i and j; S56, generating a weighted 3D distance matrix; converting the Gaussian basis function to The information is integrated into the matrix to generate a weighted 3D distance matrix ; Each element in the matrix represents the weighted distance between the atom pair (i, j); S57, perform nonlinear transformation of the matrix; map the weighted 3D distance matrix through a multi-layer perceptron , generate the final 3D distance matrix; S58, constructing the final representation of the 3D distance matrix; transforming the transformed 3D distance matrix Input into the model for training and used to adjust the self-attention weight of the model.

7. The template-free molecular retrosynthesis method according to claim 1, characterized in that: The step F specifically comprises: S61, combines spatial distance with the standard attention mechanism; when calculating the attention score between atoms, in addition to the original similarity calculation based on vector representation, As an additional term, the standard attention score is adjusted so that spatially closer atom pairs can receive larger attention weights; S62, construct the final attention score; combine the standard attention score and weighted spatial information to construct the final attention score ; Calculation method: ,in, and are the query vector and bond vector of atom i and atom j respectively, dim is the dimension of the vector, It is a hyperparameter used to adjust the influence of 3D spatial weight on attention; S63, use the Softmax function to normalize the attention score; apply the Softmax function to Normalization is performed to ensure that the sum of all attention weights is 1; through the Softmax function, the attention weights between each atom pair will be assigned according to their relative distance and vector similarity; S64, calculate the weighted attention output; use the normalized attention weights , perform weighted summation on the values ​​of each atom to obtain the final attention output ; Calculation formula: , in, is the value vector of atom j, N(i) is the set of atoms adjacent to atom i, is the attention weight of atom i to atom j; the final Will be passed as input to the next layer of the model.

8. The template-free molecular retrosynthesis method according to claim 1, characterized in that: The step G specifically comprises: S71, performing atomic mapping to generate a SMILES alignment graph; atomic mapping refers to the pairing relationship between corresponding atoms in reactants and products; the SMILES alignment graph is generated by analyzing the atomic mapping relationship between the products and reactants; S72, design an attention-guided loss function; this loss function guides the model to allocate attention to the correct atomic correspondence during training by comparing the difference between the model's predicted attention distribution and the actual SMILES alignment map.

9. The template-free molecular retrosynthesis method according to claim 1, characterized in that: The step H specifically comprises: S81, apply Root-align SMILES enhancement; select the root atom in the product and find the corresponding atom in the reactant; generate a reactant sequence with high similarity through the alignment of the root atoms; S82, use early stopping strategy to prevent overfitting; monitor the loss value on the validation set during training, and stop training early if it is found that the performance of the model on the validation set has not improved further; prevent the model from over-learning on the training set, resulting in a decrease in generalization ability.

10. The template-free molecular retrosynthesis method according to claim 1, characterized in that: The step I specifically comprises: S91, beam search initialization; in beam search, the beam size is defined as 10, which means that in each generation step, the model retains the top 10 optimal candidate results; initially, the beam search generates the reactant sequence starting from the first symbol of the product sequence, and each step selects the next symbol according to the model's prediction until the complete reactant sequence is generated.

Citation Information

Patent Citations

  • Deep learning-based inverse synthesis prediction method and device, medium and equipment

    CN114220496A

  • Molecular multi-step inverse synthesis prediction method and device based on template-free

    CN117292763A

  • Inverse synthesis prediction method and device based on general molecular graph representation learning model

    CN117316333A

  • Molecular large model based on multi-dimensional molecular information, construction method and application

    CN117524353A

  • Multi-model ensemble learning-based method for improving confidence of retrosynthesis

    WO2023193259A1