Reverse synthesis method and device of cyclization reaction, electronic equipment and storage medium

By performing rationality verification and minimum ring system analysis on multi-source chemical reaction data, and training a dedicated cyclic reaction model using the OpenNMT framework and dynamic masking mechanism, the problem of low prediction accuracy of general models for cyclic reactions was solved, and higher-precision cyclic skeleton synthesis was achieved.

CN121075461APending Publication Date: 2025-12-05SOUTHWEAT UNIV OF SCI & TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511182077.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

General retrosynthetic models struggle to accurately capture the unique mechanisms of cyclization reactions, resulting in low prediction accuracy and an inability to effectively cover the entropy-enthalpy balance, strain effects, and multi-step cooperative mechanisms of cyclization reactions.

Method used

By collecting multi-source chemical reaction data and verifying their chemical rationality, we use the minimum ring system to analyze and identify cyclic reactions, construct a data augmentation base set, and use the reaction center offset maximum common substructure alignment algorithm and the OpenNMT framework for model training. The dynamic masking mechanism enhances model learning, and a dedicated model for cyclic reactions is constructed. The model accuracy is evaluated by the weighted Top-k index.

Benefits of technology

It significantly improves the prediction accuracy of cyclization reactions, enhances the reliability and feasibility of cyclic skeleton synthesis, and solves the shortcomings of general models in predicting cyclization reactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121075461A_ABST
    Figure CN121075461A_ABST
Patent Text Reader

Abstract

The invention discloses an inverse synthesis method and device of a ring formation reaction, electronic equipment and a storage medium, and aims to solve the problems that a general inverse synthesis model is low in prediction precision of the ring formation reaction and cannot accurately capture a unique mechanism of the ring formation reaction, the method comprises the following steps: collecting and processing multi-source chemical reaction data to obtain an initial data set; the method comprises the following steps: analyzing and identifying a ring formation reaction by adopting a minimum ring system, screening a special data set, generating an aligned SMILES sequence through a reaction center offset type maximum common substructure alignment algorithm, constructing a data enhancement basic set, pre-training a model by utilizing general data on the basis of an OpenNMT framework, and constructing a ring formation reaction special model through fine tuning, so as to weight Top-k index evaluation, and finally, constructing a ring formation reaction model. A heterogeneous interface is constructed to realize bidirectional data interaction between an interaction interface and a special model, and the inverse synthesis prediction precision of the ring formation reaction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of deep learning, in particular to a ring-forming reaction inverse synthesis method and device, electronic equipment and storage medium. BACKGROUND

[0002] Currently, the mainstream direction of inverse synthesis prediction research focuses on constructing general inverse synthesis reaction models. Such models attempt to cover a wide range of chemical reaction types in a unified framework, aiming to provide a comprehensive solution for complex molecule synthesis route design.

[0003] However, the general model has defects in accurately capturing deep rules of reactions with unique mechanisms, resulting in deviations in the prediction results of these types of reactions. During the reaction path prediction, the continuous accumulation of errors in the iterative prediction process of the single-step model (even prediction errors for some key steps) easily leads to the lack of feasibility and reliability of the final deduced synthesis route. Ring-forming reactions, which focus on building ring skeletons, play an important role in the exploration of new compound structure categories. However, the prediction of ring-forming reactions is more difficult than other reaction categories, as it requires precise control of the spatial arrangement and electronic interaction of multiple reaction sites within the molecule to form ring structures of specific size and conformation. The prediction of ring-forming reactions is more challenging than linear reactions or simple functional group transformations in terms of reaction site identification, transition state modeling, and product distribution prediction.

[0004] The uniqueness of ring-forming reactions makes it difficult for general inverse synthesis models to effectively cover them. General models often fail to fully capture the entropy enthalpy balance, tension effect, and multi-step coordination mechanism in the ring-forming process, resulting in low prediction accuracy. SUMMARY

[0005] In view of the above problems, the present application provides a ring-forming reaction inverse synthesis method, device, electronic equipment and storage medium, which can solve the technical problems of low prediction accuracy of general inverse synthesis models for ring-forming reactions and inability to accurately capture the unique mechanisms of ring-forming reactions.

[0006] In a first aspect, the present application provides a ring-forming reaction inverse synthesis method, which comprises: Collecting multi-source chemical reaction data, and obtaining an initial data set after data processing, the data processing including chemical reasonableness verification; Based on the initial data set, identifying ring-forming reactions using minimum ring system analysis, and screening a special data set for ring skeleton synthesis, the analysis including judging new ring formation by comparing the number of rings and ring characteristics of reactants and products; Using a reaction center offset maximum common substructure alignment algorithm, generating multiple aligned SMILES sequences for the reactions in the special data set, and constructing a data augmentation base set; Based on the OpenNMT framework, an initial model is pre-trained using general chemical reaction data to obtain a pre-trained model, and a dynamic mask mechanism is used to enhance the learning of the pre-trained model on molecular topology correlation; The pre-trained model is loaded, fine-tuned using the data enhancement base set, a ring-forming reaction special model is constructed, and a weighted Top-k index is used to evaluate the prediction accuracy of the model; Based on the ring-forming reaction special model, an isomeric system call interface is constructed to enable bidirectional data interaction between the interactive interface and the ring-forming reaction special model.

[0007] In some embodiments, the data processing includes: Merging multiple source data, and realizing cross-library field alignment through customized data mapping rules; Converting molecules in the reaction sequence to Canonical SMILES based on the rdkit toolkit, and eliminating isomerism of molecular structure representation based on SMILES specification algorithm; Based on the Canonical SMILES, using a hash comparison algorithm for deduplication, calculating the hash value of the reaction SMILES sequence combination, and retaining unique reaction samples.

[0008] In some embodiments, the chemical plausibility verification includes: Using molecular structure analysis of the rdkit toolkit to screen multiple source chemical reaction data to identify and eliminate entries that cannot generate valid molecular objects and SMILES parsing failures, wherein the invalid molecular objects include bond errors and abnormal atomic valence structures.

[0009] In some embodiments, the judgment of new ring formation by comparing the number of rings and ring characteristics of reactants and products includes: Using the rdkit toolkit to extract the symmetric independent minimum ring set of reactants and products; Statistically analyzing the total number of rings of reactants and products, and if the total number of rings of products is greater than that of reactants, it is determined to be a ring-forming reaction; If the number of rings is the same, compare the ring characteristics, including ring size, atom type in the ring, and bond type in the ring, and if the product has new ring characteristics not included in the reactant, it is determined to be a ring-forming reaction.

[0010] In some embodiments, the reaction center offset maximum common substructure alignment algorithm is used to generate multiple sets of aligned SMILES sequences for the reactions in the special data set, and a data enhancement base set is constructed, including: Determine the reaction center by analyzing the bond type change value and the atomic valence change value, wherein the bond type change value is the sum of the difference of the bond type of atoms in the reactants and the products, and the atomic valence change value is the absolute difference of the atomic valence; Screen the maximum common substructure of the reactants and the products, select the chain end atoms far from the reaction center in the maximum common substructure as the starting nodes of the molecular graph traversal, and generate the aligned SMILES sequence of the reactants and the products; Perform multiple data augmentation on the special data set, and realize it by selecting the atoms corresponding to the second and third levels of the distance from the reaction center in the maximum common substructure as the starting points of traversal.

[0011] In some embodiments, the OpenNMT framework is used to pre-train an initial model based on general chemical reaction data to obtain a pre-trained model, and a dynamic mask mechanism is used to enhance the learning of the pre-trained model on molecular topology correlation, including: The general chemical reaction data is divided into a training set, a validation set, and a test set according to a target proportion; Randomly select a preset percentage of character positions in the SMILES of the reactants in the general chemical reaction data for mask processing; wherein the mask positions of the target percentage complementary to the preset percentage are replaced by symbols, half of the preset percentage is replaced by random characters, and the other half of the preset percentage is kept as the original characters, to simulate noise interference in molecular structure analysis; A Transformer model is constructed based on the OpenNMT framework, an Adam optimizer and a Noam learning rate schedule are used for training, the learning rate is dynamically adjusted based on the performance of the validation set, and an early stopping mechanism is triggered when the model performance does not significantly improve for a continuous preset number of cycles, to obtain a pre-trained model.

[0012] In some embodiments, the pre-trained model is fine-tuned using the data augmentation base set, including: The data augmentation base set is divided into a training set, a validation set, and a test set according to a target proportion, and a stratified sampling strategy is used to ensure that the distribution deviation of different types of ring-forming reactions in the training set, the validation set, and the test set is controlled within a preset deviation threshold; Load the encoder and decoder weight parameters of the pre-trained model, only initialize the top layer network, and after training, use a sliding window parameter averaging method to fuse the model parameters of the last preset number of checkpoints.

[0013] In a second aspect, an embodiment of the present application provides a ring-forming reaction retrosynthesis device, including: A collection module is configured to collect multi-source chemical reaction data, and after data processing, an initial data set is obtained, wherein the data processing includes chemical reasonableness verification; The screening module is configured to identify a ring-forming reaction by using a minimal ring system analysis based on the initial data set, and to screen a special data set for cyclic skeleton synthesis, wherein the analysis identification comprises judging new ring formation by comparing the number and characteristics of rings of reactants and products; The construction module is configured to generate a plurality of aligned SMILES sequences for reactions in the special data set by using a reaction center offset maximum common substructure alignment algorithm, and to construct a data augmentation base set; The training module is configured to pre-train an initial model to obtain a pre-trained model by using general chemical reaction data based on an OpenNMT framework, and to enhance learning of the pre-trained model on molecular topology correlation by using a dynamic mask mechanism. The evaluation module is configured to load the pre-trained model, fine-tune the pre-trained model by using the data augmentation base set, construct a ring-forming reaction special model, and evaluate prediction accuracy of the model by using a weighted Top-k index. The interaction module is configured to construct a heterogeneous system call interface based on the ring-forming reaction special model, so that a data bidirectional interaction between an interaction interface and the ring-forming reaction special model is realized.

[0014] In a third aspect, an embodiment of the present application provides an electronic device, including a memory and a processor, and program code stored in the memory and executable on the processor, and the program code is executed by the processor to implement the reverse synthesis method of the ring-forming reaction as introduced in any of the embodiments of the first aspect.

[0015] In a fourth aspect, an embodiment of the present application provides a computer storage medium, which stores one or more programs executable by an electronic device as introduced in the third aspect to implement the reverse synthesis method of the ring-forming reaction as introduced in any of the embodiments of the first aspect.

[0016] The embodiment of the application provides a ring formation reaction reverse synthesis method, device, electronic equipment and storage medium, which comprises the following steps: collecting multi-source chemical reaction data, obtaining an initial data set through data processing, wherein the data processing comprises chemical rationality verification; based on the initial data set, a minimum ring system analysis is used to identify ring formation reactions, and a special data set for cyclic skeleton synthesis is obtained through screening; new ring formation is judged by comparing the number and characteristics of the rings of reactants and products; a reaction center offset type maximum common substructure alignment algorithm is used to generate a plurality of aligned SMILES sequences from the reactions in the special data set; a data enhancement base set is constructed; an initial model is pre-trained based on an OpenNMT framework and general chemical reaction data to obtain a pre-trained model; the pre-trained model is enhanced by a dynamic mask mechanism to learn the molecular topological correlation; the pre-trained model is loaded, and the data enhancement base set is used for fine tuning to construct a ring formation reaction special model, and a weighted Top-k index is used to evaluate the prediction accuracy of the model; based on the ring formation reaction special model, a heterogeneous system call interface is constructed to realize the bidirectional interaction of the interactive interface and the data of the ring formation reaction special model, wherein the heterogeneous system call interface builds a system interface containing compound input, parameter setting and result display, so as to solve the technical problems of low prediction accuracy of the general reverse synthesis model for ring formation reactions and the inability to accurately capture the unique mechanism of ring formation reactions.

[0017] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the application, nor is it used to limit the scope of the application. Other features of the application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0018] The application will be described in more detail below based on the embodiments and with reference to the accompanying drawings.

[0019] Figure 1 An exemplary ring formation reaction reverse synthesis method flowchart is shown in an embodiment of the application; Figure 2 An exemplary ring formation reaction reverse synthesis method specific implementation flowchart is shown in an embodiment of the application; Figure 3 An exemplary model training flowchart suitable for single-step reverse synthesis of ring formation reactions is shown in an embodiment of the application; Figure 4 A ring formation reaction identification diagram is shown in an embodiment of the application; Figure 5 A Top-N accuracy evaluation diagram of a ring formation reaction special model and an existing general model is shown in an embodiment of the application; Figure 6An output result sample diagram of the same input under test of a ring-forming reaction special model proposed in an embodiment of the present application and an existing template method-based special model is shown. Figure 7 A system architecture diagram proposed in an embodiment of the present application is shown. Figure 8 A system use embodiment result display diagram proposed in an embodiment of the present application is shown. Figure 9 A reactant content generation diagram proposed in an embodiment of the present application is shown. Figure 10 A structure block diagram of a ring-forming reaction reverse synthesis device proposed in an embodiment of the present application is shown. Figure 11 A structure block diagram of an electronic device for performing the ring-forming reaction reverse synthesis method according to an embodiment of the present application is shown. Figure 12 A computer readable storage medium for saving or carrying the ring-forming reaction reverse synthesis method according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0020] To make the objectives, technical solutions and advantages of the present application clearer, further detailed description will be given below in combination with embodiments and drawings, and the illustrative embodiments of the present application and the description thereof are only used to explain the present application and do not limit the present application.

[0021] In the prior art, the uniqueness of the ring-forming reaction makes it difficult to be effectively covered by a general reverse synthesis model, and the general model often cannot fully capture the entropy enthalpy balance, tension effect and multi-step coordination mechanism in the ring-forming process, resulting in low prediction accuracy.

[0022] The applicant has found that due to these unique reaction mechanisms of the ring-forming reaction, a special model can be trained to learn and optimize these specific information.

[0023] Therefore, compared with the general model, the single-step reverse synthesis model special for the ring-forming reaction can provide higher accuracy and more specific prediction results when predicting the corresponding reaction. For the construction process of the ring skeleton, the key ring-forming step plays a core role in the entire synthesis process. Whether the ring-forming reaction is successful directly determines the feasibility and efficiency of the entire synthesis route. Therefore, in the single-step skeleton reverse synthesis prediction, developing a special model for the ring-forming reaction has a very high value.

[0024] To solve the above problems, the applicant provides a ring-forming reaction inverse synthesis method, device, electronic equipment and storage medium. The special model is trained by using the molecular node alignment method without atomic mapping to improve the ring-forming reaction prediction accuracy, and the difficulty of the current new type of cyclic compound skeleton inverse synthesis disassembly problem is solved to a certain extent.

[0025] The ring-forming reaction inverse synthesis method is described in detail in the subsequent embodiments.

[0026] The application scenarios of the ring-forming reaction inverse synthesis method provided in the embodiments of the application are introduced as follows: Please refer to Figure 1 , Figure 1 A flowchart of the ring-forming reaction inverse synthesis method provided in the embodiments of the application is shown in the figure. In this embodiment, the ring-forming reaction inverse synthesis method can be applied to, for example, Figure 10 The ring-forming reaction inverse synthesis device 300 and Figure 11 The electronic equipment 200, which includes a computer, a mobile terminal and the like. In addition, the ring-forming reaction inverse synthesis method can also be specifically applied to Figure 7 The system architecture and Figure 8 The system interface module. The ring-forming reaction inverse synthesis method will be described in detail with reference to Figure 1 The flowchart is shown in the figure. The ring-forming reaction inverse synthesis method can include S110 to S160.

[0027] S110: Collect multi-source chemical reaction data, and obtain an initial data set after data processing. The data processing includes chemical reasonableness verification.

[0028] The data processing includes merging, cleaning and deduplication of the collected multi-source chemical reaction data.

[0029] In some embodiments, the data processing in S110 includes S111 to S113.

[0030] S111: Merge the multi-source data, and realize cross-database field alignment through customized data mapping rules.

[0031] In the embodiments of the application, in order to construct the basic data set for the research of the ring-forming reaction single-step inverse synthesis, reaction-related raw data can be obtained in batches from the industry-recognized authoritative chemical reaction database. These data cover reaction substrates, product structures and reaction conditions and other key information. Due to the differences in the format specifications of different data sources, cross-database data merging needs to be completed through customized data mapping rules to ensure field alignment and information integrity.

[0032] In some embodiments, the chemical reasonableness verification in S110 includes: Screening of multi-source chemical reaction data by molecular structure analysis of rdkit toolkit to identify and eliminate entries that cannot generate valid molecular objects and SMILES parsing failures, including valence bond errors and abnormal atomic valence structures.

[0033] In the embodiments of the present application, by introducing a chemical reasonableness checking mechanism, all reaction data quality screening is performed using the rdkit molecular structure analysis module to automatically identify and eliminate entries that cannot generate valid molecular objects (such as structures with bond errors and abnormal atomic valence) and SMILES parsing failures.

[0034] S112: Convert molecules in the reaction sequence to Canonical SMILES based on the rdkit toolkit, and eliminate isomerism of molecular structure representation based on SMILES specification algorithm.

[0035] In the embodiments of the present application, with the help of rdkit toolkit, the Canonical SMILES generation function is called to convert all molecules (including substrates, intermediates and products) in the reaction sequence to unique canonical codes, and SMILES specification algorithm is used to eliminate isomerism of molecular structure representation, realizing global consistency of molecular identification.

[0036] S113: Based on Canonical SMILES, use hash comparison algorithm for deduplication, calculate the hash value of the reaction SMILES sequence combination, and retain unique reaction samples.

[0037] In the embodiments of the present application, based on the standardized Canonical SMILES sequence, the hash comparison algorithm is used for reaction data deduplication, the hash value of the SMILES sequence combination of each reaction is calculated, the redundant data with completely matched substrate and product SMILES are quickly located, and only single sample is retained to reduce redundancy. Finally, an initial data set sample is formed which can guarantee the uniqueness and accuracy of reaction information and provide high quality support for subsequent model training.

[0038] S120: Based on the initial data set, use the smallest ring system analysis to identify ring-forming reactions, and screen to obtain a special data set for cyclic skeleton synthesis, and analyze and identify the formation of new rings by comparing the number of rings and ring characteristics of reactants and products.

[0039] In some embodiments, the formation of new rings by comparing the number of rings and ring characteristics of reactants and products in S120 includes steps S121 to S123, wherein: S121: Extract the symmetric independent minimum ring set of reactants and products using the rdkit toolkit.

[0040] In the embodiments of the present application, to realize accurate identification of ring-forming reactions, it is necessary to strictly judge whether a new ring is formed in combination with minimum ring system (SSSR) analysis. First, the minimum ring system of the reactants and products, i.e., the smallest set of rings that cannot be further split, is extracted using RDKit.

[0041] For example, the minimum ring system of benzene is one 6-membered ring, and the minimum ring system of naphthalene is two fused 6-membered rings.

[0042] In S122, the total number of rings of the reactants and the total number of rings of the products are counted, and if the total number of rings of the products is greater than the total number of rings of the reactants, it is determined that a ring-forming reaction has occurred.

[0043] In the embodiments of the present application, by counting the total number of rings of the reactants (the sum of the number of rings in all minimum ring systems) and the total number of rings of a single product, if the total number of rings of the product is greater than the total number of rings of the reactant, it is directly determined that a ring-forming reaction has occurred.

[0044] In S123, if the number of rings is the same, the ring characteristics are compared, including ring size, atom type in the ring, and bond type in the ring. If the product has a new ring characteristic that is not included in the reactant, it is determined that a ring-forming reaction has occurred. In the embodiments of the present application, if the number of rings is the same, the ring characteristics of the minimum ring systems of the two are further compared, including ring size (number of atoms), atom type in the ring, and bond type in the ring (single bond / double bond distribution). If there is a characteristic of a ring in the minimum ring system of the product that does not match the characteristics of all the rings of the reactant (i.e., the ring is a new ring that does not exist in the reactant), it is determined that a ring-forming reaction has occurred. Otherwise, no ring-forming reaction has occurred.

[0045] In the embodiments of the present application, through the fine feature comparison of the minimum ring system, both the change in the number of rings and the replacement of the ring structure (such as the formation of a new ring system after ring opening and ring closure) can be captured, significantly improving the accuracy of ring-forming reaction identification.

[0046] In S130, a reaction center offset type maximum common substructure alignment algorithm is used to generate a plurality of aligned SMILES sequences for the reactions in the special data set, and a data enhancement base set is constructed.

[0047] In the embodiments of the present application, the atom nodes of the reactant and product SMILES in all reaction data need to be aligned. S130 includes S131 to S133. In S131, the reaction center is determined by analyzing the bond type change value and the atom valence change value. The bond type change value is the sum of the difference in the bond type of an atom between the reactant and the product, and the atom valence change value is the absolute difference in the valence of the atom.

[0048] In the embodiments of the present application, the reaction center is determined by analyzing the bond type change and the difference in the atom valence.

[0049] S132: Screening the maximum common substructure (MCS) of the reactant and the product, selecting the chain end atom far from the reaction center in the maximum common substructure as the starting node of the molecular graph traversal, and generating the aligned reactant and product SMILES sequences.

[0050] In the embodiments of the application, the maximum common substructure (MCS) containing the same large fragment is screened; the chain end atom far from the reaction center in the MCS is selected as the starting node of the graph traversal, and the molecular graph depth traversal is performed from this starting point, the continuity of the skeleton is preferentially retained, and a new product SMILES is generated.

[0051] S133: Multiple data augmentation is performed on the special data set, which is realized by selecting the atoms corresponding to the second and third levels of distance from the reaction center in the maximum common substructure as the traversal starting points.

[0052] In the embodiments of the application, based on the root node reverse mapping of the product SMILES, the corresponding MCS atoms in the reactant SMILES are matched as the root node, and the same traversal rule is used to generate the aligned reactant SMILES.

[0053] Specifically, the general reaction data with a large amount of data is subjected to 5-fold data augmentation, and the ring-forming reaction data set with a small amount of data is subjected to 10-fold data augmentation, which is realized by selecting the atoms in the MCS with different levels of distance from the reaction center as the starting points of the molecular graph traversal, that is, the topological distances of all atoms in the MCS from the reaction center are first calculated, and then the atoms ranked second and third in distance are preferentially selected as the root nodes, and 4 and 9 additional aligned sequences are generated for each original reaction. Finally, the processed reactant and product SMILES are split according to the order of atom appearance and stored in Reactant.txt and Product.txt, respectively, as input sequences for model pre-training.

[0054] S140: Based on the OpenNMT framework, an initial model is pre-trained using general chemical reaction data to obtain a pre-trained model, and the learning of the pre-trained model on the molecular topological correlation is enhanced through a dynamic masking mechanism.

[0055] S140 includes S141 to S143: S141: The general chemical reaction data is divided into a training set, a validation set, and a test set according to a target ratio.

[0056] In the embodiments of the application, the general data set is divided into a training set, a validation set, and a test set in a ratio of 8:1:1.

[0057] S142: randomly select a preset percentage of character positions in the reactant SMILES in the general chemical reaction data for mask processing; wherein the preset percentage of target percentage of mask positions are replaced with symbols, the preset percentage of half are replaced with random characters, and the preset percentage of the other half remain the original characters, to simulate noise interference in molecular structure analysis.

[0058] In the embodiments of the application, for the SMILES representation of the reactants, dynamic random mask covering processing is implemented, and 20% of the character positions are randomly selected for mask replacement in each iteration (covering key elements such as atomic symbols, chemical bonds and branch markers), wherein 80% of the mask positions are replaced with a special symbol [MASK], 10% of the positions are replaced with random characters, and 10% of the positions remain original characters. This mask mechanism simulates the noise interference scenario in molecular structure analysis, forcing the model to learn the topological correlation and chemical semantic logic between molecular fragments, and laying a foundation for subsequent generation of SMILES that meet chemical rationality.

[0059] S143: build a Transformer model based on the OpenNMT framework, train using the Adam optimizer and Noam learning rate scheduling, dynamically adjust the learning rate based on the performance of the validation set, trigger the early stopping mechanism when the model performance does not significantly improve for a continuous preset number of cycles, and obtain a pre-trained model.

[0060] In the embodiments of the application, after pre-training, the initial model parameters are loaded using the transfer learning strategy, and the general reaction data set is used for multi-stage training fine-tuning. During the training, the learning rate decay coefficient is dynamically adjusted based on the reaction prediction accuracy of the validation set, and the early stopping mechanism is triggered when the model performance does not significantly improve for a continuous 5 cycles. The final general reaction reverse synthesis model can be obtained.

[0061] S150: load the pre-trained model, fine-tune using the data augmentation base set, build a ring reaction dedicated model, and use the weighted Top-k index to evaluate the prediction accuracy of the model.

[0062] In some embodiments, the fine-tuning using the data augmentation base set in S150 includes S151 to S152.

[0063] S151: divide the data augmentation base set into a training set, a validation set and a test set according to a target ratio, and use a stratified sampling strategy to ensure that the distribution deviation of different types of ring reactions in the training set, the validation set and the test set is controlled within a preset deviation threshold.

[0064] In the embodiments of the present application, the data division ratio consistent with the general model training is adopted, i.e., 8:1:1, and the stratified sampling strategy is adopted to divide the ring formation reaction data set, so as to ensure that the distribution deviation of different types of ring formation reactions (covering the dimensions of ring formation atom type, ring system size and cyclization mode) in the training set, the validation set and the test set is controlled within 3%, thereby eliminating the interference of uneven sample distribution on the special capability of the model.

[0065] In S152, the encoder and decoder weight parameters of the pre-trained model are loaded, only the top layer network is initialized, and after the training is completed, the model parameters of the last preset number of checkpoints are fused by using the sliding window parameter average method.

[0066] In the embodiments of the present application, the vocabulary tables of reactants and products are constructed for the special data set. Before the special model training is started, the parameter migration strategy is adopted to load the encoder and decoder weight parameters of the general pre-trained model, and only the top layer network is initialized. In order to further improve the stability of the model and reduce the interference of noise, after the training is completed, the model parameters of the last 5 checkpoints are fused by using the sliding window parameter average method: by calculating the arithmetic mean of each parameter tensor in the window, the parameter jitter introduced by the random gradient fluctuation in the later training stage is effectively suppressed.

[0067] In addition, in S150 of the present application, the weighted Top-K test is performed on the averaged model by using the test set to view the performance of the model.

[0068] In S160, based on the ring formation reaction special model, a heterogeneous system call interface is constructed to realize the bidirectional interaction between the interactive interface and the ring formation reaction special model.

[0069] In some embodiments, the heterogeneous system call interface includes: The input interface function receives the molecule SMILES and the model parameters, and verifies the legality of the input by using the regular expression and the RDKit; The output interface function stores the reactant SMILES list output by the model in the CSV format, and the file naming includes the timestamp and the MD5 check value of the SMILES; The compound input module of the system interface supports the SMILES format input and real-time syntax checking, the parameter setting module supports adjusting the output quantity and the beam size, and the result display module supports the molecule structure visualization and the result export.

[0070] In the embodiments of the present application, reference is made to Figure 7The system architecture shown builds a system call interface to realize data flow between the interactive interface and the model. Specifically, the input and output interface functions need to be customized, which take molecular SMILES and model parameters as input, then call the start model, output the corresponding number of reactant molecular SMILES list, and store it as a CSV format file; at the same time, the above interface functions are encapsulated and packaged into an EXE executable file. Then the visual interactive interface realizes the function of transmitting input data to the model by calling the interface built by the model execution file, and captures the output results of the model and visualizes them.

[0071] It should be noted that the software system interface is built so that the interface can directly call the interface to complete the use of the model.

[0072] Reference is made to Figure 8 The system interface module reference diagram includes at least the following three modules: compound input module (for inputting target products that need to be predicted), parameter setting module (through inputting modifiable parameters to adjust the output of the model), and result display module (for displaying the SMILES molecular structure output by the model). These modules are combined and information is transmitted to each other to complete the construction of the system interface.

[0073] Reference is made to the following content in the specific embodiments of the present application: Reference is made to Figure 2 The specific implementation flowchart of the reverse synthesis method of the ring-forming reaction includes the following steps: S1: Collect chemical reaction data from multiple data sources and perform operations such as merging, cleaning, and deduplication.

[0074] In this step, chemical reaction related data is downloaded from Open Reaction Database (ORD) and Hugging Face. It should be noted that the downloaded data file format and data structure have diversity, but ultimately they need to be merged into a csv file for subsequent processing.

[0075] For the data downloaded on ORD (in compressed.pb.gz files), the data needs to be converted to JSON format through the API call of the ord_schema package, and then the chemical reaction data part is read and converted to csv for easy merging with other data. The final merged csv data table needs at least one column to represent the chemical reaction, and its format is expressed as "Reactant >> Product", where "Reactant" and "Product" are both molecular SMILES sequences, and ">" is the actual symbol. For data with multiple Reactants or Products, the SMILES are separated by the symbol ".".

[0076] A python script is used to read the data and delete all data with missing reactants or products, i.e. reaction data represented directly by data lines starting and ending with ">>". All SMILES are represented, identified and normalized using the Rdkit tool to become Canonical SMILES, and SMILES that cannot be identified are deleted.

[0077] Among them, the reactants and products can refer to the ring-forming reaction identification schematic diagram shown in Figure 4 .

[0078] S2: Adopting a ring-forming reaction identification strategy based on the analysis of the smallest ring system to obtain a ring skeleton synthesis related reaction special data set from the reaction data set.

[0079] In this step, based on strict algorithm logic, it is mainly realized through the RDKit tool package. First, the number of rings of the reactant and product molecules is calculated, i.e. the Symmetrically Independent Smallest Set of Smallest Rings (SSSR) is obtained. If the total number of rings of the product exceeds the total number of rings of the reactant (which can be represented as: , where and represent the ring structure set of the product and the reactant, respectively), it is directly determined as a ring-forming reaction. As shown in Figure 6 , this judgment condition is intuitive and efficient, and can quickly identify obvious ring-forming reactions.

[0080] If the number of rings has not increased, a more detailed ring structure comparison analysis is needed. Through RDKit, the atomic composition set of each ring is obtained, and all ring structures of the reactant and product are stored as ordered atomic tuple sets. Specifically, the atomic index in each ring is sorted to form a tuple, ensuring the unique representation of the ring structure. Then, through set operation, it is judged whether there is a ring structure in the product that does not appear in the reactant (i.e. If such a new ring structure exists, a cyclization reaction is considered to have occurred; otherwise, no cyclization is considered to have occurred. This strategy can not only identify reactions with an increase in the number of rings, but also capture complex reaction types such as ring rearrangement, ring expansion, or ring contraction, significantly improving the accuracy of cyclization reaction identification. The specific reaction discrimination method is as follows: 1. Ring rearrangement reaction: When the number of rings in the product and the reactant is the same but the ring structure is different, the changes in the atomic connection paths can be compared. If there is a recombination of atomic bonding relationships (such as the migration between nitrogen and carbon atoms in the Beckmann rearrangement), it is determined to be a ring rearrangement.

[0081] 2. Ring expansion reaction: If the product contains a ring with a similar ring structure to the reactants but with an increased number of atoms (e.g., cyclohexanone expands under specific conditions to form cycloheptanone), the difference in the number of ring atoms can be calculated. (where n is the number of ring atoms), when It was identified as an expanded ring at the time.

[0082] 3. Ring-condensation reaction: The opposite of ring-expansion reaction, if the number of ring atoms in the product decreases... If there is a corresponding relationship between the ring structures, it is determined to be a ring-condensation reaction, such as the ring-opening of a bicyclic compound to form a monocyclic structure.

[0083] S3: Employ a reaction center offset-based maximum common substructure alignment algorithm to obtain multiple aligned reaction SMILES sequences for each reaction in the dataset, and construct a data augmentation base set for subsequent model training.

[0084] In this step, the reaction center offset maximum common substructure (MCS) alignment algorithm is used to align the SMILES sequences of reactants and products, thereby reducing the edit distance of SMILES between reactants and products, and simultaneously performing data augmentation on the reaction data. Considering that the SMILES sequence characterization originates from a depth-first search on the molecular graph, and that most substructures between reactants and products are highly consistent, for a given DFS order of any product, there must exist a corresponding DFS order on the reactant molecular graph, making the order of corresponding atoms on reactants and products as similar as possible. This greatly constrains the edit distance between input and output, making them highly similar.

[0085] Specifically in the molecular sequence alignment and data augmentation process of S3 stage, the core objective is to shorten the edit distance between reactants and products by optimizing the sequence structure of reaction SMILES, providing high-quality input data for model training. This process relies on the reaction center biased MCS alignment algorithm: for each cleaned reaction SMILES, first locate the reaction center by analyzing the bond type change and atomic valence difference. Let the reactant molecular graph be , and the product molecular graph be , for each atom , calculate its: 1. Bond type change value:

[0086] where and are the types of the bond (single bond = 1, double bond = 2, triple bond = 3, etc.) in the reactant and product, respectively, is the set of adjacent atoms of .

[0087] 2. Valence change value:

[0088] where and are the valence of v in the reactant and product, respectively.

[0089] The reaction center set C is defined as:

[0090] where and are pre-set thresholds (usually ).

[0091] After locating the reaction center, the MCS of the reactant and product is screened by the following steps: (1) Construct the atomic compatibility matrix , where:

[0092] (2) Find the maximum common substructure MCS, which satisfies:

[0093] and for any , if , then , and vice versa.

[0094] During sequence alignment, the chain-end atoms in the MCS that are far from the reaction center are used as the starting point for graph traversal (such as the terminal methyl carbon of an alkane chain). The molecular graph is traversed according to the depth-first rule, prioritizing the preservation of the continuity of the skeletal structure such as benzene rings and carbon chains to generate new products SMILES. Then, using the root node of the product SMILES as the reference, the corresponding MCS atoms in the reactant SMILES are matched in reverse as the root node, and the same traversal rule is used to generate aligned reactant SMILES, ensuring that the common substructures of the two correspond in position in the sequence (for example, the benzene ring fragment of the aligned reactant and the benzene ring fragment of the product are in the same interval in the string).

[0095] The data augmentation phase employs a differentiated strategy based on the dataset size: Let the augmentation factor be k. For large-volume general reaction datasets, k=5; for scarce cyclic reaction datasets, k=10. Specifically, this is achieved by adjusting the starting point of the molecular graph traversal. First, the topological distance from all atoms within the MCS to the reaction center is calculated: For atom u in the MCS and reaction center c, the topological distance is defined as: (e.g., the distance between adjacent atoms is 1, and the distance between meta atoms is 2).

[0096] After sorting by distance from farthest to nearest, let the sorting result be... (in, ), select the distance from the second ranked The third As a new starting point, all processed reactants and products are ultimately separated according to the order of atomic appearance and stored in Reactant.txt and Product.txt respectively, forming the basic input sequence for model pre-training and laying the data foundation for the subsequent model to learn the molecular structure changes before and after the reaction.

[0097] S4: Based on the OpenNMT framework, a pre-trained model is constructed using general chemical reaction data.

[0098] In this step, first, the pretreated general dataset is randomly divided into training set, validation set and test set according to the ratio of 8:1:1, ensuring uniform data distribution. Then, for the reactant SMILES in the training set, molecular graph analysis is performed using the RDKit library, and 20% of the atomic positions are randomly selected for mask processing. The mask strategy uses a composite replacement mechanism: 80% probability of replacing [MASK] label, 10% probability of replacing random atom symbol (such as replacing C with N), and 10% probability of keeping the original symbol. For example, the SMILES of anisole: COc1ccccc1, may be masked as [MASK]Oc1[MASK] ccc[MASK]1. Note that when processing: filter invalid SMILES (cannot be converted into molecules); keep the key reaction center for complex molecules; generate the SMILES pair before and after masking as pre-training data.

[0099] The model construction and training adopts a progressive learning strategy, which is divided into two steps of model construction and training, namely pre-training and fine-tuning. In the pre-training stage, the onmt_build_vocab command is used to construct a bidirectional mapping SMILES vocabulary based on the full training data, generating source (masked SMILES) and target (original SMILES) vocabulary files containing atom symbols, bond types, and branch markers. The model architecture adopts an enhanced Transformer design: a standard 6-layer encoder-decoder structure, 512-dimensional word vectors and hidden layer dimensions, 8 attention heads providing multi-scale structure information capture capabilities, and 2048-dimensional feedforward networks enhancing non-linear representation capabilities. In particular, a position encoding matrix is introduced in the attention mechanism to encode the topological distance information between atoms; a molecular fragment memory module is added to the decoder to preserve key substructure information. The training parameter configuration is as follows: use the Adam optimizer (β1=0.9, β2=0.998) with Noam learning rate scheduling, initial learning rate 2.0, and warmup_steps set to 8000; use gradient accumulation technology (accum_count=4) to achieve an equivalent batch size of 16384 tokens; disable gradient clipping (max_grad_norm=0.0) to preserve weak gradient signals; train continuously for 200k steps, saving model checkpoints every 5k steps.

[0100] Fine-tuning stage, adopt transfer learning strategy to load initial model parameters, use general reaction dataset for multi-stage training fine-tuning, after loading pre-trained model parameters, switch training data to product-reactant pairs, adjust learning rate to 1 / 2 (i.e. 1.0) of the pre-training stage, perform 100k steps of domain adaptation training. Dynamic learning rate decay mechanism is implemented during training: evaluate model performance on validation set every 10k steps, monitor Top-1 SMILES reconstruction accuracy and Mean Rank indicators, when the indicators improve less than 0.5% for 3 consecutive evaluation periods, the learning rate is decayed by 0.95 times. At the same time, introduce early stopping mechanism, when the Exact Match accuracy on the validation set stagnates for 5 consecutive periods, terminate training.

[0101] S5: Load the pre-trained model and train it using ring-forming reaction data to complete the construction of a model specialized for the synthesis of cyclic compound skeletons and evaluate the model using a weighted Top-k index.

[0102] In this step, after completing the pre-training of the general model, a specialized model needs to be further constructed and its performance optimized for specific types of retrosynthesis prediction tasks such as ring-forming reactions. First, perform key preprocessing operations on the specialized dataset: split it into training, validation, and test sets in a strict ratio of 8:1:1 to provide diverse learning samples for the model. Stratified sampling strategy can be used to ensure balanced distribution of different types of ring-forming reactions in each subset. Let the original dataset contain types of ring-forming reactions, with samples for each type, then the total number of samples . Stratified sampling is achieved through the following steps: Calculate the number of samples each type should allocate to the training, validation, and test sets, respectively , and to ensure that the sample proportion of each subset is strictly 8:1:1.

[0103] For each reaction type , use random sampling method to extract samples from samples to form the training set, then extract samples from the remaining samples to form the validation set, and finally the remaining samples form the test set.

[0104] Repeat this operation for all types of reactions, and finally obtain the training set , validation set , and test set In each reaction type, the proportion of each reaction type is consistent with the original data set, so as to achieve the goal of stratified sampling.

[0105] Then, a special and independent vocabulary table is constructed for the reactants and products respectively to train the ring-forming reaction model. The independent vocabulary table can encode the structural information more meticulously and avoid confusion with general field expressions. Subsequently, the special model training phase is entered. To make full use of the chemical knowledge learned by the general pre-trained model, its parameters are loaded before training to serve as the initialization basis for the special model. Then, the ring-forming reaction special data is used to fine-tune the model. To prevent excessive deviation from general knowledge, a small learning rate of 0.9 is used during training, and a layer-by-layer unfreezing strategy is combined: first, only the top layer of the model is trained to allow the model to adapt to the characteristics of the special data, and then the bottom layer parameters are gradually unfrozen to allow the model to learn the unique rules of ring-forming reactions while retaining general chemical understanding.

[0106] After training, to improve model stability and reduce noise interference caused by parameter fluctuations during training, the parameters of the last 5 checkpoints are averaged. By calculating the mean of each element in the parameter matrix corresponding to the 5 checkpoints, the smoothed final model parameters are generated. This operation effectively reduces the local optimal fluctuations generated by the model during a specific training phase and improves the model's generalization ability in the special field.

[0107] Finally, the weighted Top-K test is performed on the averaged model using the test set to evaluate its performance in the ring-forming reaction prediction task.

[0108] Due to the data augmentation strategy, the model outputs 10*10 prediction results for each predicted molecule. To select reliable predictions from these results, a comprehensive weighting scheme is used: the weighted score of each prediction result is which is calculated by the following formula:

[0109] wherein, represents the number of occurrences of the ith prediction result, and the higher the occurrence frequency, the stronger the reliability; represents the ranking of the result in the jth prediction, which is calculated by

[0110] The results with higher rankings obtain higher weights, is an adjustable weight decay rate; is the position of the first occurrence of the result, which controls the degree of its impact on the score, and the earlier the position, the higher the weight.

[0111] Through the above weighting strategy, combined with the sequential ranking of model output, the number of occurrences of the same result and the first occurrence position, the information redundancy brought by data enhancement can be effectively integrated, the prediction randomness can be reduced, and finally the result with the highest score or the Top-K results are obtained, which are used to calculate professional indicators such as accurate matching accuracy, structural similarity, ring closure accuracy, etc. to comprehensively measure the accuracy and reliability of the model in the ring formation reaction prediction task. Referring to Figure 3 An exemplary model training process diagram applicable to single-step reverse synthesis of ring formation reactions is shown. It can be seen that a general model is trained using existing template-free generation-based methods using general reaction data (which includes ring formation reaction-specific training set data). The comparison results of the two models under the same independent test set (ring formation reaction data constitute the independent test set) are shown in Figure 6 The output result example diagram of the ring formation reaction specific model and the existing template-based method specific model under the same input is shown.

[0112] In Figure 6 , it is obvious that the model trained using the method of the present patent is completely superior to the existing general model based on the template-free method in the Top-10 index comparison in the skeleton synthesis and disassembly task. Especially for Top-1, the method of the present patent can reach 80.588%, while the Top-1 accuracy of the general model can only reach 48.321%. For the final result Top-10 accuracy, the model trained by the method of the present patent can reach 95.222%, while the general model can only reach Top-10 accuracy 85.951%. From Figure 5 and the above analysis, it can be seen that in the task of reverse synthesis and disassembly of cyclic compound skeletons, it is very necessary to train a specific model. Due to the specificity of the data, the specific model can better learn the synthesis mechanism of the ring skeleton, and can more accurately predict the reverse synthesis of the ring formation reaction compared to the general model.

[0113] The application of the present embodiment to the actual task of reverse synthesis and disassembly of cyclic skeleton, through Figure 3 the process (the model training process of the present patent) is completed, for inputting a product SMILES with a cyclic skeleton: O=c1c2cc3scnc3cc2[nH]c2ncccc12, the output quantity is set to 10, as shown in Figure 6As shown, the model will output 10 different reactants, and the third result predicted by the patented method model is the label result for that reactant. In contrast, existing template-based methods, due to limitations in the template library, can only use a limited number of templates for this product, ultimately outputting only two results without a correct label. Clearly, the dedicated model trained using this patented method, compared to existing template-based models, ensures significantly greater output diversity, and all its outputs correspond to cyclic reactions. This demonstrates that the patented model training method effectively captures and learns the transformation pattern from product to cyclic product in cyclic reactions.

[0114] In summary, the dedicated retrosynthetic model trained by this patented method demonstrates significant advantages in the retrosynthetic disassembly of cyclic compound skeletons, with its Top 1-10 accuracy far exceeding that of general-purpose models. Furthermore, compared to traditional template-based methods, the dedicated model is not limited by the types of template libraries, enabling it to output more diverse and accurate prediction results. This means that the dedicated model trained using this patented method has high accuracy in predicting the retrosynthetic disassembly of cyclic compound skeletons, and can alleviate the current difficulties in retrosynthetically disassembling novel cyclic compound skeletons to some extent.

[0115] This embodiment also provides a single-step retrosynthetic system suitable for cyclization reactions. To facilitate model invocation, a heterogeneous system invocation interface needs to be constructed to achieve bidirectional data interaction between the interactive interface and the retrosynthetic model. The specific implementation methods for its interface function customization and encapsulation mechanism and interface construction are as follows (the overall system architecture is as follows). Figure 7 (as shown) 1. Input interface function (InputHandler) The function prototype is defined as def InputHandler(target_smiles:str,output_num: int, beam_size:int) ->Tuple[Union[str,None], int, int], where: input parameter target_smiles must strictly follow the OpenSMILES specification established by IUPAC, supporting molecular structure encoding containing stereochemical configuration markers (such as @, @@) and isotope information; output_num is an unsigned integer, with a value domain limited to [1, 50], and the default value is set to 5; beam_size is the data augmentation multiple, and the default value is 10. The function has built-in double-layer verification logic: first-level verification passes regular expression matching SMILES syntax rules (such as bracket nesting legality, charge marker format), and second-level verification calls RDKit's MolFromSmiles interface for chemical validity verification (excludes molecules that cannot be parsed). When verification fails, return error code (0x001 indicates syntax error, 0x002 indicates chemical invalid structure) and detailed diagnostic information. 2. Output interface function (OutputHandler) The function prototype is def OutputHandler(candidate_list:List[Tuple[str, float]])->str, which receives the precursor molecule candidate set returned by the model (each element contains a SMILES string and a generation probability value), and performs data persistence according to the following rules: stored in CSV format, field definition as index (INT), precursor_smiles (STR), probability (FLOAT, 4 decimal places), UTF0-8-BOM encoding is used to ensure cross-platform compatibility. The file naming uses the format RetroResult_{timestamp}_{hash}.csv, where timestamp is the Unix timestamp (accurate to milliseconds) at the time of generation, and hash is the MD5 checksum of the target molecule SMILES with the first 8 bits, and the storage path is mapped to the %USERPROFILE%\RingRetro\Results directory by default (supports customizing the path through the registry key HKCU\Software\RingRetro\OutputPath).

[0116] 3. Packaging and executable program construction PyInstaller is used for packaging, and the packaging process is configured through the --onefile --noupx --clean parameters, and the key implementation includes: Dependency management: Explicitly declare core libraries such as RDKit and Pandas using the hiddenimports parameter.

[0117] Runtime environment: Resources are dynamically extracted through sys._MEIPASS, missing VC++ runtime components are automatically detected and repaired during program initialization, and the environment runs with zero configuration. 4. Interface Component Design This is an interactive system built on the Pyside5 framework using an MVC architecture. The main window is divided into: Target reactant input module: includes QLineEdit control, QSlider control (adjusts output_num parameter, step size 1), QPushButton control (binds model call event), and enables real-time syntax highlighting in the input area (illegal characters are highlighted in red). Model parameter input module: includes QLineEdit control, QSlider control (adjusts output_num parameter, step size 1), and QPushButton control (binds model call event). Input results need to be validated, and a pop-up reminder will be used for invalid input.

[0118] Results display module: Uses QTableView to display the candidate precursor list (supports sorting by probability and filtering by SMILES); integrates RDKit's 2D molecular rendering engine, and uses QGraphicsView control to visualize molecular structures; provides QAction collection to support result export (PNG format molecular graphs with a resolution of 300dpi, supporting batch export; CSV files can specify the storage path via QFileDialog).

[0119] 5. Inter-process communication mechanisms The user interface and the EXE program use a two-pipe communication mode: Input channel: Serialization parameters are passed via subprocess.PIPE, using | as the separator (e.g., C1=CC=CC=C1|5), and flow control is enabled (8KB buffer size) to prevent data overflow.

[0120] Output channel: After the model calculation is completed, the absolute path of the CSV file (prefixed with FILE: / / ) is returned via stdout. The interactive interface calls QFileSystemWatcher to monitor the file generation event and triggers the OutputParser thread to load the data asynchronously. Error handling: capture model runtime exceptions (such as memory overflow, calculation timeout) through stderr, exception information is encapsulated in JSON format (including error type, stack trace, recommended solution), and a modal dialog box is popped up after interface parsing to display.

[0121] In one possible embodiment, the specific example is as follows: By inputting the target product SMILES: CCOC(=O)C1=CC=CC=C1 that needs to be subjected to inverse synthesis analysis, setting the parameter Beam Size to 10, outputting the result as 5, and clicking to run the model, 5 corresponding reactants of the reaction will be obtained, and the result is as shown in Figure 9 .

[0122] As can be seen from the above description and in combination with the embodiment effect diagram, the single-step inverse synthesis system suitable for ring-forming reactions realized in the embodiment can easily call the model through simple interface operation, realize the prediction of ring-forming reactions, and visualize the output SMILES molecular character result as a corresponding molecular structure, facilitating the operator to view the result.

[0123] Please refer to Figure 10 , Figure 10 A structure block diagram of an inverse synthesis device for ring-forming reactions provided by the present application, the device comprising: a collection module 310, a screening module 320, a construction module 330, a training module 340, an evaluation module, and an interaction module 350, wherein: The collection module 310 is used to collect multi-source chemical reaction data, and after data processing, an initial data set is obtained, and the data processing includes chemical reasonableness verification; The screening module 320 is used to identify ring-forming reactions based on the initial data set, and a special data set for ring skeleton synthesis is screened out by using the smallest ring system analysis, and the analysis and identification include judging new ring formation by comparing the number of rings and the ring characteristics of reactants and products; The construction module 330 is used to generate a plurality of aligned SMILES sequences for the reactions in the special data set by using a reaction center offset type maximum common substructure alignment algorithm, and to construct a data enhancement base set; The training module 340 is used to pre-train an initial model based on an OpenNMT framework to obtain a pre-trained model by using general chemical reaction data, and to enhance the learning of the pre-trained model on molecular topology correlation by using a dynamic mask mechanism; The evaluation module 350 is used to load the pre-trained model, fine-tune the data enhancement base set, construct a ring-forming reaction special model, and evaluate the prediction accuracy of the model by using a weighted Top-k index; The interaction module 360 is configured to construct a heterogeneous system call interface based on the ring reaction special model, so that the interaction interface and the data of the ring reaction special model interact bidirectionally.

[0124] It should be noted that the device embodiments in the present application correspond to the foregoing method embodiments, and the specific principles in the device embodiments can be referred to the content in the foregoing method embodiments, which will not be described here again.

[0125] In several embodiments provided in the present embodiment, the coupling between the modules can be electrical, mechanical or other forms of coupling.

[0126] In addition, each functional module in each embodiment of the present application can be integrated in one processing module, or each module can exist physically alone, or two or more modules can be integrated in one module. The integrated module can be realized in the form of hardware or in the form of a software functional module.

[0127] Please refer to Figure 11 , Figure 11 A structural block diagram of an electronic device 200 that can perform the reverse synthesis method of the ring reaction provided in the embodiments of the present application is provided, and the electronic device 200 can be a smart phone, a tablet computer, a computer, a portable computer or the like.

[0128] The electronic device 200 further includes a processor 202 and a memory 204. The memory 204 stores a program that can perform the content in the foregoing embodiments, and the processor 202 can execute the program stored in the memory 204.

[0129] The processor 202 can include one or more cores for processing data and a message matrix unit. The processor 202 connects various parts within the entire electronic device 200 by various interfaces and lines, executes various functions of the electronic device 200 and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 204, and calling data stored in the memory 204. Alternatively, the processor 202 can be implemented in at least one of a hardware form of a digital signal processing (DSP), a field-programmable gate array (FPGA), a programmable logic array (PLA). The processor 202 can be integrated with one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes an operating system, a user interface, and an application program; the GPU is responsible for rendering and drawing display content; and the modem is used for processing wireless communication. It can be understood that the above-mentioned modem can also not be integrated into the processor, but can be implemented by a separate communication chip.

[0130] The memory 204 can include a random access memory (RAM) and a read-only memory (ROM). The memory 204 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 204 can include a program storage area and a data storage area, wherein the program storage area can store instructions for implementing an operating system, instructions for implementing at least one function (such as instructions for obtaining a random number by a user), and instructions for implementing various method embodiments described below. The data storage area can also store data created by the terminal in use (such as random numbers) and the like.

[0131] The electronic device 200 can further include a network module for receiving and sending electromagnetic waves, and a screen for displaying interface content and interacting with data. The network module can include various circuit elements for performing these functions, such as an antenna, a radio frequency transceiver, a digital signal processor, a cryptographic chip, a subscriber identity module (SIM) card, a memory, and the like. The network module can communicate with various networks, such as the Internet, an intranet, a wireless network, or other devices via the wireless network. The wireless network can include a cellular telephone network, a wireless local area network, or a metropolitan area network. The screen can display interface content and interact with data.

[0132] Reference is made to Figure 12 , Figure 12 A structure block diagram of a computer readable storage medium provided by an embodiment of the present application is shown. The computer readable storage medium 400 stores program codes 410, which can be invoked by a processor to execute the methods described in the above method embodiments.

[0133] The computer readable storage medium 400 can be an electronic storage such as a flash memory, an EEPROM (electrically erasable programmable read-only memory), an EPROM, a hard disk, or a ROM. Alternatively, the computer readable storage medium includes a non-transitory computer readable medium. The computer readable storage medium 400 has a storage space for the program codes 410 for executing any of the above methods. These program codes 410 can be read from or written to one or more computer program products. The program codes 410 can be compressed in an appropriate form, for example.

[0134] The embodiments of the present application further provide a computer program product or a computer program, which includes computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to cause the computer device to perform the inverse synthesis method of the ring formation reaction described in the various optional implementation manners above.

[0135] The above examples are only used to illustrate the technical solutions of the present application, and are not intended to limit the present application; although the present application has been described in detail with reference to the foregoing examples, those skilled in the art will understand that the technical solutions recorded in the foregoing examples can be modified, or some technical features can be replaced by equivalent features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A retro-synthetic approach to ring forming reactions characterized by, The method comprises: Collecting multi-source chemical reaction data, and obtaining an initial data set after data processing, wherein the data processing comprises chemical reasonableness checking; Based on the initial data set, a minimum ring system analysis is used to identify ring-forming reactions, and a special data set for cyclic skeleton synthesis is screened, wherein the analysis and identification comprises judging new ring formation by comparing the number of rings and ring characteristics of reactants and products; A reaction center offset maximum common substructure alignment algorithm is used to generate a plurality of aligned SMILES sequences for the reactions in the special data set, and a data enhancement base set is constructed; Based on the OpenNMT framework, an initial model is pre-trained using general chemical reaction data to obtain a pre-trained model, and a dynamic mask mechanism is used to enhance the learning of the pre-trained model on molecular topology association; The pre-trained model is loaded, fine-tuned using the data enhancement base set, a ring-forming reaction special model is constructed, and a weighted Top-k index is used to evaluate the prediction accuracy of the model; Based on the ring-forming reaction special model, an isomeric system call interface is constructed to enable bidirectional interaction between the interactive interface and the ring-forming reaction special model.

2. The retro-synthetic approach to ring forming reactions according to claim 1, characterized in that, The data processing comprises: Merging multi-source data and aligning cross-database fields through customized data mapping rules; Based on the rdkit toolkit, the molecules in the reaction sequence are converted into Canonical SMILES, and the isomerism of the molecular structure representation is eliminated based on the SMILES specification algorithm; Based on the Canonical SMILES, a hash comparison algorithm is used for deduplication, the hash value of the reaction SMILES sequence combination is calculated, and only unique reaction samples are retained.

3. The retro-synthetic approach to ring forming reactions according to claim 1, characterized in that, The chemical reasonableness checking comprises: Using the molecular structure analysis of the rdkit toolkit to screen multi-source chemical reaction data to identify and eliminate entries that cannot generate valid molecular objects and SMILES analysis failures, wherein the inability to generate valid molecular objects includes bond errors and abnormal atomic valence structures.

4. The retro synthetic approach to ring assembly reaction according to claim 1, wherein, The judgment of new ring formation by comparing the number of rings and ring characteristics of reactants and products comprises: Using the rdkit toolkit to extract the symmetric independent minimum ring set of reactants and products; Statistically analyzing the total number of rings of reactants and products, and determining that it is a ring-forming reaction if the total number of rings of the product is greater than that of the reactant; If the number of rings is the same, compare the ring characteristics, including ring size, atom type in the ring, and bond type in the ring, and determine that it is a ring-forming reaction if the product has new ring characteristics that the reactant does not contain.

5. The retro synthetic approach to ring assembly reaction according to claim 1, wherein, The reaction center offset maximum common substructure alignment algorithm is used to generate a plurality of aligned SMILES sequences for the reactions in the special data set, and a data enhancement base set is constructed, comprising: Determining the reaction center by analyzing the bond type change value and the atomic valence change value, wherein the bond type change value is the sum of the difference between the bond types of atoms in the reactants and products, and the atomic valence change value is the absolute difference of the atomic valence; Screening the maximum common substructure of the reactants and products, selecting the chain end atoms far from the reaction center in the maximum common substructure as the starting nodes of the molecular graph traversal, and generating the aligned reactant and product SMILES sequences; The special data set is subjected to multiple data augmentation, and the atoms corresponding to the second and third levels of the reaction center distance in the maximum common substructure are selected as the starting points of traversal to realize the alignment.

6. The retro synthetic ring assembly reaction method according to claim 1, wherein, The OpenNMT framework is used to pre-train an initial model based on general chemical reaction data to obtain a pre-trained model, and a dynamic mask mechanism is used to enhance the learning of the pre-trained model on the molecular topological correlation, including: The general chemical reaction data is divided into a training set, a validation set and a test set according to a target proportion; The reactant SMILES in the general chemical reaction data is randomly selected to mask a preset percentage of character positions; wherein, the mask positions of the target percentage complementary to the preset percentage are replaced with symbols, half of the preset percentage is replaced with random characters, and the other half of the preset percentage is kept as the original characters, to simulate noise interference in molecular structure analysis; The Transformer model is constructed based on the OpenNMT framework, the Adam optimizer and the Noam learning rate scheduling are used for training, the learning rate is dynamically adjusted based on the performance of the validation set, and the early stopping mechanism is triggered when the model performance does not significantly improve for a continuous preset number of cycles, to obtain the pre-trained model.

7. The retro synthetic approach to ring assembly reaction according to claim 1, wherein, The data augmentation base set is used for fine-tuning, including: The data augmentation base set is divided into a training set, a validation set and a test set according to a target proportion, and a stratified sampling strategy is used to ensure that the distribution deviation of different types of ring-forming reactions in the training set, the validation set and the test set is controlled within a preset deviation threshold; The encoder and decoder weight parameters of the pre-trained model are loaded, only the top network is initialized, and after training, the model parameters of the last preset number of checkpoints are fused by using a sliding window parameter averaging method.

8. A retrosynthetic device for ring forming reactions, characterized in that, The device comprises: A collection module is configured to collect multi-source chemical reaction data, and obtain an initial data set after data processing, wherein the data processing includes chemical reasonableness verification; A screening module is configured to identify ring-forming reactions by using a minimum ring system analysis based on the initial data set, and screen a special data set for cyclic skeleton synthesis, wherein the analysis includes judging new ring formation by comparing the number of rings and ring characteristics of reactants and products; A construction module is configured to generate a plurality of aligned SMILES sequences for reactions in the special data set by using a reaction center offset maximum common substructure alignment algorithm, and construct a data augmentation base set; A training module is configured to pre-train an initial model based on the OpenNMT framework using general chemical reaction data to obtain a pre-trained model, and enhance the learning of the pre-trained model on the molecular topological correlation by using a dynamic mask mechanism; An evaluation module is configured to load the pre-trained model, fine-tune the data augmentation base set, construct a ring-forming reaction special model, and evaluate the prediction accuracy of the model by using a weighted Top-k index; An interaction module is configured to construct a heterogeneous system call interface based on the ring-forming reaction special model, so that the data of the interaction interface and the ring-forming reaction special model can be interacted bidirectionally.

9. An electronic device, comprising: The electronic device comprises a memory and a processor, the memory stores program codes executable on the processor, and the program codes are executed by the processor to implement the inverse synthesis method of the ring-forming reaction according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores program codes, and the program codes can be called and executed by one or more processors to implement the inverse synthesis method of the ring-forming reaction according to any one of claims 1-7.