Deep chemical reaction prediction method and system based on domain knowledge editing

By introducing domain knowledge editing and chemical reaction constraints in chemical reaction prediction, using neural networks to generate chemical reaction prediction results, the problems of high dependence on experts and ineffectiveness of the generated results are solved, and effective prediction of specific chemical reactions is achieved.

CN119993301AActive Publication Date: 2025-05-13ZHEJIANG LAB
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510467532.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-13
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

The existing chemical reaction prediction methods are highly dependent on experts in the field, and the production results are highly ineffective, making it difficult to predict specific chemical reactions.

Method used

A chemical reaction prediction method based on domain knowledge editing is proposed. By adding chemical reaction constraints to the reaction data of the chemical reaction to be predicted, a molecular map is constructed, and a training neural network deep learning model is used to generate SMILES sequences that meet the constraints, and finally post-process and validate to obtain the prediction results.

Benefits of technology

It effectively reduces the hallucination problem of generative models, improves the accuracy of prediction of specific chemical reactions, and reduces the dependence on experts in the field of chemistry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993301A_ABST
    Figure CN119993301A_ABST
Patent Text Reader

Abstract

The invention provides a deep chemical reaction prediction method and system based on domain knowledge editing. The method comprises the following steps: adding a chemical reaction constraint to reaction data of a to-be-predicted chemical reaction as a dynamic constraint mark, wherein the reaction data of the to-be-predicted chemical reaction comprises SMILES character strings of reactants or products; constructing a molecular map of the reactant or the product; inputting the molecular map into a trained neural network deep learning model to obtain an SMILES sequence conforming to the chemical reaction constraint; and performing post-processing and verification on the SMILES sequence conforming to the chemical reaction constraint to obtain a prediction result of the chemical reaction to be predicted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of chemical reaction prediction, and in particular to a deep chemical reaction prediction method and system based on domain knowledge editing. Background Art

[0002] Chemical reaction prediction is a key research direction in the field of cheminformatics, which aims to predict the products of chemical reactions through computer algorithms or to predict the reactants based on the products. Traditional methods mainly rely on expert systems and rule-based reasoning to make predictions through predefined reaction templates and chemical transformation rules, but their generalization ability and ability to handle new reactions are limited. In recent years, with the development of machine learning technology, data-driven prediction methods have gradually become mainstream.

[0003] The methods used for molecular reaction prediction in the prior art can be roughly divided into three categories: template-based methods, semi-template methods, and template-free methods. The basic principle of the template-based method is to simulate the process of chemists selecting suitable reactions for target molecules based on known reaction types. This type of method constructs reaction templates by extracting core transformation rules from chemical reaction data sets. After establishing a reaction template library, the relevant algorithm matches the target molecule with these templates to generate corresponding prediction results. The semi-template-based method adopts a strategy that combines template-driven and deep learning methods. This method predicts the output results by considering the intermediates or synthetic intermediates produced in multiple steps. Typical technical solutions include MEGAN and Graph2Edits. The template-free method avoids the need to build an external template database. This method uses a deep learning model to directly encode the reaction mechanism, thereby realizing the prediction from reactants to products. In this technical field, this task is generally analogized to the problem of neural machine translation. To achieve this conversion, the prior art solutions usually use a Transformer model equipped with a graph structure or sequence encoder and decoder. Among them, the sequence encoder and decoder are used to process and generate SMILES (Simplified Molecular Linear Input Specification System) representations of chemical substances. Template-free methods are not only dependent on training datasets, but also lack constraint mechanisms in the model generation process, which is prone to hallucination problems caused by the underlying generation model, resulting in invalid or erroneous chemical reaction formulas. Summary of the invention

[0004] In order to solve the problems that current methods are highly dependent on domain experts, generate highly invalid results, and are difficult to predict specific chemical reactions, the present invention proposes a chemical reaction prediction method and system based on domain knowledge editing.

[0005] Specifically, the present application is implemented through the following technical solutions: On the one hand, the present application provides a chemical reaction prediction method based on domain knowledge editing, comprising: Adding chemical reaction constraints as dynamic constraint tags to reaction data of a chemical reaction to be predicted, wherein the reaction data of the chemical reaction to be predicted includes a SMILES string of a reactant or a product; and constructing a molecular graph of the reactant or the product; Inputting the molecular graph into a trained neural network deep learning model to obtain a SMILES sequence that meets the chemical reaction constraints; The SMILES sequence that meets the chemical reaction constraints is post-processed and verified to obtain a prediction result of the chemical reaction to be predicted.

[0006] In one embodiment, inputting the molecular graph into a trained neural network deep learning model to obtain a SMILES sequence that meets the chemical reaction constraints comprises: Performing graphic encoding on the molecular graph and performing sequence decoding on the encoded molecular graph features; The sequence decoding of the encoded molecular graph features includes: using a beam search strategy to calculate the cumulative probability of all candidate paths, retaining the first n optimal paths, where n is the beam width; and dynamically filtering the candidate paths in combination with chemical reaction constraints to exclude tokens that do not meet the chemical reaction constraints.

[0007] In one embodiment, inputting the molecular graph into a trained neural network deep learning model to obtain a SMILES sequence that meets the chemical reaction constraints comprises: Performing graphic encoding on the molecular graph and performing sequence decoding on the encoded molecular graph features; The sequential decoding of the encoded molecular graph features includes: generating tokens by using a Top-k sampling strategy, and limiting the Top-k candidate range by a mask mechanism based on chemical reaction constraints.

[0008] In one embodiment, inputting the molecular graph into a trained neural network deep learning model to obtain a SMILES sequence that meets the chemical reaction constraints comprises: Performing graphic encoding on the molecular graph and performing sequence decoding on the encoded molecular graph features; The sequential decoding of the encoded molecular graph features includes: generating tokens by using a Top-p sampling strategy, and adding chemical descriptors based on chemical reaction constraints to the Top-p candidate set.

[0009] In one embodiment, graph encoding the molecular graph includes: utilizing a multi-head attention mechanism and position encoding molecular topology information.

[0010] In one embodiment, post-processing and verifying the SMILES sequence that meets the chemical reaction constraint to obtain the prediction result of the chemical reaction to be predicted includes: Through any one or more of chemical reaction conservation verification, chemical rationality verification and atom mapping backtracking, SMILES sequences that do not conform to the essence of chemical reactions are eliminated to obtain the prediction results of the chemical reaction to be predicted.

[0011] In one embodiment, after acquiring the data of the chemical reaction to be predicted and before adding chemical reaction constraints to the reaction data of the chemical reaction to be predicted as dynamic constraint tags, the method further includes expanding the data volume of the chemical reaction data to be predicted by data enhancement.

[0012] On the other hand, the present application provides a chemical reaction prediction system based on domain knowledge editing, comprising: A data processing module, used for adding chemical reaction constraints as dynamic constraint tags to reaction data of a chemical reaction to be predicted, wherein the reaction data of the chemical reaction to be predicted includes a SMILES string of a reactant or a product; and constructing a molecular graph of the reactant or the product; A chemical reaction prediction module, used for inputting the molecular graph into a trained neural network deep learning model to obtain a SMILES sequence that meets the chemical reaction constraints; The post-processing module is used to post-process and verify the SMILES sequence that meets the chemical reaction constraints to obtain the prediction result of the chemical reaction to be predicted.

[0013] On the other hand, the present application provides a computer-readable storage medium, wherein the storage medium stores a computer program, wherein the computer program implements the above method when executed by a processor.

[0014] On the other hand, the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the above method is implemented when the processor executes the program.

[0015] The technical solution provided by this application can achieve the following beneficial effects: Although the chemical reaction prediction method based on domain knowledge editing provided by this application has a certain dependence on training data, by introducing chemical reaction constraints in the stages of training data construction, decoder reasoning and post-generation processing, it does not require in-depth participation of chemical experts and can effectively reduce the hallucination problem of the generated model. In particular, the method of the present invention is based on the vectorized reaction mechanism obtained by the model, combined with domain knowledge, and can effectively predict specific chemical reactions and improve the effectiveness of reaction types covered by the model prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 A schematic diagram of a process flow of a deep chemical reaction prediction method based on domain knowledge editing provided in this specification is shown; Figure 2 A schematic diagram of a deep chemical reaction prediction system based on domain knowledge editing provided in this specification is shown; Figure 3 This specification provides a method corresponding to Figure 1 Schematic diagram of electronic equipment. DETAILED DESCRIPTION

[0017] In order to make the purpose, technical solutions and advantages of this specification more clear, the technical solutions of this specification will be clearly and completely described below in combination with the specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this specification.

[0018] The technical solutions provided by the embodiments of this specification are described in detail below in conjunction with the accompanying drawings.

[0019] Chemical reaction prediction includes forward prediction of products from reactants and reverse prediction of possible reaction paths and reactants from target products. It can be applied to the fields of new drug molecule design, material synthesis path planning, etc. Figure 1 The flowchart of the deep chemical reaction prediction method based on domain knowledge editing provided in this specification includes the following steps S101-S103: S101: adding chemical reaction constraints as dynamic constraint tags to reaction data of a chemical reaction to be predicted, wherein the reaction data of the chemical reaction to be predicted includes a SMILES string of reactants or products; and constructing a molecular graph of the reactants or products.

[0020] The reaction data of the chemical reaction to be predicted can be extracted from chemical databases (such as USPTO, Reaxys), including SMILES strings of reactants, products, reaction conditions (temperature, solvent, etc.). SMILES (Simplified Molecular Linear Input System) is a standard method for representing chemical structures using ASCII strings, which is used to convert three-dimensional molecular structure information into a one-dimensional string format. The reaction data of the chemical reaction to be predicted can be cleaned, such as using RDKit to verify the legality of SMILES and filter out molecules that cannot be parsed or have abnormal valence states.

[0021] In this application, chemical reaction constraints refer to physical, chemical or engineering restrictions that must be followed during the reaction process, covering multiple dimensions such as thermodynamics, kinetics, control strategies and calculation methods. More specifically, chemical reaction constraints can be chemical reaction conservation constraints, which refer to a set of constraints constructed based on basic laws such as atomic number conservation, charge conservation, and mass conservation in chemical reactions. Reaction conservation knowledge can be used to verify the rationality of prediction results and improve the accuracy of model predictions.

[0022] In one embodiment provided in this specification, the reaction conservation knowledge is the charge quantity and the atomic quantity of the reactants or products, for example, 0&C~3N~1H~1, "&" separates the two parts of information, charge and atomic quantity, and "~" separates each element and quantity, and each atom and quantity. In the current embodiment, "0&C~3N~1H~1" represents 0 charge, 3 C atoms, 1 N atom, and 1 H atom.

[0023] In chemical reaction prediction, when processing data related to chemical reactions, the encoder can add the charge number and the number of atoms of the reactants or products as dynamic prefixes. For example, for each input reactant or product, the encoder can count the charge number and the number of atoms and add this information as a prefix to the input data. In this way, the encoder can take into account the knowledge of reaction conservation when processing the data, which may improve the accuracy of data processing and generate output that is more consistent with the knowledge of reaction conservation. The decoder can also take the knowledge of reaction conservation into consideration when generating output.

[0024] In an embodiment provided in the present specification, before adding chemical reaction constraints to the reaction data of the chemical reaction to be predicted as dynamic constraint tags, the method may further include expanding the data volume of the chemical reaction data to be predicted by data enhancement.

[0025] Expanding the amount of data through data augmentation is a common method in the field of deep learning. Since it starts from a certain starting atom (any atom), the entire molecular graph is traversed in a depth-first manner, and the corresponding characters are recorded according to the atoms and bond types encountered to construct the SMILES of the molecule, so that when the starting atom is selected differently, the same molecule can have multiple SMILES writing methods. The multiple writing methods of SMILES make data augmentation possible. The amount of data can be expanded by hydrogenation in SMILES. The following is a simplified example showing that [H]C#C[N+]#[C-]>>[H]C#[C].[C]#N is obtained by hydrogenation of [C-]#[N+]C#C>>[C]#C.[C]#N. Generating equivalent or deformed SMILES can also expand the amount of data. As a simplified example, [H]C#C[N+]#[C-]>>[H]C#[C].[C]#N is expressed as C([H])#C[N+]#[C-]>>[C]#C[H].[C]#N; [H]C#C[N+]#[C-]>>C(#[C])[H].N#[C]; [C-]#[N+]C#C[H]>>C(#[C])[H].[C]#N, etc. By increasing the amount of data through data augmentation, the generalization ability of the model can be improved.

[0026] S102: Inputting the molecular graph into a trained neural network deep learning model to obtain a SMILES sequence that meets the chemical reaction constraints.

[0027] In one embodiment provided in the present specification, the molecular graph is input into a trained neural network deep learning model to obtain a SMILES sequence that meets the chemical reaction constraints, including: graphically encoding the molecular graph and sequentially decoding the encoded molecular graph features.

[0028] In one embodiment provided in this specification, the molecular graph is graphically encoded, including: using a multi-head attention mechanism and position encoding molecular topological information. For example, using a graph attention network (GAT) or a graph transformer, a multi-head attention mechanism is used to capture key interactions between atoms (such as nucleophilic / electrophilic sites); and random walks or Laplace position encoding are used to retain the topological relationship of the three-dimensional conformation of the molecule.

[0029] The multi-head attention mechanism allows the model to focus on different parts of the input data at the same time, including local attention and global attention. Local attention focuses on local molecular structural features, such as specific atomic groups or functional groups. By limiting the scope of attention, local features in the molecule can be understood more finely. Global attention integrates the overall information of the entire molecule and captures global features at the molecular level. This helps the model understand the overall configuration and properties of the molecule. The multi-head attention mechanism improves the model's ability to capture complex relationships by processing multiple attention heads in parallel, with each head focusing on different aspects of the input data.

[0030] Position encoding is used to provide the model with information about the position of atoms in the graph. Position encoding includes centrality encoding, spatial encoding, and edge encoding. Centrality encoding reflects the relative importance or centrality of atoms in a molecule. Spatial encoding provides information about the relative position of atoms in molecular space. Edge encoding captures the type and nature of chemical bonds between atoms. Combining these three encoding modes can obtain key information of 2D molecular graphs, including the type, position, connection mode of atoms, and the configuration and properties of the entire molecule. This information is crucial for subsequent tasks such as chemical reaction prediction and molecular property prediction.

[0031] In an embodiment provided in the present specification, sequential decoding of the encoded molecular graph features includes: using a beam search strategy to calculate the cumulative probability of all candidate paths, retaining the top n optimal paths, where n is the beam width; and dynamically filtering the candidate paths in combination with chemical reaction constraints to exclude tokens that do not meet the chemical reaction constraints.

[0032] During the decoding process, beam search maintains a fixed-size set of candidate paths (i.e., beam width, such as beam_size=2), and only retains the top beam_size candidate sequences with the highest cumulative probability at each generation step. In this way, the local optimal trap of greedy search is avoided, and the high computational cost of exhaustive search is reduced. Each time a new token is generated, the cumulative probability of all possible paths is calculated (i.e., the product of the probabilities of each step), and only some of the paths with the highest probability are retained, and the remaining paths are pruned and discarded. The beam search strategy is used in conjunction with chemical reaction constraints (such as atomic conservation) to dynamically filter candidate paths, for example, excluding tokens that do not meet the valence legitimacy (such as pentavalent carbon). Through the above mechanism, beam search significantly improves the rationality and diversity of the generated results while maintaining high efficiency. ‌‌‌

[0033] In an embodiment provided in this specification, sequential decoding of the encoded molecular graph features includes: generating tokens using a Top-k sampling strategy, and limiting the Top-k candidate range through a mask mechanism based on chemical reaction constraints.

[0034] The Top-k sampling strategy only retains the top k highest probability tokens (such as k=50) in the model prediction probability distribution when generating tokens at each step, and the remaining tokens are directly filtered to form a candidate pool. The selected k tokens are probabilistically renormalized (i.e., their sum of probabilities is recalculated to 1) to ensure that the sampling probability distribution is legal. While utilizing the Top-k sampling strategy, the Top-k candidate range is limited through a mask mechanism based on chemical reaction constraints. For example, when generating ring structures, only bond types that meet ring strain (such as conjugated double bonds in aromatic rings) are allowed to avoid abnormal structures. Through the above mechanism, a controllable balance between quality and diversity can be achieved during the generation process.

[0035] In an embodiment provided in this specification, sequential decoding of the encoded molecular graph features includes: generating tokens using a Top-p sampling strategy, and adding chemical descriptors based on chemical reaction constraints to the Top-p candidate set.

[0036] The Top-p sampling strategy samples from the minimum token set whose cumulative probability exceeds a threshold p (such as p=0.95), and dynamically adjusts the number of candidates to adapt to the uncertainty of different generation stages. It is suitable for dealing with scenarios with ambiguous reaction paths, such as adaptively selecting reasonable atomic connections when predicting intermediates in catalytic reactions. By using the Top-p sampling strategy and adding chemical descriptors based on chemical reaction constraints (such as reaction activation energy) to the Top-p candidate set, thermodynamically stable structures can be given priority. Through the above mechanism, low-probability noise interference (such as illegal chemical structures) is avoided while retaining reasonable diversity (such as different reaction mechanism paths).

[0037] In an embodiment provided in this specification, the above three methods: beam search strategy combined with chemical reaction constraints, Top-k sampling strategy combined with chemical reaction constraints, and Top-p sampling strategy combined with chemical reaction constraints can work together. Beam search provides global optimal path guidance, Top-k and Top-p enhance local diversity, and the combination of the three methods improves the coverage and accuracy of the generated results.

[0038] In one embodiment provided in the present specification, training a neural network deep learning model includes generating a molecular graph for all reactants in a training set and a validation set, wherein the molecular graph generated for all products includes the structure, node features, and edge features of the graph; inputting the structure, node features, and edge features of the graph into the neural network deep learning model to obtain a SMILES sequence that conforms to the chemical reaction constraints; optimizing the parameters of the neural network deep learning model using the training set as input, performing a preliminary evaluation of the capabilities of the neural network deep learning model using the validation set as input, and saving the neural network deep learning model with the smallest validation set loss (cross entropy loss) during the training process.

[0039] In an embodiment provided in this specification, the encoded molecular graph features are sequence decoded to obtain a token_ids sequence, for example, [25, 7, 61, 24, 0], and the corresponding tokens are 'O', '.', '[CH2+]', 'N', '[EOS]'.

[0040] Step S103: post-processing and verifying the SMILES sequence that meets the chemical reaction constraints to obtain a prediction result of the chemical reaction to be predicted.

[0041] In an embodiment provided in this specification, post-processing and verifying the SMILES sequence that meets the chemical reaction constraint to obtain the prediction result of the chemical reaction to be predicted includes: Through any one or more of chemical reaction conservation verification, chemical rationality verification and atom mapping backtracking, SMILES sequences that do not conform to the essence of chemical reactions are eliminated to obtain the prediction results of the chemical reaction to be predicted.

[0042] Chemical reaction conservation verification includes counting the types and quantities of atoms in reactants and products (such as C, O, H, etc.) to ensure the total number is consistent, and verifying whether the total ion charge before and after the reaction is equal to avoid illegal reactions that do not conserve charge (such as H in an acidic solution). + With OH - Chemical rationality verification includes checking whether the valence of atoms in the product conforms to chemical laws (such as carbon atoms not exceeding four valences, oxygen atoms not having positive valences, etc.), and judging whether the SMILES sequence conforms to the reaction type characteristics based on known reaction mechanisms (such as nucleophilic substitution, redox). Atom mapping backtracking includes establishing a mapping relationship between reactant and product atoms to ensure that each product atom has a clear source, and determining sequences that cannot establish a complete mapping (such as oxygen atoms that appear out of thin air) as illegal generation results. The combination of the three methods can improve the accuracy of the generation results.

[0043] In an embodiment provided in this specification, post-processing and verifying the SMILES sequence that meets the chemical reaction constraint to obtain the prediction result of the chemical reaction to be predicted includes: Use the pre-trained ChemBERTa to encode the SMILES of reactants and products as vectors; Using contrastive learning to fine-tune the pre-trained ChemBERTa to shorten the semantic distance of positive sample pairs and lengthen the semantic distance of negative sample pairs; the positive sample pairs are verified reactant-product pairs, and the negative sample pairs are combinations of reactants and products that do not match the reaction mechanism; Use beam search to generate Top-K products; The model generation probability, comparison similarity, atomic conservation, valence state legitimacy, and mechanism matching are comprehensively considered, and weighted sorting is performed to select the optimal result.

[0044] By bringing positive sample pairs closer together and negative sample pairs further apart, the model can more accurately capture the essential relationship between reactants and products. Supervised comparative learning enables the deep transfer of chemical knowledge and significantly improves the practicality and reliability of the generated results.

[0045] The deep chemical reaction prediction method based on domain knowledge editing in this application significantly reduces the hallucination problem of the generated model by imposing chemical reaction conservation constraints, effectively avoiding the generation of invalid and erroneous chemical reaction formulas. At the same time, the prediction of specific chemical reactions has good robustness and can maintain a high prediction accuracy even when the training data is sparse.

[0046] The above is the deep chemical reaction prediction method based on domain knowledge editing in this specification. Based on the same idea, this specification also provides a corresponding deep chemical reaction prediction system based on domain knowledge editing, such as Figure 2 As shown, including: A data processing module 201 is used to add chemical reaction constraints as dynamic constraint tags to reaction data of a chemical reaction to be predicted, wherein the reaction data of the chemical reaction to be predicted includes a SMILES string of reactants or products; and to construct a molecular graph of the reactants or products; A chemical reaction prediction module 202, for inputting the molecular graph into a trained neural network deep learning model to obtain a SMILES sequence that meets the chemical reaction constraints; The post-processing module 203 is used to post-process and verify the SMILES sequence that meets the chemical reaction constraints to obtain the prediction result of the chemical reaction to be predicted.

[0047] This specification also provides a computer-readable storage medium, which stores a computer program, which can be used to execute the above Figure 1A deep chemical reaction prediction method based on domain knowledge editing is provided.

[0048] This manual also provides Figure 3 The one shown corresponds to Figure 1 A schematic diagram of the electronic device. Figure 3 As mentioned above, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 The deep chemical reaction prediction method based on domain knowledge editing. Of course, in addition to the software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0049] For the improvement of a technology, it can be clearly distinguished whether it is a hardware improvement (for example, improvement of the circuit structure of diodes, transistors, switches, etc.) or a software improvement (improvement of the method flow). However, with the development of technology, many improvements of the method flow today can be regarded as direct improvements of the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that the improvement of a method flow cannot be implemented with a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming themselves, without having to ask chip manufacturers to design and make dedicated integrated circuit chips. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages ​​and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.

[0050] The controller may be implemented in any suitable manner, for example, the controller may take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (e.g., software or firmware) executable by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller, examples of which include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320, and the memory controller may also be implemented as part of the control logic of the memory. It is also known to those skilled in the art that, in addition to implementing the controller in a purely computer-readable program code manner, the controller may be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, such a controller may be considered as a hardware component, and the devices for implementing various functions included therein may also be considered as structures within the hardware component. Or even, the devices for implementing various functions may be considered as both software modules for implementing the method and structures within the hardware component.

[0051] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0052] For the convenience of description, the above device is described in various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0053] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0054] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0055] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0056] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0057] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0058] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0059] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0060] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0061] It should be understood by those skilled in the art that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0062] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0063] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0064] The above description is only an embodiment of the present specification and is not intended to limit the present specification. For those skilled in the art, the present specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present specification shall be included in the scope of the claims of the present specification.

Claims

1. A chemical reaction prediction method, characterized in that: The method comprises: Adding chemical reaction constraints as dynamic constraint tags to reaction data of a chemical reaction to be predicted, wherein the reaction data of the chemical reaction to be predicted includes a SMILES string of a reactant or a product; and constructing a molecular graph of the reactant or the product; Inputting the molecular graph into a trained neural network deep learning model to obtain a SMILES sequence that meets the chemical reaction constraints; The SMILES sequence that meets the chemical reaction constraints is post-processed and verified to obtain a prediction result of the chemical reaction to be predicted.

2. The chemical reaction prediction method according to claim 1, characterized in that: The step of inputting the molecular graph into a trained neural network deep learning model to obtain a SMILES sequence that meets the chemical reaction constraints comprises: Performing graphic encoding on the molecular graph and performing sequence decoding on the encoded molecular graph features; The sequence decoding of the encoded molecular graph features includes: using a beam search strategy to calculate the cumulative probability of all candidate paths, retaining the first n optimal paths, where n is the beam width; and dynamically filtering the candidate paths in combination with chemical reaction constraints to exclude tokens that do not meet the chemical reaction constraints.

3. The chemical reaction prediction method according to claim 1 or 2, characterized in that: The step of inputting the molecular graph into a trained neural network deep learning model to obtain a SMILES sequence that meets the chemical reaction constraints comprises: Performing graphic encoding on the molecular graph and performing sequence decoding on the encoded molecular graph features; The sequential decoding of the encoded molecular graph features includes: generating tokens by using a Top-k sampling strategy, and limiting the Top-k candidate range by a mask mechanism based on chemical reaction constraints.

4. The chemical reaction prediction method according to claim 3, characterized in that: The step of inputting the molecular graph into a trained neural network deep learning model to obtain a SMILES sequence that meets the chemical reaction constraints comprises: The sequential decoding of the encoded molecular graph features includes: generating tokens by using a Top-p sampling strategy, and adding chemical descriptors based on chemical reaction constraints to the Top-p candidate set.

5. The chemical reaction prediction method according to claim 4, characterized in that: The graphical encoding of the molecular graph includes: utilizing a multi-head attention mechanism and position encoding of molecular topological information.

6. The chemical reaction prediction method according to claim 4, characterized in that: The post-processing and verifying the SMILES sequence that meets the chemical reaction constraint to obtain the prediction result of the chemical reaction to be predicted includes: Through any one or more of chemical reaction conservation verification, chemical rationality verification and atom mapping backtracking, SMILES sequences that do not conform to the essence of chemical reactions are eliminated to obtain the prediction results of the chemical reaction to be predicted.

7. The chemical reaction prediction method according to claim 6, characterized in that: Before adding chemical reaction constraints to the reaction data of the chemical reaction to be predicted as dynamic constraint tags, the method further includes expanding the data volume of the chemical reaction data to be predicted by data enhancement.

8. A chemical reaction prediction system, characterized in that: The system comprises: A data processing module, adding chemical reaction constraints as dynamic constraint tags to reaction data of a chemical reaction to be predicted, wherein the reaction data of the chemical reaction to be predicted includes a SMILES string of a reactant or a product; and constructing a molecular graph of the reactant or the product; A chemical reaction prediction module, used for inputting the molecular graph into a trained neural network deep learning model to obtain a SMILES sequence that meets the chemical reaction constraints; The post-processing module is used to post-process and verify the SMILES sequence that meets the chemical reaction constraints to obtain the prediction result of the chemical reaction to be predicted.

9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of claims 1 to 7 is implemented.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Systems and method for designing organic synthesis pathways for desired organic molecules

    CA3153469A1

  • Systems and methods for designing organic synthesis pathways for desired organic molecules

    CN114730618A

  • Modeling method and device for synthesis mechanism of 2-mercaptobenzothiazole and application

    CN115206445A

  • Generating organic synthesis procedures from simplified molecular input linear entry system reactions

    CN116075899A

  • Drug molecule optimization method and device based on pre-training fine tuning

    CN117116383A