Chemical reaction graph coding software, corresponding methods, and related data applications

The CRS format addresses inefficiencies in existing chemical reaction encoding by compactly representing bond changes, improving machine learning applications with enhanced speed and accuracy for multi-step and equilibrium reactions.

JP7846081B2Active Publication Date: 2026-04-14FIRMENICH SA
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
FIRMENICH SA
Filing Date
2021-10-26
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing chemical reaction encoding formats, such as SMARTS and SMIRKS, require excessive memory, are inefficient for machine learning, and fail to compactly represent multi-step reactions, equilibrium reactions, reaction mechanisms, stereochemistry, and biochemical pathways, leading to reduced signal-to-noise ratio and limited machine learning performance.

Method used

A novel chemical reaction graph compression format (CRS) encodes changing bonds using a set of bijective characters, focusing on bond changes to reduce memory usage and enhance machine learning efficiency by targeting relevant reaction parts.

Benefits of technology

The CRS format allows for compact encoding of multi-step and equilibrium reactions, supports reaction reversibility, and improves machine learning applications by enhancing speed and accuracy through focused encoding on relevant reaction details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007846081000016
    Figure 0007846081000016
  • Figure 0007846081000017
    Figure 0007846081000017
  • Figure 0007846081000018
    Figure 0007846081000018
Patent Text Reader

Abstract

A chemical reaction encoding method (100) for single-step, multi-step, and equilibrium reactions includes a step (105) of receiving, on a computer interface, a chemical reaction graph including at least one chemical reaction reagent and at least one chemical reaction product; a first step (110) of encoding, by a computer device, the chemical reaction graph describing the structure of the at least one reagent and the product; a step (115) of determining, by a computer device, variable bonds in the encoding representing the chemical structure of the at least one reagent and the product; and, for the determined at least one variable bond, at least one character representing an atom undergoing the bond change; The method includes a second step (120) of encoding, by a computer device, at least one character representing the determined type of variable bond and at least one character representing an atom resulting from the bond change into a single string, wherein the variable bond is encoded by a set of two characters representing the determined variable bond, a first character representing a reagent bond and a second character representing a product bond, each character being selected from a library of bijective characters, one character representing a variable bond type; and a step (125) of providing, on a computer interface, a string of characters corresponding to the encoding of the variable bonds of the chemical reaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to chemical reaction graph compression software, corresponding methods, chemical reaction graph formats, methods for expanding chemical reaction datasets, methods for preprocessing chemical reaction datasets, training methods for classifiers, transformers or regressors, methods for predicting chemical reaction bond progress, chemical reaction generation methods, computer-implemented classifiers, transformers or regressors, and related computer programs. The present invention is particularly applicable to the field of organic chemistry including, but not limited to, pharmaceuticals, fragrances, flavorings, cleaning products, fragrance design and olfactory testing, air fresheners, fine fragrance ingredients and scent design.

[0002] Background of the Invention One important encoding system in the field of digital modeling of chemical species and chemical reactions is the line notation, for example, the Simplified molecular-input line-entry system (SMILES) format. Such formats are well documented in many sources, including, for example, the collaborative encyclopedia Wikipedia.

[0003] Such formats, such as SMARTS and SMIRKS, have served the understanding and ability to model chemical interactions, but drawbacks are beginning to appear.

[0004] - To encode chemical reactions, an excessive number of characters are required, occupying the corresponding physical memory space. This means longer transmission and processing times for systems using such formats.

[0005] - The performance of such formats degrades in machine learning applications due to the excessive amount of potentially irrelevant information stored within the format.

[0006] - In older formats, such as the SMARTS or SMIRKS string, this string consists of reagents separated by dots, activators (those enabling the reaction, reaction conditions) separated by dots, and products separated by dots, requiring explicit atomic mapping to define the reaction. Our novel short CRS format requires a large amount of information.

[0007] - Reversibility cannot be coded simply and compactly.

[0008] - Multi-step reactions, i.e., A>B>C, cannot be coded simply and compactly.

[0009] - Equilibrium reactions, i.e., A<>B or A>B>A, cannot be coded simply and compactly.

[0010] - The reaction mechanism, i.e., A>T>B (where T defines the transition state of the reaction), cannot be coded simply and compactly.

[0011] - Reaction classification and data cleaning are impossible, which reduces the signal-to-noise ratio when the data is used.

[0012] - The characters used are ambiguous, which reduces the signal-to-noise ratio when the data is used.

[0013] - Biochemical pathways composed of multiple intermediates cannot be encoded simply and compactly.

[0014] - Stereochemistry cannot be encoded simply and compactly.

[0015] - It is not possible to display the stereoisomerism change around the tetravalent chiral center.

[0016] Furthermore, modern chemical reaction research and development cycles require more sophisticated tools than typical trial-and-error approaches or other approaches that rely solely on existing knowledge within the organization. In this context, machine learning is considered a foundation for optimizing this research and development cycle. However, the performance of machine learning models is limited by the quality of the input data. Currently, there is no satisfactory method for generating machine learning models to predict chemical reaction behavior or to autonomously generate new chemical reactions.

[0017] Summary of the Invention This invention aims to improve all or some of these shortcomings.

[0018] For this purpose, according to a first aspect, the present invention is chemical reaction graph compression software for one-step reactions, multi-step reactions and equilibrium reactions, - A step of receiving a chemical reaction graph containing at least one chemical reaction reagent and at least one chemical reaction product on a computer interface, - A first step of encoding the chemical reaction graph describing the structure of at least one reagent and the product using a computer device, - A step of determining, using a computer device, the changing bonds within the encoding that represent the chemical structure of at least one reaction reagent and the product, - A second step of encoding, using a computer device, for at least one determined changing bond, at least one character representing the atom undergoing the bond change, at least one character representing the type of determined changing bond, and at least one character representing the atom resulting from the bond change, wherein the changing bond is encoded by a set of two characters representing the determined changing bond, the first character representing the reagent bond, and the second character representing the product bond, and each character is selected from a library of bijective characters where one character represents one type of bond. - A step of providing the string corresponding to the encoding of changing bonds in a chemical reaction on a computer interface. This is intended for software that executes the corresponding instructions.

[0019] This type of provision focuses on the location of the reagent where the binding change occurs, thus enabling highly executable encoding by limiting the substances to be encoded. The resulting code is more compact and limits physical memory usage. Furthermore, by focusing on the binding change, machine learning applications can target only the relevant parts of the chemical reaction, thus enabling improvements in speed and accuracy.

[0020] In addition, this formatting makes it possible to modularize multi-step reactions or chemical equilibrium reactions, i.e., A<>B as a pseudo-two-step reaction by describing individual reactions A>B and B>A, or multi-step reactions A>B>A.

[0021] Furthermore, the resulting format is reversible, enabling the definition of equilibrium reactions, the encoding of reaction mechanisms, unambiguity, reaction classification and data cleaning, the encoding of stereochemical changes, and the indication of changes to tetravalent chiral centers.

[0022] In a particular embodiment, the second encoding step is configured to embed two characters representing the determined changing combination between two neutral tag characters representing the presence of the encoding of the changing combination.

[0023] In certain embodiments, a multi-step reaction represented by a series of bonding changes between two atoms is encoded by a series of single letters, each single letter representing a sequential state of bonding between the two atoms, and the order of the letters represents the sequence of bonding changes between the two atoms.

[0024] In such an embodiment, it is possible to automatically recognize by means of software elements that two characters representing a changing bond are separated from those representing the atoms themselves.

[0025] According to a second aspect, the present invention is a method for compressing chemical reaction graphs for single-step reactions, multi-step reactions and equilibrium reactions, comprising: - receiving on a computer interface a chemical reaction graph comprising at least one chemical reaction reagent and at least one chemical reaction product; - a first step of encoding by a computer device the chemical reaction graph describing the structures of the at least one reagent and the product; - a step of determining by a computer device changing bonds within the encoding representing the chemical structures of the at least one reaction reagent and the product; - for at least one determined changing bond, a second step of encoding by a computer device into a single character string at least one character representing an atom undergoing a bond change, at least one character representing the type of the determined changing bond, and at least one character representing an atom resulting from the bond change, the changing bond being encoded by a set of two characters representing the determined changing bond, the first character representing a bond in the reagent and the second character representing a bond in the product, each character being selected from within a library of bijective characters in which one character represents one change of bond type; - providing on a computer interface a character string corresponding to the encoding of the changing bonds of the chemical reaction; and is aimed at a method.

[0026] The benefits and advantages of this method correspond to those of the software which is the object of the first aspect of the present invention.

[0027] In certain embodiments, a multi-step reaction represented by a series of bonding changes between two atoms is encoded by a series of single letters, each single letter representing a sequential state of bonding between the two atoms, and the order of the letters represents the sequence of bonding changes between the two atoms.

[0028] In certain embodiments, the first encoding step is configured to encode a chemical reaction graph into line notation, and the method further includes a step of extending the line notation encoding before the second encoding step.

[0029] Such embodiments allow for starting with a single chemical reaction graph and gradually increasing the sample size. This is particularly useful in machine learning applications.

[0030] In certain embodiments, the second encoding step includes extracting a reagent-product binding table from computer memory using a computer device, and the encoding is performed as a function of the binding table.

[0031] In a particular embodiment, the second encoding step includes removing at least one atomic identifier from at least one reagent and / or product from the first encoding obtained from the first encoding step, wherein each atom is removed as a result of the determination step if the atom and associated bond are located in the product and / or reagent that remain unchanged from the reagent reaction step to the product step of the chemical reaction.

[0032] In such embodiments, restricting the notation of the reaction to the reaction sites allows for greater compression of the chemical reaction format.

[0033] In certain embodiments, the method that is the object of the present invention includes the step of obtaining the encoded chemical reaction product by carrying out the chemical reaction within a physical device.

[0034] According to a third aspect, the present invention aims at an encoded chemical reaction comprising a string of characters obtained by the method which is the object of the second aspect of the present invention.

[0035] The benefits and advantages of this formatted chemical reaction graph correspond to the benefits of the method which is the object of a second aspect of the present invention.

[0036] According to a fourth aspect, the present invention is - The process of receiving a string on a computer interface according to the encoding which is the objective of the third aspect of the present invention, - A step of rearranging a string by a computer system in order to shift at least one string of characters, each string consisting of at least one character representing an atom and at least one character representing a bond change associated with the corresponding atom. - The process of outputting an extended string on a computer interface that corresponds to the response initially encoded from the received string. The aim is to provide a method for extending chemical reaction datasets, including [specific data points].

[0037] This type of provision allows you to start with a single chemical reaction graph and gradually increase the sample size. This is particularly useful in machine learning applications.

[0038] In certain embodiments, the method which is the object of the present invention includes the step of associating at least two strings by a computer system according to the format which is the object of a third aspect of the present invention, each of which strings represents the same chemical reaction graph.

[0039] This type of provision makes it possible to create multidimensional inputs that are particularly useful in machine learning applications.

[0040] According to a fifth aspect, the present invention is - A step of receiving a dataset of at least two chemical reaction graphs, each containing at least one chemical reaction reagent and at least one chemical reaction product, on a computer interface, - A step of compressing at least two chemical reaction graphs according to a method which is the object of a second aspect of the present invention, - A process of determining the distribution of chemical reaction classes within an encoded dataset using a computer system, - A step of extending the dataset for at least one chemical reaction class as a function of the determined distribution, according to a method which is the object of a fourth aspect of the present invention, - The process of outputting the preprocessed dataset on a computer interface. The aim is to develop a method for preprocessing chemical reaction datasets, including [specific data points].

[0041] This type of provision enables dynamic and intelligent scaling of datasets to optimize machine learning applications.

[0042] According to a sixth aspect, the present invention is - An object of the third aspect of the present invention is to input a dataset of chemical reaction graphs encoded in compressed encoding onto a computer interface, - To operate a computer system with a recurrent neural network architecture configured to classify the progression of chemical reaction bonds as a function of the input, using a dataset of chemical reaction graphs as input. - To output a trained classifier, transformer, or regressor on a computer interface. This is intended for training methods for classifiers, transformers, or regressors, including those mentioned above.

[0043] This type of provision enables the optimal creation of trained classifiers, transformers, or regressors, as the chemical graph reaction format used significantly improves the quality of the generated models.

[0044] According to the seventh aspect, the present invention aims to provide a method for predicting the progression of chemical reaction bonds, which operates a classifier, transformer, or regressor obtained by the method that is the objective of the sixth aspect of the present invention.

[0045] This type of provision makes it possible to accurately predict the bonding progress of arbitrarily entered chemical reagents.

[0046] According to the eighth aspect, the present invention aims to provide a method for generating a chemical reaction that operates a classifier, transformer, or regressor obtained by the method that is the objective of the sixth aspect of the present invention.

[0047] This type of provision allows for the autonomous generation of chemical reactions using corresponding graphs and / or row notations.

[0048] According to the ninth aspect, the present invention aims to provide a computer-implemented classifier, transformer, or regressor, wherein the classifier, transformer, or regressor is obtained by the method described in the sixth aspect of the present invention.

[0049] The benefits and advantages of this computer-implemented classifier, transformer, or regressor correspond to the benefits of the method which is the object of the sixth aspect of the present invention.

[0050] According to the tenth aspect, the present invention relates to a computer program that includes instructions for operating a method which is an object of any one of the sixth, seventh, or eighth aspects of the present invention.

[0051] The benefits and advantages of this computer program correspond to the benefits of the methods which are the objectives of the corresponding sixth, seventh, or eighth aspect of the present invention.

[0052] Other advantages, purposes, and specific features of the present invention will become apparent from the following non-extensive description of at least one specific embodiment of the invention in relation to the accompanying drawings. [Brief explanation of the drawing]

[0053] [Figure 1] This figure schematically represents a first specific sequence of steps that constitute the method for which the present invention is aimed. [Figure 2] This figure schematically represents a chemical reaction graph encoded by the method that is the objective of the present invention. [Figure 3] This figure schematically represents a second specific sequence of steps that constitute the method for which the present invention is aimed. [Figure 4] This figure schematically represents a third specific sequence of steps that constitute the method for which the present invention is aimed. [Figure 5] This figure schematically represents a fourth specific sequence of steps that constitute the method for which the present invention is aimed. [Figure 6] This diagram schematically represents the state in which a chemical reaction is encoded by the software that is the objective of the present invention. [Figure 7] This diagram schematically represents the state in which an equilibrium chemical reaction is encoded by the software that is the objective of the present invention. [Figure 8] This diagram schematically represents the instructions of a specific instruction set of software, which is the objective of this invention. [Figure 9] This diagram schematically represents the state in which a multi-step chemical reaction is encoded by the software that is the objective of the present invention. [Figure 10] This diagram schematically represents the state in which an equilibrium chemical reaction is encoded by the software that is the objective of the present invention. [Figure 11] This figure schematically represents a specific sequence of steps in the expansion method, which is the objective of the present invention. [Figure 12] This figure schematically represents a specific sequence of steps in a method for generating a chemical reaction graph, which is the objective of the present invention. [Figure 13] This figure schematically represents a first specific sequence of steps in a method for training a classifier, which is the objective of the present invention. [Figure 14] This figure schematically represents a second specific sequence of steps in a method for training a classifier, which is the objective of the present invention. [Figure 15] This figure schematically represents a specific example of the production method, which is the objective of the present invention, and related results. [Figure 16] This figure schematically represents a specific example of the production method, which is the objective of the present invention, and related results. [Figure 17] This figure schematically represents a specific example of the production method, which is the objective of the present invention, and related results. [Figure 18] This figure schematically represents a specific example of the production method, which is the objective of the present invention, and related results. [Figure 19] This figure schematically represents a specific example of the production method, which is the objective of the present invention, and related results. [Figure 20] This figure schematically represents a specific example of the production method, which is the objective of the present invention, and related results. [Figure 21] This figure schematically represents a specific example of the production method, which is the objective of the present invention, and related results. [Figure 22] This figure schematically represents a specific example of the production method, which is the objective of the present invention, and related results. [Figure 23] This figure schematically represents a specific example of the production method, which is the objective of the present invention, and related results. [Figure 24] This figure schematically represents a specific example of the production method, which is the objective of the present invention, and related results. [Figure 25] This figure schematically represents a specific example of the production method, which is the objective of the present invention, and related results. [Figure 26] This figure schematically represents a specific example of the production method, which is the objective of the present invention, and related results. [Figure 27] This figure schematically represents a specific example of the production method, which is the objective of the present invention, and related results.

[0054] Detailed description of the invention This description is not exhaustive, as each feature of one embodiment can be advantageously combined with any other feature of any other embodiment.

[0055] Furthermore, various inventive concepts can be embodied as one or more methods for which examples are provided. The actions performed as part of the method can be ordered in any suitable manner. Thus, embodiments can be constructed in which actions are performed in a different order than those shown. This may include performing several actions simultaneously, even if they are shown as sequential actions in the exemplary embodiments.

[0056] In this specification and in the claims, the indefinite articles "a" and "an" should be understood to mean "at least one" unless explicitly indicated otherwise.

[0057] In this specification and in the claims, the expression “and / or” as used herein should be understood to mean “either or both” of the elements thus combined, that is, elements that exist in some cases as a combination and in other cases as separate. Multiple elements listed with “and / or” should be interpreted in the same manner, that is, “one or more” of the elements thus combined. Other elements other than those specifically identified by the “and / or” clause may exist, whether related to or unrelated to those specifically identified elements. For this reason, as a non-restrictive example, a reference to “A and / or B” when used in conjunction with an open-ended expression, for example, “comprising,” may refer in one embodiment to A only (which may include elements other than B), in another embodiment to B only (which may include elements other than A), in yet another embodiment to both A and B (which may include other elements), and so on.

[0058] As used herein and in the claims, “or” should be understood to have the same meaning as “and / or” as defined above. For example, when separating items within a list, “or” or “and / or” should be interpreted as inclusive, that is, including at least one of a number of elements or a list of elements, and possibly including additional non-listed items. Conversely, only the explicitly indicated terms, such as “one of” or “exactly one of” or, as used in the claims, “consisting of,” would refer to including exactly one element of a number of elements or a list of elements. In general, as used herein, the term “or” should be interpreted as indicating exclusive substitution (i.e., “one or the other, but not both”) only when modified by terms of exclusivity, such as “either,” “one of,” “one of” or “exactly one of.” “Consisting of essentially” should have the usual meaning as used in the field of patent law when used in the claims.

[0059] As used herein and in the claims, the expression “at least one” with respect to a list of one or more elements means at least one element selected from any one or more elements in the list of elements, but not necessarily including at least one of each and all of the elements specifically listed in the list of elements, and should be understood not to exclude any combination of elements in the list of elements. Furthermore, this definition may allow for the presence of elements other than those specifically identified in the list of elements to which the expression “at least one” refers, whether related to those specifically identified elements or not. Therefore, as a non-limiting example, "at least one of A and B" (or equivalently, "at least one of A or B" or equivalently, "at least one of A and / or B") may mean, in one embodiment, at least one A, possibly including two or more A's, and no B (and possibly including elements other than B); in another embodiment, at least one B, possibly including two or more B's, and no A (and possibly including elements other than A); and in yet another embodiment, at least one A, possibly including two or more A's, and at least one B, possibly including two or more B's (and possibly including other elements).

[0060] In the claims and the above specification, all transitional clauses, such as “comprising,” “including,” “carrying,” “having,” “containing,” “involving,” “holding,” and “composed of,” should be understood to be open-ended, meaning “including but not limited to these.” Only the transitional clauses “consisting of” and “consisting essentially of” are closed or semi-closed transitional clauses, respectively, as described in Section 2111.03 of the U.S. Patent and Trademark Office's Manual of Patent Examination Procedures.

[0061] Please note that the diagram is not to an accurate scale.

[0062] It should be noted that the term “computer interface” should be understood as any type of human-machine interface, such as a graphical user interface (GUI) associated with an input means, such as a keyboard, mouse, or touchscreen. Furthermore, these terms also refer to any software or digital interface, such as an application programming interface ("API"), or any other type of digital input / output means or software.

[0063] It should be noted that the terms “computer device” or “computer system” should be understood as any electronic computing means, for example, a microprocessor associated with computer memory and the necessary input / output subsystems. The specific architecture of the computer system used in the following description is not important in consideration of the present invention. That is, such a computer system can be distributed and integrated using a client-server architecture or using local and / or remote computer resources. The data to be stored and accessed can be stored in traditional databases, computer memory or distributed databases.

[0064] It should be noted that the term "chemical reaction graph" specifies a modeling of a chemical reaction in a graph format, where each molecule (reagent and product) is modeled on a graph, with its vertices corresponding to atoms of the compound and its edges corresponding to chemical bonds. In other words, a chemical reaction graph models the structural formula of a compound from the perspective of graph theory. Typically, a molecular graph includes digital identifiers for atoms and bonds that enable the construction of the graph. These digital identifiers can be graphically translated into labels and vertices. Such digital identifiers can be stored in digital storage devices, such as computer memory, server databases, or distributed databases.

[0065] The term “character” should be understood to refer to any symbol (whether or not it is an alphabet) that can be used to generate a code from the input. Typically, a character can be an ASCII ("American Standard Code for Information Interchange") code representing a character. However, this is not to limit the invention.

[0066] Figure 1 shows a series of steps corresponding to the instructions for chemical reaction graph compression software, for example, for one-step reactions, multi-step reactions, and equilibrium reactions. This software is, - Step 105 of receiving a chemical reaction graph containing at least one chemical reaction reagent and at least one chemical reaction product on a computer interface, - A first step 110 in which a computer device encodes the chemical reaction graph describing the structure of at least one reagent and the product, - Step 115, in which a computer device determines the changing bonds in the coding representing the chemical structure of at least one reaction reagent and the product, - A second step 120 of encoding, for at least one determined changing bond, a computer device encodes, in which case at least one character representing the atom undergoing the bond change, at least one character representing the type of determined changing bond, and at least one character representing the atom resulting from the bond change into a single string, wherein the changing bond is encoded by a set of two characters representing the determined changing bond, the first character representing the reagent bond, and the second character representing the product bond, and each character is selected from a library of bijective characters where one character represents one type of bond. - Step 125 provides the string corresponding to the encoding of the changing bonds in the chemical reaction on a computer interface. Executes the corresponding instruction.

[0067] The receiving process 105 is performed, for example, using any type of computer interface. During this receiving process 105, a digital resource is received, which represents a chemical reaction graph. The “digital resource” should be understood in the broadest possible sense, i.e., a structured set of data. Such a digital resource can be a file stored in computer memory or generated when needed. Alternatively, instead of the file itself, a digital address about the file can be received.

[0068] Alternatively, during receiving step 105, a digital identifier corresponding to at least one reagent and at least one product is received. Such a digital identifier can be either a digital resource representing the reagent or product, or any pointer to said digital resource. Such a digital identifier can be, for example, an address in a database or a natural language string representing said reagent or product. In another variation, the digital identifier is a GUI component that is user-executable and, once activated, triggers the input of the associated resource and / or the address of said resource.

[0069] The receiving process 105 can be triggered by the user or by automated input.

[0070] The first encoding step 110 is performed, for example, by a computer system configured to run dedicated software. This encoding step 110 can be performed, for example, in the same way as the SMILES format of the chemical reaction graph is generated. During this encoding step 110, the chemical reaction graph is encoded, preferably, into a string of characters in ASCII format.

[0071] Alternatively, the first step 110 encoding is configured to provide a line representation using a SMARTS ("SMILES arbitrary target specification") variant of the SMILES encoding format. The SMARTS encoding format is a language for identifying substructure patterns within a molecule. Figure 6 shows the results of such a first step 110 encoding for references 630 and 640 for reactions 605 and 610, respectively.

[0072] The determination step 115 is performed, for example, by a computer system configured to run dedicated software. Several options can be implemented during this determination step 115.

[0073] - The need for human input on a computer interface to map atoms and bonds in a chemical reaction graph or - Either a computer system automatically maps atoms and bonds in a chemical reaction graph, and in either case, - To detect changes in molecular structure due to either changes in atoms at specific mapped positions in the molecular graph or changes in bonding to the mapped atoms or any other atoms in the related molecule, a computer system compares the molecular chemistry graph of the product with the molecular chemistry graph of the reagent and - Classifying specific combination changes within a predefined list of types, as a function of the results of the comparison process, using a computer system.

[0074] Such embodiments using superposition comparison are typically used in modern solutions. However, these approaches typically lack certainty in the resulting mappings because they seek minimal structural commonalities. This commonality would not detect, for example, the breakdown-production process if oxygen molecules were used as both a reagent and a product.

[0075] In more advanced embodiments, transformer machine learning algorithms are used.

[0076] Such models can be trained with data including the USPTO-50 set (or a portion thereof) from the paper by Schneider et al. (Schneider, N.; Stiefl, N.; Landrum, GA, What's What: The - Nearly - Definitive Guide to Reaction Role Assignment. J Chem Inf Model 2016, 56, 2336-2346), and some calculations can also use the training set data from Jaworksi et al. (Jaworski, W., Szymkuc, S., Mikulak-Klucznik, B. et al. Automatic mapping of atoms across both simple and complex chemical reactions. Nat Commun 10, 1434-2019).

[0077] Such models can be tested against test sets, which may be part of the USPTO-50 set not used for training, or against manually selected reactions. In addition, the performance of the developed method can be tested using the 857 reaction test set by Jaworksi et al.

[0078] Such data can be refined before input. Furthermore, the data can be compressed and encoded according to Method 100, which is the objective of this invention, before being used as training / test data.

[0079] For example, one of the following transformer architectures described in the following publications can be used. - Vaswani, A., et al. Attention Is All You Need. Preprint at https: / / arxiv.org / abs / 1706.03762 (2017) - Schwaller, P., et al. Molecular transformer: a model for uncertainty-calibrated chemical reaction prediction. ACS Cent. Sci. 5, 1572-1583 (2019) and / or - Tetko, IV, Karpov, P., Van Deursen, R. et al. State-of-the-art augmented NLP transformer models for direct and single-step retrosynthesis. Nat Common 11, 5575 (2020)

[0080] Specifically, the transformer consists of six layers and eight heads (6×8). Model training was limited to 100 epochs, using a batch size of 3000 characters. The input data was reaction data (both reagents and products) in SMIRKS format, and the target was each chemical reaction graph compressed and encoded according to Method 100, which is the objective of this invention. Both the input and target sequences can be augmented, as shown in Figure 12. This improves data diversity and eliminates the effects of overfitting in the neural network. The data for model training and testing can be augmented, for example, by 5× and 20× times, respectively.

[0081] The transformer model allows for the generation of multiple predictions for given input data using beam search. Using beam search, n=10, and thus, in accordance with the objectives of Method 100, which is an objective of the present invention, 10 predicted compressed and coded chemical reaction graphs (CRS) are received for each input reaction. Since the 20 × data augmentation used can be applied to each reaction, a total number of up to 200 predicted CRSs can be calculated for each analyzed reaction.

[0082] Further post-processing, for example, - Due to obvious formatting errors, some calculated CRS will be filtered before further analysis. - Mass equilibrate the reactants and / or products to confirm that all reactants and reagents generated by the decomposition of CRS are present during the initial reaction. It is possible to do so.

[0083] Such a transformer model can provide the following results. [Table 1]

[0084] By implementing such a transformer, superior performance was demonstrated when trained using this data. A model developed using the 43.8k (43,800) USPTO-50k training set demonstrated 99.9% coverage and 100% precision for a test set of 4,885 reactions. Therefore, the transformer was able to accurately predict the atomic mapping of all reactions from that test set. The model's performance was lower for manually annotated set A, achieving 96.7% coverage and 96.9% precision for this set.

[0085] For the NatureTest set, the coverage was much lower, at only 67.3%. This lower coverage indicated that the Nature set contained reaction species for which no patents existed and / or were more complex, and the model was unable to generate one or more valid CRSs for them. However, the same very high accuracy was calculated. Therefore, the transformer model was able to reproduce the correct mapping precisely, provided that the generated CRSs included all components of the initial reaction data.

[0086] The addition of the NatureTrain set (n=548) improved data diversity, resulting in an approximately 7.3% improvement in coverage for NatureTest and an improvement of over 1% for set A. Adding simulation data generated using the NatureTrain set provided an additional boost to coverage for NatureTest. This data included 10 generated reactions for each initial reaction. This generation better represented rarer reactions and improved model accuracy. However, even after adding the simulated reactions, coverage for NatureTest remained below 80%. This indicates that some reactions from this set are not adequately represented in both the USPTO patents and the NatureTrain set.

[0087] To address this issue, it is possible to include simulated responses to the NatureTest. This increases coverage for this setting to 95% without sacrificing accuracy. The latter expansion of the dataset also provided the best overall results for set A. Coverage improved to 98.9%, and accuracy reached 97.4%. For set B, the results remained unchanged, with all three accuracy measures approximately 100%.

[0088] The second encoding step 120 is performed, for example, by a computer system configured to run dedicated software. During this second encoding step 120, at least one of the determined changing combinations is encoded into a set of ASCII characters representing the type of determined changing combination.

[0089] This second encoding step 120 may include, for example, the following steps:

[0090] - A step of syntactically analyzing the encoded response graph obtained from the first encoding step 110, - Steps to extract binding tables for reagents and products, - A step of generating a second encoding by assembling the reaction graph of the reagents and products, and - Depending on the circumstances, the generated second encoding can be exported to a canonical linear notation string and the concatenation can be written using the specified symbols shown below.

[0091] The second step 120 of encoding the changing bonds in a single string can be performed by associating each changing bond with a symbol, for example, a set of at least one character describing the type of bond change. This symbol is preferably associated with the adjacent atoms during which the bond change is occurring.

[0092] Such symbols can be a string of four characters consisting of a single binding character for the bond species of the reagent and product, enclosed by curly braces or other characters defined as neutral. A change from a single bond to a double bond in a reaction is described, for example, using the character permutation "{-=}". The relationship linking a character to the bond change expressed is preferably bijective. The term "bijective" as used herein refers to a one-to-one relationship linking a character to the bond change expressed. The term "character" is to be understood as any symbol in a dictionary of symbols and is not limited to alphanumeric characters. This means that a library of characters can be set up before the encoding step. In this library, each character represents a type of bond change. This library can be constructed manually or automatically. In certain embodiments, an algorithm can be trained to learn its own symbols. During the subsequent encoding step, the appropriate character or symbol is selected from the library as a function of the determined bond change.

[0093] Apart from formats where both reagents and products are single strings, this format stands out for its extremely short format, which does not require explicit atomic numbers to mark reaction sites. In fact, reactions are implicitly defined by changing bonds. In SMARTS chemical reactions, all reagents and products are defined by a new SMILES string. The order of atoms can vary widely, including in canonical form. Consequently, explicit indices must be defined in the SMILES string to define which atoms are identical in the reagent and product, e.g., [CH3:1][CH2:2][CH3:3] (where :1, :2, and :3 define atomic indices). Agents are typically not included, as they do not contribute to net chemical modification. Agents and conditions vary by reaction and can be selected by the user based on the reaction species. Such agents and conditions can be adjusted at the user's discretion, as illustrated in Figure 7. A further important advantage of this proposal is its compressibility for large datasets. In fact, this format is the shortest format known for describing reactions. While this application focuses on reactions involving bond cleavage, formation, and changes in bond order, other types of bond changes can be encoded in this manner. Reactions involving ionic bond formation and cleavage, as well as purification, such as chiral separation, are not considered here. In the latter group of reactions, the graph connectivity of atoms does not change. Such separations can be described as AB > B, with AB > B being an example of purification producing B.

[0094] The corresponding tables represent possible symbolic choices for different types of join changes. [Table 2]

[0095] In the reagent, bonds indicated as "none" are bonds in the product formed during the reaction. In the product, bonds indicated as "none" are reagent bonds that were cleaved during the reaction.

[0096] In certain embodiments, the second encoding step 120 is configured to encode the changing bond in a set of two characters representing the determined changing bond. The first character represents the reagent bond, and the second character represents the product bond.

[0097] In a particular embodiment, the second encoding step 120 is configured to embed two characters representing the determined changing combination between two neutral tag characters representing the presence of the encoding of the changing combination.

[0098] An example of such output 205 is shown in Figure 2.

[0099] The provided process 125 is performed, for example, on a GUI or using an API.

[0100] Figure 1 further illustrates the method 100 implemented by the software disclosed above. This chemical reaction graph compression method 100 for one-step reactions, multi-step reactions and equilibrium reactions is: - Step 105 of receiving a chemical reaction graph containing at least one chemical reaction reagent and at least one chemical reaction product on a computer interface, - A first step 110 in which a computer device encodes the chemical reaction graph describing the structure of at least one reagent and / or the product, - Step 115, in which a computer device determines the changing bonds within the encoding representing the chemical structure of at least one reaction reagent, - A second step 120 in which, for at least one determined changing bond, the changing bond in a string of characters associated with at least one character representing the atom undergoing the bond change and at least one character representing the atom resulting from the bond change is encoded into a single string by a computer device, - Step 125 provides a string of characters on a computer interface that corresponds to the coding of the changing bonds in a chemical reaction. Includes.

[0101] In a particular embodiment, the first encoding step 110 is configured to encode the chemical reaction graph into line notation. The method further includes a step 130 to extend the line notation encoding before the second encoding step 120.

[0102] The extension process 130 is performed, for example, by a computer system configured to run dedicated software. During this extension process 130, the row representation of the chemical reaction graph is rearranged so as not to change the properties of the encoded chemical reactions, while still providing alternative encodings for those chemical reactions.

[0103] An example of the results of the extended process 130 is shown in Figure 12.

[0104] In certain variants, the reaction can be reduced to a reaction site in a reaction coding or chemical reaction graph reduction step (not shown). Such a reduction step of reaction coding is performed, for example, based on a row notation obtained from a first coding step 110 or a second coding step 120. This removes all atomic identifiers that remain inert during the modeled chemical reaction. For example, if the atomic identifier and associated bond are located in a molecule that remains unchanged from the reagent reaction step to the product step of the chemical reaction, the atomic identifier is removed.

[0105] This reduction step of reaction coding further compresses the chemical reaction graph into a useful set of symbols. Figure 6 shows reaction 615 reduced to the reaction site for product formation. The reaction is shown including the first neighboring atom.

[0106] In a modified form in which the reaction (615 in Figure 6) is reduced to a reaction site, a corresponding linear representation limited to the atoms and bonds being modified can be obtained during the chemical reaction result 620. Such a linear representation can be labeled "SiteSMARTS". Such a result can correspond to the output of the first encoding step 110, or to the output of a dedicated reduction step of reaction encoding that may be located upstream or downstream of the first encoding step 110.

[0107] In fact, in a reaction, it is possible to distinguish between reagents, which are compounds that interact with each other to produce the product, and other chemicals that do not change during the reaction, such as catalysts and solvents. For this reason, atomic mapping is only necessary for chemicals whose bonding information changes, while non-interacting / unchanging parts can be skipped.

[0108] "SiteCRS" is a tool that supports compressed chemical reaction graphs limited to reaction sites. - A step to identify the reaction site, - Depending on the case, a step of flagging all site atoms that are "related" to the reaction, including adjacent atoms and / or atoms in the ring or ring system of atoms up to the user-specified topology depth, - A process to remove all atoms from the molecule that are not flagged as "related," - A step to export a reduced graph of the reaction with the reaction site, in some cases, as a canonical string of characters. It can be calculated using

[0109] In the current dataset, reactions can be described using the SMARTS format, i.e., "reagent > activator > product". For a complete reaction, the SMARTS format is: - A step of identifying atoms having an altered environment based on atomic and bonding changes, - Depending on the case, a step of flagging all atoms that have a change, including adjacent atoms and / or atoms in a ring or ring system of atoms up to a user-specified topology depth, - A process to remove all atoms from the molecule that are not flagged as "related," - A process of renumbering the map numbers on atoms by arbitrary canonicalization, - The process of exporting SMARTS, in some cases, to create canonical SiteSMARTS. The process is applied and converted to SiteSMARTS.

[0110] A subset of the reactions is analyzed, for example, from the NextMove Pistachio dataset 5. These reactions can be divided and analyzed according to their published classes. For example, class 1.1.1 defines Chan-ram alkylamine couplings. For example, - A step to identify the reagents and products involved in the net reaction, - A step to remove the active ingredients, solvent and non-participating reactions, - The process of calculating SiteSMARTS and SiteCRS in order to characterize the response. This can be applied.

[0111] The generated SiteCRS can be used to cluster reaction transformations using string tags instead of fingerprints. There are two main advantages: chemists can understand the tags, and therefore can check whether the resulting tags are related to the reactants in the selection process.

[0112] The compressed chemical graph of a reaction obtained from Method 100, which is the object of the present invention, can be extended to include changes for stereochemistry. For example, {- / } and {-\} define a change from a single bond to an upright or downright single bond, and {-^} and {-_} for single bonds define a change to “single up” or “single down” for relative stereochemistry on the tetrahedral center. Similarly, it is possible to transition from a double bond to a single up or single down. For this purpose, the symbols {=^} and {=_} can be used. For this reason, the reverse direction can be used, for example, {^=} from single up to a double bond and {_=} from single down to a double bond. An example of such a reaction is the hydrogenation of alkynes. Depending on the reaction conditions, chemists can carry out reactions involving syn-hydrogenation or anti-hydrogenation to prepare cis- and trans-alkenes from alkynes, respectively. An example of a stereochemical reaction having a tetrahedral center is the biocatalytic reduction of ketones to secondary alcohols by the enzyme class alcohol dehydrogenase. An example is the reduction of raspberry ketone to 4-3R-hydroxybutylphenol.

[0113] For reference, Figure 2 shows examples of formatted chemical reaction graphs 205 and 210 obtained by the method 100 disclosed above.

[0114] Figure 3 shows a specific embodiment of Method 300, which is the objective of the present invention. This chemical reaction dataset expansion method 300 is - Step 305 of receiving a string on a computer interface according to a format obtained from any variation of the implementation of Method 100 disclosed with respect to Figure 1, - A step 310 in which a computer system rearranges a string of characters, in order to shift at least one string of characters, each of which represents at least one character representing an atom and at least one character representing a bond change associated with the corresponding atom. - Step 320 outputs an extended string of characters corresponding to the response initially encoded by the received string on the computer interface. Includes.

[0115] Functionally and structurally, the receiving process 305 is similar to any variation of the receiving process 105 disclosed with respect to Figure 1.

[0116] The sorting step 310 is functionally and structurally similar to the expanding step 130 disclosed with respect to Figure 1. During this step 310, the symbols or characters of the chemical reaction graph formatted and compressed according to method 100 are formally rearranged to provide alternative encodings that represent a single chemical reaction graph. An example of this can be seen in Figure 2, where the chemical reaction graph is formatted and compressed in two alternative encodings 205 and 210.

[0117] In a particular variant, Method 300, which is an object of the present invention, includes a step 315 in which a computer system associates at least two strings according to a format obtained from any variant of the implementation of Method 100 disclosed with respect to Figure 1, wherein each string represents the same chemical reaction graph.

[0118] This association step 315 is performed, for example, by a computer system configured to run dedicated software. During this association step 315, alternative compressed encodings for the chemical reaction graph can be concatenated into a single string, which can preferably be separated by a neutral symbol or character, such as a dot in Example 215 shown in Figure 2.

[0119] The output process 320 is functionally and structurally similar to the provided process 125 disclosed with respect to Figure 1.

[0120] A broader view of the extensions 1100 achievable by using the present invention can be seen in Figure 11. Figure 11 shows several possible extension inputs 1105, 1110, and 1115, as follows:

[0121] - A compressed and formatted chemical reaction graph 1105 according to the format that is the object of the present invention, such a format is abbreviated as CRS ("chemical reaction string").

[0122] - An input (e.g., a file) 1110 defining one or more valid reagents and products in any machine-readable chemical format, this input including, for example, .mol ​​("Molfile"), .sdf ("Structure Data File"), .xyz ("XYZ File Format") files, and / or - Input representing chemical reactions, for example, the line notation of SMARTS-encoded chemical reactions, such a format is abbreviated as RxnSmarts.

[0123] Figure 11 shows several possible extended outputs 1125, 1130, 1135, 1140, and 1145, as follows:

[0124] - Alternative compressed and formatted chemical reaction graphs 1125-canonical form can be used to standardize atomic order, describing the same reaction with changes in atomic order.

[0125] - A finite list of compressed and formatted chemical reaction graphs of [1,N] that define the same reaction, this list will presumably reduce to a unique set of compressed and formatted chemical reaction graphs.

[0126] - For example, a finite list of compressed and formatted chemical reaction graphs in [1,N] separated by the dot character ".", can be reduced to a unique set of compressed and formatted chemical reaction graphs describing the same reaction.

[0127] - A list or set of finite lists of [1,N] delimited, compressed, and formatted chemical reaction graphs. - A finite matrix having [1-N] rows and [1-M] columns that defines a single or concatenated, compressed, and formatted chemical reaction graph for the same reaction, and / or - A list or set of [1,N] finite matrices with [1-N] rows and [1-M] columns that define a single or concatenated, compressed, and formatted chemical reaction graph.

[0128] Such extensions 1120 can be achieved in the same way as the extension step 130 or the sorting step 310 disclosed above.

[0129] Data augmentation can be used in various applications.

[0130] - Data augmentation for training models on small datasets - Balancing and / or balancing imbalanced datasets - Training a neural network or model using ensemble representations.

[0131] Figure 4 shows a specific embodiment of Method 400, which is the objective of the present invention. This chemical reaction dataset pretreatment method 400 is - Step 405 of receiving a dataset of at least two chemical reaction graphs, each containing at least one chemical reaction reagent and at least one chemical reaction product, on a computer interface, - Step 100 of compressing at least two chemical reaction graphs according to the method disclosed with respect to Figure 1, - Step 410, in which a computer system determines the distribution of chemical reaction classes within an encoded dataset, - Step 300 of extending the dataset for at least one chemical reaction class as a function of the determined distribution according to the method disclosed with respect to Figure 3, - Step 415 to output the preprocessed dataset on a computer interface. Includes.

[0132] Functionally and structurally, the receiving process 405 is similar to any variation of the receiving process 105 disclosed with respect to Figure 1. This receiving process 405 can be implemented by implementing several sequential or serial instances of the receiving process 105, or by implementing a single receiving process 105 configured to receive several datasets with a single input.

[0133] The compression process or compression method 100 is disclosed in several modified forms with respect to Figure 1.

[0134] The determination step 410 is performed, for example, by a computer system configured to run dedicated software. During this determination step 410, statistical analysis is performed on the dataset and compared to static or dynamic tolerance thresholds. Such thresholds may be, for example, absolute or relative values ​​for samples per reaction class with respect to samples for other reaction classes in the dataset. The term “chemical reaction class” is also called “chemical reaction species” (e.g., synthesis, decomposition, and substitution).

[0135] The process or method 300 for extending the dataset is disclosed in several variations with respect to Figure 3. Alternatively, this process 300 for extending the dataset may be performed in place of, or in parallel with, the execution of the dataset extension process 130 prior to the encoding process 120 for extending the dataset.

[0136] The output process 415 is functionally and structurally similar to the provided process 125 disclosed with respect to Figure 1.

[0137] Figure 5 shows a specific embodiment of Method 500, which is an objective of the present invention. This training method 500 for classifiers, transformers or regressors is - Step 505 involves inputting a dataset of chemical reaction graphs encoded in a compressed format, such as one obtained by any variation of Method 100 disclosed with respect to Figure 1, onto a computer interface. - Step 510 involves a computer system operating a recurrent neural network architecture configured to classify the progression of chemical reaction bonds as a function of the input, using a dataset of chemical reaction graphs as input. - Step 515 outputs a trained classifier, transformer, or regressor on a computer interface. Includes.

[0138] Functionally and structurally, the receiving process 505 is similar to any variation of the receiving process 105 disclosed with respect to Figure 1. This receiving process 405 can be implemented by implementing several sequential or serial instances of the receiving process 105, or by implementing a single receiving process 105 configured to receive several datasets with a single input.

[0139] The operation step 510 is performed, for example, by running a recurrent neural network architecture and associated software on a computer system based on a training set.

[0140] The output process 515 is functionally and structurally similar to the provided process 125 disclosed with respect to Figure 1.

[0141] Regressa can be trained according to the target "reaction yield," "reaction equilibrium constant," or "transition state energy."

[0142] Such a regressor can be trained according to one of the following examples.

[0143] - "Predicting reaction performance in CN cross-coupling using machine learning" by DT Ahneman, JG Estrada, S. Lin, SD Dreher, AG Doyle - April 13, 2018 or - Schwaller, Philippe; Vaucher, Alain C.; Laino, Teodoro; Reymond, Jean-Louis (2020): Prediction of Chemical Reaction Yields using Deep Learning. ChemRxiv. Preprint. (https: / / doi.org / 10.26434 / chemrxiv.12758474.v2).

[0144] Furthermore, the present invention also aims to provide a method for predicting the progression of chemical reaction bonds, which operates a classifier, transformer, or regressor obtained by the training method disclosed with respect to Figure 5.

[0145] These embodiments enable the detection of reaction sites. Such embodiments are disclosed above.

[0146] Furthermore, the present invention also aims to provide a method for generating chemical reactions that operate a classifier, transformer, or regressor obtained by the training method disclosed with respect to Figure 5.

[0147] This method of generating chemical reactions uses the following as input.

[0148] - A vector of length N containing discrete values ​​to identify the type of character in a portion of a compressed and formatted chemical reaction graph obtained by method 100, which is an objective of the present invention, for example, a compressed and formatted chemical reaction graph tokenized into a one-hot encoder. - A compressed and formatted chemical reaction graph obtained by method 100, which is the object of the present invention, is tokenized into an N × M dimensional one-hot coding matrix having N rows defining possible characters and M columns describing the length of a portion of the compressed and formatted chemical reaction graph. - A one-hot encoded vector of size M that defines the following character, - Obtained from Method 100, which is the object of the present invention, and can optionally be used as a set of characters, for example, "{!-}" which defines a single position in a vector, flexible reaction bonds and / or in a compressed and formatted chemical reaction graph - A tokenizer that adds a stop character to the end of a sentence.

[0149] In such a chemical reaction generation method, for example, a four-layer architecture including the following is used as the network.

[0150] - An input layer that takes a tokenized vector or matrix for characters of possible sequence lengths N and M, - One or more recurrent neural networks (RNNs) with sequence lengths from 2 to 1024, - A dropout layer for a portion of the RNN's output (from 0 to less than 100%) and a dense layer of size M vectors with probabilities for the next character.

[0151] Such a model can be trained to chemically accurate predict the next most likely character in the network. The network predicts the probability of all possible characters and randomly selects the next character. Writing is a recursive process of writing: select-predict-select-predict until a finite number N valid responses are generated.

[0152] The network output continuously writes CRS both inside and outside the learned reaction space, depending on how deeply the generative model has been trained.

[0153] Figure 12 further illustrates an architecture 1200 that performs two steps: a step 1205 for training a generative neural network and a step 1210 for generating responses, as well as related steps: a step 1215 for inputting sample data to train the generative neural network and a step 1220 for outputting the generated responses.

[0154] Figure 13 further illustrates the training method 1300 disclosed above. In this method 1300, - The compressed and formatted (encoded) chemical reaction graph is input to tokenizer 1310 (1305), - The tokenizer 1310 is, - A process of tokenizing a compressed and formatted (encoded) chemical reaction graph into a network input 1315 which is either a discrete vector or a one-hot matrix, - Step 1325 of pairing each token with the next character in the input, compressed, and formatted (encoded) chemical reaction graph. The system is configured to operate. The token is used as a training target 1320 for the RNN, and the training target 1320 is edited, for example, into a one-hot vector.

[0155] Figure 14 shows an alternative method 1400 to that in Figure 13. In this method 1400, a string of characters that encodes the bond changes between atoms is encoded as a specific unitary token.

[0156] Figure 6 shows a specific embodiment of state 600 in which a chemical reaction is encoded by the software that is the objective of the present invention.

[0157] For example, chemical reaction graphs for reference numerals 605 and 610 can be seen in Figure 6. Figure 6 shows Williamson ether synthesis as an example. Reference numeral 605 specifies the ether synthesis between ethyl alcohol and ethyl bromide to form diethyl ether, and 610 specifies the ether synthesis between cyclohexanol and ethyl bromide to form ethoxycyclohexane.

[0158] Figure 7 shows a specific embodiment of state 700, in which the equilibrium chemical reaction is encoded by the software that is the objective of the present invention. These states 700 are - Equilibrium reaction 705, which has a reaction constant K and exhibits complete atomic mapping and net chemical equilibrium, - SMARTS and compressed chemical reaction graphs for forward reaction 710, - SMARTS and compressed chemical reaction graphs for reverse reaction 715 Includes.

[0159] The compressed and formatted combination between forward and reverse reactions demonstrates the easy reversibility of the format, which is the objective of this invention, by changing the order of combinations within a string of characters. This can be easily seen in the character swap of "=" and "!" shown between reactions 710 and 715, which represent reverse actions such as synthesis and retrosynthesis. This significantly reduces the amount of data to be stored and the ability to use fewer samples for machine learning applications.

[0160] For example, any reaction, such as reaction 705, can be formally represented by an equilibrium in which a constant K can define the ratio between the product and the reagent. The value of K can vary from zero to infinity. This phenomenon can be used to extend reaction data by using both CRS representations (preferably combining the CRS for forward reactions and the CRS for reverse reactions).

[0161] A further important advantage of the format of the present invention is its compressibility for large datasets. The compressed chemical reaction graph defines the shortest format for defining the net chemical reactions currently available.

[0162] Figure 7 also shows the ability of the format to add reaction conditions, such as solvents and / or catalysts, to a CRS letter string. The example shown here is a Grignard reaction, which is carried out using magnesium Mg in the solvent diethyl ether. Water, chemically described as "O" in the CRS, is used to terminate the reaction by hydrolysis. This type of CRS can be considered a "conditional CRS" for proposing reaction conditions for a given CRS.

[0163] Figure 8 schematically shows the instructions of a specific embodiment 800 of the software that is the objective of the present invention. These instructions are - For example, inputting the RxnSMARTS format including alkaline conditions (KOH) and solvent (Me2SO) (805), - Cleaning with only the reagents and products involved in the RxnSMARTS reaction (810) (this step also neutralizes the reagents and / or products and defines the net chemical transformation), - Completing the atomic map number (815) to define a complete net chemical reaction, - To generate CRS, SiteCRS and / or SiteSMARTS (820) That is the case.

[0164] Figure 9 schematically shows the sequential reaction steps (A and B) encoded within the multi-step reaction coding 900 by the software which is the objective of the present invention.

[0165] In certain embodiments, a multi-step reaction represented by a series of bonding changes between two atoms is encoded by a series of single letters, each single letter representing a sequential state of bonding between the two atoms, and the order of the letters represents the sequence of bonding changes between the two atoms.

[0166] In this embodiment, the bond change between two atoms is encoded as follows: "atomic symbol 1" "{" (neutral character) "reagent bond character" "reaction bond character of the first-step product" "reaction bond character of the second-step product" "reaction bond character of the nth-step product" "}" (neutral character) "atomic symbol 2".

[0167] Figure 10 schematically shows the equilibrium reaction encoded within the equilibrium reaction encoding 1000 by the software which is the objective of the present invention.

[0168] The novel reaction formats disclosed herein represent the shortest possible syntax for describing net chemical transformations. Indeed, the newly generated compressed chemical reaction graphs are approximately 20% shorter in length compared to the corresponding RxnSMARTS for the same reactions (Figures 6–10).

[0169] Furthermore, this method of generating chemical reactions can also be understood from the perspective of Figures 15 to 27.

[0170] In recent years, generative neural networks have become a powerful deep learning method for generating realistic in silico data from real-world examples. Generative neural networks are successfully used to generate deepfakes for images and audio, and to create realistic computer-generated images and videos. An example of a deep generative model is a latent space, typically Z(μ, σ). 2 This includes variational autoencoders (VAEs) that use a compressed set of parameters based on sampling from ) or use a generative adversarial network in which two networks, namely a generator G and a discriminator D, repeatedly compete to produce a realistic synthetic solution that the discriminator can no longer distinguish from real data.

[0171] In chemistry, generative models are highly useful for molecular discovery, using the techniques described above to generate novel molecules. In particular, generative neural networks trained to describe the chemical language SMILES have been used using methodologies known from natural language processing. These approaches are limited to molecular-level processing. This invention proposes a verification mechanism involving probabilistic sampling. This novel strategy introduces a generative verification network that defines an adaptation of an early stopping function to maintain the highest level of creativity. In this verification mechanism, the model generates a statistically reasonable sample to evaluate the model's success in describing chemically correct SMILES strings, i.e., SMILES that can be processed by a chemical toolkit without error. As illustrated, training of the neural network is stopped after the network is statistically stable with respect to the generated entries.

[0172] The format that is the object of this invention provides a syntax for defining a single-line representation of a chemical reaction graph. This syntax, sometimes referred to as a “Chemical Reaction String” (CRS) in a non-limiting sense, introduces reaction bonds into a line representation. This syntax defines a large compression of currently known reaction SMARTS and is free from any need for explicit atomic indexing. CRS can be extended to include auxiliary unmodified molecules. The CRS syntax has two main advantages: 1) easy reversibility of reactions by inverting the bond symbols used; and 2) easy extension for multi-step reactions by adding additional steps to flexible bonds. In this specification, these capabilities are exemplified for the following set of reactions: 1) a set of eight substitution reactions with iodine such as eliminating a group; and 2) multi-step hydrogenation and dehydrogenation between alkynes, alkenes, and alkanes. Finally, the main advantage of producing single-step or multi-step reactions by CRS strings is the immediate generation of multiple tasks. First, any single reaction CRS simultaneously defines the reagents, products, and reaction. Therefore, the reaction conditions can then be included, for example, invariant molecules or solvents. Example: For the Grignard reaction, "CC({-!}Br)C(C){=-}O.CCOCC.[Mg].O". In this string, CCOCC and [Mg] are auxiliary reagents.

[0173] One such example utilizes the following technical considerations.

[0174] - Datasets: This study used datasets generated using molecules publicly available on PubChem. Reaction datasets were obtained later based on well-known reactions. For example of a one-step reaction, substitution reactions for the strong leaving group iodine were used. For a multi-step reaction, the hydrogenation of alkynes to alkanes via alkenes was used.

[0175] - Substitution Reactions: Aliphatic and aromatic iodine molecules containing a single iodine were selected from PubChem. Eight substitutions for iodine were applied, defining eight different one-step substitution reactions (Figure 15). In these reactions, iodine is the stronger leaving group, and the reactions are considered non-equilibrium reactions.

[0176] - Hydrogenation reactions: Molecules with a single aliphatic carbon-carbon triple bond were selected from PubChem. This bond was transformed in a multi-step reaction to define reaction 1600: alkynes > alkenes > alkanes (Figure 16). All forward reactions were reversed by replacing the hydrogenation bond, i.e., {#=-}, with a multi-step dehydrogenation bond, i.e., {-=#}.

[0177] - Neural Network: In this example, a recurrent neural network was used to predict the next possible character. Thus, such a network defines an iterative writer that samples the next possible character based on a sequence of previously written characters. The neural network used herein consists of the following layers (Figure 17). [Table 3]

[0178] - The exemplary neural network is trained using categorical cross-entropy. Training of the neural network was stopped using a check mechanism. The check mechanism is an early stopping function that generates statistically relevant samples of tens or hundreds of generated entries and measures the number of valid entries. The early stopping function stops training when the model shows results that are statistically stable based on a user-specified percentage of valid entries. The percentage of valid entries is considered statistically stable if that percentage is within a 90% confidence interval for the sample size used for at least 10 epochs. A generator mechanism, such as the one described below, is also used as a generator for this early stopping function.

[0179] The neural network used for generation is used herein to predict the next possible character based on the previously written character. Therefore, the network is an iterative writer. Figure 17 shows the network layout. The network used herein to illustrate this application is a network that takes a one-hot encoder matrix describing a sequence as input. Figure 18 shows a monitoring plot of the learning process showing the progress of the categorical cross-entropy loss function. Figure 19 shows the early stopping feature used in the generative check network, showing the percentage of valid responses generated by the generative neural network. The thick and dashed lines show the mean percentage with associated 90% confidence intervals for a sample size of 100 generated responses. Training is stopped early if the results are statistically stable within the 90% confidence interval. Therefore, in the above example, training was stopped after 65 epochs.

[0180] - Generation: The generation process begins when the neural network is fully trained, i.e., when the neural network has achieved statistically stable results for generating valid responses. The generator is an iterative writer that predicts the next possible character based on the last character count "n". If few characters have been written, the generator uses all characters. The initial seed used is "\n" to define the end of the previous numerator. During generation, this method repeatedly writes characters, e.g., "\nC", "\nCC", "\nCCC", etc. When it reaches size n+1, this method predicts the next character using only the last n characters of the word.

[0181] - Evaluation: Model evaluation for the 180 reaction sets is performed by counting the number of correct reactions and extracting SiteCRS, i.e., reaction site keys that define the reaction species. Based on the results, it is evaluated whether some reactions are generated more frequently, less frequently, or at approximate ratios than the ratios in the dataset. For this calculation, the number of invalid reactions is ignored in the calculation and is listed separately in the table.

[0182] Such examples lead to the results disclosed below.

[0183] Figures 20-22 show the results for substitution reactions. Reactions flagged with (#) are defined as readable but invalid due to valency errors. Reactions flagged with (^) are composed of combinations of multiple substitutions. The term "publicly known" is the opposite of a reaction known from the literature. The term "possible" indicates that the reaction may occur. The term "possible" commented with "two-step" indicates that the reaction probably consists of two independent steps, and "one-pot" indicates that both steps can occur in a single step, even if the reactions are different.

[0184] Figure 23 shows examples of reactions generated for input reactions. The reactions proposed by the generator define the reactions for effective reagents and effective products. The generator was trained exclusively using knowledge to assume possible chemical reactions, and not using information on yields. A) Substitution of aliphatic iodine with chlorine. B) Substitution of aliphatic iodine with bromine. C) Substitution of aromatic iodine with chlorine. D) Substitution of aromatic iodine with bromine. E) Substitution of aliphatic iodine with amine. F) Methyl ether formation using a Williamson reaction. G) Aromatic methoxylation by iodine substitution. H) Aromatic substitution of iodine with primary amine.

[0185] Figure 24 shows the results for multi-stage hydrogenation and dehydrogenation.

[0186] Figure 25 shows examples of reactions generated for the model's input reactions. All reactions were generated as multi-step reactions using SiteCRS, shown on the right. For illustrative purposes, the multi-step reactions were broken down into their first and second steps. The reactions shown here were generated in silico, and their synthetic feasibility was not evaluated.

[0187] To make it understandable, the generation method, which is the objective of this invention, creates a single reaction dataset consisting of eight different substitution reactions for aromatic and aliphatic iodines. All substitutions have in common that the strong leaving group iodine is replaced by another nucleophile. As a result, the reaction generator can be found to be capable of generating examples of generators for all reactions available in the training set. Note that all proportions of valid reactions are calculated excluding the number of invalid reactions. As a result, the displayed density values ​​can be compared to the reaction densities in the input set. In all samples, it is observed that the density in the generated set may clearly differ from the density in the input set. Nevertheless, the majority of reactions fall within the class of reactions presented to the generative neural network. Statistical variability is clearly a significant advantage of this generative neural network, as the generator can freely generate based on the selection of the next letter within the boundaries of the predicted probability. As a result, the distribution of generated reactions may vary among the sets of generated molecules. In addition, the degrees of freedom of the generator is a significant advantage for creating new reactions. These new reactions involve substitutions at multiple sites, but the reactions may define new ideas previously unknown to the generator. An example of such a reaction is the substitution of N-iodopyrrole to N-aminopyrrole. This example is noteworthy because the input dataset contained only substitutions on carbon atoms. Thus, in summary, reaction generators can propose both reactions within the same reaction space and generate new reactions based on their acquired knowledge of writing chemically correct molecules. An essential mechanism for maintaining the creativity of the generator is the use of a probabilistic testing mechanism that periodically tests the generator's knowledge to produce valid chemical reactions.

[0188] Figure 26 shows examples of novel reactions generated by the generator. Although the input set was unknown, i.e., consisted of the eight reactions initially defined, the generator produced new reactions. All reactions are indicated by the reagents, products, and SiteCRS above the reaction arrows. Examples include A) dehalogenation of an alkane, B) substitution of iodine with chlorine on an amine, C) substitution of aliphatic + aromatic iodine with bromine, D) substitution of iodine with a carbanion, E) substitution of N-iodopyrrole with N-aminopyrrole, and F) double aromatic substitution of iodine with bromine.

[0189] As an example of a multi-step reaction, the generator was trained for multi-step hydrogenation, i.e., hydrogenation from alkynes to alkenes and from the resulting alkenes to alkanes. Within the dataset, dehydrogenation was also defined as a multi-step reaction, i.e., a reaction from alkanes to alkenes and alkynes. For this reason, hydrogenation and dehydrogenation are written as SiteCRS C{#=-}C and C{-=#}, respectively. The CRS syntax was chosen for its flexibility and to accommodate multiple reaction steps. The SiteCRS for multi-step hydrogenation, i.e., C{#=-}C, is an internal bond of two hydrogenation reactions: 1) alkyne to alkane written as C{#-}C and 2) alkene to alkane written as C{=-}. The example of a two-step reaction is the first extension of the single reaction shown earlier. At the user's discretion, this flexible bond species can be extended to include additional letters to define third, fourth, etc. reactions. Compared to the previous results, it is clear that the reaction generator has a higher success rate in generating valid reactions for these reactions, even though it had to consider multi-step reactions. The main difference is the reduced diversity in the set of molecules, namely, all molecules used in this dataset are aliphatic alkanes, alkenes, and alkynes, while the substitution dataset includes both aromatic and aliphatic compounds. Figure 24 summarizes the generation results for three runs of 180 generated reactions and three runs of 180 generated examples. This reduced diversity in the set of molecules can also be seen along with a reduced level of creativity. In fact, the set of new reactions proposed is very limited. Nevertheless, the generator demonstrates the generation of new chemistry and hypothesizes new reactions. First, it can be seen that the generator generated molecules with multiple reaction sites, e.g., molecules with "{#=-}.{#=-}" which define molecules with two triple bonds (Figures 27C-27D). This is noteworthy because the model was trained on a dataset consisting of single sites. Next, the generator introduced equilibrium reactions, such as "C{-=-}C" and "C{#=#}C".These SiteCRSs define equilibrium reactions for the dehydrogenation of alkanes to alkenes and the hydrogenation of alkynes to alkenes (Figures 27A-27B). This network is open to accommodate any type of one-step, two-step, or multi-step reaction. The equilibrium generated by the neural network (Figures 27A-27B) is a special type of two-step reaction, and therefore can be handled using the CRS format.

[0190] Figure 27 shows the generation of a new reaction for a multi-step reaction. The above reaction is not shown in the dataset. This example includes the generation of two equilibrium reactions (A and B) and two multi-step reactions, even though the training only included a single reaction. A) Equilibrium for alkane-alkene dehydrogenation. B) Equilibrium for alkyne-alkene hydrogenation. C) Hydrogenation reaction at two sites. D) Dehydrogenation reaction at two sites.

[0191] In other embodiments targeting reaction generation, an AI algorithm can be configured for mining the chemical space. This introduces a statistical inspection mechanism to select the earliest possible stage of a model that ensures diversity is maintained and reliably writes the chemistry. The same algorithm can be applied to generate reactions such as those disclosed above.

[0192] The main advantages of the generated "CRS" are 1) the products that can be extracted from the generated CRS, and 2) the reagents that can be extracted from the generated CRS. The generated pathway can then be examined to see if it is possible to use the existing starting material.

[0193] Therefore, the main advantage is that the products and pathways are generated by a single synthesis. The possibility of generating a reaction rather than a single molecule is quite different from the current approach. In the current approach, 1) generate / define the molecule; 2) consider possible synthesis.

[0194] The application "Generation" includes CRS generation. Patents for this application must be defended, and additional patent applications may need to be considered as backup. It has been shown that CPU computers can rapidly mine chemical spaces, and that chemical space mining is an essential tool for identifying new molecules (public data sources, e.g., PubChem, provide very few candidate molecules).

[0195] In other embodiments targeting regression and classification applications, the applications defined below apply to both the prediction of the molecule itself and the generated CRS string. The CRS defines the reaction that produces the molecule of interest. Consequently, any predictive target for a molecule is also a matter of interest for prediction in the CRS.

[0196] for example, - Regression / Classification of Renewable Carbon: From the proposed reaction, the algorithm will determine whether the pathway is a renewable carbon pathway. If all starting materials used are "renewable," the product can be said to be "renewable." The higher the content, the better the future acceptance.

[0197] - Regression / Classification of Enzyme Reactions: Regression / classification can predict whether a reaction is an enzymatic reaction. The advantage of enzymatic reactions is that the products are considered "natural." This will also likely drive future acceptance.

[0198] - Regression / classification of reaction yield: Even when reaction yields are rarely reported, it is possible to roughly estimate whether the reaction will work.

[0199] - Regression / classification of thermodynamic properties and transition states: Such predictions are energy predictions that may be useful in determining the ease of synthesis or the yield of synthesis.

[0200] - Regression / classification of relevant targets for olfaction or gustatory perception: For the generated products, it may be possible to identify: 1) whether the product can be introduced to market ("evaluation fate"); 2) olfactory descriptors; 3) relevant sensory and physicochemical properties, e.g., odor detection threshold, odor value, Henry, solubility, logP, volatility and / or vapor pressure; 4) activity on olfactory receptors; 4) gustatory receptor activity (e.g., allosteric modulators that enhance sweetness); 5) top-heart-based annotation classification: this is a metric that defines intensity. The mechanism for prediction can be varied and may include knowledge-based methods in chemical formats, classical machine learning methods and deep learning methods.

[0201] - Regression / classification of MS or NMR spectra: The identity of new molecules can be confirmed using predictions of MS and NMR spectra.

[0202] - Regression / Classification for Impurity Prediction: Such applications can be useful in predicting impurities produced by a reaction and their quantities. Here, we primarily consider mixtures of stereoisomers (e.g., R-limonene or S-limonene) and positional isomers (para-lyral and meta-lyral). However, the prediction algorithm can also predict other impurities produced.

[0203] - Regression / classification for predicting hazards: Here, the stability of the product, all types of toxicity, and all types of accumulation (soil, water, etc.) need to be evaluated.

[0204] - Regression / Classification of Production Costs: This method examines the generation of reactions.

[0205] - Regression / Classification of Changing Bonds: Predict changing bonds in the product (SMILES in => CRS out) to obtain reaction predictions. Such predictions can be reinforced (reinforcement learning) with quantitative rewards for any of the following characteristics: 1) ingredients in the market, 2) renewable carbon, 3) enzymatic reactions, or 4) high-yield reactions. In reinforcement learning, rewards are given for particularly good solutions that meet several selection criteria.

[0206] As should be understood, any embodiment can be used to encode, classify, or generate any one of the following non-limiting list of chemical reactions: [Table 4-1] [Table 4-2] [Table 4-3] [Table 4-4] [Table 4-5] [Table 4-6] [Table 4-7] [Table 4-8] [Table 4-9] [Table 4-10] [Table 4-11] [Table 4-12]

Claims

1. A chemical reaction coding program for one-step reactions, multi-step reactions, and equilibrium reactions, - A first step (105) of receiving a chemical reaction graph including at least one chemical reaction reagent and at least one chemical reaction product, - A second step (110) of encoding the chemical reaction graph describing the structure of at least one reagent and the product into a string of characters, - A third step (115) to determine the changing bonds within the string representing the chemical structure of at least one reaction reagent and the product, - A fourth step (120) of encoding, for at least one determined changing bond, at least one character representing the atom undergoing the bond change, at least one character representing the type of determined changing bond, and at least one character representing the atom resulting from the bond change into a single string, - A fifth step (125) of outputting a string of characters corresponding to the encoding of the changing bonds in the chemical reaction, Have the computer run it, A program characterized in that, in the fourth step (120), the changing bond is encoded by a set of two characters representing the determined changing bond, the first character representing the reagent bond and the second character representing the product bond, and each character is selected from a library of bijective characters, each character representing one type of bond.

2. The program according to claim 1, wherein the fourth encoding step (120) is configured to embed two characters representing the determined changing combination between two neutral tag characters representing the existence of the encoding of the changing combination.

3. The program according to claim 1 or 2, wherein a multi-step reaction represented by a series of bond changes between two atoms is encoded by a series of single characters, each single character representing a continuous state of bonding between the two atoms, and the order of the characters represents the order of bond changes between the two atoms.

4. A method for encoding chemical reactions for one-step reactions, multi-step reactions and equilibrium reactions (100), - A first step (105) in which a computer receives a chemical reaction graph including at least one chemical reaction reagent and at least one chemical reaction product, - A second step (110) in which the computer encodes the chemical reaction graph describing the structure of the at least one reagent and the product into a string of characters, - A third step (115) in which the computer determines the changing bonds in the string representing the chemical structure of the at least one reaction reagent and the product, - A fourth step (120) in which the computer encodes into a single string at least one character representing the atom undergoing the bond change, at least one character representing the type of bond that has been determined to change, and at least one character representing the atom resulting from the bond change, for the determined at least one changing bond, - A fifth step (125) in which the computer outputs a string of characters corresponding to the encoding of the changing bonds in the chemical reaction, Includes, Method (100), wherein in the fourth step (120), the changing bond is encoded by a set of two characters representing the determined changing bond, the first character representing the reagent bond and the second character representing the product bond, and each character is selected from a library of bijective characters, each character representing a change in one type of bond.

5. The method (100) according to claim 4, wherein a multi-step reaction represented by a series of bond changes between two atoms is encoded by a series of single letters, each single letter representing a continuous state of bonding between the two atoms, and the order of the letters represents the order of bond changes between the two atoms.

6. The method (100) according to claim 4 or 5, wherein the second encoding step (110) is configured to encode the chemical reaction graph into a row notation, and the method further includes a step (130) in which the computer extends the row notation encoding before the fourth encoding step (120).

7. The method (100) according to claim 4 or 6, wherein the fourth coding step (120) includes the step (121) of the computer extracting a reagent-product binding table from computer memory, and the coding is performed as a function of the binding table.

8. The method (100) according to any one of claims 4 to 7, wherein the fourth encoding step (120) includes the step of the computer removing at least one atomic identifier from at least one reagent and / or product from a first string obtained from the second encoding step (110), and each atom is removed as a result of the third determining step (115) if the atom and associated bond are located in the product and / or reagent which remain unchanged from the reagent reaction step to the product step of the chemical reaction.

9. The method (100) according to any one of claims 4 to 8, comprising the step (135) of obtaining an encoded chemical reaction product by performing the chemical reaction within a physical device.

10. A recording medium on which an encoded chemical reaction containing the string (205, 210) is recorded, A recording medium characterized in that the encoded chemical reaction is obtained by the method (100) described in any one of claims 4 to 9.

11. A method for extending a chemical reaction dataset (300), - A step (305) in which the computer receives a string obtained by the method described in any one of claims 4 to 9, - The computer rearranges the strings (310) in order to shift at least one string of characters, each string consisting of at least one character representing an atom and at least one character representing a bond change associated with the corresponding atom. - The computer outputs an extended string corresponding to the response initially encoded by the received string (320), An extension method (300) characterized by including the following.

12. The extension method (300) according to claim 11, further comprising the step (315) of the computer relating at least two strings obtained by the method according to any one of claims 4 to 9, wherein each of the strings represents the same chemical reaction graph.

13. A method for preprocessing a chemical reaction dataset (400), - The computer receives a dataset of at least two chemical reaction graphs, each containing at least one chemical reaction reagent and at least one chemical reaction product (405), - The computer compresses at least two chemical reaction graphs according to the method described in any one of claims 4 to 9 (100), - The computer determines the distribution of chemical reaction classes within the encoded dataset (410), - The computer extends the dataset (300) as a function of the determined distribution according to the method of claim 7 or 8, for at least one chemical reaction class, - The computer outputs the pre-processed dataset (415), A pretreatment method (400) characterized by including the following.

14. A training method (500) for a classifier, transformer or regressor, - A step (505) of inputting a dataset of chemical reaction graphs encoded by the method described in claim 8 onto a computer interface, - A step (510) of operating a recurrent neural network architecture on a computer system, which is configured to classify the progression of chemical reaction bonds as a function of the input, using the aforementioned chemical reaction graph dataset as input. - A step (515) of outputting a trained classifier, transformer, or regressor on a computer interface. A training method (500) characterized by including the following.

15. A method for predicting the progression of chemical reaction bonds, A prediction method characterized by operating a classifier, transformer, or regressor obtained by the method (500) described in claim 14.

16. A method for generating chemical reactions, A method for generating a chemical reaction, characterized by operating a classifier, transformer, or regressor obtained by the method (500) described in claim 14.

17. A computer-implemented classifier, A classifier characterized in that the classifier, transformer, or regressor is obtained by the method (500) described in claim 14.

18. It is a computer program, A computer program characterized by including instructions for causing a computer to perform the method (500) described in any one of claims 14 to 16.

Citation Information

Patent Citations

  • Processing method for chemical reaction information

    JP1987057017A