Systems for chemistry tasks based on artificial intelligence methods with improved reasoning using chemistry feedback

By integrating a Chemistry Feedback Dataset and aligning AI systems with chemical reasoning, the trustworthiness of AI predictions for chemical reactions is enhanced, addressing the limitations of existing AI systems in chemistry applications.

WO2025215117A1PCT designated stage Publication Date: 2025-10-16MOLECULE ONE SP ZOO
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/059797
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-10
Filing Date
2025-04-09
Publication Date
2025-10-16

AI Technical Summary

Technical Problem

The lack of trust in artificial intelligence (AI) predictions for chemical reactions due to potential hallucinations and insufficient rationale for predictions limits its impact on chemistry applications.

Method used

Implementing a Chemistry Feedback Dataset and methods to adapt AI systems by aligning their reasoning with chemistry feedback, incorporating explicit chemical reasoning into AI model training, and using structured workflows to refine models with user interactions and laboratory experiments.

Benefits of technology

Enhances the trustworthiness of AI predictions by providing rationale and evidence, improving the reliability and accuracy of chemical reaction outcomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025059797_16102025_PF_FP_ABST
    Figure EP2025059797_16102025_PF_FP_ABST
Patent Text Reader

Abstract

Methods and systems are disclosed in which the trustworthiness of predictive models, e.g., machine learning models, is enhanced by incorporating feedback regarding training set reactions is used to train the model so that the model is adapted so that subsequent predictions align with or account for the feedback. The feedback may include, e.g., process level reasoning, mechanism level reasoning, outlines of mechanistic reasoning, suggestions of reference reactions, or estimates of a probability of success of a given reaction. The feedback may itself be generated or proposed by a machine learning model. The model may direct an automated laboratory to perform reactions from which feedback is extracted and used to train the model.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS FOR CHEMISTRY TASKS BASED ON ARTIFICIAL INTELLIGENCE METHODS WITH IMPROVED REASONING USING CHEMISTRY FEEDBACKCROSS-REFERENCE TO RELATED CASES

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 632,180, entitled “Systems For Chemistry Tasks Based on Artificial Intelligence Methods Adapted Based on Chemistry Feedback,” filed on April 10, 2024; and U.S. Provisional Patent Application No. 63 / 645,332, entitled “Systems For Chemistry Tasks Based on Artificial Intelligence Methods Adapted Based on Chemistry Feedback," filed on May 10, 2024, both of which are hereby incorporated by reference. This application is related to U.S. Pat. App. No. 17 / 060,765, entitled “ Systems And Method For Designing Organic Synthesis Pathways For Desired Organic Molecules,” filed on October 01, 2020; and to U.S. Pat. App. No. 18 / 048,981, entitled “Systems And Methods For Predicting Outcomes And Conditions Of Chemical Reactions With High Reliability Based On A Highly Diverse And Accurate Dataset,” filed on October 24, 2022, both of which are hereby incorporated by reference.BACKGROUND

[0002] Predicting outcomes of chemical reactions is a central task to a plethora of industries that use chemistry such as drug discovery, agriculture or cosmetics. Consider drug discovery. For every drug in the market, usually thousands have to be synthesized and tested in the laboratory. Any inefficiency in the chemical processes used - for example failing to obtain desired products in the laboratory - can have a significant downstream effect, increasing the prices of goods and services.

[0003] Unfortunately, predicting outcomes of many classes of chemical reactions is hard for humans and computers alike. One approach to predicting reaction outcomes is based on machine learning, and deep neural networks (DNNs) in particular. Such Artificial Intelligence (Al) — and generative Al in particular — has a growing impact on chemistry applications such as synthesis planning, condition recommendation, or de novo drug discovery.

[0004] An issue is that users often lack trust in Al predictions. This limits the impact Al can have on many chemistry applications.

[0005] The lack of trust arises from two main sources: the potential for Al to hallucinate and propose incorrect predictions about chemistry, and Al not providing sufficient evidence or rationale for its predictions, even if those predictions might be correct.

[0006] Thus, what is needed are systems and methods for increasing the trustworthiness of predicted outcomes of chemical reactions.BRIEF DESCRIPTION OF THE DRAWINGS

[0007] The embodiments are illustrated by way of example and not limitation in the figures of the accompanying drawings, in which like references indicate similar elements, and in which:

[0008] FIG. 1 illustrates an embodiment of a synthesis planning system;

[0009] FIG. 2 illustrates an embodiment of an adapted synthesis planning system;

[0010] FIG. 3 is a screenshot from a user interface in an embodiment of a synthesis planning system;

[0011] FIG. 4 is a screenshot from a user interface in an embodiment of a synthesis planning system;

[0012] FIG. 5 is a screenshot from a user interface in an embodiment of a synthesis planning system;

[0013] FIG. 6 is a screenshot from a user interface in an embodiment of a synthesis planning system;

[0014] FIG. 8 is a screenshot from a user interface in an embodiment of a synthesis planning system;

[0015] FIG. 9 is a screenshot from a user interface in an embodiment of a synthesis planning system;

[0016] FIG. 10 is a screenshot from a user interface in an embodiment of a synthesis planning system;

[0017] FIG. 11 is a screenshot from a user interface in an embodiment of a synthesis planning system;

[0018] FIG. 12 is a screenshot from a user interface in an embodiment of a synthesis planning system;

[0019] FIG. 13 is a screenshot from a user interface in an embodiment of a synthesis planning system;

[0020] FIG. 14 is a screenshot from a user interface in an embodiment of a synthesis planning system;

[0021] FIG. 15A - FIG. 15C are a flowchart of an embodiment of an evaluation process for creating chemistry feedback;

[0022] FIG. 16 is a flowchart of an embodiment of a method for refining a machine learning model using chemistry feedback;

[0023] FIG. 17A is a chart showing results from different evaluations of an embodiment of a synthesis planning system;

[0024] FIG. 17B is a chart showing results from before and after fine-tuning an embodiment of a reaction scorer in an embodiment of a synthesis planning system;

[0025] FIG. 18 is an exemplary block diagram depicting an embodiment of a system for implementing embodiments of methods of the disclosure; and

[0026] FIG. 19 is an exemplary block diagram depicting a computing device.DETAILED DESCRIPTION

[0027] 1. Overview

[0028] Embodiments described within provide for an increase in the trustworthiness of predicted outcomes of chemical reactions by introducing Al Systems that use one or both of: 1) a process to create a database of Chemistry Feedback (defined below); or 2) methods for adapting Al Systems for predicting or proposing chemistry tasks that align the Reasoning of the Al System with the Chemistry Feedback.

[0029] This disclosure describes embodiments implementing Chemistry Feedback, a Chemistry Feedback Dataset, and examples of machine-learning approaches to using a system refined by using feedback. In turn, it describes embodiments of systems employing Chemistry Feedback, Chemistry Feedback Datasets, and machine-learning, which include a synthesisplanning system and a Natural Language-based Al Assistant for Chemists. Such embodiments improve user trust by incorporating explicit chemical reasoning into Al model training.

[0030] 2. Features

[0031] Below is a brief summary of features in one or more embodiments:

[0032] Machine Learning Models With Reasoning Adapted by Chemistry Feedback Systems may include machine learning models first trained on a broad dataset of historical chemistry knowledge (e.g., textbooks, patents, historically performed reactions) and then adapted using a smaller, more expensive dataset generated by validating aspects of reasoning of the model.

[0033] Chemistry Feedback GenerationSystems may include machine learning models that can generate or propose (parts of) Chemistry Feedback. For instance, the model might outline its mechanistic reasoning or suggest relevant historical references. A feature includes the training of models to reason “like chemists” when Chemistry Feedback includes process-level or mechanism-level rationale.

[0034] Structured Process for Creating and Using Chemistry FeedbackEmbodiments of the approach may involve designing and enforcing a structured workflow for collecting high-value Chemistry Feedback, typically via specialized user interfaces (Reaction Evaluation UI, Synthesis Planning UI) and using that data to iteratively refine the Al models.

[0035] 3. Definitions

[0036] Chemical Reaction: A fully or partially specified chemical reaction, including at least the main substrates and the main product. A fully specified reaction has sufficient information for a person to carry it out in the lab without further context.

[0037] Chemistry Reasoning: The sequence of steps by which a given Model or Human arrived at the prediction regarding a Chemical Reaction, and / or explanation or evidence supporting the prediction. The majority of historical datasets don't disclose the process by which the prediction has been arrived at, which necessitates developing methods for training models to improve their reasoning. Reasoning may regard or address, for example, one or more of: a reaction mechanism, regioselectivity, chemoselectivity, stereoselectivity, an outcome of a chemical reaction at least one atom or bond removed or added, a relevance of a historical reaction, an electronic effect, chemical correctness, feasibility, stability, toxicity, yield, complexity, synthetic accessibility, practical utility, cost of starting materials or intermediates, reaction temperature, solvent identity, catalyst identity, reagent identity, or reaction atmosphere, mechanistic plausibility, likelihood of elementary reaction steps, or reliability of predicted reaction pathways, chemical impact of functional group addition, removal, or substitution inreactants or products, confidence, uncertainty, or prediction reliability, or chemical similarity to historically validated reactions.

[0038] Chemistry Question: A question about a chemistry problem that reflects a step of the Chemistry Reasoning, often related to a particular Chemical Reaction. Examples include: “Is this reaction mechanism plausible?”; “How likely is the reaction to occur at another site?”; “How confident is the prediction?”; and “Does the fact this historical reaction has worked make this reaction more plausible?”

[0039] Evaluator: An algorithm, (automated) laboratory, or individual capable of evaluating answers to Chemistry Questions (e.g., a human chemist, quantum simulation algorithm). An Evaluator may label predictions as correct, partially correct, or incorrect, and can provide uncertainty estimates.

[0040] Chemistry Feedback: An answer to a specific Chemistry Question plus an evaluation of that answer by an Evaluator. This feedback can include rationale, references, confidence measures, or explicit step-by-step reasoning. Collecting Chemistry Feedback is often expensive or time-consuming, as it may require prompting experts for their judgment or directing experiments to be performed in a laboratory, e.g., a semi-automated high-throughput laboratory.

[0041] Chemistry Feedback Dataset (CFD): A dataset composed of individual pieces of Chemistry Feedback. The CFD can include feedback gathered from human experts, laboratory experiments, or computational models (quantum mechanical simulations, advanced Al systems, etc.).

[0042] Reaction Generator: A machine learning model that generates descriptions of chemical reactions, e.g., proposing sets of substrates for a desired product (retrosynthesis).

[0043] Reaction Discriminator: A machine learning model that predicts the feasibility or yield of a given reaction, effectively classifying whether the reaction is likely to succeed.

[0044] Synthesis Path: A plan for synthesizing a target molecule, preferably starting from commercially available compounds or simpler intermediates, and typically through a series of reaction steps.

[0045] Synthesis Planning: The process (manual or computational) of generating a Synthesis Path for a target compound.

[0046] Reference Reactions: A set of historically performed chemical reactions relevant to a newly proposed reaction or Synthesis Path, used to increase confidence in that proposed reaction.

[0047] Reference Reaction Generator: A component or algorithm that identifies and ranks relevant historical reactions in support of a newly proposed reaction or Synthesis Path.

[0048] 5. Chemistry Feedback Dataset

[0049] In some embodiments, Chemistry Feedback consists of list of predictions regarding a given chemical reaction (e.g., whether the yield is above 10%) associated with one or more pieces of information regarding the reasoning used to arrive at the answer (e.g., prediction regarding occurrence of some elementary steps, confidence level, etc.).

[0050] 5.1. Chemistry Questions

[0051] In some embodiments of this kind, generated reactions are automatically recorded during user interactions. In some embodiments of this kind, users of the applications based on a Reaction Generator are able to select risky or unfeasible reactions using a graphical interface and record specific Chemistry Questions about them in order to be later evaluated by an Evaluator. In some embodiments of this kind, reactions generated by a Reaction Generator are recorded for evaluation only if they pass a filter based on some threshold on a prediction of the Reaction Discriminator of their yield. In some embodiments of this kind, Chemistry Questions can ask whether a given Chemical Reaction follows a plausible mechanism (i.e., is there a rational explanation for why it occurs) or whether there is a related Reference Reaction supporting regioselectivity (the atomic region of a compound at which the reaction takes place). In some embodiments of this kind, the series of Chemistry Questions can include asking the same Chemistry Question about each reaction in the Synthesis Path, for example, the plausibility of each of the reactions.

[0052] In some embodiments, Chemistry Questions about reactions generated by a Reaction Generator include questions about specific chemical properties of a given Chemical Reaction, such as its correctness, reactant stability, reaction yield, usefulness in Synthesis Planning, or toxicity of the reaction substrates or obtained product.

[0053] In some embodiments, Chemistry Questions regard a Synthesis Path generated by a Synthesis Planning System. In some embodiments of this kind, all the Synthesis Paths generated by the system are automatically recorded for evaluation. In some embodiments of thiskind, users of the Synthesis Planning System are able to select risky or unfeasible Synthesis Paths using a graphical interface and record specific Chemistry Questions about them in order to be evaluated by an Evaluator.

[0054] In some embodiments, Chemistry Questions about a Synthesis Path include questions about the chemical feasibility of the reactions in the Synthesis Path, the cost of the starting materials involved in a Synthesis Path, the complexity of a Synthesis Path, or the estimated total cost of performing all the reactions in the Synthesis Path.

[0055] In some embodiments, Chemistry Questions include questions about the reasoning process of a chemist. In some embodiments of this kind, Chemistry Questions include questions to specify the sequence of steps that a chemist used to arrive at a given Chemistry Feedback. In some embodiments of this kind, one of the steps might be to look up relevant historical reactions in some database of chemical knowledge using a graphical interface. In some embodiments of this kind, one of the questions asks about the source of chemical knowledge used in a given step.

[0056] In some embodiments, Chemistry Questions consider a Reference Reaction supporting a Chemical Reaction or a Reference Reaction supporting a Synthesis Path. In some embodiments of this kind, all the Reference Reactions generated by a Reference Reactions Generator for Chemical Reactions or Synthesis Paths during the interaction of a user with an application are automatically recorded for evaluation. In some embodiments of this kind, users of the applications that use a Reference Reaction Generator are able to select untrustworthy or irrelevant Reference Reactions using a graphical interface and record specific Chemistry Questions about them in order to be evaluated by an Evaluator.

[0057] In some embodiments, Chemistry Questions about a Reference Reaction include questions about the applicability of a given Reference Reaction as supporting evidence for a corresponding Chemical Reaction or Synthesis Path, or questions about the chemical correctness of the Reference Reaction itself. In some embodiments of this kind, Chemistry Questions about a Reference Reaction include the evaluation of how much the Reference Reaction increased the trustworthiness of the Chemical Reaction or Synthesis Path that it supports.

[0058] 5.2. Chemistry Feedback

[0059] In some embodiments, the Chemistry Feedback (CF) is generated from a structured process that consists of evaluating predictions of the system for a series of Chemistry Questions.

[0060] In some embodiments, a Chemistry Feedback Dataset (CFD) includes CF that evaluates answers (e.g., predictions) to Chemistry Questions about one or more Chemical Reactions, Synthesis Paths, or Reference Reactions. In some embodiments of this kind, a Chemistry Feedback Dataset includes Chemistry Feedback that evaluates the reasoning process behind a given Chemical Reaction, Synthesis Path, or Reference Reaction.

[0061] In some embodiments, a CFD includes CF that is generated during the interaction of a human Evaluator with a Graphical User Interface designed specifically for evaluating answers to Chemistry Questions about chemical reactions. In some embodiments of this kind, such Graphical User Interface includes a description of the reaction, descriptions of Reference Reactions generated for the reaction, and options for the Evaluator (human or otherwise) to input feedback about the reaction, such as its feasibility or what kind of chemical issues this reaction may have. The responses of the human evaluator are then recorded and saved on a server as part of a CFD. The structure of this process allows for building a high quality CFD that is consistent between humans allows for training better machine learning models.

[0062] In some embodiments, the CFD includes the specification of some or all of the steps (automatic and manual) performed to answer a given Chemical Question. In some embodiments of this kind, CFD includes the specification of how relevant (influencing the provided Chemistry Feedback) chemical knowledge was searched and what information was found in which source of chemical knowledge.

[0063] In some embodiments, CF includes information about the confidence in a given piece of CF. In some embodiments of this kind, an Evaluator is asked about confidence concerning a given piece of CF. In some embodiments of this kind, an automatic measure of certainty is created. In some embodiments of this kind, the automatic measure of certainty is based on provided information about the used chemical knowledge when creating the CF.

[0064] In some embodiments, the CFD includes CF generated by directing chemical reactions or other experiments in a laboratory. In some embodiments of this kind, Chemical Reactions for which an Evaluator assigned high uncertainty to their CF are performed in a laboratory (e.g., by the system 1810a, 1810b (FIG. 18) contacting and directing an automatedlaboratory 1822 (FIG. 18), and a new CF is generated that answers the Chemistry Questions using lab results (e.g., by the automated laboratory supplying the system with test results). In some embodiments of this kind, Synthesis Paths that were assigned high uncertainty by the Evaluator with regards to their CF are performed in a chemical laboratory in order to answer Chemistry Questions about the Synthesis Paths (e.g., by the system directing the automated laboratory and being supplied with the test results).

[0065] In some embodiments in which experiments are directed to be performed in a laboratory, e.g., automated laboratory 1822 (FIG. 18), one or more pieces of specialized equipment are employed to generate empirical data for Chemistry Feedback. By way of nonlimiting example, such equipment may include: (a) a liquid handling robot (e.g., an Opentrons OT-2 system) to measure, mix, and dispense reactants with precision, (b) auto-sampling equipment to facilitate systematic and automated sample collection for downstream analyses, (c) analytical instrumentation, such as an LC / MS machine, for determining reaction yields, side products, and other relevant chemical properties, (d) robotic platform to transfer material between stations, (e) automated warehouse for storage and retrieval of reagents. In certain embodiments, these tools are integrated in a fully automated workflow that orchestrates reagent preparation, reaction execution, quenching, and analytical evaluation with minimal manual intervention. In other embodiments, some or all of these steps may be semi-automated or include manual oversight by a chemist for confirmation and quality control. In some embodiments, data obtained from these automated or semi-automated laboratory processes — covering yields, byproducts, reaction rates, or mechanism-verifying details — are captured as Chemistry Feedback and incorporated into the Chemistry Feedback Dataset (CFD). This integration of high- throughput, experimentally validated results bolsters the reliability of the Al-driven models and further refines their predictive accuracy through iterative retraining and adaptation.

[0066] In some embodiments, the CFD includes CF fully or partially based on computational methods. This approach enables improving aspects of reasoning that might be challenging to improve using laboratory or human feedback, e.g., due to the expensive nature of the required experiment, complexity of the problem (e.g., nuanced prediction of energy of a transition state). In some embodiments, a quantum mechanical simulation is performed to determine the Chemistry Feedback (e.g., the correctness of a given Chemical Reaction). In some embodiments of this kind, quantum mechanical simulation is performed on a reaction for whichCF with high uncertainty has been generated by an Evaluator. In some embodiments of this problem, a quantum calculation package (such as ORCA or Schrodinger) is used to calculate the energy of a transition state that is relevant for the reactivity that is addressed by a given Chemical Feedback.

[0067] In some embodiments, the CF includes a new set of Chemical Reactions with associated information about their success, yield, or other chemical properties. In some embodiments of this kind, the new set of Chemical Reactions is used as evidence to answer a Chemical Question about another Chemical Reaction or a Synthesis Path.

[0068] 6. Methods for Adapting a Machine Learning Model using the ChemistryFeedback Dataset

[0069] In some embodiments, a machine learning model that will be adapted based on Chemistry Feedback (CF) is first trained on historical chemistry knowledge such as patents, textbooks, scientific publications, laboratory experiments summaries, or other sources of chemical information. In some embodiments of this kind, the training of a machine learning model includes the task of generative modeling of (e.g., autoregressive modeling of a reaction as a list of characters or tokens) Chemical Reactions that are present in the historical dataset. In some embodiments of this kind, the task of generating Chemical Reactions can be formulated as: a task of generating reaction product based on reaction substrates (forward synthesis prediction); a task of generation reaction substrates based on reaction product (retrosynthesis prediction); or a task of generating de novo a feasible reaction from scratch (generative reaction modeling, which may be accomplished by training a model to generate reactions conditioned on random noise, and sampling the noise when using the model). In some embodiments of this kind, a Chemical Reaction is represented in a way that can be processed by a machine learning model within the class of graph-based neural networks (reaction is represented using the adjacency matrix encoding substrates and product) or sequence-to-sequence models such as Transformer (in which a reaction is represented as a sequence of tokens or characters). In some embodiments of this kind, Chemical Reactions are represented as transformations of graphs of atoms and bonds. In other embodiments of this kind, Chemical Reactions are represented as a transformation between textual representations of substrates, products, and other information that specifies a particular Chemical Reaction. In other embodiments of this kind that perform generative reaction modeling, a Chemical Reaction is generated using a diffusion model that produces a reaction insequential diffusion steps starting from an empty reaction graph or an empty sequence, which enables faster generation due to parallel (occurring at all nodes of the graph) generation of the Chemical Reaction.

[0070] In some embodiments, the Chemistry Feedback Dataset (CFD) is used to further adapt a machine learning model for chemistry (starting from the weights of the model trained on historical data). In some embodiments of this kind, a Reasoning Scoring Function is defined that evaluates the agreement between model outputs and Chemistry Feedback. In some embodiments of this kind, the scoring function is differentiable which allows the application of standard training procedures. In some embodiments of this kind, a search procedure (e.g., a random search) is performed to identify parameters (e.g., weights that control interaction of two modules) that configure the model to achieve a better value of the scoring function. In some embodiments of this kind, a reinforcement learning procedure (with reward configured to result in optimization of the scoring function) is used to adjust weights, which is particularly useful in cases when the scoring function is nondifferentiable. In some embodiments, a mixture of different approaches for adapting is used, which can be for example useful when the scoring function can be divided into a part that is differentiable and a part that is not. In some embodiments of this kind, a Reaction Generator, a Reaction Discriminator, or a Reference Reaction Generator may be adapted in this fashion based on CFD.

[0071] In some embodiments, for the purpose of adapting the model, a process is used to determine or approximate the reasoning process that the model uses, output of which is used to determine how to adapt the machine learning model to better agree with Chemistry Feedback. In some embodiments, the Chemistry Feedback is converted into virtual chemical reactions that are input into the model and the outputs are used by a mathematical function to quantify the consistency of the model reasoning with the Chemistry Feedback; for example, if an evaluator deems a reaction unlikely to proceed due to steric hindrance, a number of virtual reactions characterized by steric hindrance could be created and input into the model. In some embodiments, the base machine learning model is based at least in part on a language model and can be directly prompted to output or approximate the reasoning process used. In some embodiments of this kind, the machine learning model is configured in part by weights that influence aspects of reasoning quantified in the Chemistry Feedback and these weights are tuned towards improving the value of the Reasoning Scoring Function. In some embodiments of thiskind, by testing the model on imaginary Chemical Reactions, the reasoning process is inferred before adapting the model. In some embodiments of this kind, the model is tested on Chemical Reactions which are modified to contain a functional group so that the impact of the presence or absence of the functional group is quantified (and then Chemistry Feedback might contain examples quantifying how important the functional group is for the output).

[0072] In some embodiments of this kind, a CFD is converted to a format that allows continuing the same training process as previously but on a new dataset, for example by extracting all or some of Chemical Reactions from Chemistry Feedback that have an associated label indicating correctness and continuing training of the Reaction Discriminator on those and potentially historical reactions, which has the benefit of making the optimization process more efficient by reducing the need of adapting the model to handle a new type of task. In some embodiments of this kind, the machine learning model is reconfigured to train on a new dataset in the second stage of the training (e.g., see FIG. 2 and related text). In some embodiments of this kind, parts of the historical chemistry knowledge dataset are additionally included in the second training stage so that the model avoids forgetting the historical chemistry knowledge.

[0073] In some embodiments, an ensemble of machine learning models is trained on different types of CFDs to be able to answer specific Chemistry Questions, such as chemical reaction feasibility, molecule stability, toxicity, reaction yield, or usefulness of a chemical reaction in synthesis planning. In some embodiments of this kind, each model from the ensemble is first trained on the historical chemistry knowledge dataset and further trained in a second stage on a CFD about a specific Chemistry Question. In some embodiments of this kind, a final model is created by combining predictions of the models for specific Chemistry Questions as a weighted average.

[0074] In some embodiments, a machine learning model is trained to provide the reasoning process, which, in particular, facilitates adapting it to better align with the desired reasoning encoded in CFD. In some embodiments of this kind, a model is trained in two stages, first, on the historical dataset of chemistry knowledge, and second on a CFD with the reasoning process provided by Evaluators. In other embodiments of this kind, a model is trained to develop a reasoning process in an unsupervised manner (i.e., without direct supervision on what reasoning is expected in Chemistry Feedback) so that the developed reasoning process maximizes the model accuracy on a target task. The reasoning process can include a series ofadditional Chemistry Questions that a human Evaluator answered to better figure out the answer to a specific Chemistry Question, such as: “What reference reactions can I find to learn this information for this reaction?”, or “Is there a place in the substrates where the reaction would occur more likely,” etc.

[0075] In some embodiments, a Language Model (e.g., a Large Language Model (LLM)) is trained (e.g. autoregressively) in a second stage on the CED, which in particular allows to directly answer natural language questions about chemistry. In some embodiments of this kind, the LLM training includes a stage in which the CFD is expressed in natural language. In some embodiments of this kind, the LLM is trained on examples of reasoning processes recorded in a CFD to provide reasoning behind its predictions using natural language.

[0076] In some embodiments, the machine learning models trained on a CFD can be used to generate synthetic CF for other models in the form of virtual reactions and / or natural language explanations, which leverages the abilities of the machine learning model to infer a broader set of CF. In some embodiments of this kind, a Language Model is used to generate the synthetic CF.In other embodiments of this kind, an ensemble of many machine learning models is used to evaluate predictions of a single model.In some embodiments, the historical data of chemistry knowledge or a CFD can be used to evaluate a machine learning model for chemistry. In some embodiments of this kind, an algorithm automatically extracts from a dataset of historically performed chemical reactions a set of reaction templates that describe transformations that have been historically performed and evaluates reactions generated by a Reaction Generator based on the existence of the corresponding reaction template in this set. In some embodiments, reactions are scored by an Evaluator to be probably infeasible when all of the reactions with the same reaction template in CFD were evaluated to be chemically infeasible.

[0077] 7. Al Software Systems based on Chemistry Feedback Dataset

[0078] 7.1. Synthesis Planning based on Chemistry Feedback Dataset

[0079] In some embodiments, Software for Synthesis Planning (SFSP) provides outputs, e.g., Synthesis Paths, using one or more machine learning models trained based at least in part on a Chemistry Feedback Dataset (CFD). In some embodiments of this kind, a machine learning model included in the SFSP (e.g., in RRGs 120, 220 and RSs 124, 224) is first trained based onhistorical data of chemistry knowledge. For example, machine learning models in RRG 120 and RS 124 are first trained on historical data of chemistry knowledge, e.g., Reaction Database 110.

[0080] In some embodiments, the SFSP is based on one or more modules for: generating reactions (Reaction Generator), scoring reactions (Reaction Scorer), and a search algorithm (Search), such as A* or MCTS, that uses these modules to find a Synthesis Path. In some embodiments of the first stage of the training (e.g., FIG. 1 and related text), the Reaction Generator and Reaction Scorer are trained on the historical data of chemistry knowledge. This provides the modules required to operate the first version of the Synthesis Planning system. In some embodiments, a CFD geared towards improving modules comprising the system can be built by generating Synthesis Paths for a selected test set of compounds, extracting Chemical Reactions, Synthesis Paths, and / or Reference Reactions from these generated Plans, and evaluating their correctness by Evaluators.

[0081] In some embodiments, SFSP outputs are also based on a module capable of generating the whole or partial synthesis tree without performing an explicit search. In some embodiments, SFSP outputs are also based on a module that can predict the price of synthesis, or another score relevant to the software, of compound or synthesis (sub)tree.

[0082] In some embodiments, SFSP outputs include reference (i.e., historical) reactions that accompany other predictions of the SFSP concerning reactions (Reference Reaction Generator). In some embodiments of this kind, the Reference Reaction Generator is trained in part on Chemistry Feedback. In some embodiments of this kind, the Chemistry Feedback includes scores of usefulness of reference reactions for some chemical reactions. In some embodiments of this kind, Reference Reaction Generator is trained or adapted to suggest Reference Reactions that would be historically scored highly by evaluators. In some embodiments, the Reference Reaction Generator outputs a text that indicates the reason why a given reference reaction was given. In some embodiments of this kind, Reference Reaction Generator is trained to predict Chemistry Feedback that includes examples of reasons why a given reference reaction was given.

[0083] In some embodiments, an optimization method is used to determine optimal configuration parameters of one or more of: Reaction Generator, Reaction Discriminator, Search, or other modules that determine the output of software that maximizes the alignment between the outputs of the SFSP and the CFD, which results in synthesis plans that are found more likely tosucceed to humans and / or are more likely to succeed in general. In some embodiments, the configuration includes one or more of: the relative penalty for synthesis plan depth (which determines the preference of the search to find shallow synthesis trees), the relative penalty for reaction infeasibility, etc. In some embodiments of this kind, different sets of parameters are searched (e.g., through enumeration or random search) and each set of parameters is evaluated using Chemistry Feedback Dataset. In some embodiments of this kind, outputs of SFSP configured using these sets of parameters are used to develop Chemistry Questions. In some embodiments, this configuration is selected so that the incorrect reactions in CFD occur in generated Synthesis Paths rarely, while the correct reactions from CFD occur in the generated Synthesis Paths more often. In some embodiments, a CFD contains fully evaluated Synthesis Paths, and in some embodiments of this kind the Scoring Function penalizes Synthesis Paths that were given low evaluation scores (in the CFD as one of the CQ) and promotes Synthesis Paths that were given high evaluation scores.

[0084] In some embodiments, configuration parameters include parameters that steer the behavior of the search, such as the way the price of compounds is estimated during the search. In some embodiments, each reaction generated by Reaction Generator is given a score by the Reaction Generator, and is also additionally scored by the Reaction Discriminator. In some embodiments of this kind, SFSP is parametrized by a weight that controls how strongly the output of the Reaction Discriminator influences the designed pathway and this weight is optimized based on the CFD. In some embodiments, these scores are combined using a function (a Joining Function) to obtain a single reaction score. In these embodiments, a CFD is used to find the optimal Joining Function so that the incorrect reactions are given the lowest possible scores and correct reactions are given the highest possible scores. In these embodiments, the Joining Function is found within different types of mathematical functions, including but not limited to the such as weighted average, harmonic mean, geometric mean, etc.

[0085] In some embodiments, the user can influence the weights of the system that impact the Scoring Function. In some embodiments of this kind, the user can decide on the relative weight in predictions of a model trained on historical data only and a model trained also on a CFD, which allows balancing between a system that is more grounded in historical data and a system that can address specific preferences of the evaluators that gathered the CFD.

[0086] In some embodiments, the Synthesis Planning software outputs include conditions required to perform such a reaction, such as choice of catalysts, solvents, temperature, or reaction environment. In some embodiments, these suggestions are done by a machine learning model trained on historical data and / or a CFD to predict reaction conditions or to predict the yield of a reaction when reaction conditions are specified. In some embodiments, the model for predicting reaction conditions is trained in a two-stage fashion as described in Section 5 and with reference to FIG. 1 and FIG. 2, first training on historical data of reactions with conditions, and second training on a CFD about reaction conditions proposed by the model trained only in a single stage.

[0087] 7.2. Natural Language-Based Al Assistant for Chemists

[0088] In some embodiments, a natural language processing system based on a Large Language Model (LLM), capable of answering chemistry-related questions stated in natural language, is trained on a dataset of historical chemistry knowledge and a Chemistry Feedback Dataset (CFD). In some embodiments, an LLM is first trained on a dataset of natural language texts, using standard LLM training techniques. The model might be then trained in a second stage on a mixture of natural language texts and historical chemistry knowledge, represented as text, to incorporate historical chemistry knowledge better, while retaining natural language abilities. In the third stage, the model might be trained on a mixture of natural language texts, historical chemistry knowledge, and CFD(s). In some embodiments, the CFD(s) used to improve the LLM could be gathered from CF or a CFD from a different product, such as a different SFSP, and repurposed to improve chemistry-related abilities for the LLM by, for example, finetuning the LLM on CFD represented as natural language text.

[0089] In some embodiments, an LLM-based Al Assistant for Chemists is refined to improve reasoning using methods from Section 6, which allows for the Assistant to be a conversational engine (of the kind of ChatGPT) that is better aligned with how chemists reason and more accurate in general, which addresses the technical problem of hallucination present in LLMs. In some embodiments, a CFD specific for an LLM-based Al Assistant is gathered. In some embodiments of this kind, the Al Assistant is deployed for a human user to interact with it as a chatbot, and the user asks chemistry-related questions to the Assistant and rates the responses. The questions and answers, together with user ratings, are then recorded as a CFD. In some embodiments, this CFD is then used for multi-stage improvement of LLMs described in the previous paragraphs.

[0090] In some embodiments, the Assistant based on CFD is accessible through a GUI or an API of a chemical software, such as a Synthesis Planning system, which allows the user to use the features of the system through a human-like interface of speaking or chatting with the agent who is directed to use features of the system. In some embodiments of this kind, the assistant is adapted to support executing on behalf of the user any functionality that the SFSP supports. In some embodiments of this kind, the assistant is adapted using any standard adaptation techniques used in the field.

[0091] 8. Example Implementation of Synthesis Planning Software Adapted Based onChemistry Feedback

[0092] This section describes an exemplary embodiment of an Al-based synthesis planning system adapted based on Chemistry Feedback.

[0093] 8.1. Main system components

[0094] FIG. 1 illustrates and embodiment of a system 100 and method for synthesis planning. In FIG. 1, SFSP 102 consists of two main User Interface components: 1) a Synthesis Planning UI 118, which is a Graphical User Interface for submitting queries for generating synthesis plans for target compounds, as well as browsing and reading the generated results; and 2) a Reaction Evaluation UI 104, which is a Graphical User Interface for submitting Chemistry Feedback about example reactions generated by the system.

[0095] In the embodiment, SFSP 102 is built using the following databases: 1) a Reaction Database 110, which is a database of historically performed reactions, which can contain publicly available descriptions of reactions from literature or patents and reactions from proprietary collections; and 2) a Starting Materials Database 114, which is a database of purchasable compounds that are available on the market for a reasonable price. SFSP is used to produce a third database - a Chemistry Feedback Dataset (or database) (CFD) 106, which is a database of Chemistry Feedback that is gathered for use with the models used in the system, using some of the methods described .

[0096] SFSP 110 consists of the following computational components: 1) a Retrosynthesis Reaction Generator (RRG) 120, a type of Reaction Generator that proposes the main substrates of a reaction for a given reaction product (an RRG generates a ranked list of candidate sets of substrates with scores between 0 and 1, estimating their conditional probability based on the input product); 2) a Reaction Scorer (RS) 124, a type of Reaction Discriminator thatassigns a reaction with a score between 0 and 1 which estimates the probability that the reaction has a sufficiently high yield to be useful for the purpose of synthesis planning; 3) a Reference Reaction Generator 126, an algorithm that, for a given input reaction, outputs a set of reference reactions from the Reaction Database 110 that are similar to the input reaction and can support its correctness; 4) a Single-Step Reaction Generator (SSRG) 112, which combines RRG 120, RS 124, and Reference Reaction Generator 126 to generate, for an input product, a ranked list of candidate sets of substrates (which are assigned a Combined Score between 0 to 1), together with a set of Reference Reactions for each set of substrates (the Combined Score is obtained from the RS and RRG scores using a Joining Function 122, which is found using Chemistry Feedback. The details of Joining Function 122 are described in Section 8.6.2); and 5) a Synthesis Planning Search (SPS) 116, a graph search algorithm that uses an SSRG (Query 1.5 136 in FIG. 1) and a Starting Materials database 114 (Query 1.4 134 in FIG. 1) in order to find a Synthesis Path that leads from an input molecule to the Starting Materials. To efficiently navigate the vast chemical space and identify viable synthetic pathways, the SPS 116 algorithm iteratively applies the SSRG 112 to generate potential reactants, expanding the search space while prioritizing pathways with reactions with high Combined Score and low cost of Starting Materials, until a complete synthetic route to Starting Materials is identified. In other words, until a Synthesis Path is identified.

[0097] FIG. 1 shows how the initial version of the synthesis planning system 100 is built and used to obtain Chemistry Feedback Dataset 106. In FIG. 2, we show how the synthesis planning system 100 is adapted into a synthesis planning system 200 using the Training Chemistry Feedback Dataset 108. In the next sections, we will describe and refer to the steps depicted in the diagrams (for example: Train 1.1 128, etc.).

[0098] FIG. 1. The initial version of the Synthesis Planning System 100 is deployed and accessible by the users using the Synthesis Planning UI 118. The users can export selected results to be evaluated using Reaction Evaluation UI 104. From the evaluation, a Chemistry Feedback Dataset 106 is generated, which can be used to improve the system.

[0099] FIG. 2. illustrates an embodiment of process for how a Synthesis Planning system 200 may be created by refining system 100 based on TCFD 108. The system can be further used to create a refined Chemistry Feedback Dataset 206, which can be used subsequently to improve upon the previous versions of the system using a refined training chemistry dataset (RefinedTCFD 208). Further iterations of the process depicted in FIG. 2 may be used to iteratively refine the system as many times as desired.

[0100] In FIG. 2, system 100 has been refined into system 200 by using TCFD 108 to adapt the elements of SFSP 102 into SFSP 202. SFSP 202 consists of two main User Interface components: 1) Synthesis Planning UI 118; and 2) Reaction Evaluation UI 104. In the embodiment, SFSP 202 is built using the following databases: 1) Reaction Database 110; 2) Starting Materials Database 114; and 3) Training Chemistry Feedback Dataset (or database) (TCFD) 108. SFSP 110 consists of the following computational components: 1) Retrosynthesis Reaction Generator (RRG) 220, a refined Reaction Generator that proposes the main substrates of a reaction for a given reaction; 2) a Reaction Scorer (RS) 224, a refined Reaction Discriminator that assigns a reaction with a score between 0 and 1 which estimates the probability that the reaction has a sufficiently high yield to be useful for the purpose of synthesis planning; 3) a Reference Reaction Generator 126, an algorithm that, for a given input reaction, outputs a set of reference reactions from the Reaction Database 110 that are similar to the input reaction and can support its correctness; 4) a refined Single-Step Reaction Generator (SSRG) 212, which combines RRG 220, RS 224, and Reference Reaction Generator 226 to generate, for an input product, a ranked list of candidate sets of substrates (which are assigned a Combined Score between 0 to 1), together with a set of Reference Reactions for each set of substrates (the Combined Score is obtained from the RS and RRG scores using a refined Joining Function 222, which is found using Chemistry Feedback. The details of Joining Function 222 are described in Section 8.6.2); and 5) a Synthesis Planning Search (SPS) 116, a graph search algorithm that uses an SSRG (Query 2.7 in FIG. 2) and a Starting Materials database 114 (Query 2.6234 in FIG. 2) in order to find a Synthesis Path that leads from an input molecule to the Starting Materials. To efficiently navigate the vast chemical space and identify viable synthetic pathways, the SPS 116 algorithm iteratively applies the SSRG 212 to generate potential reactants, expanding the search space while prioritizing pathways with reactions with high Combined Score and low cost of Starting Materials, until a complete synthetic route to Starting Materials is identified.

[0101] 8.2 First Stage of Model Training

[0102] The system for Synthesis Planning 100 is first built using models trained on the Reaction Database 110 as shown in FIG. 1. System 100 is refined into Systems 200 as shown in FIG. 2. In the following, description may apply to steps from both FIG. 1 and FIG. 2.

[0103] The Retrosynthesis Reaction Generators (RRG) 120, 220 are trained on a dataset of reactions 110 for a conditional generation task in order to predict reaction substrates from product (Train 1.1 128 and Train 2.1 228). Two different models may be used, which can be chosen by the user in the Synthesis Planning UI 118. A Transformer model may be chosen, which is based on the standard Transformer architecture that represents reaction generation as a translation between input and output SMILES. A METRO model (Molecule-Edit Templates for RetroSynthesis) may be chosen, which learns to predict reactions represented as reaction templates, which encode all the transformations within the reaction as molecular graph modifications.

[0104] Two different model architectures may be used for the Reaction Scorers (RS) 124, 224: A Graph Attention Network (GAT) and a Reaction Prior. A GAT is a type of Graph Neural Network, may be trained on Reaction Database 110 as positive reaction samples and artificial negative reactions based on reaction templates applied to reactions from Reaction Database 110, which is a standard procedure in training reaction scoring models in the literature (Train 1.2 130 and Train 2.2230). The GAT model is particularly strong in capturing the structural relationships between reactants and products in chemical reactions, leveraging graph-based representations to model how atoms and bonds interact in a reaction. This allows it to learn complex chemical patterns and make more accurate predictions. Furthermore, the GAT can process both the direct chemical context and indirect interactions, offering a more holistic understanding of reaction mechanisms.

[0105] The second Reaction Scorer RS 124, 224 option, called Reaction Prior, is a decoder-only Transformer network that models reaction probability from scratch by autoregressive decoding of SMILES tokens in a manner similar to models used for Large Language Models (Train 1.2 130 and Train 2.2230). The Reaction Prior is trained exclusively on Reaction Database 110, utilizing only positive reaction data. This approach has the advantage of learning the likelihood of reactions without the need to generate negative examples. By focusing solely on positive reactions, the Reaction Prior avoids the potential issues of noise or misrepresentation that can arise when generating artificial negative data. This leads to a more accurate and reliable model that is trained on the true distribution of reactions. The model benefits from an autoregressive training process, which allows it to capture the complex patterns in chemical reactions, making it particularly effective for reaction prediction tasks. Additionally,because it does not require negative data, the Reaction Prior simplifies the training process, resulting in faster training times and reducing the risk of overfitting to imperfectly generated negative examples.

[0106] The final Reaction Scorer (RS) 124, 224 may be an ensemble of these two models, with equal weights in the initial system. The combination of the GAT's ability to model structural chemical relationships and the Reaction Prior's focus on reaction likelihood from positive data provides a powerful and flexible approach to reaction scoring.

[0107] The Single-Step Reaction Generators (SSRG) 112, 212 combine the respective RRG 120, 220 and the RS 124, 224 to generate for an input product a ranked list of candidate sets of substrates, which are each assigned a Combined Score between 0 and 1 (Query 1.5 136 and Query 2.7236). In the initial system, the RRG 120 is used to generate reaction candidates by predicting multiple sets of possible substrates for a given product. These candidates are scored using RRG scores, which reflect the likelihood or feasibility of each predicted transformation based on the model's learned chemical rules and training data. The reactions are then additionally scored with the RS 124. The final score for each generated reaction is computed using the default Joining Function 122, which is the product of its RRG score and its RS score. Additionally, a Reference Reaction Generator 126 adds a set of Reference Reactions for each of the reactions generated by the SSRG. These reactions are searched in the Reaction Database using a metric for reaction similarity based on chemical reaction fingerprints (Query 1.3 132).

[0108] 8.3. Synthesis Planning UI

[0109] The Synthesis Planning UI 118 allows the user to submit batches of compounds for which synthesis plans will be generated (Query 1.6 138 and Query 2.8 238), inspect the status of the submitted batches, and view the generated synthesis plans. Upon entering the Synthesis Planning UI 118, the user sees the login view. After logging in, the user is redirected to an input view (FIG. 3).

[0110] The user can submit a batch of input molecules using a CSV file with SMILES (FIG. 3) or a graphical editor (FIG. 4). The user can browse the status of all their batches in a separate view, view detailed results for a selected batch (FIG. 5), or see the top synthesis plan proposed for a single molecule from a batch (FIG. 6).

[0111] FIG. 3 is a screenshot 300 from an embodiment of a Synthesis Planning UI 118 to input Query 1.6 138 or Query 2.8 238. In FIG. 3, the user can input target molecules 302 as aCSV file 304 containing SMILES representing the compounds 306. FIG. 4 is a screenshot of graphical editor 400 within Synthesis Planning UI 118 illustrating the option of inputting target molecules 402, 404 using a graphical editor toolbar 406, with tools for adding elements and structures.

[0112] One of the key functionalities of the UI 118 is the ability to export the results of synthesis planning to a machine-readable JSON format (Step 1.7 140 in FIG. 1 and Step 2.9240 in FIG. 2). The user can select the compounds for which the results should be exported. Such exported results can then be sent to the developer of the application and used to gather Chemistry Feedback about them. In particular, the user can export results that seem incorrect or of high risk in order to evaluate them with the help of a chemistry expert.

[0113] Such results could also be used to guide the operations of an automated or semiautomated laboratory by providing a precise sequence of chemical reactions for synthesis. This could facilitate the execution of synthesis plans, enabling automation in experimental workflows and improving reproducibility. Additionally, the availability of alternative reaction sequences allows for exploring different synthetic routes, potentially identifying more efficient or feasible options.

[0114] In an embodiment, a Synthesis Planning UI may provide a view allowing the user to see the status of all the batches of compounds submitted for synthesis planning. Each line may contain information about a batch, such as the status of the computation, batch name, number of compounds, cost in credits, and creation date. The view may also allow a user to remove batches or see a detailed description of a batch by clicking on it. This view is useful in order to select the results to export in Step 1.7 140 in FIG. 1 or Step 2.9240 in FIG. 2.

[0115] FIG. 5 is a screenshot 500 from an embodiment of a Synthesis Planning UI. A detailed summary of results for a completed batch 502 of compounds 502. The main table displays information for each of the compounds 504 from the batch, such as compound SMILES 506, synthesis planning status 508, score of the synthesis plan 510, total estimated price of performing the synthesis 512, and number of reactions involved in the synthesis 514. For example, target molecule 532 has a “completed” status with a score of 5.70 534, a price of 536, and 4 reactions 538 in the synthesis plan. On the right-hand side, an additional interface 520 allows filtering the displayed compounds 504 based on properties such as synthesis planning status 522, score 524, number of reactions 526, or the estimated synthesis cost 528. The interfacealso allows the user to export of all the results to a machine-readable JSON format using the “Export results” button 530 (Step 1.7 140 in FIG. 1 and Step 2.9240 in FIG. 2).

[0116] FIG. 6 is a screenshot 600 from an embodiment of a Synthesis Planning UI. Summary of the synthesis plan 602 found for one of the molecules from an input batch. The displayed synthesis plan starts from the left of the target molecule 532 and ends in the starting materials 604... 610 (indicated by the buttons 612...618 with a shopping card, which allows finding a supplier to buy a starting material). Each step 620... 626 in the displayed graph 602 represents a chemical reaction proposed by the Al models, with Reference Reactions being accessible when clicking on the button 628...634 with an opened book. This view is useful in order to select the results to export in Step 1.7 140 in FIG. 1 or Step 2.9240 in FIG. 2.

[0117] 8.4. Reaction Evaluation UI

[0118] The Reaction Evaluation UI 104 causes the system to perform step 1.8 142 from FIG. 1 and step 2.10242 from FIG. 2. In an embodiment, the Reaction Evaluation UI is accessed by human experts who evaluate the results provided by the Software For Synthesis Planning (SFSP). It is accessed by a login page, which after logging in redirects to the Imports view (FIG. 7), with an option 702, 704 to switch to the Reactions view (FIG. 9).

[0119] FIG. 7 is a screenshot 700 from an embodiment of a Reaction Evaluation UI 104 for different human Evaluators. The Imports view allows batches of reactions to be imported 706 for evaluation. These reactions can come from synthesis plans that were generated and exported using the Synthesis Panning UI. Clicking on the “Import from file” button 706 opens a window 800 (FIG. 8) that allows the user to select the file 802 to upload, passing it with a name 804 and tags 806 that will be assigned to reactions from this batch (tags can be used for filtering, depicted in FIG. 13). After the batch is imported, its reactions are saved in a database of reactions to evaluate.

[0120] FIG. 9 is a screenshot 900 from an embodiment of a Reaction Evaluation UI 104. The Reactions view displays a list 902 of all the reactions that need evaluation by an Evaluation in order to obtain Chemistry Feedback. Each reaction is depicted graphically 904a, 904b. The Evaluator can input their assessment of the reaction using the Verdict 906 and Verdict weight 908 pull-down menus (depicted in more detail in FIG. 10). The Evaluator can also view the Reference Reaction proposed by the Reference Reaction Generator for the given reaction by clicking on the “References” button 910, which opens another view, depicted in FIG. 11.Optionally, the Evaluator can add comments 912 about their decision process when evaluating the reaction (FIG. 12). The Evaluator can also filter the reactions in the Reactions view using the UI filter feature 1302 depicted in FIG. 13. All the evaluation results can be exported to a single CSV file using the “Export results” button 914. This CSV file can be used as a CFD. A detailed process for evaluating reactions is described in Section 7.5.

[0121] FIG. 10 is a screenshot 1000 from an embodiment of a Reaction Evaluation UI 104. It depicts the pull-down menu options for Verdict 906 and Verdict Weight 908. The Evaluator can select a type of chemical problem with a proposed reaction using Verdict menu 906 and the severity of the problem using Verdict weight menu 908. A detailed process for choosing appropriate verdicts for reactions is described in Section 7.5.

[0122] FIG. 11 is a screenshot 1100 from an embodiment of a Reaction Evaluation UI 104. It shows the UI that presents Reference Reactions 1102 found for a given evaluated reaction 1104. Optionally, the Evaluator can show additional information about any of the Reference Reactions, such as the original reaction description from its source patent (“Show patent” 1106 option), or extract SMILES 1108 of a reference reaction.

[0123] FIG. 12 is a screenshot 1200 from an embodiment of a Reaction Evaluation UI 104, which depicts a UI feature that allows the Evaluator to input comments 912, for example, about their decision process when evaluating the reaction. These comments are saved as part of CFD exported from the evaluation results. Natural language comments, including decision process explanations, can be used particularly to finetune the LLM agent described in section 6.2.

[0124] FIG. 13 is a screenshot 1300 from an embodiment of a Reaction Evaluation UI, which depicts a UI filter feature 1302 that allows filtering out reactions shown in the Reactions view (FIG. 9) based on their properties 1304, such as Status 1306 (Unclassified / Classified) or tags 1308. This allows the Evaluator to focus, for example, only on the reactions that are assigned to them (with specific tags) or reactions that are not yet evaluated (unclassified).

[0125] 8.5. Process for generating Chemistry Feedback about Chemical Reactions

[0126] Each human Evaluator who takes part in creating Chemistry Feedback using the Reaction Evaluation UI (step 1.8 142 from FIG. 1 and step 2.10 242 from FIG. 2) is instructed to follow a strict procedure in order to guarantee the quality and replicability of the evaluationresults. An example process that could be followed during the evaluation is depicted in decision tree 1500, FIG. 15A - FIG. 15C.

[0127] FIG. 15A - FIG. 15C illustrate an embodiment of a decision tree 1500 for evaluating reactions in the Reaction Evaluation UI 104. Process 1500 contains questions that the human evaluator is directed to ask to decide on the Verdict and Verdict Weight for a reaction. The evaluator starts in the START node, and answers a series of chemistry-related questions1504... 1522 (rhombs) to decide about the reaction Verdict and Verdict Weight. Nodes1524... 1538 contain final decisions about reaction Verdicts 1542... 1552 and optionally Verdict Weights 1560... 1572. Nodes may contain decision rules for assigning Verdict Weight whenever such rules are necessary, e.g., as in nodes 1564, 1568... 1572. The procedure from decision tree1500 is conducted by the Evaluator for each of the reactions in the Reaction Evaluation UI, in order to select the Verdict and Verdict Weight field values. Optionally, the Evaluator can add comments which explain the decisions they made during the evaluation process (FIG. 12). Following decision tree 1500 guarantees that each evaluated reaction has a non-empty Verdict and a non-empty Verdict Weight.

[0128] A Chemistry Feedback Dataset 106, 206 may be created by gathering all the reactions that have non-empty Verdicts and Verdict Weights in the Reaction Evaluation UI. The creation of CFD is done by clicking the “Export results” button in the Reactions view (FIG. 8), which saves evaluation results in a CSV file. Section 7.6 describes how this dataset may be used to improve the models for synthesis planning.

[0129] 8.6. Adapting the System Using Chemistry Feedback

[0130] This section describes how Chemistry Feedback gathered using the Reaction Evaluation UI 104 is used to refine Synthesis Planning System 102, resulting in Synthesis Planning System 202.

[0131] 8.6.1. Constructing a Training Dataset Based on Chemistry Feedback

[0132] This process description applies to Process 1.9 144 in FIG. 1 and Process 2.11 244 in FIG. 2. After obtaining the Chemistry Feedback Dataset (CFD) 106, 206 using the Reaction Evaluation UI, the reactions from the CFD are automatically split into positive and negative examples based on evaluators verdicts. Positive examples are the reactions with the Verdict Weight “safe bet” or “worthwhile”, while negative examples are the reactions with the Verdict Weight “rather not” or “nonsense”. Optionally, Verdict Weight can be used as sampleweight for each reaction for the purpose of model training or model evaluation. For example, negative reactions with “nonsense” Verdict Weight have higher weight than “rather not” reactions, while positive “safe bet” reactions have higher weight than “worthwhile” reactions. Each of the reactions is also assigned to a training or validation dataset, by assigning reactions from 50% randomly selected synthesis plans to the training set, and the other 50% of the reactions to the validation set. Such a dataset is called a Training Chemistry Feedback Dataset (TCFD) 108.

[0133] 8.6.2. Determining Joining Functions for a Single-Step Reaction Generator

[0134] In an embodiment, a set of parameterized Joining Functions 122, 222 is defined, each of which takes as input the relative RRG score and RS score and outputs a single number between 0 and 1 (Find 2.4, FIG. 2). A set of Joining Functions include operations such as minimum, maximum, normalized sum of logarithms, and weighted multiplication, among others. Each function can have multiple possible parameter sets, defining how the input scores are combined. For every combination of a Joining Function and its parameter set, the combined score is computed for all reactions in the Training Chemistry Feedback Dataset (TCFD) 108, 208 which is treated as a dataset for binary classification. The ROC AUC (Receiver Operating Characteristic Area Under Curve) metric is then calculated for each case, and the final Joining Function is selected based on the highest ROC AUC value. This selected function is used in the version of the system adapted using chemistry feedback.

[0135] 8.6.3. Finetuning the Reaction Scorer with Chemistry Feedback

[0136] TCFD 108 is used to conduct a second stage of training for the two Reaction Scorer models (Finetune 2.5): Graph Attention Network (GAT) and Reaction Prior (RP).

[0137] To fine-tune GAT, which is a binary classification model, the model trained on the Reaction Database is loaded, and its training is continued using positive and negative reactions from TCFD.

[0138] RP is a decoder-only Transformer, originally trained exclusively on positive reactions from the Reaction Database. During this training, the model learns to predict the likelihood of a given reaction based on the relationships between reactants, products, and reagents. It captures complex patterns such as reaction mechanisms, functional group transformations, reagent compatibility, and molecular structure dependencies, allowing it to generalize across a variety of chemical reactions.

[0139] To finetune it on a dataset containing both positive and negative reactions from TCFD, its architecture must be adapted for binary classification. To achieve this, a technique called Linear Probing is used, which introduces special linear mapping layers into the hidden representations of the RP model. These layers are injected at specific depths within the Transformer to modify its output without altering the underlying learned features. Only these additional layers are trained on the TCFD, while the core Transformer remains frozen. This approach ensures that the model retains the rich representation it learned from the Reaction Database — capturing complex reaction patterns, reagent dependencies, molecular transformations, and other features like stereochemistry and regiochemistry — while allowing it to adjust its outputs to provide accurate classification predictions based on TCFD.

[0140] 8.6.4. Training the Adapted RS

[0141] After finetuning both GAT and RP, the final Reaction Scorer used in the version of the system adapted using Chemistry Feedback is a random forest model trained on the predictions of GAT and RP. Random forest is an ensemble learning method that constructs multiple decision trees during training and combines their outputs to improve predictive accuracy and reduce overfitting. Each tree in the random forest is trained on a random subset of the training data, and the final prediction is obtained by aggregating the outputs of all trees, typically through majority voting for classification tasks or averaging for regression.

[0142] In the embodiment, the random forest is trained to optimize the classification of reactions by leveraging the predictions from GAT and RP as input features. The model learns to assign an importance to each predictor based on its contribution to classification performance, effectively capturing complex interactions between the two models. The training process involves evaluating the random forest on the TCFD dataset and selecting hyperparameters that maximize the ROC AUC score. By using a random forest, the system ensures a more robust and adaptive combination of GAT and RP, reducing overreliance on any single model and enhancing overall classification performance.

[0143] 8.6.5. Evaluating a Single-Step Reaction Generator using the Chemistry FeedbackDataset

[0144] Training Chemistry Feedback Dataset (TCFD) can be used to evaluate the ability of a specific Single-Step Reaction Generator (SSRG) to generate correct chemical reactions. For each of the evaluated reactions in the TCFD, the following procedure can be applied. The SSRGcan be asked to predict a list of ranked reaction substrates from the product of the reaction. If the reaction does not appear in this list, we assume that the SSRG gave it a score of 0. Otherwise, the Combined Score that the SSRG gives for this reaction is calculated. After acquiring such scores for each of the reactions in the TCFD, performance metrics such as Accuracy, Cross-Entropy loss, or Receiver Operating Characteristics can be calculated between the predicted labels and ground truth labels. The predicted label is created by applying a threshold to the Combined Score: if the score is above the threshold, the reaction is considered positive (label = 1), otherwise negative (label = 0). The ground truth label is determined by the Evaluators' verdicts: a reaction is labeled as positive (1) if it is deemed successful by the evaluators, and negative (0) if it is not. In particular, to measure the performance of SSRG on specific types of chemistry problems, these metrics can be calculated on subsets of TCFD corresponding to each of the possible Verdicts. For example, to calculate the performance of the SSRG on Chirality problems, we calculate the metrics on the subset of TCFD that contains all positive reactions and all negative reactions with Verdict: Chirality. These verdict-specific metrics can inform the developers of the SFSP about the most urgent directions of model improvements. Similarly, TCFD 208 may be used as described above to refine System 202. Furthermore, subsequent systems may generate additional TCFDs, which may be used in refining the systems as many times as desired.

[0145] 8.6.6. Displaying information about the improvement based on ChemistryFeedback Datasets in the Synthesis Planning UI

[0146] Synthesis Planning UI 118 can have an additional UI view, accessible from the main view, screenshot 1400, which allows a user to inspect which Chemistry Feedback Datasets (CFDs) 1402a, 1402b have been used to improve the system (FIG. 14). This view can include information about the date 1404 that the CFD was used and the number of positive 1406 and negative 1408 reactions used to improve the system. After each iteration of improving the Software for Synthesis Planning using a CFD, the table in this view is appended with a new entry.

[0147] 8.7. Example for Improving Software for Synthesis Planning Using ChemistryFeedback

[0148] This section describes an embodiment for generating a Chemistry Feedback Dataset using Software for Synthesis Planning (SFSP) and Reaction Evaluation UI 104 and used to improve the SFSP.

[0149] 8.7.1. Obtaining Results for Evaluation from Software for Synthesis Planning1. A user of SFSP logs in to the Synthesis Planning UI 118.2. The user is in the Searches view. The user uploads a CSV file with a batch of target molecules for synthesis planning (FIG. 3).3. The user instructs the Synthesis Planning UI to send a query to search for synthesis plans for the input batch to the Synthesis Planning Search (Query 1.6 138 on FIG. 1, i.e., the user selects the “Start search” link). Meanwhile, the user can see the information about the running batch and other running or finished batches in the UI (the running batch is named “Test Batch 2” 502 (FIG. 5)).4. The Synthesis Planning Search (SPS) finds the Top 1 synthesis path for each of the molecules from the batch, querying the Starting Materials database and Single-Step Reaction Generator (SSRG) 112 in the process (Query 1.4 134 and Query 1.5 136 on FIG. 1).5. After SPS finishes the calculation, it sends the results to the Synthesis Planning UI 118 (Query 1.6 138 on FIG. 1).6. The user may now see that the batch has the Status “Completed” in the UI (FIG. 5). The user clicks on the batch name (Test Batch 2) to see a detailed summary of synthesis planning results for each molecule in the batch (FIG. 5).7. The user clicks on one of the targets in the batch (e.g., CC(Nclcc(F)cc(C(N2)=NC3(CCN(CCS(c4ccccc4)=O)CC3)C2=O)cl l)=CCl=O, 532, FIG. 5) and is redirected to a view displaying the synthesis path 602 found for this target (FIG. 6).8. The user decides that they want to export reactions from this synthesis plan for evaluation. The user clicks on “Back to results” to go back to the view displayed in FIG. 5. The user selects the relevant batch by clicking on a checkbox 540 on the left of the SMILES column. The user clicks on “Export results”, which downloads a CSV file with the reactions from the selected synthesis path to the user’s machine.9. Using other channels, such as email, dropbox, and other methods of sharing data files, the user sends the exported CSV to the maintainer of SFSP in order to be considered for evaluation to generate Chemistry Feedback.

[0150] 8.7.2. Generating Chemistry Feedback Dataset

[0151] In an embodiment of a process for generating a CFD:1. An Evaluator receives the CSV file with reactions to evaluate from the user of SFSP.2. The Evaluator logs in to the Reaction Evaluation UI and is redirected to the Imports view (FIG. 7).3. The Evaluator clicks on the option “Import from file” to import reactions from the CSV file using the “Import from file” view (FIG. 8). The Evaluator assigns the imported file with an “Import name” and some “Reaction default tags” - for example, “2024_03_retro_search_l 5”.4. The Evaluator is redirected back to the Imports view, where the newest Import is visible (FIG. 7). The Evaluator clicks on the “Reactions” button to be redirected to the Reactions view (FIG. 9), where the Evaluator can now see all the reactions uploaded so far for evaluation, including the ones from the most recent upload (listed on pages: 1... 1143, Next).5. The Evaluator wants to select only the reactions from the batch that they just uploaded. To do so, they click on the filter feature 1302 (FIG. 13) and select only reactions with the tag “2024_03_retro_search_l 5” that are not yet evaluated (Status: Unclassified).6. The Evaluator now sees the Reactions view with only reactions that they want to evaluate (FIG. 9). The evaluator proceeds to evaluate the first reaction displayed in the view.7. In a separate browser window, the Evaluator opens a view displaying a decision tree 1500 (FIG. 15A - FIG. 15C), which instructs how to evaluate the reaction. The Evaluator follows steps from decision tree 1500 to input relevant evaluation results using the view from FIG. 9: Verdict, Verdict Weight (FIG. 10), and optionally proof reference reactions (e.g., a reaction that supports the Evaluator’s verdict) and Comments (FIG. 12). To better answer the questions from decision tree 1500 and to select proof reference reactions, the Evaluator can analyze the Reference Reactions1102 suggested by SFSP 102, 202 for the evaluated reaction (FIG. 11). Each of these evaluation results (Verdict, Verdict Weight, Proof reference reactions, Uncertainty, Comment(s)) constitutes Chemistry Feedback (CF) about the evaluated reaction.8. The Evaluator can repeat the process from step 7 to evaluate more reactions displayed in the Reactions view (FIG. 9). When finishing the evaluation, the Evaluator exports the results of the evaluation using the “Export results” button, which downloads a CSV file with the reactions together with their evaluation labels as a CSV to the Evaluator’s machine. This CSV file constitutes a Chemistry Feedback Dataset (CFD).

[0152] 8.7.3. Using Chemistry Feedback Dataset to improve Software for SynthesisPlanning

[0153] In an embodiment of a method for using CFD to improve SFSP:1. The developer of the SFSP receives a CSV with evaluated reactions, which constitutes a CFD.2. The developer uses the CFS to generate a Training Chemistry Feedback Dataset (TCFD), as described in section 7.6.1.3. The developer uses the CFD, together with a Reaction Database, to train the Reaction Scorer (RS) model (Step Train 2.2 on FIG. 2). This step is described in detail in section 7.6.3.4. The developer uses the CFD to find an optimal Joining Function between the Retrosynthesis Reaction Generator (RRG) and RS (step Find 2.4 on FIG. 2). This step is described in detail in section 7.6.2.5. The developer deploys the improved system as a new refined version of SFSP.6. For the refined version of SFSP, a new CFD can be generated, following steps similar to the example from sections 7.7.1 and 7.7.3 (Steps 2.8, 2.9, 2.10, and 2.11 from FIG. 2). Such a CFD can be used again to generate the next refined version of SFSP.

[0154] FIG. 16 is a flowchart of an embodiment of a computer- implemented method 1600, including steps 1602 - 1616, for optimizing a machine learning model configured to output predictions regarding chemical reactions. In method 1600, step 1602 requires computationally generating, by a computing system including at least one processor and memory, at least one machine learning model including adjustable weights and being configured to output predictions regarding chemical reactions. Step 1604 requires training, by thecomputing system, the at least one machine learning model on historical chemical data. Step 1606 requires generating, by a reaction evaluation module executed by the computing system, at least one example chemical reaction. Step 1608 requires, directing, by the reaction evaluation module, the evaluation of each generated example chemical reaction through at least one of: a performance of a laboratory experiment, a quantum mechanical simulation, or using a user interface to extract at least one answer from a human. Step 1610 requires, receiving, by the computing system, chemistry feedback for each example chemical reaction comprising: a prediction regarding the example chemical reaction, and evidence supporting the prediction. Step 1612 requires, generating, by the computing system, a chemistry feedback dataset using the chemistry feedback. Step 1614 requires, evaluating, by the computing system, the machine learning model by computing an evaluation score configured to measure agreement between outputs of the machine learning model and the chemistry feedback dataset, wherein computing a first evaluation score includes: deriving virtual chemical reactions from the chemistry feedback dataset; inputting the virtual chemical reactions into the machine learning model; receiving output from the machine learning model based on the virtual chemical reactions; comparing the output to the chemistry feedback dataset to determine consistencies and / or inconsistencies between the output and the chemistry feedback dataset; and quantifying the determined consistencies and / or inconsistencies as the evaluation score. And, step 1616 requires, refining, by an optimization module executed by the computing system, the machine learning model to create a refined machine learning model by changing at least one numerical parameter of the machine learning model, wherein the at least one numerical parameter is changed to cause a predicted evaluation score to be greater than the first evaluation score, indicating higher measured agreement between outputs of the refined machine learning model and the chemistry feedback dataset. For example, the predicted evaluation score may be based on a hypothetical re-computation of the first evaluation score using the refined machine learning module.

[0155] 8.8. Experimental validation of Synthesis Planning Software Adapted Based onChemistry Feedback

[0156] The objective of this section is to empirically validate the effectiveness of the methods described herein, specifically demonstrating that adapting the Al-based Synthesis Planning Software (SFSP) 102 using a Chemistry Feedback Dataset (CFD) 106 improves the performance of the refined SFSP 202. The primary hypothesis is that such adaptation enhancesthe system's ability to discriminate between chemically correct and incorrect reactions, particularly concerning specific error types identified by expert evaluators, thereby aligning the system's output more closely with expert chemical judgment compared to a system trained solely on historical data.

[0157] 8.8.1. Comparing baseline open-source SFSP with adapted SFSP based onChemistry Feedback

[0158] Baseline SFSPs

[0159] FIG. 17A and FIG. 17B are charts showing results from different evaluations of an embodiment of a synthesis planning system. In FIG. 17A, the comparison involves SFSP 202 and two leading open-access synthesis planning tools: AiZynthFinder (https: / / github.com / MolecularAI / aizynthfinder) and IBM RXN(https: / / rxn.app. accelerate. science / ). AiZynthFinder is executed locally based on the implementation provided in its GitHub repository, while IBM RXN is accessed through its web application.

[0160] Chemistry Feedback Dataset (CFD)

[0161] For the comparison in FIG. 17A, the Chemistry Feedback Dataset (CFD) was constructed following the procedures outlined in Sections 8.5 and 8.7.2. It consisted of reactions initially proposed by various Reaction Reasoning Generators (RRGs) for a diverse set of target molecules. These reactions were subsequently evaluated by expert chemists, forming a curated dataset of accepted and rejected reactions.

[0162] Dataset Composition

[0163] The TCFD, constructed as described in Section 8.6.1. with reference to FIG. 1, was used to fine-tune the adapted synthesis planning system as discussed with respect to FIG. 2. The TCFD contained approximately 3,000 evaluated reactions, of which approximately 2,000 were classified as ‘Correct’ and 700 as ‘Incorrect.’

[0164] Adapted SFSP

[0165] The Adapted Reaction Scorer was incorporated into an Adapted SFSP by integrating an RRG based on a Transformer model architecture and a Reaction Scorer utilizing a Graph Attention Network (GAT). These models were trained on a publicly available dataset from the USPTO, consisting of reactions extracted from patents, as detailed in Section 8.2.Together, these components formed the SSRG, which incorporated a Joining Function optimizedusing the methodology described in Section 8.6.2. The adapted SFSP was constructed by combining the SSRG with an SPS based on the Retro* algorithm.

[0166] Drug-Like Molecule Set

[0167] To evaluate the performance of the synthesis planning systems, a set of 26 druglike molecules was used that exhibited a wide range of structural complexity and diverse chemical motifs. These molecules were selected to assess the robustness and generalizability of the synthesis planning software, ensuring that the evaluation covers real-world chemical challenges.

[0168] Evaluation Procedure

[0169] Following the procedure described in Section 8.7.1, each Baseline Reaction Scorer, the Adapted Reaction Scorer, and the resulting SFSPs were used to provide a Synthesis Path for each of the molecules from the Drug-Like Molecule Set. For each target molecule, the top-1 Synthesis Path generated by the system was analyzed using the methodology outlined in Section 8.7.2. Evaluators assessed the validity of each reaction within the path, classifying it as either ‘Correct’ or ‘Incorrect.’ A Synthesis Path was considered correct if all its reactions were correct; if at least one reaction was incorrect, the entire path was deemed incorrect.

[0170] Results

[0171] FIG. 17A presents the number of correct and incorrect pathways for each of the Baseline Reaction Scorers, the Adapted Reaction Scorer, and their associated SFSPs. The results demonstrate a significant improvement in pathway correctness for the Adapted system, which incorporated the Joining Function adapted using Chemistry Feedback. These findings highlight that leveraging chemistry feedback resulted in a SFSP with dramatically enhanced performance over the Baselines, and illustrate the improvement in the SFSP’s ability to generate synthetically viable pathways.

[0172] 8.8.2.Enhancing Reaction Scorer with Fine-Tuning on the Training Chemistry Feedback Dataset

[0173] This section demonstrates, with reference to FIG. 17B, the improvement in reaction correctness within the SSRG 112 when the Reaction Scorer 124 is fine-tuned on the Training Chemistry Feedback Dataset (TCFD) 108, becoming refined RS 224.

[0174] Two versions of the Reaction Scorers are compared:1. Baseline GAT Reaction Scorer: This model employs a Graph Attention Network (GAT)-based Reaction Scorer 124.2. Baseline RP Reaction Scorer: This model employs a Reaction Prior (RP)-based Reaction Scorer 124.3. Fine-Tuned Reaction Scorer 212: It contains GAT-based Reaction Scorer and a Reaction Prior Scorer fine-tuned and integrated using a Random Forest algorithm, as detailed in Section 8.6.3 and 8.6.4.

[0175] Construction of the Training Chemistry Feedback Dataset (TCFD) 108

[0176] Similar to Section 8.6.1, TCFD 108 is derived from the Chemistry Feedback Dataset 106. However, beyond the simple classification of reactions as ‘Correct’ or ‘Incorrect,’ this dataset 108 was also caused to capture the primary chemical flaw identified during expert review. Incorrect reactions are categorized based on their underlying issues, including: Unknown Mechanism, Selectivity Issue, and Other Issue.

[0177] Evaluation Procedure

[0178] To assess the performance of the two Reaction Scorers, a validation split of TCFD is used. The evaluation focuses on the models’ ability to reject incorrect reactions while retaining correct ones. The rejection threshold is set to ensure that 80% of correct reactions from the validation set are preserved.

[0179] Results

[0180] FIG. 17B presents the Acceptance Rate for correct reactions and the Rejection Rate for incorrect reactions, analyzed per Verdict Category. Across all categories, fine-tuning the Reaction Scorer significantly improves its ability to reject incorrect reactions compared to the baseline models.

[0181] FIG. 18 is an exemplary block diagram depicting an embodiment of system for implement embodiments of methods of the disclosure, e.g., as described with reference to the previous figures. In FIG. 18, computer network 1800 includes a number of computing devices 1810a-1810b, and one or more server systems 1820 and / or automated laboratories 1822 coupled to a communication network 1860 via a plurality of communication links 1830. Communication network 1860 provides a mechanism for allowing the various components of distributed network 1800 to communicate and exchange information with each other.

[0182] Communication network 1860 itself is comprised of one or more interconnected computer systems and communication links. Communication links 1830 may include hardwire links, optical links, satellite or other wireless communications links, wave propagation links, or any other mechanisms for communication of information. Various communication protocols may be used to facilitate communication between the various systems shown in FIG. 18. These communication protocols may include TCP / IP, UDP, HTTP protocols, wireless application protocol (WAP), BLUETOOTH, Zigbee, 802.11, 802.15, 6L0WPAN, L1F1, Google Weave, NFC, GSM, CDMA, other cellular data communication protocols, wireless telephony protocols, Internet telephony, IP telephony, digital voice, voice over broadband (VoBB), broadband telephony, Voice over IP (VoIP), vendor-specific protocols, customized protocols, and others. While in one embodiment, communication network 1860 is the Internet, in other embodiments, communication network 1860 may be any suitable communication network including a local area network (LAN), a wide area network (WAN), a wireless network, a cellular network, a personal area network, an intranet, a private network, a near field communications (NFC) network, a public network, a switched network, a peer-to-peer network, and combinations of these, and the like.

[0183] In an embodiment, server 1820 and automated laboratory 1822 may not be located near a user of a computing device, and are communicated with over network 1860. In a different embodiment, the server 1820 is a device that a user can carry upon his person, or can keep nearby. In an embodiment, the server 1820 has a large battery to power long distance communications networks such as a cell network or Wi-Fi. The server 1820 communicates with the other components of the system via wired links or via low powered short-range wireless communications such as BLUETOOTH. In an embodiment, one of the other components of the system plays the role of the server, e.g., the PC 1810b.

[0184] In an embodiment, automated laboratory 1822 includes a semi-automated high- throughput laboratory is used, which enables generating large datasets of chemical reactions.

[0185] Distributed computer network 1800 in FIG. 18 is merely illustrative of an embodiment incorporating the embodiments and does not limit the scope of the invention as recited in the claims. One of ordinary skill in the art would recognize other variations, modifications, and alternatives. For example, more than one server system 1820 may be connected to communication network 1860. As another example, a number of computingdevices 1810a-1810b may be coupled to communication network 1860 via an access provider (not shown) or via some other server system.

[0186] Computing devices 1810a-1810b typically request information from a server system that provides the information. Server systems by definition typically have more computing and storage capacity than these computing devices, which are often such things as portable devices, mobile communications devices, or other computing devices that play the role of a client in a client-server operation. However, a particular computing device may act as both a client and a server depending on whether the computing device is requesting or providing information. Aspects of the embodiments may be embodied using a client-server environment or a cloud-cloud computing environment.

[0187] Server 1820 is responsible for receiving information requests from computing devices 1810a-1810b, for performing processing required to satisfy the requests, and for forwarding the results corresponding to the requests back to the requesting computing device. The processing required to satisfy the request may be performed by server system 1820 or may alternatively be delegated to other servers connected to communication network 1860 or to other communications networks. A server 1820 may be located near the computing devices 1810 or may be remote from the computing devices 1810. A server 1820 may be a hub controlling a local enclave of things in an internet of things scenario.

[0188] Computing devices 1810a-1810b enable users to access and query information or applications stored by server system 1820. Some example computing devices include portable electronic devices (e.g., mobile communications devices) such as the Apple iPhone®, the Apple iPad®, the Palm Pre™, or any computing device running the Apple iOS™, Android™ OS, Google Chrome OS, Symbian OS®, Windows 10, Windows Mobile® OS, Palm OS® or Palm Web OS™, or any of various operating systems used for Internet of Things (loT) devices or automotive or other vehicles or Real Time Operating Systems (RTOS), such as the RIOT OS, Windows 10 for loT, WindRiver VxWorks, Google Brillo, ARM Mbed OS, Embedded Apple iOS and OS X, the Nucleus RTOS, Green Hills Integrity, or Contiki, or any of various Programmable Logic Controller (PLC) or Programmable Automation Controller (PAC) operating systems such as Microware OS-9, VxWorks, QNX Neutrino, FreeRTOS, Micrium pC / OS-II, Micrium pC / OS-III, Windows CE, TI-RTOS, RTEMS. Other operating systems may be used. In a specific embodiment, a “web browser” application executing on a computingdevice enables users to select, access, retrieve, or query information and / or applications stored by server system 1820. Examples of web browsers include the Android browser provided by Google, the Safari® browser provided by Apple, the Opera Web browser provided by Opera Software, the BlackBerry® browser provided by Research In Motion, the Internet Explorer® and Internet Explorer Mobile browsers provided by Microsoft Corporation, the Firefox® and Firefox for Mobile browsers provided by Mozilla®, and others.

[0189] FIG. 19 is an exemplary block diagram depicting a computing device 1900 of an embodiment. Computing device 1900 may be any of the computing devices 1810a, 1810b, 1820 from FIG. 18. Computing device 1900 may include a display, screen, or monitor 1905, housing 1910, and input device 1915. Housing 1910 houses familiar computer components, some of which are not shown, such as a processor 1920, memory 1925, battery 1930, speaker, transceiver, antenna 1935, microphone, ports, jacks, connectors, camera, input / output (I / O) controller, display adapter, network interface, mass storage devices 1940, various sensors, and the like.

[0190] Input device 1915 may also include a touchscreen (e.g., resistive, surface acoustic wave, capacitive sensing, infrared, optical imaging, dispersive signal, or acoustic pulse recognition), keyboard (e.g., electronic keyboard or physical keyboard), buttons, switches, stylus, or combinations of these.

[0191] Mass storage devices 1940 may include flash and other nonvolatile solid-state storage or solid-state drive (SSD), such as a flash drive, flash memory, or USB flash drive. Other examples of mass storage include mass disk drives, floppy disks, magnetic disks, optical disks, magneto-optical disks, fixed disks, hard disks, SD cards, CD-ROMs, recordable CDs, DVDs, recordable DVDs (e.g., DVD-R, DVD+R, DVD-RW, DVD+RW, HD-DVD, or Blu-ray Disc), battery-backed-up volatile memory, tape storage, reader, and other similar media, and combinations of these.

[0192] Embodiments may also be used with computer systems having different configurations, e.g., with additional or fewer subsystems. For example, a computer system could include more than one processor (i.e., a multiprocessor system, which may permit parallel processing of information) or a system may include a cache memory. The computer system shown in FIG. 19 is but an example of a computer system suitable for use with the embodiments. Other configurations of subsystems suitable for use with the embodiments will be readilyapparent to one of ordinary skill in the art. For example, in a specific implementation, the computing device is a mobile communications device such as a smartphone or tablet computer. Some specific examples of smartphones include the Droid Incredible and Google Nexus One, provided by HTC Corporation, the iPhone or iPad, both provided by Apple, and many others. The computing device may be a laptop or a netbook. In another specific implementation, the computing device is a non-portable computing device such as a desktop computer or workstation.

[0193] A computer-implemented or computer-executable version of the program instructions useful to practice the embodiments may be embodied using, stored on, or associated with computer-readable medium. A computer-readable medium may include any medium that participates in providing instructions to one or more processors for execution, such as memory 1925 or mass storage 1940. Such a medium may take many forms including, but not limited to, nonvolatile, volatile, transmission, non-printed, and printed media. Nonvolatile media includes, for example, flash memory, or optical or magnetic disks. Volatile media includes static or dynamic memory, such as cache memory or RAM. Transmission media includes coaxial cables, copper wire, fiber optic lines, and wires arranged in a bus. Transmission media can also take the form of electromagnetic, radio frequency, acoustic, or light waves, such as those generated during radio wave and infrared data communications.

[0194] For example, a binary, machine-executable version, of the software useful to practice the embodiments may be stored or reside in RAM or cache memory, or on mass storage device 1940. The source code of this software may also be stored or reside on mass storage device 1940 (e.g., flash drive, hard disk, magnetic disk, tape, or CD-ROM). As a further example, code useful for practicing the embodiments may be transmitted via wires, radio waves, or through a network such as the Internet. In another specific embodiment, a computer program product including a variety of software program code to implement features of the embodiment is provided.

[0195] Computer software products may be written in any of various suitable programming languages, such as C, C++, C#, Pascal, Fortran, Perl, Matlab (from MathWorks, www.mathworks.com), SAS, SPSS, JavaScript, CoffeeScript, Objective-C, Swift, Objective-J, Ruby, Rust, Python, Erlang, Lisp, Scala, Clojure, and Java. The computer software product may be an independent application with data input and data display modules. Alternatively, thecomputer software products may be classes that may be instantiated as distributed objects. The computer software products may also be component software such as Java Beans (from Oracle) or Enterprise Java Beans (EJB from Oracle).

[0196] An operating system for the system may be the Android operating system, iPhone OS (i.e., iOS), Symbian, BlackBerry OS, Palm web OS, Bada, MeeGo, Maemo, Limo, or Brew OS. Other examples of operating systems include one of the Microsoft Windows family of operating systems (e.g., Windows 95, 98, Me, Windows NT, Windows 2000, Windows XP, Windows XP x64 Edition, Windows Vista, Windows 10 or other Windows versions,, Windows CE, Windows Mobile, Windows Phone, Windows 10 Mobile), Linux, HP-UX, UNIX, Sun OS, Solaris, Mac OS X, Alpha OS, AIX, IRIX32, or IRIX64, or any of various operating systems used for Internet of Things (loT) devices or automotive or other vehicles or Real Time Operating Systems (RTOS), such as the RIOT OS, Windows 10 for loT, WindRiver VxWorks, Google Brillo, ARM Mbed OS, Embedded Apple iOS and OS X, the Nucleus RTOS, Green Hills Integrity, or Contiki, or any of various Programmable Logic Controller (PLC) or Programmable Automation Controller (PAC) operating systems such as Microware OS-9, VxWorks, QNX Neutrino, FreeRTOS, Micrium pC / OS-II, Micrium pC / OS-III, Windows CE, TI-RTOS, RTEMS. Other operating systems may be used.

[0197] Furthermore, the computer may be connected to a network and may interface to other computers using this network. The network may be an intranet, internet, or the Internet, among others. The network may be a wired network (e.g., using copper), telephone network, packet network, an optical network (e.g., using optical fiber), or a wireless network, or any combination of these. For example, data and other information may be passed between the computer and components (or steps) of a system useful in practicing the embodiments using a wireless network employing a protocol such as Wi-Fi (IEEE standards 802.11, 802.1 la, 802.1 lb, 802.11e, 802.11g, 802.11 i, and 802. Unjust to name a few examples), or other protocols, such as BLUETOOTH or NFC or 802.15 or cellular, or communication protocols may include TCP / IP, UDP, HTTP protocols, wireless application protocol (WAP), BLUETOOTH, Zigbee, 802.11, 802.15, 6L0WPAN, L1F1, Google Weave, NFC, GSM, CDMA, other cellular data communication protocols, wireless telephony protocols or the like. For example, signals from a computer may be transferred, at least in part, wirelessly to components or other computers.

[0198]

[0199] The following paragraphs include enumerated embodiments.

[0200] Embodiment 1. A computer- implemented method comprising: computationally generating, by a computing system including at least one processor and memory, at least one machine learning model including adjustable weights and being configured to output predictions regarding chemical reactions; training, by the computing system, the at least one machine learning model on historical chemical data; generating, by a reaction evaluation module executed by the computing system, at least one example chemical reaction; directing, by the reaction evaluation module, the evaluation of each generated example chemical reaction through at least one of: a performance of a laboratory experiment, a quantum mechanical simulation, or using a user interface to extract at least one answer from a human; receiving, by the computing system, chemistry feedback for each example chemical reaction comprising: a prediction regarding the example chemical reaction, and evidence supporting the prediction; generating, by the computing system, a chemistry feedback dataset using the chemistry feedback; evaluating, by the computing system, the at least one machine learning model by computing an evaluation score configured to measure agreement between outputs of the at least one machine learning model and the chemistry feedback dataset, wherein computing a first evaluation score includes: deriving virtual chemical reactions from the chemistry feedback dataset; inputting the virtual chemical reactions into the at least one machine learning model; receiving output from the at least one machine learning model based on the virtual chemical reactions; comparing the output to the chemistry feedback dataset to determine consistencies and / or inconsistencies between the output and the chemistry feedback dataset; and quantifying the determined consistencies and / or inconsistencies as the evaluation score; andrefining, by an optimization module executed by the computing system, the at least one machine learning model to create at least one refined machine learning model by changing at least one numerical parameter of the at least one machine learning model, wherein the at least one numerical parameter is changed to cause a predicted evaluation score to be greater than the first evaluation score, indicating higher measured agreement between outputs of the at least one refined machine learning model and the chemistry feedback dataset.

[0201] Embodiment 2. The computer- implemented method of embodiment 1, or any other embodiment, further comprising: receiving, by the computing system, input including a target molecule; proposing, by the computing system using the at least one refined machine learning model, a synthesis path for the target molecule; and directing, by the computing system, a laboratory to synthesize the target molecule using the synthesis path.

[0202] Embodiment 3. The computer-implemented method of embodiment 2, or any other embodiment, wherein: the at least one machine learning model: generates a fully specified chemical reaction based on a partially specified chemical reaction; or outputs one or more predictions when input a chemical reaction, the one or more predictions indicating predicted properties of the chemical reaction.

[0203] Embodiment 4. The computer-implemented method of embodiment 1 , or any other embodiment, wherein the evidence used to arrive at the prediction regards at least one of: a reaction mechanism, regioselectivity, chemoselectivity, stereoselectivity, an outcome of a chemical reaction at least one atom or bond removed or added, a relevance of a historical reaction, an electronic effect, chemical correctness, feasibility, stability, toxicity, yield, complexity, synthetic accessibility, practical utility, cost of starting materials or intermediates, reaction temperature, solvent identity, catalyst identity, reagent identity, or reaction atmosphere, mechanistic plausibility, likelihood of elementary reaction steps, or reliability of predicted reaction pathways, chemical impact of functional group addition, removal, or substitution in reactants or products, confidence, uncertainty, or prediction reliability, or chemical similarity tohistorically validated reactions.

[0204] Embodiment 5. The computer-implemented method of embodiment 1, or any other embodiment, wherein the refining, by the optimization module, the at least one machine learning model, further includes the optimization module changing the at least one machine learning model by changing how strongly two modules within the at least one machine learning model interact.

[0205] Embodiment 6. The computer-implemented method of embodiment 1 , or any other embodiment, wherein the evaluation of each generated example chemical reaction is through only the performance of a laboratory experiment or quantum mechanical simulation.

[0206] Embodiment 7. The computer-implemented method of embodiment 6, or any other embodiment, wherein the evidence used to arrive at the prediction regards at least one of: a reaction mechanism, regioselectivity, chemoselectivity, stereoselectivity, an outcome of a chemical reaction at least one atom or bond removed or added, a relevance of a historical reaction, an electronic effect, chemical correctness, feasibility, stability, toxicity, yield, complexity, synthetic accessibility, practical utility, cost of starting materials or intermediates, reaction temperature, solvent identity, catalyst identity, reagent identity, or reaction atmosphere, mechanistic plausibility, likelihood of elementary reaction steps, or reliability of predicted reaction pathways, chemical impact of functional group addition, removal, or substitution in reactants or products, confidence, uncertainty, or prediction reliability, or chemical similarity to historically validated reactions.

[0207] Embodiment 8. The computer-implemented method of embodiment 1, or any other embodiment, wherein the at least one machine learning model is configured to include, with each output prediction, evidence supporting the output prediction.

[0208] Embodiment 9. The computer-implemented method of embodiment 1 , or any other embodiment, wherein the at least one machine learning model includes a module that outputs one or more predictions when input a chemical reaction, the one or more predictions indicating predicted properties of the chemical reaction, and wherein refining includes training the module to output predictions that are more similar to predictions included in the chemistry feedback dataset.

[0209] Embodiment 10. A system for providing predictions regarding chemical reactions, the system comprising at least one processor and memory provided with instruction,which when executed configure the system to perform actions including: computationally generating at least one machine learning model including adjustable weights and being configured to output predictions regarding chemical reactions; training the at least one machine learning model on historical chemical data; generating, by a reaction evaluation module, at least one example chemical reaction; directing, by the reaction evaluation module, the evaluation of each generated example chemical reaction through at least one of: a performance of a laboratory experiment, a quantum mechanical simulation, or using a user interface to extract at least one answer from a human; receiving chemistry feedback for each example chemical reaction comprising: a prediction regarding the example chemical reaction, and evidence supporting the prediction; generating a chemistry feedback dataset using the chemistry feedback; evaluating the at least one machine learning model by computing an evaluation score configured to measure agreement between outputs of the machine learning model and the chemistry feedback dataset, wherein computing the evaluation score includes: deriving virtual chemical reactions from the chemistry feedback dataset; inputting the virtual chemical reactions into the at least one machine learning model; receiving output from the at least one machine learning model based on the virtual chemical reactions; comparing the output to the chemistry feedback dataset to determine consistencies and / or inconsistencies between the output and the chemistry feedback dataset; and quantifying the determined consistencies and / or inconsistencies as the evaluation score; and refining, by an optimization module, the at least one machine learning model to create at least one refined machine learning model by changing at least one numerical parameter of the at least one machine learning model, wherein the at least one numerical parameter is changed so that a predicted evaluation score is greater than the firstevaluation score, indicating higher measured agreement between outputs of the at least one refined machine learning model and the chemistry feedback dataset.

[0210] Embodiment 11. The system of embodiment 10, or any other embodiment, the actions further comprising: receiving input including a target molecule; proposing, using the at least one refined machine learning model, a synthesis path for the target molecule; and directing a laboratory to synthesize the target molecule using the synthesis path.

[0211] Embodiment 12. The system of embodiment 11, or any other embodiment, the actions further including at least one of: generating a fully specified chemical reaction based on a partially specified chemical reaction; or outputting one or more predictions when input a chemical reaction that indicate observed properties of the chemical reaction.

[0212] Embodiment 13. The system of embodiment 10, or any other embodiment, wherein the evidence used to arrive at the prediction regards at least one of: a reaction mechanism, regioselectivity, chemoselectivity, stereoselectivity, an outcome of a chemical reaction at least one atom or bond removed or added, a relevance of a historical reaction, an electronic effect, chemical correctness, feasibility, stability, toxicity, yield, complexity, synthetic accessibility, practical utility, cost of starting materials or intermediates, reaction temperature, solvent identity, catalyst identity, reagent identity, or reaction atmosphere, mechanistic plausibility, likelihood of elementary reaction steps, or reliability of predicted reaction pathways, chemical impact of functional group addition, removal, or substitution in reactants or products, confidence, uncertainty, or prediction reliability, or chemical similarity to historically validated reactions.

[0213] Embodiment 14. The system of embodiment 10, or any other embodiment, wherein the refining, by the optimization module, the at least one machine learning model, further includes the optimization module changing the at least one machine learning model by changing how strongly two modules within the at least one machine learning model interact.

[0214] Embodiment 15. The system of embodiment 10, or any other embodiment, wherein the evaluation of each generated example chemical reaction is through only theperformance of a laboratory experiment or quantum mechanical simulation.

[0215] Embodiment 16. The system of embodiment 10, or any other embodiment10, further comprising: a data storage unit containing the historical chemical data and accessible by the at least one processors for using in training the at least one machine learning model; and a communication link between the computing system and one or both of a quantum mechanical simulation engine, or laboratory instrumentation for creating chemistry feedback.

[0216] Embodiment 17. A non-transitory computer-readable medium provided with instructions, which when executed by at least one processor of a system, cause the system to perform actions including: computationally generating at least one machine learning model including adjustable weights and being configured to output predictions regarding chemical reactions; training the at least one machine learning model on historical chemical data; generating, by a reaction evaluation module, at least one example chemical reaction; directing, by the reaction evaluation module, the evaluation of each generated example chemical reaction through at least one of: a performance of a laboratory experiment, a quantum mechanical simulation, or using a user interface to extract at least one answer from a human; receiving chemistry feedback for each example chemical reaction comprising: a prediction regarding the example chemical reaction, and evidence supporting the prediction; generating a chemistry feedback dataset using the chemistry feedback; evaluating the at least one machine learning model by computing an evaluation score configured to measure agreement between outputs of the at least one machine learning model and the chemistry feedback dataset, wherein computing the evaluation score includes: deriving virtual chemical reactions from the chemistry feedback dataset;inputting the virtual chemical reactions into the at least one machine learning model; receiving output from the at least one machine learning model based on the virtual chemical reactions; comparing the output to the chemistry feedback dataset to determine consistencies and / or inconsistencies between the output and the chemistry feedback dataset; and quantifying the determined consistencies and / or inconsistencies as the evaluation score; and refining, by an optimization module, the at least one machine learning model to create at least one refined machine learning model by changing at least one numerical parameter of the at least one machine learning model, wherein the at least one numerical parameter is changed so that a predicted evaluation score is greater than the first evaluation score, indicating higher measured agreement between outputs of the at least one refined machine learning model and the chemistry feedback dataset.

[0217] Embodiment 18. The non-transitory computer-readable medium of embodiment 17, or any other embodiment, the actions further comprising: receiving input including a target molecule; proposing, using the at least one refined machine learning model, a synthesis path for the target molecule; and directing a laboratory to synthesize the target molecule using the synthesis path.

[0218] Embodiment 19. The non-transitory computer-readable medium of embodiment 18, or any other embodiment, the actions further including at least one of: generating a fully specified chemical reaction based on a partially specified chemical reaction; or outputting one or more predictions when input a chemical reaction that indicate observed properties of the chemical reaction.

[0219] Embodiment 20. The non-transitory computer-readable medium of embodiment 17, or any other embodiment, wherein the refining, by the optimization module, the at least one machine learning model, further includes the optimization module changing the at least one machine learning model by changing how strongly two modules within the at least one machine learning model interact.

[0220] While the embodiments have been described with regards to particular embodiments, it is recognized that additional variations may be devised without departing from the inventive concept.

[0221] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the claimed subject matter. As used herein, the term “and / or” includes any and all combinations of one or more of the associated listed items. As used herein, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well as the singular forms, unless the context clearly indicates otherwise. It will further be understood that the terms “comprises” and / or “comprising,” when used in this specification, specify the presence of states features, steps, operations, elements, and / or components, but do not preclude the present or addition of one or more other features, steps, operations, elements, components, and / or groups thereof.

[0222] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one having ordinary skill in the art to which the embodiments belong. It will further be understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and the present disclosure and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.

[0223] In describing the embodiments, it will be understood that a number of elements, techniques, and steps are disclosed. Each of these has individual benefit and each can also be used in conjunction with one or more, or in some cases all, of the other disclosed elements, or techniques. The specification and claims should be read with the understanding that such combinations are entirely within the scope of the embodiments and the claimed subject matter.

[0224] In the description above and throughout, numerous specific details are set forth in order to provide a thorough understanding of an embodiment of this disclosure. It will be evident, however, to one of ordinary skill in the art, that an embodiment may be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form to facilitate explanation. The description of the preferred embodiments is not intended to limit the scope of the claims appended hereto. Further, in the methods disclosed herein, various steps are disclosed illustrating some of the functions of an embodiment. These steps are merely examples and are not meant to be limiting in any way. Other steps andfunctions may be contemplated without departing from this disclosure or the scope of an embodiment.

Claims

CLAIMSWhat is claimed is:

1. A computer- implemented method comprising: computationally generating, by a computing system including at least one processor and memory, at least one machine learning model including adjustable weights and being configured to output predictions regarding chemical reactions; training, by the computing system, the at least one machine learning model on historical chemical data; generating, by a reaction evaluation module executed by the computing system, at least one example chemical reaction; directing, by the reaction evaluation module, the evaluation of each generated example chemical reaction through at least one of: a performance of a laboratory experiment, a quantum mechanical simulation, or using a user interface to extract at least one answer from a human; receiving, by the computing system, chemistry feedback for each example chemical reaction comprising: a prediction regarding the example chemical reaction, and evidence supporting the prediction; generating, by the computing system, a chemistry feedback dataset using the chemistry feedback; evaluating, by the computing system, the at least one machine learning model by computing an evaluation score configured to measure agreement between outputs of the at least one machine learning model and the chemistry feedback dataset, wherein computing a first evaluation score includes: deriving virtual chemical reactions from the chemistry feedback dataset; inputting the virtual chemical reactions into the at least one machine learning model; receiving output from the at least one machine learning model based on the virtual chemical reactions; comparing the output to the chemistry feedback dataset to determine consistencies and / or inconsistencies between the output and the chemistry feedback dataset; andquantifying the determined consistencies and / or inconsistencies as the evaluation score; and refining, by an optimization module executed by the computing system, the at least one machine learning model to create at least one refined machine learning model by changing at least one numerical parameter of the at least one machine learning model, wherein the at least one numerical parameter is changed to cause a predicted evaluation score to be greater than the first evaluation score, indicating higher measured agreement between outputs of the at least one refined machine learning model and the chemistry feedback dataset.

2. The computer-implemented method of claim 1, further comprising: receiving, by the computing system, input including a target molecule; proposing, by the computing system using the at least one refined machine learning model, a synthesis path for the target molecule; and directing, by the computing system, a laboratory to synthesize the target molecule using the synthesis path.

3. The computer-implemented method of claim 2, wherein: the at least one machine learning model: generates a fully specified chemical reaction based on a partially specified chemical reaction; or outputs one or more predictions when input a chemical reaction, the one or more predictions indicating predicted properties of the chemical reaction.

4. The computer-implemented method of claim 1, wherein the evidence used to arrive at the prediction regards at least one of: a reaction mechanism, regioselectivity, chemoselectivity, stereoselectivity, an outcome of a chemical reaction at least one atom or bond removed or added, a relevance of a historical reaction, an electronic effect, chemical correctness, feasibility, stability, toxicity, yield, complexity, synthetic accessibility, practical utility, cost of starting materials or intermediates, reaction temperature, solvent identity, catalyst identity, reagent identity, or reaction atmosphere, mechanistic plausibility, likelihood of elementary reaction steps, or reliability of predicted reaction pathways, chemical impact of functional group addition,removal, or substitution in reactants or products, confidence, uncertainty, or prediction reliability, or chemical similarity to historically validated reactions.

5. The computer-implemented method of claim 1, wherein the refining, by the optimization module, the at least one machine learning model, further includes the optimization module changing the at least one machine learning model by changing how strongly two modules within the at least one machine learning model interact.

6. The computer-implemented method of claim 1 wherein the evaluation of each generated example chemical reaction is through only the performance of a laboratory experiment or quantum mechanical simulation.

7. The computer-implemented method of claim 6, wherein the evidence used to arrive at the prediction regards at least one of: a reaction mechanism, regioselectivity, chemoselectivity, stereoselectivity, an outcome of a chemical reaction at least one atom or bond removed or added, a relevance of a historical reaction, an electronic effect, chemical correctness, feasibility, stability, toxicity, yield, complexity, synthetic accessibility, practical utility, cost of starting materials or intermediates, reaction temperature, solvent identity, catalyst identity, reagent identity, or reaction atmosphere, mechanistic plausibility, likelihood of elementary reaction steps, or reliability of predicted reaction pathways, chemical impact of functional group addition, removal, or substitution in reactants or products, confidence, uncertainty, or prediction reliability, or chemical similarity to historically validated reactions.

8. The computer-implemented method of claim 1, wherein the at least one machine learning model is configured to include, with each output prediction, evidence supporting the output prediction.

9. The computer-implemented method of claim 1 , wherein the at least one machine learning model includes a module that outputs one or more predictions when input a chemical reaction, the one or more predictions indicating predicted properties of the chemical reaction, and wherein refining includes training the module to output predictions that are more similar to predictionsincluded in the chemistry feedback dataset.

10. A system for providing predictions regarding chemical reactions, the system comprising at least one processor and memory provided with instruction, which when executed configure the system to perform actions including: computationally generating at least one machine learning model including adjustable weights and being configured to output predictions regarding chemical reactions; training the at least one machine learning model on historical chemical data; generating, by a reaction evaluation module, at least one example chemical reaction; directing, by the reaction evaluation module, the evaluation of each generated example chemical reaction through at least one of: a performance of a laboratory experiment, a quantum mechanical simulation, or using a user interface to extract at least one answer from a human; receiving chemistry feedback for each example chemical reaction comprising: a prediction regarding the example chemical reaction, and evidence supporting the prediction; generating a chemistry feedback dataset using the chemistry feedback; evaluating the at least one machine learning model by computing an evaluation score configured to measure agreement between outputs of the machine learning model and the chemistry feedback dataset, wherein computing the evaluation score includes: deriving virtual chemical reactions from the chemistry feedback dataset; inputting the virtual chemical reactions into the at least one machine learning model; receiving output from the at least one machine learning model based on the virtual chemical reactions; comparing the output to the chemistry feedback dataset to determine consistencies and / or inconsistencies between the output and the chemistry feedback dataset; and quantifying the determined consistencies and / or inconsistencies as the evaluation score; and refining, by an optimization module, the at least one machine learning model to create at least one refined machine learning model by changing at least one numerical parameter of the at least one machine learning model, wherein the at least one numerical parameter is changed sothat a predicted evaluation score is greater than the first evaluation score, indicating higher measured agreement between outputs of the at least one refined machine learning model and the chemistry feedback dataset.

11. The system of claim 10, the actions further comprising: receiving input including a target molecule; proposing, using the at least one refined machine learning model, a synthesis path for the target molecule; and directing a laboratory to synthesize the target molecule using the synthesis path.

12. The system of claim 11, the actions further including at least one of: generating a fully specified chemical reaction based on a partially specified chemical reaction; or outputting one or more predictions when input a chemical reaction that indicate observed properties of the chemical reaction.

13. The system of claim 10, wherein the evidence used to arrive at the prediction regards at least one of: a reaction mechanism, regioselectivity, chemoselectivity, stereoselectivity, an outcome of a chemical reaction at least one atom or bond removed or added, a relevance of a historical reaction, an electronic effect, chemical correctness, feasibility, stability, toxicity, yield, complexity, synthetic accessibility, practical utility, cost of starting materials or intermediates, reaction temperature, solvent identity, catalyst identity, reagent identity, or reaction atmosphere, mechanistic plausibility, likelihood of elementary reaction steps, or reliability of predicted reaction pathways, chemical impact of functional group addition, removal, or substitution in reactants or products, confidence, uncertainty, or prediction reliability, or chemical similarity to historically validated reactions.

14. The system of claim 10, wherein the refining, by the optimization module, the at least one machine learning model, further includes the optimization module changing the at least one machine learning model by changing how strongly two modules within the at least one machine learning model interact.

15. The system of claim 10, wherein the evaluation of each generated example chemical reaction is through only the performance of a laboratory experiment or quantum mechanical simulation.

16. The system of claim 10, further comprising: a data storage unit containing the historical chemical data and accessible by the at least one processors for using in training the at least one machine learning model; and a communication link between the computing system and one or both of a quantum mechanical simulation engine, or laboratory instrumentation for creating chemistry feedback.

17. A non-transitory computer-readable medium provided with instructions, which when executed by at least one processor of a system, cause the system to perform actions including: computationally generating at least one machine learning model including adjustable weights and being configured to output predictions regarding chemical reactions; training the at least one machine learning model on historical chemical data; generating, by a reaction evaluation module, at least one example chemical reaction; directing, by the reaction evaluation module, the evaluation of each generated example chemical reaction through at least one of: a performance of a laboratory experiment, a quantum mechanical simulation, or using a user interface to extract at least one answer from a human; receiving chemistry feedback for each example chemical reaction comprising: a prediction regarding the example chemical reaction, and evidence supporting the prediction; generating a chemistry feedback dataset using the chemistry feedback; evaluating the at least one machine learning model by computing an evaluation score configured to measure agreement between outputs of the at least one machine learning model and the chemistry feedback dataset, wherein computing the evaluation score includes: deriving virtual chemical reactions from the chemistry feedback dataset; inputting the virtual chemical reactions into the at least one machine learning model; receiving output from the at least one machine learning model based on the virtual chemical reactions;comparing the output to the chemistry feedback dataset to determine consistencies and / or inconsistencies between the output and the chemistry feedback dataset; and quantifying the determined consistencies and / or inconsistencies as the evaluation score; and refining, by an optimization module, the at least one machine learning model to create at least one refined machine learning model by changing at least one numerical parameter of the at least one machine learning model, wherein the at least one numerical parameter is changed so that a predicted evaluation score is greater than the first evaluation score, indicating higher measured agreement between outputs of the at least one refined machine learning model and the chemistry feedback dataset.

18. The non-transitory computer-readable medium of claim 17, the actions further comprising: receiving input including a target molecule; proposing, using the at least one refined machine learning model, a synthesis path for the target molecule; and directing a laboratory to synthesize the target molecule using the synthesis path.

19. The non-transitory computer-readable medium of claim 18, the actions further including at least one of: generating a fully specified chemical reaction based on a partially specified chemical reaction; or outputting one or more predictions when input a chemical reaction that indicate observed properties of the chemical reaction.

20. The non-transitory computer-readable medium of claim 17, wherein the refining, by the optimization module, the at least one machine learning model, further includes the optimization module changing the at least one machine learning model by changing how strongly two modules within the at least one machine learning model interact.

Citation Information

Patent Citations

  • Intelligent personalized chemical synthesis planning

    US20200050947A1

  • System and method for feedback-driven automated drug discovery

    US20220351053A1

  • Systems and methods for predicting outcomes and conditions of chemical reactions with high reliability based on a highly diverse and accurate dataset

    US20230131234A1