Systems and methods for optimizing chemical reactions using machine learning

The integration of machine learning and cheminformatics for catalyst selection in chemical reactions addresses the inefficiencies of human intuition, enabling rapid identification of optimal catalysts and ligands, thus enhancing the efficiency and cost-effectiveness of catalyst discovery.

JP2025536310APending Publication Date: 2025-11-05MERCK PATENT GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025522192
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-17
Filing Date
2023-10-16
Publication Date
2025-11-05

AI Technical Summary

Technical Problem

Current catalyst selection in chemical reactions relies heavily on human intuition and limited experimentation, often resulting in suboptimal choices due to the neglect of a vast number of potential catalysts, leading to prolonged development times and increased costs.

Method used

A method utilizing machine learning and cheminformatics to parameterize catalysts, group them into clusters based on chemical features, and select representative catalysts for a test kit, followed by machine learning algorithms to predict optimal catalysts or ligands for chemical reactions.

Benefits of technology

This approach efficiently identifies optimal catalysts and ligands, reducing the need for extensive experimentation and accelerating the discovery of cost-effective, environmentally friendly catalysts for chemical reactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025536310000030
    Figure 2025536310000030
  • Figure 2025536310000031
    Figure 2025536310000031
  • Figure 2025536310000032
    Figure 2025536310000032
Patent Text Reader

Abstract

The present disclosure provides methods and systems for optimizing chemical reactions through machine learning. A chemical space is defined by grouping potential chemicals based on selected features. Representative chemicals are selected from each group and test kits are assembled. The test kits are then used to identify optimal catalysts or ligands for the catalyzed chemical reaction. In some embodiments, the system can recommend test kits, receive results from experiments, generate distance matrices, and rank potential chemicals based on scores obtained for the representative chemicals. These methods include reducing the dimensionality of chemical features and normalizing distances. The disclosed embodiments can suggest potential chemicals for optimizing chemical reactions by classifying chemicals based on scores obtained from the representative chemicals.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] background The present disclosure relates to the optimization of chemical reactions. In particular, the present disclosure relates to systems and methods that utilize test kits and machine learning to optimize catalytic chemical reactions, for example, by optimizing the selection of one or more catalysts and / or ligands using a design of experiments approach. [Background technology]

[0002] Related Fields 2. Description of Related Art Catalysis is the process of adding a non-consumable material to a chemical reaction to increase its rate, efficiency, or otherwise modify it to achieve desired process parameters or results. A catalyst is a substance that accelerates a chemical reaction and / or reduces the activation energy, temperature, or pressure required to initiate a chemical reaction. Catalysts are not consumed during the reaction and typically remain unchanged after the chemical reaction is complete. Small amounts of catalyst are often sufficient to promote a chemical reaction.

[0003] Approximately 90% of commercially produced chemical products involve a catalyst at some stage in the manufacturing process. Chemists in academia and industry worldwide routinely optimize catalytic reactions by varying catalyst materials, reactants, solvents, or reaction conditions to increase yield, efficiency, or cost-effectiveness.

[0004] In a process chemistry lab, a research chemist or research engineer designs a synthetic route to optimize a single catalytic chemical reaction for large-scale use in a manufacturing plant by conducting multiple test experiments with approximately 20 different catalysts. Finding the right catalyst can take weeks, resulting in high costs and further delays in bringing the product to market. In most cases, catalysts are selected to be safe, cost-effective, and environmentally friendly while increasing the yield of the reaction and suppressing unwanted side reactions.

[0005] Currently, catalyst selection relies primarily on human intuition and is reduced to the most common materials stocked by laboratories and chemical suppliers, ignoring the vast number of other materials in chemical space, some of which may be more effective than traditional choices. Thus, typically, only a small fraction of potential catalysts are screened, which does not necessarily lead to an optimal choice.

[0006] Therefore, to comply with and advance the state of the art, it is desirable to find new approaches for selecting optimal catalysts, ligands, or other chemicals from among a large catalog of chemicals without the need to experimentally test each chemical. Summary of the Invention

[0007] overview According to the present disclosure, a method for assembling a test kit for optimizing a catalytic chemical reaction via a computer using machine learning is disclosed. The method may include the following acts: parameterizing catalysts for the catalytic chemical reaction via a computer with respect to respective chemical features specific to the catalytic chemical reaction; grouping the parameterized catalysts into a predetermined number of clusters spanning the entire chemical space of the catalysts based on their chemical features and molecular descriptors; selecting one representative catalyst from all clusters using a computer according to certain defined criteria; and assembling a test kit using the selected representative catalyst as a component.

[0008] This approach differs from known prior art in particular in the following respects: 1. Integrating physical test kits into workflows to identify optimized ligands and catalysts without prior knowledge of the reaction. 2. Test kits can optionally be standardized to support a variety of reactions. 3. A clustering algorithm involves identifying the best set of catalysts or ligands for the test kit through a combination of chemical features and commercial feasibility and / or availability.

[0009] Advantageous and therefore exemplary further developments of the disclosure emerge from the relevant dependent claims as well as from the description and the associated drawings.

[0010] A further exemplary embodiment of the disclosed method includes that each chemical feature is determined via cheminformatics and computational modeling on a computer.

[0011] These exemplary further aspects of the disclosed methods include that the grouping is performed computationally via k-means clustering, density-based spatial clustering for noisy applications ("DBSCAN"), spectral clustering, Gaussian mixture models, or other clustering algorithms known to those skilled in the relevant art. For example, density-based, distribution-based, centroid-based, or hierarchical-based clustering algorithms can be used.

[0012] In another exemplary embodiment of the present disclosure, chemicals are selected based on a combination of chemical characteristics, commercial viability, and / or availability of the catalyst or ligand.

[0013] Another exemplary embodiment of the disclosed method includes storing all available catalysts or ligands in a database connected to a computer.

[0014] In another exemplary embodiment, a test kit having a specific, defined number of catalyst or ligand components assembled using the methods disclosed herein is used.

[0015] Another aspect of the present disclosure includes a method of optimizing a catalytic chemical reaction using a computer-supported test kit, comprising any of the following operations: conducting standardized experiments of a catalytic chemical reaction using components in the test kit; inputting result data from the conducted experiments into a computer; using a machine learning algorithm, such as, but not limited to, a clustering algorithm, a regression algorithm, a classification algorithm, an unsupervised machine learning algorithm, or a regression model, executed on the computer to interpolate between a predetermined number of clusters in the spanned chemical space of all available catalysts or ligands; using a machine learning regression model to predict an optimal catalyst for the catalytic chemical reaction within the interpolated parameter space of all available catalysts or ligands; and performing the catalytic chemical reaction using the predicted catalyst or ligand.

[0016] One further exemplary embodiment of the disclosed method includes providing, via the computer, a web interface for users to upload chemical reaction yield values ​​from conducted experiments.

[0017] It is understood that all aspects of the preferred further developments may be combined, even if not explicitly stated, unless this is clearly impossible due to the nature of the respective features.

[0018] A system including one or more computers can be configured to perform specific operations or actions by installing software, firmware, hardware, or a combination thereof. Such software, firmware, or hardware, when operated, causes the system to perform desired actions. Additionally, one or more computer programs can be designed to perform specific operations or actions by containing instructions that, when executed by a data processing device, cause the device to perform the desired actions.

[0019] In one general aspect, the method may include selecting a chemical space for grouping. Each chemical space may be defined by a plurality of potential chemicals. The method also includes selecting a plurality of chemical features, each chemical feature corresponding to a plurality of potential chemicals. Further, the method may include grouping the plurality of potential chemicals within the grouping space based on the plurality of chemical features. In addition, the method may include selecting a plurality of representative chemicals from the potential chemicals. Each representative chemical corresponds to a group of the plurality of potential chemicals grouped within the grouping space. Moreover, the method may include assembling a test kit having a plurality of test chemicals. Each test chemical corresponds to a representative chemical of the plurality of representative chemicals. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the method.

[0020] In one general aspect, the method may include an implementation that includes one or more of the following features as part of the method. The method may include a case where a grouping space is defined by a plurality of chemical features. The grouping space may be a dimensionally reduced space of the plurality of chemical features. The method also includes an act of generating a plurality of reduced chemical spaces, each reduced chemical space being a dimensionally reduced space of chemical space. The method may include calculating a plurality of distances, each distance of the plurality of distances being a distance between a prospective chemical of the plurality of prospective chemicals and a test chemical of the plurality of test chemicals. The plurality of test chemicals may be a subset of the plurality of prospective chemicals. The plurality of prospective chemicals may not include the plurality of test chemicals. The method may include calculating a plurality of distance metrics, each distance metric of the plurality of distance metrics being 1 minus a distance between one of the plurality of prospective chemicals and one of the plurality of test chemicals.

[0021] In one general aspect, the method may include an implementation that may include one or more of the following features as part of the method. The method may include averaging each of a plurality of distance metrics across all of the reduced chemical space to generate a plurality of averaged distance metrics. The method may include generating a plurality of weights, each weight of the plurality of weights corresponding to one of a plurality of outcomes, and each result of the plurality of outcomes corresponding to one of a plurality of test chemicals. The plurality of weights may be determined according to an exponential function. Each of the plurality of distances of the distance metrics may be multiplied by a respective weight of the plurality of weights. The maximum value of each prospective chemical may be taken across all of the multiplied plurality of distance metrics between each prospective chemical and the plurality of test chemicals. The method may include ranking the maximum value of each prospective chemical. The method may include testing a plurality of test chemicals to determine a plurality of outcomes, each of the plurality of outcomes corresponding to a respective outcome of the test chemical of the plurality of test chemicals.

[0022] In one general aspect, the method may include an implementation that may include one or more of the following features as part of the method: The plurality of results may be uploaded to a server. The method may include selecting a prediction space utilizing a plurality of predictive features to provide a prediction in the outcome space, the plurality of results defining data points in the prediction space that are mapped to the outcome space. The outcome space may include a chemical yield. The outcome space may be a by-product metric. The outcome space may be a single parameter. The prediction space may include a grouping space. The method may include selecting a chemical from a plurality of likely chemicals that correspond to an optimized value in the outcome space. The prediction space, in some aspects thereof, may include a plurality of chemical features therein. The computer may be configured to fit the plurality of results using regression fitting to map the plurality of predictive features to the outcome space. The computer may be configured to train a learner using the plurality of results.

[0023] In one general aspect, the method may include an implementation that may include one or more of the following features as part of the method. The learner may include at least one of an artificial neural network, a k-nearest neighbor, a decision tree, a random forest, a support vector machine, a Bayesian regressor, and / or an ensemble. The plurality of predictive features may include at least one of space-filling features, bulk features, orientation features, electrical features, Vin, frontier Mos, Fukui function, NBO analysis, NMR tensors, steric properties, Sterimol L, B1, B5, B1, B5, quadrant analysis, octant analysis, total volume, buried volume, dipole moment, solvation energy, and potential variance. The plurality of chemical features may include a plurality of chemical featurizations. The plurality of chemical features may include a plurality of molecular descriptors. The chemical space may include a plurality of catalysts. The chemical space may include a plurality of ligands. The plurality of chemical features may be determined via at least one of cheminformatics and computational modeling on a computer. The plurality of chemical signatures may be stored in a computer-implemented database.

[0024] In one general aspect, the method may include an implementation that may include one or more of the following features as part of the method: testing a plurality of test chemicals to determine a plurality of outcomes, each of the plurality of outcomes corresponding to an outcome for a respective test chemical among the plurality of test chemicals; selecting a prediction space utilizing a plurality of prediction features to provide a prediction in an outcome space, the plurality of outcomes defining data points in the prediction space; mapping the plurality of prediction features to the outcome space using regression; determining a best catalyst or best ligand from among all catalysts or ligands available for the catalytic chemical reaction in the prediction space according to the prediction; and / or performing the catalytic chemical reaction using the best catalyst or best ligand according to the prediction.

[0025] In one general aspect, the method may include an implementation that may include one or more of the following features as part of the method: The plurality of results may be uploaded to a server. The method may include selecting a prediction space utilizing a plurality of prediction features to provide a prediction in the outcome space, the plurality of results defining data points within the prediction space. The computer may be configured to map the plurality of prediction features to the outcome space using regression by fitting the plurality of results to the prediction features and the outcome space. The grouping space may be generated by performing dimensional reduction on the plurality of chemical features. The dimensional reduction may be implemented on a computer using principal component analysis. The chemical space may be defined by a plurality of phosphine ligands for cross-coupling reactions. The cross-coupling reaction may include either a Suzuki catalyst or a Buchwald catalyst. The plurality of test chemicals may include 24 chemicals. The plurality of chemical features can include at least one of space-filling features, bulk features, orientation features, electrical features, Vin, frontier Mos, Fukui functions, NBO analysis, NMR tensors, steric properties, Sterimol L, B1, B5, B1, B5, quadrant analysis, octant analysis, total volume, buried volume, dipole moment, solvation energy, and / or potential dispersion. Implementations of the described techniques can include hardware, methods, or processes, or a computer tangible medium.

[0026] In one general aspect, the method can include assembling a test kit having a plurality of test chemicals, each test chemical corresponding to a representative chemical of a plurality of representative chemicals; testing the plurality of test chemicals to determine a plurality of results, each of the plurality of results corresponding to a respective result for a respective test chemical of the plurality of test chemicals; and determining a best catalyst or best ligand of all catalysts or ligands available for catalyzing a chemical reaction within the prediction space according to the prediction. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs stored on one or more computer storage devices, each configured to perform the method.

[0027] Implementations may include one or more of the following features: the method may include generating a plurality of reduced chemical spaces, each reduced chemical space being a dimensionally reduced space of the chemical space; calculating a plurality of distances, each distance of the plurality of distances being a distance between a prospective chemical of the plurality of prospective chemicals and a test chemical of the plurality of test chemicals; calculating a plurality of distance metrics, each distance metric of the plurality of distance metrics being 1 minus the distance between one of the plurality of prospective chemicals and one of the plurality of test chemicals; and / or averaging each of the plurality of distance metrics across all of the reduced chemical spaces to generate a plurality of averaged distance metrics.

[0028] In one general aspect, the method may include an implementation that may include one or more of the following features as part of the method. The method may include generating a plurality of weights, each weight of the plurality of weights corresponding to one of a plurality of outcomes, and each outcome of the plurality of outcomes corresponding to one of a plurality of test chemicals. The plurality of weights may be determined according to an exponential function. Each of the plurality of distances of the distance metrics may be multiplied by a respective weight of the plurality of weights. The maximum value for each prospective chemical may be taken over all of the multiple multiplied distance metrics between each prospective chemical and the plurality of test chemicals.

[0029] In one general aspect, the method may include an implementation that may include one or more of the following features as part of the method. The method may include ranking the maximum value of each potential chemical to thereby find the best catalyst or best ligand; selecting a prediction space utilizing a plurality of prediction features to provide a prediction in an outcome space, where the plurality of outcomes define the data points in the prediction space; and / or using regression to fit the data points in the plurality of prediction features to the outcome space. The plurality of prediction features may include at least one of space-filling features, bulk features, orientation features, electrical features, Vin, Frontier Mos, Fukui function, NBO analysis, NMR tensors, steric properties, Sterimol L, B1, B5, B1, B5, quadrant analysis, octant analysis, total volume, buried volume, dipole moment, solvation energy, and potential dispersion, etc. The method may include performing the catalytic chemical reaction using the best catalyst or best ligand according to the prediction. Implementations of the described techniques may include hardware, methods or processes, or computer tangible media.

[0030] In one general aspect, the method can include parameterizing, via a computer, catalysts or ligands of each catalytic chemical reaction with respect to respective chemical features specific to the catalytic chemical reaction. The method can also include grouping the parameterized catalysts or ligands into a predetermined number of clusters spanning the chemical space of catalysts or ligands based on their chemical features. Further, the method can include using a computer to select one representative catalyst or ligand from each cluster according to predetermined criteria. In addition, the method can also include assembling a test kit using the selected representative catalysts or ligands as components. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs stored on one or more computer storage devices, each configured to perform the actions of the method.

[0031] Implementations may include one or more of the following features: Each chemical feature may be determined via cheminformatics and computational modeling on a computer. The method may use a computer to perform grouping using k-means clustering. Selection may be based on a combination of the chemical features of the catalyst or ligand and commercial feasibility and / or availability as predetermined criteria. All available catalysts and / or ligands may be stored in a database connected to the computer.

[0032] In one general aspect, the method may include an implementation that includes one or more of the following features as part of the method. The method may further include conducting standardized experiments of catalytic chemical reactions using components in a test kit; inputting result data from the conducted experiments into a computer; interpolating between a predetermined number of clusters within a spanned chemical space of all available catalysts or ligands using a machine learning regression model executed on the computer; using the machine learning regression model to predict an optimal catalyst for the catalytic chemical reaction within the spanned chemical space of all available catalysts or ligands; and / or conducting the catalytic chemical reaction using the predicted catalyst or ligand. A web interface may be provided via the computer through which a user uploads result data from the conducted experiments. Implementations of the described techniques may include hardware, methods, or processes, or tangible computer media.

[0033] In one general aspect, the system disclosed herein can: select chemical spaces for grouping, each chemical space defined by a plurality of potential chemicals; select a plurality of chemical features, each chemical feature corresponding to a plurality of potential chemicals, and group the plurality of potential chemicals in the grouping space based on the plurality of chemical features; select a plurality of representative chemicals from the potential chemicals, each representative chemical corresponding to a group of the plurality of potential chemicals grouped in the grouping space; and / or recommend a test kit having a plurality of test chemicals, each test chemical corresponding to a representative chemical of the plurality of representative chemicals. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the method.

[0034] In one general aspect, the disclosed system can receive multiple results from multiple experiments performed using multiple test chemicals, each of the multiple results corresponding to a respective result for each of the multiple test chemicals. The system can also recommend a best catalyst or best ligand from among all catalysts or ligands available for the catalyzed chemical reaction within the prediction space according to the prediction. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs stored on one or more computer storage devices, each configured to perform the actions of the method.

[0035] In one general aspect, the method can include conducting an initial screening experiment on a set of test chemicals to generate results. The method can also include obtaining chemical features that describe the properties of the potential chemicals. The method can further include reducing the dimensionality of the chemical features to generate a predetermined number of distinct chemical spaces. In addition, the method can include calculating, for a given potential chemical, a distance between each of the predetermined number of chemical spaces and each of the test chemicals. Moreover, the method can include normalizing the distances for each of the predetermined number of chemical spaces to a [0, 1] interval. The method can also include subtracting the normalized distances from 1 to generate distance metrics for each potential chemical and each test chemical within each chemical space.

[0036] In one general aspect, the method may include an implementation that may include one or more of the following features as part of the method. Further, the method may include averaging distance metrics across all chemical space for each prospective chemical. The method may further include normalizing results obtained from the initial screening experiments to [0,1] and converting them to weights. Moreover, the method may include multiplying the weights by the distance metrics to obtain a weighted distance metric, the distance metric being generated via the averaging operation and having a prospective chemical axis and a test chemical axis. The method may also include taking the maximum value along the prospective chemical axis to obtain a score for each prospective chemical. Furthermore, the method may include ranking the N prospective chemicals from highest to lowest based on the obtained scores. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the method.

[0037] Implementations may include one or more of the following features: The chemical features may be DFT-based features. The chemical features may be reduced using PCA, spatial PCA, kernel PCA with an RBF kernel, kernel PCA with a cosine kernel, Fast ICA, spectral embedding, Isomap, or local linear embedding. The results obtained from the initial screening experiments may be yield or enantioselectivity. To obtain a weighted distance metric, the weights obtained from the normalized results are multiplied column-wise by the distance metric. The chemical features are properties of the phosphine ligand. Implementations of the described techniques may include hardware, methods, or processes, or a tangible computer medium.

[0038] In one general aspect, the method can include obtaining a list of N potential chemicals; obtaining chemical features configured to describe properties of the N potential chemicals; clustering the N potential chemicals based on their chemical features to obtain a set of representative chemicals, where the set of representative chemicals defines a set of test chemicals; calculating a distance between each test chemical and each representative chemical based on the chemical features using a distance metric; normalizing the distances to a [0,1] interval; subtracting the normalized distances from 1 to generate distance metrics for each representative chemical and each test chemical; normalizing results obtained from an initial screening experiment to a [0,1] interval and converting them to weights; multiplying the weights by the distance metric to obtain weighted distance metrics; taking the maximum value along the representative chemical axis to obtain a score for each representative chemical; and / or ranking the N potential chemicals based on the scores obtained for their representative chemicals. Other embodiments of this aspect include corresponding computer systems, apparatuses, and computer programs stored on one or more computer storage devices, each configured to perform the actions of the method.

[0039] Implementations may include one or more of the following features: A computer program product may be implemented as the systems or methods described herein (such as a computer-readable medium and / or a non-transitory computer-readable medium). In some embodiments, distances may be calculated after dimensional reduction of the chemical features. Implementations of the described techniques may include hardware, methods, or processes, or tangible computer media.

[0040] In one general aspect, the method can include obtaining a list of N potential chemicals. The method can also include receiving results for the M potential chemicals, thereby defining test chemicals; obtaining chemical signatures configured to describe properties of the N potential chemicals; determining a score for each representative chemical; and ranking the N potential chemicals based on the scores obtained for the representative chemicals. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the method.

[0041] An implementation may include one or more of the following features: determining a score for each representative chemical may include calculating a distance between each test chemical and each representative chemical using a distance metric for each of a plurality of dimensionally reduced spaces of the chemical feature; normalizing the distance to a [0, 1] interval for each of the plurality of dimensionally reduced spaces; subtracting the normalized distance from 1 to generate a distance metric for each representative chemical and each test chemical for each of the plurality of dimensionally reduced spaces; averaging the distance metric for each representative chemical and each test chemical across all dimensionally reduced spaces; normalizing the results obtained from the received results of the M potential chemicals to [0, 1] and converting them to weights; multiplying the weight by the distance metric to obtain a weighted distance metric, the distance metric having the representative chemical axis and the test chemical axis; and / or taking the maximum value along the representative chemical axis to obtain a score for each representative chemical. Test chemicals may be removed from the representative chemical axis. Implementations of the described techniques may include hardware, methods or processes, or computer tangible media. [Brief explanation of the drawings]

[0042] These and other aspects will become more apparent from the detailed description of the various embodiments of the present disclosure, taken in conjunction with the drawings.

[0043] [Figure 1] FIG. 1 shows a schematic diagram of a system for optimizing chemical reactions using machine learning, according to an embodiment of the present invention.

[0044] [Figure 2] FIG. 2 illustrates the clustering of phosphine ligands according to an embodiment of the present disclosure.

[0045] [Figure 3] FIG. 3 illustrates the clustering of phosphine ligands based on molecular characterization and a set of molecular descriptors, according to an embodiment of the present disclosure.

[0046] [Figure 4] FIG. 4 illustrates the use of a physics kit containing 24 phosphine ligands from each cluster representing the entire chemical space, according to an embodiment of the present disclosure.

[0047] [Figure 5] FIG. 5 illustrates an artificial intelligence-based web platform for suggesting one or more optimal catalysts for use in a reaction with a particular set of reactants, according to an embodiment of the present disclosure.

[0048] [Figure 6] FIG. 6 illustrates a system for selecting one or more optimal chemicals according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0049] Detailed Description FIG. 1 shows a schematic diagram of a system 7 for using assembled test kits to find optimal catalysts for a desired catalytic chemical reaction. The catalyst selection system 7 comprises a web-based platform 5, realized for example as a homepage, which runs on a computer 3 in the form of a server and is accessible by a user 1 via a web browser on a remote device, such as a mobile phone or personal computer. The web-based platform 5 thus provides a user interface 4, preferably a graphical user interface (GUI), through which the user 1 can input the results of experiments performed using the assembled test kit. Furthermore, the web-based platform can be configured to provide functionality for initially assembling a test kit via the homepage depending on information provided by the user 1, e.g., the desired catalytic chemical reaction. Using this information and further data regarding the availability of specific catalysts or ligands, the system 7 can calculate the components of the test kit and submit a request for assembly of the test kit.

[0050] For example, in one embodiment of the present disclosure, the "Phosphine Predictor" suggests phosphine ligands for cross-coupling reactions, e.g., Suzuki and Buchwald catalysts. It suggests which commercially available monodentate phosphine ligands to try for CC or CN cross-coupling reactions of specific reaction substrates. Users run reactions using the provided and specified universal training set ligands (test kits) and enter the results (which may be yield, efficiency, or cost-effectiveness) into a secure Phosphine Predictor portal (e.g., via user interface 4). The optimal ligand suggestions are uploaded to the user account.

[0051] System 7 also includes a machine learning algorithm 6, as described herein. The machine learning algorithm 6 may be designed to learn patterns and make predictions or decisions based on the data provided. The machine learning algorithm 6 may use a data set of various chemicals, as described herein. This model may be used to make predictions for hypothetical chemical reactions. For new inputs, the machine learning algorithm may make predictions based on previously learned patterns or using any algorithm, as described herein. System 7 may use a CPU and / or GPU to parallelize calculations on computer 3.

[0052] This exemplary embodiment is further introduced by describing method acts that further illustrate specific examples of the disclosed method. In one embodiment, the method is included in the design of an AI-based experimentation platform. This embodiment is not limited to the hardware disclosed and used herein.

[0053] FIG. 2 provides an overview of all three steps of the disclosed method: clustering of phosphine ligands based on a set of molecular characterizations and molecular descriptors (Step 1), a physical kit containing 24 phosphine ligands from each cluster (Step 2), and an AI-based web platform that suggests the best catalyst for reactions with a specific set of reactants (Step 3).

[0054] Figure 3 shows the results of clustering phosphine ligands via a k-means clustering approach to screen all ligands available on the web platform or homepage used. Catalysts / ligands are parameterized by chemical features resulting from cheminformatics and computational modeling spanning chemical space. The catalysts or ligands are then grouped into 24 clusters based on their chemical features and molecular descriptors.

[0055] In FIG. 4, one representative catalyst / ligand from each of the 24 clusters may be included in a physical kit that is available for purchase via a web-based platform 5. Selection is based on availability. A potential customer acquires the kit and performs test experiments using all 24 representative catalysts / ligands from the physical kit for a particular catalytic chemical reaction of interest, while holding all other reaction conditions constant if possible.

[0056] As seen in Figure 5, the customer inputs yield values ​​for 24 chemical reactions into a web portal or web-based platform interface. Based on the input, the system performs a regression, interpolates between 24 clusters in the catalyst / ligand chemical space, and recommends the best possible catalyst / ligand for the particular reaction.

[0057] FIG. 6 shows a system 600 that suggests promising chemicals that can potentially optimize a chemical reaction based on a combination of distance metrics and performance results. System 600 may be implemented on computer 3 of FIG. 1. System 600 can be implemented in hardware, software, software executed by a processor and / or GPU, etc. In some embodiments, system 600 may be implemented in the cloud, as a software-as-a-service platform, as a distributed system, etc. System 600 includes an interface component 614 that retrieves chemical signatures 616 from a database 612 and a web interface component 602 that receives results 604 from test chemicals uploaded by a user.

[0058] The interface component 614 obtains chemical features 616 that describe a predetermined set of properties of the potential chemical compound. These chemical features 616 can be obtained using a standard interface, such as a web interface or a REST API. In a specific embodiment, the chemical features used were DFT-based features described in Gensch, T.; dos Passos Gomes, G.; Friederich, P.; Peters, E.; Gaudin, T.; Pollice, R.; Jorner, K.; Nigam, A.; Lindner-D'Addario, M.; Sigman, MS; Aspuru-Guzik, A.; A Comprehensive Discovery Platform for Organophosphorus Ligands for Catalysis. J. Am. Chem. Soc. 2022, 144, 3, 1205-1217, the contents of which are incorporated herein by reference in their entirety. However, any other type of chemical feature that describes the properties of the potential chemical compound can also be used. For example, the properties of the phosphine ligands involved can also be considered to optimize chemical reactions utilizing these chemicals (such as catalysts).

[0059] Due to the high dimensionality of the chemical features 616, dimensionality reduction is applied to the chemical features by a dimensionality reduction component 618. In certain embodiments of the present disclosure, dimensionality reduction is applied nine times using different techniques to generate nine different chemical spaces 609. For ease of visualization, two chemical spaces 608, 610 are displayed spanning space p to space q. Any suitable number of chemical spaces 609 can be used. In alternative embodiments, other dimensionality reduction techniques, such as autoencoders and t-SNE, may also be employed. The goal of dimensionality reduction is to reduce the number of features to a manageable number that can be used to efficiently and effectively compare the properties of potential chemicals.

[0060] As mentioned, the chemical features 616 are reduced to nine different embeddings 620 that can be mapped to nine different chemical spaces 609 using a spatial mapping component 622. Dimensionality reduction can be performed using PCA, spatial PCA, kernel PCA with an RBF kernel, kernel PCA with a cosine kernel, Fast ICA, spectral embedding, Isomap, local linear embedding, and multidimensional scaling. In some particular embodiments, each of the embeddings 620 may be 20-dimensional.

[0061] Once the chemical features 616 have been reduced and mapped via the space mapping component 622, distances to a given set of test chemicals are calculated for each potential chemical in each generated chemical space 609. Distances are calculated using the reduced chemical features that describe properties of the potential chemical relevant to optimizing a chemical reaction. If there are 24 test chemicals and 9 spaces 609, then 24 x 9 distances are calculated, for a total of 216 distances for each potential chemical.

[0062] The distances in each of the nine chemical spaces are then normalized to the [0,1] interval (e.g., Euclidean distance, etc.). The normalization may be per space 609 (e.g., because absolute values ​​are highly dependent on the space). For example, in certain embodiments, there may be nine sets of MxN distances, and each set of MxN distances may be normalized to 0,1. Any normalization may be used, such as linear normalization, min-max normalization, etc. This allows for standardization of the distances and more easily comparable analysis of the results 604. All distances are then subtracted from 1 to produce a distance metric within each chemical space where the shortest distance is 1 and the longest distance is 0. In an alternative embodiment, other types of distance metrics may be used, such as Mahalanobis distance, Manhattan distance, Euclidean distance, etc. The choice of distance metric depends on the nature of the chemical properties being compared and the particular application. Additionally, any post-normalization technique, such as rescaling or normalization of the distances, may alternatively or optionally be employed.

[0063] Next, all distance metrics for the different embeddings (i.e., chemical space 609) of each potential chemical are averaged. This generates an NxM matrix, where N is the number of potential chemicals for which recommendations are made and M is the number of test chemicals (e.g., M=24). Each location in the matrix may represent the average (mean value) of the nine distance metrics across all spaces 609 from a particular potential chemical to a particular test chemical. In some alternative embodiments, this step can be performed using other averaging techniques, such as weighted averaging or geometric averaging.

[0064] Results 604 of experiments performed using test chemicals may be uploaded by the web interface component 602. Results 604 (e.g., yields) are also normalized to [0,1] and converted to weights using an exponential function of the form 2^(x-1), where x is the result. These weights are then multiplied column by column by the distance metrics. The maximum value along the prospective chemical axis is taken, resulting in a score ranging from 0 to 1 for each prospective chemical, representing the best distance / performance combination for the test chemical. These values ​​are considered scores 624. In an alternative embodiment, different score calculation methods, such as a weighted sum or product of distance and performance scores, can also be employed.

[0065] These scores 624 are then used to rank the N potential chemicals from highest to lowest, with higher values ​​indicating better predicted outcomes. The disclosed method provides a reliable and efficient way to suggest potential chemicals that can potentially optimize a chemical reaction. Taking the potential chemicals with the highest scores suggests another set of test chemicals, and the process can be repeated.

[0066] In further alternative embodiments, the disclosed methods can be employed in a variety of industries, including pharmaceuticals, materials science, and agrochemicals. The methods can also be used to predict the activity of chemical compounds other than potential chemicals, such as drug candidates and natural products. These chemical features can be obtained using a variety of sources, including in silico or in vitro assays. Furthermore, in addition to the nine different dimensionality reduction techniques used in this disclosure, other techniques, such as UMAP or nonnegative matrix decomposition, can also be employed.

[0067] Furthermore, an alternative to the distance metric calculation step is the use of machine learning models, such as regression models or neural networks. These models can predict the distance between prospective and test chemicals based on chemical features. The predictive accuracy of these models can be assessed using cross-validation or by testing on holdout data.

[0068] Another alternative aspect of the score calculation step is the incorporation of uncertainty measures such as confidence intervals or probability distributions for the scores. These measures can provide additional information about the reliability and robustness of the score predictions.

[0069] In conclusion, the disclosed method and system provide a versatile and global approach to predict the most promising prospective chemicals for optimizing chemical reactions. The method allows for the use of various chemical features and dimensionality reduction techniques, as well as alternative distance metrics and scoring methods. These aspects make the method applicable to various chemical industries and can improve the accuracy and reliability of predictions.

[0070] The present invention is further illustrated by the following examples, which should not be construed as limiting in any way. Those skilled in the art will recognize that various modifications, additions, and variations can be made to the present invention without departing from the spirit and scope of the invention, as defined in the appended claims. example

[0071] General information

[0072] Unless otherwise noted, all reagents and solvents were purchased from MilliporeSigma and used as received. 2'-Dicyclohexylphosphino-2-methoxy-1-phenylnaphthalene, cBRIDP, and di-tert-butyl(2',6'-dimethoxy-[1,1'-biphenyl]-2-yl)phosphine were purchased from Ambeed. VPhos, tri(m-tolyl)phosphine, 9-[2-(dicyclohexylphosphino)phenyl]-9H-carbazole, CPhos, trioctylphosphine, and bis(3,5-bis(trifluoromethyl)phenyl)(2',6'-bis(isopropoxy)-3,6-dimethoxybiphenyl-2-yl)phosphine were purchased from STREM. 3-(Diphenylphosphino)phenol and 2-(dicyclohexylphosphino)-2'-methoxybiphenyl were purchased from Combi-Blocks. 2-Diphenylphosphino-6-methylpyridine and tris(diethylamino)phosphine were purchased from TCI.

[0073] Flash chromatography was performed on either a Biotage Isolera™ or a Biotage Selekt system. 1 H NMR and 13 The compounds were characterized by C NMR. The NMR spectra were recorded on either a Varian 500 MHz instrument or a Bruker 500 MHz instrument. All 1 H NMR experiments are reported in δ units, parts per million (ppm), and were referenced to the signal of residual chloroform-d (7.26 ppm). 13 C NMR spectra are reported in ppm relative to chloroform-d (77.23 ppm) and all 1All GC analyses were performed on an Agilent 7820A gas chromatograph equipped with an FID detector using an SPB-1 fused silica column, 30µm x 250µm x 1µm (Cat. No. 24029). All reaction vials were prepared in a positive-pressure Vac Omni-Lab glovebox and reactions were performed in a Radley Mya4 reaction station under positive nitrogen pressure. WARNING! Pure phosphine may react with air and moisture and may be pyrophoric; however, pyrophoricity can be minimized when used as a solution. Follow all precautions outlined in the SDS.

[0074] Phosphine Kit Ligands [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4]

[0075] Ligand prediction model

[0076] The following describes a particular embodiment of the system 600 described above with reference to Figure 6. To understand how the ligand prediction / recommendation model described herein works, and with particular reference to Figure 6, we will use the term "kit ligand" (particular aspect of a test chemical) for any ligand that is in the set of 24 ligands on which initial screening tests are performed and that is ranked through the model for new experimental suggestions for this set (particular aspect of a prospective chemical), but will now use the term simply "ligand" for any ligand that is not part of this set.

[0077] The model (e.g., system 600) suggests ligands based on two criteria: distance to the kit ligand (how similar the ligand is to the kit ligand) and performance of that kit ligand (in this case, yield / conversion, but could also be enantioselectivity or other criteria).

[0078] Distances are obtained from molecular features / descriptors that describe the properties of the ligand. In this case, we used DFT-based features described in the Kraken paper (Gensch, T.; dos Passos Gomes, G.; Friederich, P.; Peters, E.; Gaudin, T.; Pollice, R.; Jorner, K.; Nigam, A.; Lindner-D'Addario, M.; Sigman, MS; Aspuru-Guzik, A.; A Comprehensive Discovery Platform for Organophosphorus Ligands for Catalysis. J. Am. Chem. Soc. 2022, 144, 3, 1205-1217). These features are available from the website (https: / / kraken.cs.toronto.edu / ) using a REST API or can be calculated by following the public code in the corresponding GitHub repository (https: / / github.com / aspuru-guzik-group / kraken). In principle, any other set of descriptors describing the properties of phosphine ligands relevant to catalysis can be used. Due to the high dimensionality of the features, we applied dimensionality reduction to generate an embedding with 20 components (e.g., embedding 620). Various techniques can be used for this purpose. A total of nine dimensionality reduction techniques were used: PCA, sparse PCA, kernel PCA with RBF and cosine kernels, Fast ICA, spectral embedding, Isomap, locally linear embedding, and MDS. Within each of these embeddings, the distances from each ligand to each kit ligand are calculated. For each embedding, these distances are normalized to the [0,1] interval as described above, and then the distances are subtracted from 1 as described above. The resulting distance is such that 1 corresponds to the closest ligand (the kit ligand identity) and 0 is the furthest ligand within the kit ligand. Finally, the average of all distances for the different embeddings is taken.This results in an NxM distance matrix, where N is the number of recommended ligands and M is the number of kit ligands (M=24).

[0079] Performance (herein yield / conversion) is also normalized to [0,1], where x is the performance, resulting in 24 weights for the 24 kit ligands, 2 x-1 These weights are then multiplied column-wise by the distance matrix and then the maximum value along the ligand axis is taken, resulting in a score for each ligand between 0 and 1 for the best distance / performance combination with respect to the kit ligand.

[0080] The final ranking of the N ligands is then used to suggest new experiments by taking the ligands with the highest scores from the above procedure. In principle, any list of N ligands for which features can be obtained or calculated can be taken. In this case, a list of approximately 400 monodentate phosphine ligands was used, which was also the basis for the initial clustering for the creation of the ligand kit.

[0081] Cross-coupling screening reaction

[0082] Example 1

[0083] General procedure for Buchwald-Hartwig CN cross-coupling reaction 1: [ka]

[0084] N-[2,6-bis(1-methylethyl)phenyl]-2,4,6-tris(1-methylethyl)benzenamine (1): Place a 4 mL screw-cap vial in a glovebox and add Pd(dba) (4.6 mg, 0.5 mmol, 0.5 mol%), phosphine ligand (0.1 mmol, 1.0 mol%), 4,4′-di-tert-butylbiphenyl (internal standard, 80 mg, 0.3 mmol, 0.3 equiv), NaO-t-Bu (144 mg, 1.5 mmol, 1.5 equiv), 2,4,6-triisopropylbromobenzene (0.25 mL, 1.0 mmol, 1.0 equiv), 2,6-diisopropylanaline (0.23 mL, 1.2 mmol, 1.2 equiv), and toluene (2.0 mL, 0.5 M). The vial was sealed with a rubber / Teflon septum, removed from the glovebox, and placed in a Radley Mya4 reaction station equipped with a stirrer preheated to 80 °C and set at 300 rpm. After 1 or 20 h of reaction time, a 50 μL sample was passed through a syringe filter and diluted with 1 mL of ethyl acetate, and the reaction mixture was analyzed by GC. After the reaction was judged complete by GC, the reaction mixture was diluted with ethyl acetate and filtered through a short plug of silica gel. After drying, the crude reaction mixture was dry loaded onto a Biotage flash chromatography column on silica gel (0–20% EtOAc / heptane) to obtain N-[2,6-bis(1-methylethyl)phenyl]-2,4,6-tris(1-methylethyl)benzenamine as a colorless solid.The product was confirmed by comparison with literature NMR spectral data (Raders, SM; Moore, JN; Parks, JK; Miller, AD; Leissing, TM; Kelley, SP; Rogers, RD; Shaughnessy, KH Trineopentylphosphine: A Conformationally Flexible Ligand for the Coupling of Sterically Demanding Substrates in the Buchwald-Hartwig Amination and Suzuki-Miyaura Reaction. J. Org. Chem. 2013, 78, 4649-4664).

[0085] 1 H NMR (CDCl3): δ 7.7 (d, J = 7.6 Hz, 2H), 6.98-6.93 (m, 3H), 4.77 (s, 1H), 3.16-3.2 (m, 4H), 3.7 (sept, J = 6.7 Hz, 1H), 1.24 (d, J = 6.9 Hz, 6H), 1.8 (t, J = 7.1 Hz, 24H).

[0086] 13 C NMR (CDCl3): δ 143.8, 141.9, 141.2, 140.1, 138.3, 124.0, 122.3, 121.8, 34.2, 28.1, 27.9, 24.5, 23.9, 23.8.

[0087] Ligand Screening Kit for Buchwald-Hartwig CN Cross-Coupling Reaction 1 a [ka] [Table 2]

[0088] Predicted ligand structure for Buchwald-Hartwig CN cross-coupling reaction 1 [Table 3-1] [Table 3-2] [Table 3-3]

[0089] Buchwald-Hartwig C-N cross-coupling reaction 1 a Predicted ligand screening results [Table 4]

[0090] Example 2

[0091] General procedure for Buchwald-Hartwig CN cross-coupling reaction 2: [ka]

[0092] 2-(1H-indol-1-yl)benzoxazole (2): A 4 mL screw-cap vial was placed in a glovebox and Pd(dba) (4.6 mg, 0.5 mmol, 0.5 mol%), phosphine ligand (0.1 mmol, 1.0 mol%), 4,4′-di-tert-butylbiphenyl (internal standard, 80 mg, 0.3 mmol, 0.3 equiv), NaO-t-Bu (144 mg, 1.5 mmol, 1.5 equiv), indole (141 mg, 1.2 mmol, 1.2 equiv), KPO (318 mg, 1.5 mmol, 1.5 equiv), toluene (2.0 mL, 0.5 M), and 2-chlorobenzoxazole (114 μL, 1.0 mmol, 1.0 equiv) were added. The vial was sealed with a rubber / Teflon septum, removed from the glovebox, and the reaction was placed in a Radley Mya4 reaction station equipped with a stirrer preheated to 80 °C and set at 300 rpm. After 16 h, a 50 μL sample was passed through a syringe filter and diluted with 1 mL of EtOAc, and the reaction mixture was analyzed by GC. After the reaction was judged complete by GC, the reaction mixture was diluted with ethyl acetate and filtered through a short plug of silica gel. After drying, the crude reaction mixture was dry loaded onto a Biotage flash chromatography column on silica gel (0–5% EtOAc / heptane) to yield 2-(1H-indol-1-yl)benzoxazole as a colorless solid. The product was confirmed by comparison with literature NMR spectral data (Li, D.-H.; Lan, X.-B.; Song, A.-X. Rahman, M.M.; Xu, C.; Huang, F.-D.; Szostak, R.; Szostak, M.; Liu, F.-S. Buchwald-Hartwig Amination of Coordinating Heterocycles Enabled by Large-but-Flexible Pd-BIAN-NHC Catalysts. Chem. Eur. J. 2022, 28, e202103341).

[0093] 1H NMR (CDCl3) δ 8.57 (d, J = 8.3 Hz, 1H), 7.88 (d, J = 3.6 Hz, 1H), 7.73 - 7.64 (m, 2H), 7.55 (d, J = 7.8 Hz, 1H), 7.49 - 7.41 (m, 1H), 7.33 (m, 3H), 6.79 (d, J = 3.6 Hz, 1H).

[0094] 13 C NMR (CDCl3) δ 154.8, 148.5, 141.5, 134.7, 130.2, 124.9, 124.7, 124.6, 123.6, 123.0, 121.3, 118.9, 114.6, 109.9, 108.5.

[0095] Ligand Screening Kit for Buchwald-Hartwig C-N Cross-Coupling Reaction 2 a [ka] [Table 5]

[0096] Predicted ligand structure for Buchwald-Hartwig CN cross-coupling reaction 2. [Table 6-1] [Table 6-2] [Table 6-3]

[0097] Buchwald-Hartwig C-N cross-coupling reaction 2 a Predicted ligand screening results [Table 7]

[0098] RuPhos a Optimization of reaction 2 by [ka] [Table 8]

[0099] Example 3

[0100] General procedure for Suzuki CC cross-coupling reaction 3: [ka]

[0101] 2-(2-thienyl)quinoxaline (3): Place a 4 mL screw-cap vial in a glovebox and add Pd(dba) (4.6 mg, 0.5 mmol, 0.5 mol%), phosphine ligand (0.1 mmol, 1.0 mol%), 4,4′-di-tert-butylbiphenyl (internal standard, 80 mg, 0.3 mmol, 0.3 equiv), 2-thienylboronic acid (192 mg, 1.5 mmol, 1.5 equiv), 2-chloroquinoxaline (165 mg, 1.0 mmol, 1.0 equiv), toluene (2.0 mL, 0.5 M), and EtN (420 μL, 3.0 mmol, 3.0 equiv). The vial was sealed with a rubber / Teflon septum, removed from the glovebox, and placed in a Radley Mya4 reaction station preheated to 100 °C with stirring set at 300 rpm. After 20 h, a 50 μL sample was passed through a syringe filter, diluted with 1 mL of EtOAc, and the reaction mixture was analyzed by GC. After the reaction was judged complete by GC, the reaction mixture was diluted with ethyl acetate and filtered through a short plug of silica gel. After drying, the crude reaction mixture was dry loaded onto a Biotage flash chromatograph on silica gel (0-30% EtOAc / heptane) to obtain 2-(2-thienyl)quinoxaline as a white solid. The product was confirmed by comparison with literature NMR spectral data (Knapp, DM; Gillis, EP; Burke, MD. A General Solution for Unstable Boronic Acids: Slow-Release Cross-Coupling from Air-Stable MIDA Boronates. J. Am. Chem. Soc. 2009, 131, 20, 6961-6963).

[0102] 1 H NMR (500.1 MHz, CDCl3): δ 9.25 (s, 1H), 8.9-8.7 (m, 2H), 7.88-7.87 (m, 1H), 7.78-7.69 (m, 2H), 7.56-7.55 (m, 1H), 7.23-7.21 (m, 1H).

[0103] 13 C NMR (CDCl3): δ 147.3, 142.2, 142.1, 142.0, 141.3, 130.4, 129.7, 129.1, 129.1, 128.4, 126.

[0104] Ligand Screening Kit 3 for Suzuki CC Cross-Coupling Reactions a [ka] [Table 9]

[0105] Predicted Ligand Structure for Suzuki CC Cross-Coupling Reaction 3 [Table 10-1] [Table 10-2] [Table 10-3]

[0106] Suzuki CC cross-coupling reaction 3 a Predicted ligand screening results. [Table 11]

[0107] Suzuki CC Cross-Coupling Reaction 3 RuPhos a Optimization by [ka] [Table 12]

[0108] Each of the features and examples described above, and combinations thereof, is said to be encompassed by the present disclosure. Accordingly, the present disclosure is drawn to the following non-limiting aspects:

[0109] (1) A method for optimizing chemical reactions using machine learning, comprising: (1) selecting a chemical space for grouping, each chemical space being defined by a plurality of potential chemicals; selecting a plurality of chemical features, each chemical feature corresponding to a plurality of potential chemicals; grouping the plurality of potential chemicals in a grouping space based on the plurality of chemical features; selecting a plurality of representative chemicals from the potential chemicals, each representative chemical corresponding to a group of the plurality of potential chemicals grouped in the grouping space; and assembling a test kit having a plurality of test chemicals, each test chemical corresponding to a representative chemical of the plurality of representative chemicals.

[0110] (2) The method of aspect 1, wherein the grouping space is defined by a plurality of chemical features.

[0111] (3) The method of aspect 1, wherein the grouping space is a dimensionally reduced space of multiple chemical features.

[0112] (4) The method of aspect 1, further comprising generating a plurality of reduced chemical spaces, each reduced chemical space being a dimensionally reduced version of the chemical space.

[0113] (5) The method of aspect 4, further comprising calculating a plurality of distances, each distance of the plurality of distances being a distance between a prospective chemical of the plurality of prospective chemicals and a test chemical of the plurality of test chemicals.

[0114] (6) The method of claim 5, wherein the plurality of test chemicals is a subset of the plurality of prospective chemicals.

[0115] (7) The method of aspect 5, wherein the plurality of prospective chemicals does not include a plurality of test chemicals.

[0116] (8) The method of aspect 4, further comprising calculating a plurality of distance metrics, each distance metric of the plurality of distance metrics being 1 minus the distance between one of the plurality of prospective chemicals and one of the plurality of test chemicals.

[0117] (9) The method of aspect 8, further comprising averaging each of the plurality of distance metrics over the reduced chemical space to generate a plurality of averaged distance metrics.

[0118] (10) The method of aspect 1, further comprising generating a plurality of weights, each weight of the plurality of weights corresponding to one of a plurality of outcomes, and each outcome of the plurality of outcomes corresponding to one of a plurality of test chemicals.

[0119] (11) The method of aspect 10, wherein the plurality of weights are determined according to an exponential function.

[0120] (12) The method of aspect 10, wherein each of the plurality of distances of the distance metric is multiplied by a respective weight of the plurality of weights.

[0121] (13) The method of aspect 12, wherein the maximum value for each prospective chemical is taken across multiple multiplied distance metrics between each prospective chemical and multiple test chemicals.

[0122] (14) The method of aspect 13, further comprising ranking the maximum value of each potential chemical.

[0123] (15) The method of aspect 1, further comprising testing a plurality of test chemicals to determine a plurality of outcomes, each of the plurality of outcomes corresponding to a respective outcome for a test chemical among the plurality of test chemicals.

[0124] (16) The method of claim 15, wherein the plurality of results are uploaded to a server.

[0125] (17) The method of aspect 15, further comprising selecting a prediction space utilizing a plurality of prediction features to provide a prediction in the outcome space, wherein the plurality of outcomes defines data points in the prediction space that are mapped to the outcome space.

[0126] (18) The method of aspect 17, wherein the result space includes a chemical yield.

[0127] (19) The method of aspect 17, wherein the result space is a by-product metric.

[0128] (20) The method of aspect 17, wherein the result space is a single parameter.

[0129] (21) The method of aspect 17, wherein the prediction space includes a grouping space.

[0130] (22) The method of aspect 17, wherein the prediction space includes multiple chemical features therein.

[0131] (23) The method of aspect 17, wherein the computer is configured to fit the plurality of outcomes using regression fitting to map the plurality of predictive features to an outcome space.

[0132] (24) The method of aspect 21, further comprising selecting a chemical from the plurality of potential chemicals that correspond to the optimized value in the outcome space.

[0133] (25) The method of aspect 17, wherein the computer is configured to train the learner using the multiple results.

[0134] (26) The method of aspect 25, wherein the learner is at least one of an artificial neural network, a k-nearest neighbor, a decision tree, a random forest, a support vector machine, a Bayesian regressor, and an ensemble.

[0135] (27) The method of aspect 1, wherein the plurality of chemical features comprises a plurality of chemical characterizations.

[0136] (28) The method of aspect 1, wherein the plurality of chemical features comprises a plurality of molecular descriptors.

[0137] (29) The method of aspect 1, wherein the chemical space comprises a plurality of catalysts.

[0138] (30) The method of aspect 1, wherein the chemical space includes a plurality of ligands.

[0139] (31) The method of aspect 1, wherein the plurality of chemical features are determined via at least one of cheminformatics and computational modeling on a computer.

[0140] (32) The method of aspect 1, wherein the plurality of chemical features is stored in a database implemented by a computer.

[0141] (33) The method of aspect 1, further comprising: testing a plurality of test chemicals to determine a plurality of outcomes, each of the plurality of outcomes corresponding to an outcome for a respective test chemical among the plurality of test chemicals; selecting a prediction space utilizing a plurality of predictive features to provide a prediction in an outcome space, the plurality of outcomes defining data points in the prediction space; mapping the plurality of predictive features to the outcome space using regression; determining a best catalyst or best ligand of all catalysts or ligands available for the catalyzed chemical reaction in the prediction space according to the prediction; and performing the catalyzed chemical reaction using the best catalyst or best ligand according to the prediction.

[0142] (34) The method of aspect 1, wherein the grouping space is generated by performing dimensionality reduction on a plurality of chemical features.

[0143] (35) The method of aspect 1, wherein the dimensionality reduction is performed on a computer using principal component analysis.

[0144] (36) The method of aspect 1, wherein the chemical space is defined by multiple phosphine ligands for the cross-coupling reaction.

[0145] (37) The method according to aspect 36, wherein the cross-coupling reaction involves one of a Suzuki catalyst and a Buchwald catalyst.

[0146] (38) The method of aspect 1, wherein the plurality of test chemicals consists of 24 chemicals.

[0147] (39) The method of aspect 1, wherein the plurality of chemical features includes at least one of space-filling features, bulk features, orientation features, electrical features, Vin, frontier Mos, Fukui function, NBO analysis, NMR tensors, steric properties, Sterimol L, B1, B5, B1, B5, quadrant analysis, octant analysis, total volume, buried volume, dipole moment, solvation energy, and potential dispersion.

[0148] (40) The method of aspect 1, wherein the plurality of predicted features include at least one of space-filling features, bulk features, orientation features, electrical features, Vin, frontier Mos, Fukui function, NBO analysis, NMR tensors, steric properties, Sterimol L, B1, B5, B1, B5, quadrant analysis, octant analysis, total volume, buried volume, dipole moment, solvation energy, and potential dispersion.

[0149] (41) A method of utilizing an experimental test kit, comprising: assembling a test kit having a plurality of test chemicals, each test chemical corresponding to a representative chemical of a plurality of representative chemicals; testing the plurality of test chemicals to determine a plurality of results, each of the plurality of results corresponding to a respective result for a respective test chemical of the plurality of test chemicals; and determining, according to the prediction, a best catalyst or a best ligand of all catalysts or ligands available for catalyzing a chemical reaction within the prediction space.

[0150] (42) The method of aspect 41, further comprising generating a plurality of reduced chemical spaces, each reduced chemical space being a dimensionally reduced version of the chemical space.

[0151] (43) The method of aspect 42, further comprising calculating a plurality of distances, each distance of the plurality of distances being a distance between a prospective chemical of the plurality of prospective chemicals and a test chemical of the plurality of test chemicals.

[0152] (44) The method of aspect 43, further comprising calculating a plurality of distance metrics, each distance metric of the plurality of distance metrics being 1 minus the distance between one of the plurality of prospective chemicals and one of the plurality of test chemicals.

[0153] (45) The method of aspect 44, further comprising averaging each of the plurality of distance metrics across the reduced chemical space to generate a plurality of average distance metrics.

[0154] (46) The method of aspect 41, further comprising generating a plurality of weights, each weight corresponding to one of a plurality of outcomes, and each outcome of the plurality of outcomes corresponding to one of a plurality of test chemicals.

[0155] (47) The method of aspect 46, wherein the plurality of weights are determined according to an exponential function.

[0156] (48) The method of aspect 46, wherein each of the plurality of distances of the distance metric is multiplied by a respective weight of the plurality of weights.

[0157] (49) The method of aspect 48, wherein the maximum value for each prospective chemical is taken across all of the multiplied distance metrics between each prospective chemical and the plurality of test chemicals.

[0158] (50) The method of aspect 49, further comprising ranking the maximum value of each potential chemical to thereby find the best catalyst or the best ligand.

[0159] (51) The method of aspect 41, further comprising selecting a prediction space utilizing a plurality of predictive features to provide predictions in an outcome space, wherein a plurality of outcomes define data points within the prediction space.

[0160] (52) The method of aspect 51, further comprising using regression to fit data points within a plurality of predictive traits to an outcome space.

[0161] (53) The method of aspect 51, wherein the plurality of predicted features include at least one of space-filling features, bulk features, orientation features, electrical features, Vin, frontier Mos, Fukui function, NBO analysis, NMR tensors, steric properties, Sterimol L, B1, B5, B1, B5, quadrant analysis, octant analysis, total volume, buried volume, dipole moment, solvation energy, and potential dispersion.

[0162] (54) The method of aspect 41, further comprising carrying out the catalytic chemical reaction using the best catalyst or best ligand according to the prediction.

[0163] (55) The method of aspect 33, wherein the plurality of results are uploaded to a server.

[0164] (56) The method of aspect 33, further comprising selecting a prediction space utilizing a plurality of prediction features to provide predictions in an outcome space, wherein a plurality of outcomes define data points in the prediction space.

[0165] (57) The method of aspect 56, wherein the computer is configured to map a plurality of predictor traits to an outcome space using regression by fitting the plurality of outcomes to the predictor traits and outcome space.

[0166] (58) A method for assembling a test kit for optimizing a catalytic chemical reaction via a computer, comprising: parameterizing, via a computer, catalysts or ligands of each catalytic chemical reaction with respect to respective chemical features specific to the catalytic chemical reaction; grouping the parameterized catalysts or ligands into a predetermined number of clusters spanning the chemical space of catalysts or ligands based on their chemical features; selecting, using a computer, one representative catalyst or ligand from each cluster according to predetermined criteria; and assembling the test kit using the selected representative catalysts or ligands as building blocks.

[0167] (59) The method of aspect 58, wherein each chemical feature is determined by cheminformatics and computational modeling on a computer.

[0168] (60) The method of aspect 58, further comprising using a computer to perform a grouping operation using k-means clustering.

[0169] (61) The method of aspect 58, wherein the selection is based on a combination of chemical characteristics of the catalyst or ligand and commercial feasibility and / or availability as predetermined criteria.

[0170] (62) The method of aspect 61, wherein all available catalysts or ligands are stored in a database connected to a computer.

[0171] (63) A test kit having a specific, predetermined number of catalyst or ligand components assembled using the method described in aspect 58.

[0172] (64) The method of aspect 58, further comprising: conducting a standardized experiment of a catalyzed chemical reaction using components in the test kit; inputting result data from the conducted experiment into a computer; interpolating between a predetermined number of clusters in a spanned chemical space of all available catalysts or ligands using a machine learning regression model executed on the computer; using the machine learning regression model to predict an optimal catalyst for the catalyzed chemical reaction in the spanned chemical space of all available catalysts or ligands; and conducting the catalyzed chemical reaction using the predicted catalyst or ligand.

[0173] (65) The method of aspect 64, wherein a web interface is provided via the computer through which a user uploads chemical reaction result data from conducted experiments.

[0174] (66) A system for optimizing chemical reactions implemented by an operating set of processor-executable instructions configured to execute on at least one processor, the at least one processor and the operating set of processor-executable instructions being configured to: select a chemical space for grouping, each chemical space defined by a plurality of potential chemicals; select a plurality of chemical features, each chemical feature corresponding to a plurality of potential chemicals, and group the plurality of potential chemicals in the grouping space based on the plurality of chemical features; select a plurality of representative chemicals from the potential chemicals, each representative chemical corresponding to a group of the plurality of potential chemicals grouped in the grouping space; and recommend a test kit having a plurality of test chemicals, each test chemical corresponding to a representative chemical of the plurality of representative chemicals.

[0175] (67) A system implemented by an operational set of processor-executable instructions configured to execute on at least one processor, the operational set of processor-executable instructions being configured to: recommend a test kit having a plurality of test chemicals, each test chemical corresponding to a representative chemical of a plurality of representative chemicals; receive a plurality of results from a plurality of experiments conducted using the plurality of test chemicals, each of the plurality of results corresponding to a respective result of a respective test chemical of the plurality of test chemicals; and recommend a best catalyst or a best ligand among all catalysts or ligands available for a catalyzed chemical reaction within a prediction space according to the prediction.

[0176] (68) Conducting an initial screening experiment on a set of test chemicals to generate results; obtaining chemical features characteristic of the potential chemicals; reducing the dimensionality of the chemical features to generate a predetermined number of distinct chemical spaces; calculating, for a given potential chemical, the distance between each of the predetermined number of chemical spaces and each of the test chemicals; normalizing the distances for each of the predetermined number of chemical spaces to the [0, 1] interval; subtracting the normalized distances from 1 to generate distance metrics for each potential chemical and each test chemical within each chemical space; and for each potential chemical, calculating the distance between all of the chemical spaces and each of the test chemicals. a chemical space of N potential chemicals; ...

[0177] (69) The method of aspect 68, wherein the chemical features are DFT-based features.

[0178] (70) Aspect 68 methods in which chemical signatures are reduced using PCA, spatial PCA, kernel PCA with an RBF kernel, kernel PCA with a cosine kernel, Fast ICA, spectral embedding, Isomap, or local linear embedding.

[0179] (71) The method of aspect 68, where the results obtained from the initial screening experiments are yield or enantioselectivity.

[0180] (72) The method of aspect 68, in which the weights obtained from the normalized results are multiplied column-wise by the distance metric to obtain a weighted distance metric.

[0181] (73) The method of aspect 68, wherein the chemical feature is a property of a phosphine ligand.

[0182] (74) A method for suggesting promising chemicals for optimizing a chemical reaction, comprising: obtaining a list of N promising chemicals; obtaining chemical features configured to describe the properties of the N promising chemicals; clustering the N promising chemicals based on their chemical features to obtain a set of representative chemicals, wherein the set of representative chemicals defines a set of test chemicals; calculating a distance between each test chemical and each representative chemical based on the chemical features using a distance metric; normalizing the distance to a [0, 1] interval; subtracting the normalized distance from 1 to generate a distance metric for each representative chemical and each test chemical; normalizing results obtained from an initial screening experiment to [0, 1] and converting them to weights; multiplying the weights by the distance metric to obtain a weighted distance metric; taking the maximum value along the representative chemical axis to obtain a score for each representative chemical; and ranking the N promising chemicals based on the scores obtained for the representative chemicals.

[0183] (75) A computer program product including a non-transitory computer-readable recording medium encoded with instructions for performing the steps of the method of aspect 74.

[0184] (76) A system for suggesting potential chemicals that can be used to optimize a chemical reaction, comprising a computer system configured to implement the method of aspect 74.

[0185] (77) The method of aspect 74, wherein the distance is calculated after dimensionality reduction of the chemical features.

[0186] (78) A computer program product including a non-transitory computer-readable storage medium encoded with instructions for implementing the method of aspect 74.

[0187] (79) A system for identifying potential chemicals for optimizing a chemical reaction, comprising a computer system configured to implement the method of aspect 74.

[0188] (80) A method for suggesting potential chemicals for optimizing a chemical reaction, comprising: obtaining a list of N potential chemicals; receiving results for M potential chemicals and thereby defining test chemicals; obtaining chemical signatures configured to describe properties of the N potential chemicals; determining a score for each representative chemical; and ranking the N potential chemicals based on the scores obtained for the representative chemicals.

[0189] (81) The method of aspect 80, wherein the act of determining a score for each representative chemical includes: calculating a distance between each test chemical and each representative chemical using a distance metric for each of a plurality of dimensionally reduced spaces of the chemical feature; normalizing the distance to a [0, 1] interval for each of the plurality of dimensionally reduced spaces; subtracting the normalized distance from 1 to generate a distance metric for each representative chemical and each test chemical for each of the plurality of dimensionally reduced spaces; averaging the distance metric for each representative chemical and each test chemical across all dimensionally reduced spaces; normalizing the results obtained from the received results of M potential chemicals to [0, 1] and converting them to weights; multiplying the weights by the distance metric to obtain a weighted distance metric, the distance metric having a representative chemical axis and a test chemical axis; and taking the maximum value along the representative chemical axis to obtain a score for each representative chemical.

[0190] (82) The method of aspect 81, wherein the test chemical is removed from the representative chemical axis.

[0191] (83) A computer program product including a non-transitory computer-readable recording medium encoded with instructions for performing the steps of the method of aspect 80.

[0192] (84) A system for suggesting potential chemicals that can be used to optimize a chemical reaction, comprising a computer system configured to implement the method of aspect 80.

Claims

1. 1. A method for optimizing a chemical reaction using machine learning, comprising: selecting chemical spaces for grouping, each chemical space being defined by a plurality of potential chemical substances; selecting a plurality of chemical features, each chemical feature corresponding to a plurality of potential chemical substances; grouping a plurality of potential chemical substances in a grouping space based on a plurality of chemical features; selecting a plurality of representative chemicals from the potential chemicals, each representative chemical corresponding to a group of the plurality of potential chemicals grouped in the grouping space; and Assembling a test kit having a plurality of test chemicals, each test chemical corresponding to a representative chemical of a plurality of representative chemicals. The method comprising:

2. The method of claim 1 , wherein the grouping space is defined by a plurality of chemical features.

3. 10. The method of claim 1, wherein the grouping space is a dimensionally reduced space of multiple chemical features.

4. 10. The method of claim 1, further comprising generating a plurality of reduced chemical spaces, wherein each reduced chemical space is a dimensionally reduced version of the chemical space.

5. 5. The method of claim 4, further comprising calculating a plurality of distances, wherein each distance of the plurality of distances is a distance between a prospective chemical of the plurality of prospective chemicals and a test chemical of the plurality of test chemicals.

6. 6. The method of claim 5, wherein the plurality of test chemicals is a subset of the plurality of prospective chemicals.

7. 6. The method of claim 5, wherein the plurality of prospective chemicals does not include a plurality of test chemicals.

8. 5. The method of claim 4, further comprising calculating a plurality of distance metrics, wherein each distance metric of the plurality of distance metrics is one minus the distance between one of the plurality of prospective chemicals and one of the plurality of test chemicals.

9. 10. The method of claim 8, further comprising averaging each of the plurality of distance metrics across the reduced chemical space to generate a plurality of averaged distance metrics.

10. 10. The method of claim 1, further comprising generating a plurality of weights, wherein each weight of the plurality of weights corresponds to one of a plurality of outcomes, and each outcome of the plurality of outcomes corresponds to one of a plurality of test chemicals.

11. The method of claim 10 , wherein the plurality of weights are determined according to an exponential function.

12. The method of claim 10 , wherein each of the plurality of distances of the distance metric is multiplied by a respective weight of the plurality of weights.

13. 13. The method of claim 12, wherein the maximum value for each potential chemical is taken across multiple multiplied distance metrics between each potential chemical and multiple test chemicals.

14. 14. The method of claim 13, further comprising ranking the maximum value of each potential chemical.

15. 10. The method of claim 1, further comprising testing a plurality of test chemicals to determine a plurality of outcomes, wherein each of the plurality of outcomes corresponds to an outcome for a respective test chemical of the plurality of test chemicals.

16. The method of claim 15, wherein the plurality of results are uploaded to a server.

17. 16. The method of claim 15, further comprising selecting a prediction space utilizing a plurality of prediction features to provide a prediction in the outcome space, wherein the plurality of outcomes define data points in the prediction space that are mapped to the outcome space.

18. The method of claim 17 , wherein the result space comprises chemical yield.

19. The method of claim 17 , wherein the result space is a by-product metric.

20. 18. The method of claim 17, wherein the result space is a single parameter.

21. The method of claim 17 , wherein the prediction space comprises a grouping space.

22. 18. The method of claim 17, wherein the prediction space includes multiple chemical features therein.

23. 18. The method of claim 17, wherein the computer is configured to fit the plurality of outcomes using regression fitting to map the plurality of predictive features to an outcome space.