Closed-Loop Optimization of General Reaction Conditions for Heteroaryl Suzuki-Miyaura Coupling
A closed-loop workflow using data-guided matrix refinement, uncertainty minimization, and robotic experimentation effectively identifies generic reaction conditions for heteroaryl Suzuki-Miyaura cross-coupling, significantly improving yield efficiency.
Patent Information
- Application Number
- JP2025524669
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-26
- Filing Date
- 2023-10-25
- Publication Date
- 2025-12-16
AI Technical Summary
Existing methods for discovering general reaction conditions in organic chemistry typically consider a small region of chemical space and are impractical for exhaustive experimentation due to the large matrix of substrates and reaction conditions, limiting the efficiency and generality of reaction optimization.
A closed-loop workflow leveraging data-guided matrix refinement, uncertainty minimization through machine learning, and robotic experimentation to identify generic reaction conditions for heteroaryl Suzuki-Miyaura cross-coupling, using a robotic system to automate reactions and optimize conditions based on yield predictions.
The method doubles the average yield of heteroaryl Suzuki-Miyaura cross-coupling reactions compared to traditional benchmarks, demonstrating efficient and reproducible discovery of general reaction conditions.
Smart Images

Figure 2025540581000001_ABST
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 419,702, filed October 26, 2022. GOVERNMENT SUPPORT
[0002] This invention was made with government support under HR00111920027 awarded by the Defense Advanced Research Projects Agency (DARPA). The government has certain rights in this invention. [Background technology]
[0003] General conditions for organic reactions are important but rare, and attempts to discover them typically consider only a small region of chemical space. Finding more general reaction conditions requires considering a large region of chemical space derived from a large matrix of substrates intersecting with a high-dimensional matrix of reaction conditions, making exhaustive experimentation impractical.
[0004] Embodiments of the present disclosure discover general reaction conditions using a simple closed-loop workflow that leverages data-guided matrix refinement, uncertainty minimization, machine learning, and robotic experimentation. Application to a challenging and important problem, heteroaryl Suzuki-Miyaura cross-coupling, identified conditions that doubled average yields compared to widely used benchmarks previously developed using traditional approaches. The present disclosure provides a practical roadmap for solving multidimensional chemical optimization problems with large search spaces. Summary of the Invention [Means for solving the problem]
[0005] In certain embodiments, selecting a reaction pair comprising a first molecule and a second molecule, the first molecule being selected from a first matrix and the second molecule being selected from a second matrix; selecting one or more conditions for reactions of the selected reaction pair, the selection being based on past use of the conditions and the structural and functional diversity of the selected reaction pair; automatically carrying out an initial round of reactions between selected reaction pairs under one or more selected conditions by a robotic system; optimizing one or more reaction conditions associated with an initial round of reactions based on the yield of each reaction in the initial round of reactions using a machine learning model; determining an optimized sequence of reactions to be performed; performing a reaction in the optimized sequence of reactions, thereby forming a product and predicting its yield; and outputting an optimized set of generic reaction conditions for each optimized reaction between the selected reaction pairs.
[0006] In certain embodiments, A robot system, A system is provided that includes a computing node including a computer-readable storage medium having program instructions embodied therein, the program instructions being executable by a processor of the computing node to cause the processor to perform methods, including those described herein. [Brief explanation of the drawings]
[0007] [Figure 1] 1 is a flow chart of an exemplary process for generating general reaction conditions according to the technology presented herein. [Figure 2] 1 shows a general workflow for discovering generic reaction conditions. [Figure 3A] 1 shows T-distributed stochastic neighbor embedding (t-SNE) mapping of substrate combinations. [Figure 3B] t-SNE mapping of the product space synthesized in the training and test sets compared to the overall reaction space is shown. [Figure 3C] The reaction schemes and chemical structures of the initial training set are shown. [Figure 3D] A robotic system for automatically performing all reactions is shown. [Figure 3E] 1 is a table showing all couplings performed between pairs of substrates under conditions corresponding to each JACS2009 benchmark. [Figure 3F] 1 is a graph showing a Spearman rank correlation matrix. [Figure 4A] Shows convergence of model uncertainty. [Figure 4B] A comparison of ML guided search and random search for general conditions is presented. [Figure 4C] 1 shows a comparison of yield distributions between reactions reported in the literature and reactions carried out in this disclosure. [Figure 4D] We show that the model gains the ability to accurately classify these conditions into high, medium, and low overall average yields, and in subsequent rounds establishes correct rankings within these categories. [Figure 4E] Shows the rankings for each general condition per round as recognized by the ML model. [Figure 4F] This shows that the uncertainty in the model rankings decreases between rounds, especially in the top condition. [Figure 4G] The first three rounds test several substrates per round across many conditions, followed in later rounds by selection of models primarily for filling in the top conditions. [Figure 4H] shows that by the fifth round, the model had explored almost all of the top seven conditions, corresponding to all conditions for which the overall average yield estimated by the model was >50%. [Figure 4I] The yields of the reactions that the model requires and that were analyzed to gain more information about the reaction-condition space are shown. [Figure 5A]A diverse set of 20 compounds from outside the training set was selected to test whether the discovered general reaction conditions translate to other diverse heteroaryl product classes. JACS 2009: 5:1 dioxane:water, 60°C, K3PO4, Pd SPhos G4. ML General Conditions 1: 5:1 dioxane:water, 100°C, Na2CO3, Pd XPhos G4. ML General Conditions 2: 5:1 dioxane:water, 100°C, Na2CO3, Pd SPhos G4. ML General Conditions 3: 5:1 dioxane:water, 100°C, Na2CO3, Pd(PPh3)4. [Figure 5B] Jitter plot of the performance of the top ML conditions against the benchmark. Brackets indicate 95% confidence intervals. [Figure 5C] Jitter plot of the relative performance of the top ML conditions against the benchmark in terms of yield variation. Brackets indicate 95% confidence intervals. [Figure 5D] The number of products per typical condition measured with a yield >10% is shown. [Figure 5E] Relative protodeboronation per condition, measured by integrated UV peak area (UVPDB), normalized to an internal standard (UVSTD) is shown. [Figure 5F] The relative remaining halide for each condition is shown as measured by integrated UV peak area (UVHAL) normalized to an internal standard (UVSTD). [Figure 5G] Relative product formation per condition (UVPDT) compared with by-product formation (UVBYPDT). [Figure 6] 1 illustrates an exemplary computing node. [Figure 7] A schematic diagram of an automated synthesizer is shown. [Figure 8] A photograph of the automatic synthesizer is shown. [Figure 9] The valve tube connection diagram is shown. 1-A to 5-D indicate the reaction vial connections 1 to 36 in order. S1 to S20 indicate solvent reservoirs. [Figure 10]A schematic diagram of the argon / vacuum manifold is shown. Sole = solenoid valve. Opening and closing the solenoid valves connected to the argon and vacuum pumps allows automation of the Schlenk line vacuum / fill process. [Figure 11] A schematic diagram of a four-channel solenoid driver is shown. [Figure 12] Top: Circular heating block design shown (units are inches). Bottom: In situ measured reaction temperatures per vial between two 12-reaction vial heating block configurations. Each vial was filled with 8 mL of dimethyl sulfoxide (DMSO), and the hot plate was turned on, set to 85°C, and allowed to equilibrate (approximately 30 minutes). Temperatures were measured by immersing a thermometer in the heated DMSO solution in each vial until a constant temperature was recorded (approximately 1 minute). "X" indicates the placement of the heating probe. [Figure 13] Shows the "front panel" of the LabVIEW code. The code is started by clicking the arrow in the top left. [Figure 14] An automated test reaction probing boron speciation is shown. The experiment was performed in eight parallel replicates. Data are presented as box plots showing the maximum and minimum observed values. Solid components were weighed into reaction vials in an argon-filled glovebox (due to the air sensitivity of the Pd catalyst) and sealed under argon before loading into the automated synthesizer. Reactions were performed with a 0.1 mmol scale of halide, 3 equivalents of boron species, 1 equivalent of phenanthrene internal standard, 7.5 equivalents of Na2CO3, and 5 mol% Pd(PPh3)4 catalyst in 8 mL of argon-sparged 5:1 dioxane:water solvent. [Figure 15] A manual test reaction to probe the inert environment is shown. The reaction was performed with a 0.1 mmol scale of halide, 3 equivalents of boron species, 1 equivalent of phenanthrene internal standard, 7.5 equivalents of KPO, and 5 mol% Pd SPhos G4 catalyst in 8 mL of 5:1 dioxane:water solvent. [Figure 16]The performance of the automated Schlenk line is demonstrated. The automated Schlenk cycle consisted of 5 minutes of vacuum followed by 1 minute of argon. Reactions were performed with a 0.1 mmol scale of halide, 3 equivalents of boron species, 1 equivalent of phenanthrene internal standard, 7.5 equivalents of K3PO4, and 5 mol% Pd SPhos G4 catalyst in 8 mL of 5:1 dioxane:water solvent. Reactions were performed in parallel by alternately introducing reaction vials into the instrument and running the Schlenk cycle. [Figure 17] 1 shows the schematic architecture of the neural components used in the GP model. [Figure 18A] The dependence of the prediction error and uncertainty of the models (NNE and GPE) on the relative size of the datasets used for training is shown. For each training set size, the training data were randomly selected from the entire Santanilla (16) dataset. The remainder of the dataset was used as a test set. The dataset selection and model training were repeated 60 times, with mean values indicated by dots and lines, and standard deviations from these means indicated by shaded areas. Orange is assigned to GPE and blue to NNE. A) Mean Absolute Error (MAE), solid line = test set, dotted line = training set. [Figure 18B] Kendall's τ, which measures the rank correlation between the absolute error and uncertainty of the model, is shown. [Figure 18C] Model uncertainty is shown (solid line = test set, dotted line = training set). [Figure 18D] It shows the z-score, i.e., the absolute error divided by the model uncertainty. [Figure 19] Show the next reaction selection. [Figure 20] Comparison of GP(NN), NNE, and GPE(NN) using different acquisition functions and substrate selection strategies. The vertical axis shows the actual / experimental yield under conditions predicted as optimal by a model. [Figure 21]Comparison of GP(NN), NNE, and GPE(NN) using different acquisition functions and substrate selection strategies is shown. The rank indicates where the actual best condition is classified by a method (see text for details). [Figure 22] The probability that a GPE(NN) model of a certain size will make the same selection as a reference GPE(NN) with 2000 models is shown. For each acquisition function (EI and PI), two probabilities are considered: the probability of selecting the same condition as the reference model (EI_cond and PI_cond), and the probability of simultaneously selecting the same condition and substrate (PI_maxUnc and EI_maxUnc). [Figure 23] Comparison of GPE(NN) with different numbers of models and different substrate selection schemes is shown. The rank indicates where a method classifies the best real / experimental conditions (see text for details). [Figure 24] Comparison of GPE(NN) for different models and different substrate selection schemes. The vertical axis shows the actual / experimental yield under conditions predicted as optimal by a given model. [Figure 25A] Information gap
number
[0008] In certain embodiments, selecting a reaction pair comprising a first molecule and a second molecule, the first molecule being selected from a first matrix and the second molecule being selected from a second matrix; selecting one or more reaction conditions for the reaction pair, the selection being based on past use of the one or more reaction conditions and the structural and functional diversity of the selected reaction pair; automatically carrying out an initial round of reaction between selected reaction pairs under one or more selected reaction conditions by a robotic system; optimizing one or more reaction conditions associated with an initial round of reactions based on the yield of each reaction in the initial round of reactions using a machine learning model; determining an optimized sequence of reactions to be performed; performing a reaction in the optimized sequence of reactions, thereby forming a product and predicting its yield; and outputting an optimized set of generic reaction conditions for each optimized reaction between the selected reaction pairs.
[0009] In certain embodiments, the method further includes minimizing uncertainty in the machine learning model.
[0010] In a further embodiment, minimizing the uncertainty comprises: constructing a surrogate model for predicting one or more reaction yields from a reaction pair; estimating an objective function for the performed reaction based on the predicted output from the surrogate model; and estimating the objective function for the unexecuted reactions.
[0011] In certain embodiments, selecting a reaction pair comprises: clustering the first matrix by common ring structures and pendant functional groups, the clustering producing a first centroid, the first centroid including the closest representative of the first matrix; selecting a second molecule from the second matrix; identifying all combinations comprising the first centroid and the second molecule, thereby generating a chemical space; comparing the chemical space to a corpus of chemical products, thereby generating a product space.
[0012] In a further embodiment, the method comprises: applying a greedy algorithm to the product space; and identifying a set of pairs of the first centroid and the second molecule to maximize the mutual dissimilarity of the products of the resulting reaction pairs.
[0013] In certain embodiments, selecting one or more reaction conditions comprises: considering at least one of a solvent, a base, a catalyst, and a temperature as a variable for one or more reaction conditions; determining a set of initial conditions based on an analysis of previous use of one or more reaction conditions.
[0014] In a further embodiment, the solvent is 5:1 dioxane:water.
[0015] In still further embodiments, the base is K3PO4 or Na2CO3.
[0016] In still further embodiments, the catalyst is a palladium catalyst, optionally further comprising a ligand.
[0017] In certain embodiments, the ligand is selected from the group consisting of SPhos, XPhos, and triphenylphosphine (PPh3).
[0018] In a further embodiment, the catalyst is selected from the group consisting of Pd(SPhos)G4, Pd(PPh3)4, and Pd(XPhos)G4.
[0019] In still further embodiments, the temperature is from about 50°C to about 150°C.
[0020] In still further embodiments, the temperature is about 60°C or about 100°C.
[0021] In certain embodiments, optimizing one or more reaction conditions includes: carrying out a reaction under one or more reaction conditions, thereby forming a product; and determining one or more reaction condition parameters that output the highest yield.
[0022] In a further embodiment, optimizing one or more reaction conditions comprises: Identifying one or more catalysts that have similar yields for different substrates; and eliminating one or more identified catalysts from the possible reaction conditions, thereby reducing redundancy.
[0023] In certain embodiments, the method comprises: iteratively optimizing one or more reaction conditions to thereby form a product; Determining the yield of the product; determining that a threshold yield is met.
[0024] In certain embodiments, the method comprises: The method further includes generating a small dataset for optimization by the machine learning model, the small dataset including negative data.
[0025] In certain embodiments, the first molecule comprises a halo-substituted aryl or heteroaryl.
[0026] In a further embodiment, the first molecule has the formula (Ia): [ka] wherein: A is aryl or heteroaryl; each R1 is independently selected from the group consisting of alkyl, alkoxyl, alkenyl, alkynyl, aralkyl, heteroaryl(alkyl), aryl, heteroaryl, halo, haloalkyl, hydroxyl, carboxyl, acyl, ester, amino, amido, cyano, cycloalkyl, and heterocycloalkyl; n1 is 0, 1, 2, 3, 4, or 5; X1 is a halo.
[0027] In still further embodiments, X1 is bromo.
[0028] In certain embodiments, the first molecule is [ka] [ka] [ka] is selected from the group consisting of:
[0029] In certain embodiments, the second molecule comprises an aryl or heteroaryl that further comprises a boronic acid, a boronic ester, or a tetrafluoroborate.
[0030] In a further embodiment, the second molecule has the formula (Ib): [ka] wherein: B is aryl or heteroaryl; each R2 is independently selected from the group consisting of alkyl, alkoxyl, alkenyl, alkynyl, aralkyl, heteroaryl(alkyl), aryl, heteroaryl, halo, haloalkyl, hydroxyl, carboxyl, acyl, ester, amino, amido, cyano, cycloalkyl, and heterocycloalkyl; n2 is 0, 1, 2, 3, 4, or 5; X2 is selected from the group consisting of N-methylimidodiacetic acid boronic acid ester, tetramethyl N-methyliminodiacetic acid boronic acid ester, pinacol boronic acid ester, boronic acid, or tetrafluoroborate.
[0031] In still further embodiments, the second molecule is [ka] [ka] [ka] wherein BMIDA is N-methylimidodiacetic acid boronic acid ester.
[0032] In certain embodiments, A robot system, a computing node including a computer-readable storage medium having program instructions embodied therein, the program instructions being executable by a processor of the computing node to cause the processor to perform a method described herein; A system is provided that includes:
[0033] The development of automated synthetic methods for peptides, nucleic acids, and polysaccharides requires the discovery of highly general reaction conditions applicable to a wide range of building block combinations. In contrast, in the synthesis of small organic molecules, custom-tailored reaction conditions are typically developed to maximize the yield of each target molecule, minimize by-products, and / or minimize the cost of the corresponding process. This is often necessary because synthetic methods are typically optimized for only one or a few pairs of substrates and then applied to a wider range of substrate combinations, with the largely unfulfilled hope that the same conditions will generally produce higher yields. Even the application of machine learning to optimization protocols does not guarantee the generality necessary to automate, accelerate, and ultimately democratize the small molecule production process. Identifying such general conditions is challenging because the search space (encompassing all possible combinations of reaction conditions multiplied by all possible combinations of substrates) is enormous and impractical to navigate using standard approaches.
[0034] Heteroaryl molecular fragments are ubiquitous in many industrially important functional molecules, including pharmaceuticals, materials, catalysts, dyes, and natural products. Synthesis remains a significant bottleneck in all of these areas. Therefore, finding general conditions for (hetero)aryl Suzuki-Miyaura cross-coupling (SMC) is an important problem. This is also a challenging and largely unsolved problem, primarily due to the very large and broad range of potential heteroaryl and aryl substrates, with varying degrees of both desired and undesired reactivity. Embodiments of the present disclosure use machine learning (ML) to discover general reaction conditions by leveraging the extensive chemical literature on (hetero)aryl SMC.
[0035] Here, we report a simple, closed-loop workflow that can efficiently navigate vast substrate condition spaces to discover generic reaction conditions. This approach leverages: i) data-guided matrix refinement to maintain global validity while making the vast search space tractable; ii) uncertainty-minimization ML to efficiently drive predictive optimization; and iii) robotic experimentation to increase the throughput, accuracy, and reproducibility of on-demand, recursively generated datasets (Figure 2). This workflow successfully identifies generic reaction conditions for (hetero)aryl SMC reactions. This solution doubles the average yield compared to benchmark generic conditions previously developed by traditional human-guided experimentation and has since been widely used in academic and industrial laboratories worldwide (cited in >590 papers and patent applications). Thus, this approach can uncover impactful solutions in vast, multidimensional search spaces, accelerating the progress toward automation and democratization of small molecule synthesis, which is critically in need of more generic reaction conditions in organic chemistry.
[0036] 1 is a flowchart 100 of an exemplary process for generating generic reaction conditions. For example, process 100 (i.e., steps 102-112) can be automated or initiated in response to user input. In step 102, process 100 can include selecting a set of molecules from a first matrix. The molecules can be strategically narrowed down to define a chemical space.
[0037] In step 104, process 100 may include selecting one or more conditions for reaction of the set of selected molecules, where the selection is based on past use of the conditions and the structural and functional diversity of the set of selected molecules. In step 106, process 100 may include automatically performing an initial round of reaction between the set of selected molecules under the one or more selected conditions by a robotic system. The robotic system may be an automated synthesizer as described herein.
[0038] In step 108, process 100 may include using a machine learning model to optimize one or more reaction conditions associated with an initial round of reactions based on the yield of each reaction in the initial round of reactions. Optimizing the conditions may include minimizing uncertainty in the machine learning model. In step 110, process 100 may include determining an optimized series of reactions to be performed and predicting their yields. In some embodiments, the process may include minimizing uncertainty in the machine learning model.
[0039] In step 112, process 100 may include outputting an optimized set of generic reaction conditions for each optimized reaction between the set of selected molecules.
[0040] Figure 2 shows an embodiment of a simple closed-loop workflow that can efficiently navigate a vast substrate condition space to discover generic reaction conditions. In step 202, data-guided matrix refinement is performed to create a chemical space for the generation of generic reaction conditions. In step 204, a closed loop is created between the uncertainty minimization machine learning model and the robotic experiments. In step 206, generic reaction conditions are determined and output from the iterative closed-loop process.
[0041] Data-Guided Substrate Refinement. To enable practical exploration of common hetero(aryl)SMC reaction conditions, both the matrix of potential building block combinations and the matrix of potential reaction conditions are strategically refined in a manner that maintains the relevance of these subsets to the whole (as shown in step 202 in Figure 2). Specifically, we data-mined the inventories of common fine chemical suppliers to compile a list of approximately 5,400 (hetero)aryl halide building blocks that are actually commercially available and therefore accessible. To define a representative subset of this chemical space, we applied a stratified clustering strategy (Figure 26) to algorithmically cluster the building blocks by common (hetero)aromatic ring substructures and pendant functional groups, and refined 54 "centroid" molecules that best represent each section of the available chemical space. These molecules were combined with 54 selected commercially available (hetero)aryl MIDA boronates to define a refined substrate range consisting of 2,688 representative cross-coupling products (Figures S22 and S23). Mapping this potential product space and comparing it to all heteroaryl products previously reported in the literature reveals substantial overlap between both sets, suggesting that it is representative of the entire heteroaryl chemical space (Figure 3A). Figure 3A shows a T-distributed stochastic neighbor embedding (t-SNE) mapping of the investigated substrate combinations (here, 2688 heteroaryl products) compared to all previously reported (hetero)aryl products. Blue circles represent products reported in the literature, yellow stars represent products that belong exclusively to the reported search space, and green triangles represent products present in both sets.
[0042] However, testing even this initially narrowed set of cross-coupling products against many possible reaction conditions is technically infeasible. Therefore, a second layer of narrowing was pursued. Specifically, a Tanimoto similarity-based greedy algorithm was used to identify a set of 11 representative substrate pairs from this larger set that maximized the mutual dissimilarity of the resulting products (Figure 3B). LCMS-UV / Vis response factor curves were determined for all these products, allowing for the automated determination of the yields of the automatically performed reactions. Figure 3B shows t-SNE mapping of the product space synthesized during the training and test sets compared to the overall reaction space. Blue circles indicate products belonging to the reported search space, yellow stars indicate products belonging to the test set, and green triangles indicate products belonging to the training set.
[0043] Data-guided refinement of conditions. Four variables were considered for the conditions: solvent, base, catalyst / ligand, and temperature. Because the goal was to test a wide range of conditions, a representative class of conditions was initially refined based on structural and functional diversity as well as the degree of previous use derived from our previous comprehensive literature analysis. For example, the two most commonly used solvents in the literature are dioxane and dimethoxyethane, but they both belong to the same solvent class of ethers, and therefore only one, dioxane, was selected. Similar reasoning led to the retention of only one carbonate base. The temperature of 100°C was the most frequently used temperature in the literature, and 60°C was the temperature used in a previously developed benchmark protocol. Finally, three solvents (dioxane, toluene, and dimethylformamide, all used in a 5:1 mixture with water), two bases (sodium carbonate and potassium phosphate), two temperatures (60 °C and 100 °C), and seven catalysts (Pd SPhos G4, Pd(PPh3)4, Pd XPhos G4, Pd P(tBu)3 G4, Pd PCy3 G4, Pd2(dba)3, and Pd(dppf)Cl2; G4 refers to the fourth-generation Buchwald precatalyst for evaluation). The 11 building block combinations selected above were tested under an initial set of conditions to "seed" the ML optimization (Figure 3C, showing the reaction scheme and chemical structures of the initial training set), and then iteratively tested under a broader set of conditions during the ML-guided optimization phase.
[0044] Seeding Experiments, Reaction Standardization, and Condition Space. All reactions were performed automatically with the robotic system shown in Figure 3D. Before solvent addition, heating, and stirring, the reaction mixture was purged with 10 automated vacuum / argon cycles, which resulted in highly reproducible reaction yields (Figure 16). This automated Schlenk process was necessary for reproducibility, even when using air-stable precatalysts and building blocks. To "seed" the optimization procedure, all couplings were performed between the 11 pairs of substrates mentioned above, each under seven different conditions: those corresponding to the JACS 2009 benchmark (5:1 dioxane:water, 60 °C, K3PO4, Pd SPhos G4), the same base and solvent but with other selected palladium catalysts (Pd XPhos G4, Pd P(tBu)3G4, Pd The results were obtained using PCyG4, Pd2(dba)3, and Pd(dppf)Cl2. The conditions used were the most common catalyst (Pd(PPh3)4), base (Na2CO3), temperature (100 °C), and solvent (dioxane:water) used in the literature (Figure 3E). Each reaction was repeated twice, and the yields showed only a ±2% deviation, highlighting one of the key advantages of automated experiments (even when the same skilled expert repeats the same reaction, a variability of approximately 10–15% has been reported).
[0045] This initial round of experiments also allowed for the identification of catalysts that systematically exhibited similar yields for different substrate pairs and thus may be redundant. Such functional similarity, rather than structural similarity, is quantified by the Spearman rank correlation matrix shown in Figure 3F, which gives correlated yields for all 11 substrate pairs using two different catalytic ligands. In this representation, redundant catalysts correspond to highly correlated off-diagonal elements (e.g., XPhos and dppf, PCy3 and SPhos). Based on this analysis, PCy3 and dppf were removed from the ligand pool to reduce redundancy, and Pd2(dba)3 was removed due to its poor performance (<5% yield for 8 / 11 substrates), resulting in a total space to be explored of 528 reactions (11 substrates × 2 temperatures × 2 bases × 3 solvents × 4 catalysts).
[0046] Uncertainty-Minimizing ML for Generality. Reaction conditions can be considered maximally general if they provide the highest average yield across the widest range of chemical space. Optimizing generality is an open and underexplored challenge in the evolving field of ML. An alternative approach was explored, where small sets of highly reproducible data were generated on-demand during ML-guided closed-loop optimization, including negative data that are largely underrepresented in existing datasets. ML algorithms were also strategically focused on reducing model uncertainty, thereby maximizing the efficiency of the learning process.
[0047] Let C = {c} be the set of possible reaction conditions, S = {s} be the set of substrate pairs, and y(s,c) be the reaction yield. The objective is to maximize the objective function given by:
number
[0048] The general condition cgeneral is given as follows:
number
[0049] At first glance, with a minimal number of experiments, general The problem of identifying f(c) is similar to standard Bayesian optimization (BO). However, there is a substantial difference between all BO algorithms, in that each experiment / measurement performed immediately provides information about the objective function desired for optimization. In contrast, experimental evaluation of f(c) for a given problem requires multiple experiments (because the summation in Eq. 1 is over the entire set S), i.e., determining f(c) for given conditions requires experiments with all pairs of substrates in the set S. To address this issue, the standard BO approach is extended to include the reaction yield
number
number
[0050] To estimate uncertainty, we used a model that provides prediction uncertainty commensurate with the prediction error. For example, highly reliable predictions with large errors are undesirable. Based on the analysis of numerous neural network (NN) and Gaussian process (GP) models (SM Section 3), we selected the ensemble GPE(NN) of GPs supplemented with an NN kernel component. Such models are particularly attractive due to their flexibility (similarity metrics between different conditions are learned from the data) and reliability of prediction uncertainty (e.g., predictions of test samples are guaranteed not to be more confident than those of training samples). Regarding the selection of conditions to be tested, we selected the probability of improvement (PI) as the gain function.
[0051] Closed-loop, ML-guided optimization with robotic experimentation. A GPE(NN) / PI model guided automated experiments across a narrowly selected search space. Before updating the theoretical model, multiple experiments were performed to create "work batches." Within each batch, the algorithm formed a "priority queue" of unexplored reactions by sorting and selecting conditions according to the calculated PI and substrate pairs according to their predicted uncertainty. The batch sizes for optimization rounds 1-2 were 36 duplicate reactions, followed by rounds 3-4 and 5, which contained 72 and 84 reactions, respectively, with no duplicates.
[0052] Throughout the closed-loop rounds, the model uncertainty decreased and converged to the threshold obtained during the calibration simulations by the fifth round (Figures 4A, 25A, and 25B), suggesting that the model had gained sufficient knowledge of the entire space and that the optimization was terminated at this point. Figure 4A shows the convergence of the model uncertainty. The dashed horizontal line indicates the threshold obtained during the calibration simulations.
number
[0053] When the algorithm explores the reaction-condition space, the reaction yields for our dataset were more or less uniformly distributed across the range of expected values (Figure 4C, comparison of yield distributions between reactions reported in the literature and those performed in this disclosure). In other words, the protocol was learned by probing both low- and high-yield conditions. This should contrast with the distribution of yields in published reaction sets, where such yields are heavily biased toward positive outcomes, limiting the usefulness of approaches aimed at learning from public datasets.
[0054] In this dataset, the top 1 condition discovered yielded an average yield of 72% across all 11 substrates, while the benchmark condition (also found to be the top 5 [fifth-best] condition) yielded an average yield of 64%. To understand how the model reached this optimal condition, we examined the model's recognition of the average yield and ranking of each general condition per round, as shown in Figure 4D-E. As shown in Figure 4D, within the first two rounds, the model acquired the ability to accurately classify these conditions into high, medium, and low overall average yields, and in subsequent rounds, it established correct rankings within these categories. Figure 4E shows the ranking of each general condition per round as recognized by the ML model. The improvement in model accuracy over the course of the experiment is reproduced in Figure 4F. This improvement in accuracy indicates a decrease in the model's ranking uncertainty between rounds, particularly evident for the top conditions. The model chose to test a few substrates per round across many conditions in the first three rounds, followed by filling in primarily the top conditions in later rounds (Figure 4G). By the fifth round, the model had explored nearly all of the top seven conditions, corresponding to all conditions for which the overall average yield estimated by the model was >50% ( Figure 4H ).
[0055] Finally, the yields of the reactions predicted by the model were analyzed to gain more information about the reaction-condition space (Figure 3I). Because the yield of a single experiment was not our goal, these values are not expected to increase as the optimization progressed. Considering the uncertainty-guided selection of substrates, the opposite might be expected: once a suitable set of conditions was identified, further development should include reactions with lower yields to verify that the candidate conditions found were indeed better and to increase the confidence in the estimate of f(c). The results shown in Figure 3I indicate that after searching for good reactions in the second iteration, the model gradually shifted its attention to parts of the reaction-condition space that could be considered "negative examples" (and, in doing so, improved its prediction accuracy). These results reveal that i) relatively good candidate solutions were identified early, ii) the model initially sought reactions with better yields (to find better candidates), and iii) as the "loop" progressed, more attention was paid to reducing the uncertainty of its estimates.
[0056] Quantifying Generality. After discovering the higher-yielding general conditions within the training set, we next sought to examine whether learning could be transferred to substrates outside the scope of optimization, specifically, over 20 substrate pairs selected (by the Butina (37) algorithm) to maximize dissimilarity to the training set while ensuring coverage of heterocyclic substructure and functional group space (Figure 3B). We then focused on synthesizing and purifying all of the computer suggestions and testing them against benchmark conditions and the top three highest-yielding general reaction conditions discovered during the closed-loop optimization (Figure 5A) (ranked by the model after completing round 5). Despite including several very challenging building block combinations, this process achieved a 95% success rate, with only one product failing to have a measurable yield across all four conditions.
[0057] The general reaction conditions discovered by ML were significantly better than previously reported and widely used benchmark conditions (14). The top two conditions resulted in a statistically significant increase in average yield compared to the benchmark, with the top condition doubling the overall average yield from 21% to 46% (Figure 4B). Comparing the relative increases in yield reveals statistically significant differences between the top one condition and both the top two and top three conditions (Figure 5C). Notably, experimental yield correlates with the predicted ranking of the conditions, such that the top one yield is higher than the top two, which in turn is higher than the top three. In the context of functional discovery efforts, the dual ability to isolate or not isolate testable quantities of purified targeted compounds is arguably more important than a specific yield percentage. We estimate that a 10% yield is a practical limit for isolating purified product. In the benchmark condition, only 11 / 20 targeted products cleared this bar, but this percentage increased to 19 / 20 in the top-1 condition ( Figure 4 D).
[0058] Note that extending the reaction time for low-yielding couplings under benchmark conditions did not increase yields (Figures 39A-39E). A comprehensive analysis of by-product and product formation for all 20 reactions indicated that a favorable shift from the former to the latter was accompanied by a shift from the benchmark to the ML-discovered reaction conditions. Specifically, the ML-discovered conditions were associated with a trend toward decreased protodeboronation (Figure 5E), increased halide conversion (Figure 5F), and an overall statistically significant increase in the ratio of product to total by-product formation (Figure 5G) (0.30 ± 0.12 in JACS vs. 0.58 ± 0.12 in ML conditions, p = 0.0005).
[0059] The straightforward workflow developed here allows for accelerated discovery of improved general reaction conditions for challenging C-C bond-forming reactions, representing an important step toward increasing the efficiency, generality, and accessibility of small molecule synthesis. The results also highlight the power of refinement selection as an entry point into large, multidimensional search spaces, the distinct advantages of de novo ML approaches for navigating such spaces by generating data sets that equally reflect the realities of positive and negative data during optimization, and the unique suitability of robotic chemistry for generating high-quality, reproducible data.
[0060] Computing Node. Referring now to Figure 6, a schematic diagram of an example computing node is shown. Computing node 10 is only one example of a suitable computing node and is not intended to suggest any limitation as to the scope of use or functionality of the embodiments described herein. In any event, computing node 10 may implement and / or perform any of the set of functions described above.
[0061] Computing node 10 includes computer system / server 12, which operates in conjunction with numerous other general-purpose or special-purpose computing system environments or configurations. Examples of well-known computing systems, environments, and / or configurations that may be suitable for use with computer system / server 12 include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments that include any of the above systems or devices.
[0062] The computer system / server 12 may be described in the general context of computer system-executable instructions, such as program modules, being executed by the computer system. Generally, program modules may include routines, programs, objects, components, logic, data structures, etc. that perform particular tasks or implement particular abstract data types. The computer system / server 12 may be practiced in a distributed cloud computing environment where tasks are performed by remote processing devices that are linked through a communications network. In a distributed cloud computing environment, program modules may be located in both local and remote computer system storage media, including memory storage devices.
[0063] 6, computer system / server 12 within computing node 10 is shown in the form of a general-purpose computing device. Components of computer system / server 12 may include, but are not limited to, one or more processors or processing units 16, a system memory 28, and a bus 18 that couples various system components including system memory 28 to processor 16.
[0064] Bus 18 may be one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of bus architectures, including, by way of example and not limitation, an Industry Standard Architecture (ISA) bus, a MicroChannel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe), and an Advanced Microcontroller Bus Architecture (AMBA).
[0065] Computer system / server 12 typically includes a variety of computer system-readable media, which may be any available media that can be accessed by computer system / server 12 and includes both volatile and nonvolatile media, removable and non-removable media.
[0066] The system memory 28 may include computer system-readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The computer system / server 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 34 may be provided for reading from and writing to non-removable, non-volatile magnetic media (not shown, commonly referred to as a "hard drive"). Although not shown, a magnetic disk drive may be provided for reading from and writing to removable, non-volatile magnetic disks (e.g., "floppy disks"), and an optical disk drive may be provided for reading from or writing to removable, non-volatile optical disks, such as CD-ROMs, DVD-ROMs, or other optical media. In such cases, one or more data media interfaces may be respectively connectable to the bus 18. As further illustrated and described below, the memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of embodiments of the present disclosure.
[0067] A program / utility 40 having a set of program modules 42 (at least one) may be stored in memory 28, as well as, by way of example and not limitation, an operating system, one or more application programs, other program modules, and program data. Each of the operating system, one or more application programs, other program modules, and program data, or some combination thereof, may include an implementation of a network environment. The program modules 42 generally perform the functions and / or methodologies of the embodiments described herein.
[0068] The computer system / server 12 may also communicate with one or more external devices 14, such as a keyboard, pointing device, and display 24; one or more devices that allow a user to interact with the computer system / server 12; and / or any device (e.g., a network card, modem, etc.) that allows the computer system / server 12 to communicate with one or more other computing devices. Such communication may occur via an input / output (I / O) interface 22. Additionally, the computer system / server 12 may communicate with one or more networks, such as a local area network (LAN), a general wide area network (WAN), and / or a public network (e.g., the Internet), via a network adapter 20. As shown, the network adapter 20 communicates with other components of the computer system / server 12 via a bus 18. Although not shown, it should be understood that other hardware and / or software components may be used in conjunction with the computer system / server 12. Examples include, but are not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems.
[0069] The present disclosure may be embodied as a system, method, and / or computer program product, which may include computer-readable storage medium(s) having computer-readable program instructions thereon for causing a processor to perform aspects of the present disclosure.
[0070] A computer-readable storage medium may be a tangible device that can hold and store instructions for use by an instruction execution device. The computer-readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific, non-exhaustive examples of computer-readable storage media include: portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD), memory sticks, floppy disks, mechanically encoded devices such as punch cards or ridge-in-groove structures with instructions recorded thereon, and any suitable combination of the above. As used herein, computer-readable storage media should not be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses passing through a fiber optic cable), or electrical signals transmitted over electrical wires.
[0071] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device or to an external computer or external storage device over a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface on each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage on a computer-readable storage medium within the respective computing / processing device.
[0072] The computer-readable program instructions for carrying out the operations of the present disclosure may be either assembler instructions, instruction set architecture (ISA) instructions, machine language instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, or the like, and conventional procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) can execute computer-readable program instructions by utilizing state information of the computer-readable program instructions to personalize the electronic circuitry to implement aspects of the present disclosure.
[0073] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0074] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, executed by the processor of the computer or other programmable data processing apparatus, create means for performing the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams. These computer-readable program instructions may also be stored on a computer-readable storage medium, which has instructions stored therein, can direct a computer, programmable data processing apparatus, and / or other device to function in a particular manner, including an article of manufacture containing instructions that implement aspects of the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.
[0075] Furthermore, computer-readable program instructions can be loaded into a computer, other programmable data processing apparatus, or other device to execute a series of operational steps on the computer, other programmable data processing apparatus, or other device to create a computer-implemented process, such that the instructions executing on the computer, other programmable data processing apparatus, or other device perform the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.
[0076] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, apparatuses, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of instructions, including one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may in fact be executed substantially simultaneously, or the blocks may sometimes be executed in reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, may be implemented by special-purpose hardware-based systems that perform the specified functions or acts, or a combination of special-purpose hardware and computer instructions.
[0077] The description of various embodiments of the present disclosure has been presented for purposes of illustration and is not intended to be exhaustive or limiting of the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terms used herein have been selected to best explain the principles of the embodiments, practical applications or technical improvements to technology found in the market, or to enable those skilled in the art to understand the embodiments disclosed herein.
[0078] definition As used herein, the term "greedy algorithm" refers to any algorithm that follows a problem-solving heuristic that makes locally optimal choices at each stage.
[0079] As used herein, the term "organic moiety" refers to a monovalent group containing one or more carbon atoms. The organic moiety may be aromatic or hydrocarbon-derived. The organic moiety may contain one or more heteroatoms, one or more units of unsaturation, and / or one or more functional groups. The organic moiety may be substituted.
[0080] As used herein, the term "alkyl" is a term of art and refers to saturated aliphatic groups, including straight-chain alkyl groups, branched-chain alkyl groups, cycloalkyl (alicyclic) groups, alkyl-substituted cycloalkyl groups, and cycloalkyl-substituted alkyl groups. In certain embodiments, a straight-chain or branched-chain alkyl has about 30 or fewer carbon atoms in its backbone (e.g., C1-C for straight chain). 30 , C3 to C for branched chains 30 ), or having about 20 or less, or 10 or less carbon atoms. In certain embodiments, the term "alkyl" refers to a C1-C 10 In certain embodiments, the term "alkyl" refers to a C1 to C6 alkyl group, for example, a C1 to C6 straight chain alkyl group. In certain embodiments, the term "alkyl" refers to a C3 to C6 alkyl group. 12 "Alkyl" refers to a branched chain alkyl group. In certain embodiments, the term "alkyl" refers to a C3 to C8 branched chain alkyl group. Representative examples of alkyl include, but are not limited to, methyl, ethyl, n-propyl, isopropyl, n-butyl, sec-butyl, isobutyl, tert-butyl, n-pentyl, isopentyl, neopentyl, and n-hexyl.
[0081] The term "cycloalkyl" refers to monocyclic, bicyclic or bridged saturated carbocyclic rings, each having from 3 to 12 carbon atoms. Some cycloalkyls have from 5 to 12 carbon atoms in their ring structure, and some have 6 to 10 carbons in the ring structure. Preferably, cycloalkyl is a (C3-C7)cycloalkyl, which refers to a monocyclic saturated carbocyclic ring having from 3 to 7 carbon atoms. Examples of monocyclic cycloalkyls include cyclopropyl, cyclobutyl, cyclopentyl, cyclopentenyl, cyclohexyl, cyclohexenyl, cycloheptyl, and cyclooctyl. Bicyclic cycloalkyl ring systems include bridged monocyclic rings and fused bicyclic rings. Bridged monocyclic rings are rings in which two non-adjacent carbon atoms of a monocyclic ring are joined by an alkylene bridge of 1 to 3 additional carbon atoms (i.e., of the form -(CH2) w - (where w is 1, 2, or 3) bridging group). Representative examples of bicyclic systems include, but are not limited to, bicyclo[3.1.1]heptane, bicyclo[2.2.1]heptane, bicyclo[2.2.2]octane, bicyclo[3.2.2]nonane, bicyclo[3.3.1]nonane, and bicyclo[4.2.1]nonane. Fused bicyclic cycloalkyl ring systems contain a monocyclic cycloalkyl ring fused to either a phenyl, a monocyclic cycloalkyl, a monocyclic cycloalkenyl, a monocyclic heterocyclyl, or a monocyclic heteroaryl. Bridged or fused bicyclic cycloalkyls are attached to the parent molecular moiety through any carbon atom contained within the monocyclic cycloalkyl ring. The cycloalkyl group is optionally substituted. In certain embodiments, the fused bicyclic cycloalkyl is a 5- or 6-membered monocyclic cycloalkyl ring fused to either a phenyl ring, a 5- or 6-membered monocyclic cycloalkyl, a 5- or 6-membered monocyclic cycloalkenyl, a 5- or 6-membered monocyclic heterocyclyl, or a 5- or 6-membered monocyclic heteroaryl, wherein the fused bicyclic cycloalkyl is optionally substituted.
[0082] As used herein, the term "(cycloalkyl)alkyl" refers to an alkyl group substituted with one or more cycloalkyl groups. An example of a (cycloalkyl)alkyl is a cyclohexylmethyl group.
[0083] The term "heterocycloalkyl," as used herein, refers to radicals of non-aromatic ring systems, including but not limited to monocyclic, bicyclic, and tricyclic rings, which may be fully saturated or which may contain one or more units of unsaturation, where, for the avoidance of doubt, the degree of unsaturation does not result in an aromatic ring system and has from 3 to 12 atoms including at least one heteroatom, e.g., nitrogen, oxygen, or sulfur. For illustrative purposes, and not to be construed as limiting the scope of the invention, examples of heterocycles are listed below: aziridinyl, azirinyl, oxiranyl, thiiranyl, thiirenyl, dioxiranyl, diazirinyl, diazepanyl, 1,3-dioxanyl, 1,3-dioxolanyl, 1,3-dithiolanyl, 1,3-dithianyl, imidazolidinyl, isothiazolinyl, isothiazolidinyl, isoxazolinyl, isoxazolidinyl, azetyl, oxetanyl, oxetyl, thietanyl, thiethyl, diazetidinyl, dioxetanyl, dioxetenyl, dithietanyl, dithiethyl, dioxalanyl, oxazolyl, thiazolyl, triazinium, Heterocycloalkyl groups include aryl, isothiazolyl, isoxazolyl, azepine, azetidinyl, morpholinyl, oxadiazolinyl, oxadiazolidinyl, oxazolinyl, oxazolidinyl, oxopiperidinyl, oxopyrrolidinyl, piperazinyl, piperidinyl, pyranyl, pyrazolinyl, pyrazolidinyl, pyrrolinyl, pyrrolidinyl, quinuclidinyl, thiomorpholinyl, tetrahydropyranyl, tetrahydrofuranyl, tetrahydrothienyl, thiadiazolinyl, thiadiazolidinyl, thiazolinyl, thiazolidinyl, thiomorpholinyl, 1,1-dioxidethiomorpholinyl (thiomorpholinesulfone), thiopyranyl, and trithianyl. Heterocycloalkyl groups are optionally substituted with one or more substituents, described below.
[0084] As used herein, the term "(heterocycloalkyl)alkyl" refers to an alkyl group substituted with one or more heterocycloalkyl (ie, heterocyclyl) groups.
[0085] As used herein, the term "alkenyl" refers to a straight- or branched-chain hydrocarbon radical containing 2 to 10 carbons and containing at least one carbon-carbon double bond formed by the removal of two hydrogens. Representative examples of alkenyl include, but are not limited to, ethenyl, 2-propenyl, 2-methyl-2-propenyl, 3-butenyl, 4-pentenyl, 5-hexenyl, 2-heptenyl, 2-methyl-1-heptenyl, and 3-decenyl. The unsaturated bond(s) in an alkenyl group can be located at any position in the moiety and can have either the (Z) or (E) configuration about the double bond(s).
[0086] The term "alkynyl," as used herein, refers to a straight- or branched-chain hydrocarbon radical containing 2 to 10 carbon atoms and containing at least one carbon-carbon triple bond. Representative examples of alkynyl include, but are not limited to, acetylenyl, 1-propynyl, 2-propynyl, 3-butynyl, 2-pentynyl, and 1-butynyl.
[0087] The term "alkylene" is art-recognized and, as used herein, refers to a diradical obtained by removing two hydrogen atoms from an alkyl group, as defined above. In one embodiment, alkylene refers to a disubstituted alkane, i.e., an alkane substituted at two positions with substituents such as halogen, azide, alkyl, aralkyl, alkenyl, alkynyl, cycloalkyl, hydroxyl, alkoxyl, amino, nitro, sulfhydryl, imino, amido, phosphonate, phosphinate, carbonyl, carboxyl, silyl, ether, alkylthio, sulfonyl, sulfonamide, ketone, aldehyde, ester, heterocyclyl, aromatic or heteroaromatic moiety, fluoroalkyl (such as trifluoromethyl), cyano, and the like. That is, in one embodiment, "substituted alkyl" is "alkylene."
[0088] The term "amino" is a term of art and as used herein refers to both unsubstituted and substituted amines, for example, those of the general formula: [ka] (In the formula, R a , R b , and R c are each independently hydrogen, alkyl, alkenyl, -(CH2) x -R d or R a and R b together with the N atom to which they are attached complete a heterocycle with 4 to 8 atoms in the ring structure, and R d represents aryl, cycloalkyl, cycloalkenyl, heterocyclyl, or polycyclyl, and x is zero or an integer ranging from 1 to 8. In certain embodiments, R a or R b Only one of R a , R b and nitrogen together do not form an imide. a and R b (and optionally Rc ) are each independently hydrogen, alkyl, alkenyl, or -(CH2) x -R d In certain embodiments, the term "amino" refers to -NH2.
[0089] In certain embodiments, the term "alkylamino" refers to -NH(alkyl).
[0090] In certain embodiments, the term "dialkylamino" refers to -N(alkyl).
[0091] The term "amide," as used herein, means -NHC(=O)-, where the amide group is attached to the parent molecular moiety through a nitrogen. Examples of amides include alkylamides, such as CHC(=O)N(H)- and CHCHC(=O)N(H)-.
[0092] The term "acyl" is a term of art and, as used herein, refers to any group or radical of the form RCO-, where R is any organic group, such as alkyl, aryl, heteroaryl, aralkyl, and heteroaralkyl. Representative acyl groups include acetyl, benzoyl, and malonyl.
[0093] As used herein, the term "aminoalkyl" refers to an alkyl group substituted with one or more amino groups. In one embodiment, the term "aminoalkyl" refers to an aminomethyl group.
[0094] The term "aminoacyl" is a term of art and as used herein refers to an acyl group substituted with one or more amino groups.
[0095] The term "aminothionyl" as used herein refers to the analog of aminoacyl in which the O of RC(O)-- is replaced by sulfur, thus the form RC(S)--.
[0096] The term "phosphoryl" is a term of art and as used herein generally refers to a group having the formula: [ka] where Q50 represents S or O, and R59 represents hydrogen, lower alkyl, or aryl, such as -P(O)(OMe)- or -P(O)(OH). For example, when used to substitute an alkyl, the phosphoryl group of the phosphorylalkyl has the general formula: [ka] (wherein Q50 and R59 are each independently defined above, and Q51 represents O, S, or N), for example, -OP(O)(OH)OMe or -NH-P(O)(OH)2. When Q50 is S, the phosphoryl moiety is a "phosphorothioate."
[0097] The term "aminophosphoryl," as used herein, refers to a phosphoryl group substituted with at least one amino group, as defined herein, for example, -P(O)(OH)NMe2.
[0098] The terms "azide" or "azido" as used herein refer to an -N3 group.
[0099] The term "carbonyl" as used herein refers to -C(=O)-.
[0100] The term "thiocarbonyl" as used herein refers to -C(=S)-.
[0101] The term "alkylphosphoryl," as used herein, refers to a phosphoryl group substituted with at least one alkyl group, as defined herein, e.g., -P(O)(OH)Me.
[0102] The term "alkylthio" as used herein refers to alkyl-S-. The term "(alkylthio)alkyl" refers to an alkyl group substituted with an alkylthio group.
[0103] The term "carboxy" as used herein means a -CO2H group.
[0104] The term "aryl" is a term of art and, as used herein, refers to monocyclic, bicyclic, and polycyclic aromatic hydrocarbon groups, including, for example, benzene, naphthalene, anthracene, and pyrene. Typically, aryl groups contain 6 to 10 carbon ring atoms (i.e., (C6-C8) 10 ) aryl). The aromatic ring can be substituted at one or more ring positions with one or more substituents such as halogen, azide, alkyl, aralkyl, alkenyl, alkynyl, cycloalkyl, hydroxyl, alkoxyl, amino, nitro, sulfhydryl, imino, amido, phosphonate, phosphinate, carbonyl, carboxyl, silyl, ether, alkylthio, sulfonyl, sulfonamide, ketone, aldehyde, ester, heterocyclyl, aromatic or heteroaromatic moiety, fluoroalkyl (such as trifluoromethyl), cyano, and the like. The term "aryl" also includes polycyclic ring systems having two or more cyclic rings, where two or more carbons are common to two adjacent rings (the rings are "fused rings"), at least one of the rings is aromatic hydrocarbon, and the other cyclic rings can be, for example, cycloalkyl, cycloalkenyl, cycloalkynyl, aryl, heteroaryl, and / or heterocyclyl. In certain embodiments, the term "aryl" refers to a phenyl group.
[0105] The term "alkylene" refers to a diradical obtained by removing two hydrogen atoms from an aryl group as defined above. In certain embodiments, arylene refers to a disubstituted arene, i.e., an arene in which two positions are substituted with substituents such as halogen, azide, alkyl, arylalkyl, alkenyl, alkynyl, cycloalkyl, hydroxyl, alkoxyl, amino, nitro, sulfhydryl, imino, amide, phosphonate, phosphinate, carbonyl, carboxyl, silyl, ether, alkylthio, sulfonyl, sulfonamide, ketone, aldehyde, ester, heterocyclyl, aromatic or heteroaromatic moiety, fluoroalkyl (e.g., trifluoromethyl), cyano, etc. That is, in certain embodiments, "substituted aryl" is "arylene."
[0106] The term "heteroaryl" is a term of art and as used herein refers to monocyclic, bicyclic, and polycyclic aromatic groups having a total of 3 to 12 atoms and including one or more heteroatoms, such as nitrogen, oxygen, or sulfur, in the ring structure. Exemplary heteroaryl groups include azaindolyl, benzo(b)thienyl, benzimidazolyl, benzofuranyl, benzoxazolyl, benzothiazolyl, benzothiadiazolyl, benzotriazolyl, benzoxadiazolyl, furanyl, imidazolyl, imidazopyridinyl, indolyl, indolinyl, indazolyl, isoindolinyl, isoxazolyl, isothiazolyl, isoquinolinyl, oxadiazolyl, oxazolyl, purinyl, pyranyl, pyrazinyl, pyrazolyl, pyridinyl, pyrimidinyl, pyrrolyl, pyrrolo[2,3-d]pyrimidinyl, pyrazolo[3,4-d]pyrimidinyl, quinolinyl, quinazolinyl, triazolyl, thiazolyl, thiophenyl, tetrahydroindolyl, tetrazolyl, thiadiazolyl, thienyl, thiomorpholinyl, triazolyl, or tropanyl, and the like. "Heteroaryl" refers to an alkyl group, aryl group, or heteroaryl group, which may be substituted at one or more ring positions with one or more substituents such as halogen, azide, alkyl, aralkyl, alkenyl, alkynyl, cycloalkyl, hydroxyl, alkoxyl, amino, nitro, sulfhydryl, imino, amido, phosphonate, phosphinate, carbonyl, carboxyl, silyl, ether, alkylthio, sulfonyl, sulfonamide, ketone, aldehyde, ester, heterocyclyl, aromatic or heteroaromatic moiety, fluoroalkyl (such as trifluoromethyl), cyano, etc. The term "heteroaryl" also includes polycyclic ring systems having two or more cyclic rings, where two or more carbons are common to two adjacent rings (the rings are "fused rings"), at least one of the rings is an aromatic group having one or more heteroatoms in the ring structure, and the other cyclic rings may be, for example, cycloalkyl, cycloalkenyl, cycloalkynyl, aryl, heteroaryl, and / or heterocyclyl.
[0107] The term "heteroarylene" refers to a diradical obtained by removing two hydrogen atoms from a heteroaryl group as defined above. In certain embodiments, heteroarylene refers to a disubstituted heteroarene, i.e., a heteroarene in which two positions are substituted with substituents such as halogen, azide, alkyl, arylalkyl, alkenyl, alkynyl, cycloalkyl, hydroxyl, alkoxyl, amino, nitro, sulfhydryl, imino, amide, phosphonate, phosphinate, carbonyl, carboxyl, silyl, ether, alkylthio, sulfonyl, sulfonamide, ketone, aldehyde, ester, heterocyclyl, aromatic or heteroaromatic moiety, fluoroalkyl (e.g., trifluoromethyl), cyano, etc. That is, in certain embodiments, "substituted heteroaryl" is "heteroarylene."
[0108] The term "aralkyl" or "arylalkyl" is a term of art and, as used herein, refers to an alkyl group substituted with an aryl group, where the moiety is attached to the parent molecule through the alkyl group.
[0109] The term "heteroaralkyl" or "heteroarylalkyl" is a term of art and, as used herein, refers to an alkyl group substituted with a heteroaryl group, appended to the parent molecular moiety through an alkyl group.
[0110] The term "alkoxy" as used herein means an alkyl group, as defined herein, appended to the parent molecular moiety through an oxygen atom. Representative examples of alkoxy include, but are not limited to, methoxy, ethoxy, propoxy, 2-propoxy, butoxy, tert-butoxy, pentyloxy, and hexyloxy.
[0111] The term "alkoxyalkyl" refers to an alkyl group substituted with an alkoxy group.
[0112] The term "alkoxycarbonyl" means an alkoxy group, as defined herein, appended to the parent molecular moiety through a carbonyl group, as defined herein, represented by -C(=O)-. Representative examples of alkoxycarbonyl include, but are not limited to, methoxycarbonyl, ethoxycarbonyl, and tert-butoxycarbonyl.
[0113] The term "alkylcarbonyl," as used herein, means an alkyl group, as defined herein, appended to the parent molecular moiety through a carbonyl group, as defined herein. Representative examples of alkylcarbonyl include, but are not limited to, acetyl, 1-oxopropyl, 2,2-dimethyl-1-oxopropyl, 1-oxobutyl, and 1-oxopentyl.
[0114] The term "arylcarbonyl," as used herein, means an aryl group, as defined herein, appended to the parent molecular moiety through a carbonyl group, as defined herein. Representative examples of arylcarbonyl include, but are not limited to, benzoyl and (2-pyridinyl)carbonyl.
[0115] The terms "alkylcarbonyloxy" and "arylcarbonyloxy" as used herein mean an alkylcarbonyl group or an arylcarbonyl group, as defined herein, appended to the parent molecular moiety through an oxygen atom. Representative examples of alkylcarbonyloxy include, but are not limited to, acetyloxy, ethylcarbonyloxy, and tert-butylcarbonyloxy. Representative examples of arylcarbonyloxy include, but are not limited to, phenylcarbonyloxy.
[0116] The terms "alkenoxy" or "alkenoxyl" mean an alkenyl group, as defined herein, appended to the parent molecular moiety through an oxygen atom. Representative examples of alkenoxyl include, but are not limited to, 2-propene-1-oxyl (i.e., CH=CH-CH-O-) and vinyloxy (i.e., CH=CH-O-).
[0117] The term "aryloxy" as used herein, means an aryl group, as defined herein, appended to the parent molecular moiety through an oxygen atom.
[0118] The term "heteroaryloxy" as used herein, means a heteroaryl group, as defined herein, appended to the parent molecular moiety through an oxygen atom.
[0119] The term "carbocyclyl," as used herein, refers to a monocyclic or polycyclic (e.g., bicyclic, tricyclic, etc.) hydrocarbon radical containing 3 to 12 carbon atoms and which is fully saturated or has one or more unsaturated bonds, and for the avoidance of doubt, the degree of unsaturation does not result in an aromatic ring system (e.g., phenyl). Examples of carbocyclyl groups include 1-cyclopropyl, 1-cyclobutyl, 2-cyclopentyl, 1-cyclopentenyl, 3-cyclohexyl, 1-cyclohexenyl, and 2-cyclopentenylmethyl.
[0120] The term "cyano" is a term of art and as used herein refers to --CN.
[0121] The term "halo" is a term of art and, as used herein, refers to -F, -Cl, -Br, or -I.
[0122] As used herein, the term "haloalkyl" refers to an alkyl group, as defined herein, in which some or all of the hydrogens have been replaced with halogen atoms.
[0123] The term "hydroxy" is a term of art and as used herein refers to --OH.
[0124] The term "hydroxyalkyl," as used herein, means at least one hydroxy group, as defined herein, appended to the parent molecular moiety through an alkyl group, as defined herein. Representative examples of hydroxyalkyl include, but are not limited to, hydroxymethyl, 2-hydroxyethyl, 3-hydroxypropyl, 2,3-dihydroxypentyl, and 2-ethyl-4-hydroxyheptyl.
[0125] The term "silyl," as used herein, includes hydrocarbyl derivatives of the silyl (HSi-) group (i.e., (hydrocarbyl)Si-), where the hydrocarbyl group is a monovalent group formed by removing a hydrogen atom from a hydrocarbon, e.g., ethyl, phenyl. The hydrocarbyl group may be a combination of different groups, such as trimethylsilyl (TMS), tert-butyldiphenylsilyl (TBDPS), tert-butyldimethylsilyl (TBS / TBDMS), triisopropylsilyl (TIPS), and [2-(trimethylsilyl)ethoxy]methyl (SEM), and can be varied to provide multiple silyl groups.
[0126] As used herein, the term "silyloxy" means a silyl group, as defined herein, appended to the parent molecule through an oxygen atom.
[0127] It is understood that "substituted" or "substituted with" includes the implicit proviso that such substitution is in accordance with the valences allowed for the substituted atom and substituent, and that the substitution results in a stable compound, e.g., a compound that does not spontaneously undergo transformation, such as by rearrangement, fragmentation, decomposition, cyclization, elimination, or other reaction.
[0128] The term "substituted" is also contemplated to include all permissible substituents of organic compounds. In a broad aspect, the permissible substituents include acyclic and cyclic, branched and unbranched, carbocyclic and heterocyclic, aromatic and nonaromatic substituents of organic compounds. Illustrative substituents include, for example, those described herein above. The permissible substituents can be one or more and the same or different for appropriate organic compounds. For purposes of this invention, heteroatoms such as nitrogen can have hydrogen substituents and / or any permissible substituents of organic compounds described herein that satisfy the valence of the heteroatom. This invention is not intended to be limited in any manner by the permissible substituents of organic compounds.
[0129] In certain embodiments, optional substituents may include, for example, halogen, haloalkyl, hydroxy, carbonyl (e.g., carboxyl (—COOH), alkoxycarbonyl, formyl, or acyl), thiocarbonyl (e.g., thioester, thioacetate, or thioformate), alkoxy, alkenyloxy, alkynyloxy, phosphoryl, phosphate, phosphonate, phosphinate, amino (including alkylamino and dialkylamino), amido, amidine, imine, cyano, nitro, oxo (═O), azido, sulfhydryl, alkylthio, sulfate, sulfonate, sulfamoyl, sulfonamido, sulfonyl, silyl, silyloxy, heterocycloalkyl, cycloalkyl, alkyl, alkenyl, alkynyl, aryl, heteroaryl, arylalkyl, or heteroarylalkyl groups.
[0130] For purposes of this invention, chemical elements are identified according to the CAS version of the Periodic Table of the Elements (Handbook of Chemistry and Physics, 67th Ed., 1986-87, inside cover).
[0131] Other chemical terms herein are used in accordance with conventional usage in the art as exemplified by The McGraw-Hill Dictionary of Chemical Terms (ed. Parker, S., 1985), McGraw-Hill, San Francisco (incorporated herein by reference). Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0132] In the compounds of the present invention, atoms that are not specifically designated as a particular isotope are meant to represent any stable isotope of that atom.Unless otherwise stated, when a position is specifically designated as "H" or "hydrogen", it is understood that the position has hydrogen at its natural abundance isotopic composition.Also, unless otherwise stated, when a position is specifically designated as "D" or "deuterium", it is understood that the position has deuterium at an abundance of at least 3340 times the natural abundance of deuterium (0.015%) (i.e., at least 50.1% deuterium incorporation).
[0133] As used herein, the term "isotopic enrichment factor" means the ratio between the isotopic abundance and the natural abundance of a specified isotope.
[0134] In various embodiments, when compounds of the invention are designated as deuterium, such compounds have an isotopic enrichment factor for each designated deuterium atom of at least 3500 (52.5% deuterium incorporation at each designated deuterium atom), at least 4000 (60% deuterium incorporation), at least 4500 (67.5% deuterium incorporation), at least 5000 (75% deuterium), at least 5500 (82.5% deuterium incorporation), at least 6000 (90% deuterium incorporation), at least 6333.3 (95% deuterium incorporation), at least 6466.7 (97% deuterium incorporation), at least 6600 (99% deuterium incorporation), or at least 6633.3 (99.5% deuterium incorporation). [Example]
[0135] The invention will now be generally described, but will be more readily understood by reference to the following examples, which are included solely for the purpose of illustrating certain aspects and embodiments of the invention and are not intended to be limiting thereof.
[0136] Example 1 - Exemplary Materials and Methods Materials. Commercial reagents were purchased from Sigma-Aldrich, Combi-Blocks, TCI America, Fisher Scientific, OakWood Products, or Strem and used without further purification unless otherwise noted. Solvents (1,4-dioxane, DMF, and toluene) were purchased from Fisher Scientific and purged with argon for 1 hour before use in the reactions. Deionized water was purified via a Milli-Q water filtration system. Na2CO3 and K3PO4 were both anhydrous and used as purchased from the supplier.
[0137] Characterization: Proton nuclear magnetic resonance ( 1 H NMR spectra were recorded at room temperature on one of the following instruments: a Bruker 500 MHz spectrometer equipped with a broadband cryoprobe (CB500) or a Bruker Avance Neo 600 MHz spectrometer equipped with a Prodigy BBO-BB cryoprobe (B600). Chemical shifts (δ) are reported in parts per million (ppm) downfield from tetramethylsilane and are referenced to residual protium in the NMR solvent (CDCl, δ = 7.26 ppm, center line; (CD)CO, δ = 2.05 ppm, center line; (CD)SO, δ = 2.50 ppm, center line; CDOD, δ = 3.31 ppm, center line). Data are reported as follows: chemical shift, multiplicity (s = singlet, d = doublet, t = triplet, q = quartet, quint = quintet, sept = septet, m = multiplet, dd = doublet of doublet, dt = triplet doublet, dq = quartet doublet), coupling constant (J) in Hertz (Hz), and integral. 13C NMR spectra were recorded at room temperature on one of the following instruments: CB500 and B600. Chemical shifts (δ) are reported in ppm downfield from tetramethylsilane and are referenced to the carbon resonance in the NMR solvent (CDCl3, δ = 77.0 ppm, centerline; (CD3)2CO, δ = 29.84 ppm, centerline; (CD3)2SO, δ = 39.52 ppm, centerline; CD3OD, δ = 49.00 ppm, centerline). Fluorine nuclear magnetic resonance ( 19Chemical shifts for F NMR are reported in parts per million downfield from chlorotrifluoromethane (δ = 0). High-resolution mass spectra (HRMS) using TOF-MS were performed by Furong Sun and Haijun Yao at the University of Illinois School of Chemical Sciences Mass Spectrometry Laboratory. Liquid chromatography-mass spectrometry (LCMS) was performed using an Agilent 1260 Infinity II HPLC equipped with a diode array UV-Vis detector (DAD) coupled to an Agilent 6530C quadrupole time-of-flight (QTOF) MS detector. For LCMS, a YMC America Triart C18 150 mm long x 4.6 mm 120 Å internal diameter, 5 μm particle size column (model number: TA12S05-1546WT) was utilized, and 1 μL of the undiluted crude reaction mixture was injected for quantification under standard separation conditions. The standard separation conditions for all molecules (except for product 27, which coeluted with a peak) were: column compartment temperature set at 25 °C, gradient A: water containing 0.1% formic acid and B: LCMS-grade acetonitrile, flow rate 1 mL / min: 0 → 1 min 100% A 0% B, 1 → 20 min 0% A 100% B, 20 → 22 min 0% A 100% B, 20 → 22 min 0% A 100% B, 22 → 23 min 95% A 5% B, 23 → 24 min 100% A 0% B, DAD measurement signals: 200 nm, 230 nm, 250 nm, 270 nm, 290 nm, 310 nm, 380 nm, 430 nm (each with a 4 nm bandwidth). The separation conditions for product 27 were the same as above except for the gradient changes: 0 → 1 min 100% A 0% B, 1 → 20 min 60% A 40% B, 20 → 21 min 0% A 100% B, 21 → 27 min 0% A 100% B, 27 → 28 min 100% A 0% B, 28 → 29 min 100% A 0% B.
[0138] Purification. Preparative high performance liquid chromatography (prep HPLC) was performed on an Agilent 1200 series instrument equipped with a Waters Sunfire Prep C18 OBD 5 μM 30×150 mm column.
[0139] Automated synthesis. All reactions were carried out in 40 mL glass I-Chem vials with PTFE-lined septum caps (Fisher Scientific catalog number 05-719-106) equipped with rare-earth stir bars (oval, 5 mm diameter, 10 mm length, Sigma-Aldrich catalog number Z671622) using an automated synthesizer under positive argon pressure. Solid chemicals (catalysts, bases, building blocks, internal standards) were weighed into I-Chem vials under air using a Mettler-Toledo Quantos automated powder dispensing analytical balance set to a tolerance of ±2% by weight. Liquid chemicals (some of the halide building blocks) were added via microliter pipette. The reaction vials were loaded into the synthesizer, and the automated synthesis procedure was performed. Typically, this involved atmosphere exchange of the reaction vials using an automated Schlenk line procedure (applying vacuum for 5 min, followed by backfilling with argon for 1 min (10 6-min cycles, total time 1 h) to degas 36 reaction vials), followed by solvent addition via automated syringe pump (pre-programmed selection of up to 20 different solvents per vial), heating and stirring for a given reaction time, and cooling. Reactions were then purified by preparative HPLC for molecular characterization and generation of LCMS response factor curves, or quantified by direct injection into LCMS.
[0140] Reaction Data Analysis. The reaction data sets shown in the figures of this application were visualized and analyzed using GraphPad Prism 9.3.1. Statistical comparisons of common reaction conditions utilized repeated measures one-way ANOVA with Geisser-Greenhouse correction and Sidak's multiple comparison test.
[0141] Code and Data Availability. All data and code generated as part of this study are freely accessible, either in the Supplementary Materials or in an open repository. Code for example simulations, model selection, and substrate clustering, as well as machine-readable and human-readable versions of the reaction data, are freely accessible on Zenodo. The automated synthesis code, machine parts list and assembly guide, data-mined commercially available building block set, checklist for reporting and evaluating machine learning models, and tabular numerical data underlying Figures 2–4 have been deposited on Zenodo. The code used in the building block selection process is available on Zenodo. All code related to the closed-loop optimization is freely available on Zenodo.
[0142] Example 2 - Automated Synthesizer Automated Synthesizer Description. The automated synthesizer was built using the same reconfigurable valves and syringe pumps as our previously reported iterative small molecule synthesizer in a unique configuration suitable for parallel reaction testing and driven by custom LabVIEW software. Each reaction vial was connected to a solvent reservoir via a syringe needle, PEEK tubing, and a Luer-lock fitting via two computer-controlled 9-port 4-valve modules and a 10 mL computer-controlled syringe pump (Figure 9). Solvent was stored in a Pyrex glass media bottle and connected to both the argon manifold and valve module via PEEK tubing, a Luer-lock fitting, and a 3-port Cole-Parmer VapLock solvent delivery cap. The third port of the solvent cap was sealed with a Luer plug fitting, which was temporarily removed when refilling solvent through the system to allow argon sparging. Each reaction vial was also connected to an argon / vacuum manifold via a passive 9-port valve and solenoid valve via a syringe needle, PEEK tubing, and Luer-lock fitting (Figure 10). The argon cylinder was maintained at a regulator setting of approximately 5 PSI and did not need to be replaced throughout the study. The solenoid valve was controlled by a National Instruments USB-6001 Multifunction I / O Device and powered by a custom 4-Channel Solenoid Driver from University of Illinois Electronic Services (Figure 11). Similar solenoid valve drivers are commercially available. Each of the three computer-controlled hot plates contained a circular heating block with 12 positions for reaction vials, evenly spaced around a centrally located heating probe. This heating block effectively eliminated any "edge effects" associated with typical square or rectangular heating blocks, which can result in differences in stirring and heating depending on reaction location through different relative distances from the heating probe (heating) and magnetic field (stirring) (Figure 12).The height of the heating block was also configurable and could be increased by stacking aluminum heating blocks or decreased by filling each reaction position with an equal amount of sand. For the automated synthesis in this study, the heating block was as shown in Figure 8 and heated to a reaction volume height of 8 mL for each reaction vial. This allowed the reaction solvent to be heated to reflux without redistribution across the manifold or among the reaction vials, as the remainder of each vial exposed to ambient air acted as an air-cooled reflux condenser. Custom LabVIEW software was designed as a single general code run with the following controllable parameters: number of parallel reactions (1–36 reactions), solvent selection per reaction position (1–20 solvents), organic solvent volume (mL), water volume (mL for varying organic solvent:water ratios), reaction temperature (°C), reaction time (h), and stirring speed (rpm) (Figure 13). The automated synthesizer user simply loaded the relevant chemicals into the synthesizer, selected reaction parameters from the "front panel" of the LabVIEW code, and ran the program. The code then prompts the user to ensure the instrument is online and that there is sufficient solvent in the reservoirs: if the user clicks "OK" on this prompt, the automation code will run the complete synthesis. A complete parts list required to build this synthesizer is available in Table 1. [Table 1]
[0143] Example 3 - Platform Verification To validate the platform's performance and reproducibility in the context of heteroaryl cross-coupling, we investigated its performance for the synthesis of furanyl indole 1, with reaction yield as the readout. We first investigated the role of the boron species and found that in situ release of MIDA boronate, compared with boronic acids and esters, resulted in the highest overall levels of reproducibility and yield (Figure 14). This test reaction featured the most statistically common palladium, base, solvent, and temperature in the heteroaryl cross-coupling literature (Reaxys database, references). This result suggests that both the stability and reactivity of the boron species and halide are important for reproducible high yields. When the boronic acid is added slowly, the relative concentration of the boronic acid in the reaction mixture is kept low, and decomposition is prevented by its slow addition to the reaction: the resulting large variability suggests that the halide coupling partner may be decomposed during the slow addition time. Rapid addition of the boronic acid leads to higher yields and more reproducible results, but yields potentially decrease due to the decomposition of the unstable boronic acid over time at high concentrations. The use of pinacol boronate esters prevents decomposition of the boron species, improving yield and reproducibility, but potentially allows the slower, more hindered pinacol ester to decompose the halide over time, resulting in a capped yield of nearly 80%. The in situ hydrolysis kinetics of the MIDA boronate appear to be commensurate with ensuring high yields and highly reproducible automated reactions. While we did not measure the hydrolysis kinetics in this study, our previous studies of the hydrolysis kinetics of 2-furanyl MIDA boronate under similar conditions suggest that it is likely to be converted to the boronic acid over a 1-hour period. Our previous use of Pd(PPh3)4 required a glovebox due to the air sensitivity of the catalyst, but we were surprised to find that the use of an air-stable palladium precatalyst (e.g., SPhos Pd G4) weighed in air reduced yields despite passive argon flow in the automated synthesizer.We manually conducted a series of head-to-head control experiments (Figure 15) to probe the environmental control necessary to prevent oxygen-mediated yield loss. These experiments revealed that residual air in the headspace of the reaction vial was a contributing factor to the yield loss and that passive argon purging was ineffective at displacing the atmosphere. We also noted that a manual Schlenk line vacuum / fill cycle replicated similar performance to a reaction set up entirely inside a glove box. This new understanding provided us with the option of either manually performing the Schlenk line procedure on the reaction vial before introduction into the synthesizer or automating the process after introduction into the instrument. We designed a simple Schlenk line integrated into the machine (see Figure 10), consisting of a solenoid valve isolating the argon cylinder from the vacuum pump. We measured the time required to reach maximum vacuum with the vacuum connected to a 36-well reaction gas manifold by recording (approximately 4 minutes; the vacuum pump is quietest at maximum vacuum) and tested cycles of 5 minutes of vacuum followed by 1 minute of argon backfill for their effectiveness in promoting heteroaryl cross-coupling (Figure 16). From this data, it is clear that there is a direct correlation between the number of automated Schlenk cycles and the corresponding reaction yield, exceeding that of manual Schlenk cycles and rendering the yield statistically indistinguishable from that of a glovebox system using ≥9 cycles. For the remainder of this study, we utilized 10 Schlenk cycles as a conservative approach, requiring only 1 hour of automation time. The design and optimization of this equipment resulted in a ±2% yield reproducibility between replicates throughout the optimization.
[0144] Example 4 - Problem definition and AI framework for reaction generality optimization The goal of this study is to discover a general set of reaction conditions for a class of small molecules, specifically the heteroaryl-(hetero)aryl Suzuki-Miyaura cross-coupling, where at least one coupling partner is a heteroaryl ring. Heteroaryl cross-coupling is a problem of great impact because the motif is highly representative in materials science, natural products, and drug molecules. While the literature demonstrates a great deal of structural diversity in heteroaryl Suzuki coupling, there is a surprising lack of generality in reaction conditions. Scientists in this subfield typically test a small set of conditions based on their own knowledge or experience with catalysts, bases, solvents, ligands, and temperatures. The problem is that only the highest-yielding conditions per substrate pair are typically reported. Conversely, reaction optimization studies report a wide range of reaction conditions for a very limited range of substrates. To solve this problem generally for a broad class of molecules, a diverse and large set of representative substrate pairs must be tested under many general reaction conditions. The reaction conditions with the highest average yield across a range of representative substrates are considered the most general. Optimizing generality in this manner is novel.
[0145] To mathematically formulate this problem, we denote the set of possible reaction conditions as C = {c}, the set of substrate pairs as S = {s}, and the reaction yield as a function of the substrates and reaction conditions y(s,c). The objective function f(c) that we aim to maximize is given by the following equation:
number
number
[0146] At first glance, with a minimal number of experiments, generalThe problem of identifying f(c) is similar to standard Bayesian optimization (BO). However, there are substantial differences between all BO algorithms, with each experiment / measurement performed immediately providing information about the objective function desired to be optimized. In contrast, the experimental evaluation of f(c) in our problem requires multiple experiments (because the summation in Eq. 1 is over the entire set S), i.e., determining f(c) for given conditions requires experiments with all pairs of substrates in the set S. To address this problem, we propose a surrogate model for predicting reaction yields.
number
number
[0147] In selecting an appropriate model, we considered two factors in the above scheme: a) uncertainty estimation and b) estimation performance. Regarding (a), we sought a model that provided prediction uncertainty commensurate with the prediction error. For example, highly reliable predictions with large errors are undesirable. Regarding (b), we wanted to find a means to test various possible models before actually starting our own closed-loop optimization. These aspects are detailed in the following subsections.
[0148] Benchmarking Prediction Uncertainty. We first decided to use a previously published dataset of 5,760 Suzuki coupling reactions to evaluate the uncertainty of various models in predicting reaction yields. This set was chosen because (i) the reactions were obviously directly relevant to our current study and (ii) they were performed automatically (in a flow reactor) in a highly standardized manner, and therefore the reported yields were expected to be free of random error. Each reaction in this set was run under approximately 300 different conditions, varying in the solvent, ligand, and base used. A drawback was that the substrates were based on only two structural motifs: quinoline and indazole, which differed in the type of boronic acid moiety (boronic acid, trifluoroborate, boronic acid pinacol ester) and coupling partner (iodide, bromide, chloride, triflate). This meant that this dataset was not suitable for training a model to suggest "next" substrates. For each combination of structural motifs, we decided to include only one type of halide and boronic acid functional group to avoid redundancy, leaving 588 reactions that were used to explore various uncertainty estimation techniques.
[0149] All of the models we examined were based on neural networks (NNs). There are several different approaches for determining the uncertainty of NN predictions. We tested the following: i) traditional ensembles, ii) bootstrap ensembles, iii) Monte Carlo dropout, iv) Bayesian neural networks with flip-out layers, v) NNs with direct uncertainty estimation using negative log-likelihood loss (NLL), and vi) a combination of neural networks and Gaussian processes (GPs). The last method tested (vi) uses GPs as a model for statistical estimation, where the NN part can estimate (1) the mean value of the function (which the GP then estimates the uncertainty), (2) the covariance between different points (how the uncertainty varies with distance from the observation point), or (3) both. The inclusion of a GP supplemented with a NN component seemed particularly attractive, since in the absence of meaningful metric quantification differences between reaction conditions, we could learn it from the data by co-training the GP with a NN kernel (a similar approach has recently been used in Bayesian optimization).
[0150] The NN model was implemented using Keras and Tensorflow with GPflow to model Gaussian processes. Uncertainty estimation methods (i)–(v) were tested on two feedforward neural networks with two hidden layers. The first network had two dropout layers, and the second had no dropout regularization. For each network, the architecture hyperparameters were first optimized using hyperas. In contrast, for the neural component of the GP model (vi), a grid search was performed over a range of architectures (shown in Figure S12) defined by the following parameters: 1) the embedding layer N e dimensionality (10, 50, 100, or 200 neurons), 2) the next hidden layer N h dimension of (50, 100, 150, or 200 neurons with ReLU activation), 3) hidden layer LH (1, 2, or 3), 4) the dropout probability D (ranges from 0.1 to 0.9 with 0.1 intervals), 5) the output layer N o The embedding component was built on top of the dimensionality (1, 10, 50, 100, 150, or 200 neurons with ReLU activation), as well as 6) an L2 regularization weight (either 0.001, 0.01, or 0.05). The kernel function built on top of this embedding component was chosen as either a radial basis function, Matern32, or Matern52.
[0151] For all AI models considered here, we use the same input representation of the reaction, which is a concatenation of the following vectors: 1) the FCFP6 fingerprint representation of the bromide coupling partner, 2) the FCFP6 fingerprint representation of the BMIDA coupling partner, 3) a scalar indicating the temperature, 4) a one-hot encoding of the solvent, 5) a one-hot encoding of the base, and 6) a one-hot encoding of the ligand. For the fingerprints, we represent them in a "dense format" created as follows: First, we collect all identifiers (numbers indicating specific substructures of the fingerprint vector present in all substrates in our range (e.g., bromide)) and count the number of occurrences of each identifier. Next, we sort these identifiers according to their frequency of occurrence in the dataset and assign each to a component of a "dense" fingerprint vector (the length of that vector therefore corresponds to the total number of identifiers). Finally, for a given molecule, we set the value of the vector's component using the count of the corresponding identifier in the molecule's FCFP6 fingerprint. Note that with this representation, the resulting uncertainty in the AI model depends not only on the similarity between substrates in the training / test set but also on the conditions, including assumed non-additive effects (e.g., the same change in conditions may affect substrate pair A more than substrate pair B). Furthermore, our model before training does not rely at all on the similarity between reaction conditions (i.e., solvent)—everything is inferred (learned) directly from the data.
[0152] Since prediction uncertainty will provide the basis for our acquisition algorithm and is expected to indicate, in particular, the "lack of knowledge" of the model, we analyzed the distribution of uncertainty generated by different models. In doing so, we were interested in models in which the prediction uncertainty varies monotonically with the prediction error, i.e., models in which data points with large prediction errors have higher uncertainty than data points with small prediction errors (predictions of relative values).
[0153] With this in mind, we compared the mean absolute error and the mean prediction uncertainty for both the training and test sets. Specifically, we calculated the difference between the two: the difference in mean error, ΔMAE = MAE test -MAE train , and the difference in mean uncertainty ΔU=U test -U train We evaluated our criteria by examining the relationship between the ΔMAE and ΔU. We required that both ΔMAE and ΔU a) have the same sign and b) have an absolute magnitude greater than 1% (this was true for all ΔMAEs in our comparisons). In all experiments, a common randomly selected 20% of the entire dataset was used as the training set, and the remaining 80% served as the test set. This division was intended to mimic the early stages of closed-loop optimization, where the model has knowledge of only a small portion of the entire space.
[0154] Among all trained NN models (Table 2), only NLL and flip-out do not meet our criteria, since the mean uncertainties of the training and test sets are virtually the same, whereas the test MAE is generally significantly higher than that for training. Among the remaining non-GP models, the best results are obtained from the ensemble approach. However, because their performance is comparable, we decided to conduct additional tests on a smaller training set, 5% of the entire dataset (Table 3). In this test, both the bootstrap model and the traditional ensemble model met our criteria, with the latter achieving a lower MAE. test was provided.
[0155] As for the GP-based models, they all met our criteria (Table 2), with the GP with an NN-based kernel providing a lower test error (24.7%). While this value is significantly higher than that of the "pure" NN model (15.2% for the conventional ensemble mentioned above), the training MAE for the GP (NN) is 7.3%, lower than that for the NNE (8.2%). By definition, the GP calculates a probability distribution taking into account the observed data points, so the combined model is guaranteed to converge to the ground truth with dataset expansion. Therefore, we choose the GP with an NN kernel for further testing, along with the conventional neural network ensemble. [Table 2] [Table 3]
[0156] Above, we performed model selection based on a criterion that aims to reflect, on average, the relationship between model uncertainty and prediction error, or, intuitively, whether the model can "predict" the extent of its own error. In this experiment, we focused on two aspects of the correlation between model error and predicted uncertainty: qualitative agreement, reflected by the Kendall τ coefficient, and z-score.
number
[0157] Selection algorithms for closed-loop experiments. Next, we trained a model, the reaction space N substrate pairs ×N conditions (which reactions have already been performed), as well as the selection strategy, T conditions and T substratesThe proposed experimental selection was considered taking into account the above. The entire protocol is outlined in Figure 19. Briefly, we 1) use the model to predict yields across the entire reaction space (i.e., the space of all unperformed reactions) along with the prediction uncertainty, 2) convert this result into an objective function, and 3) select substrate pairs and reaction conditions to test in the next experiment according to a selection strategy. In general, we expect the selection strategy to return a "score" for each candidate, with the best candidate being indicated by the highest score. T conditions With respect to the objective f(c), we obtain an estimate of the condition acquisition function a c In the described study, we examine two well-established BO selection strategies: a) Probability of Improvement (PI), i.e., the probability of improvement in the new condition c new is better than the best known, and b) the expected improvement (EI), i.e., c new Expected improvement over the best case to date, calculated for a probability distribution of . Both (a) and (b) are based on the assumption that, for each c, the distribution of the objective function is normal, with a mean given by f(c) and a standard deviation given by the uncertainty in f(c). T substrates For , it is considered to score substrates that have not yet been explored under given conditions. In particular, we consider two such strategies: i) random selection, and ii) selection with maximum uncertainty (discussed at the beginning of Section 3).
[0158] Simulation of closed-loop optimization for general conditions Next, we simulated a "closed-loop" optimization post hoc using previously published experimental data in anticipation of our own experiments. The dataset we used for uncertainty calibration involved only two distinct structural motifs for each of the coupling partners, and therefore could not be used to investigate substrate selection strategies. To our knowledge, there are no publicly available datasets of Suzuki reactions performed in a standardized manner and covering a broader substrate range. Therefore, we decided to explore other Pd-catalyzed reactions and selected the dataset reported by Santanilla et al. for C-N couplings (Buchwald-Hartwig aminations). The new dataset contains 1,536 robotically executed reactions, covering 48 different conditions for each of 32 different substrate pairs.
[0159] Next, we performed simulations of the closed-loop protocol as follows: First, we randomly selected 5% of the total data as an initial training set and used it to train various models for predicting the yield y(s,c) (these predictions were then used to calculate f(c)). Then, we performed closed-loop iterations until the entire data set was explored. During each iteration, we (i) trained a given model on the data already seen, (ii) selected the "next substrate pair" according to the procedure in Figure 19, and (iii) expanded the training set by including the proposed reaction. To account for the random component of this procedure (arising from different choices of the initial set and random initialization of the NN parameters), we performed 100 independent simulations for each model and collected statistics.
[0160] performance metrics We used two metrics to describe the quality of the prediction: the rank of the best condition and the yield of the condition predicted to be the best. The first metric specifies the position of the best condition (actual, experimental) in the ranking generated according to the model's prediction; for example, a value of 0 means that the model correctly predicts a particular condition to be better than any other condition. The second metric indicates the experimental average yield of the condition predicted to be the best. In other words, the first metric describes how good the model is at finding the best condition, and the second metric describes how good the estimation of the corresponding average yield is.
[0161] Closed-Loop Simulation In the following experiments, we focus on two models from the previous section: Gaussian Process with Neural Network Kernel (hereafter denoted as GP(NN)), and Neural Network Ensemble (NNE). Note that the ensemble type used here refers to a "traditional" ensemble without dropout regularization (as opposed to a "bootstrap" ensemble). That is, each model in the ensemble was trained on the same data but with different values of the initial parameters. A combination of the following selection strategies: conditions For , the expected improvement (EI) and probability of improvement (PI) were examined, while for T substrates For , we used either a random strategy (randSbs) or the selection of the substrate with the highest prediction uncertainty (maxUnc). In the GP(NN) model, all four possible combinations lead to similar model performance (Figures 20 and 21).
[0162] A comparison between GP(NN) and NNE shows that NNE outperforms GP(NN) in both ranking and proposal terms, providing a high average yield during the first approximately 150 iterations, after which both methods eventually converge to comparable results.
[0163] The relative poor performance of GP(NN) can be attributed to the neural network used as the kernel. In the current implementation, stochastic gradient descent cannot be used during training (parameter gradients and updates are calculated for the entire training set at once), which effectively causes the NN component to get stuck in the nearest local minimum, whose location depends on the initial parameter values. To overcome this limitation, we decided to create an ensemble of GP(NN)s (abbreviated as GPE(NN)) (black curves in Figures 20 and 21), which proved to be superior to NNEs (similar modifications are already known in the literature). With these improvements in place, we moved on to optimizing the size of the ensemble and the type of acquisition function. We note that because extensive (brute-force) testing of GPE(NN) models of different sizes would be quite costly, we decided to train a large GPE(NN) model consisting of 2000 independent GP(NN) models and record all intermediate results. Next, we could simulate the results of smaller GPE(NN)s by selecting a subset of the desired size from the recorded results. For each size range, from 2 to 1200, we randomly selected 500 such samples and used such ensembles to calculate the probability that the smaller GPE(NN) would make the same selection as the reference 2000-model GPE(NN) in the very first step of the closed-loop simulation (Figure 22). For the PI acquisition function, an ensemble of as few as 10 models was able to select the same condition as the reference with a probability greater than 0.98, while for the PI+maxUnc selection strategy, an ensemble consisting of 100 models was able to predict the same condition and substrate with a probability of approximately 0.8. For EI, the probability of selecting the same substrate and condition (as the reference, larger ensemble) converged very slowly; even for a 1000-member ensemble, the corresponding probability was approximately 0.78–0.79. As a result, we selected the PI acquisition function for further testing and concluded that an ensemble of 100 models is a reasonable compromise between computational time and accuracy.
[0164] Next, we compared GPE(NN) with 20 and 100 models using the PI+Random and PI+maxUnc selection strategies (Figures 23 and 24). For ensembles consisting of 20 GPE(NN), maxUnc yields slightly worse results than random substrate selection, but ultimately (after step 150) converges to essentially the same results. For ensembles of 100 members, uncertainty-based substrate selection yielded slightly better results, while random substrate selection yielded the same results as for smaller ensembles. These results suggest that the maxUnc substrate selection procedure outperforms random selection only for sufficiently large NN ensembles.
[0165] Simulation of stopping criteria In this section, we consider uncertainty-based stopping criteria. In particular, we consider the average uncertainty over the entire search space.
number
number
number
[0166] Example 5 - Data mining, clustering, and reaction scope selection of building blocks Data mining of substrate-scope building blocks To define generality in the context of heteroaryl Suzuki coupling, we needed to prospectively determine the substrate scope with the largest representation of accessible heteroaryl chemical space. This required the creation of a large in silico molecular LEGO-kit consisting of all commercially available (hetero)aryl halides and (hetero)aryl MIDA boronates. To data mine this library, we used a web and database scraper to find commercially available building block molecules from common chemical suppliers. This scraper cataloged all aryl halides, heteroaryl halides, and MIDA boronates reported in the PubChem database. These structures were then scraped and cross-referenced with actual pricing data from the back-end and front-ends of the world's largest and most trusted fine chemical suppliers (Sigma-Aldrich, Oakwood Chemical, Combi-blocks, and others). To achieve this, despite the large amount of data involved and its relative inaccessibility, we explored the substructures of all known heteroarenes. These final datasets were filtered to ensure they were low-priced (minimum $150 per bottle), not made on demand, in stock, not backordered, and chemically compatible with Suzuki coupling (examples of compatible functional groups: hydroxyl, amine, ester, amide, etc.; examples of chemically incompatible functional groups: ketene, isocyanate, multiple reactive halogens). Through this process, building blocks were narrowed from millions in databases and literature to hundreds of thousands of chemicals listed by chemical suppliers, priced at tens of thousands, and ultimately to >5,000 currently commercially available, chemically compatible building blocks. This final list is highly diverse and represents the entire chemical space currently accessible via heteroaryl Suzuki coupling. Importantly, this list includes examples of all heteroaryl substructures found in drug, natural product, and materials databases, with various functional groups, with or without protecting groups.
[0167] Clustering and selection of representative substrates With the establishment of an in silico building block library, we needed to define a representative subset of this chemical space that could be practically purchased and stored in our laboratory for use in automated experiments. To do this, we algorithmically clustered commercially available (hetero)aryl halides by their common (hetero)aromatic ring substructure and pendant functional groups and applied a stratified clustering strategy (Figure S21) to select molecules that best represented each section of available chemical space. In doing so, we selected 54 (hetero)aryl halides, which we then purchased (Figure 27). Next, we selected an equally sized set of 54 (hetero)aryl MIDA boronates, which we then purchased (Figure 28). Because commercially available (hetero)aryl MIDA boronates are scarce compared to halides, we did not apply a clustering strategy to the selection of these building blocks; instead, we selected them based on maximizing the representation of the heteroaryl substructure. 2-Pyridyl MIDA boronate was excluded from this selection process because it requires an additive that was not considered in subsequent optimization. Overall, the combination of (hetero)aryl halides and MIDA boronates building blocks yields a substrate range containing >2900 unique products. Searching this entire space (multiplying dozens of possible conditions for each substrate pair; see below) is technically infeasible, and therefore we pursued the following strategy. First, we selected a small set of 11 substrate pairs to minimize the mutual similarity of the resulting products (achieved with a greedy algorithm based on Tanimoto similarity) (Figure 3C). For these products, we performed (time-consuming) normalization of the UV-Vis spectra (LCMS-UV / Vis response factor curves, see sections 8 and 9), which subsequently allowed us to automatically determine the yields of the reactions. These 11 pairs, and a subset thereof, that we tested under the initial set of conditions (Figure 3E) were tested under different conditions during the AI-guided optimization phase.Once this step was complete, we then selected another set of 20 substrate pairs that were (i) maximally different from each other and (ii) maximally different from the 11 substrate pairs used during optimization (Figure 5A). In other words, we tested the optimized generic conditions on a diverse test set of molecules not seen during model training.
[0168] To compare this selection strategy with that employed during traditional reaction optimization, we examined the substrate scope of the widely used heteroaryl cross-coupling report (Figure 29), which was used as a benchmark in this study. We plotted the products from JACS 2009, along with the training and test sets examined in this study, using T-distributed stochastic neighbor embedding (t-SNE) mapping (Figure 30). In this plot, the substrate scope from JACS 2009 appears to form a cluster primarily due to the commonly occurring smaller selection of chemical building blocks, while the training and test sets examined in this report remain spread out throughout the entire space. This result suggests that the reference JACS 2009 represents a relatively small chemical space relative to the training and test sets examined in this study, and that the latter more accurately represent the entire chemical space. This example primarily highlights the inherent shortcomings of traditional reaction optimization processes, namely, that the qualitative aspects of substrate scope selection do not consider or guarantee the generality of the reaction. To further compare these two sets, a similarity metric was devised that can distinguish pairs of building blocks that differ in the arrangement of the halide and MIDA boronate (i.e., halide in substrate 1 and MIDA boronate in substrate 2 versus halide in substrate 2 and MIDA boronate in substrate 1; this is of great chemical importance because the reactivity of these moieties differs depending on the molecular structure).
[0169] This metric worked as follows, using Tanimoto similarity with the Morgan fingerprint with radius=3 nBits=2048: 1. The similarity between the two reactions was measured as follows: (Similarity between the first bromide and the second bromide) * (Similarity between the first MIDA boronate and the second MIDA boronate) 2. Such similarity calculated for each combination of pairs of substrates in the set 3. The average value calculated from all the calculated similarities This approach yielded the following results: The average similarity of all product spaces examined in this study: 0.0314 The average similarity of the training set examined in this study was 0.0427. Average similarity of substrate coverage from reference JACS2009: 0.1136 According to this data, the training set examined in this study is similar in diversity to that of the full set (which it is supposed to represent) and more diverse than the conventional substrate coverage from the reference JACS2009.
[0170] Importantly, the JACS2009 reference is a useful benchmark for the state of the art in the field, given its citations and the widespread use of its optimized protocols. According to SCOPUS, it is cited by 440 scientific publications (98th percentile), with a field-weighted citation impact of 7.67 (where 1 is the average). Google Scholar, including patents, lists 594 citations. To illustrate its current relevance, we have included a Web of Science graph (Figure 31) (reproduced below) showing both the number of publications per year citing this work and the number of citations those papers receive. Overall, this indicates sustained interest in this publication and increasing relevance over time.
[0171] Range of reaction conditions Regarding conditions, we considered four variables: solvent, base, catalyst / ligand, and temperature. In selecting these, we considered not only prevalence in the literature but also diversity. For example, the two most common solvents in the literature are dioxane and dimethoxyethane, but they both belong to the same solvent class of ethers, and we chose only one of them, dioxane. By similar reasoning, we were able to retain only one carbonate group (we previously showed that the nature of the cation does not alter yields), and the functional similarity of certain catalysts, detailed in the next section, allowed us to eliminate some catalysts. Regarding temperature, we chose 100 °C, which is the most common temperature in the literature, and 60 °C, which was used in the optimal conditions from the JACS 2009 paper on which the current results are benchmarked. Finally, we selected three solvents (dioxane, toluene, and dimethylformamide (each used in a 5:1 mixture with water)), two bases (sodium carbonate and potassium phosphate), two temperatures (60 °C and 100 °C), and seven catalysts (Pd SPhos G4, Pd(PPh3)4, Pd XPhos G4, Pd P(tBu)3G4, Pd PCy3G4, Pd2(dba)3, and Pd(dppf)Cl2). As a result, we considered a 3 × 2 × 2 × 7 = 84 condition space.
[0172] Other reaction variables not explicitly varied in this optimization are the molar ratio of MIDA boronate to halide, catalyst loading, molar ratio of base, concentration, ratio of organic solvent to water, reaction time, stirring rate, and metal-to-ligand ratio of the catalyst. These variables were not explored primarily due to practical considerations: each additional variable increases the total search space by approximately an order of magnitude, as the reaction condition space is multiplied by the full substrate space (11 substrate pairs). In general, the variables selected for testing (solvent, base, catalyst / ligand, and temperature) are often important in influencing the outcome of the reaction and are the most commonly varied by bench chemists when optimizing individual reactions, often important in the context of discovery chemistry. Each of these necessary components affects the fundamental bond-forming steps of the reaction in some way: 1. The solvent solubilizes the components, fills open coordination sites on the catalyst, influences speciation, stabilizes charge, etc. 2. The base influences the formation of boronic acid species, which is necessary for the subsequent transmetallation to the palladium oxidative addition complex. 3. Ligands influence the steric and electronic environment around the palladium complex, affecting the relative reaction rates of different steps in the catalytic cycle. 4. Temperature provides energy to overcome energy barriers in a process, including influencing the relative rates of productive and non-productive pathways. Variables such as reaction time, concentration, and molar equivalents of reagents and catalysts are important in the context of scale-up / process optimization to limit waste. In discovery situations, where the relative reaction rates of newly synthesized compounds are unknown a priori, it is more beneficial to use longer reaction times and slight excesses of less stable reaction partners to ensure complete conversion of starting materials, which generally maximizes the yield of the desired product.
[0173] Example 6 - Automated synthesis Training Set - General Procedure 1 (GP1) According to the general method for automated synthesis, a 40 mL I-Chem vial equipped with a stir bar was charged with halide (0.1 mmol, 1 equiv.), MIDA boronate (0.12 mmol, 1.2 equiv.), Pd(PPh3)4 (5.8 mg, 0.005 mmol, 5 mol%), Na2CO3 (80 mg, 0.75 mmol, 7.5 equiv.), and phenanthrene (internal standard, 17.8 mg, 0.1 mmol, 1 equiv.). The vial(s) were loaded into the synthesizer, and the automated synthesis procedure was run. The synthesizer then ran 10 cycles of the automated Schlenk line procedure, adding solvent (8 mL of 5:1 dioxane:water), heating (100 °C), and stirring (300 rpm) for a 12-hour reaction time, after which the hotplate temperature was reduced (20 °C) and stirring was stopped (0 rpm) to terminate the reaction. The reactions were then purified by preparative HPLC for molecular characterization and generation of response factor curves.
[0174] [ka] 5-(furan-3-yl)-1H-indole. Following GP1, furan-3-MIDA boronate (27 mg, 0.12 mmol) and 5-bromoindole (20 mg, 0.1 mmol) were used to give 1 (80%) as a beige solid. 1 H-NMR (500 MHz, acetone-d6) δ 10.26 (br s, 1H), 7.93 (m, 1H), 7.81 (m, 1H), 7.61 (t, J = 1.7 Hz, 1H), 7.45 (d, J = 8.4 Hz, 1H), 7.38 (dd, J = 8.4, 1.6 Hz, 1H), 7.35 (t, J = 2.8 Hz, 1H), 6.90 (dd, J = 1.8, 0.8 Hz, 1H), 6.49 (m, 1H); 13 C-NMR (125 MHz, acetone-d6) δ 143.5, 137.7, 135.7, 128.6, 127.8, 125.3, 123.6, 119.9, 117.3, 111.6, 109.0, 101.6; HRMS (ESI+) C 12 H10 Calculated value for NO [M+H] + m / z 184.0757, observed value 184.0754.
[0175] [ka] 3,6-Dimethoxy-4-(2,3,4,5,6-pentamethylphenyl)pyridazine. According to GP1, 3,6-dimethoxypyridazine-4-MIDA boronate (35 mg, 0.12 mmol) and 1-bromo-2,3,4,5,6-pentamethylbenzene (23 mg, 0.1 mmol) were used to obtain 2 (84%) as a colorless crystalline solid. 1 H-NMR (500 MHz, acetone-d6) δ 6.77 (s, 1H), 4.05 (s, 3H), 3.93 (s, 3H), 2.26 (s, 3H), 2.22 (s, 6H), 1.89 (s, 6H); 13 C-NMR (125 MHz, acetone-d6) δ 162.3, 160.4, 135.8, 134.9, 132.2, 131.1, 130.6, 120.7, 53.8, 53.6, 17.1, 15.9, 15.6; HRMS (ESI+) C 17 H 23 Calculated value for N2O2 [M+H] + m / z 287.1754, observed value 287.1746.
[0176] [ka] Ethyl 2-(1-(tetrahydro-2H-pyran-2-yl)-1H-pyrazol-3-yl)-4,5,6,7-tetrahydrobenzo[d]thiazole-4-carboxylate. According to GP1, 1-(tetrahydro-2H-pyran-2-yl)-3-pyrazole MIDA boronate (37 mg, 0.12 mmol) and 2-bromo-4,5,6,7-tetrahydrobenzothiazole-4-carboxylic acid ethyl ester (29 mg, 0.1 mmol) were used to obtain 3 (31%) as a colorless oil. 1H-NMR (500 MHz, acetone-d6) δ 7.52 (d, J = 1.8 Hz, 1H), 6.71 (d, J = 1.8 Hz, 1H), 6.34 (dd, J = 10.1, 2.3 Hz, 1H), 4.25 - 4.15 (m, 2H), 3.92 - 3.86 (m, 2H), 3.62 (td, J = 11.3, 3.1 Hz, 1H), 2.97 - 2.86 (m, 2H), 2.49 (m, 1H), 2.19 (m, 1H), 2.15 - 2.08 (m, 3H), 1.99 - 1.86 (m, 2H), 1.75 - 1.53 (m, 3H), 1.28 (t, J = 7.1 Hz, 3H); 13 C-NMR (125 MHz, acetone-d6) δ 172.8, 153.4, 148.2, 138.6, 135.9, 132.3, 108.0, 84.7, 67.1, 60.3, 43.4, 28.9, 26.7, 25.0, 22.9, 22.7, 21.0, 13.7; HRMS (ESI+) C 18 H 24 Calculated value for N3O3S [M+H] + m / z 362.1533, observed value 362.1523.
[0177] [ka] 5-(Phenantren-9-yl)-1H-pyrrole-2-carbaldehyde. Following GP1, 9-phenanthrenyl MIDA boronate (40 mg, 0.12 mmol) and 5-bromo-1H-pyrrole-2-carbaldehyde (17 mg, 0.1 mmol) were used to give 4 (88%) as an off-white solid. 1H-NMR (500 MHz, acetone-d6) δ 11.49 (br s, 1H), 9.68 (s, 1H), 8.94 (d, J = 8.3 Hz, 1H), 8.87 (d, J = 8.3 Hz, 1H), 8.26 (d, J = 8.2 Hz, 1H), 8.04 (d, J = 5.8 Hz, 2H), 7.79 - 7.74 (m, 2H), 7.72 - 7.67 (m, 2H), 7.24 (d, J = 3.7 Hz, 1H), 6.70 (d, J = 3.7 Hz, 1H); 13 C-NMR (125 MHz, acetone-d6) δ 178.5, 138.3, 133.9, 131.3, 130.7, 130.3, 130.3, 129.0, 128.6, 128.5, 127.5, 127.2, 127.0, 127.0, 126.2, 123.2, 122.7, 120.7, 112.4; HRMS (ESI+) C 19 H 14 Calculated value for NO [M+H] + m / z 272.1070, observed value 272.1071.
[0178] [ka] 5-Methyl-2-(phenanthrene-9-yl)pyridine. Following GP1, 9-phenanthrenyl MIDA boronate (40 mg, 0.12 mmol) and 2-bromo-5-methylpyridine (17 mg, 0.1 mmol) were used to obtain 5 (43%) as a pale yellow oil. 1H-NMR (500 MHz, acetone-d6) δ 8.93 (dd, J = 8.3, 0.5 Hz, 1H), 8.88 (d, J = 8.3 Hz, 1H), 8.65 (m, 1H), 8.19 (dd, J = 8.3, 0.8 Hz, 1H), 8.05 (dd, J = 7.8, 0.8 Hz, 1H), 7.91 (s, 1H), 7.81 (dd, J = 7.9, 2.0 Hz, 1H), 7.77 - 7.71 (m, 2H), 7.69 (m, 1H), 7.65 - 7.59 (m, 2H), 2.48 (s, 3H); 13 C-NMR (125 MHz, acetone-d6) δ 156.4, 149.7, 137.5, 137.0, 131.8, 131.5, 130.7, 130.5, 130.3, 128.9, 128.0, 127.1, 127.0, 126.9, 126.6, 126.5, 124.3, 123.0, 122.6, 17.3; HRMS (ESI+) C 20 H 16 Calculated value for N [M+H] + m / z 270.1277, observed value 270.1281.
[0179] [ka] 2-(Phenantren-9-yl)-1H-benzo[d]imidazole. Following GP1, 9-phenanthrenyl MIDA boronate (40 mg, 0.12 mmol) and 2-bromo-1H-benzimidazole (20 mg, 0.1 mmol) were used to give 6 (87%) as an off-white solid. 1H-NMR (500 MHz, DMSO d6) δ 13.04 (s, 1H), 9.11 (dd, J = 8.2, 1.2 Hz, 1H), 8.98 (d, J = 8.0 Hz, 1H), 8.93 (d, J = 8.3 Hz, 1H), 8.38 (s, 1H), 8.13 (d, J = 7.0 Hz, 1H), 7.83 - 7.78 (m, 3H), 7.78 - 7.73 (m, 2H), 7.61 (d, J = 7.4 Hz, 1H), 7.33 - 7.25 (m, 2H); 13 C-NMR (125 MHz, DMSO d6) δ 151.7, 144.3, 134.9, 130.9, 130.7, 130.7, 129.8, 129.6, 129.6, 128.6, 127.9, 127.7, 127.7, 127.6, 127.0, 123.7, 123.5, 123.2, 122.1, 119.6, 111.8;HRMS (ESI+) C 21 H 15 Calculated value for N2 [M+H] + m / z 295.1230, observed value 295.1229.
[0180] [ka] 4-(Benzyloxy)-5-(4-(benzyloxy)phenyl)-N,N-dimethylpyrimidin-2-amine. According to GP1, using 4-(benzyloxy)phenyl MIDA boronate (41 mg, 0.12 mmol) and 4-benzyloxy-5-bromo-2-(N,N-dimethylamino)pyrimidine (31 mg, 0.1 mmol), 7 (88%) was obtained as a beige solid. 1 H-NMR (500 MHz, acetone-d6) δ 8.17 (s, 1H), 7.52 - 7.47 (m, 6H), 7.43 - 7.30 (m, 6H), 7.05 (dd, J = 8.7, 1.8 Hz, 2H), 5.51 (s, 2H), 5.15 (s, 2H), 3.19 (s, 6H); 13C-NMR (125 MHz, acetone-d6) δ 165.6, 161.2, 157.8, 157.4, 137.6, 137.6, 129.5, 128.4, 128.3, 127.7, 127.7, 127.6, 127.5, 127.3, 114.6, 109.3, 69.5, 67.0, 36.2; HRMS (ESI+) C 26 H 26 Calculated value for N3O2 [M+H] + m / z 412.2020, measured value 412.2015.
[0181] [ka] 2-(2-(2-(benzyloxy)ethoxy)phenyl)benzo[b]thiophene. Following GP1, benzothiophene-2-MIDA boronate (35 mg, 0.12 mmol) and 1-(2-(benzyloxy)ethoxy)-2-bromobenzene (31 mg, 0.1 mmol) were used to obtain 8 (65%) as a light brown oil. 1 H-NMR (500 MHz, acetone-d6) δ 8.03 (s, 1H), 7.90 (m, 1H), 7.79 (dd, J = 7.7, 1.6 Hz, 1H), 7.74 (m, 1H), 7.41 (d, J = 7.1 Hz, 2H), 7.38 - 7.26 (m, 6H), 7.21 (d, J = 8.3 Hz, 1H), 7.08 (td, J = 7.5, 1.3 Hz, 1H), 4.70 (s, 2H), 4.40 (t, J = 4.7 Hz, 2H), 4.01 (t, J = 4.7 Hz, 2H); 13 C-NMR (125 MHz, acetone-d6) δ 155.8, 140.4, 140.0, 139.6, 138.8, 129.4, 129.1, 128.2, 127.5, 127.3, 124.2, 124.2, 123.5, 123.0, 122.8, 121.7, 121.1, 113.1, 72.8, 68.7, 68.2; HRMS (ESI+) C 23 H21 Calculated value for O2S [M+H] + m / z 361.1257, observed value 361.1250.
[0182] [ka] 5-(3,6-Dimethoxypyridazin-4-yl)-4,6-dimethylpyrimidin-2-amine. According to GP1, 3,6-dimethoxypyridazine-4-MIDA boronate (35 mg, 0.12 mmol) and 5-bromo-4,6-dimethylpyrimidin-2-amine (20 mg, 0.1 mmol) were used to obtain 9 (24%) as an off-white solid. 1 H-NMR (500 MHz, DMSO d6) δ 7.20 (s, 1H), 6.65 (s, 2H), 3.98 (s, 3H), 3.93 (s, 3H), 1.96 (s, 6H); 13 C-NMR (125 MHz, DMSO d6) δ 164.8, 163.0, 162.3, 160.2, 131.5, 122.2, 114.4, 54.9, 54.6, 22.6;HRMS (ESI+) C 12 H 16 Calculated value for N5O2 [M+H] + m / z 262.1299, observed value 262.1295.
[0183] [ka] [3,3'-Bithiophene]-5-carbonitrile. Following GP1, 2-thiophene MIDA boronate (29 mg, 0.12 mmol) and 4-bromothiophene-2-carbonitrile (19 mg, 0.1 mmol) were used to give 10 (71%) as an off-white solid. 1H-NMR (500 MHz, acetone-d6) δ 8.24 (t, J = 1.2 Hz, 1H), 8.13 (t, J = 1.2 Hz, 1H), 7.86 (m, 1H), 7.60 (dd, J = 5.0, 2.9 Hz, 1H), 7.57 (m, 1H); 13 C-NMR (125 MHz, acetone-d₆) δ 138.0, 136.8, 135.2, 127.2, 127.0, 126.1, 121.8, 113.8, 109.9; HRMS (EI+) calculated for C₈H₅NS₂ [M] + m / z 190.9863, observed value 190.9864.
[0184] [ka] tert-Butyl 4-(4-(1-(tert-butoxycarbonyl)-1H-pyrrol-2-yl)-1H-pyrazol-1-yl)piperidine-1-carboxylate. According to GP1, using N-Boc-pyrrole-2-MIDA boronate (39 mg, 0.12 mmol) and 1-(4-Boc-piperidino)-4-bromopyrazole (33 mg, 0.1 mmol), 11 (47%) was obtained as a colorless oil. 1 H-NMR (500 MHz, acetone-d6) δ 7.82 (s, 1H), 7.55 (s, 1H), 7.30 (dd, J = 3.3, 1.9 Hz, 1H), 6.20 (m, 2H), 4.40 (m, 1H), 4.21 (d, J = 10.7 Hz, 2H), 2.96 (s, 2H), 2.09 (m, 2H), 1.94 (m, 2H), 1.51 (s, 9H), 1.48 (s, 9H); 13 C-NMR (125 MHz, acetone-d6) δ 154.1, 149.1, 138.9, 127.1, 126.9, 121.9, 113.8, 113.3, 110.4, 83.2, 78.7, 58.7, 42.6, 32.3, 27.7, 27.1; HRMS (ESI+) C 22 H 33Calculated for N4O4 [M+H] + m / z 417.2496, observed value 417.2483.
[0185] Test Set - General Procedure 2 (GP2) Following the general method for automated synthesis (Section 1), a 40 mL I-Chem vial equipped with a stir bar was charged with halide (0.1 mmol, 1 equiv.), MIDA boronate (0.12 mmol, 1.2 equiv.), Pd XPhos G4 (4.3 mg, 0.005 mmol, 5 mol%), NaCO3 (80 mg, 0.75 mmol, 7.5 equiv.), and phenanthrene (internal standard, 17.8 mg, 0.1 mmol, 1 equiv.). The vial(s) were loaded into the synthesizer and the automated synthesis procedure was run. The synthesizer then ran 10 cycles of the automated Schlenk line procedure, adding solvent (8 mL of 5:1 dioxane:water), heating (100 °C), and stirring (300 rpm) for a 12-hour reaction time, after which the reaction was stopped by cooling the hotplate (20 °C) and stopping the stirring (0 rpm). The reactions were then purified by preparative HPLC for molecular characterization and generation of response factor curves.
[0186] [ka] 3'-Methyl-1,1'-bis(tetrahydro-2H-pyran-2-yl)-1H,1'H-[3,4'-bipyrazole]-5'-carbaldehyde. According to GP2, 1-(tetrahydro-2H-pyran-2-yl)-3-pyrazole MIDA boronate (37 mg, 0.12 mmol) and 4-bromo-5-methyl-2-(oxan-2-yl)pyrazole-3-carbaldehyde (27 mg, 0.1 mmol) were used to obtain 12 as a colorless oil. 1H-NMR (500 MHz, acetone-d6) δ 9.55 (s, 1br H), 7.62 (s, 1H), 6.43 (s, 1H), 6.17 - 6.11 (m, 1H), 5.07 (s, 1H), 4.00 - 3.97 (m, 1H), 3.84 (d, J = 8.4 Hz, 1H), 3.75 - 3.70 (m, 1H), 3.45 - 3.31 (m, 1H), 2.49 - 2.36 (m, 2H), 2.14 (s, 3H), 2.12 - 2.08 (m, 1H), 2.04 - 1.96 (m, 2H), 1.89 - 1.83 (m, 1H), 1.80 - 1.72 (m, 1H), 1.65 - 1.47 (m, 5H); 13 C-NMR (125 MHz, acetone-d6) δ 180.1, 138.8, 137.3, 132.4, 109.2, 85.4, 85.3, 67.4, 67.4, 66.8, 24.9, 24.9, 24.8, 24.8, 22.5, 22.5, 22.3, 11.0; HRMS (ESI+) C 18 H 24 Calculated value for N4NaO3 [M+Na] + m / z 367.1741, observed value 367.1737.
[0187] [ka] 5-(Thiophen-3-yl)-1H-pyrazole. Following GP2, 3-thiophene MIDA boronate (29 mg, 0.12 mmol) and 5-bromo-1H-pyrazole (15 mg, 0.1 mmol) were used to give 13 as an off-white solid. 1 H-NMR (500 MHz, acetone-d6) δ 12.16 (s, 1H), 7.73 (dd, J = 2.8, 1.0 Hz, 1H), 7.68 (d, J = 1.7 Hz, 1H), 7.56 (dd, J = 5.0, 1.1 Hz, 1H), 7.51 (dd, J = 5.0, 2.9 Hz, 1H), 6.61 (d, J = 2.2 Hz, 1H);13 C-NMR (125 MHz, CDCl) δ 145.3, 133.5, 132.8, 126.3, 125.9, 121.1, 103.0; HRMS (ESI+) calculated for C7H7N2S [M+H] + m / z 151.0324, observed value 151.0326.
[0188] [ka] 4-(3-Nitrophenyl)-1,3-dihydro-2H-benzo[d]imidazol-2-one. According to GP2, 3-nitrophenyl MIDA boronate (33 mg, 0.12 mmol) and 4-bromo-1,3-dihydro-2H-benzimidazol-2-one (21 mg, 0.1 mmol) were used to give 14 as a yellow solid. 1 H-NMR (500 MHz, DMSO d6) δ 10.90 (s, 2H), 8.29 (s, 1H), 8.21 (d, J = 8.0 Hz, 1H), 7.98 (d, J = 7.5 Hz, 1H), 7.75 (t, J = 7.9 Hz, 1H), 7.08 (d, J = 4.1 Hz, 2H), 7.01 (m, 1H); 13 C-NMR (125 MHz, DMSO d6) δ 156.1, 148.6, 139.2, 135.2, 130.9, 130.8, 127.7, 123.5, 122.5, 121.7, 121.1, 120.8, 109.1;HRMS (ESI+)C 13 H 10 Calculated for N3O3 [M+H] + m / z 256.0717, observed value 256.0711.
[0189] [ka] 4-(2,3-Difluorophenyl)-1,3-dihydro-2H-benzo[d]imidazol-2-one. Following GP2, (2,3-difluorophenyl)MIDA boronate (32 mg, 0.12 mmol) and 4-bromo-1,3-dihydro-2h-benzimidazol-2-one (21 mg, 0.1 mmol) were used to give 15 as an off-white solid. 1 H-NMR (500 MHz, acetone-d6) δ 10.08 (s, 2H), 7.39–7.33 (m, 1H), 7.33–7.28 (m, 2H), 7.16–7.11 (m, 2H), 7.07–7.03 (m, 1H). 13 C-NMR (125 MHz, acetone-d6) δ 155.5, 151.0 (dd, J = 246.2, 13.0 Hz), 148.0 (dd, J = 247.5, 13.1 Hz), 130.1, 128.3, 127.6 (d, J = 12.3 Hz), 126.4 - 126.1 (m), 124.8 (dd, J = 7.4, 4.7 Hz), 122.0 (d, J = 0.7 Hz), 121.1, 116.5 (d, J = 17.4 Hz), 116.2 (d, J = 2.7 Hz), 108.9; 19 F-NMR (471 MHz, acetone-d6) δ -139.7 (m, 1F), -140.9 (m, 1F);HRMS (ESI+) C 13 Calculated value for H9F2N2O [M+H] + m / z 247.0677, observed value 247.0678.
[0190] [ka] 3-(Pyridin-4-yl)pyridazine. Following GP2, 4-pyridine MIDA boronate (28 mg, 0.12 mmol) and 3-bromopyridazine (16 mg, 0.1 mmol) were used to give 16 as a white solid. 1H-NMR (500 MHz, acetone-d6) δ 9.31 (dd, J = 4.9, 1.5 Hz, 1H), 8.79 (dd, J = 4.5, 1.7 Hz, 2H), 8.31 (dd, J = 8.6, 1.5 Hz, 1H), 8.15 (dd, J = 4.5, 1.7 Hz, 2H), 7.86 (dd, J = 8.6, 5.0 Hz, 1H); 13 C-NMR (125 MHz, acetone-d₆) δ 157.0, 151.4, 150.6, 143.8, 127.4, 124.3, 120.9; HRMS (ESI+) calculated for C₈H₈N₃ [M+H] + m / z 158.0713, observed value 158.0717.
[0191] [ka] 6-(4-(Trifluoromethoxy)phenyl)-1H-indole-3-carboxylic acid. Following GP2, 4-(trifluoromethoxy)phenyl MIDA boronate (38 mg, 0.12 mmol) and 6-bromo-1H-indole-3-carboxylic acid (24 mg, 0.1 mmol) were used to give 17 as a beige solid. 1 H-NMR (500 MHz, CD3OD) δ 8.28 (d, J = 8.4 Hz, 1H), 7.81 (s, 1H), 7.73 (d, J = 8.6 Hz, 2H), 7.60 (d, J = 1.1 Hz, 1H), 7.38 (dd, J = 8.3, 1.6 Hz, 1H), 7.31 (d, J = 8.4 Hz, 2H); 13 C-NMR (125 MHz, CD3OD) δ 173.3, 147.9 (q, J = 1.9 Hz), 141.6, 137.3, 133.2, 130.4, 128.2, 126.8, 121.8, 120.9, 120.7 (q, J = 255.2 Hz, CF3), 119.2, 114.2, 109.3; 19F-NMR (471 MHz, CD3OD) δ -59.5;HRMS (ESI+) C 16 H 11 Calculated value for F3NO3 [M+H] + m / z 322.0686, observed value 322.0680.
[0192] [ka] tert-Butyl 4-(5-(1-(tert-butoxycarbonyl)-1H-pyrrol-2-yl)-7H-pyrrolo[2,3-d]pyrimidin-4-yl)piperazine-1-carboxylate. Following GP2, N-Boc-pyrrole-2-MIDA boronate (39 mg, 0.12 mmol) and 4-(4-Boc-1-piperazinyl)-5-bromo-7H-pyrrolo[2,3-d]pyrimidine (38 mg, 0.1 mmol) were used to give 18 as an off-white solid. 1 H-NMR (500 MHz, acetone-d6) δ 11.05 (s, 1H), 8.37 (s, 1H), 7.46 (dd, J = 3.3, 1.9 Hz, 1H), 7.34 (s, 1H), 6.32 (t, J = 3.3 Hz, 1H), 6.28 (dd, J = 3.2, 1.9 Hz, 1H), 3.30 (s, 4H), 3.18 (s, 4H), 1.44 (s, 9H), 1.04 (s, 9H); 13 C-NMR (125 MHz, acetone-d6) δ 160.5, 154.0, 153.1, 150.6, 149.4, 128.6, 122.6, 121.8, 113.8, 110.4, 108.7, 106.5, 82.7, 78.9, 50.1, 49.2, 27.6, 26.3; HRMS (ESI+) C 24 H 33 Calculated for N6O4 [M+H] + m / z 469.2558, observed value 469.2557.
[0193] [ka] [2,3′-Bithiophene]-5,5′-dicarbonitrile. Following GP2, 5-cyanothiophene-2-MIDA boronate (32 mg, 0.12 mmol) and 4-bromothiophene-2-carbonitrile (19 mg, 0.1 mmol) were used to give 19 as an off-white solid. 1 H-NMR (500 MHz, acetone-d6) δ 8.35 (d, J = 1.5 Hz, 1H), 8.29 (d, J = 1.5 Hz, 1H), 7.88 (d, J = 4.0 Hz, 1H), 7.65 (d, J = 4.0 Hz, 1H); 13 C-NMR (125 MHz, acetone-d6) δ 143.6, 139.2, 136.3, 134.0, 129.8, 125.7, 113.5, 113.2, 111.3, 108.2; HRMS (EI+) C 10 Calculated value for H4N2S2 [M] + m / z 215.9816, observed value 215.9821.
[0194] [ka] 2-(3-Nitrophenyl)furan-3-carboxylic acid. Following GP2, 3-nitrophenyl MIDA boronate (33 mg, 0.12 mmol) and 2-bromofuran-3-carboxylic acid (19 mg, 0.1 mmol) were used to give 20 as a yellow solid. 1 H-NMR (500 MHz, CD3OD) δ 8.96 (t, J = 1.9 Hz, 1H), 8.45 (m, 1H), 8.15 (m, 1H), 7.62 (t, J = 8.1 Hz, 1H), 7.54 (d, J = 1.8 Hz, 1H), 6.78 (d, J = 1.8 Hz, 1H); 13 C-NMR (125 MHz, CD3OD) δ 170.6, 149.2, 148.3, 141.4, 132.6, 132.4, 128.9, 123.4, 121.6, 121.2, 113.8;HRMS (ESI+) C 11Calculated value for H6NO5 [MH] + m / z 232.0251, observed value 232.0241.
[0195] [ka] 4-(Phenantren-9-yl)isoquinoline. Following GP2, 9-phenanthrenyl MIDA boronate (40 mg, 0.12 mmol) and 4-bromoisoquinoline (21 mg, 0.1 mmol) were used to give 21 as an off-white solid. 1 H-NMR (500 MHz, CDCl3) δ 9.43 (s, 1H), 8.84 (dd, J = 13.9, 8.3 Hz, 2H), 8.67 (s, 1H), 8.14 (d, J = 8.2 Hz, 1H), 7.95 (dd, J = 7.8, 0.7 Hz, 1H), 7.83 (s, 1H), 7.77 (m, 1H), 7.72 - 7.64 (m, 3H), 7.56 (m, 1H), 7.49 (d, J = 8.4 Hz, 1H), 7.46 - 7.43 (m, 2H); 13 C-NMR (125 MHz, CDCl3) δ 152.5, 143.7, 135.7, 133.2, 131.9, 131.8, 131.4, 130.7, 130.5, 130.4, 129.2, 128.8, 128.2, 127.9, 127.4, 127.2, 127.1, 127.1, 126.8, 125.5, 123.0, 122.7;HRMS (ESI+) C 23 H 16 Calculated value for N [M+H] + m / z 306.1277, observed value 306.1272.
[0196] [ka] 5-(5-Fluoropyrazin-2-yl)thiazole. Following GP2, 5-thiazole MIDA boronate (29 mg, 0.12 mmol) and 2-bromo-5-fluoropyrazine (18 mg, 0.1 mmol) were used to give 22 as an off-white solid. 1 H-NMR (500 MHz, acetone-d6) δ 9.15 (s, 1H), 8.93 (t, J = 1.4 Hz, 1H), 8.64 (s, 1H), 8.60 (dd, J = 8.3, 1.3 Hz, 1H); 13 C-NMR (125 MHz, acetone-d6) δ 159.5 (d, J = 250.6 Hz), 155.8, 144.7 (d, J = 4.6 Hz), 141.6 (d, J = 1.4 Hz), 138.1 (d, J = 10.1 Hz), 135.9 (d, J = 1.8 Hz), 132.9 (d, J = 38.9 Hz); 19 F-NMR (471 MHz, acetone-d) δ -84.57 (d, J = 7.8 Hz); HRMS (ESI+) calculated for C7H5FN3S [M+H] + m / z 182.0183, observed value 182.0184.
[0197] [ka] 4-(3,6-Dimethoxypyridazin-4-yl)-2,3,5,6-tetramethylaniline. Following GP2, 3,6-dimethoxypyridazine-4-MIDA boronate (35 mg, 0.12 mmol) and 4-bromo-2,3,5,6-tetramethylaniline (23 mg, 0.1 mmol) were used to obtain 23 as a white solid. 1 H-NMR (500 MHz, acetone-d6) δ 6.74 (s, 1H), 4.30 (s, 2H), 4.03 (s, 3H), 3.91 (s, 3H), 2.11 (s, 6H), 1.87 (s, 6H); 13C-NMR (125 MHz, acetone-d6) δ 162.3, 161.0, 144.0, 136.4, 130.7, 122.9, 121.2, 117.3, 53.7, 53.6, 16.9, 12.7; HRMS (ESI+) C 16 H 22 Calculated value for N3O2 [M+H] + m / z 288.1707, observed value 288.1702.
[0198] [ka] Ethyl 2-(thiazol-5-yl)-4,5,6,7-tetrahydrobenzo[d]thiazole-4-carboxylate. Following GP2, 5-thiazole MIDA boronate (29 mg, 0.12 mmol) and 2-bromo-4,5,6,7-tetrahydro-benzothiazole-4-carboxylic acid ethyl ester (29 mg, 0.1 mmol) were used to give 24 as an off-white solid. 1 H-NMR (500 MHz, CDCl3) δ 8.81 (s, 1H), 8.22 (s, 1H), 4.29 - 4.19 (m, 2H), 3.92 (t, J = 5.8 Hz, 1H), 2.92 (m, 1H), 2.83 (m, 1H), 2.22 (m, 1H), 2.15 - 2.02 (m, 2H), 1.91 (m, 1H), 1.32 (t, J = 7.1 Hz, 3H); 13 C-NMR (125 MHz, CDCl3) δ 173.3, 155.1, 153.7, 148.0, 141.5, 133.5, 132.2, 61.0, 43.1, 26.7, 23.4, 20.9, 14.3;HRMS (ESI+) C 13 H 15 Calculated value for N2O2S2 [M+H] + m / z 295.0569, observed value 295.0563.
[0199] [ka] 6-Fluoro-1-methyl-5-(2,4,6-trifluorophenyl)-1H-benzo[d][1,2,3]triazole. Following GP2, 2,4,6-trifluorophenyl MIDA boronate (34 mg, 0.12 mmol) and 5-bromo-6-fluoro-1-methyl-1,2,3-benzotriazole (23 mg, 0.1 mmol) were used to give 25 as a beige solid. 1 H-NMR (600 MHz, acetone-d6) δ 8.03 (dd, J = 6.1, 0.3 Hz, 1H), 7.64 (dd, J = 9.0, 0.3 Hz, 1H), 7.05 - 6.99 (m, 2H), 4.26 (s, 3H); 13 C{ 1 H}-NMR (150 MHz, acetone-d6) δ 163.8 (t, J = 15.7 Hz), 162.1 (t, J = 15.7 Hz), 161.5 (dd, J = 15.7, 9.5 Hz), 159.9 (dd, J = 15.7, 9.4 Hz), 158.3, 142.4, 134.3 (d, J = 14.2 Hz), 123.1 (d, J = 4.5 Hz), 113.7 (d, J = 22.0 Hz), 109.0 (td, J = 21.0, 4.8 Hz), 100.7 - 100.3 (m), 96.4 (d, J = 29.3 Hz), 33.9; 13 C{ 19 F}-NMR (150 MHz, acetone-d6) δ 162.9 (t, J C-H = 6.3 Hz), 160.7 (t, J C-H = 3.5 Hz), 159.1 (dd, J C-H = 10.1, 5.9 Hz), 142.4 (d, J C-H = 5.2 Hz), 134.3 (dt, J C-H = 7.7, 2.1 Hz), 123.7 (d, J C-H = 0.9 Hz), 122.6 (d, J C-H = 0.9 Hz), 113.7 (d, J C-H= 4.8 Hz), 109.0 (q, J C-H = 4.9 Hz), 101.1 (d, J C-H = 4.2 Hz), 100.0 (d, J C-H = 4.2 Hz), 97.0 (d, J C-H = 1.3 Hz), 95.9 (d, J C-H = 1.3 Hz), 33.9 (q, J C-H = 141.8 Hz); 19 F-NMR (471 MHz, acetone-d6) δ -108.6 (m, 1F), -110.9 (m, 2F), -115.3 (m, 1F);HRMS (ESI+) C 13 Calculated value for H8F4N3 [M+H] + m / z 282.0649, observed value 282.0643.
[0200] [ka] Ethyl 7-(benzo[c][1,2,5]oxadiazol-5-yl)-1H-benzo[d][1,2,3]triazole-5-carboxylate. Following GP2, 5-benzofurazan MIDA boronate (33 mg, 0.12 mmol) and ethyl 7-bromo-1H-1,2,3-benzotriazole-5-carboxylate (27 mg, 0.1 mmol) were used to give 26 as a pale yellow solid. 1 H-NMR (500 MHz, acetone-d6) δ 8.83 (s, 1H), 8.68 (s, 1H), 8.39 (d, J = 5.5 Hz, 1H), 8.35 (d, J = 9.3 Hz, 1H), 8.15 (m, 1H), 4.47 (qd, J = 7.1, 1.8 Hz, 2H), 1.44 (t, J = 7.1 Hz, 3H); 13C-NMR (125 MHz, acetone-d6) δ 165.4, 149.7, 148.8, 142.0, 140.1, 138.6, 133.6, 128.5, 127.4, 124.1, 117.2, 116.5, 115.6, 61.2, 13.7; HRMS (ESI+) C 15 H 12 Calculated value for N5O3 [M+H] + m / z 310.0935, observed value 310.0927.
[0201] [ka] 6-(Thiazol-5-yl)-[1,2,4]triazolo[1,5-a]pyrazin-2-amine. Following GP2, 5-thiazole MIDA boronate (29 mg, 0.12 mmol) and 6-bromo-[1,2,4]triazolo[1,5-a]pyrazin-2-amine (21 mg, 0.1 mmol) were used to give 27 as a white solid. 1 H-NMR (500 MHz, DMSO d6) δ 9.46 (d, J = 1.0 Hz, 1H), 9.12 (s, 1H), 8.85 (d, J = 1.0 Hz, 1H), 8.56 (s, 1H), 6.61 (s, 2H); 13 C-NMR (125 MHz, DMSO d₆) δ 167.7, 155.3, 146.1, 139.7, 137.5, 137.2, 133.2, 117.9; HRMS (ESI+) calculated for C₈H₇N₆S [M+H] + m / z 219.0447, observed value 219.0452.
[0202] [ka] Methyl 4-(3,6-dimethoxypyridazin-4-yl)-5-methyl-1H-pyrazole-3-carboxylate. Following GP2, 3,6-dimethoxypyridazine-4-MIDA boronate (35 mg, 0.12 mmol) and methyl 4-bromo-5-methyl-1H-pyrazole-3-carboxylate (22 mg, 0.1 mmol) were used to obtain 28 as a white solid. 1 H-NMR (500 MHz, CD3OD) δ 7.05 (s, 1H), 4.04 (s, 3H), 3.95 (s, 3H), 3.77 (s, 3H), 2.25 (s, 3H); 13 C-NMR (125 MHz, CD3OD) δ 162.2, 160.1, 127.6, 121.1, 53.7, 53.7, 50.9, 8.5;* HRMS (ESI+) C 12 H 15 Calculated for N4O4 [M+H] + m / z 279.1088, observed value 279.1084. *Probably due to tautomerism of pyrazole. 13 The C resonance is missing.
[0203] [ka] 5-(4-Amino-2,3,5,6-tetramethylphenyl)thiophene-2-carbonitrile. Following GP2, 5-cyanothiophene-2-MIDA boronate (32 mg, 0.12 mmol) and 4-bromo-2,3,5,6-tetramethylaniline (23 mg, 0.1 mmol) were used to give 29 as an off-white solid. 1 H-NMR (500 MHz, acetone-d6) δ 7.85 (d, J = 3.7 Hz, 1H), 6.90 (d, J = 3.7 Hz, 1H), 4.46 (s, 2H), 2.12 (s, 6H), 2.00 (s, 6H); 13C-NMR (125 MHz, acetone-d6) δ 153.6, 145.0, 138.3, 133.1, 128.3, 119.7, 117.3, 113.9, 108.6, 17.3, 12.8; HRMS (ESI+) C 15 H 17 Calculated value for N2S [M+H] + m / z 257.1107, observed value 257.1100.
[0204] [ka] 2-(Pyrrolidin-1-yl)-5-(1-(tetrahydro-2H-pyran-2-yl)-1H-pyrazol-3-yl)thiazole. According to GP2, 1-(tetrahydro-2H-pyran-2-yl)-3-pyrazole MIDA boronate (37 mg, 0.12 mmol) and 5-bromo-2-pyrrolidinothiazole (23 mg, 0.1 mmol) were used to obtain 30 as a pale yellow oil. 1 H-NMR (500 MHz, CDCl3) δ 7.55 (d, J = 1.8 Hz, 1H), 7.33 (s, 1H), 6.31 (d, J = 1.8 Hz, 1H), 5.34 (dd, J = 10.3, 2.4 Hz, 1H), 4.10 (m, 1H), 3.70 (td, J = 11.6, 2.4 Hz, 1H), 3.54 - 3.50 (m, 4H), 2.58 (m, 1H), 2.11 - 2.08 (m, 4H), 1.95 (m, 1H), 1.76 (m, 2H), 1.68 - 1.57 (m, 2H); 13 C-NMR (125 MHz, CDCl3) δ 168.7, 140.0, 139.4, 135.2, 112.0, 107.1, 84.4, 67.6, 49.5, 29.3, 25.7, 24.9, 23.0;HRMS (ESI+) C 15 H 21 Calculated value for N4OS [M+H] + m / z 305.1431, observed value 305.1429.
[0205] Closed-Loop Experiments - General Procedure 3 (GP3) Following the general method for automated synthesis (Section 1), a 40 mL I-Chem vial equipped with a stir bar was charged with halide (0.1 mmol, 1 equiv.), MIDA boronate (0.12 mmol, 1.2 equiv.), catalyst (0.005 mmol, 5 mol%), base (0.75 mmol, 7.5 equiv.), and phenanthrene (internal standard, 17.8 mg, 0.1 mmol, 1 equiv.). The catalyst, base, solvent, and temperature for each substrate were selected using the ML model. The vial(s) were loaded into the synthesizer, and the automated synthesis procedure was run. The synthesizer then ran 10 cycles of the automated Schlenk line procedure, adding solvent (8 mL of solvent), heating, and stirring (300 rpm) for a 12-hour reaction time, after which the hotplate temperature was reduced (20 °C) and stirring was stopped (0 rpm) to terminate the reaction. The reaction was then quantified by direct LCMS injection.
[0206] Example 7 - Closed-loop optimization and post-mortem analysis Initial pool of reactions To "seed" the optimization procedure, an initial set of reactions was selected without AI guidance as described above. We decided to perform couplings of all 11 substrate pairs (representing the most common class in the literature) that define our scope in a 1,4-dioxane:water mixture as the solvent and the following base and catalyst combinations: i) sodium carbonate and Pd(PPh3)4 at 100 °C (the most common base, most common catalyst, and most common temperature in the literature), and ii) potassium phosphate (representative of the second most common base in the literature and used in the benchmark conditions) in combination with the following catalysts: PdSPhos G4, PdXphos G4, PdP(tBu)3G4, PdPCy3G4, Pd2(dba)3, and Pd(dppf)Cl2, each set at 60 °C (used in the benchmark conditions). With these selections, our robotic system was tasked with performing 11 × 7 reactions (each run in duplicate). The average yields obtained in these experiments are tabulated in Figure 3E of the present application.
[0207] Remarkably, after this initial round, we observed that the yields obtained for a given substrate with different pairs of ligands were very similar. In other words, certain catalysts appear to systematically produce similar results and may therefore be redundant. To quantify such functional similarity, we calculated a Spearman rank correlation matrix (Figure 3F of the present application), showing the correlated yields obtained for all 11 substrate pairs using two different catalytic ligands. In this representation, redundant catalysts correspond to highly correlated off-diagonal elements. In this particular case, the XPhos catalyst was found to be "functionally similar" to Dppf, and PCy3 to SPhos. Note that these "functional similarities" could not be assessed simply based on structural similarities, which are largely absent. Due to this result, we decided to eliminate PCy3 and Dppf from our ligand pool to reduce redundancy. We also eliminated Pd2(dba)3 from our catalytic pool due to its poor performance (<5% yield for 8 / 11 substrates).
[0208] Selection of batches of reactions in closed-loop experiments With the model, selection strategy, and initial dataset selected, we proceeded with the actual experiments. The search space included 11 substrate pairs multiplied by 48 conditions (2 bases × 2 temperatures × 3 solvents × 4 catalysts). Navigation throughout this substrate-condition space was guided by the algorithm detailed in Section 3. One important aspect is that we conducted our studies in experimental batches, meaning that multiple experiments were performed before updating the theoretical model. Within each batch, an experiment was selected as the model's best suggestion, second-best suggestion, third-best suggestion, and so on, until the number of slots in the batch was filled. More specifically, the algorithm first formed a "priority queue" of unexplored reactions by i) sorting reaction conditions according to their calculated PI (in descending order) and ii) sorting reactions (i.e., substrate pairs) under the same conditions according to their predicted uncertainty (again in descending order). Next, we selected an arbitrary number of Nsel Using the preferred condition, we iteratively: a) extract top-N sel and b) adding the top reaction from each of the selected conditions to the proposed batch and removing it from the queue. These selection iterations continue until the batch reaches the desired batch size B. Since our experimental setup allowed us to perform 36 and later 72 reactions each week, we created a set N sel It was decided to set the number of probes to 18, i.e., each batch probed 18 different conditions and 2 or 4 pairs of substrates for each condition.
[0209] End of loop To determine when to terminate the loop, we monitored the model's prediction uncertainty for both the run ("known") and unexplored portions of the reaction-condition space. By general inference, if the average uncertainty for the unexplored space falls to a level comparable to that of the set explored prior to a given generation (the "training set") (corresponding to measurement uncertainty in GP), the model is likely to have gained sufficient knowledge of the entire space. Specifically, we monitored a) the average uncertainty for the entire reaction space (both known and unexplored), b) the average uncertainty for the "training" set, and c) the average uncertainty for the unexplored portions of the reaction space (Figures 4A and 4F). As expected, we observed a gradual decrease in the model's uncertainty, indicating an increase in knowledge of the reaction-condition space. This progress slows down after the third iteration, when the average uncertainty for the unexplored portions of the space approaches the value characterizing the training set.
[0210] At the end of the loop, the model explored a total of 329 unique reactions: 77 duplicates from seed round 1, 36 duplicates from round 2, 72 from round 3, 72 from round 4, and 84 from round 5. Of these experiments, 33 experiments from the seed round used conditions outside the set considered in the closed-loop optimization. This was because those conditions contained catalysts that were functionally redundant (Figure 3F). The total space to explore was 48 conditions × 11 substrates = 528 reactions, so the model explored 308 / 528 (58%) of the reactions. The model's selection of substrates to test per round is shown in Figure S27.
[0211] Post-hoc analysis - comparison with randomly selected baseline After our model uncertainty plateaued at approximately 3% (consistent with measurement uncertainty), we terminated the closed-loop experiment and performed a post-hoc comparison of our algorithm's performance to a random-selection baseline using the simulation approach described in Section 3.3. To do so, we assumed that the model had collected enough data to make reliable extrapolations to unexplored portions of the reaction-condition space; therefore, we interpreted its predictions across the entire space as ground truth. Because the actual experiment adjusted the batch size after the first two iterations, we performed simulations mimicking these conditions, and the results are shown in Figure 4B. These results suggest that identifying optimal conditions through random selection of the "next" experiment requires nearly twice as many reaction runs as the AI-guided scheme.
[0212] Post-mortem analysis - comparison of yield distributions between our data set and published sets of Suzuki couplings It is also instructive to compare the distribution of reaction yields recorded in our experiments with the distribution of yields for Suzuki couplings reported in the literature performed under the same conditions (i.e., the solvent, base, and ligand combinations used in at least one of our experiments). As can be seen in Figure 4C of this application, the reaction yield is more or less uniformly distributed across the range of expected values for our dataset, whereas for the reactions reported in the literature, it is strongly biased toward higher yields (i.e., closer to the average yield of all literature-reported reactions), with a peak around 70-80%. In other words, our protocol learned by probing both unsuccessful and high-yielding conditions, whereas the published literature is dominated by positive results only (this limits the usefulness of approaches aimed at learning from literature data, as we have discussed in much of our previous work on computerized synthesis). Next, we examined the average yields for the top-k predicted "general conditions" at each iteration (Figure 33 and Table 4). Although the model-predicted ranking of the best "general" conditions was initially inaccurate, the relative performance of the top conditions gradually improved throughout optimization, eventually surpassing the third-best condition, which consistently produced good yields. The average yield for the JACS2009 benchmark conditions was 64%, while the top ML-identified conditions after round 5 had an average yield of 72% (Figures 34A and 34B). The superiority of the top ML conditions was finally confirmed by an additional 20 reactions in the out-of-box test set (Figure 5). [Table 4]
[0213] Post-mortem analysis - generalization of conditions optimized for individual reactions Here, we wish to evaluate a commonly used protocol that involves the optimization of conditions for individual reaction(s) followed by generalization to a wider range of substrates (we refer to this "baseline" method as the "independent BO," or iBO, optimization strategy). In particular, we aim to (1) generate N from a pool of envisaged starting materials; subs We test strategies that (1) select substrate pairs (we assume random selection, as there may be no pre-existing knowledge to drive the selection); (2) independently optimize conditions for each selected pair of starting materials (preferably using some Bayesian optimization approach); and (3) evaluate the best-yielding conditions against other substrates in the pool. In testing this strategy, we used the data generated in our experiments and 1536 data points a posteriori. In step (2), we tested a recent random forest (RF) algorithm using 10 randomly selected conditions for RF initialization, and, for example, the best solution found up to a given point was selected after its "discovery" by N wait A stopping criterion is chosen to terminate the optimization if there is no improvement within a step. For step (3), we subs Aggregate all data collected during the optimization of individual reactions (i.e., all condition-yield pairs) and estimate the average yield under given conditions. This is done by (i) grouping the collected data by reaction condition (whereby each reaction condition has a value between 1 and N depending on the number of optimization runs that occurred); subs (ii) the yields within each group are averaged. subs Because only a subset of substrate pairs can be examined in optimization, there may be a different number of measurements to score each reaction condition (e.g., one yield measurement under condition A versus three measurements under condition B). However, in the absence of a statistical model aimed at addressing this issue (which was the goal of our work), we interpret these yield estimates as the best data available.
[0214] Note that steps 1-3 involve several sources of randomness: (i) the N subs (ii) selection of the pairs (e.g., if seven out of 11 substrate pairs are selected, 330 possible choices are possible); (iii) selection of the initial conditions to seed the BO optimizer for each reaction (we have 48 conditions, so approximately 6 10 of the initial 10 reaction conditions used to seed the system). 9 (iii) the inherent randomness of the random forest algorithm (therefore, the same initial reaction set may generate different optimization trajectories). To average out these factors, we repeated steps (1)-(3) 1000 times, which allowed us to determine whether the "true" general conditions (i.e., the conditions found by the algorithm described in the text) were within the range of N subs It is then possible to calculate the probability that the top condition identified according to the empirical ranking derived from the independent optimization of the reaction is the same as the top-ranked condition (the statistical value of the top 1 from steps (1) to (3)).
[0215] Figure 35 shows the results of both the iBO baseline in our calibration dataset (1536 Buchwald-Hartwig (BH) reactions) and the heteroaryl Suzuki coupling (hetSMC) data collected in this study (missing values imputed with the predictions of the final model), with two control variables: N subs and N wait The heatmaps in panels A and C quantify the top-1 statistics for the BH and hetSMC datasets, respectively. In both cases, the iBO strategy is substantially inferior to the optimization protocol described in this paper. Indeed, for the BH set, N wait Regardless of subs = 1, approximately 5 to 9%, N subs For the hetSMC dataset (a "narrower" condition space), the performance is good: N subs = 1 and N waitFor =3, the top 1 is simply 15.7%, and N subs =7 and N wait For N = 20, it reaches 53.5% (which is consistent with our insight that the use of more substrates and reactions should increase the probability of finding the global optimum). subs =7;N wait Even with the 297 reactions performed during the optimization protocol (=20), these probabilities are roughly a coin toss.
[0216] These results are N subs and N wait It is instructive to add the context of the total number of reactions involved in each combination of N (Figures 35B and 35D). Since in the iBO scheme we have no control over the total number of reactions (optimization may terminate earlier or later depending on which reactions are drawn in the initial pool), each tile actually corresponds to a distribution of the number of reactions, from which we select this maximum (upper bound) for visualization. As expected, this upper bound increases monotonically from the bottom left to the top right for both the BH and hetSMC datasets. For hetSMC, N subs =7 and N wait At = 20, the number of reactions performed (297) matches the total number of experiments performed in our study (approximately 300), but as we discussed above, the iBO baseline has a roughly even probability of a top 1, whereas in our model this converges to 1 (Figure 4B).
[0217] Overall, these results indicate that independent optimization of individual reactions is significantly worse at finding "generally optimal" conditions than the "coupling" method we have described.
[0218] Example 9 - LCMS quantification of reaction yield Calculating Percent Yield After completion of the reaction, the reaction mixture was injected (undiluted, 1 μL) into the LCMS. The relative ratio of the total wavelength count (TWC) UV / Vis integrated peak area between the product and the internal standard (phenanthrene) was determined. This peak area ratio was multiplied by the slope obtained from the response factor curve of the corresponding compound to obtain the percent yield.
[0219] Example 10 - Distribution of Test Set Products and By-Products To understand the cause of the increased yields of the top-ranked ML general reaction conditions compared to the benchmark conditions, we first probed the extended reaction time of the benchmark conditions to rule out incomplete reaction profiles due to the lower temperature of the benchmark conditions (60 °C) compared to the top-ranked ML conditions (100 °C). To do so, we conducted direct comparison experiments at reaction times of 12, 24, and 36 hours for five of the test set compounds that showed the largest yield differences between the benchmark and top-ranked ML conditions (Figure S33). Interestingly, we observed that the reaction yields did not increase when longer reaction times were used. In fact, all yields decreased slightly, likely due to product decomposition. This implies that incomplete conversion at lower reaction temperatures is not the cause of the differences in reaction outcomes.
[0220] To understand why yields increased under the top ML general conditions, we quantified and compared the formation of all observable by-products and remaining starting material under the benchmark and top ML conditions for each of the substrates in the test set, and this information is included in Figure 5E–G. By-products were identified by comparison to purchased authentic standards (if commercially available), ionization by ESI+LCMS (if observed with high mass accuracy and isotopic scoring), or independently synthesized authentic standards (see synthetic details at the end of this section and structural characterization data in Section 7). A comparison of product distributions for exemplary test set substrates under the JACS benchmark and top ML conditions is shown in Figure 40.
[0221] The length of the HPLC gradient and the wide polarity range allowed for the general separation and quantification of these compounds. This data has several limitations, notably the volatility of some byproducts (e.g., thiophenes, thiazoles), which can result in evaporative losses in the argon manifold over the reaction time, especially at elevated temperatures. Furthermore, in some cases, the retention times of byproducts cannot be distinguished from other reaction components. This is most evident in the case of product #19, where protodeboronation and protodehalogenation result in the same chemical structure. Furthermore, MIDA boronate, boronic acid, and homocoupling products are not observed in these reactions. The explanation for these observations is relatively straightforward: MIDA boronate is entirely hydrolyzed during the reaction conditions, boronic acid is entirely consumed or protodeboronated by the reaction conditions, and homocoupling is eliminated by the rigorous air exclusion of the robotic experimental platform. With these aspects formally acknowledged, this analysis demonstrated that the ML general conditions provide higher yields, primarily due to a shift in product distribution away from byproducts and toward targeted product formation.
[0222] Differences between the general conditions for protodeboronation (Figure 5E), residual halide starting material (Figure 5F), and the ratio of product formation to all by-products (Figure 5G) are apparent. For the latter, "all by-products" is defined as the sum of all peaks that cannot be assigned to either internal standards, products, or halides. Differences between the general conditions for protodehalogenation are not apparent, likely due to its relatively low incidence (Figure 41).
[0223] Example 11 - Synthesis of By-Products Because protodehalogenated standards for test set reactions 12 and 18 were not commercially available and were not detectable by ESI+LCMS, these molecules were independently synthesized for comparison as described above. Synthetic details are included below.
[0224] [ka] 12-PDH was prepared according to a modified literature procedure. To a 40 mL I-Chem vial equipped with a septa cap and a rare-earth Teflon-coated stir bar (10 mm diameter) was added 4-bromo-3-methyl-1-(tetrahydro-2H-pyran-2-yl)-1H-pyrazole-5-carbaldehyde (546 mg, 2 mmol), Pd(dtbpf)Cl2 (13.4 mg, 0.02 mmol, 1 mol%), di-H2O (8 mL), and N,N-diisopropylethylamine (1 mL, 6 mmol). The reaction was stirred at room temperature for 2 minutes, followed by the dropwise addition of 1,1,3,3-tetramethyldisiloxane (530 μL, 3 mmol). The vial was sealed and stirred at 500 rpm. After 1 hour, the reaction was extracted with ethyl acetate (2 × 10 mL). The combined organic layers were dried over Na2SO4 and concentrated in vacuo in the presence of Celite (bath temperature: 35 °C). The crude reaction mixture adhered to the Celite was loaded onto a 10 g silica gel column with additional hexane (1 mL) and purified by flash column chromatography (elution: 10 → 20% ethyl acetate / hexane) to give 12-PDH in the presence of siloxane impurities. 12-PDH was further purified by solid-phase extraction: dry loading onto C18 silica gel (dissolving solvent: acetone) followed by elution from the C18 silica gel (eluent: 1:1 water:acetonitrile) afforded pure 12-PDH as a colorless oil (47.4 mg, 0.244 mmol, 12% yield). 1 H-NMR (500 MHz, CDCl3) δ 9.85 (s, 1H), 6.70 (s, 1H), 6.07 (dd, J = 10.3, 2.5 Hz, 1H), 4.12 - 4.04 (m, 1H), 3.72 (td, J = 11.5, 2.5 Hz, 1H), 2.42 - 2.34 (m, 1H), 2.32 (s, 3H), 2.12 - 2.02 (m, 1H), 1.98 - 1.90 (m, 1H), 1.78 - 1.65 (m, 2H), 1.61 - 1.53 (m, 1H). 13C-NMR (125 MHz, CDCl3) δ 179.9, 149.0, 139.8, 115.1, 85.6, 68.3, 29.9, 24.9, 22.7, 13.4.;HRMS (EI+) C 10 H 14 Calculated value for N2O2 [M+] + m / z 194.1055, measured value 194.1051.
[0225] [ka] To a flame-dried 40 mL I-Chem vial equipped with a septa cap and a rare-earth Teflon-coated stir bar (10 mm diameter) was added tert-butyl 4-(5-bromo-7H-pyrrolo[2,3-d]pyrimidin-4-yl)piperazine-1-carboxylate (197 mg, 0.515 mmol) and anhydrous THF (5 mL). The reaction was cooled to -78 °C, and n-BuLi (1.6 M in hexanes, 0.708 mL, 1.13 mmol) was added dropwise. The reaction was stirred for 5 minutes, and methanol (50 μL, 1.24 mmol) was added. The reaction was stirred for 1 minute and quenched by the addition of 10 mL of phosphate buffer (pH 5). The reaction was extracted with ethyl acetate (2 × 10 mL). The combined organic layers were dried over NaSO and concentrated in vacuo in the presence of Celite (bath temperature: 35 °C). The crude reaction mixture adhered to Celite was loaded onto a 10 g silica gel column and purified by flash column chromatography (elution: 50→75% ethyl acetate / hexane [starting material elutes], then 75→90% ethyl acetate / hexane [product elutes]) to give 18-PDH as a white solid (42 mg, 0.138 mmol, 27% yield). 1 H-NMR (500 MHz, CDCl3) δ 12.12 (s, 1H), 8.36 (s, 1H), 7.15 (d, J = 3.7 Hz, 1H), 6.50 (d, J = 3.7 Hz, 1H), 4.02 - 3.93 (m, 4H), 3.68 - 3.55 (m, 4H), 1.49 (s, 9H). 13C-NMR (125 MHz, CDCl) δ 157.1, 154.8, 152.2, 150.6, 121.1, 103.1, 101.0, 80.1, 45.3 (broad), 43.8 (broad), 42.8 (broad), 28.4. HRMS (ESI+) C 15 H 22 Calculated value for N5O2 [M+H] + m / z 304.1773, observed value 304.1774.
[0226] Example 12 - Statistical Comparison Dataset: Figure 5B TIFF2025540581000063.tif148165TIFF2025540581000064.tif160165TIFF2025540581000065.tif174165
[0227] Further, with reference to Figure 5B, the p-values for each comparison were calculated as a function of sample size using the pMoSS:p-value model with the sample size method by Gomez-de-Mariscal and colleagues, along with associated code using standard parameters (Figure 42). In essence, this method uses Monte Carlo cross-validation to model the p-value of the Mann-Whitney U test as a sample-size-dependent function. In this context, a faster decay of the sample-size-dependent p-value function p(n) toward the significance threshold α indicates stronger evidence against the null hypothesis, meaning that statistical significance is revealed at smaller sample sizes. Figure 42 shows that the statistically significant comparisons in Figure 5B remain (specifically, reaching the α threshold of 0.05; General Condition 1 vs. Benchmark, General Condition 2 vs. Benchmark) and exhibit significant differences in the relative slope of the exponential decay function compared to comparisons that lack statistical significance (General Condition 3 vs. Benchmark). These differences are more apparent at sample sizes as small as 5 and, more specifically, at sample sizes of 10. These data can also be examined in terms of a binary index θα,γ (from the same reference), where θα,γ = 1 indicates a difference between the datasets being tested, and θα,γ = 0 indicates failure to reject the null hypothesis based on an area-under-curve comparison between a p-value function dependent on sample size, a constant function of the significance level α, and the regularization parameter γ. In this context, comparisons shown to be statistically significant in Figure 4B have θα,γ = 1, while comparisons lacking statistical significance have θα,γ = 0. Interestingly, only 10 data points were required to observe a statistically significant change between the general and benchmark conditions found by top-rank ML, whereas up to 16 data points were required for statistically significant changes of smaller magnitude. For comparisons lacking statistical significance, the exponential decay function requires approximately >100 data points to reach the threshold α = 0.05. In this regard, in retrospect, these data suggest that the selection of 20 cases in the test set was relatively optimal, and that significant effects of sample size governing p-values may not occur until approximately 100 cases. Dataset: Figure 5C TIFF2025540581000066.tif148165TIFF2025540581000067.tif160165TIFF2025540581000068.tif109165Dataset: Figure 5E TIFF2025540581000069.tif137165 Paired t-test results: TIFF2025540581000070.tif141165 Dataset: Figure 5F TIFF2025540581000071.tif136165 Paired t-test results: TIFF2025540581000072.tif141165 dataset: Figure 5G TIFF2025540581000073.tif136165 Paired t-test results: TIFF2025540581000074.tif147165 dataset: Figure 41 TIFF2025540581000075.tif136165 Paired t-test results: TIFF2025540581000076.tif141165
[0228] Incorporation by Reference All U.S. patents and U.S. and PCT patent application publications referred to herein are incorporated by reference in their entirety as if each individual publication or patent was specifically and individually indicated to be incorporated by reference. In case of conflict, the present application, including any definitions herein, will control.
[0229] equivalent While specific embodiments of the subject disclosure have been discussed, the above specification is illustrative and not restrictive. Many variations of the present disclosure will become apparent to those skilled in the art upon review of this specification and the following claims. The full scope of the disclosure should be determined by reference to the claims, along with their full scope of equivalents, and the specification, along with such variations.
Claims
1. selecting a reaction pair comprising a first molecule and a second molecule, wherein the first molecule is selected from a first matrix and the second molecule is selected from a second matrix; selecting one or more reaction conditions for the reaction pair, said selection being based on past use of the one or more reaction conditions and the structural and functional diversity of the selected reaction pair; automatically carrying out an initial round of reactions between the selected reaction pairs under the one or more selected reaction conditions by a robotic system; optimizing the one or more reaction conditions associated with the initial round of reactions based on the yield of each reaction in the initial round of reactions using a machine learning model; determining an optimized sequence of reactions to be performed; performing a reaction in the optimized series of reactions, thereby forming a product and predicting its yield; outputting an optimized set of generic reaction conditions for each optimized reaction between the selected reaction pairs; A method comprising:
2. The method of claim 1 , further comprising minimizing uncertainty in the machine learning model.
3. Minimizing uncertainty is constructing a surrogate model for predicting one or more reaction yields from said reaction pairs; estimating an objective function for the executed reaction based on the predicted output from the surrogate model; estimating the objective function for unperformed reactions; The method of claim 2 , comprising:
4. Selecting a reaction pair clustering the first matrix by common ring structures and pendant functional groups, the clustering producing a first centroid, the first centroid containing the closest representatives of the first matrix; selecting the second molecule from the second matrix; identifying all combinations that include the first centroid and the second molecule, thereby generating a chemical space; comparing said chemical space with a corpus of chemical products, thereby generating a product space; The method according to any one of claims 1 to 3, comprising:
5. applying a greedy algorithm to the product space; Identifying a set of pairs of first centroids and second molecules that maximize the mutual dissimilarity of the resulting products of said reaction pairs; The method of claim 4 further comprising:
6. selecting the one or more reaction conditions considering at least one of a solvent, a base, a catalyst, and a temperature as a variable for the one or more reaction conditions; determining a set of initial conditions based on an analysis of previous uses of the one or more reaction conditions; The method of any one of claims 1 to 5, further comprising:
7. 7. The method of claim 6, wherein the solvent is 5:1 dioxane:water.
8. The base is K 3 P.O. 4 or Na 2 CO 3 The method according to claim 6 or 7, wherein
9. The method of any one of claims 6 to 8, wherein the catalyst is a palladium catalyst, optionally further comprising a ligand.
10. The ligands include SPhos, XPhos, and triphenylphosphine (PPh 3 10. The method of claim 9, wherein the compound is selected from the group consisting of:
11. The catalyst is Pd(SPhos)G4, Pd(PPh 3 ) 4 The method of any one of claims 6 to 10, wherein the hydroxybenzoate is selected from the group consisting of Pd(XPhos)G4, and Pd(XPhos)G4.
12. The method of any one of claims 6 to 11, wherein the temperature is from about 50°C to about 150°C.
13. The method of any one of claims 6 to 12, wherein the temperature is about 60°C or about 100°C.
14. optimizing the one or more reaction conditions carrying out a reaction under said one or more reaction conditions, thereby forming a product; determining the parameters of said one or more reaction conditions that output the highest yield; and The method according to any one of claims 1 to 13, comprising:
15. optimizing the one or more reaction conditions Identifying one or more catalysts that have similar yields for different substrates; removing said identified catalyst(s) from the envisioned reaction conditions, thereby reducing redundancy; and The method according to any one of claims 1 to 13, comprising:
16. iteratively repeating said optimization of said one or more reaction conditions, thereby forming a product; determining the yield of said product; determining that a threshold yield is met; The method of any one of claims 1 to 15, further comprising:
17. 16. The method of claim 1, further comprising generating a small dataset for optimization by the machine learning model, the small dataset including negative data.
18. The method of any one of claims 1 to 17, wherein the first molecule comprises a halo-substituted aryl or heteroaryl.
19. The first molecule has the formula (Ia): 【Chemistry 1】 wherein: A is aryl or heteroaryl; Each R 1 is independently selected from the group consisting of alkyl, alkoxyl, alkenyl, alkynyl, aralkyl, heteroaryl(alkyl), aryl, heteroaryl, halo, haloalkyl, hydroxyl, carboxyl, acyl, ester, amino, amido, cyano, cycloalkyl, and heterocycloalkyl; n1 is 0, 1, 2, 3, 4, or 5; X 1 is a halo, The method according to any one of claims 1 to 18.
20. X 1 20. The method of claim 19, wherein is bromo.
21. the first molecule is 【Chemistry 2-1】 【Chemistry 2-2】 [Chemistry 2-3] The method of any one of claims 18 to 20, selected from the group consisting of:
22. 22. The method of any one of claims 1 to 21, wherein the second molecule comprises an aryl or heteroaryl further comprising a boronic acid, a boronic ester, or a tetrafluoroborate.
23. The second molecule has the formula (Ib): 【Transformation 3】 wherein: B is aryl or heteroaryl; Each R 2 is independently selected from the group consisting of alkyl, alkoxyl, alkenyl, alkynyl, aralkyl, heteroaryl(alkyl), aryl, heteroaryl, halo, haloalkyl, hydroxyl, carboxyl, acyl, ester, amino, amido, cyano, cycloalkyl, and heterocycloalkyl; n2 is 0, 1, 2, 3, 4, or 5; X 2 is selected from the group consisting of N-methylimidodiacetic acid boronic acid ester, tetramethyl N-methyliminodiacetic acid boronic acid ester, pinacol boronic acid ester, boronic acid, or tetrafluoroborate.
24. the second molecule is 【Chemistry 4-1】 【Chemistry 4-2】 and BMIDA is N-methylimidodiacetic acid boronic acid ester.
25. A robot system, a computing node including a computer readable storage medium having program instructions embodied therein, the program instructions being executable by a processor of the computing node to cause the processor to perform a method including the method of any one of claims 1 to 24; Including, the system.