Machine Learning Catalyst Prediction via Variational Autoencoder

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current catalyst design in asymmetric catalysis faces challenges due to the lack of mechanistic understanding, limitations in recognizing patterns in large data sets, and the absence of quantitative guidelines, making it difficult to accurately predict selective catalysts, especially beyond the bounds of training data.

Innovation Solution

A chemoinformatics-guided workflow using average steric occupancy (ASO) and electronic descriptors to create a universal training set of catalysts, enabling the prediction of enantioselective reactions by selecting optimal catalysts through machine learning methods, which covers a broad range of feature space and captures subtle catalyst structures responsible for enantioinduction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If machine learning methods are used to predict enantioselectivity, then the ability to identify selective catalysts is improved, but the accuracy of predictions beyond training data bounds deteriorates

Engineering Contradiction:
Improveprediction accuracyVSAvoidextrapolation capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent transforms the prediction problem from selecting among existing catalysts to generating entirely new catalyst structures by adding a generative dimension. The variational autoencoder learns the latent space of catalyst structures and enables navigation to regions corresponding to high enantioselectivity that were not present in the training data, thus resolving the contradiction between prediction accuracy and extrapolation capability.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent performs preliminary generation and screening of catalyst candidates in silico before experimental validation. The generative model pre-identifies promising catalyst structures with predicted high enantioselectivity, allowing experimentalists to focus resources on the most promising candidates and achieving better extrapolation beyond training data.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If a large number of catalyst candidates are evaluated, then the probability of finding highly selective catalysts is improved, but the time and resources required for screening increase

Engineering Contradiction:
Improvecatalyst selection accuracyVSAvoidscreening time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent creates a virtual copy of the catalyst screening process through the variational autoencoder model. Instead of physically synthesizing and testing numerous catalyst candidates, the model generates and evaluates virtual catalyst structures in silico, dramatically reducing the time and resources required while maintaining the ability to identify highly selective catalysts.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary in silico screening of generated catalyst candidates before experimental validation. The variational autoencoder pre-identifies promising catalyst structures based on learned structure-activity relationships, allowing experimentalists to prioritize the most promising candidates and significantly reducing overall screening time and resources.

Inventive Principle:
Principle #10Preliminary action

3Loss of information

If traditional Edisonian empiricism is used for catalyst design, then mechanistic understanding is maintained, but the ability to find patterns in large data sets deteriorates

Engineering Contradiction:
Improvepattern recognition capabilityVSAvoidmethodological complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent replaces the mechanical process of human pattern recognition with a computational variational autoencoder model. The neural network automatically learns complex patterns in catalyst structure-enantioselectivity relationships from large datasets, overcoming the limitations of human cognitive processing while maintaining interpretability through visualization of the latent space and generated structures.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11664093B2Extrapolative prediction of enantioselectivity enabled by computer-driven workflow, new molecular representations and machine learning
Publication Date: 2023.05.30 THE BOARD OF TRUSTEES OF THE UNIV OF ILLINOIS
  • US11664093B2 patent drawing
  • US11664093B2 patent drawing
  • US11664093B2 patent drawing

AI summary

Catalyst design in asymmetric reaction development has traditionally been driven by empiricism, wherein experimentalists attempt to qualitatively recognize structural patterns to improve selectivity. Machine learning algorithms and chemoinformatics can potentially accelerate this process by recognizing otherwise inscrutable patterns in large datasets. Herein we report a computationally guided workflow for chiral catalyst selection using chemoinformatics at every stage of development. Robust molecular descriptors that are agnostic to the catalyst scaffold allow for selection of a universal training set on the basis of steric and electronic properties. This set can be used to train machine learning methods to make highly accurate predictive models over a broad range of selectivity space. Using support vector machines and deep feed-forward neural networks, we demonstrate accurate predictive modeling in the chiral phosphoric acid-catalyzed thiol addition to N-acylimines.