Transformer Model Predicts Enzyme-Substrate Interactions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Engineering enzymes with novel functionality from scratch remains a significant challenge due to the dependency on identifying a suitable starting point with measurable activity, as existing methods lack efficiency in predicting enzyme-substrate interactions and amino acid mutations.
Innovation Solution
A machine learning model, specifically a multi-headed self-attention transformer-based model, is used to predict enzyme-substrate pairs and their interaction probabilities by embedding enzyme primary sequences, substrate representations, and interaction flags, allowing for the prediction of amino acid mutations and alterations in substrates, thereby assessing the likelihood of enzyme-substrate interactions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If directed evolution is used to engineer enzymes, then functional enzymes can be obtained with measurable activity, but the process heavily depends on identifying a suitable starting point which requires domain expert intuition and serendipity
Solution Approach 1:
The patent replaces the manual, expert-intuition-based mechanical process of identifying parent enzymes with an automated machine learning system. The ML model predicts enzyme-substrate interactions and suggests mutations without requiring domain expert intuition or serendipity, thereby reducing identification process complexity while maintaining reliable enzyme activity outcomes.
Solution Approach 2:
The system enables enzymes to be engineered through self-service by using computational predictions to identify suitable parent enzymes and guide mutation strategies. The ML model autonomously analyzes enzyme sequences, predicts substrate interactions, and generates mutation suggestions, eliminating the need for expert intervention in the identification process.
2Productivity
If existing methods are used for predicting enzyme-substrate interactions, then some predictions can be made, but the methods lack efficiency in predicting enzyme-substrate interactions and amino acid mutations
Solution Approach 1:
The patent employs a composite approach by integrating multiple machine learning models (sequence encoders, graph neural networks for substrates, interaction predictors) into a unified system. This composite ML architecture combines different computational techniques to simultaneously achieve high prediction efficiency and accurate interaction predictions, overcoming the limitations of individual methods.
Solution Approach 2:
The machine learning system performs multiple functions simultaneously: it predicts enzyme-substrate interactions, suggests amino acid mutations, and identifies suitable parent enzymes. This multi-functional approach increases overall prediction efficiency while maintaining or improving accuracy across different prediction tasks through a single integrated platform.
Data Source
AI summary
Techniques for predicting a pair of an enzyme primary sequence and a substrate, and interaction probability for the pair are described. An exemplary method includes receiving a request to predict a pair of an enzyme primary sequence and a substrate, and interaction probability for the pair; combining an enzyme vector, a substrate vector, and an interaction indication for the enzyme and substrate to form a machine learning model input; applying a machine learning model to the machine learning model input to predict the pair of an enzyme primary sequence and a substrate, and interaction probability for the pair; and outputting a result of the application of the machine learning model including the predicted pair of an enzyme primary sequence and a substrate, and interaction probability for the pair.


