ProGen Transformer Model for Controllable Protein Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current protein engineering methods rely heavily on expensive and time-consuming structural annotations and traditional heuristics or random mutations, resulting in limited success in generating functional proteins with desired properties.
Innovation Solution
A neural network-based protein generation model, ProGen, uses transformer architectures from natural language processing to generate amino acid sequences conditioned on target protein properties, leveraging large datasets of protein sequences without structural annotations for efficient and controlled protein generation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional heuristics and random mutations are used for protein engineering, then the process is simple and requires minimal computational resources, but the success rate in generating functional proteins is very limited
Solution Approach 1:
The patent introduces a language model as an intermediary between protein sequences and structural/functional properties. The model is trained on large datasets of protein sequences with annotated properties, enabling it to predict structural and functional characteristics without requiring expensive experimental structural annotations for each new protein design. This intermediary system bridges the gap between simple sequence generation and accurate property prediction.
Solution Approach 2:
The language model is pre-trained on extensive protein sequence data with structural and functional annotations before being used for protein engineering tasks. This preliminary training phase allows the model to learn complex relationships between sequences and properties, so that during actual protein engineering applications, predictions can be made quickly without requiring new experimental structural data for each case.
2Measurement precision
If structural annotations are obtained through traditional methods, then accurate structural information is available, but the process is expensive and time consuming
Solution Approach 1:
Instead of obtaining structural information through expensive and time-consuming experimental methods for each new protein, the patent uses a language model that has learned structural patterns from a dataset of annotated proteins. The model generates predictions that copy the accuracy of experimental structural annotations by leveraging patterns learned from training data, effectively creating a computational surrogate for experimental structural determination.
Solution Approach 2:
The patent replaces expensive, time-consuming experimental structural annotation methods with a computational language model that provides comparable accuracy at minimal cost and time. The model can be trained once on a comprehensive dataset and then used repeatedly for countless protein sequence evaluations without requiring additional experimental resources for each prediction.
3Adaptability or versatility
If protein sequences are generated using random mutations, then diverse sequences can be explored, but the generation process lacks control over desired protein properties
Solution Approach 1:
The language model provides feedback mechanisms that enable controlled protein sequence generation. By conditioning the generation process on desired structural and functional properties, the model can generate sequences that are both diverse and targeted toward specific property goals. The model learns from training data which sequence features correlate with desired properties, allowing it to guide the generation process toward achieving target characteristics while maintaining sequence diversity.
Data Source
AI summary
The present disclosure provides systems and methods for controllable protein generation. According to some embodiments, the systems and methods leverage neural network models and techniques that have been developed for other fields, in particular, natural language processing (NLP). In some embodiments, the systems and methods use or employ models implemented with transformer architectures developed for language modeling and apply the same to generative modeling for protein engineering.


