Protein Sequence Design with Dynamic AlphaFold Structure Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for de novo protein design struggle to reliably predict accurate high-resolution protein structures from amino acid sequences, particularly in the context of thermodynamic equilibrium and conformational changes, and lack effective integration of structural and sequence information.
Innovation Solution
A computer-implemented method that uses AlphaFold for structural prediction, integrates fitness functions considering both structural and biological properties, and employs evolutionary algorithms for optimizing amino acid sequences to achieve native-like structures with enhanced solubility and expression.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If AlphaFold is used for structure prediction, then structural prediction accuracy is improved, but the ability to predict proteins under thermodynamic equilibrium and conformational changes deteriorates
Solution Approach 1:
The patent applies dynamics by training AlphaFold to predict ensembles of structures representing different conformational states rather than a single static structure. This enables the model to capture proteins in thermodynamic equilibrium and during conformational changes, making the prediction system adaptive to dynamic biological realities while maintaining high structural accuracy for each predicted state.
2Adaptability or versatility
If de novo protein design methods are used, then novel protein structures can be created, but reliability of accurate high-resolution structure prediction deteriorates
Solution Approach 1:
The patent implements feedback by using the trained AlphaFold model to evaluate and refine designed protein sequences iteratively. The model predicts structures for designed sequences, and these predictions feed back into the design process to guide optimization toward sequences that fold into accurate, native-like high-resolution structures. This closed-loop approach ensures both novelty and reliability.
3Measurement precision
If traditional structure determination methods are used, then experimental accuracy is improved, but productivity and speed of protein design deteriorates
Solution Approach 1:
The patent applies copying by training AlphaFold on experimental protein structure data from databases like PDB. The model learns to replicate the accuracy of experimental structure determination by copying known structure-sequence relationships. Once trained, it can rapidly predict structures for novel sequences without performing actual experiments, achieving both experimental-level accuracy and computational speed.
4Stability of the object's composition
If protein design focuses on single stable structure, then structural definition is improved, but ability to model conformational changes and intrinsically disordered proteins deteriorates
Solution Approach 1:
The patent applies dynamics by modifying AlphaFold to predict multiple conformations and ensembles rather than a single static structure. This enables the model to represent proteins undergoing conformational changes and intrinsically disordered proteins that sample multiple states, while still providing well-defined structural information for each predicted conformational state through the dynamics module.
Data Source
AI summary
Provided is a computer implemented method for designing at least one protein includes creating at least one amino acid sequence to be tested, wherein ones of the amino acids, contained in the amino acid sequence to be tested, are selected according to a probability distribution; predicting, from the aligned at least one amino acid sequence, structural properties of the at least one protein; calculating, based on the structural properties of the at least one protein, a value of a fitness function for the at least one amino acid sequence to be tested; selecting or deselecting, dependent on the value of the fitness function, the at least one amino acid sequence to be tested. The amino acid sequence may be post-processed to be more native-like, have enhanced solubility, and/or improved expression.


