Surrogate Model for Protein Mutation Screening
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for directed protein evolution are inefficient due to the challenges of multi-parameter optimization, limited mutation rate ranges, combinatorial explosion, and the need for iterative and costly experimental evaluation, leading to slow convergence and high resource consumption.
Innovation Solution
The integration of in silico detection methods for single mutations, intelligent library construction for wide mutation rates, diverse screening, and optimal library design using a search model to identify combinatorial substitutions that improve protein properties, enabling faster convergence on desired protein candidates.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If conventional combinatorial library construction methods are used to target single residue mutations, then the method is simple to implement, but it can only target a limited number of mutations in limited regions and requires multiple iterative rounds
Solution Approach 1:
The patent segments the protein sequence into multiple regions and systematically targets mutations in each region by dividing the combinatorial library into subsets, each focusing on specific residue positions. This segmentation enables comprehensive coverage of multiple mutation regions simultaneously rather than requiring sequential iterative rounds
Solution Approach 2:
The patent transitions from conventional single-point mutation libraries to multi-dimensional combinatorial libraries that simultaneously explore multiple residue positions and mutation rates. This dimensional expansion allows the library to represent complex mutation combinations that would traditionally require multiple evolution rounds to discover
2Reliability
If iterative directed evolution is performed to obtain final candidates meeting desired criteria, then the method is proven successful, but it is expensive, time-consuming, and path dependent
Solution Approach 1:
The patent performs preliminary computational screening and in silico evaluation of combinatorial mutants before experimental construction. By pre-filtering and prioritizing promising candidates through computer-based prediction, the method reduces the number of experimental iterations needed and eliminates time-consuming trial-and-error processes
Solution Approach 2:
The patent implements feedback mechanisms where experimental results from initial screening are fed back into computational models to refine predictions for subsequent library construction. This iterative feedback loop optimizes the search strategy and reduces the total number of evolution rounds required
3Device complexity
If linear methods are used to acquire working models for protein evolution direction, then the method is simple, but it cannot adequately consider interactions between different residues
Solution Approach 1:
The patent combines multiple computational approaches and data types into a composite predictive model. This integrates sequence-based predictions, structure-based analyses, and experimental data into a unified framework that can capture residue-residue interactions and epistatic effects that linear methods miss
4Quantity of substance
If AI inference is used to generate combined mutants, then a large number of mutants with comparable functions are produced, but painstaking efforts are required to prioritize a limited number for synthesis and evaluation
Solution Approach 1:
The patent replaces manual, labor-intensive prioritization processes with automated computational scoring and ranking systems. Machine learning algorithms automatically evaluate and rank generated mutants based on predicted functional properties, eliminating the need for painstaking manual review and selection
Data Source
AI summary
Systems and methods improving a target property of a target protein are provided. Each single point mutation in a first plurality of single point mutations of the target protein is obtained, each defining a corresponding single point substituted protein with respect to a reference sequence. A set of values for a set of properties of each point substituted protein is used to filter the first plurality of single point mutations into a reduced second plurality of single point mutations. The second plurality of mutations informs the selection of combinatorially substituted proteins and the target property is measured for each of them. The combinatorially substituted proteins and their measured values serve to train a surrogate model that, in turn, serves to update a search model. The updated search mode informs on which of the second plurality point mutations are to be used in future mutants of the target protein.


