Sequence Activity Models for Protein Library Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Protein design is challenging due to the vastness of sequence space, making it difficult to identify functional proteins with desired properties efficiently, as existing methods are hindered by the NP-hard nature of the protein design problem and the complexity of exploring all possible protein variants.

Innovation Solution

The development of methods and systems for identifying amino acid residues to vary in protein sequences using sequence activity models, which predict activity based on amino acid residue types and positions, allowing for the systematic variation and optimization of protein libraries to impact desired activities such as stability or catalytic activity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If exhaustive exploration of sequence space is performed to identify functional proteins, then functional proteins with desired properties can be identified, but the computational time and resources required become prohibitively large due to the NP-hard nature of the problem

Engineering Contradiction:
Improveidentification of functional proteinsVSAvoidcomputational time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by using machine learning models to predict protein function and properties before performing exhaustive computational exploration. The system pre-trains models on known protein structures and functions, then uses these models to guide the search through sequence space, identifying promising candidates before full analysis. This preliminary computational preparation significantly reduces the time required for subsequent exhaustive searches while maintaining high reliability in identifying functional proteins.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces machine learning models as intermediary components between the sequence space exploration and functional protein identification. These models act as mediators that predict protein properties and guide the search process, enabling efficient navigation through the vast sequence space without requiring exhaustive computation of all possible sequences. The intermediary models filter and prioritize candidate sequences, reducing computational time while maintaining identification accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If directed evolution with high throughput screening is used to design proteins, then functional proteins can be identified, but the process becomes complex and time-consuming due to iterative recombination formats

Engineering Contradiction:
Improveprotein function identificationVSAvoidevolution process complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent replaces the mechanical iterative recombination and screening process with computational machine learning models that can predict protein functions directly from sequences. Instead of physically performing iterative evolution experiments, the system uses AI models to simulate and predict outcomes, substituting the complex mechanical evolution process with computational prediction. This significantly reduces process complexity while maintaining reliable protein function identification.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the fundamental parameters of the protein design process from physical iterative recombination to computational prediction. Rather than performing physical evolution cycles with screening, the system uses machine learning models to predict and optimize protein properties directly. This parameter change from physical to computational domain dramatically reduces complexity while maintaining functional identification reliability.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If systematic variation of amino acid residues is performed to optimize protein activity, then desired protein activities can be achieved, but the search space becomes too large to explore exhaustively

Engineering Contradiction:
Improveprotein activity optimizationVSAvoidnumber of protein variants
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent applies partial action by using machine learning models to predict protein activities and identify only the most promising protein variants for further investigation. Instead of exhaustively generating and analyzing all possible protein variants, the system performs partial exploration guided by predictive models that filter out unlikely candidates. This approach achieves reliable protein activity optimization while dramatically reducing the number of variants that need to be systematically varied and tested.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS9996661B2Methods, systems, and software for identifying functional bio-molecules
Publication Date: 2018.06.12 CODEXIS INC
  • US9996661B2 patent drawing
  • US9996661B2 patent drawing
  • US9996661B2 patent drawing

AI summary

The present invention generally relates to methods of rapidly and efficiently searching biologically-related data space. More specifically, the invention includes methods of identifying bio-molecules with desired properties, or which are most suitable for acquiring such properties, from complex bio-molecule libraries or sets of such libraries. The invention also provides methods of modeling sequence-activity relationships. As many of the methods are computer-implemented, the invention additionally provides digital systems and software for performing these methods.