Sequence Activity Model for Protein Design

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for protein design are hindered by the vast combinatorial explosion of possible molecules in sequence space, making it impractical to exhaustively explore and identify proteins with desired properties using existing high-throughput screening and recombination formats, which are time- and cost-intensive.

Innovation Solution

The development of methods using genetic algorithms and support vector machines to build sequence activity models, filtering out uninformative data, and training models with structural data to guide directed evolution, enabling the identification of proteins with beneficial properties from complex libraries.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If high throughput screening and recombination formats are used to explore protein sequence space, then the efficiency of identifying useful polypeptides is improved, but the time and cost required increase significantly

Engineering Contradiction:
Improveefficiency of identifying useful polypeptidesVSAvoidtime required for screening and sequencing
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent creates computational models (digital twins) of protein sequences and structures that replicate physical screening results. By training machine learning models on existing experimental data, the system can predict protein properties and activity without physically testing each variant, effectively copying the outcome of exhaustive screening through computational simulation rather than physical experimentation.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces physical high-throughput screening mechanisms with computational modeling and machine learning algorithms. Instead of using robotic systems to physically assay thousands of protein variants, the system uses software to simulate and predict protein behavior, substituting mechanical screening processes with digital analysis that is faster and more scalable.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If exhaustive exploration of sequence space is performed to identify proteins with desired properties, then the completeness of protein discovery is improved, but the computational and experimental resources required become impractical

Engineering Contradiction:
Improvecompleteness of protein discoveryVSAvoidcomplexity of screening and sequencing systems
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent performs preliminary computational filtering and prioritization before physical experimentation. By using machine learning models to pre-rank and select the most promising protein variants based on predicted activity and properties, the system reduces the scope of exhaustive exploration to only the most likely candidates, making complete discovery feasible without requiring examination of every possible sequence.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces computational models as intermediaries between sequence space exploration and physical protein testing. These models act as mediators that translate theoretical sequence possibilities into predicted functional properties, allowing the system to navigate the complexity of sequence space intelligently rather than exhaustively, reducing the apparent complexity while maintaining discovery completeness.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20220238179A1Methods and systems for engineering biomolecules
Publication Date: 2022.07.28 CODEXIS INC
  • US20220238179A1 patent drawing
  • US20220238179A1 patent drawing
  • US20220238179A1 patent drawing

AI summary

Disclosed are methods for building a sequence activity model with reference to structural data, which model can be used to guide directed evolution of proteins having beneficial properties. Some embodiments use genetic algorithms and structural data to filter out uninformative data. Some embodiments use a support vector machine to train the sequence activity model. The filtering and training methods can generate a sequence activity model having higher predictive power than conventional modeling methods. Systems and computer program products implementing the methods are also provided.