Deep Representation Learning for Protein Engineering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional protein engineering methods rely on random mutagenesis and screening without modeling the relationship between protein sequence and function, limiting the efficiency and predictability of developing new protein variants with improved or novel functions.

Innovation Solution

An unsupervised deep representation learning model is used to learn fundamental protein characteristics from large collections of protein sequences, enabling the prediction and optimization of protein stability and function, and guiding the design of new protein variants through the use of artificial neural networks and recurrent neural networks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If random mutagenesis and screening methods are used for protein engineering, then new protein variants can be generated, but the efficiency and predictability of developing new protein variants with improved or novel functions is limited

Engineering Contradiction:
Improveefficiency of developing new protein variantsVSAvoidpredictability of protein function
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent replaces the mechanical/random approach of traditional mutagenesis and screening with a computational/intellectual system based on deep learning models. The system uses neural networks to predict protein function and stability, substituting random experimental variation with rational, model-guided design. This allows researchers to computationally predict which mutations will achieve desired functions before performing experiments, dramatically improving both efficiency and predictability.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the fundamental parameters of protein engineering by introducing quantitative models that predict protein properties as continuous functions of sequence. Instead of relying on discrete, random mutations followed by screening, the system uses deep learning to evaluate and predict the effects of mutations on protein stability and function, enabling rational design based on predicted parameter changes rather than random sampling.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If traditional random mutagenesis approaches are used, then protein variants can be explored, but extensive experimental characterization is required

Engineering Contradiction:
Improveability to explore protein variantsVSAvoidtime for experimental characterization
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent performs preliminary computational analysis using deep learning models to predict protein function and stability before conducting experiments. The system evaluates numerous potential variants in silico, identifying promising candidates for experimental validation. This preliminary computational screening reduces the number of variants that require extensive experimental characterization, saving time and resources while maintaining the ability to explore diverse protein sequence space.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates computational copies and models of protein sequences and their predicted functions. Instead of physically creating and testing every possible variant, the system generates virtual representations of protein variants and their expected properties through deep learning models. These computational copies allow researchers to screen and evaluate numerous variants in silico before selecting a small subset for physical experimentation.

Inventive Principle:
Principle #26Copying

3Reliability

If model-based rational design is used, then quantitative models of protein properties can be built, but a generalizable framework has not been consolidated to date

Engineering Contradiction:
Improvequantitative modeling of protein propertiesVSAvoidcomplexity of generalizable framework
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent develops a universal deep learning framework that can predict multiple protein properties (stability, function, binding affinity) from sequence data alone. The same neural network architecture and training approach can be applied across different protein families and functions, providing a generalizable solution rather than requiring separate models for each specific property or protein type. This multi-functional framework consolidates various prediction tasks into a unified system.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent segments the complex task of protein design into distinct predictive components handled by specialized neural network models. Different models or model components are trained to predict specific properties (e.g., stability, enzymatic activity, binding affinity) independently. This segmentation allows each model to be optimized for its specific function while maintaining a modular architecture that can be combined and applied generally across different protein engineering tasks.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12040050B1Systems and methods for rational protein engineering with deep representation learning
Publication Date: 2024.07.16 NABLA BIO INC
  • US12040050B1 patent drawing
  • US12040050B1 patent drawing
  • US12040050B1 patent drawing

AI summary

A dataset describing a collection of proteins is loaded, which identifies, for each protein, a respective value of a characteristic of interest. The dataset is provided as one or more inputs to a trained unsupervised representation model to cause the trained unsupervised representation model to generate a representation for each protein in the collection. The representation for each protein is input into a supervised top model to train the supervised top model to obtain a predicted characteristic and the trained supervised top model is used to obtain a predicted characteristic for a particular protein.