Machine Learning Protein Sequence Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing protein engineering methods, such as directed evolution, are resource-intensive and time-consuming, requiring extensive experimental efforts to generate and screen protein variants, while molecular dynamics simulations are computationally expensive and often lack necessary reference structures.

Innovation Solution

A machine learning-guided approach using a trained model to predict optimized protein or nucleic acid sequences by evaluating substitutions and mutations based on a set of training data, including native and engineered sequences, to generate sequences with improved functionality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If directed evolution is used to generate protein variants, then protein functionality can be improved, but the experimental burden and time required increase significantly

Engineering Contradiction:
Improveprotein functionalityVSAvoidtime required for variant generation and screening
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent creates computational copies (in silico models) of protein variants through machine learning algorithms that predict protein properties and guide sequence optimization, eliminating the need for physical laboratory screening of each variant and dramatically reducing the time required compared to traditional directed evolution

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical experimental process of generating and physically screening protein variants in the laboratory with a computational machine learning system that predicts protein properties and identifies optimized sequences through algorithmic analysis of training data

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If molecular dynamics simulations are used to predict protein variant changes, then structural predictions can be made, but computational resources and processor hours required increase significantly

Engineering Contradiction:
Improvestructural prediction accuracyVSAvoidcomputational resources required
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent uses machine learning models trained on existing protein structure and property data to create predictive copies that can rapidly estimate protein properties without requiring computationally intensive molecular dynamics simulations, reducing processor hours from hundreds to minimal levels

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent changes the approach from physics-based molecular dynamics simulations to data-driven machine learning models that use trained parameters and patterns from training data to predict protein properties, fundamentally altering how computational resources are consumed from high-energy simulations to low-energy pattern recognition

Inventive Principle:
Principle #35Parameter changes

3Manufacturing precision

If full molecular dynamics simulations are performed for each variant, then detailed structural changes can be predicted, but the complexity and resource intensity increase

Engineering Contradiction:
Improveprediction detail levelVSAvoidcomputational complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent extracts essential predictive capabilities from complex molecular dynamics simulations by training machine learning models on representative data, separating the core predictive function from the computationally intensive simulation process, and enabling detailed predictions without the full computational complexity

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS20250210146A1Sequence optimization
Publication Date: 2025.06.26 UAB BIOMATTER DESIGNS
  • US20250210146A1 patent drawing
  • US20250210146A1 patent drawing
  • US20250210146A1 patent drawing

AI summary

Provided herein is an apparatus for generating an optimized protein or nucleic acid sequence from a target protein or nucleic acid sequence, wherein the optimized protein or nucleic acid sequence has an improved function over the target sequence. The apparatus comprises at least one processor and at least one memory including a computer program code, the at least one memory and computer program code configured to, with the at least one processor, cause the apparatus to at least: operate a machine learning model configured to receive the target protein sequence or a nucleic acid sequence, and to generate therefrom one or more corresponding optimized sequences, wherein the machine learning model has been trained on a set of training data comprising native or engineered protein or nucleic acid sequences, and additionally, at least a subset of the sequences comprising one or more masked portions and/or at least a subset of the sequences comprising one or more mutations introduced therein.