Generative ML Architecture for Protein Sequence Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The development of proteins for therapeutic purposes is resource-intensive and limited by the need for extensive synthesis and testing, and existing techniques for classifying amino acid sequences based on structural features are computationally expensive and inaccurate due to reliance on first principles-based approaches and large datasets.

Innovation Solution

The use of generative machine learning architectures, such as generative adversarial networks and autoencoders, to produce amino acid sequences for training inferential models, which can classify proteins based on structural features with reduced computational resources and increased accuracy by iteratively refining the models using datasets with specific characteristics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If first principles-based approaches and large datasets are used for classifying amino acid sequences, then measurement precision may be improved, but computational cost and device complexity increase significantly

Engineering Contradiction:
Improveclassification accuracyVSAvoidcomputational resources
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent uses generative models to create synthetic amino acid sequences that copy the statistical and structural properties of real protein sequences. These synthetic copies serve as training data, replacing the need for extensive real experimental data while maintaining classification accuracy. The generative models learn from real sequences and generate realistic variants that preserve key biological features.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces traditional first principles-based computational approaches with machine learning models. Instead of using computationally intensive physics-based simulations to analyze amino acid sequences, the system uses trained neural networks that have learned patterns from data, significantly reducing computational requirements while maintaining or improving accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If extensive synthesis and testing of proteins is performed, then reliability of therapeutic protein development is improved, but loss of time and resources increase

Engineering Contradiction:
Improvetherapeutic protein developmentVSAvoiddevelopment time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary classification and prediction of protein characteristics using trained machine learning models before actual synthesis and testing. By predicting which amino acid sequences are likely to produce desired therapeutic properties, the system filters out poor candidates early, preventing wasted time and resources on unsuccessful synthesis attempts.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The generative models are trained to understand the relationship between amino acid sequences and protein properties, enabling the system to self-evaluate and predict outcomes without requiring extensive external experimental validation at each step. The models internally capture biological rules and patterns, allowing autonomous prediction of protein behavior.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12080380B2Implementing a generative machine learning architecture to produce training data for a classification model
Publication Date: 2024.09.03 JUST EVOTEC BIOLOGICS INC
  • US12080380B2 patent drawing
  • US12080380B2 patent drawing
  • US12080380B2 patent drawing

AI summary

Amino acid sequences of proteins can be produced using one or more generative machine learning architectures. The amino acid sequences produced by the one or more generative machine learning architectures can be used to train a classification model architecture. The classification model architecture can classify amino acid sequences according to a number of classifications. Individual classifications of the number of classifications can correspond to at least one of a structural feature of proteins, a range of values of a structural feature of proteins, a biophysical property of proteins, or a range of values of a biophysical property of proteins.