Generative Adversarial Networks for Antibody Sequence Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The development of proteins with therapeutic benefits is a resource-intensive and time-consuming process, limited by the need for extensive synthesis and testing, and conventional computer-implemented techniques are limited in scope, accuracy, and complexity, particularly in generating protein sequences with specified characteristics and lengths.

Innovation Solution

The use of generative adversarial networks (GANs) and autoencoder architectures to efficiently generate amino acid sequences of proteins and antibodies, trained on diverse datasets to produce sequences with specific biophysical properties and structures, and the separate generation and combination of heavy and light chain sequences to improve efficiency and reduce computational resources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If conventional computer-implemented techniques are used to generate protein sequences, then the process is simpler, but the accuracy and complexity of generated sequences are limited

Engineering Contradiction:
Improveaccuracy of generated protein sequencesVSAvoidcomplexity of generation system
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent replaces conventional computer-implemented techniques with a generative adversarial network (GAN) system that uses neural network architectures (including CNNs, RNNs, LSTMs, and attention mechanisms) to generate protein sequences. This substitution of traditional computational methods with advanced machine learning systems enables significantly improved accuracy and complexity in generated sequences while maintaining computational feasibility through the distributed architecture of multiple neural network components.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If extensive synthesis and testing are performed to develop therapeutic proteins, then the reliability of therapeutic outcomes is improved, but the process becomes more resource-intensive and time-consuming

Engineering Contradiction:
Improvereliability of therapeutic protein developmentVSAvoidspeed of developing therapeutic proteins
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent performs preliminary computational generation and filtering of protein sequences using GANs before physical synthesis. The system pre-generates multiple candidate sequences, applies computational filters to identify promising candidates based on desired properties, and only then proceeds to synthesis and testing. This preliminary computational action significantly reduces the number of sequences requiring resource-intensive wet lab validation, thereby improving productivity while maintaining reliability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements a feedback mechanism where generated protein sequences are evaluated against desired properties (such as binding affinity, solubility, and stability), and this evaluation feedback is used to refine and improve subsequent generations. The system learns from previous generation results and adjusts its generation process to produce higher-quality sequences, reducing the need for extensive iterative synthesis and testing while maintaining high reliability in therapeutic outcomes.

Inventive Principle:
Principle #23Feedback

3Manufacturing precision

If GANs are used to generate protein sequences with specified characteristics, then the manufacturing precision is improved, but the computational resources required increase

Engineering Contradiction:
Improveaccuracy of generated protein sequencesVSAvoidcomputational resources for sequence generation
Core Design Contradiction:
Manufacturing precisionVSUse of energy by moving object

Solution Approach 1:

The patent segments the protein sequence generation task into separate GAN models for different protein components (such as heavy chains and light chains for antibodies). Each segmented model specializes in generating specific portions of the protein structure, allowing for more efficient computational processing compared to a single monolithic model. This segmentation reduces the computational burden on any single system component while maintaining high accuracy in the generated sequences.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent generates a large number of candidate protein sequences beyond what is immediately needed, using the excess computational capacity to explore the sequence space more thoroughly. This partial action approach generates more sequences than strictly necessary, but the high volume of generated data improves the statistical quality and accuracy of the results. The system accepts excessive computational action in the generation phase to achieve superior manufacturing precision in the final therapeutic protein development.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentEP3956896B1Generation of protein sequences using machine learning techniques
Publication Date: 2024.05.01 JUST EVOTEC BIOLOGICS INC
  • EP3956896B1 patent drawingFigure 1
  • EP3956896B1 patent drawingFigure 2
  • EP3956896B1 patent drawingFigure 3

AI summary

Amino acid sequences of antibodies can be generated using a generative adversarial network that includes a first generating component that generates amino acid sequences of antibody light chains and a second generating component that generates amino acid sequences of antibody heavy chains. Amino acid sequences of antibodies can be produced by combining the respective amino acid sequences produced by the first generating component and the second generating component. The training of the first generating component and the second generating component can proceed at different rates. Additionally, the antibody amino acids produced by combining amino acid sequences from the first generating component and the second generating component may be evaluated according to complentarity-determining regions of the antibody amino acid sequences. Training datasets may be produced using amino acid sequences that correspond to antibodies have particular binding affinities with respect to molecules, such as binding affinity with major histocompatibility complex (MHC) molecules.