IgLM Generative Model for Antibody Sequence Design
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for generating synthetic antibody libraries face challenges such as poor expression, low solubility, low thermal stability, and high aggregation, and require large, diverse libraries with substantial fractions of non-functional antibodies.
Innovation Solution
The development of a bidirectional generative immunoglobulin language model (IgLM) that leverages bidirectional context for designing antibody sequence spans of varying lengths, trained on a large-scale natural antibody dataset to generate full-length antibody sequences and diversify loops, resulting in high-quality libraries with favorable biophysical properties.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If massive synthetic libraries on the order of 10^10-10^11 variants are constructed to discover antibodies with high affinity, then the diversity and potential functionality of the library is improved, but the fraction of non-functional antibodies increases and the complexity of library construction increases
Solution Approach 1:
The patent uses generative AI models to create synthetic antibody sequences by copying and adapting from natural antibody sequences in training data. The model learns patterns from observed antibody repertoires and generates novel sequences that replicate successful structural and functional characteristics, eliminating the need to physically construct massive libraries.
Solution Approach 2:
The patent replaces the mechanical process of physically constructing and screening massive antibody libraries with a computational approach using generative AI. Instead of synthesizing 10^10-10^11 physical antibody variants, the system uses machine learning algorithms to generate equivalent or superior diversity through computational sequence design.
2Reliability
If traditional hybridoma technology or phage display technology is used to obtain monoclonal antibodies, then the process is experimentally validated, but the antibodies exhibit poor expression, low solubility, low thermal stability, and high aggregation
Solution Approach 1:
The patent applies parameter changes by modifying specific sequence parameters in the antibody CDR regions through AI-generated variations. The model adjusts amino acid sequences, loop lengths, and structural parameters to optimize biophysical properties such as solubility, thermal stability, and aggregation resistance while maintaining antigen-binding functionality.
Solution Approach 2:
The patent incorporates feedback mechanisms where the generative AI model uses observed data from natural antibody repertoires and experimental results to iteratively refine its predictions. The model learns from successful antibody designs and adjusts its generation process to produce sequences with improved biophysical properties in subsequent iterations.
3Adaptability or versatility
If synthetic DNA is introduced into regions of the antibody sequences that define the complementarity determining regions (CDRs) to create synthetic antibody libraries, then new antigen-binding sites are created, but the space of possible synthetic antibody sequences becomes extremely large yielding substantial fractions of non-functional antibodies
Solution Approach 1:
The patent applies local quality by focusing the generative AI model's variations specifically on the CDR regions where antigen binding occurs, while maintaining the framework regions intact. The model introduces diversity locally in the CDR loops rather than randomly mutating the entire antibody sequence, ensuring that only the necessary regions are optimized for antigen binding while preserving overall structural integrity and functionality.
Data Source
AI summary
Provided herein are methods of producing a trained model for generating peptide or protein sequence information and infilling of targeted residue spans. In some embodiments, the methods include training a model using a training dataset comprising a population of reference amino acid sequence representations in which a given amino acid sequence representation in the population is conditioned on one or more conditioning tags that provide a controllable generation of selected amino acid sequence representation types. Related methods, systems, and computer program products are also provided.


