Antibody Classification via Deep Mutational Scanning and Machine Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The process of optimizing therapeutic antibodies in antibody drug discovery is hindered by the need to address multiple parameters such as expression level, viscosity, pharmacokinetics, and immunogenicity, which is challenging due to the low-throughput nature of mammalian cell expression systems and the limited screening of minor changes, often leading to unintended consequences like diminished antigen binding.

Innovation Solution

A method combining directed evolution with machine learning to classify amino acid sequences of binding proteins, generating variant sequences through deep mutational scanning and using machine learning models to predict antigen-specificity and optimize properties like affinity and developability, enabling the identification of thousands of optimized lead candidates from a vast protein sequence space.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If mammalian cell expression systems are used to screen antibody libraries, then full-length IgG antibodies can be expressed and screened for multiple parameters, but the screening throughput is limited to about 10^3 antibody molecules due to low-throughput cloning, transfection and purification strategies

Engineering Contradiction:
Improvemulti-parameter optimization capabilityVSAvoidscreening throughput
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent uses deep sequencing to create digital copies of antibody sequences from pooled libraries, eliminating the need for physical cloning and individual well transfection. By sequencing DNA from pooled mammalian cell transfections, the system can screen 10^5-10^6 sequences while maintaining the ability to assess multiple parameters including antigen binding, expression level, and developability properties

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent combines multiple screening objectives into a single pooled library screening approach. Multiple parameters (antigen binding affinity, expression level, solubility, immunogenicity) are assessed simultaneously from the same library through deep sequencing and machine learning analysis, rather than screening separate libraries for each parameter

Inventive Principle:
Principle #5Merging (Combining)

2Ease of manufacture

If only minor changes (single point mutations) are screened in mammalian cells, then cloning and transfection can be managed at low-throughput, but the fraction of protein sequence space interrogated is extremely small

Engineering Contradiction:
Improvecloning and transfection feasibilityVSAvoidsequence space coverage
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent adds the dimension of library diversity by enabling screening of 10^5-10^6 sequences instead of 10^3. This is achieved through pooled library transfection followed by deep sequencing, which allows comprehensive interrogation of protein sequence space while maintaining manageable cloning and transfection processes through high-throughput molecular biology techniques

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Reliability

If lead optimization addresses multiple parameters in parallel (expression level, viscosity, pharmacokinetics, solubility, immunogenicity), then comprehensive antibody evaluation is achieved, but the time and costs take up the majority of the drug preclinical discovery and development cycle

Engineering Contradiction:
Improvecomprehensive parameter evaluationVSAvoidoptimization cycle duration
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary assessment of multiple parameters (antigen binding, expression level, solubility, immunogenicity) during the initial library screening phase rather than sequentially during optimization. Deep sequencing and machine learning models evaluate all these parameters simultaneously from pooled libraries, allowing early identification of candidates that meet multiple criteria and reducing iterative optimization cycles

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20220157403A1Systems and methods to classify antibodies
Publication Date: 2022.05.19 ALLOY THERAPEUTICS INC
  • US20220157403A1 patent drawing
  • US20220157403A1 patent drawing
  • US20220157403A1 patent drawing

AI summary

The present disclosure describes systems and methods to make predictions classifying one or more properties of a binding protein such as an antibody, for example, antibody affinity or specificity for an antigen. The system can include one or more machine learning models that can extrapolate complex relationships between amino acid sequence and function. The system can be trained on high-quality training data generated through a two-step single-site and combinatorial deep mutational scanning approach. The trained models can then make predictions on novel variant sequences generated in silico. The present disclosure describes amino acid sequences generated by the systems and methods provided, and uses of the generated sequences to produce proteins for therapeutic and diagnostic use.