Computational Antibody Variant Generation via Epistatic Modeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for identifying monoclonal antibodies in drug discovery, such as hybridoma technologies and display technologies, face challenges in efficiently exploring the molecular coevolution of antibody sequences, particularly due to the high divergent sequence identity of the CDRH3 region.

Innovation Solution

A novel computational pipeline that utilizes machine learning models to generate a candidate pool of antibody variants by analyzing molecular coevolutionary landscapes, iteratively generating variants, and computationally screening them based on structural and biophysical properties.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional laboratory methods (hybridoma technologies or display technologies) are used to identify monoclonal antibodies, then antibody identification can be performed, but the efficiency of exploring molecular coevolution of antibody sequences is limited

Engineering Contradiction:
Improveefficiency of exploring molecular coevolutionVSAvoidtime required for antibody identification
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent replaces conventional laboratory methods (hybridoma technologies, display technologies) with a computational pipeline that uses machine learning models to analyze molecular coevolution. The system uses sequence data, multiple sequence alignment, and epistatic models to predict antibody variants, eliminating the need for time-consuming wet lab experiments while maintaining antibody identification capability.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent performs preliminary computational analysis by collecting homologous sequences, creating multiple sequence alignments, and training epistatic models before actual antibody identification. This preliminary action prepares the computational framework in advance, enabling rapid prediction of antibody variants without requiring time-consuming iterative laboratory work.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If direct couplings analysis is applied to antibody CDR regions, then molecular coevolution can be analyzed, but the requirement for large multiple sequence alignment with coverage over CDRH3 region becomes unachievable due to high sequence divergence

Engineering Contradiction:
Improveaccuracy of direct couplings analysisVSAvoidnumber of relevant sequences available
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent changes the approach from requiring large numbers of sequences with high sequence identity to using a targeted search strategy that collects homologous sequences based on specific structural and functional parameters. The system uses a novel search approach that prioritizes sequences with relevant structural features over purely sequence-based similarity, enabling accurate coevolution analysis with fewer sequences.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces an intermediary search approach that mediates between the goal of obtaining CDRH3 coverage and the reality of sequence divergence. Instead of directly searching for sequences with high CDRH3 identity (which yields few results), the system uses a multi-stage search that incorporates germline gene information, structural constraints, and functional annotations to identify relevant sequences.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If a novel search approach is used to collect homologous sequences and represent them in multiple sequence alignment, then coverage over CDRH3 region can be achieved, but the complexity of sequence collection and alignment increases

Engineering Contradiction:
Improvecoverage over CDRH3 regionVSAvoidcomplexity of sequence collection and alignment
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the sequence collection process into distinct functional stages: (1) input sequence data, (2) homologous sequence collection using novel search, (3) multiple sequence alignment, (4) epistatic model computation, and (5) variant generation. This segmentation allows each stage to be optimized independently, managing overall system complexity while achieving comprehensive CDRH3 coverage.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal computational pipeline that handles multiple tasks within the same framework: sequence collection, alignment, coevolution analysis, and variant prediction. The epistatic model serves multiple purposes by capturing both direct couplings and higher-order interactions, reducing the need for separate analysis tools and simplifying the overall system.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250131983A1Computationally Directed Protein Sequence Evolution
Publication Date: 2025.04.24 EVQLV INC
  • US20250131983A1 patent drawing
  • US20250131983A1 patent drawing
  • US20250131983A1 patent drawing

AI summary

Sequence data is received that specifies at least one sequence of interest. Thereafter, homologous sequence are collected based on the sequence data and are represented in a multiple sequence alignment using a novel search approach. Next, an epistatic model is computed by a first machine learning model that represents a revolutionary landscape of the multiple sequence alignment. Later, a second machine learning model is used to iteratively generate statistical inferences based upon the epistatic model, to result in a candidate pool of sequences comprising variants of the sequence of interest. Data can then be provided which characterizes the candidate pool of sequences.