Deep Learning Antigen Chain Pairing via Transformer Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for identifying B cell receptor (BCR) heavy-light chain pairs and T cell receptor (TCR) αβ chain pairs are limited in throughput, specificity, and applicability, particularly in bulk sequencing approaches where pairing information is lost, and existing computational methods are restricted to specific datasets and sequences, making it challenging to generate viable light chains for therapeutic antibody discovery.

Innovation Solution

A deep learning method, termed 'Matchmaker,' uses a Transformer model to predict light chains from heavy chains, framing the problem as a neural machine translation task, and is expanded to tackle TCR chain pairing, enabling the generation of viable light chains for any given heavy chain and predicting functional pairings with improved accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If bulk B cell sequencing is used to increase throughput, then the number of B cell sequences recovered is improved, but heavy-light chain pairing information is lost

Engineering Contradiction:
ImprovethroughputVSAvoidpairing information
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent uses computational pairing methods as an intermediary to reconstruct heavy-light chain pairing information from bulk sequencing data. The system processes bulk sequencing results through algorithms that infer pairings based on sequence characteristics and statistical models, thereby recovering the lost pairing information without requiring single-cell sequencing

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates virtual representations of paired chains by generating paired heavy and light chain sequences through computational methods. These virtual pairings are constructed by matching sequences based on their statistical properties and structural constraints, effectively copying the pairing relationships that would exist in single-cell data

Inventive Principle:
Principle #26Copying

2Loss of information

If single-cell sequencing is used to preserve pairing information, then heavy-light chain pairing information is maintained, but throughput is limited

Engineering Contradiction:
Improvepairing informationVSAvoidthroughput
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The patent replaces the mechanical single-cell sequencing approach with computational pairing methods. Instead of physically isolating and sequencing individual cells to preserve pairing information, the system uses algorithmic processing of bulk sequencing data to reconstruct pairing relationships, thereby achieving high throughput while maintaining information integrity

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Loss of information

If computational pairing methods are used to recover pairing information, then pairing data is reconstructed, but accuracy and applicability are limited to specific datasets

Engineering Contradiction:
Improvepairing informationVSAvoidaccuracy
Core Design Contradiction:
Loss of informationVSReliability

Solution Approach 1:

The patent develops computational pairing methods that are designed to work across diverse datasets and sequence types. The system incorporates multiple pairing strategies and statistical models that can adapt to different biological contexts, making the pairing reconstruction accurate and reliable for various antibody and T cell receptor datasets beyond the specific training data

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20240203523A1Engineering of antigen-binding proteins
Publication Date: 2024.06.20 ALCHEMAB THERAPEUTICS LTD
  • US20240203523A1 patent drawing
  • US20240203523A1 patent drawing
  • US20240203523A1 patent drawing

AI summary

Methods of identifying an antigen-binding protein comprising a pair of chains are described. The methods comprise providing a query sequence comprising a first chain sequence, and identifying a corresponding chain sequence by providing the query sequence to a deep learning model configured to take as input a query first chain sequence and to produce as output at least one corresponding chain sequence, thereby identifying a corresponding chain sequence for the query sequence, wherein the deep learning model has been trained using training first and corresponding chain sequences from known chain pairs. The first chain sequence may be a heavy/light chain of an antibody or B cell receptor or β/α/δ/γ chain of a T cell receptor, and the corresponding chain may be a light/heavy chain of an antibody or B cell receptor or an a β/α/δ/γ chain of a T cell receptor. The methods find uses in any context where it is desirable to identify chain pairings for antigen-binding molecules, such as e.g. in the context of identifying antigen-binding molecules that have a desired (e.g. therapeutic or functional) property. Related methods, systems and products are described.