Contact Point Pair Encoding for Interpretable Binding Affinity Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for predicting peptide-MHC binding affinity lack interpretability and provide uncertainly estimates, limiting their application in biomedical fields, particularly in vaccine development, where mechanistic understanding and human intervention are essential.

Innovation Solution

A computer-implemented method that encodes amino acid sequences as contact point pairs and applies a machine learning or statistical model to predict binding affinity, providing interpretable results and uncertainty estimates, facilitating rational decision-making in vaccine development.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If machine learning methods are used to predict binding affinity, then prediction accuracy is improved, but interpretability deteriorates

Engineering Contradiction:
Improveprediction accuracyVSAvoidinterpretability
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent segments the binding interface into discrete contact points and encodes amino acid pairs at each contact point as separate features. This segmentation allows the model to process complex binding interactions through modular, interpretable units while maintaining high prediction accuracy through comprehensive feature coverage.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms the binding affinity prediction problem by changing the parameter representation from raw sequence data to encoded amino acid pair features at contact points. This parameter transformation enables the use of interpretable linear models while capturing the essential binding determinants through carefully selected contact point encodings.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If complex machine learning models are applied, then prediction quality is improved, but mechanistic interpretation is lost

Engineering Contradiction:
Improveprediction qualityVSAvoidmodel complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts the essential binding determinants by identifying and encoding only the amino acid pairs at contact points, separating the critical binding information from the rest of the sequence data. This extraction enables simplified linear models to achieve high prediction quality by focusing on the most relevant features.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the model complexity parameter by transitioning from complex non-linear machine learning models to simple linear models. This parameter change is compensated by transforming the input parameters into encoded amino acid pair features that capture binding determinants in a linearly separable representation.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If allele-specific models are used, then prediction accuracy for specific alleles is improved, but applicability to novel alleles deteriorates

Engineering Contradiction:
Improveprediction accuracyVSAvoidallele generalizability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal binding affinity prediction model that functions across multiple MHC alleles by encoding amino acid pairs at contact points in a allele-agnostic manner. This universal encoding scheme allows the same linear model to accurately predict binding for both known and novel alleles, achieving multi-functionality without sacrificing accuracy.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20220208301A1Method and system for binding affinity prediction and method of generating a candidate protein-binding peptide
Publication Date: 2022.06.30 NEC ONCOIMMUNITY AS
  • US20220208301A1 patent drawing
  • US20220208301A1 patent drawing
  • US20220208301A1 patent drawing

AI summary

According to a first aspect of the present invention there is provided a computer-implemented method of predicting a binding affinity value of a query binder molecule to a query target molecule, the query binder molecule having a first amino acid sequence and the query target molecule having a second amino acid sequence, the method comprising: encoding the first and second amino acid sequences together as a plurality of data elements to generate an encoded pair of amino acids, each data element of the encoded pair representing which amino acids from the first and second amino acid sequences are paired at a respective contact point between the first amino acid sequence and the second amino acid sequence to form a contact point pair, wherein a contact point pair is a pairing of amino acids from a binder molecule and a target molecule which are proximal to one another to influence binding; and, applying a machine learning or statistical model to the encoded pair of amino acids to predict a binding affinity value, wherein the machine learning model or statistical model is trained by: accessing, with at least one processor, a reference data store of reference binder-target pairs comprising respective paired reference binder sequences and reference target sequences, each reference binder-target pair having an associated measured binding value; and, encoding each reference binder-target pair as a plurality of data elements, each data element of the encoded reference binder-target pair representing which amino acids from the respective paired reference binder sequences and reference target sequences are paired at a respective contact point to form a contact point pair, such that the predicted binding affinity value is representative of a contribution to binding of each contact point pair of the query binder molecule and the query target molecule.