Multi-Chain Protein Encoding with Linker-Based Chain Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional computational methods for predicting properties of multi-chain proteins are inconsistent and unreliable due to sensitivity in encoding and presenting multi-chain proteins, failing to distinguish between different chains and being limited by the need for extensive experimental data.

Innovation Solution

A method involving generating concatenated amino acid sequences with linkers, encoding them using a protein language model, and processing with a trained machine learning model to accurately predict properties such as aggregation, stability, and viscosity, while augmenting training data with permutations and linkers to enhance model training efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional computational methods are used to predict multi-chain protein properties, then the prediction process is simple, but the prediction accuracy and reliability are low

Engineering Contradiction:
Improveprediction accuracyVSAvoidmethod complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments multi-chain protein sequences by inserting unique linker tokens between chains, allowing the model to distinguish individual chains within concatenated sequences. This segmentation approach transforms the complex multi-chain prediction problem into manageable chain-level representations while maintaining overall protein property prediction accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces linker tokens as intermediary elements between different protein chains in the sequence representation. These linkers serve as mediators that enable the machine learning model to identify chain boundaries and relationships, improving prediction reliability without requiring complex separate processing for each chain.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If extensive experimental data is collected for training, then model prediction reliability improves, but time consumption and cost increase

Engineering Contradiction:
Improvemodel reliabilityVSAvoidtraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies data augmentation techniques as a preliminary action during training, generating synthetic variations of existing protein sequences through permutations and linker insertions. This preliminary data preparation expands the training dataset without requiring additional experimental measurements, thereby improving model reliability while avoiding the time and cost of collecting more experimental data.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If multi-chain proteins are encoded without distinguishing chains, then encoding is simple, but prediction consistency deteriorates

Engineering Contradiction:
Improveprediction consistencyVSAvoidencoding complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies local quality differentiation by inserting unique linker tokens at specific locations between chains in the sequence. This local modification enables the encoding to distinguish between different chains while maintaining the overall simplicity of concatenated sequence representation, thereby achieving prediction consistency without excessive encoding complexity.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20250335825A1Data augmentation and encoding of multi-chain protein structures
Publication Date: 2025.10.30 AMGEN INC
  • US20250335825A1 patent drawing
  • US20250335825A1 patent drawing
  • US20250335825A1 patent drawing

AI summary

Described herein are techniques for predicting one or more properties of a multi-chain protein, the multi-chain protein including at least a first chain and a second chain. In some embodiments, the techniques include: obtaining sequence data for the multi-chain protein, the sequence data indicating a first amino acid sequence specifying at least a portion of the first chain and a second amino acid sequence specifying at least a portion of the second chain; generating a concatenated amino acid sequence by concatenating the first amino acid sequence, a linker, and the second amino acid sequence; encoding the concatenated amino acid sequence to obtain a numeric representation of the concatenated amino acid sequence; and processing the numeric representation of the concatenated amino acid sequence using a trained machine learning model to obtain an output indicative of the one or more properties of the multi-chain protein.