Variable-Length Peptide Folding for Neoantigen Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing cancer vaccines face challenges due to the diversity of neoantigen sequences and lengths, leading to data shortages and information loss in AI-based prediction models, limiting their effectiveness in predicting immunogenicity and binding forces with HLA alleles.

Innovation Solution

A method that processes peptides of varying lengths into a unified unit length by folding and combining sequence values, using a model to predict immunogenicity and binding force, determining neoantigens based on HLA class I and II alleles.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If AI-based prediction models use diverse neoantigen sequences of various lengths, then the coverage of neoantigen types is improved, but data shortages and information loss occur leading to reduced prediction accuracy

Engineering Contradiction:
Improvecoverage of neoantigen typesVSAvoidprediction accuracy
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent applies parameter changes by transforming the length parameter of neoantigen sequences. Instead of treating sequences of various lengths as separate categories, the invention folds them into a unified fixed-length format (e.g., 9-mer or 10-mer), thereby changing the length parameter while preserving the essential immunogenicity and binding information. This resolves the contradiction by maintaining data completeness (avoiding information loss) while achieving universal processing (improving coverage).

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent segments longer neoantigen sequences into smaller fixed-length units through folding operations. For example, a 15-mer sequence is divided and folded into multiple 9-mer or 10-mer segments. This segmentation allows the model to process diverse length sequences using a unified framework, improving adaptability without losing critical binding information, as each segment retains the essential features for immunogenicity and HLA binding prediction.

Inventive Principle:
Principle #1Segmentation

2Manufacturing precision

If conventional methods process peptides of various lengths separately, then the precision for each length is maintained, but the complexity of the prediction system increases and data efficiency decreases

Engineering Contradiction:
Improveprediction precision for each lengthVSAvoidsystem complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent implements universality by creating a single prediction model that can handle neoantigen sequences of any length through the folding mechanism. Instead of developing separate models for 8-mers, 9-mers, 10-mers, etc., the invention uses one unified model that processes all lengths by folding them into a standard format. This reduces system complexity while maintaining prediction precision, as the model learns generalizable features that apply across different lengths.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent inverts the conventional approach by not processing each length separately but rather transforming all lengths into a common format. Instead of adapting the model to various lengths, the invention adapts the input data (peptide sequences) to a fixed length through folding. This inversion simplifies the system architecture while preserving the ability to accurately predict immunogenicity and binding affinity across different peptide lengths.

Inventive Principle:
Principle #13The other way round (Inversion)

Data Source

PatentEP4625421A1Method and computer program for predicting neoantigens by processing lengths of peptides of various lengths by folding
Publication Date: 2025.10.01 THERAGEN BIO CO LTD
  • EP4625421A1 patent drawingFigure 1
  • EP4625421A1 patent drawingFigure 2
  • EP4625421A1 patent drawingFigure 3a

AI summary

Disclosed is a method for processing peptide sequences exceeding a unit length through a process of folding when predicting neoantigens using peptide sequences and HLA class I and/or HLA class II allele sequences. According to the method, peptide sequences contained in cancer tissue can determine neoantigens in the cancer tissue, regardless of the diversity in length of the peptide sequences. Accordingly, it is possible to overcome the imbalance and lack of information in training data on length and more accurately predict a binding force to determine the neoantigens in the cancer tissue.