Variable-Length Peptide Folding for Neoantigen Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing cancer vaccines face challenges due to the diversity of neoantigen sequences and lengths, leading to data shortages and information loss in AI-based prediction models, limiting their effectiveness in predicting immunogenicity and binding forces with HLA alleles.
Innovation Solution
A method that processes peptides of varying lengths into a unified unit length by folding and combining sequence values, using a model to predict immunogenicity and binding force, determining neoantigens based on HLA class I and II alleles.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If AI-based prediction models use diverse neoantigen sequences of various lengths, then the coverage of neoantigen types is improved, but data shortages and information loss occur leading to reduced prediction accuracy
Solution Approach 1:
The patent applies parameter changes by transforming the length parameter of neoantigen sequences. Instead of treating sequences of various lengths as separate categories, the invention folds them into a unified fixed-length format (e.g., 9-mer or 10-mer), thereby changing the length parameter while preserving the essential immunogenicity and binding information. This resolves the contradiction by maintaining data completeness (avoiding information loss) while achieving universal processing (improving coverage).
Solution Approach 2:
The patent segments longer neoantigen sequences into smaller fixed-length units through folding operations. For example, a 15-mer sequence is divided and folded into multiple 9-mer or 10-mer segments. This segmentation allows the model to process diverse length sequences using a unified framework, improving adaptability without losing critical binding information, as each segment retains the essential features for immunogenicity and HLA binding prediction.
2Manufacturing precision
If conventional methods process peptides of various lengths separately, then the precision for each length is maintained, but the complexity of the prediction system increases and data efficiency decreases
Solution Approach 1:
The patent implements universality by creating a single prediction model that can handle neoantigen sequences of any length through the folding mechanism. Instead of developing separate models for 8-mers, 9-mers, 10-mers, etc., the invention uses one unified model that processes all lengths by folding them into a standard format. This reduces system complexity while maintaining prediction precision, as the model learns generalizable features that apply across different lengths.
Solution Approach 2:
The patent inverts the conventional approach by not processing each length separately but rather transforming all lengths into a common format. Instead of adapting the model to various lengths, the invention adapts the input data (peptide sequences) to a fixed length through folding. This inversion simplifies the system architecture while preserving the ability to accurately predict immunogenicity and binding affinity across different peptide lengths.
Data Source
Figure 1
Figure 2
Figure 3a
AI summary
Disclosed is a method for processing peptide sequences exceeding a unit length through a process of folding when predicting neoantigens using peptide sequences and HLA class I and/or HLA class II allele sequences. According to the method, peptide sequences contained in cancer tissue can determine neoantigens in the cancer tissue, regardless of the diversity in length of the peptide sequences. Accordingly, it is possible to overcome the imbalance and lack of information in training data on length and more accurately predict a binding force to determine the neoantigens in the cancer tissue.