Protein Sequence Humanness Scoring With Virtual Peptide Mutations
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for evaluating immunogenicity of protein sequences are time-consuming and incomplete, failing to comprehensively predict the immunogenicity of all possible peptide chains.
Innovation Solution
A computer-assisted method using a deep learning neural network, integrating a protein Large Language Model (LLM) with architectures like CNN, RNN, GNN, VAE, and Transformer, to predict the humanness score of peptide chains and modify them through single-point virtual mutations to reduce immunogenicity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If experimental approaches are used to evaluate immunogenicity, then measurement precision may be improved, but productivity deteriorates due to time-consuming and expensive processes
Solution Approach 1:
The patent creates a computational copy of the experimental evaluation process through deep learning models. The protein language model and supervised learning algorithms replicate the immunogenicity assessment function that would otherwise require physical experiments, enabling rapid in silico screening of peptide chains without actual laboratory testing.
Solution Approach 2:
The patent replaces the mechanical and biological experimental system with an information-processing computational system. Instead of physically synthesizing and testing peptides in the lab, the system uses neural networks and language models to predict immunogenicity from sequence data alone, substituting wet-lab mechanics with digital computation.
2Measurement precision
If experimental approaches are used to evaluate immunogenicity, then measurement precision may be improved, but loss of time increases due to lengthy experimental procedures
Solution Approach 1:
The patent performs preliminary computational screening of all possible peptide chains before actual experimental work. The deep learning model pre-evaluates immunogenicity risk for numerous sequences in silico, allowing researchers to prioritize only the most promising candidates for physical experimentation, thereby reducing overall evaluation time.
Solution Approach 2:
The computational model creates a virtual replica of the immunogenicity testing process, allowing parallel evaluation of multiple peptide chains simultaneously without the sequential time constraints of laboratory experiments. This digital copying enables rapid assessment that cannot be achieved through physical means.
3Measurement precision
If comprehensive screening of all peptide chains is performed, then measurement precision improves, but device complexity increases due to multiple deep learning models
Solution Approach 1:
The patent segments the complex evaluation task into distinct functional components: a protein language model for sequence understanding, a supervised learning model for immunogenicity prediction, and separate processing pipelines for different peptide chain lengths. This modular segmentation manages complexity while maintaining comprehensive screening capability.
Solution Approach 2:
The deep learning system is designed as a universal platform that can evaluate immunogenicity across diverse protein sequences and peptide lengths using the same core architecture. The model handles multiple functions (sequence tokenization, embedding, prediction) within a unified framework, reducing overall system complexity despite comprehensive capabilities.
4Productivity
If deep learning models are used to predict immunogenicity, then productivity improves through rapid screening, but measurement precision may deteriorate compared to experimental methods
Solution Approach 1:
The system incorporates feedback mechanisms where prediction results guide further analysis and modification. The immunogenicity scores from the deep learning model feed into peptide chain modification suggestions, creating a closed-loop system that continuously refines predictions and improves accuracy through iterative optimization based on initial screening results.
Solution Approach 2:
The patent combines multiple computational approaches (protein language models, supervised learning, unsupervised learning) into a composite predictive system. This composite model integrates strengths of different algorithms to achieve higher accuracy than any single method alone, bridging the gap between rapid screening and precise measurement.
Data Source
AI summary
The embodiment discloses a computer-assisted method for evaluating and modifying immunogenicity, as well as related computer systems and storage media. The method includes using unsupervised learning of a protein large language model on all human sequences and a supervised deep learning neural networks to establish a predictive scoring model for the humanness score of peptide chains based on the data from the model training dataset to achieve classification between human and non-human species. This involves cutting protein sequences into all possible peptide chains of a preset length using a dynamic window method and importing these chains into the predictive scoring, thereby evaluating their immunogenicity in terms of humanness score. Peptide chains with scores above a certain threshold undergo all possible single-point virtual mutations to generate a set of modified peptide chains which are then reassessed using the model, selecting those with scores above the threshold of humanness for further consideration.


