Protein Sequence Humanness Scoring With Virtual Peptide Mutations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for evaluating immunogenicity of protein sequences are time-consuming and incomplete, failing to comprehensively predict the immunogenicity of all possible peptide chains.

Innovation Solution

A computer-assisted method using a deep learning neural network, integrating a protein Large Language Model (LLM) with architectures like CNN, RNN, GNN, VAE, and Transformer, to predict the humanness score of peptide chains and modify them through single-point virtual mutations to reduce immunogenicity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If experimental approaches are used to evaluate immunogenicity, then measurement precision may be improved, but productivity deteriorates due to time-consuming and expensive processes

Engineering Contradiction:
Improveimmunogenicity prediction accuracyVSAvoidevaluation speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent creates a computational copy of the experimental evaluation process through deep learning models. The protein language model and supervised learning algorithms replicate the immunogenicity assessment function that would otherwise require physical experiments, enabling rapid in silico screening of peptide chains without actual laboratory testing.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical and biological experimental system with an information-processing computational system. Instead of physically synthesizing and testing peptides in the lab, the system uses neural networks and language models to predict immunogenicity from sequence data alone, substituting wet-lab mechanics with digital computation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If experimental approaches are used to evaluate immunogenicity, then measurement precision may be improved, but loss of time increases due to lengthy experimental procedures

Engineering Contradiction:
Improveimmunogenicity prediction accuracyVSAvoidevaluation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary computational screening of all possible peptide chains before actual experimental work. The deep learning model pre-evaluates immunogenicity risk for numerous sequences in silico, allowing researchers to prioritize only the most promising candidates for physical experimentation, thereby reducing overall evaluation time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The computational model creates a virtual replica of the immunogenicity testing process, allowing parallel evaluation of multiple peptide chains simultaneously without the sequential time constraints of laboratory experiments. This digital copying enables rapid assessment that cannot be achieved through physical means.

Inventive Principle:
Principle #26Copying

3Measurement precision

If comprehensive screening of all peptide chains is performed, then measurement precision improves, but device complexity increases due to multiple deep learning models

Engineering Contradiction:
Improvecomprehensive immunogenicity predictionVSAvoidmodel architecture complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the complex evaluation task into distinct functional components: a protein language model for sequence understanding, a supervised learning model for immunogenicity prediction, and separate processing pipelines for different peptide chain lengths. This modular segmentation manages complexity while maintaining comprehensive screening capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The deep learning system is designed as a universal platform that can evaluate immunogenicity across diverse protein sequences and peptide lengths using the same core architecture. The model handles multiple functions (sequence tokenization, embedding, prediction) within a unified framework, reducing overall system complexity despite comprehensive capabilities.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Productivity

If deep learning models are used to predict immunogenicity, then productivity improves through rapid screening, but measurement precision may deteriorate compared to experimental methods

Engineering Contradiction:
Improvescreening speedVSAvoidprediction accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system incorporates feedback mechanisms where prediction results guide further analysis and modification. The immunogenicity scores from the deep learning model feed into peptide chain modification suggestions, creating a closed-loop system that continuously refines predictions and improves accuracy through iterative optimization based on initial screening results.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent combines multiple computational approaches (protein language models, supervised learning, unsupervised learning) into a composite predictive system. This composite model integrates strengths of different algorithms to achieve higher accuracy than any single method alone, bridging the gap between rapid screening and precise measurement.

Inventive Principle:
Principle #40Composite materials

Data Source

PatentUS20250279158A1Computer-assisted method and system for evaluating and modifying immunogenicity of protein sequences using a protein large language model
Publication Date: 2025.09.04 AINNOCENCE LLC
  • US20250279158A1 patent drawing
  • US20250279158A1 patent drawing
  • US20250279158A1 patent drawing

AI summary

The embodiment discloses a computer-assisted method for evaluating and modifying immunogenicity, as well as related computer systems and storage media. The method includes using unsupervised learning of a protein large language model on all human sequences and a supervised deep learning neural networks to establish a predictive scoring model for the humanness score of peptide chains based on the data from the model training dataset to achieve classification between human and non-human species. This involves cutting protein sequences into all possible peptide chains of a preset length using a dynamic window method and importing these chains into the predictive scoring, thereby evaluating their immunogenicity in terms of humanness score. Peptide chains with scores above a certain threshold undergo all possible single-point virtual mutations to generate a set of modified peptide chains which are then reassessed using the model, selecting those with scores above the threshold of humanness for further consideration.