Protein Landscape Mapping for High-Coverage Proteoform Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional mass-spectrometry-based proteomics techniques provide limited protein sequence coverage, making it difficult to distinguish various forms of proteins (proteoforms) and detect modifications, splicing events, and single nucleotide polymorphisms, which are essential for understanding protein presence and regulation in cells and therapeutic proteins.
Innovation Solution
The use of multiple proteases for sample digestion combined with high-resolution mass-spectrometry and bioinformatic analysis to achieve increased sequence coverage, allowing for the identification of proteoforms and modifications across the full length of proteins, with up to 80% mean sequence coverage and improved amino acid resolution.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional proteases (e.g., trypsin) are used for protein digestion, then the digestion process is simple and efficient, but the sequence coverage is limited to 15-20%, which is insufficient for distinguishing proteoforms and detecting modifications
Solution Approach 1:
The patent applies segmentation by dividing the protein digestion process into multiple independent digestions using different proteases (e.g., trypsin, chymotrypsin, Lys-N, Lys-C, Glu-C, Asp-N). Each protease generates complementary peptide fragments, and combining results from multiple digestions achieves comprehensive sequence coverage exceeding 80%, enabling detection of proteoforms and modifications that single proteases miss
Solution Approach 2:
The patent employs a composite digestion strategy combining multiple proteases with different specificities. This composite approach creates a synergistic effect where each protease contributes unique peptide fragments, collectively covering the entire protein sequence and enabling comprehensive proteoform characterization
2Loss of information
If multiple proteases are used to increase sequence coverage, then proteoform identification and modification detection improve, but the analysis time and data processing complexity increase
Solution Approach 1:
The patent implements preliminary action by performing all protease digestions simultaneously on replicate samples before mass spectrometry analysis. This parallel processing approach allows comprehensive peptide library generation from multiple proteases to be completed in advance, with data integration performed through computational algorithms that efficiently combine results and identify proteoforms
Solution Approach 2:
The patent uses replicate samples for each protease digestion, creating multiple copies of the protein sample that can be processed in parallel. This copying strategy enables comprehensive data collection without extending the overall analysis timeline, as multiple samples are analyzed concurrently through high-throughput mass spectrometry
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach enables comprehensive analysis of protein sequences and proteoforms, providing detailed insights into protein modifications and variations, which is crucial for characterizing therapeutic proteins and ensuring quality assurance in pharmaceutical applications.
Implementation Method 1
proteins in a sample undergo proteolytic digestion, breaking the proteins into smaller pieces (peptides)
Implementation Method 2
the resulting data is often processed using a search engine in conjunction with a sequence database containing data of known peptides and proteins
Implementation Method 3
generating tandem mass spectrometry data on each digested polypeptide sample
Data Source
AI summary
In shotgun proteomics, generally only a fraction of peptides from a parent protein are actually detected. Because a large portion of the protein sequence is not detected, it is often impossible to determine whether the expressed protein is present in a modified, spliced, or truncated form. Provided herein are methods and systems for analyzing polypeptides which allow for the increase of the mean sequence coverage of a protein concomitant with bioinformatics analysis in order to distinguish putative proteoforms with improved amino acid resolution. Aspects of the invention include (1) a deep sequencing strategy to provide more protein sequence coverage than is typically achieved, and (2) a computational approach to view protein expression across its full length and identify regions of the protein that are potentially subject to such regulation. This technology has global utility in proteomics and will be of particular use for the analysis of biosimilar protein drug therapeutics.

