Protein Landscape Mapping for High-Coverage Proteoform Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional mass-spectrometry-based proteomics techniques provide limited protein sequence coverage, making it difficult to distinguish various forms of proteins (proteoforms) and detect modifications, splicing events, and single nucleotide polymorphisms, which are essential for understanding protein presence and regulation in cells and therapeutic proteins.

Innovation Solution

The use of multiple proteases for sample digestion combined with high-resolution mass-spectrometry and bioinformatic analysis to achieve increased sequence coverage, allowing for the identification of proteoforms and modifications across the full length of proteins, with up to 80% mean sequence coverage and improved amino acid resolution.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional proteases (e.g., trypsin) are used for protein digestion, then the digestion process is simple and efficient, but the sequence coverage is limited to 15-20%, which is insufficient for distinguishing proteoforms and detecting modifications

Engineering Contradiction:
Improvesequence coverageVSAvoiddigestion process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the protein digestion process into multiple independent digestions using different proteases (e.g., trypsin, chymotrypsin, Lys-N, Lys-C, Glu-C, Asp-N). Each protease generates complementary peptide fragments, and combining results from multiple digestions achieves comprehensive sequence coverage exceeding 80%, enabling detection of proteoforms and modifications that single proteases miss

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent employs a composite digestion strategy combining multiple proteases with different specificities. This composite approach creates a synergistic effect where each protease contributes unique peptide fragments, collectively covering the entire protein sequence and enabling comprehensive proteoform characterization

Inventive Principle:
Principle #40Composite materials

2Loss of information

If multiple proteases are used to increase sequence coverage, then proteoform identification and modification detection improve, but the analysis time and data processing complexity increase

Engineering Contradiction:
Improveproteoform information completenessVSAvoidanalysis time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent implements preliminary action by performing all protease digestions simultaneously on replicate samples before mass spectrometry analysis. This parallel processing approach allows comprehensive peptide library generation from multiple proteases to be completed in advance, with data integration performed through computational algorithms that efficiently combine results and identify proteoforms

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses replicate samples for each protease digestion, creating multiple copies of the protein sample that can be processed in parallel. This copying strategy enables comprehensive data collection without extending the overall analysis timeline, as multiple samples are analyzed concurrently through high-throughput mass spectrometry

Inventive Principle:
Principle #26Copying

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach enables comprehensive analysis of protein sequences and proteoforms, providing detailed insights into protein modifications and variations, which is crucial for characterizing therapeutic proteins and ensuring quality assurance in pharmaceutical applications.

Implementation Method 1

proteins in a sample undergo proteolytic digestion, breaking the proteins into smaller pieces (peptides)

Methodology Applied
Scientific EffectProteolytic digestion: Hydrolysis

Implementation Method 2

the resulting data is often processed using a search engine in conjunction with a sequence database containing data of known peptides and proteins

Methodology Applied
Scientific EffectMass spectrometry:

Implementation Method 3

generating tandem mass spectrometry data on each digested polypeptide sample

Methodology Applied
Scientific EffectTandem mass spectrometry:

Data Source

PatentUS12061204B2Method to map protein landscapes
Publication Date: 2024.08.13 WISCONSIN ALUMNI RES FOUND
  • US12061204B2 patent drawing
  • US12061204B2 patent drawing

AI summary

In shotgun proteomics, generally only a fraction of peptides from a parent protein are actually detected. Because a large portion of the protein sequence is not detected, it is often impossible to determine whether the expressed protein is present in a modified, spliced, or truncated form. Provided herein are methods and systems for analyzing polypeptides which allow for the increase of the mean sequence coverage of a protein concomitant with bioinformatics analysis in order to distinguish putative proteoforms with improved amino acid resolution. Aspects of the invention include (1) a deep sequencing strategy to provide more protein sequence coverage than is typically achieved, and (2) a computational approach to view protein expression across its full length and identify regions of the protein that are potentially subject to such regulation. This technology has global utility in proteomics and will be of particular use for the analysis of biosimilar protein drug therapeutics.