Combinatorial Barcoding for Low-Abundance Protein Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for protein identification, such as mass spectrometry and peptide sequencing, require large amounts of proteins and cannot detect low-abundance proteins, limiting their effectiveness in analyzing complex protein mixtures.
Innovation Solution
A method involving attaching nucleic acid molecules to target amino acid residues, followed by split-pool barcoding to generate unique barcode sequences, which are then sequenced to determine amino acid residue frequencies and identify proteins in complex mixtures using standard laboratory equipment.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If mass spectrometry or peptide sequencing is used for protein identification, then reliable protein identification can be achieved, but large amounts of proteins are required and low-abundance proteins cannot be detected
Solution Approach 1:
The patent segments the protein identification process into two distinct stages: (1) protein fragmentation into peptides through proteolytic digestion, and (2) peptide sequencing followed by computational reconstruction of the parent protein sequence. This segmentation allows low-abundance proteins to be detected through their peptide fragments, which can be enriched and sequenced with higher sensitivity, while still achieving complete protein identification through bioinformatic assembly.
Solution Approach 2:
The patent introduces peptide sequences as an intermediary between the target protein and direct detection. Instead of directly sequencing or detecting intact low-abundance proteins, the method uses peptide fragments as mediators that can be enriched, sequenced with higher sensitivity, and then computationally reconstructed to identify the parent protein sequence, thereby enabling detection of proteins at very low abundances.
2Quantity of substance
If single-molecule protein sequencing is used, then low-abundance proteins can be detected, but sophisticated instrumentation is required and throughput is limited
Solution Approach 1:
The patent creates multiple copies of peptide sequences through combinatorial pooling and barcoding strategies. By fragmenting proteins into peptides and using combinatorial indexing to generate multiple representative copies of each peptide sequence across different libraries, the method enables detection of low-abundance proteins through amplified signal representation without requiring single-molecule sensitivity or sophisticated instrumentation.
Solution Approach 2:
The patent employs universal proteolytic enzymes (such as trypsin) that can cleave a broad range of proteins at predictable sites, making the method applicable to any protein in the proteome. This universal approach, combined with combinatorial barcoding, allows simultaneous processing and identification of numerous different proteins through a single standardized workflow, achieving high throughput without specialized instrumentation for each protein type.
3Measurement precision
If conventional protein identification methods are used, then protein sequences can be determined, but the process is time-consuming and requires large amounts of protein
Solution Approach 1:
The patent performs preliminary proteolytic digestion of proteins into peptides before sequencing. This preliminary fragmentation action breaks down complex proteins into smaller, more easily sequenced peptide fragments that can be processed in parallel. By pre-digesting the protein sample and using combinatorial barcoding to tag multiple peptides simultaneously, the method accelerates the overall identification process while maintaining accurate sequence determination through subsequent computational assembly.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach allows for the rapid and inexpensive identification of proteins in complex mixtures, overcoming the limitations of existing methods by enabling the analysis of full-length proteins and achieving high throughput without the need for sophisticated instrumentation.
Implementation Method 1
attaching a first nucleic acid molecule to target amino acid residues in the plurality of proteins in the sample to provide nascent nucleic acid tags
Implementation Method 2
ligating a barcode nucleic acid molecule to the first nucleic acid molecules attached to the proteins, wherein the barcode nucleic acid molecule in each partition in the plurality of partitions comprise a barcode sequence unique to that partition
Implementation Method 3
sequencing the mature nucleic acid tags; and determining a frequency of the target amino acid residues
Data Source
AI summary
Methods and kits for assessing and partially or uniquely identifying proteins, such as in complex mixtures of proteins and other molecules/structures, are described. In an embodiment, the methods include combinatorially barcoding target amino acid residues on the protein and, in certain embodiments, enzymatically cleaving the protein, such as before and after rounds of combinatorial barcoding. In an embodiment, the methods include (a) attaching a first nucleic acid molecule to target amino acid residues in the plurality of proteins in the sample to provide nascent nucleic acid tags at the target amino acid residues; (b) performing one or more rounds of split-pool barcoding to provide mature nucleic acid tags at the target amino acid residues (c) sequencing the mature nucleic acid tags; and (d) determining a frequency of the target amino acid residues in one or more protein of the plurality of proteins.


