Single-Molecule Protein Sequencing Using ClickP Terminal Labeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current protein sequencing technologies lack the ability to detect low-copy number proteins, provide single-molecule sensitivity, and offer spatial information, with methods like Mass Spectrometry and Edman degradation being low throughput and lacking scalability, while immunohistochemistry does not provide sequence information.

Innovation Solution

A method using ClickP compounds to bind to terminal amino acids of peptides, tether them to a substrate, cleave them from the peptide, and detect the resulting ClickP-amino acid complexes, enabling high-resolution protein sequencing and identification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If Mass Spectrometry is used for protein identification, then protein quantification is achieved, but single-molecule sensitivity and detection of low-copy number proteins is lost

Engineering Contradiction:
Improveprotein quantification accuracyVSAvoiddetection sensitivity for low-copy proteins
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The method segments the protein sequencing problem into single-molecule analysis units. Individual proteins are immobilized on separate locations of a substrate array, allowing each molecule to be analyzed independently. This segmentation enables detection of single-copy proteins while maintaining quantification accuracy through statistical analysis of multiple replicated measurements across the array.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The invention creates multiple copies of the same protein sample across different locations on a substrate array. Each location contains an immobilized protein molecule that serves as a template for sequencing reactions. These replicated copies enable ensemble averaging and statistical analysis, improving detection sensitivity for low-abundance proteins while maintaining measurement precision.

Inventive Principle:
Principle #26Copying

2Measurement precision

If Edman degradation is used for protein sequencing, then N-terminal amino acid identification is achieved, but high-throughput capability is lost

Engineering Contradiction:
Improveamino acid identification accuracyVSAvoidsequencing throughput
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The sequencing process is segmented across thousands of parallel reaction sites on a substrate array. Each site contains an immobilized protein molecule undergoing Edman degradation independently. This spatial segmentation transforms a sequential single-molecule process into a massively parallel high-throughput system, maintaining amino acid identification accuracy while increasing productivity by orders of magnitude.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The invention transitions Edman degradation from a temporal sequential process to a spatially parallel process by distributing reaction sites across a two-dimensional substrate array. This dimensional transformation allows simultaneous processing of thousands of protein molecules, achieving high throughput without compromising the precision of amino acid identification at each site.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Use of energy by moving object

If Immunohistochemistry is used for protein visualization, then spatial localization is achieved, but sequence information is lost

Engineering Contradiction:
Improvespatial resolution for protein localizationVSAvoidamino acid sequence information
Core Design Contradiction:
Use of energy by moving objectVSLoss of information

Solution Approach 1:

The invention merges the spatial localization capability of immunohistochemistry with the sequence information capability of Edman degradation. Proteins are immobilized on a substrate array with preserved spatial coordinates, then subjected to sequencing reactions. The combination maintains the spatial information (where the protein was located) while adding sequence information (what the protein is), eliminating the information loss of traditional immunohistochemistry.

Inventive Principle:
Principle #5Merging (Combining)

4Reliability

If ensemble measurements from many cells are used for protein analysis, then statistical significance is achieved, but cell-to-cell variations are masked

Engineering Contradiction:
Improvestatistical significance of resultsVSAvoidsingle-cell heterogeneity information
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The method segments the protein analysis into single-cell single-molecule units immobilized on a substrate array. Each immobilized protein represents an individual molecular entity that can be analyzed independently. This segmentation preserves cell-to-cell and molecule-to-molecule heterogeneity information while enabling statistical analysis across thousands of replicated measurements, providing both reliability and resolution of biological variability.

Inventive Principle:
Principle #1Segmentation

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enables high-resolution, single-molecule protein sequencing with spatial information, allowing for ultrasensitive diagnostics and protein expression profiling in complex samples.

Implementation Method 1

contacting the peptide with a ClickP compound, wherein the ClickP compound binds to a terminal amino acid or a terminal amino acid derivative of the peptide to form a ClickP-peptide complex

Methodology Applied
Scientific EffectChemical Bonding: Chemical Bonding

Implementation Method 2

tethering the ClickP-peptide complex to a substrate

Methodology Applied
Scientific EffectChemical Bonding: Chemical Bonding

Implementation Method 3

cleaving the ClickP-peptide complex from the peptide to form a ClickP-amino acid complex

Methodology Applied
Scientific EffectChemical Bonding: Chemical Bonding

Data Source

PatentUS20260016480A1Single-molecule protein and peptide sequencing
Publication Date: 2026.01.15 MASSACHUSETTS INST OF TECH
  • US20260016480A1 patent drawing
  • US20260016480A1 patent drawing
  • US20260016480A1 patent drawing

AI summary

The present description provides methods, assays and reagents useful for sequencing proteins. Sequencing proteins in a broad sense involves observing the plausible identity and order of amino acids, which is useful for sequencing single polypeptide molecules or multiple molecules of a single polypeptide. In one aspect, the methods are useful for sequencing multiple polypeptides. The methods and reagents described herein can be useful for high resolution interrogation of the proteome and enabling ultrasensitive diagnostics critical for early detection of diseases.