Single-Molecule Protein Sequencing via ClickP Terminal Capture

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current protein sequencing technologies lack the ability to detect low-copy number proteins with single-molecule sensitivity and provide spatial information, and existing methods are limited by scalability and inefficiencies in identifying terminal amino acids.

Innovation Solution

The use of ClickP compounds to bind to terminal amino acids of peptides, tether them to a substrate, cleave them from the peptide, and detect the resulting ClickP-amino acid complexes, allowing for high-resolution sequencing of peptides and proteins.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If mass spectrometry is used for protein identification and quantification, then attomole detection sensitivity is achieved, but low-copy number proteins remain undetected due to sensitivity limitations

Engineering Contradiction:
Improvedetection sensitivityVSAvoidlow-copy number proteins
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The invention segments the protein sequencing process into single-molecule analysis units, where individual proteins are captured and sequenced separately on a substrate. This segmentation enables detection of low-copy number proteins by analyzing each molecule individually rather than requiring ensemble measurements, thereby achieving single-molecule sensitivity for proteins that were previously undetectable by mass spectrometry.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The invention introduces an intermediary capture system consisting of substrate-bound reagents that specifically bind to target proteins. This intermediary layer enables the detection of low-abundance proteins by concentrating and immobilizing them on the substrate before sequencing, effectively bridging the gap between low protein abundance and detection capability.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If Edman degradation is used for protein sequencing, then 98% sequencing efficiency is achieved, but the method is inherently low throughput and requires highly purified protein

Engineering Contradiction:
Improvesequencing efficiencyVSAvoidthroughput
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The invention segments the sequencing process into parallel single-molecule reactions on a substrate surface. Multiple protein molecules are captured and sequenced simultaneously at different locations on the substrate, enabling high-throughput parallel processing while maintaining the high efficiency of terminal amino acid identification. This eliminates the need for highly purified protein samples required by traditional Edman degradation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The invention creates multiple copies of the sequencing reaction on the substrate surface, where each captured protein molecule undergoes independent sequencing. This copying approach allows simultaneous analysis of many protein molecules, dramatically increasing throughput compared to sequential Edman degradation while maintaining high sequencing efficiency through parallel terminal amino acid identification.

Inventive Principle:
Principle #26Copying

3Loss of information

If immunohistochemistry is used for protein identification, then spatial information is provided, but protein sequence information is excluded and scalability is limited

Engineering Contradiction:
Improvespatial informationVSAvoidscalability to entire proteome
Core Design Contradiction:
Loss of informationVSAdaptability or versatility

Solution Approach 1:

The invention segments the proteome analysis into individual protein molecule events captured at specific spatial locations on a substrate. Each captured protein is sequentially identified through terminal amino acid detection, with spatial coordinates recorded. This segmentation enables both spatial information preservation and scalability to the entire proteome by processing multiple protein types in parallel at different substrate locations.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The invention creates a universal platform that combines the spatial visualization capability of immunohistochemistry with the sequencing capability of Edman degradation. The substrate-based capture system can identify any protein through terminal amino acid detection while maintaining spatial location information, making the method universally applicable to the entire proteome rather than being limited to specific proteins requiring custom antibodies.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Quantity of substance

If ensemble measurements from many cells are used for protein analysis, then protein identification is achieved, but cell-to-cell variations are masked

Engineering Contradiction:
Improveprotein abundanceVSAvoidcell-to-cell variations
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The invention segments the protein analysis from ensemble measurements to single-molecule measurements. By capturing and sequencing individual protein molecules from single cells on a substrate, the method preserves cell-to-cell variations that are masked in ensemble measurements. Each protein molecule is analyzed independently, maintaining information about heterogeneity in protein expression across different cells.

Inventive Principle:
Principle #1Segmentation

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enables high-throughput, single-molecule sequencing of proteins with improved sensitivity and spatial resolution, facilitating ultrasensitive diagnostics and comprehensive proteomic analysis.

Implementation Method 1

contacting the peptide with a ClickP compound, wherein the ClickP compound binds to a terminal amino acid or a terminal amino acid derivative of the peptide to form a ClickP-peptide complex

Methodology Applied
Scientific EffectChemical conjugation: Chemical Bonding

Implementation Method 2

the ClickP-peptide complex is tethered to a physical substrate

Methodology Applied
Scientific EffectCovalent bonding: Chemical Bonding

Implementation Method 3

the ClickP-peptide complex is cleaved from the peptide resulting in a ClickP-amino acid complex

Methodology Applied
Scientific EffectChemical cleavage: Chemical Bonding

Data Source

PatentUS12379380B2Single-molecule protein and peptide sequencing
Publication Date: 2025.08.05 MASSACHUSETTS INST OF TECH
  • US12379380B2 patent drawing
  • US12379380B2 patent drawing
  • US12379380B2 patent drawing

AI summary

The present description provides methods, assays and reagents useful for sequencing proteins. Sequencing proteins in a broad sense involves observing the plausible identity and order of amino acids, which is useful for sequencing single polypeptide molecules or multiple molecules of a single polypeptide. In one aspect, the methods are useful for sequencing multiple polypeptides. The methods and reagents described herein can be useful for high resolution interrogation of the proteome and enabling ultrasensitive diagnostics critical for early detection of diseases.