Methods for sequencing polypeptides and related compositions
The INDEED method addresses the limitations of existing protein sequencing techniques by using DNA-encoded Edman degradation and binding-dependent primer extension to achieve single-amino acid resolution and high throughput for single-cell protein sequencing.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-23
- Publication Date
- 2026-03-10
AI Technical Summary
Current methods for single-cell protein sequencing lack the generality, sensitivity, and throughput required to accurately sequence proteins and their post-translational modifications (PTMs) at the single-cell level, with existing techniques like nanopore-based sequencing, Edman degradation-based peptide fluorescent fingerprinting, and real-time dynamic protein sequencing facing limitations in charge dependence, labeling efficiency, and sequence interference.
A method involving intramolecular DNA-encoded Edman degradation (INDEED) that labels N-terminal amino acids with unique molecular identifiers and cycle number barcodes, degrades them sequentially, and uses binding-dependent primer extension to generate DNA-barcoded fragments, enabling accurate sequencing of polypeptides through DNA sequencing platforms.
The method achieves generalizable, single-amino acid resolution with single-molecule sensitivity and the throughput necessary for single-cell protein sequencing, overcoming limitations of existing techniques by using DNA barcoding and primer extension to decode peptide sequences in a massively parallel manner.
Smart Images

Figure 2026508216000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 448,131, filed February 24, 2023, which is incorporated herein by reference in its entirety.
[0002] Introduction Over the past decade, advances in DNA sequencing have enabled mRNA sequencing at the single-cell level, revolutionizing our understanding of the heterogeneity of biological systems. The medical impact of these insights ranges from fundamental understanding of early organ development to identifying rare, treatment-resistant cell populations within complex tumors. Historically, researchers have assumed that mRNA and protein expression levels are directly correlated. However, gene-gene correlation analysis has shown that mRNA levels explain only approximately 40% of the variation in protein levels. This is because protein levels are influenced by many factors, including translation rate, translational regulation, and protein degradation. Furthermore, proteins often undergo post-translational modifications (PTMs) after synthesis, such as phosphorylation, glycosylation, and methylation. It is well known that PTMs can dramatically affect protein activity, localization, and interactions with other biomolecules. Importantly, PTMs are regulated by enzymatic processes that are not directly encoded by genes and therefore cannot be predicted from the transcriptome. Therefore, there is an urgent unmet need to go beyond mRNA sequencing and directly sequence proteins (and their PTMs) at the single-cell level.
[0003] Protein sequencing has been performed by mass spectrometry (MS) since 1970, and the sensitivity of MS instruments has increased dramatically over the past 50 years. For example, current advanced MS instruments can detect approximately 1,000 different proteins in 0.8 ng of cell lysate. While this is impressive, it represents only a small fraction of the approximately 20,000 proteins known to exist in cells. Importantly, even with further advances, it is uncertain whether MS will have sufficient dynamic range to achieve single-cell protein sequencing in the future.
[0004] Several techniques have been proposed for single-cell protein sequencing, and their potential can generally be calibrated using the following three indicators: (1) generality, (2) sensitivity, and (3) throughput. Generality relates to whether the method can be used to identify any sequence of amino acids, regardless of the chemical composition of the amino acids, such as charge, hydrophobicity, and length. Generality also relates to whether the method can distinguish between natural and PTM-modified amino acids. Sensitivity relates to whether the method retains the potential to ultimately measure a single amino acid in a single protein. Throughput relates to whether the method retains the potential to ultimately sequence all proteins from a single human cell, where one human cell typically contains 10,000 to 20,000 different types of proteins. 8 ~10 9 It contains protein molecules.
[0005] Current technological developments toward single-molecule protein sequencing can be broadly divided into three categories: 1) nanopore-based peptide fingerprinting, 2) Edman degradation-based peptide fluorescence fingerprinting, and 3) real-time dynamic protein sequencing.
[0006] Nanopore sequencing platforms use either pore-forming proteins or nanofabricated pores to measure changes in electrical current as peptides translocate through the pore, creating a peptide "fingerprint." Theoretically, the strength of nanopore sequencing platforms is their ability to read a wide variety of amino acids with high accuracy. For example, aerolysin nanopores have been shown to distinguish all 20 amino acids when located at the C-terminus of polyarginine peptides. More recently, MspA nanopores have been shown to be capable of fingerprinting peptides carrying single amino acid substitutions. However, in practice, current implementations of nanopore sequencing are not general enough to sequence arbitrary peptides. For example, the aforementioned studies were performed using only highly charged model peptides, while the uneven charge of natural peptides hinders their translocation through the nanopore. In addition, accurate identification of peptide sequences at single-amino acid resolution by nanopores is extremely challenging, as up to eight amino acids can contribute to changes in ionic current as the peptide passes through the nanopore. Importantly, perhaps the biggest drawback of nanopore-based sequencing is its throughput: for example, it can take 30 minutes to measure a single 25 amino acid peptide. 8 Considering that it contains ~10 peptides, it is uncertain whether nanopore sequencing can reach the throughput required for single-cell proteomics, even with massively parallel operation.
[0007] The second strategy, called Edman degradation-based peptide fluorescent fingerprinting, combines Edman chemistry with single-molecule microscopy. In this approach, cysteine and lysine side chains are fluorescently labeled, and the peptide is immobilized on a solid support via its C-terminus. The N-terminal amino acids are then sequentially removed, one residue at a time, by Edman degradation. Digestion of the fluorescently labeled amino acids causes a decrease in fluorescence intensity, which serves as a unique fingerprint for a given peptide. The strength of this method is its potential to fingerprint millions of peptides in parallel, and thanks to the robustness of Edman degradation, this approach is generalizable to most peptide sequences. The main weakness of this method is that spectrally distinguishable fluorophores serve as amino acid surrogates; therefore, specific labeling of amino acids with high efficiency is essential for this technique. Unfortunately, fluorescent labeling of only lysine and cysteine has been demonstrated to date. To date, only a small set of amino acids (e.g., lysine, cysteine, and tyrosine) can be labeled with sufficient specificity and efficiency required by this technique. Furthermore, photobleaching, chemical degradation of fluorophores, and energy transfer between fluorophores (in the context of peptides) are all likely to limit the sensitivity and generalizability of this technique.
[0008] Finally, real-time dynamic protein sequencing utilizes the continuous degradation of surface-immobilized peptides driven by aminopeptidases. During degradation, newly generated N-terminal amino acids are recognized in real time by a mixture of dye-labeled N-terminal amino acid binders evolved from the adaptor protein ClpS. N-terminal amino acids are identified not only based on binding affinity but also on binding kinetics, allowing for the identification of multiple amino acids with a single binder protein. The advantage of this method is that it is not limited by the charge state of the peptide or the chemical functionality of the amino acid side chain, and therefore has the potential to be generalizable to most peptide sequences. In addition, the use of N-terminal amino acid binders greatly expands the number of sequenceable amino acids. Previously, the ClpS protein used in real-time dynamic protein sequencing has been shown to distinguish between seven different N-terminal amino acids. However, a major weakness of this strategy is that ClpS binding to N-terminal proteins is affected by neighboring amino acids, which is a significant problem. Recently, it has been shown that the affinity and binding kinetics of ClpS proteins vary dramatically even within a small subset of possible downstream sequences. Therefore, the feasibility of this methodology depends on the availability of novel reagents whose binding affinity and kinetics are independent of the amino acid attached to the N-terminal amino acid. Unfortunately, such reagents have not been discovered to date. In addition, the enzymatic digestion used in real-time dynamic protein sequencing does not proceed stepwise. This leads to inconsistent lifetimes of ClpS-N-terminal amino acid complexes and an inability to measure the length of unspecified peptide segments, both of which limit the accuracy of this technique.
[0009] For the reasons set forth above, current techniques do not provide the generality, sensitivity, and throughput to ultimately reach the goal of achieving single-cell protein sequencing. Summary of the Invention
[0010] Methods for sequencing a polypeptide are provided. In certain embodiments, the methods include labeling the N-terminal amino acid of the polypeptide with a nucleic acid label comprising a unique molecular identifier and a cycle number barcode; degrading the N-terminal amino acid from the polypeptide; annealing a primer to the nucleic acid label of the degraded amino acid, where the primer comprises a barcode corresponding to the identity of the degraded N-terminal amino acid; and extending the primer annealed to the nucleic acid label to produce extension products comprising the unique molecular identifier, the cycle number barcode, and the barcode corresponding to the identity of the degraded N-terminal amino acid. The foregoing steps are performed sequentially to produce multiple such extension products, which are then sequenced, allowing the amino acid sequence of the polypeptide to be determined. Compositions and kits for use in carrying out the methods are also provided. [Brief explanation of the drawings]
[0011] [Figure 1A] Overview of single molecule polypeptide sequencing according to an embodiment of the present disclosure. In this example, the approach combines the chemistry of Edman degradation with the massively parallel nature of DNA sequencing-by-synthesis technology. [Figure 1B] Schematic of an embodiment in which a polypeptide is labeled with a nucleic acid label comprising a unique molecular identifier (UMI) and a cycle number barcode. [Figure 1C] Schematic of an embodiment in which biotinylated primers are used, allowing barcoded DNA to be pulled down by streptavidin beads on which binding-dependent primer extension occurs. In this example, the primer extension products are also indexed with cycle number barcodes. [Figure 2A]Figure 1 shows a reaction scheme for intramolecular DNA-encoded Edman degradation (INDEED) according to some embodiments of the present disclosure. In this example, a polypeptide is conjugated to a DNA unique molecular identifier (UMI)-functionalized bead (Step 1). A primer used to record the UMI is conjugated to the N-terminus of the polypeptide via a modified PTC and click reaction (Steps 2 and 3). The DNA UMI is then transcribed via a primer extension reaction (Step 4). Finally, the PTC amino acid is generated by cleavage and hydrolysis (Step 5). The remaining immobilized polypeptide is subjected to the next cycle of INDEED. [Figure 2B] Reaction scheme for binding-dependent primer extension (BD-PEX). PTC amino acids are recognized by a primer and an antibody tagged with an antibody-specific barcode (Step 1). The binding event is recorded by primer extension. The resulting DNA is sequenced to reveal the polypeptide sequence (Step 2). [Figure 2C] Figure 1 shows the reaction scheme for INDEED, according to some embodiments of the present disclosure. In this example, the C-terminal side of a peptide is conjugated to a DNA UMI-functionalized bead (Step 1). The N-terminus of the peptide is then derivatized with an azide-modified PITC, and a biotinylated primer is incorporated via a proximity-promoted click reaction (Steps 2 and 3). In the fourth step, after primer extension, the UMI is copied onto the nascent DNA strand (Step 4), and finally, the DNA-barcoded PTC amino acid is released from the solid support by Lewis acid-catalyzed cleavage followed by hydrolysis under reducing basic conditions (Steps 5 and 6). [Figure 2D]Scheme of antibody-assisted proximity extension follows that described in Figure 2C. DNA barcoded PTC amino acids are pulled down by streptavidin beads. Primers and antibodies tagged with antibody-specific barcodes are introduced, and recognition events are recorded by primer extension. The resulting DNA is converted into a sequencing library, during which cycle number barcodes are introduced. This process generates DNA containing information on the position, origin, and identity of amino acids, which can be read by DNA sequencing. [Figure 3-1] a) DNA modified with 7-deazapurine nucleotides is stable under Lewis acid-catalyzed Edman degradation conditions. b) Reaction scheme for the Lewis acid-catalyzed Edman degradation reaction. [Figure 3-2] c) MS spectrum of an oligonucleotide containing 7-deazapurine deoxynucleotides after treatment with 40 mM BF3 etherate. [Figure 4] a) Preparation of solid-phase immobilized DNA-peptide conjugate. b) Mass spectrum of PTC-tryptophan in the supernatant released by Edman degradation on solid support. The inset shows the chromatogram of the supernatant. c) Edman degradation of CPG. DNA is cleaved from CPG with ammonia and analyzed by HPLC. The PTC-peptide-DNA conjugate (0 min, right trace) was completely degraded to give the peptide-DNA conjugate (10 min, left trace). [Figure 5-1] a) Synthesis of alkyne-modified PITC derivative 2. b) Preparation of DNA-peptide conjugates on beads and DNA barcoding of the N-terminal amino acid by INDEED. [Figure 5-2]c) Quantification of degradation yield by flow cytometry. Primers are hybridized with FAM-labeled complementary strands and quantified using a flow cytometer. The decrease in fluorescence intensity after degradation is used to calculate the degradation yield. d) Screening of polymerases for primer extension. Purine nucleotides in the template and primers are completely replaced by 7-deaza dA and 7-deaza dG. The dNTP mix contains dTTP, dCTP, 7-deaza dATP, and 7-deaza dGTP. [Figure 5-3] e), f) An embodiment of INDEED where magnetic beads are used and primers for barcode transfer are introduced via proximity-facilitated SPAAC. [Figure 5-4] g) LC-MS data of DNA-barcoded PTC amino acids generated by INDEED. h) Flow cytometry plots showing the step-by-step degradation yields and overall yield of the INDEED process over five degradation cycles. [Figure 6-1] a) Binding of a commercially available antibody to DNA-conjugated PTC amino acids. b) Binding of a phosphotyrosine-specific antibody (PY20) to DNA-conjugated PTC-phosphotyrosine. c) Specificity of antibodies produced by hybridomas raised against PTC-tyrosine-conjugated BSA, characterized by ELISA. d) Specificity of antibodies produced by hybridomas raised against PTC-phenylalanine-conjugated BSA, characterized by ELISA. e) Data showing antibody recognition of PTC to asymmetric dimethylarginine (ADMA), acetyllysine, and phosphoserine. [Figure 6-2] f) Data showing the identification of antibodies against PTC-modified phenylalanine, tyrosine, tryptophan, arginine, and aspartic acid. [Figure 7-1] Conversion of PTC amino acids to DNA by binding-dependent primer extension (BD-PEX), according to some embodiments of the present disclosure. [Figure 7-2]Conversion of PTC amino acids to DNA by BD-PEX according to some embodiments of the present disclosure. [Figure 7-3] c) PAGE analysis of antibody-assisted proximity extension. Primers with a length of 5 nt and 6 nt can distinguish the presence of antibody-antigen interactions. d) Quantification of DNA output of antibody-assisted proximity extension by qPCR. Primers with a length of 6 nt gave quantitative DNA output and were selected for future experiments. e) Reducing the loading density on the beads suppressed the occurrence of undesired intermolecular primer extension. [Figure 8A] Peptide sequencing by DNA sequencing. Scheme for sequencing a single peptide species with multiple readable amino acids. The peptide was subjected to five cycles of the INDEED process. The DNA-barcoded PTC amino acids obtained from each cycle were pulled down onto streptavidin beads in separate containers. Proximity primer extension was performed using a mixture of antibodies specific to the DNA-barcoded PTC amino acids (100 nM each of PTC-Arg, PTC-Phe, PTC-Asp, and PTC-Trp antibodies). After primer extension, adapter PCR was performed in each container using an adapter primer carrying the cycle number barcode. Finally, all DNA was pooled, indexed, and sequenced on a MiSeq sequencer. [Figure 8B] Heatmap summarizing read count DNA barcodes obtained for the peptide sequence RGFDW. [Figure 9]Single-cell proteoform mapping workflow. a) Single-cell peptide extraction. Single cells are isolated via FACS in multiwell plates containing lysis buffer. Proteins of interest are pulled down by antibody-coated beads. These proteins are then digested with trypsin. b) Peptides are conjugated to DNA UMIs that also contain barcodes specific to each well. The barcode peptides are sequenced by INDEED. c) The distribution of proteoforms (e.g., phosphorylation) can be mapped in each cell at single amino acid resolution. DETAILED DESCRIPTION OF THE INVENTION
[0012] Before describing the methods, compositions, and kits of the present disclosure in more detail, it is to be understood that the methods, compositions, and kits are not limited to particular embodiments described, as such may, of course, vary. Moreover, it is to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the methods, compositions, and kits will be limited only by the appended claims.
[0013] Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit, between the upper and lower limits of that range, and any other stated value or intervening value within that stated range, is encompassed within the methods, compositions, and kits, unless the context clearly dictates otherwise. The upper and lower limits of these smaller ranges may independently be included within the smaller ranges and are also encompassed within the methods, compositions, and kits, subject to any specifically excluded limit in the stated range. When a stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the methods, compositions, and kits.
[0014] Certain ranges are presented herein with numerical values preceded by the term "about." The term "about" is used herein to provide literal support for the exact number it precedes, as well as a number that is near or approximately the number it precedes. In determining whether a number is near or approximately a specifically recited number, the number that is near or approximately the recited number may be a number that, in the context provided, provides substantial equivalence to the specifically recited number.
[0015] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the method, composition, and kit belongs. Although any methods, compositions, and kits similar or equivalent to those described herein can also be used in the practice or testing of the present methods, compositions, and kits, representative illustrative methods, compositions, and kits are described here.
[0016] All publications and patents cited herein are incorporated by reference to disclose and describe in connection with the materials and / or methods for which the publications are cited, as if each individual publication or patent was specifically and individually indicated to be incorporated by reference. The citation of any publication is for its disclosure prior to the filing date, and the dates of publication provided may be different from the actual publication dates which may need to be independently confirmed, and should not be construed as an admission that the methods, compositions, and kits are not entitled to antedate such publication.
[0017] It should be noted that, as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. It should be further noted that the claims may be drafted to exclude any optional element. Accordingly, this statement is intended to serve as a predicate for use of exclusive terminology such as "solely," "only," and the like, or for use of a "negative" limitation in connection with the recitation of claim elements.
[0018] For clarity, it is understood that certain features of the methods, compositions, and kits that are described in the context of separate embodiments may also be provided in combination in a single embodiment. Conversely, for brevity, various features of the methods, compositions, and kits that are described in the context of a single embodiment may also be provided separately or in any suitable subcombination. All combinations of embodiments are specifically encompassed by the present disclosure, and to the extent such combinations encompass operable processes and / or compositions, each and every combination is disclosed herein to the same extent as if individually and explicitly disclosed. In addition, all subcombinations listed in embodiments describing such variables are also specifically encompassed by the present methods, compositions, and kits, and each and every such subcombination is disclosed herein to the same extent as if individually and explicitly disclosed herein.
[0019] As will be apparent to those skilled in the art upon reading this disclosure, each of the individual embodiments described and illustrated herein has individual elements and features which may be readily separated from or combined with the features of any of the other several embodiments without departing from the scope or spirit of the method. Any recited method may be carried out in the order of events recited or in any other order which is logically possible.
[0020] Methods for sequencing polypeptides Aspects of the present disclosure include methods for sequencing a polypeptide. According to some embodiments, the method includes labeling the N-terminal amino acid of the polypeptide with a nucleic acid label comprising a unique molecular identifier (UMI) and degrading the N-terminal amino acid from the polypeptide. In certain embodiments, such methods further include annealing a primer to the nucleic acid label of the degraded N-terminal amino acid, the primer comprising a barcode corresponding to the identity of the degraded N-terminal amino acid. In some cases, such methods further include extending the primer annealed to the nucleic acid label to produce extension products comprising the UMI and a barcode corresponding to the identity of the degraded N-terminal amino acid. The foregoing steps may be performed in successive cycles to produce multiple extension products, each of which comprises a UMI, a respective cycle number barcode, and a barcode corresponding to the identity of the degraded N-terminal amino acid. According to some embodiments, such methods further include sequencing the multiple extension products and determining the sequence of the polypeptide based on the sequences of the multiple extension products. In certain embodiments, the method comprises indexing the extension product produced in the extension step with a cycle number barcode, hi other embodiments, the labeling step comprises labeling the N-terminal amino acid of the polypeptide with a nucleic acid label comprising a UMI and a cycle number barcode.
[0021] The disclosed method addresses the shortcomings of nanopore-based polypeptide sequencing, Edman degradation-based peptide fluorescent fingerprinting, and real-time dynamic protein sequencing approaches and constitutes an improvement over them. For example, nanopore-based polypeptide sequencing requires proteins to be charged to enable translocation, and also suffers from low throughput. For Edman degradation-based peptide fluorescent fingerprinting, only fluorescent labeling of lysines and cysteines has been demonstrated to date. For real-time dynamic protein sequencing, shortcomings include the influence of protein sequence on recognition molecule binding.
[0022] An overview of an embodiment of the disclosed method is illustrated schematically in Figure 1. In this example, Edman degradation is performed to generate "degradation fragments," which are then labeled with DNA barcodes that encode the amino acid's origin (i.e., the polypeptide from which it is derived) and its position within the polypeptide. These fragments are specifically recognized by a binding agent (e.g., an antibody, small molecule, aptamer, etc.) without interference with their original downstream polypeptide sequence. Binding of the binding agent allows the DNA barcodes to be linked to amino acids, which allows the peptide sequence to be decoded in a massively parallel manner using a nucleic acid sequencer, e.g., an Illumina-type or other suitable DNA sequencer.
[0023] Thus, the disclosed method (embodiments of which are sometimes referred to herein as intramolecular DNA-encoded Edman degradation (or "INDEED")) comprises two key process modules: the first module degrades terminal amino acids to generate DNA-barcoded amino acids (e.g., phenylthiocarbamyl (PTC)-amino acids); and the second module identifies amino acids by adjacent primer extension and reads the polypeptide sequence by DNA sequencing.
[0024] A non-limiting example of the first module according to an embodiment of the present disclosure is illustrated schematically in FIG. 2A. In the first step, a polypeptide is immobilized on a solid support (e.g., beads or other suitable solid support). This can be achieved by conjugating the C-terminal region of the polypeptide to a DNA unique molecular identifier (UMI)-functionalized bead so that each peptide connects to a unique DNA sequence (FIG. 2A, Step 1). Then, in the following two steps, a primer used to record the UMI is conjugated to the N-terminus of the polypeptide (FIG. 2A, Steps 2 and 3). This can be done by reacting the polypeptide with a modified isothiocyanate bearing a click handle (e.g., modified PITC). A primer containing a barcode for the cycle number is then installed via click chemistry. In the fourth step, the DNA UMI is transcribed by a proximity primer extension reaction (FIG. 2A, Step 4). This step transfers the parent polypeptide information to the Edman degradation fragment. In the final step, the PTC amino acid is cleaved from the polypeptide (FIG. 2A, Step 5). This step can be accomplished by a tandem cleavage-hydrolysis reaction that generates DNA-barcoded PTC amino acid fragments containing the cycle number and information of the parent polypeptide. This process is made possible by a modified Edman degradation and the use of unnatural nucleotides, which allows for the preservation of nucleic acids under the harsh conditions of polypeptide sequencing.
[0025] A non-limiting example of the second module according to an embodiment of the present disclosure is illustrated schematically in FIG. 2B. In this example, the second module for reading out a polypeptide sequence by DNA sequencing is performed in two steps. First, the identity information of amino acid fragments is converted into a specific DNA sequence. To do so, PTC fragments are recognized by their corresponding binders (e.g., antibodies, small molecules, aptamers, etc.) conjugated to primers containing binder-specific barcode sequences. Binding events are recorded by binding-dependent primer extension (BD-PEX) (FIG. 2B, step 1). This process results in a DNA duplex containing information about the parent polypeptide and the order and identity of the amino acids. Second, the polypeptide sequence is read out by DNA sequencing (FIG. 2B, step 2). This is done by combining DNAs encoding the polypeptide sequences generated by successive cycles and sequencing these DNAs on a sequencing platform. The resulting DNA sequencing data is used to reconstruct the polypeptide sequence. This is achieved by attributing DNA carrying the same UMI to a single parent polypeptide and assigning amino acid order and identity using cycle number barcodes and binder-specific barcodes.
[0026] Further non-limiting examples of modules according to embodiments of the present disclosure are illustrated schematically in Figures 2C-2D.
[0027] The advantages of the disclosed method over existing methods for polypeptide sequencing are numerous. First, the method is general. That is, it can perform the well-established Edman degradation, which is compatible with polypeptide sequences of various charges and lengths. The method directly detects degradation fragments extracted from their sequence structure and recognized by a binding agent (e.g., an antibody). Affinity-based detection is independent of the amino acid's chemical properties and is generalizable to all proteinogenic amino acids and their post-translationally modified forms. Second, the method enables polypeptide sequencing at single-amino acid resolution with single-molecule sensitivity. The method removes one amino acid from the N-terminus in each cycle and barcodes the resulting fragments with DNA encoding the amino acid's origin and position. DNA-barcoded PTC amino acid readout, without interference with downstream polypeptide sequences by BD-PEX, allows polypeptide sequencing at single-amino acid resolution. The resulting DNA sequence is amplified during DNA sequencing, thereby enabling polypeptide sequencing with single-molecule sensitivity. Third, the method enables the throughput required for single-cell protein sequencing. This method achieves highly parallel Edman degradation through DNA barcoding to convert polypeptide sequence information into DNA sequences. Current DNA sequencing technologies (e.g., using Illumina's NovaSeq® sequencing platform) require 10 nucleotides per run. 10 Reads exceeding 10 ...
[0028] The terms "polypeptide," "peptide," and "protein" are used interchangeably herein to designate a linear series of amino acid residues connected to one another by peptide bonds between the alpha-amino and carboxy groups of adjacent residues. The amino acids may include the 20 "standard" genetically encodable amino acids, unnatural amino acids (e.g., amino acid analogs), or combinations thereof.
[0029] The term "amino acid" generally refers to any monomeric unit comprising a substituted or unsubstituted amino group, a substituted or unsubstituted carboxy group, and one or more side chains or groups, or analogs of any of these groups. Exemplary side chains include, for example, thiol, seleno, sulfonyl, alkyl, aryl, acyl, keto, azido, hydroxyl, hydrazine, cyano, halo, hydrazide, alkenyl, alkynyl, ether, borate, boronate, phospho, phosphono, phosphine, heterocyclic, enone, imine, aldehyde, ester, thioacid, hydroxylamine, or any combination of these groups. Naturally occurring α-amino acids are those encoded by the genetic code, as well as those amino acids that are later modified (e.g., hydroxyproline, γ-carboxyglutamate, and O-phosphoserine). Naturally occurring α-amino acids include, but are not limited to, alanine (Ala), cysteine (Cys), aspartic acid (Asp), glutamic acid (Glu), phenylalanine (Phe), glycine (Gly), histidine (His), isoleucine (Ile), arginine (Arg), lysine (Lys), leucine (Leu), methionine (Met), asparagine (Asn), proline (Pro), glutamine (Gln), serine (Ser), threonine (Thr), valine (Val), tryptophan (Trp), tyrosine (Tyr), and combinations thereof. Amino acids may be referred to herein by either their commonly known three-letter symbols or by the one-letter symbols recommended by the IUPAC-IUB Commission on Biochemical Nomenclature. For example, L-amino acids may be represented herein by their commonly known three letter symbols (e.g., Arg for L-arginine) or by the uppercase single letter amino acid symbol (e.g., R for L-arginine). D-amino acids may be represented herein by their commonly known three letter symbols (e.g., D-Arg for D-arginine) or by the lowercase single letter amino acid symbol (e.g., r for D-arginine).
[0030] Polypeptides may be present in any sample of interest, including, but not limited to, protein samples isolated from a single cell, multiple cells (e.g., cultured cells), tissue, biological fluid (e.g., whole blood or portions thereof, urine, saliva, cerebrospinal fluid, sputum, etc.), organ, or organism (e.g., bacteria, yeast, etc.). In certain embodiments, the protein sample is isolated from cells, tissues, organs, etc. of a mammal (e.g., a human, a rodent (e.g., a mouse), or any other mammal of interest). In other embodiments, the protein sample is isolated from a source other than a mammal, for example, a bacterium, yeast, insect (e.g., Drosophila), amphibian (e.g., frog (e.g., Xenopus laevis)), virus, plant, or any other non-mammalian protein sample source.
[0031] In certain embodiments, the polypeptide to be sequenced is present in a protein sample isolated from a single cell. In some such cases, the method is a single-cell protein sequencing method performed on multiple polypeptides present in a protein sample isolated from a single cell.
[0032] Approaches, reagents, and kits for isolating proteins from single cells, cell populations, tissues, etc. are known in the art. Non-limiting examples of available protein extraction kits include ReadyPrep™ Protein Extraction Kit (Bio-Rad), Qproteome Protein Isolation Kit (Qiagen), T-PER™ Tissue Protein Extraction Reagent (Thermo Scientific), M-PER™ Mammalian Protein Extraction Reagent (Thermo Scientific), B-PER™ Complete Bacterial Protein Extraction Reagent (Thermo Scientific), Pierce Plant Total Protein Extraction Kit (Thermo Scientific), RIPA buffer with Triton™ X-100 (5X) (Thermo Scientific), etc.
[0033] Protein samples used in the methods of the present disclosure can be collected by any convenient means. In some cases, useful cell samples can be or be derived from biopsies. Biopsy tissue can be obtained from healthy tissue or diseased cells or tissue, including, for example, cancer cells or tissue. Thus, in some embodiments, the biopsy sample is a tumor biopsy sample. Depending on the type of cancer and / or the type of biopsy performed, the sample can be prepared from a solid tissue biopsy or a liquid biopsy.
[0034] In some cases, protein samples may be prepared from surgical biopsies. Any convenient and suitable technique for surgical biopsy may be utilized for collection of samples used in the methods described herein, including, but not limited to, excision biopsy, incisional biopsy, wire localization biopsy, etc. In some instances, a surgical biopsy may be obtained as part of a surgical procedure that has a primary purpose other than obtaining a sample, including, but not limited to, tumor resection, mastectomy, lymph node surgery, axillary lymph node dissection, sentinel lymph node surgery, etc.
[0035] Biopsy tissue can be obtained using various other biopsy techniques for use as the protein sample described herein.As a non-limiting example, sample can be obtained by needle biopsy.Any convenient and suitable technique for needle biopsy can be used to collect sample, including but not limited to, fine needle aspiration (FNA), core needle biopsy, stereotactic core biopsy, vacuum-assisted biopsy, etc.
[0036] According to an embodiment of the disclosed polypeptide sequencing method, the method comprises labeling the N-terminal amino acid of the polypeptide with a nucleic acid label comprising a unique molecular identifier (UMI) and a cycle number barcode. As used herein in the context of the structure of a polypeptide, the terms "N-terminal amino acid" and "C-terminal amino acid" refer to the amino acid- and carboxyl-terminal amino acids of a polypeptide, respectively.
[0037] As used herein, the term "unique molecular identifier (UMI)" or "UMI" refers to a sequence of nucleotides that can be used to identify and / or distinguish a first molecule to which the UMI is attached from one or more second molecules. As used herein, a UMI can include one or more nucleotides at one or both ends of the sequence that identify / distinguish nucleotides, for example, to facilitate joining (e.g., ligation) of different entities of the UMI. UMIs are typically short, e.g., about 5-40 (e.g., about 5-20) bases in length. Generally, UMIs are used to distinguish similar types of molecules within a population or group.
[0038] As used herein, "barcode" or "barcode sequence" refers to a uniquely identifiable nucleotide sequence. In some embodiments, the barcode uniquely identifies a degradation cycle number (cycle number barcode). Barcode sequences can vary widely in length and composition. According to some embodiments, the barcode has a degenerate sequence of 4 to 120 nucleotides in length, e.g., 4 to 100, 4 to 80, 4 to 60, 4 to 40, 6 to 30, 8 to 20, or 10 to 15 nucleotides in length. In certain embodiments, the barcode has a degenerate sequence of up to 20 nucleotides in length, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides in length. The barcode may contain one or more mixed bases (e.g., every third base, every fourth base, etc.) of only three possible base combinations instead of four bases to prevent homopolymer barcodes.
[0039] In certain embodiments, prior to the labeling step, the polypeptide is immobilized to the solid support via a nucleic acid attached to the C-terminus of the polypeptide and the surface of the solid support, the nucleic acid comprising a UMI and a primer binding site 3' or 5' to the UMI. According to such embodiments, labeling may comprise conjugating a degradation moiety to the N-terminal amino acid and conjugating a primer to the degradation moiety, the primer conjugated to the degradation moiety comprising a cycle number barcode and a sequence 3' to the cycle number barcode that is complementary to the primer binding site of the nucleic acid immobilizing the polypeptide to the solid support. Such labeling may further comprise annealing the primer conjugated to the degradation moiety to the primer binding site and extending the primer conjugated to the degradation moiety using the nucleic acid immobilizing the polypeptide to the solid support as a template, thereby labeling the N-terminal amino acid with a nucleic acid label comprising the UMI and the cycle number barcode.
[0040] As described above, in some embodiments, a polypeptide is immobilized to a solid support via a nucleic acid attached to the C-terminus of the polypeptide and the surface of the solid support, the nucleic acid comprising a UMI and a primer binding site 3' or 5' to the UMI. According to such embodiments, labeling may comprise conjugating to the N-terminal amino acid a degradation moiety conjugated to a primer comprising a cycle number barcode and a sequence 3' to the cycle number barcode that is complementary to the primer binding site of the nucleic acid immobilizing the polypeptide to the solid support. Such labeling may further comprise annealing the primer conjugated to the degradation moiety to the primer binding site and extending the primer conjugated to the degradation moiety using the nucleic acid immobilizing the polypeptide to the solid support as a template, thereby labeling the N-terminal amino acid with a nucleic acid label comprising the UMI and the cycle number barcode.
[0041] The term "solid support" refers to an insoluble material having a surface to which reagents and / or materials (e.g., polypeptides) can be directly or indirectly attached. In certain embodiments, a collection of solid supports has average largest dimensions of 750 μm or less, 500 μm or less, 250 μm or less, 100 μm or less, 1 μm or less, 0.75 μm or less, 0.50 μm or less, 0.25 μm or less, or 0.1 μm or less.
[0042] A variety of materials can be used as solid supports, including any material that can act as a support for attachment of reagents and / or materials. Suitable materials include, but are not limited to, organic or inorganic polymers, natural and synthetic polymers (including, but not limited to, agarose, cellulose, nitrocellulose, cellulose acetate, other cellulose derivatives, dextran, dextran derivatives and dextran copolymers), other polysaccharides, glass, silica gel, gelatin, polyvinylpyrrolidone, rayon, nylon, polyethylene, polypropylene, polybutylene, polycarbonate, polyester, polyamide, vinyl polymers, polyvinyl alcohol, polystyrene and polystyrene copolymers, polystyrene crosslinked with divinylbenzene and the like, acrylic resins, acrylates and acrylic acid, acrylamide, polyacrylamide, polyacrylamide blends, copolymers of vinyl and acrylamide, methacrylates, methacrylate derivatives and copolymers, other polymers and copolymers with various functional groups, latex, butyl rubber and other synthetic rubbers, silicone, glass, paper, natural sponges, insoluble proteins, surfactants, metals, metalloids, magnetic materials, and any combination thereof.
[0043] The solid support may be of any suitable shape, including, but not limited to, a sphere, a spherical, a rod-shaped, a disk-shaped, a pyramidal, a cube-shaped, a cylindrical, a nanohelical, a nanospring, a nanoring, an arrow-shaped, a teardrop-shaped, a tetrapod-shaped, a prism-shaped, or any other suitable geometric or non-geometric shape.
[0044] In certain embodiments, the solid support is a bead. As used herein, the term "bead" refers to a small mass that is generally spherical or globular. According to some embodiments, the beads used herein have an average diameter of about 0.50 μm to about 500 μm, e.g., about 0.75 μm to about 250 μm, e.g., about 1 μm.
[0045] Additionally, and for purposes herein, a solid support may be magnetically responsive by including one or more paramagnetic and / or superparamagnetic materials, such as, for example, magnetite. Such paramagnetic and / or superparamagnetic materials may be embedded within the matrix of the solid support and / or disposed on the exterior and / or interior surfaces of the solid support (e.g., beads).
[0046] Various approaches can be used to immobilize polypeptides to solid supports, including those described in detail in the experimental section herein. For example, in some cases, dibenzocyclooctyne (DBCO)-modified DNA sequences are synthesized on controlled-pore glass (CPG) or polystyrene-coated carboxylic acid magnetic beads. CPG is commonly used for solid-phase DNA synthesis, and magnetic beads are generally enzyme-compatible and allow for easy separation. In one non-limiting embodiment, a PTC peptide containing a C-terminal azidolysine can be conjugated to DBCO-modified DNA via strain-promoted alkyne-azide cycloaddition (SPAAC) to form a model DNA-peptide conjugate. See, for example, Figure 4A.
[0047] As shown in the Experimental Section of this specification, the present inventors have determined that Edman degradation can be performed on immobilized DNA-peptide conjugates. However, because DNA is unstable under conventional Edman degradation conditions, alternative Edman degradation reaction conditions compatible with DNA were required. Suitable alternative Edman degradation reaction conditions identified by the present inventors include, but are not limited to, Lewis acids in aprotic solvents. In some cases, the Lewis acid is BF etherate, BCl 3 , BBr 3 , scandium(III) triflate, or any combination thereof. For example, the Lewis acid can include or consist of BF 3 etherate. According to some embodiments, the aprotic solvent is acetonitrile, N,N-dimethylformamide (DMF), dimethyl sulfoxide (DMSO), or any combination thereof. In one non-limiting example, the alternative Edman degradation reaction conditions include BF 3 etherate in an aprotic solvent, e.g., 40 mM BF 3 etherate in anhydrous acetonitrile. Suitable alternatives further include, for example, triethylamine acetate in dimethylformamide (DMF) at 70°C.
[0048] To further enhance the stability of the nucleic acids implemented in the present methods, one or more stability-enhancing non-natural nucleotides may be used in any of the DNAs utilized in the present methods. According to some embodiments, one or more of the DNAs utilized in the present methods comprises one or more thermostability-increasing nucleotides. Non-limiting examples of thermostability-increasing nucleotides include 7-deaza-8-aza-purine-triphosphate, 2-amino-2'-deoxyadenosine-5'-triphosphate (2-amino-dATP), 5-methyl-2'-deoxycytidine-5'-triphosphate (5-Me-dCTP), 5-propynyl-2'-deoxycytidine-5'-triphosphate (5-Pr-dCTP), 5-propynyl-2'-deoxyuridine-5'-triphosphate (5-Pr-dUTP), and / or halogenated deoxy-uridines (XdU), such as 5-chloro-2'-deoxyuridine-5'-triphosphate (5-Cl-dUTP), 5-bromo-2'-deoxyuridine-5'-triphosphate (5-Br-dUTP), and any combination thereof. In certain embodiments, one or more nucleic acids comprise a non-natural nucleotide that stabilizes the nucleic acid during the degradation step, a non-limiting example of which is a 7-deazapurine nucleotide. For example, the present disclosure surprisingly demonstrates that polymerases (e.g., Sequenase version 2.0, Klenow (exo-), and Bst3.0) can accept 7-deazapurine nucleotide-substituted template-primer duplexes and 7-deazapurine nucleoside triphosphates as substrates. See, e.g., Example 3 in the Examples below and Figure 5D.
[0049] In certain embodiments, the method includes conjugating a degradation moiety to the N-terminal amino acid. A "degradation moiety" refers to a moiety that, when conjugated to the N-terminal amino acid of a polypeptide, facilitates cleavage of the N-terminal amino acid from the polypeptide under conditions compatible with the degradation moiety. In one non-limiting example, the degradation moiety used is an isothiocyanate (ITC). Non-limiting examples of ITCs that can be used as degradation moieties when practicing the methods of the present disclosure include phenyl isothiocyanate (PITC), substituted phenyl isothiocyanates (e.g., para-substituted phenyl isothiocyanate, ortho-substituted phenyl isothiocyanate, meta-substituted phenyl isothiocyanate, pentafluorophenyl isothiocyanate, etc.), alkyl isothiocyanates, naphthalenyl isothiocyanates, etc. The terms "conjugation" or "conjugating" generally refer to a chemical bond, either covalent or non-covalent, usually a covalent bond, that brings one molecule of interest into close proximity with a second molecule of interest.
[0050] According to some embodiments, the method comprises conjugating a primer to a degradation moiety, wherein the primer conjugated to the degradation moiety comprises a cycle number barcode and a sequence 3' to the cycle number barcode that is complementary to the primer binding site of the nucleic acid that immobilizes the polypeptide to a solid support. Various approaches can be used to conjugate the primer to the degradation moiety. In one non-limiting example, the degradation moiety comprises an isothiocyanate (e.g., PITC) that carries a reactive group for conjugation to a nucleic acid that comprises a cycle number barcode. In certain embodiments, the reactive group is a click chemistry reactive group.
[0051] Click chemistry reactions that can be used include (i) nucleophilic substitution, (ii) addition to C-C multiple bonds (e.g., Michael addition, epoxidation, dihydroxylation aziridination), (iii) non-aldol-like chemistry (e.g., N-hydroxysuccinimide active ester coupling), and (iv) cycloaddition (e.g., Diels-Alder reaction, Huisgen cycloaddition). The Huisgen cycloaddition has been applied in various fields of chemistry. The Huisgen cycloaddition consists of the condensation of an organic azide with an alkyne group to form a 1,2,3-triazole bond. Azide and alkyne functional groups can be easily introduced into the scaffolds of large-scale biologically relevant organic constructs. This reaction can be catalyzed by the introduction of copper(I). The Cu(I) core activates the slowly reacting alkyne group, thus slowing down the kinetics of the azide-alkyne condensation to approximately 10 7 ~10 8 This reaction has the dual effect of accelerating the reaction rate and organizing reactive groups through "templating" to ensure only regiospecific 1,4-disubstituted adducts are formed. This reaction is known as copper-catalyzed azide-alkyne cycloaddition (CuAAC), and its compatibility with a wide variety of biological substrates and synthetic conditions makes CuAAC the most important of click conjugation methods. Since its discovery, Cu(I)-catalyzed azide-alkyne cycloaddition has been widely used in the fields of biology, biochemistry, and biotechnology. Click chemistry reactions that can be used include, but are not limited to, Huisgen azide-alkyne 1,3-dipolar cycloaddition, copper-catalyzed azide-alkyne cycloaddition (CuAAC), and ruthenium-catalyzed azide-alkyne cycloaddition (RuAAC). Further details regarding click chemistry using nucleic acids can be found, for example, in Fantoni et al. (2021) Chem. Rev. 121(12):7122-7154.
[0052] The disclosed methods may include one or more annealing steps, such as annealing a primer conjugated to a degraded moiety to the primer binding site of a nucleic acid that immobilizes a polypeptide to a solid support, or annealing a primer to a nucleic acid label at the degraded N-terminal amino acid. One skilled in the art can design various nucleic acids used in the disclosed methods so that they can anneal to each other as desired, for example, by designing the nucleic acids to have complementary regions as needed. As used herein, the terms "complementary" or "complementarity" refer to the nucleotide sequence of a first nucleic acid that noncovalently base pairs with a region of a second nucleic acid, or the nucleotide sequence of a first region of a nucleic acid that noncovalently base pairs with a second region (e.g., a stem region) of the nucleic acid. In canonical Watson-Crick base pairing, adenine (A) base pairs with thymine (T), similar to guanine (G) with cytosine (C) in DNA. In RNA, thymine is replaced by uracil (U). Thus, A is complementary to T, and G is complementary to C. In RNA, A is complementary to U, and vice versa. Typically, "complementary" or "complementarity" refers to nucleotide sequences that are at least partially complementary. These terms can also encompass fully complementary duplexes, such that every nucleotide in one strand is complementary to every nucleotide in the other strand at corresponding positions. In certain cases, a nucleotide sequence can be partially complementary to a target, where every nucleotide is not complementary to every nucleotide in the target nucleic acid at every corresponding position. For example, a region of a first nucleic acid can be completely (i.e., 100%) complementary to a region of a second nucleic acid, or the region of the first nucleic acid can share a degree of complementarity that is less than complete (e.g., 70%, 75%, 85%, 90%, 95%, 99%). The percent identity of two nucleotide sequences can be determined by aligning the sequences for optimal comparison purposes (e.g., gaps can be introduced into the sequence of the first sequence for optimal alignment).The nucleotides at corresponding positions are then compared, and the percent identity between the two sequences is a function of the number of identical positions shared by the sequences (i.e., % identity = number of identical positions / total number of positions × 100). When a position in one sequence is occupied by the same nucleotide as the corresponding position in the other sequence, the molecules are identical at that position. A non-limiting example of such a mathematical algorithm is described in Karlin et al., Proc. Natl. Acad. Sci. USA 90:5873-5877 (1993). Such an algorithm is incorporated into the NBLAST and XBLAST programs (version 2.0), as described in Altschul et al., Nucleic Acids Res. 25:389-3402 (1997). When using BLAST and Gapped BLAST programs, the default parameters of the respective programs (e.g., NBLAST) can be used. In some embodiments, parameters for sequence comparison can be set at score=100, word length=12, or can be varied (e.g., word length=5 or word length=20).
[0053] The conditions during the annealing step can be conditions under which a first nucleic acid (e.g., a primer) specifically hybridizes to a second nucleic acid (e.g., a template nucleic acid). Whether specific hybridization occurs depends on the degree of complementarity between the relevant portions of the nucleic acid, their length, and the temperature at which hybridization occurs (the melting temperature (T) of the relevant portions of the nucleic acid). M The melting temperature is determined by factors such as the temperature at which half of the nucleic acid remains hybridized and half dissociates into single strands. The Tm of a duplex is calculated using the following formula: Tm = 81.5 + 16.6(log10[Na + ])+0.41(fraction G+C)-(600 / N), where N is the chain length and [Na +] is less than 1 M. See Sambrook and Russell (2001; Molecular Cloning: A Laboratory Manual, 3rd ed. Cold Spring Harbor Press, Cold Spring Harbor, NY, Ch. 10). Other, more sophisticated models that depend on various parameters can also be used to predict the Tm of nucleic acid duplexes in response to various hybridization conditions. Approaches to achieving specific nucleic acid hybridization can be found, for example, in Tijssen, Laboratory Techniques in Biochemistry and Molecular Biology—Hybridization with Nucleic Acid Probes, part I, chapter 2, “Overview of principles of hybridization and the strategy of nucleic acid probe assays,” Elsevier (1993).
[0054] In certain embodiments, the polypeptide sequencing method of the present disclosure comprises annealing a primer to a nucleic acid label of the degraded N-terminal amino acid, wherein the primer comprises a barcode corresponding to the identity of the degraded N-terminal amino acid. In one non-limiting embodiment, in this annealing step, the primer comprising the barcode corresponding to the identity of the degraded N-terminal amino acid is conjugated to a binding moiety that specifically binds to the degraded N-terminal amino acid, and annealing is dependent on binding of the binding moiety to the degraded N-terminal amino acid.
[0055] A variety of suitable binding moieties may be used, non-limiting examples of which include polypeptide binding moieties (eg, antibodies), small molecules, aptamers, and the like.
[0056] According to some embodiments, the binding moiety is an antibody. The term "antibody" can include antibodies or immunoglobulins of any isotype (e.g., IgG (e.g., IgG1, IgG2, IgG3, or IgG4), IgE, IgD, IgA, IgM, etc.), whole antibodies (e.g., antibodies composed of a tetramer composed of two dimers of, in turn, a heavy and a light chain polypeptide); single-chain antibodies (e.g., scFv); single-chain Fv (scFv), Fab, (Fab')2, (scFv')2, and diabodies, as well as fragments of antibodies (e.g., whole antibodies or single-chain antibody fragments) that retain specific binding to a cell surface molecule of a target cell; chimeric antibodies; monoclonal antibodies, human antibodies, humanized antibodies (e.g., humanized whole antibodies, humanized half antibodies, or humanized antibody fragments, e.g., humanized scFv); and fusion proteins comprising an antigen-binding portion of an antibody and a non-antibody protein. According to some embodiments, the antibody is selected from an IgG, Fv, single chain antibody, scFv, Fab, F(ab')2, or Fab'. In certain embodiments, the antibody is a nanobody (an antibody fragment consisting of a single monomeric variable antibody domain, also known as a single domain antibody (sdAb)), a monobody (a synthetic binding protein constructed using a fibronectin type III domain (FN3) as a molecular scaffold), or a bispecific T cell engager (BiTE).
[0057] Immunoglobulin light or heavy chain variable regions (V L and V H) are composed of a "framework" region (FR) interrupted by three hypervariable regions, also called "complementarity-determining regions" or "CDRs." The extent of the framework region and CDRs has been defined (see E. Kabat et al., Sequences of proteins of immunological interest, 4th ed. USDept. Health and Human Services, Public Health Services, Bethesda, MD (1987), and Lefranc et al. IMGT, the international ImMunoGeneTics information system®. Nucl. Acids Res., 2005, 33, D593-D597). The sequences of framework regions of different light or heavy chains are relatively conserved within a species. The framework region of an antibody, the combined framework regions of the constituent light and heavy chains, serves to position and align the CDRs. The CDRs primarily contribute to binding to an antigen epitope.
[0058] Thus, "antibody" encompasses a protein having one or more polypeptides that may be genetically encodable, for example, by immunoglobulin genes or fragments of immunoglobulin genes. Recognized immunoglobulin genes include the kappa, lambda, alpha, gamma, delta, epsilon, and mu constant region genes, as well as the myriad immunoglobulin variable region genes. Light chains are classified as either kappa or lambda. Heavy chains are classified as gamma, mu, alpha, delta, or epsilon, which define the immunoglobulin classes IgG, IgM, IgA, IgD, and IgE, respectively.
[0059] As used herein, the term "monoclonal antibody" refers to an antibody obtained from a substantially homogeneous population of antibodies, i.e., the individual antibodies comprising the population are identical except for possible naturally occurring mutations that may be present in minor amounts. For example, a monoclonal antibody may be derived from a single clone, including any eukaryotic, prokaryotic, yeast, or phage clone, or may be produced via a cell-free expression system, regardless of the method by which it is produced. A monoclonal antibody composition exhibits a single binding specificity and affinity for a particular epitope. Monoclonal antibodies are highly specific, being directed against a single antigenic site. Furthermore, in contrast to conventional (polyclonal) antibody preparations, which typically include different antibodies directed against different determinants (epitopes), each monoclonal antibody is directed against a single determinant on the antigen. The modifier "monoclonal" indicates the character of the antibody as being obtained from a substantially homogeneous population of antibodies and should not be construed as requiring production of the antibody by any particular method. Monoclonal antibodies can be prepared using a wide variety of techniques known in the art, including, but not limited to, hybridoma, recombinant, yeast display, phage display, ribosome display, DNA display, etc. For example, monoclonal antibodies may be made by the hybridoma method first described by Kohler et al., Nature 256:495 (1975), or may be made by recombinant DNA methods (see, e.g., U.S. Pat. No. 4,816,567). "Monoclonal antibodies" may also be isolated from phage antibody libraries using, for example, the techniques described in Clackson et al., Nature 352:624-628 (1991) and Marks et al., J. Mol. Biol. 222:581-597 (1991).
[0060] According to some embodiments, the binding moiety is a small molecule. By "small molecule" compound is meant a compound having a molecular weight of 1000 atomic mass units (amu) or less. In some embodiments, the small molecule is 900 amu or less, 750 amu or less, 500 amu or less, 400 amu or less, 300 amu or less, or 200 amu or less. In some cases, the small molecule is not made up of repeating molecular units such as those found in a polymer.
[0061] In certain embodiments, the binding moiety is an aptamer. "Aptamer" refers to a nucleic acid (e.g., an oligonucleotide) that has specific binding affinity for a target cell surface molecule. Aptamers exhibit certain desirable properties, such as ease of selection and synthesis, high binding affinity and specificity, and versatile synthetic accessibility.
[0062] The phrases "specifically bind," "specifically to," "immunoreactive," and "immunoreactive," and "antigen-binding specificity," when referring to a binding moiety (e.g., an antibody, small molecule, aptamer, etc.), refer to a binding reaction with an antigen (e.g., a particular amino acid or post-translationally modified form thereof) that is highly preferential for the antigen, such that the reaction is determinative for and / or selective for the antigen in the presence of a heterogeneous population of antigens (e.g., a mixture of different amino acids). Thus, under specified conditions, a particular binding moiety will bind to a particular antigen and not bind in significant amounts to other antigens present in a sample. Specific binding to an antigen under such conditions may require a binding moiety selected for its specificity for a particular antigen. For example, a binding moiety (e.g., an antibody) may specifically bind to a particular amino acid and not exhibit comparable binding (e.g., no detectable binding) to other proteins present in a sample.
[0063] In some embodiments, the binding moiety is, for example, about 10 5 M -1 or higher affinity or K a(i.e., the equilibrium association constant of the particular binding interaction in units of 1 / M). In certain embodiments, the binding moiety "specifically binds" to or associates with a particular amino acid with a 6 M -1 , 10 7 M -1 , 10 8 M -1 , 10 9 M -1 , 10 10 M -1 , 10 11 M -1 , 10 12 M -1 , or 10 13 M -1 More than K a "High affinity" binding is defined as binding to a specific amino acid with at least 10 7 M -1 , at least 10 8 M -1 , at least 10 9 M -1 , at least 10 10 M -1 , at least 10 11 M -1 , at least 10 12 M -1 , at least 10 13 M -1 , or more K a Alternatively, affinity may be expressed in units of M (e.g., 10 -5 M~10 -13 The equilibrium dissociation constant (K D In some embodiments, specific binding can be defined as a binding moiety that binds to a molecule having a specific binding activity of about 10 -5 M or less, about 10 -6 M or less, about 10 -7 M or less, about 10 -8 M or less, or about 10 -9 M or less, 10 -10 M, 10 -11 M or 10 -12 K below M DThe binding affinity of a binding moiety for a particular amino acid can be readily determined using conventional techniques, for example, by competitive ELISA (enzyme-linked immunosorbent assay), equilibrium dialysis, by using surface plasmon resonance (SPR) technology (e.g., a BIAcore 2000 instrument using the general procedures outlined by the manufacturer), by radioimmunoassay, etc.
[0064] In certain embodiments, the binding moiety specifically binds to a degraded N-terminal amino acid containing a post-translational modification (PTM), and the barcode indicates the identity of the degraded N-terminal amino acid and the PTM. PTMs are chemical modifications that play an important role in functional proteomics because they regulate activity, localization, and interactions with other cellular molecules, such as proteins, nucleic acids, lipids, and cofactors. PTMs of interest include, but are not limited to, phosphorylation, glycosylation, ubiquitination, nitrosylation, methylation, acetylation, or lipidation.
[0065] Protein phosphorylation, primarily on serine, threonine, or tyrosine residues, is one of the most important and well-studied post-translational modifications. Phosphorylation plays a key role in regulating many cellular processes, including the cell cycle, growth, apoptosis, and signal transduction pathways. Protein glycosylation is recognized as one of the major post-translational modifications, with profound effects on protein folding, conformation, distribution, stability, and activity. Glycosylation encompasses a diverse selection of sugar moieties added to proteins, ranging from simple monosaccharide modifications of nuclear transcription factors to highly complex branched polysaccharide modifications of cell surface receptors. Carbohydrates, in the form of asparagine-linked (N-linked) or serine / threonine-linked (O-linked) oligosaccharides, are major structural components of many cell surface and secreted proteins. Ubiquitin is an 8-kDa polypeptide consisting of 76 amino acids that is attached to the ε-NH2 of lysines in target proteins via its C-terminal glycine. After the initial monoubiquitination event, the formation of ubiquitin polymers may occur, and polyubiquitinated proteins are then recognized by the 26S proteasome, which catalyzes the degradation of ubiquitinated proteins and the recycling of ubiquitin. S-nitrosylation is a key PTM used by cells to stabilize proteins, regulate gene expression, and provide NO donors. The production, localization, activation, and catabolism of SNOs are tightly regulated. S-nitrosylation is a reversible reaction, and SNOs have a short half-life in the cytoplasm due to a host of reductases, including glutathione (GSH) and thioredoxin, that denitrosylate proteins. Therefore, SNOs are often stored in membranes, vesicles, interstitial spaces, and lipophilic protein folds to protect them from denitrosylation. The transfer of a one-carbon methyl group to the nitrogen or oxygen of an amino acid side chain (N-methylation and O-methylation, respectively) increases the hydrophobicity of proteins and can neutralize the negative amino acid charge when attached to a carboxylic acid. Methylation is mediated by methyltransferases, with S-adenosylmethionine (SAM) being the initial methyl group donor.N-acetylation, or the transfer of an acetyl group to nitrogen, occurs in nearly all eukaryotic proteins through both irreversible and reversible mechanisms. N-terminal acetylation requires cleavage of the N-terminal methionine by methionine aminopeptidase (MAP) before the amino acid is replaced with an acetyl group from acetyl-CoA by N-acetyltransferase (NAT) enzymes. This type of acetylation is co-translational, in that the N-terminus is acetylated on the growing polypeptide chain while still attached to the ribosome. 80–90% of eukaryotic proteins are acetylated in this manner. Lipidation is a method of targeting proteins to membranes of organelles (endoplasmic reticulum [ER], Golgi apparatus, mitochondria), vesicles (endosomes, lysosomes), and the plasma membrane. The four types of lipidation are C-terminal glycosylphosphatidylinositol (GPI) anchors, N-terminal myristoylation, S-myristoylation, and S-prenylation. Each type of modification confers a different membrane affinity to the protein, but all types of lipidation increase the protein's hydrophobicity and therefore its affinity for membranes. Different types of lipidation are also not mutually exclusive, in that more than one lipid can be attached to a given protein.
[0066] As summarized above, the steps of labeling, decomposing, annealing, and extending are performed in successive cycles to produce a plurality of extension products, each of which includes a UMI, a respective cycle number barcode, and a barcode corresponding to the identity of the respective decomposed N-terminal amino acid. Once the plurality of extension products are produced, the method further includes sequencing the plurality of extension products. As will be appreciated from the present disclosure, the sequence of the polypeptide can be determined based on the sequences of the plurality of extension products.
[0067] Sequencing of multiple extension products can be performed using any of a variety of available high-throughput nucleic acid sequencers and systems. Illustrative sequencing systems include the Illumina iSeq 100, Miniseq, MiSeq series, NextSeq series (e.g., NextSeq 500 series, NextSeq 1000, NextSeq 2000), and NovaSeq sequencing systems (Illumina, Inc., San Diego, Calif.), Pacific Biosciences Sequel (e.g., Sequel II) sequencing system (Pacific Biosciences, Menlo Park, Calif.), Oxford Nanopore Technologies MinION™, GridIONx5™, PromethION™, or SmidgION™ nanopore-based sequencing systems (Oxford Nanopore Technologies, Oxford, UK), and other systems with similar capabilities.
[0068] On the Illumina platform, the sequencing process involves clonal amplification of adaptor-ligated DNA fragments onto the surface of a glass slide. Bases are read using a cyclic reversible termination strategy, which sequences one nucleotide of the template strand at a time through progressive rounds of base incorporation, washing, imaging, and cleavage. This strategy uses fluorescently labeled 3'-O-azidomethyl-dNTPs to temporarily halt the polymerization reaction, allowing removal of unincorporated bases and fluorescent imaging to determine the added nucleotide. Following scanning of the flow cell by a charge-coupled device (CCD) camera, the fluorescent moiety and 3' block are removed, and the process is repeated.
[0069] In zero-mode waveguide (ZMW)-based sequence analysis, the ZMW is a nanoscale well that functions as an optical trap, allowing for the observation of individual polymerase molecules. As a result, nucleotide incorporation events provide observation of the incorporated nucleotide analog, which is easily distinguishable from unincorporated nucleotide analogs. For a description of ZMWs and their application in nucleic acid sequencing, see, e.g., U.S. Patent Application Publication No. 2003 / 0044781 and U.S. Patent No. 6,917,726 (each of which is incorporated by reference in its entirety for all purposes). See also Levene et al. (2003) "Zero-mode waveguides for single-molecule analysis at high concentrations" Science 299:682-686, Eid et al. (2009) "Real-time DNA sequencing from single polymerase molecules" Science 323:133-138, and U.S. Patent Nos. 7,056,676, 7,056,661, 7,052,847, 7,033,764, and 7,907,800 (the complete disclosures of which are incorporated by reference in their entirety for all purposes).
[0070] In nanopore sequencing, the nanopore functions as a biosensor, providing the only path for the ionic solution on the cis side of the membrane to contact the ionic solution on the trans side. A constant voltage bias (trans side positive) generates an ionic current through the nanopore, driving ssDNA or ssRNA in the cis chamber through the pore to the trans chamber. A processive enzyme (e.g., helicase, polymerase, nuclease, etc.) may be bound to the polynucleotide such that its stepwise movement controls the nucleotide, ratcheting nucleobase by nucleobase through the small nanopore. Because the ionic conductivity through the nanopore is sensitive to the mass of the nucleobase and the presence of its associated electric field, the level of ionic current through the nanopore reveals the sequence of the nucleobases within the translocating strand. Patch clamping, voltage clamping, etc. may be used.
[0071] Details for obtaining raw sequencing reads of nucleic acid molecules using nanopores are described, for example, in Feng et al. (2015) Genomics, Proteomics & Bioinformatics 13(1):4-16. Nanopore-based sequencing systems are available, including the SmidgION, MinION, GridION, and PromethION nanopore-based sequencing systems available from Oxford Nanopore Technologies Limited. Detailed design considerations and protocols for performing nucleic acid sequencing are provided with such systems.
[0072] The methods of the present disclosure can be carried out in any suitable container / confinement. One or more steps of the method can be carried out in a first container, while one or more other steps can be carried out in a second container. Non-limiting examples of containers in which one or more steps of the method can be carried out include tubes, vials, plates, wells of a multi-well plate (e.g., a 6-well plate, a 12-well plate, a 24-well plate, a 48-well plate, a 96-well plate, or a 384-well plate), a confinement in a microfluidic device, etc.
[0073] Compositions and Kits Aspects of the present disclosure further include compositions. In some embodiments, compositions are provided that include one or more of any of the polypeptides and / or one or any combination of reagents for performing the polypeptide sequencing methods of the present disclosure described elsewhere herein. Non-limiting examples of reagents are those combinations that may be present in the compositions of the present disclosure, including UMI-functionalized solid supports, degradation moieties bearing reactive groups for conjugation to nucleic acids, primers comprising cycle number barcodes, Edman degradation reagents (including those for providing alternative DNA-compatible Edman degradation conditions described elsewhere herein), conjugates comprising primers conjugated to binding moieties that specifically bind to amino acids, conjugates comprising primers conjugated to binding moieties that specifically bind to amino acids bearing post-translational modifications, nucleic acid sequencing adaptors, and any combination thereof.
[0074] According to some embodiments, compositions of the present disclosure include any one or any combination of any of the polypeptides and / or reagents present in a liquid medium. The liquid medium may be an aqueous liquid medium such as water, a buffer solution, etc. One or more additives may be present in such compositions, such as salts (e.g., NaCl, MgCl, KCl, MgSO), buffers (Tris buffer, N-(2-hydroxyethyl)piperazine-N'-(2-ethanesulfonic acid) (HEPES), 2-(N-morpholino)ethanesulfonic acid (MES), 2-(N-morpholino)ethanesulfonic acid sodium salt (MES), 3-(N-morpholino)propanesulfonic acid (MOPS), N-tris[hydroxymethyl]methyl-3-aminopropanesulfonic acid (TAPS), etc.), solubilizers, detergents (e.g., non-ionic detergents such as Tween-20), nuclease inhibitors, glycerol, chelating agents, etc.
[0075] The subject compositions may be present in any suitable environment. According to one embodiment, the composition is present in a reaction tube (e.g., a 0.2 mL tube, a 0.6 mL tube, a 1.5 mL tube, etc.) or a well. In certain aspects, the composition is present in two or more (e.g., a plurality) reaction tubes or wells (e.g., a plate, e.g., a 6-well plate, a 12-well plate, a 24-well plate, a 48-well plate, a 96-well plate, or a 384-well plate). The tubes and / or plates may be made of any suitable material, such as, for example, polypropylene. In certain aspects, the tubes and / or plates in which the compositions are present provide efficient heat transfer to the composition (e.g., when placed in a heat block, water bath, thermocycler, etc.) so that the temperature of the composition can be changed within a short period of time as needed, for example, to allow a particular degradation or enzymatic reaction to occur. According to certain embodiments, the composition is present in a thin-walled polypropylene tube or a plate with thin-walled polypropylene wells.
[0076] Other suitable environments for the subject compositions include, for example, microfluidic chips (e.g., "lab-on-a-chip devices"). The composition may be present in an instrument configured to bring the composition to a desired temperature, e.g., a temperature-controlled water bath, a heating block, etc. The instrument configured to bring the composition to a desired temperature may be configured to bring the composition to a series of different desired temperatures, each for a suitable period of time (e.g., the instrument may be a thermocycler).
[0077] Aspects of the present disclosure also include kits. The kits may include, for example, one or any combination of reagents for carrying out the polypeptide sequencing methods of the present disclosure described elsewhere herein. Non-limiting examples of reagents are those combinations that may be present in the compositions of the present disclosure, including UMI-functionalized solid supports, degradation moieties bearing reactive groups for conjugation to nucleic acids, primers containing cycle number barcodes, Edman degradation reagents (including those for providing alternative DNA-compatible Edman degradation conditions described elsewhere herein), conjugates comprising primers conjugated to binding moieties that specifically bind to amino acids, conjugates comprising primers conjugated to binding moieties that specifically bind to amino acids bearing post-translational modifications, nucleic acid sequencing adaptors, and any combination thereof.
[0078] According to some embodiments, the subject kits include one or any combination of the following reagents: (i) a UMI-functionalized solid support, (ii) a degradation moiety bearing a reactive group for conjugation to a nucleic acid, (iii) a primer comprising a cycle number barcode, (iv) an Edman degradation reagent, (v) a conjugate comprising a primer conjugated to a binding moiety that specifically binds to an amino acid, (vi) a conjugate comprising a primer conjugated to a binding moiety that specifically binds to an amino acid bearing a post-translational modification, and (vii) a nucleic acid sequencing adaptor. According to some embodiments, the degradation moiety comprises a PITC bearing a reactive group for conjugation to a primer comprising a cycle number barcode. In some cases, the reactive group is a click chemistry reactive group. In certain embodiments, the Edman degradation reagent comprises BF etherate and an aprotic solvent. According to some embodiments, the Edman degradation reagent comprises triethylamine acetate and N,N-dimethylformamide (DMF). In certain embodiments, one or more of the nucleic acid-based reagents comprise a non-natural nucleotide (e.g., a 7-deazapurine nucleotide) that stabilizes the nucleic acid under Edman degradation conditions. In some cases, the binding moiety is a polypeptide (e.g., an antibody). In other cases, the binding moiety is a small molecule or an aptamer.
[0079] The components of the kit may be in separate containers, or multiple components may be in a single container, for example, two or more components of the kit may be provided in a single tube or in different tubes.
[0080] In addition to the components described above, kits of the present disclosure may further include instructions for using one or any combination of the reagents, for example, to perform any of the polypeptide sequencing methods of the present disclosure. The instructions are generally recorded on a suitable recording medium. For example, the instructions may be printed on a substrate such as paper or plastic. Thus, the instructions may be present in the kit as a package insert, on labeling of the container of the kit or its components (i.e., associated with the packaging or subpackaging), etc. In other embodiments, the instructions are present as an electronic storage data file present on a suitable computer-readable storage medium, e.g., a CD-ROM, a diskette, a hard disk drive (HDD), etc. In still other embodiments, the actual instructions are not present in the kit, but means are provided for obtaining the instructions from a remote source, e.g., via the Internet. An example of this embodiment is a kit that includes a web address at which the instructions can be viewed and / or from which the instructions can be downloaded. As with the instructions, this means that the means for obtaining the instructions is recorded on a suitable substrate.
[0081] For completeness, the present disclosure is further defined in the following numbered clauses:
[0082] 1. A method for sequencing a polypeptide, the method comprising: (a) labeling the N-terminal amino acid of a polypeptide with a nucleic acid label comprising a unique molecular identifier (UMI); (b) cleaving the N-terminal amino acid from the polypeptide; (c) annealing a primer to the nucleic acid label of the degraded N-terminal amino acid, wherein the primer comprises a barcode corresponding to the identity of the degraded N-terminal amino acid; (d) extending a primer annealed to the nucleic acid target to produce an extension product comprising a barcode corresponding to the identity of the UMI and the degraded N-terminal amino acid; (e) performing steps (a)-(d) in successive cycles to produce a plurality of extension products, each of the plurality of extension products comprising a UMI, a respective cycle number barcode, and a barcode corresponding to the identity of each resolved N-terminal amino acid; (f) sequencing the plurality of extension products; (g) determining the sequence of the polypeptide based on the sequences of the plurality of extension products.
[0083] 2. The method of clause 1, comprising indexing the extension products produced in step (d) with cycle number barcodes.
[0084] 3. The method of clause 1, wherein step (a) comprises labeling the N-terminal amino acid of the polypeptide with a nucleic acid label comprising a UMI and a cycle number barcode.
[0085] 4. The method of any one of clauses 1 to 3, wherein prior to step (a), the polypeptide is immobilized to the solid support via a nucleic acid attached to the C-terminus of the polypeptide and the surface of the solid support, the nucleic acid comprising a UMI and a primer binding site 3' or 5' to the UMI.
[0086] 5. The labeling step (a) comprises: (i) conjugating a degradation moiety to the N-terminal amino acid; (ii) conjugating a primer to a degradation moiety, wherein the primer conjugated to the degradation moiety comprises a sequence that is complementary to a primer binding site of a nucleic acid that immobilizes the polypeptide to a solid support; (iii) annealing a primer conjugated to a degradation moiety to the primer binding site; (iv) using the nucleic acid that immobilizes the polypeptide on a solid support as a template to extend a primer conjugated to a degradation moiety, thereby labeling the N-terminal amino acid with a nucleic acid label that includes a UMI.
[0087] 6. The labeling step (a) comprises: (i) conjugating to the N-terminal amino acid a degradation moiety conjugated to a primer comprising a sequence complementary to a primer binding site of a nucleic acid that immobilizes the polypeptide to a solid support; (ii) annealing a primer conjugated to a degradation moiety to the primer binding site; (iii) using the nucleic acid that immobilizes the polypeptide on a solid support as a template to extend a primer conjugated to a degradation moiety, thereby labeling the N-terminal amino acid with a nucleic acid label that includes a UMI.
[0088] 7. The method of clause 5 or 6, wherein the degrading moiety comprises phenylisothiocyanate (PITC) bearing a reactive group for conjugating to a nucleic acid comprising a cycle number barcode.
[0089] 8. The method of clause 7, wherein the reactive group is a click chemistry reactive group.
[0090] 9. The method of any one of clauses 1 to 8, wherein the decomposing step (b) is carried out under conditions comprising a Lewis acid in an aprotic solvent.
[0091] 10. The method of clause 9, wherein the Lewis acid is BF3 etherate, BCl3, BBr3, scandium(III) triflate, or any combination thereof.
[0092] 11. The method of claim 9, wherein the Lewis acid is BF3 etherate.
[0093] 12. The method of any one of clauses 9 to 11, wherein the aprotic solvent is acetonitrile, N,N-dimethylformamide (DMF), dimethyl sulfoxide (DMSO), or any combination thereof.
[0094] 13. The method of clause 12, wherein the decomposing step (b) is carried out under conditions including triethylamine acetate in N,N-dimethylformamide (DMF).
[0095] 14. The method of any one of clauses 1 to 13, wherein one or more nucleic acids used in step (a) and / or step (b) comprise non-natural nucleotides that stabilize the nucleic acid during degradation step (b).
[0096] 15. The method of clause 14, wherein the non-natural nucleotide comprises a 7-deazapurine nucleotide.
[0097] 16. The method of any one of clauses 1 to 15, wherein in step (c), a primer comprising a barcode corresponding to the identity of the degraded N-terminal amino acid is conjugated to a binding moiety that specifically binds to the degraded N-terminal amino acid, and annealing is dependent on binding of the binding moiety to the degraded N-terminal amino acid.
[0098] 17. The method of clause 16, wherein the binding moiety specifically binds to the degraded N-terminal amino acid containing the post-translational modification, and the barcode indicates the identity of the degraded N-terminal amino acid and the post-translational modification.
[0099] 18. The method of clause 16 or 17, wherein the post-translational modification is phosphorylation, glycosylation, ubiquitination, nitrosylation, methylation, acetylation, or lipidation.
[0100] 19. The method of any one of clauses 16 to 18, wherein the binding moiety is a polypeptide.
[0101] 20. The method of clause 19, wherein the polypeptide is an antibody.
[0102] 21. The method of any one of clauses 16 to 18, wherein the binding moiety is a small molecule or an aptamer.
[0103] 22. The method of any one of clauses 1 to 21, wherein the polypeptide to be sequenced is present in a protein sample isolated from a single cell.
[0104] 23. The method according to clause 22, wherein the method is a single-cell protein sequencing method performed on multiple polypeptides present in a protein sample.
[0105] 24. The method of any one of clauses 1 to 23, wherein the polypeptide to be sequenced is present in a protein sample isolated from a tissue sample.
[0106] 25. The method of clause 24, wherein the tissue sample is a biopsy sample.
[0107] 26. The method of clause 25, wherein the biopsy sample is a tumor biopsy sample.
[0108] 27. The method of any one of clauses 1 to 23, wherein the polypeptide to be sequenced is present in a protein sample isolated from a biological fluid.
[0109] 28. A composition comprising: (a) UMI-functionalized solid support; (b) a degradable moiety bearing a reactive group for conjugation to a nucleic acid; (c) primers containing cycle number barcodes; (d) Edman degradation reagent, (e) a conjugate comprising a primer conjugated to a binding moiety that specifically binds to an amino acid; (f) a conjugate comprising a primer conjugated to a binding moiety that specifically binds to an amino acid bearing a post-translational modification; and (g) nucleic acid sequencing adapter A composition comprising one or any combination of:
[0110] 29. A kit comprising: (a) the following reagents: (i) UMI-functionalized solid support; (ii) a degradable moiety bearing a reactive group for conjugation to a nucleic acid; (iii) primers containing cycle number barcodes; (iv) Edman degradation reagent, (v) a conjugate comprising a primer conjugated to a binding moiety that specifically binds to an amino acid; (vi) a conjugate comprising a primer conjugated to a binding moiety that specifically binds to an amino acid bearing a post-translational modification; and (vii) nucleic acid sequencing adapter and (b) instructions for using one or any combination of reagents to carry out the method of any one of clauses 1 to 27.
[0111] 30. The kit of clause 29, wherein the degrading moiety comprises a PITC bearing a reactive group for conjugating to a primer comprising a cycle number barcode.
[0112] 31. The kit according to clause 30, wherein the reactive group is a click chemistry reactive group.
[0113] 32. The kit of any one of clauses 29 to 31, wherein the Edman degradation reagent comprises a Lewis acid in an aprotic solvent.
[0114] 33. The kit of clause 32, wherein the Lewis acid is BF3 etherate, BCl3, BBr3, scandium(III) triflate, or any combination thereof.
[0115] 34. The kit according to clause 32, wherein the Lewis acid is BF3 etherate.
[0116] 35. The kit of any one of clauses 32 to 34, wherein the aprotic solvent is acetonitrile, N,N-dimethylformamide (DMF), dimethyl sulfoxide (DMSO), or any combination thereof.
[0117] 36. The kit of any one of clauses 29-35, wherein one or more of the nucleic acid-based reagents comprises a non-natural nucleotide that stabilizes the nucleic acid under Edman degradation conditions.
[0118] 37. The kit of clause 36, wherein the non-natural nucleotide comprises a 7-deazapurine nucleotide.
[0119] 38. The kit of any one of clauses 29-37, wherein the binding moiety is a polypeptide.
[0120] 39. The kit of clause 38, wherein the polypeptide is an antibody.
[0121] 40. The kit of any one of clauses 29 to 37, wherein the binding moiety is a small molecule or an aptamer.
[0122] The following examples are offered by way of illustration and not by way of limitation. [Example]
[0123] Example 1 - Development of an alternative DNA-compatible Edman degradation reaction Conventional Edman degradation conditions are incompatible with DNA. The cleavage and conversion reactions are carried out at high temperatures using neat trifluoroacetic acid (TFA) and TFA solutions, respectively. The use of strong protic acids is the greatest problem for DNA, which is prone to depurination under acidic conditions. Here, it was observed that the cleavage reaction conditions degraded poly(dT) oligonucleotides. Furthermore, PITC and its derivatives, which are used to modify the N-terminus of peptides, can react with the exocyclic amines of nucleobases.
[0124] This example describes the development of an alternative Edman degradation reaction compatible with DNA. It was hypothesized that DNA degradation is primarily caused by protonation of nucleobases under strongly acidic conditions. Therefore, an Edman degradation procedure using BF3 etherate in an aprotic solvent for the cleavage step was employed. Consistent with the assumption, polypyrimidine sequences were stable under these conditions for extended periods. However, natural purine nucleotides still underwent depurination under these conditions, albeit at a significantly slower rate. To further enhance DNA stability, chemically modified purine nucleotides were investigated. 7-deazapurine nucleotides, lacking a nitrogen atom at the 7th position, have been reported to be resistant to depurination.
[0125] Oligonucleotides containing 7-deazapurine nucleotides are stable under degradation conditions for 4 h, as verified by LC and MS (Figures 3A and 3C). Given the rapid N-terminal amino acid cleavage in the presence of BF3 etherate (see below), the stability of 7-deazapurine-modified DNA is sufficient for the Edman degradation process. Finally, 7-deazapurine-modified DNA was subjected to PITC in water / pyridine (1:1) at 50 °C for 16 h, but no modification of the DNA was detected. This is consistent with the low nucleophilicity of the exocyclic amines of the nucleobases.
[0126] Cleavage of the N-terminal amino acid was completed in 5 min in the presence of 40 mM BF3 etherate. BF3 etherate-induced cleavage was significantly faster than the TFA cleavage reaction, which took 30 min. Anilinothiazolinone (ATZ) amino acids were the major fragments generated under these conditions, and ATZ amino acids were converted to stable PTC amino acids under mildly basic conditions (Figure 3B). PTC amino acids can be easily synthesized by reacting PITC or its derivatives with amino acids, thus allowing easy access to target molecules for the generation of binding agents (e.g., antibodies).
[0127] Example 2 - Development of conditions for solid-phase Edman degradation of DNA-peptide conjugates This example demonstrates a DNA-compatible Edman degradation in a solid-phase format. Solid-phase reactions offer two advantages. First, they allow the use of large excesses of reagents, which can be easily removed by filtration. This greatly simplifies the design of repeated Edman degradation cycles. Furthermore, DNA-conjugated peptides have different reactivities in organic solvents compared to unconjugated peptides due to the insolubility of oligonucleotides in these organic solvents, a fact well documented in the field of DNA-encoded libraries. The initial experiments performed here are consistent with the literature and suggest that Edman degradation does not occur on DNA-conjugated PTCs in anhydrous acetonitrile. It has been reported that immobilizing DNA on a solid phase makes it accessible to DNA-encoded synthesis via chemical transformations in non-aqueous solvents. With this in mind, we investigated the Edman degradation of peptide-DNA conjugates on solid supports.
[0128] As demonstrated herein, Edman degradation can be performed on immobilized DNA-peptide conjugates. To test Edman degradation on solid supports, we first confirmed that the stability of DNA under degradation was unaffected by immobilization. Next, to test the feasibility of Edman degradation, DNA-peptide conjugates were synthesized on solid supports. This was achieved by synthesizing dibenzocyclooctyne (DBCO)-modified DNA sequences on controlled-pore glass (CPG) or polystyrene-coated carboxylic acid magnetic beads. CPG is commonly used for solid-phase DNA synthesis, and magnetic beads are generally enzyme-compatible and allow for easy separation. A PTC peptide containing a C-terminal azidolysine was conjugated to DBCO-modified DNA via strain-promoted alkyne-azide cycloaddition (SPAAC) to form a model DNA-peptide conjugate (Figure 4A). The cleavage reaction was carried out using 40 mM BF3 etherate in anhydrous acetonitrile. The supernatants were collected, and the release of PTC amino acids was confirmed by LC-MS on both solid supports (Figure 4B). The DNA-peptide conjugates were cleaved from the degraded CPG and analyzed by HPLC. The results indicated that degradation was complete within 10 min (Figure 4C).
[0129] Example 3 - Barcoding of degradation fragments of N-terminal amino acids on model peptides In this example, the first cycle of INDEED is achieved using the DNA-compatible Edman degradation reaction described above. To achieve this, first, a 7-deazapurine-modified DNA (UMI) bearing 3' amino and 5' DBCO groups is synthesized. The UMI is immobilized on 1 μm Carboxylic Acid Dynabeads™ using carbodiimide chemistry. A polypeptide containing a C-terminal azidolysine is conjugated to DNA via SPAAC. Second, the N-terminus of the polypeptide is modified with an alkyne-bearing PITC derivative (2). Third, methyltetrazine azide (3) is conjugated to an alkyne via copper(I)-catalyzed azide-alkyne cycloaddition (CuAAC), and a trans-cyclooctene (TCO)-modified primer is installed via an inverse electron demand Diels-Alder (IEDDA) reaction. In the fourth step, the DNA UMI is transcribed via a primer extension reaction. In the final step, the N-terminal amino acid is cleaved from the peptide by treatment with BF3 etherate (Figure 5B), which is accomplished by a tandem cleavage-hydrolysis reaction to generate PTC amino acid fragments barcoded with DNA UMIs.
[0130] Experiments conducted to investigate several key steps for DNA barcoding of degradation fragments determined that Edman degradation can be performed on the complete UMI-peptide-primer construct. Using the conditions established above, peptide-UMI conjugates were immobilized on beads. First, alkyne-modified PITC (2) was synthesized by treating 4-(2-aminoethyl)aniline with an alkyne NHS ester (1) at 0 °C to achieve selective acylation of aliphatic amines. The aniline moiety was then converted to an isothiocyanate under conditions reported by Scattolin et al. (Figure 5A). Modification of the peptide-UMI conjugate on beads with (2) was quantitative. The primer was successfully installed via the previously described CuAAC-IEDDA cascade to yield the complete UMI-peptide-primer construct (Figure 5B). The relative amount of primer sequence on the beads can be quantified by flow cytometry by annealing a fluorescently labeled complementary strand. The yield of Edman degradation was determined by comparing the fluorescence intensity before and after the reaction. The degradation yield was approximately 85% after 10 minutes (Figure 5C).
[0131] Next, we determined that 7-deazapurine-modified DNA is accepted by DNA polymerases. PTC amino acid barcoding requires that polymerases accept 7-deazapurine nucleotide-substituted template-primer duplexes and 7-deazapurine nucleoside triphosphates as substrates. We examined primer extension with Sequenase version 2.0, Klenow (exo-), and Bst3.0 (Figure 5D). All of these polymerases can efficiently incorporate 7-deazapurine nucleoside triphosphates, and Sequenase version 2.0 will be used in future experiments.
[0132] Example 4 - Alternative method for barcoding degradation fragments of N-terminal amino acids on model peptides In this example, a 3'-dibenzocyclooctyne (DBCO)-modified deazapurine-substituted DNA template was immobilized on magnetic beads (Figure 5F). A 7-amino acid model peptide bearing a C-terminal azidolysine was then conjugated to the DNA via strain-promoted alkyne-azide cycloaddition (SPAAC). Azide-modified PITC (4, Figure 5E) reacted with the N-terminus of the model peptide.
[0133] To introduce a primer for barcode transfer, we used conjugation of an alkyne-modified primer to the azide group on 4. SPAAC was chosen due to the sensitivity of the phenylthiocarbamyl group in 2 to oxidation, which can lead to undesired reactions with reactive oxygen species generated during copper(I)-catalyzed azide-alkyne cycloaddition (CuAAC). One drawback of SPAAC is its relatively slow reaction rate. However, it was determined that the reaction rate could be enhanced by hybridizing the primer with the template, thereby increasing the effective concentration of reactants. For example, conjugation of the primer in the presence of a 7-amino acid-long peptide was remarkably completed within 10 min.
[0134] After the primer was incorporated, the beads were extended using Klenow (Figure 5F) and subjected to Lewis acid-catalyzed Edman degradation. During this step, we took advantage of the insolubility of DNA in acetonitrile, allowing the DNA-barcoded ATZ amino acid to remain hybridized to the template DNA after the initial cleavage reaction. A basic conversion reaction then serves two purposes: (i) converting the ATZ amino acid to a stable PTC amino acid and (ii) denaturing the double-stranded DNA and releasing the DNA-barcoded PTC amino acid. Additionally, DTT was introduced to suppress the oxidative degradation of the PTC amino acid under basic conditions. This cleavage and conversion reaction yielded an overall yield of 96%, and the resulting structure was confirmed by mass spectrometry analysis (Figure 5G).
[0135] Example 5 - Characterization of binders that recognize PTC amino acids BD-PEX functions to translate binding events between DNA-barcoded PTC amino acids and binding agents (e.g., antibodies) into DNA output. PTC amino acids retain the structural features of the original amino acid and differ only in the PTC modification on the amino group. Because these amino groups are often modified to conjugate with carrier proteins during the generation of antibodies against amino acids, it is expected that antibodies generated against amino acids will also recognize PTC amino acids. It is then expected that commercially available antibodies can be used to detect these amino acids. A four-amino acid fingerprint is sufficient to identify most proteins in the human proteome.
[0136] This example demonstrates for the first time that an antibody against an amino acid recognizes a PTC amino acid. As proof of concept, we synthesized PTC-tryptophan by reacting tryptophan with the PITC derivative (1) and conjugated it to azide-modified DNA. Except for the linker distal to PTC-tryptophan, this conjugate is structurally identical to that formed by INDEED. The binding affinity of a commercially available tryptophan mAb to the DNA conjugate PTC-tryptophan was measured by biolayer interferometry (BLI) and surface plasmon resonance (SPR). The tryptophan mAb had a Kd of 280 nM for PTC-tryptophan, and its binding was highly specific, as no binding was observed with PTC-tyrosine or PTC-phenylalanine (Figure 6A). Furthermore, an antibody against phosphotyrosine (PY20) was also tested, yielding a Kd of 20 nM (Figure 6B). This indicates that PTC amino acids bearing post-translational modifications (PTMs) can also be recognized by their corresponding anti-PTM antibodies. Testing additional anti-PTM antibodies led to the identification of antibodies that recognize PTCs for asymmetric dimethylarginine (ADMA), acetyllysine, and phosphoserine (Figure 6E).
[0137] As shown above, PTC amino acids bearing click handles can be easily synthesized and conjugated to azide-modified carrier proteins such as BSA. This example demonstrates that antibodies against PTC amino acids can be generated. To obtain new PTC amino acid-specific antibodies, we initiated an antibody discovery program. PTC-tyrosine and PTC-phenylalanine were synthesized and conjugated to azide-modified BSA. Mice were challenged with the modified BSA, and strong immune responses were observed against both antigens. Antibody-producing B cells were then harvested and fused to form hybridomas. Subsequent subcloning and screening identified several highly specific candidates (Figures 6C and 6D). Using this strategy, we identified antibodies against PTC-modified phenylalanine, tyrosine, tryptophan, arginine, and aspartic acid (Figure 6F).
[0138] Example 6 - Development of Binding-Dependent Primer Extension (BD-PEX) to Convert DNA Barcoded Degraded Fragments into DNA Sequences A highly specific binding agent (e.g., an antibody) is used to recognize the PTC amino acids released by DNA-encoded Edman degradation and convert the binding event into a DNA output. This conversion to DNA output allows all antibody-antigen interactions to be revealed simultaneously, and the resulting DNA can be further amplified to increase detection sensitivity. To achieve this, referring to Figure 7A, the Fc region of the antibody is site-specifically modified using the SiteClick™ kit to introduce an azide functionality. A DBCO-modified primer carrying an antibody-specific barcode is then conjugated to the antibody. The primer sequence is designed to be short, thus disfavoring intermolecular primer extension. Binding of the antibody to its PTC amino acid target increases the effective molar concentration of the primer, promoting duplex formation between the complementary sequence on the primer and the template. The resulting complex can serve as a substrate for primer extension, converting the binding event into a sequenceable DNA output. Primer extension products can be analyzed by qPCR and / or DNA sequencing to determine reaction yield and detection limit.
[0139] Example 7 - Binding-dependent primer extension (BD-PEX) to convert DNA barcoded degradation fragments into DNA sequences using biotinylated primers This example describes the use of biotinylated primers during the INDEED process, allowing barcoded DNA to be pulled down by streptavidin beads, allowing BD-PEX to be performed on the beads (Figure 7B). To achieve this, referring to Figure 7B, the Fc region of an antibody was site-specifically modified using a kit such as the SiteClick™ kit or the oYo-Link kit to introduce a click handle (e.g., azide or tetrazine). DBCO- or TCO-modified primers bearing antibody-specific barcodes were then conjugated to the antibody. The stability of the primer-template complex plays an important role in controlling the efficiency and specificity of intramolecular primer extension. While increasing the primer length can improve the efficiency of primer extension, longer primers also tend to promote extension in the absence of antibody-antigen recognition events. Primers with a length of 6 nucleotides were found to enable efficient primer extension while simultaneously minimizing nonspecific primer extension (Figure 7C). Furthermore, intermolecular primer extension served as another source of nonspecific primer extension on the beads. This form of nonspecific primer extension can be effectively suppressed by reducing the surface density of the DNA-barcoded PTC amino acids. At a DNA density of 1 pmol per mg of magnetic beads, intermolecular primer extension was virtually eliminated (Figure 7E).
[0140] Example 8 - Short peptide sequencing / fingerprinting A peptide with the sequence RGFDWGK{N3} was subjected to five cycles of the INDEED process. The DNA-barcoded PTC amino acids obtained from each cycle were pulled down onto streptavidin beads in separate containers. Proximity primer extension was performed using a mixture of antibodies specific to the DNA-barcoded PTC amino acids (anti-PTC-Arg, anti-PTC-Phe, anti-PTC-Asp, and anti-PTC-Trp antibodies, each at 100 nM). After primer extension, adapter PCR was performed in each container using an adapter primer carrying the cycle number barcode. Finally, all DNA was pooled, indexed, and sequenced on a MiSeq sequencer (Figure 8A). The cycle number barcode and antibody barcode were extracted from the sequencing results. The read counts for all possible barcode combinations were plotted on a heatmap (Figure 8B), and the results were consistent with the peptide sequence.
[0141] Example 9 - Sequencing / fingerprinting of short peptides and quantification of single amino acid substitutions Single amino acid substitutions, caused by nonsynonymous single nucleotide polymorphisms (nsSNPs), often disrupt protein function by altering protein structure. While DNA sequencing allows for sensitive detection of SNPs, detection of single amino acid substitutions by MS is often limited by sensitivity. The method described herein is expected to enable detection of single amino acid substitutions by sequencing peptides at single amino acid resolution. Furthermore, DNA sequencing readout allows for signal amplification, thus enhancing sensitivity.
[0142] To demonstrate peptide sequencing via INDEED, a mixture of peptides carrying single amino acid substitutions is immobilized on UMI-coated magnetic beads using the chemistry established above. INDEED is performed repeatedly to collect DNA-barcoded PTC amino acids. Organic solvents from the digested fragment mixture are removed by solid-phase extraction or buffer exchange prior to enzymatic reaction. The combined DNA-barcoded PTC amino acids are then converted to DNA sequences by BD-PEX using a mixture of DNA-barcoded antibodies against the PTC amino acids. The resulting DNA is sequenced. It is expected that single amino acid substitutions can be identified regardless of their position within the polypeptide. In addition to identifying single amino acid substitutions in a highly parallel manner, the sequencing method of the present invention allows for the quantification of these mutations using read counts. When each polypeptide molecule is uniquely barcoded, the occurrence of single amino acid substitutions can be digitally quantified, and quantification is not affected by PCR bias.
[0143] Example 10 - Mapping and quantification of post-translational modifications (PTMs) in model peptides Identifying and quantifying PTMs is important for understanding protein function. Currently, PTMs are commonly studied using antibody-based techniques and mass spectrometry. However, antibody-based techniques are often not site-specific. While PTM analysis by MS can provide information about the modification site, accurate quantification of PTMs often requires the use of chemically synthesized isotope-labeled peptide standards. It is expected that most PTM-specific antibodies can be employed in the sequencing method disclosed herein. Because PTM amino acid fragments are barcoded with DNA containing information about the order of amino acids within the peptide, this method can specifically map PTM sites. Furthermore, quantification of PTMs can be achieved by DNA sequencing without the need to synthesize isotope-labeled peptide standards specific to the protein of interest.
[0144] To demonstrate this capability, peptides containing two tyrosine amino acids and all their possible phosphotyrosine derivatives are synthesized. INDEED is performed on a mixture of these peptides. Tyrosines are identified with a PTC-tyrosine-specific antibody, and phosphotyrosines are recognized with an anti-phosphotyrosine antibody, such as PY20, described above. Recognition events are recorded by BD-PEX, and the resulting DNA is sequenced. It is expected that the sites of phosphorylation are encoded in the DNA sequence, and the relative abundance of phosphorylation can be quantified using read counts. This method can be extended to other PTMs, such as serine and threonine phosphorylation, methylation, acylation, and glycosylation, depending on the stability of the PTMs during INDEED.
[0145] Example 11 - Full-length protein fingerprinting Proteins in eukaryotes are, on average, 400 amino acids long. Due to limitations in degradation efficiency, fingerprinting full-length proteins by Edman degradation may not be desirable. Therefore, to fingerprint full-length proteins, proteins may be digested with endopeptidases such as trypsin to obtain short peptides, which are then subjected to INDEED. This capability is demonstrated by fingerprinting the trypsin digestion of full-length proteins.
[0146] Cysteine-mediated peptide immobilization can be achieved thanks to a wide variety of cysteine-specific reactions, such as α-halocarbonyl and maleimide. By controlling the pH, selective modification of cysteine over other nucleophilic residues, such as lysine, histidine, and the N-terminus, can be achieved. However, the low abundance of cysteine (2%) can sometimes cause incomplete capture of tryptic peptides. C-terminal carboxylic acids are a more generalizable conjugation handle for peptide immobilization. C-terminal carboxylic acids can be selectively labeled by carboxypeptidases, whose proteolytic activity is inhibited at high pH, while transpeptidyl activity has been reported to catalyze the ligation of nucleophiles to the C-terminal carboxylic acid. More recently, photoredox-catalyzed decarboxylation of C-terminal carboxylic acids has been described. While this method has only been demonstrated for short peptides less than 10 amino acids in length, it may serve as a more general and efficient method for C-terminal immobilization. Conjugation of click handles, such as alkynes, has been demonstrated, and therefore these methods can be easily implemented into the INDEED workflow.
[0147] The side chains of cysteine and lysine can be protected before degradation. It is well documented that cysteine can be protected by alkylation. Protection of lysine can be achieved by first masking the N-terminus with a reversible modification, and then the lysine is irreversibly protected with a reagent such as an NHS ester. After protection of the lysine, the N-terminal amino group is released by removal of the reversible modification. Furthermore, these protection reactions can be used to introduce affinity tags that are recognized by existing affinity reagents, thus further expanding the range of sequenceable amino acids.
[0148] Example 12 - Proteoform mapping from single cells Mapping proteoforms at the single-cell level can reveal cellular heterogeneity beyond the gene or even protein level, potentially significantly advancing our understanding of cellular function, organismal development, and disease mechanisms. The disclosed polypeptide sequencing method can be used for single-molecule profiling of proteoforms, such as single amino acid substitutions and post-translational modifications. Here, we develop a workflow for mapping these proteoforms at the single-cell level (Figure 9). First, to isolate and enrich the protein of interest, single cells are isolated via FACS in a multiwell plate containing lysis buffer and beads coated with an antibody against the protein of interest. Second, the protein of interest is eluted from the antibody-coated beads and digested with trypsin. Finally, the resulting peptides are conjugated to DNA UMIs and sequenced using our single-molecule peptide sequencing technology. Importantly, the UMIs used in this workflow may also contain barcodes specific to each well, thus enabling the identification and quantification of proteoforms in each cell.
[0149] Accordingly, the foregoing merely illustrates the principles of the present disclosure. Those skilled in the art will recognize that, although not explicitly described or shown herein, they can devise various configurations that embody the principles of the present invention and are within its spirit and scope. Furthermore, all examples and conditional language recited herein are intended primarily to aid the reader in understanding the principles of the present invention and the concepts contributed by the inventors to further the art, and should be construed without limitation to such specifically recited examples and conditions. Furthermore, all statements herein describing principles, aspects, and embodiments of the present invention, as well as specific examples thereof, are intended to encompass both structural and functional equivalents thereof. Additionally, such equivalents are intended to include both currently known equivalents and equivalents developed in the future, i.e., any elements developed that perform the same function, regardless of structure. Therefore, the scope of the present invention is not intended to be limited to the exemplary embodiments shown and described herein.
Claims
1. 1. A method for sequencing a polypeptide, said method comprising: (a) labeling the N-terminal amino acid of a polypeptide with a nucleic acid label comprising a unique molecular identifier (UMI); (b) cleaving the N-terminal amino acid from the polypeptide; (c) annealing a primer to the nucleic acid label of the degraded N-terminal amino acid, wherein the primer comprises a barcode corresponding to the identity of the degraded N-terminal amino acid; (d) extending the primer annealed to the nucleic acid target to produce an extension product comprising the UMI and the barcode corresponding to the identity of the resolved N-terminal amino acid; (e) performing steps (a)-(d) in successive cycles to produce a plurality of extension products, each of the plurality of extension products comprising the UMI, a respective cycle number barcode, and a barcode corresponding to the identity of a respective resolved N-terminal amino acid; (f) sequencing the plurality of extension products; and (g) determining the sequence of the polypeptide based on the sequences of the plurality of extension products.
2. 2. The method of claim 1, comprising indexing the extension products produced in step (d) with the cycle number barcode.
3. 2. The method of claim 1, wherein step (a) comprises labeling the N-terminal amino acid of the polypeptide with a nucleic acid label comprising the UMI and the cycle number barcode.
4. 4. The method of claim 1, wherein, prior to step (a), the polypeptide is immobilized on the solid support via a nucleic acid connected to the C-terminus of the polypeptide and the surface of the solid support, the nucleic acid comprising the UMI and a primer binding site 3' or 5' to the UMI.
5. The labeling step (a) (i) conjugating a degradation moiety to said N-terminal amino acid; (ii) conjugating a primer to the degradation moiety, wherein the primer conjugated to the degradation moiety comprises a sequence that is complementary to the primer binding site of the nucleic acid that immobilizes the polypeptide to the solid support; (iii) annealing the primer conjugated to the degradation moiety to the primer binding site; (iv) using the nucleic acid that immobilizes the polypeptide to the solid support as a template to extend the primer conjugated to the degradation moiety, thereby labeling the N-terminal amino acid with the nucleic acid label that comprises the UMI.
6. The labeling step (a) (i) conjugating to the N-terminal amino acid a degradation moiety conjugated to a primer comprising a sequence complementary to the primer binding site of the nucleic acid that immobilizes the polypeptide to the solid support; (ii) annealing the primer conjugated to the degradation moiety to the primer binding site; (iii) using the nucleic acid that immobilizes the polypeptide to the solid support as a template to extend the primer conjugated to the degradation moiety, thereby labeling the N-terminal amino acid with the nucleic acid label that comprises the UMI.
7. 7. The method of claim 5 or 6, wherein the degrading moiety comprises phenylisothiocyanate (PITC) bearing a reactive group for conjugating to the nucleic acid comprising the cycle number barcode.
8. The method of claim 7 , wherein the reactive group is a click chemistry reactive group.
9. 9. The method of any one of claims 1 to 8, wherein the decomposing step (b) is carried out under conditions comprising a Lewis acid in an aprotic solvent.
10. The Lewis acid is BF 3 Etherate, BCl 3 , BBr 3 10. The method of claim 9, wherein the cation exchanger is scandium(III) triflate, scandium(III) triflate, or any combination thereof.
11. The Lewis acid is BF 3 10. The method of claim 9, wherein the compound is an etherate.
12. 12. The method of any one of claims 9 to 11, wherein the aprotic solvent is acetonitrile, N,N-dimethylformamide (DMF), dimethyl sulfoxide (DMSO), or any combination thereof.
13. 13. The method of claim 12, wherein the decomposing step (b) is carried out under conditions comprising triethylamine acetate in N,N-dimethylformamide (DMF).
14. 14. The method of any one of claims 1 to 13, wherein one or more nucleic acids used in step (a) and / or step (b) comprise non-natural nucleotides that stabilize the nucleic acid during degrading step (b).
15. The method of claim 14, wherein the non-natural nucleotide comprises a 7-deazapurine nucleotide.
16. 16. The method of any one of claims 1 to 15, wherein in step (c), the primer comprising the barcode corresponding to the identity of the degraded N-terminal amino acid is conjugated to a binding moiety that specifically binds to the degraded N-terminal amino acid, and wherein the annealing is dependent on binding of the binding moiety to the degraded N-terminal amino acid.
17. 17. The method of claim 16, wherein the binding moiety specifically binds to a degraded N-terminal amino acid containing a post-translational modification, and the barcode indicates the identity of the degraded N-terminal amino acid and the post-translational modification.
18. 18. The method of claim 16 or 17, wherein the post-translational modification is phosphorylation, glycosylation, ubiquitination, nitrosylation, methylation, acetylation, or lipidation.
19. The method of any one of claims 16 to 18, wherein the binding moiety is a polypeptide.
20. 20. The method of claim 19, wherein the polypeptide is an antibody.
21. The method of any one of claims 16 to 18, wherein the binding moiety is a small molecule or an aptamer.
22. The method of any one of claims 1 to 21, wherein the polypeptide to be sequenced is present in a protein sample isolated from a single cell.
23. 23. The method of claim 22, wherein the method is a single-cell protein sequencing method performed on multiple polypeptides present in the protein sample.
24. The method of any one of claims 1 to 23, wherein the polypeptide to be sequenced is present in a protein sample isolated from a tissue sample.
25. 25. The method of claim 24, wherein the tissue sample is a biopsy sample.
26. 26. The method of claim 25, wherein the biopsy sample is a tumor biopsy sample.
27. The method of any one of claims 1 to 23, wherein the polypeptide to be sequenced is present in a protein sample isolated from a biological fluid.
28. 1. A composition comprising: (a) UMI-functionalized solid support; (b) a degradable moiety bearing a reactive group for conjugation to a nucleic acid; (c) a primer comprising a cycle number barcode; (d) Edman degradation reagent; (e) a conjugate comprising a primer conjugated to a binding moiety that specifically binds to an amino acid; (f) a conjugate comprising a primer conjugated to a binding moiety that specifically binds to an amino acid bearing a post-translational modification; and (g) nucleic acid sequencing adaptor A composition comprising one or any combination of:
29. A kit comprising: (a) the following reagents: (i) a UMI-functionalized solid support; (ii) a degradable moiety bearing a reactive group for conjugation to a nucleic acid; (iii) a primer comprising a cycle number barcode; (iv) Edman degradation reagents, (v) a conjugate comprising a primer conjugated to a binding moiety that specifically binds to an amino acid; (vi) a conjugate comprising a primer conjugated to a binding moiety that specifically binds to an amino acid bearing a post-translational modification; and (vii) nucleic acid sequencing adaptors and (b) instructions for using one or any combination of reagents to carry out the method of any one of claims 1 to 27.
30. 30. The kit of claim 29, wherein the degradation moiety comprises a PITC having a reactive group for conjugating to the primer comprising the cycle number barcode.
31. 31. The kit of claim 30, wherein the reactive group is a click chemistry reactive group.
32. The kit according to any one of claims 29 to 31, wherein the Edman degradation reagent comprises a Lewis acid in an aprotic solvent.
33. The Lewis acid is BF 3 Etherate, BCl 3 , BBr 3 , scandium(III) triflate, or any combination thereof.
34. The Lewis acid is BF 3 33. The kit of claim 32, wherein the compound is an etherate.
35. 35. The kit of any one of claims 32 to 34, wherein the aprotic solvent is acetonitrile, N,N-dimethylformamide (DMF), dimethyl sulfoxide (DMSO), or any combination thereof.
36. 36. The kit of any one of claims 29 to 35, wherein one or more of the nucleic acid-based reagents comprises a non-natural nucleotide that stabilizes the nucleic acid under Edman degradation conditions.
37. 37. The kit of claim 36, wherein the non-natural nucleotide comprises a 7-deazapurine nucleotide.
38. The kit of any one of claims 29 to 37, wherein the binding moiety is a polypeptide.
39. 39. The kit of claim 38, wherein the polypeptide is an antibody.
40. The kit of any one of claims 29 to 37, wherein the binding moiety is a small molecule or an aptamer.