Methods of sequencing polypeptides and related compositions
By using the intramolecular DNA-encoded Edman degradation method, combined with Edman degradation and DNA sequencing technology, the versatility, sensitivity and throughput issues of single-cell protein sequencing were solved, and efficient sequencing of polypeptides and identification of post-translational modifications were achieved.
Patent Information
- Application Number
- CN202480014527.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-24
- Filing Date
- 2024-02-23
- Publication Date
- 2025-10-03
AI Technical Summary
Existing single-cell protein sequencing technologies lack versatility, sensitivity, and throughput, and are unable to effectively identify multiple amino acid sequences and their post-translational modifications. In particular, methods based on nanopore, Edman degradation, and real-time dynamic protein sequencing have limitations.
The intramolecular DNA-encoded Edman degradation method is used to label the N-terminal amino acid of the polypeptide with a nucleic acid marker containing a unique molecular identifier and a cycle number barcode, and combine Edman degradation and DNA sequencing technology to achieve efficient sequencing of the polypeptide.
It achieves high-throughput, single-amino acid resolution sequencing of polypeptides, can identify multiple amino acids and their post-translational modifications, and meet the needs of single-cell protein sequencing.
Smart Images

Figure CN120752349A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 448,131, filed February 24, 2023, which is incorporated herein by reference in its entirety.
[0003] introduction
[0004] Over the past decade, advances in DNA sequencing have enabled mRNA sequencing at the single-cell level, which has revolutionized our understanding of the heterogeneity of biological systems. The impact of these insights on healthcare is enormous—ranging from our fundamental understanding of early organ development to the identification of rare, therapy-resistant cell populations in complex tumors. Historically, researchers have assumed that mRNA expression directly correlates with protein expression levels. However, cross-gene correlation analyses have shown that mRNA levels explain only approximately 40% of the variation in protein levels. This is because protein levels are influenced by many factors, such as translation rate, translational regulation, and protein degradation. Furthermore, proteins often undergo post-translational modifications (PTMs) such as phosphorylation, glycosylation, methylation, and many other modifications after synthesis. It is well known that PTMs can have significant effects on a protein’s activity, localization, and interactions with other biomolecules. Crucially, PTMs are regulated by enzymatic processes that are not directly encoded by genes, and therefore they cannot be predicted from the transcriptome. Therefore, there is a pressing unmet need to go beyond mRNA sequencing and directly sequence proteins (and their PTMs) at the single-cell level.
[0005] Protein sequencing has been performed by mass spectrometry (MS) since the 1970s, and the sensitivity of MS instruments has increased significantly over the past five decades. For example, current state-of-the-art MS instruments can detect approximately 1,000 different proteins in 0.8 ng of cell lysate. While this is impressive, they represent only a small fraction of the approximately 20,000 proteins known to exist in cells. Importantly, even with further advances, it is uncertain whether MS will achieve sufficient dynamic range to enable future single-cell protein sequencing.
[0006] Many techniques have been proposed for single-cell protein sequencing, and generally, their potential can be calibrated using three metrics: (1) versatility; (2) sensitivity; and (3) throughput. Versatility refers to whether the method can be used to identify arbitrary amino acid sequences, regardless of their chemical composition, such as charge, hydrophobicity, and length. Versatility also refers to whether the method can distinguish between natural amino acids and PTM-modified amino acids. Sensitivity refers to whether the method has the potential to eventually measure a single amino acid in a single protein. Throughput refers to whether the method has the potential to eventually sequence all proteins from a single human cell, which typically contains 10 amino acids composed of 10,000 to 20,000 different species. 8 to 10 9 A protein molecule.
[0007] The current technological developments for single-molecule protein sequencing can be roughly divided into three categories: 1) nanopore-based peptide fingerprint analysis, 2) Edman degradation-based peptide fluorescence fingerprint analysis; and 3) real-time dynamic protein sequencing.
[0008] Nanopore sequencing platforms use pore-forming proteins or nano-fabricated pores and measure changes in the current as the peptide translocates through the pore, creating a "fingerprint" of the peptide. In theory, the advantage of nanopore sequencing platforms is that they can read a variety of amino acids with high accuracy. For example, when all 20 amino acids are located at the C-terminus of a polyarginine peptide, the Aeromonas lysin nanopore has been shown to distinguish them. Recently, it has been demonstrated that the MspA nanopore can be used to perform fingerprint analysis on peptides with single amino acid substitutions. However, in practice, current implementations of nanopore sequencing are not suitable for sequencing arbitrary peptides. For example, the above studies were conducted only using highly charged model peptides, while the non-uniform charge of natural peptides hinders their translocation through nanopores. In addition, accurately identifying peptide sequences with single amino acid resolution through nanopores is very challenging because up to eight amino acids can cause changes in the ion current when the peptide passes through the nanopore. Crucially, the biggest disadvantage of nanopore-based sequencing may be its throughput. For example, it may take 30 minutes to measure a single 25-amino acid peptide. Considering that a single cell contains about 10 8 to 10 9 Even with massively parallel operation, it is uncertain whether nanopore sequencing can achieve the throughput required for single-cell proteomics.
[0009] The second strategy, known as peptide fluorescence fingerprinting based on Edman degradation, combines Edman chemistry with single-molecule microscopy. In this method, cysteine and lysine side chains are fluorescently labeled, and the peptide is immobilized on a solid support via the C-terminus. Next, the N-terminal amino acid is sequentially removed by Edman degradation, removing one residue at a time. Digestion of the fluorescently labeled amino acid results in a decrease in fluorescence intensity, which serves as a unique fingerprint for a given peptide. The advantage of this method is that it may allow fingerprinting of millions of peptides in parallel, and due to the robustness of Edman degradation, the method can be extended to most peptide sequences. The key weakness of this method is that spectrally distinguishable fluorophores are used as substitutes for amino acids, and therefore specific labeling of amino acids with high efficiency is crucial for this technology. Unfortunately, only fluorescent labeling of lysine and cysteine has been demonstrated. To date, only a small set of amino acids (e.g., lysine, cysteine, and tyrosine) can be labeled with sufficient specificity and efficiency required for this technology. Furthermore, photobleaching, chemical degradation of fluorophores, and (in the context of peptides) energy transfer between fluorophores may limit the sensitivity and versatility of this technique.
[0010] Finally, real-time dynamic protein sequencing utilizes the continuous degradation of surface-immobilized peptides by aminopeptidases. During degradation, the nascent N-terminal amino acid is identified in real time by a mixture of dye-labeled N-terminal amino acid conjugates evolved from the adaptor protein ClpS. The N-terminal amino acid is identified not only based on binding affinity but also by binding kinetics, allowing the identification of multiple amino acids by a single binder protein. The advantage of this approach is that it is not limited by either the charge state of the peptide or the chemical functionality of the amino acid side chains and can therefore potentially be generalized to most peptide sequences. Furthermore, the use of N-terminal amino acid conjugates greatly expands the number of sequenceable amino acids. To date, ClpS proteins used for real-time dynamic protein sequencing have been shown to discriminate between seven different N-terminal amino acids. However, a key weakness of this strategy is that the binding of the ClpS protein to the N-terminal amino acid is influenced by neighboring amino acids—a crucial issue. Recently, it has been shown that the affinity and binding kinetics of the ClpS protein vary significantly even within a small subset of possible downstream sequences. Therefore, the feasibility of this approach will depend on the availability of new reagents whose binding affinity and kinetics are independent of the amino acid attached to the N-terminal amino acid. Unfortunately, no such reagents have been discovered to date. Furthermore, the enzymatic degradation used for real-time kinetic sequencing does not proceed in a stepwise manner. This results in inconsistent lifetimes of ClpS-N-terminal amino acid complexes and a lack of the ability to measure the lengths of uncharacterized peptides, both of which limit the accuracy of the technique.
[0011] For the reasons outlined above, no technology currently offers the versatility, sensitivity, and throughput to ultimately achieve the goal of single-cell protein sequencing. Summary of the Invention
[0012] A method for sequencing a polypeptide is provided. In certain embodiments, the method includes labeling the N-terminal amino acid of the polypeptide with a nucleic acid marker comprising a unique molecular identifier and a cycle number barcode; degrading the N-terminal amino acid from the polypeptide; annealing a primer to the nucleic acid marker of the degraded amino acid, wherein the primer comprises a barcode corresponding to the identity of the degraded N-terminal amino acid; and extending the primer annealed to the nucleic acid marker to produce an extension product, the extension product comprising a unique molecular identifier, a cycle number barcode, and a barcode corresponding to the identity of the degraded N-terminal amino acid. The aforementioned steps are performed continuously to produce a plurality of such extension products, which are then sequenced to enable determination of the amino acid sequence of the polypeptide. Compositions and kits for finding use in practicing the method are also provided. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figures 1A-1C : (1A) Overview of single-molecule polypeptide sequencing according to embodiments of the present disclosure. In this example, the method combines the chemistry of Edman degradation with the massive parallelism of DNA synthesis sequencing technology. (1B) Schematic illustration of an embodiment in which polypeptides are labeled with nucleic acid markers comprising a unique molecular identifier (UMI) and a cycle number barcode. (1C) Schematic illustration of an embodiment in which biotinylated primers are employed such that the barcoded DNA is pulled down by streptavidin beads, where binding-dependent primer extension occurs. Also in this example, the primer extension products are indexed using a cycle number barcode.
[0014] Figures 2A-2D:(2A) Reaction scheme of intramolecular DNA-encoded Edman degradation (INDEED) according to some embodiments of the present disclosure. In this example, a polypeptide is conjugated to a DNA unique molecular identifier (UMI) functionalized bead (step 1). A primer for recording the UMI is conjugated to the N-terminus of the polypeptide through a modified PITC and click reaction (steps 2 and 3). Then, the DNA UMI is transcribed through a primer extension reaction (step 4). Finally, the PTC amino acid is produced by cleavage and hydrolysis (step 5). The remaining fixed polypeptide is subjected to the next cycle of INDEED. (2B) Reaction scheme of binding-dependent primer extension (BD-PEX). The PTC amino acid is recognized by an antibody labeled with a primer and an antibody-specific barcode (step 1). The binding event is recorded by primer extension. The resulting DNA is sequenced to reveal the polypeptide sequence (step 2). (2C) Reaction scheme of INDEED according to some embodiments of the present disclosure. In this example, the C-terminal side of the peptide is conjugated to a DNA UMI functionalized bead (step 1). Subsequently, the N-terminus of the peptide is derivatized by azide-modified PITC and a biotinylated primer is incorporated via a proximity-promoted click reaction (steps 2 and 3). In the fourth step, the UMI is copied to the nascent DNA chain after primer extension (step 4). Finally, Lewis acid-catalyzed cleavage, followed by hydrolysis under reducing alkaline conditions, releases the DNA-barcoded PTC amino acids from the solid support (steps 5 and 6). (2D) The antibody-assisted proximity extension scheme following the scheme described in 2C. The DNA-barcoded PTC amino acids are pulled down by streptavidin beads. Antibodies labeled with primer- and antibody-specific barcodes are introduced, and recognition events are recorded by primer extension. The resulting DNA is made into a sequencing library, during which a cycle number barcode is introduced. This process generates DNA that contains information on the position, origin, and identity of the amino acids and can be read by DNA sequencing.
[0015] Figure 3 A-3C: (3A) DNA modified with 7-deazapurine nucleotides is stable under Lewis acid-catalyzed Edman degradation conditions. (3B) Reaction scheme for the Lewis acid-catalyzed Edman degradation reaction. (3C) MS spectrum of an oligonucleotide containing 7-deazapurine deoxynucleotides after treatment with 40 mM BF3 etherate.
[0016] Figure 4A-4C: (4A) Preparation of solid-phase immobilized DNA-peptide conjugates. (4B) Mass spectra of PTC-tryptophan in the supernatant released by Edman degradation on a solid support. Inset shows the chromatogram of the supernatant. (4C) Edman degradation on CPG. DNA was cleaved from CPG by ammonia and analyzed by HPLC. PTC-peptide-DNA conjugate (0 min, right trace) was completely degraded to produce peptide-DNA conjugate (10 min, left trace).
[0017] Figure 5 A-5H: (5A) Synthesis of alkyne-modified PITC derivatives. 2. (5B) Preparation of DNA-peptide conjugates on beads and DNA barcoding of the N-terminal amino acid by INDEED. (5C) Quantification of degradation yield by flow cytometry. Primers were hybridized to FAM-labeled complementary strands and quantified on a flow cytometer. The decrease in fluorescence intensity after degradation was used to calculate the degradation yield. (5D) Screening of polymerases for primer extension. Purine nucleotides in the template and primers were completely replaced with 7-deazadA and 7-deazadG. The dNTP mixture contained dTTP, dCTP, 7-deazadATP, and 7-deazadGTP. (5E)-(5F) An embodiment of INDEED, in which magnetic beads are employed and primers for barcode transfer are introduced via proximity-promoted SPAAC. (5G) LC-MS data of DNA-barcoded PTC amino acids generated by INDEED. (5H) Flow cytometry graphs showing the stepwise degradation yield and overall yield of the INDEED process over five degradation cycles.
[0018] Figure 6 A-6F: (6A) Binding of commercially available antibodies to DNA-conjugated PTC amino acids. (6B) Binding of a phosphotyrosine-specific antibody (PY20) to DNA-conjugated PTC-phosphotyrosine. (6C) Specificity of antibodies produced by hybridomas against PTC-tyrosine-conjugated BSA, characterized by ELISA. (6D) Specificity of antibodies produced by hybridomas against PTC-phenylalanine-conjugated BSA, characterized by ELISA. (6E) Data demonstrating antibody recognition of PTC against asymmetric dimethylarginine (ADMA), acetyl-lysine, and phosphoserine. (6F) Data demonstrating identification of antibodies against PTC-modified phenylalanine, tyrosine, tryptophan, arginine, and aspartic acid.
[0019] Figure 7A-7E: (7A) Conversion of PTC amino acids to DNA by binding-dependent primer extension (BD-PEX) according to some embodiments of the present disclosure. (7B) Conversion of PTC amino acids to DNA by BD-PEX according to some embodiments of the present disclosure. (7C) PAGE analysis of products generated by antibody-assisted proximity extension. Primers of 5 nt and 6 nt in length can distinguish the presence of antibody-antigen interactions. (7D) Quantification of DNA output of antibody-assisted proximity extension by qPCR. Primers of 6 nt in length gave quantitative DNA output and were selected for future experiments. (7E) Reducing the loading density on the beads inhibited the occurrence of undesired intermolecular primer extension.
[0020] Figures 8A-8B : Peptides are sequenced by DNA sequencing. (8A) Scheme for sequencing a single peptide species with multiple readable amino acids. The peptide is subjected to five cycles of the INDEED process. The resulting DNA-barcoded PTC amino acids from each cycle are pulled onto streptavidin beads in separate containers. Proximity primer extension is performed with a mixture of DNA-barcoded PTC amino acid-specific antibodies (PTC-Arg antibody, PTC-Phe antibody, PTC-Asp antibody, PTC-Trp antibody at 100nM each). After primer extension, adapter PCR is performed in each container using adapter primers with cycle number barcodes. Finally, all DNA is pooled, indexed, and sequenced on a MiSeq sequencer. (8B) Heat map of the read count DNA barcodes obtained for the peptide sequence RGFDW.
[0021] Figure 9 A-9C: Single-cell proteoform mapping workflow. (9A) Single-cell peptide extraction. Single cells are isolated via FACS in multiwell plates containing cleavage buffer. Proteins of interest are pulled down by antibody-coated beads. These proteins are then digested with trypsin. (9B) Peptides are conjugated to a DNA UMI that also includes a barcode specific to each well. The barcoded peptides are sequenced by INDEED. (9C) The distribution of protein variants (e.g., phosphorylation) can be mapped in each cell at single-amino acid resolution. DETAILED DESCRIPTION
[0022] Before describing the methods, compositions and kits of the present disclosure in more detail, it should be understood that the methods, compositions and kits are not limited to the specific embodiments described and, as such, may, of course, vary. It should also be understood that the terminology used herein is for the purpose of describing specific embodiments only and is not intended to be limiting, as the scope of the methods, compositions and kits will be limited only by the appended claims.
[0023] Where a range of values is provided, it is understood that unless the context clearly dictates otherwise, each intervening value between the upper and lower limits of the range, to the tenth of the unit of the lower limit, and any other stated or intervening value in the stated range are encompassed in the methods, compositions and kits. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges and are also encompassed in the methods, compositions and kits, subject to any specifically excluded limits in the stated ranges. Where a stated range includes one or both limits, ranges excluding either or both of those included limits are also encompassed in the methods, compositions and kits.
[0024] Certain ranges are given herein where a numerical value is preceded by the term "about." The term "about" is used herein to provide literal support for the exact number that follows it, as well as to provide literal support for other numbers that are close to or approximately the number that follows the term. In determining whether a number is close to or approximately a specifically recited number, the close or approximate unrecited number can be a substantially equivalent number to the specifically recited number given the context in which it appears.
[0025] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the methods, compositions, and kits belong. Although any methods, compositions, and kits similar or equivalent to those described herein can also be used in the practice or testing of the methods, compositions, and kits, representative illustrative methods, compositions, and kits are now described.
[0026] All publications and patents cited in this specification are herein incorporated by reference to the same extent as if each individual publication or patent was specifically and individually indicated to be incorporated by reference and are incorporated herein by reference to disclose and describe the materials and / or methods in connection with which the publication is cited. The citation of any publication is for its disclosure prior to the filing date and should not be construed as an admission that the methods, compositions, and kits of the present invention are not entitled to antedate such publication, as the dates of publication provided may be different from the actual publication dates, which may need to be independently confirmed.
[0027] It should be noted that, as used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. It should also be noted that claims can be drafted to exclude any optional elements. Thus, this statement is intended to serve as antecedent basis for use of exclusive terminology such as "solely," "only," and the like in connection with the recitation of claim elements, or for use of a "negative" limitation.
[0028] It should be understood that, for the sake of clarity, certain features of the methods, compositions, and kits described in the context of separate embodiments may also be provided in combination in a single embodiment. Conversely, for the sake of brevity, various features of the methods, compositions, and kits described in the context of a single embodiment may also be provided individually or in any suitable subcombination. All combinations of embodiments are specifically encompassed by this disclosure and are disclosed herein, as if each combination were individually and explicitly disclosed, to the extent that such combinations comprise operable processes and / or compositions. In addition, all subcombinations listed in the embodiments describing such variables are also specifically encompassed by the methods, compositions, and kits of the present invention and are disclosed herein, as if each such subcombination were individually and explicitly disclosed herein.
[0029] As will be apparent to those skilled in the art after reading this disclosure, each individual embodiment described and illustrated herein has discrete components and features that can be readily separated or combined with the features of any other several embodiments without departing from the scope or spirit of the method of the present invention. Any recited method can be carried out in the order of events recited or in any other order that is logically possible.
[0030] Methods for sequencing peptides
[0031] Aspects of the present disclosure include methods for sequencing a polypeptide. According to some embodiments, the method includes labeling the N-terminal amino acid of the polypeptide with a nucleic acid marker comprising a unique molecular identifier (UMI), and degrading the N-terminal amino acid from the polypeptide. In certain embodiments, such methods further include annealing a primer to the nucleic acid marker of the degraded N-terminal amino acid, wherein the primer comprises a barcode corresponding to the identity of the degraded N-terminal amino acid. In some cases, such methods further include extending the primer annealed to the nucleic acid marker to produce an extension product, wherein the extension product comprises a UMI and a barcode corresponding to the identity of the degraded N-terminal amino acid. The aforementioned steps can be performed in continuous cycles to produce a plurality of extension products, each of the plurality of extension products comprising a UMI, a corresponding cycle number barcode, and a barcode corresponding to the identity of the corresponding degraded N-terminal amino acid. According to some embodiments, such methods further include sequencing the plurality of extension products and determining the sequence of the polypeptide based on the sequences of the plurality of extension products. In certain embodiments, the method includes indexing the extension products produced in the extension step with the cycle number barcode. In other embodiments, the labeling step comprises labeling the N-terminal amino acid of the polypeptide with a nucleic acid tag comprising a UMI and a cycle number barcode.
[0032] The method of the present disclosure addresses these shortcomings and constitutes an improvement to nanopore-based peptide sequencing, Edman degradation-based peptide fluorescence fingerprinting, and real-time dynamic protein sequencing methods. For example, nanopore-based peptide sequencing requires that proteins be charged to enable translocation and also suffers from low throughput. With respect to Edman degradation-based peptide fluorescence fingerprinting, only fluorescent labeling of lysine and cysteine has been demonstrated so far. With respect to real-time dynamic protein sequencing, shortcomings include the effect of protein sequence on recognition subbinding.
[0033] An overview of an embodiment of the method of the present disclosure is schematically shown in FIG1 . In this example, Edman degradation is performed to create "degraded fragments" that are then labeled with DNA barcodes that encode the source of the amino acid (i.e., the polypeptide from which it originated) and its position in the polypeptide. These fragments are specifically recognized by a binding agent (e.g., an antibody, a small molecule, an aptamer, etc.) that is not interfered with by its original downstream polypeptide sequence. The binding of the binding agent allows the DNA barcode to be linked to the amino acid, which enables the peptide sequence to be decoded in a massively parallel manner using a nucleic acid sequencer (e.g., an Illumina type or other suitable DNA sequencer).
[0034] Thus, the methods of the present disclosure, embodiments of which are sometimes referred to herein as intramolecular DNA-encoded Edman degradation (or "INDEED"), comprise two key process modules. The first module degrades terminal amino acids and produces DNA-barcoded amino acids (e.g., anilinothiocarbamoyl (PTC)-amino acids). The second module identifies the amino acids by proximity primer extension and reads the polypeptide sequence by DNA sequencing.
[0035] exist Figure 2A A non-limiting example of a first module according to an embodiment of the present disclosure is schematically shown in FIG. In a first step, the polypeptide is immobilized on a solid support (e.g., a bead or other suitable solid support). This can be achieved by conjugating the C-terminal region of the polypeptide to a bead functionalized with a DNA unique molecular identifier (UMI), such that each peptide is attached to a unique DNA sequence ( Figure 2A , step 1). Then, in the following two steps, the primer for recording the UMI is conjugated to the N-terminus of the polypeptide ( Figure 2A , steps 2 and 3). This can be achieved by reacting the polypeptide with a modified isothiocyanate (e.g., modified PITC) with a click handle. Subsequently, a primer containing a barcode for the cycle number is installed via click chemistry. In the fourth step, the DNA UMI is transcribed by a proximity primer extension reaction ( Figure 2A, step 4). This step transfers the information of the parent polypeptide to the Edman degradation fragment. In the last step, the PTC amino acid is cleaved from the polypeptide ( Figure 2A , step 5). This step can be achieved through a tandem cleavage-hydrolysis reaction, which produces PTC amino acid fragments barcoded with DNA containing information about the cycle number and the parent polypeptide. This process enables improved Edma degradation and the use of non-natural nucleotides, allowing for the preservation of nucleic acids under the harsh conditions of peptide sequencing.
[0036] exist Figure 2B A non-limiting example of the second module according to an embodiment of the present disclosure is schematically shown in FIG. In this example, the second module for reading out the polypeptide sequence by DNA sequencing is performed in two steps. First, the identity information of the amino acid fragment is converted into a specific DNA sequence. To this end, the PTC fragment is recognized by its corresponding binder (e.g., antibody, small molecule, aptamer, etc.), which is conjugated with a primer containing a binder-specific barcode sequence. The binding event is recorded by binding-dependent primer extension (BD-PEX) ( Figure 2B , step 1). This process produces a DNA duplex containing information about the parent polypeptide and the order and identity of the amino acids. Next, the polypeptide sequence is read by DNA sequencing ( Figure 2B , step 2). This is achieved by combining the DNA encoding the polypeptide sequence generated by successive cycles and sequencing these DNAs on a sequencing platform. The resulting DNA sequencing data is used to reconstruct the polypeptide sequence. This is achieved by assigning DNA with the same UMI to a single parent polypeptide and assigning amino acid sequence and identity using a cycle number barcode and a binder-specific barcode.
[0037] exist Figures 2C-2D A further non-limiting example of a module according to an embodiment of the present disclosure is schematically shown in FIG.
[0038] The method disclosed herein has many advantages over existing methods for peptide sequencing. First, it is universal. That is, it can implement the well-established Edman degradation method, which is compatible with peptide sequences of varying charge and length. The method directly detects degraded fragments, which are extracted from their sequence context and recognized by a binding agent (e.g., an antibody). Affinity-based detection is independent of the chemical properties of the amino acids and can be extended to all proteinogenic amino acids and their post-translational modifications. Second, the method allows peptides to be sequenced with single-molecule sensitivity at single-amino acid resolution. Each cycle of the method removes one amino acid from the N-terminus and barcodes the resulting fragment with DNA encoding the source and position of the amino acid. Reading the DNA-barcoded PTC amino acids using BD-PEX without interference from downstream peptide sequences enables peptide sequencing with single-amino acid resolution. The resulting DNA sequence is amplified during DNA sequencing, enabling peptide sequencing with single-molecule sensitivity. Third, the method can achieve the throughput required for single-cell protein sequencing. This method achieves highly parallel Edman degradation through DNA barcoding and converts peptide sequence information into DNA sequence. Current DNA sequencing technology can already achieve >10 10 reads (e.g., using Illumina Sequencing platform). Given that the throughput of DNA sequencing continues to increase, the throughput of the method of the present invention will reach the throughput required to sequence all proteins from a single cell. Details of embodiments of the method of the present disclosure will now be described.
[0039] The terms "polypeptide," "peptide," and "protein" are used interchangeably herein to refer to a linear series of amino acid residues interconnected by peptide bonds between the α-amino and carboxyl groups of adjacent residues. The amino acids can include the 20 "standard" genetically encodable amino acids, non-natural amino acids (e.g., amino acid analogs), or a combination thereof.
[0040] The term "amino acid" generally refers to any monomeric unit comprising a substituted or unsubstituted amino group, a substituted or unsubstituted carboxyl group and one or more side chains or pendant groups, or analogs of any of these groups. Exemplary side chains include, for example, thiol, seleno, sulfonyl, alkyl, aryl, acyl, keto, azido, hydroxyl, hydrazine, cyano, halogen, hydrazide, alkenyl, alkynyl, ether, borate, boronate, phosphate, phosphono, phosphine, heterocyclic, enone, imine, aldehyde, ester, thioacid, hydroxylamine, or any combination of these groups. Naturally occurring α-amino acids are those encoded by the genetic code as well as those that are later modified (e.g., hydroxyproline, γ-carboxyglutamate, and O-phosphoserine). Naturally occurring α-amino acids include, but are not limited to, alanine (Ala), cysteine (Cys), aspartic acid (Asp), glutamic acid (Glu), phenylalanine (Phe), glycine (Gly), histidine (His), isoleucine (Ile), arginine (Arg), lysine (Lys), leucine (Leu), methionine (Met), asparagine (Asn), proline (Pro), glutamine (Gln), serine (Ser), threonine (Thr), valine (Val), tryptophan (Trp), tyrosine (Tyr), and combinations thereof. Amino acids may be referred to herein by their commonly known three letter symbols or by the one letter symbols recommended by the IUPAC-IUB Biochemical Nomenclature Commission. For example, L-amino acids may be referred to herein by their commonly known three letter symbols (e.g., Arg for L-arginine) or by a capital one letter amino acid symbol (e.g., R for L-arginine). D-amino acids may be referred to herein by either their commonly known three letter symbols (eg, D-Arg for D-arginine) or by the lowercase one letter amino acid symbols (eg, r for D-arginine).
[0041] The polypeptide can be present in any sample of interest, including but not limited to protein samples isolated from a single cell, a plurality of cells (e.g., cultured cells), a tissue, a biological fluid (e.g., whole blood or a fraction thereof, urine, saliva, cerebrospinal fluid, sputum, etc.), an organ, or an organism (e.g., bacteria, yeast, etc.). In certain embodiments, the protein sample is isolated from a cell, tissue, organ, etc. of a mammal (e.g., a human, a rodent (e.g., a mouse), or any other mammal of interest). In other embodiments, the protein sample is isolated from a source other than a mammal, such as bacteria, yeast, an insect (e.g., Drosophila), an amphibian (e.g., a frog (e.g., Xenopus)), a virus, a plant, or any other non-mammalian protein sample source.
[0042] In certain embodiments, the polypeptide to be sequenced is present in a protein sample isolated from a single cell. In some such cases, the method is a single-cell protein sequencing method performed on a plurality of polypeptides present in a protein sample isolated from a single cell.
[0043] Methods, reagents, and kits for isolating proteins from single cells, cell populations, tissues, and the like are known in the art. Non-limiting examples of useful protein extraction kits include ReadyPrep TM Protein extraction kit (Bio-Rad), Qproteome protein isolation kit (Qiagen), T-PER TM Tissue protein extraction reagent (ThermoScientific), M-PER TM Mammalian protein extraction reagent (Thermo Scientific), B-PER TM Complete Bacterial Protein Extraction Reagent (Thermo Scientific), Pierce Plant Total Protein Extraction Kit (ThermoScientific), Triton TM X-100 RIPA buffer (5X) (Thermo Scientific), etc.
[0044] The protein sample used in the method of the present disclosure can be collected by any convenient means. In some cases, a useful cell sample can be a biopsy or can be derived from a biopsy. Biopsy tissue can be obtained from healthy or diseased cells or tissues (including, for example, cancer cells or tissues). Therefore, in some embodiments, the biopsy sample is a tumor biopsy sample. Depending on the type of cancer and / or the type of biopsy performed, samples can be prepared from solid tissue biopsy or liquid biopsy.
[0045] In some cases, protein samples can be prepared from surgical biopsies. Any convenient and appropriate technique for surgical biopsies can be used to collect samples to be employed in the methods described herein, including, but not limited to, for example, excisional biopsies, incisional biopsies, line biopsies, and the like. In some cases, surgical biopsies can be obtained as part of a surgical procedure that has a primary purpose other than obtaining a sample, including, for example, tumor resection, mastectomy, lymph node surgery, axillary lymph node dissection, sentinel lymph node surgery, and the like.
[0046] Various other biopsy techniques can be used to obtain biopsy tissue for use as protein samples as described herein. As a non-limiting example, a sample can be obtained by needle biopsy. Any convenient and appropriate technique for needle biopsy can be used to collect the sample, including but not limited to, for example, fine needle aspiration (FNA), core needle biopsy, stereotactic core biopsy, vacuum-assisted biopsy, and the like.
[0047] According to an embodiment of the polypeptide sequencing method of the present disclosure, the method includes labeling the N-terminal amino acid of the polypeptide with a nucleic acid marker comprising a unique molecular identifier (UMI) and a cycle number barcode. As used herein in the context of the structure of a polypeptide, "N-terminal amino acid" and "C-terminal amino acid" refer to the amino acid at the terminal amino and carboxyl termini of the polypeptide, respectively.
[0048] As used herein, the term "unique molecular identifier (UMI)" or "UMI" refers to a nucleotide sequence that can be used to identify and / or distinguish a first molecule and one or more second molecules to which the UMI is attached. As used herein, a UMI can include one or more nucleotides at one or both ends of the identifying / distinguishing sequence of nucleotides, for example, to facilitate attachment (e.g., linking) of the UMI to different entities. UMIs are typically short, for example, having a length of about 5 to 40 (e.g., about 5 to 20) bases. Typically, UMIs are used to distinguish similar types of molecules within a population or group.
[0049] As used herein, "barcode" or "barcode sequence" refers to a uniquely identifiable nucleotide sequence. In some embodiments, a barcode uniquely identifies the number of degradation cycles (cycle number barcode). The length and composition of the barcode sequence may vary greatly. According to some embodiments, the barcode has a length of 4 to 120 nucleotides, for example, a degenerate sequence of 4 to 100, 4 to 80, 4 to 60, 4 to 40, 6 to 30, 8 to 20 nucleotides, or 10 to 15 nucleotides. In certain embodiments, the barcode has a length of up to 20 nucleotides, for example, a degenerate sequence of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides. The barcode can include one or more mixed bases (e.g., every 3 bases, every 4 bases, etc.) of only three possible base combinations rather than four possible base combinations to prevent homopolymer barcodes.
[0050] In certain embodiments, prior to the labeling step, the polypeptide is immobilized on a solid support via a nucleic acid attached to the C-terminus of the polypeptide and to the surface of the solid support, wherein the nucleic acid comprises a UMI and a primer binding site 3' or 5' of the UMI. According to such embodiments, labeling can include conjugating a degradation moiety to the N-terminal amino acid and conjugating a primer to the degradation moiety, wherein the primer conjugated to the degradation moiety comprises a cycle number barcode and a sequence 3' to the cycle number barcode that is complementary to the primer binding site of the nucleic acid that immobilizes the polypeptide to the solid support. Such labeling can further include annealing the primer conjugated to the degradation moiety to the primer binding site and extending the primer conjugated to the degradation moiety using the nucleic acid that immobilizes the polypeptide to the solid support as a template, thereby labeling the N-terminal amino acid with a nucleic acid label comprising a UMI and a cycle number barcode.
[0051] As described above, according to some embodiments, a polypeptide is immobilized on a solid support via a nucleic acid attached to the C-terminus of the polypeptide and to the surface of the solid support, wherein the nucleic acid comprises a UMI and a primer binding site 3' or 5' to the UMI. According to such embodiments, labeling can include conjugating a degradation moiety conjugated to a primer to the N-terminal amino acid, the primer comprising a cycle number barcode and a sequence 3' to the cycle number barcode that is complementary to the primer binding site of the nucleic acid that immobilizes the polypeptide to the solid support. Such labeling can further include annealing the primer conjugated to the degradation moiety to the primer binding site and extending the primer conjugated to the degradation moiety using the nucleic acid that immobilizes the polypeptide to the solid support as a template, thereby labeling the N-terminal amino acid with a nucleic acid label comprising the UMI and the cycle number barcode.
[0052] The term "solid support" means an insoluble material having a surface to which reagents and / or materials (e.g., polypeptides) can be attached directly or indirectly. In certain embodiments, the collection of solid supports has an average maximum dimension of 750 μm or less, 500 μm or less, 250 μm or less, 100 μm or less, 1 μm or less, 0.75 μm or less, 0.50 μm or less, 0.25 μm or less, or 0.1 μm or less.
[0053] Various materials can be used as solid supports. Support materials include any material that can serve as the support for the attachment of reagents and / or materials. Suitable materials include, but are not limited to, organic or inorganic polymers, natural and synthetic polymers, including but not limited to agarose, cellulose, nitrocellulose, cellulose acetate, other cellulose derivatives, dextran, dextran derivatives and dextran copolymers, other polysaccharides, glass, silica gel, gelatin, polyvinyl pyrrolidone, rayon, nylon, polyethylene, polypropylene, polybutene, polycarbonate, polyester, polyamide, vinyl polymer, polyvinyl alcohol, polystyrene and polystyrene copolymers, cross-linked polystyrene such as divinylbenzene, acrylic resin, acrylate and acrylic acid, acrylamide, polyacrylamide, polyacrylamide blends, copolymers of vinyl and acrylamide, methacrylate, methacrylate derivatives and copolymers, other polymers and copolymers with various functional groups, latex, butyl rubber and other synthetic rubbers, silicon, glass, paper, natural sponges, insoluble proteins, surfactants, metals, metalloids, magnetic materials and any combination thereof.
[0054] The solid support can be any suitable shape, including but not limited to a sphere, an ellipsoid, a rod, a disk, a pyramid, a cube, a cylinder, a nanohelix, a nanospring, a nanotorus, an arrowhead, a teardrop, a tetrapod, a prism, or any other suitable geometric or non-geometric shape.
[0055] In certain embodiments, the solid support is a bead. As used herein, the term "bead" refers to a small mass that is generally spherical or ellipsoidal in shape. According to some embodiments, the beads as used herein have an average diameter of about 0.50 μm to about 500 μm, such as about 0.75 μm to about 250 μm, such as about 1 μm.
[0056] In addition, and for the purposes of this document, the solid support can be magnetically responsive, for example, by comprising one or more paramagnetic and / or superparamagnetic substances (such as, for example, magnetite). Such paramagnetic and / or superparamagnetic substances can be embedded in the matrix of the solid support and / or can be provided on the outer surface and / or inner surface of the solid support (e.g., beads).
[0057] A variety of methods can be used to immobilize polypeptides to solid supports, including those described in detail in the experimental section of this article. For example, in some cases, dibenzocyclooctene (DBCO)-modified DNA sequences are synthesized on controlled pore glass (CPG) or polystyrene-coated carboxylic acid magnetic beads. CpG is commonly used for solid phase DNA synthesis, and magnetic beads are generally compatible with enzymes and allow for easy separation. In one non-limiting embodiment, a PTC-peptide containing a C-terminal azidolysine can be conjugated to DBCO-modified DNA via strain-promoted alkyne-azide cycloaddition (SPAAC) to form a model DNA-peptide conjugate. See, for example Figure 4 A.
[0058] As shown in the experimental section herein, the inventors have determined that Edman degradation can be performed on fixed DNA-peptide conjugates. However, since DNA is unstable under traditional Edman degradation conditions, it is necessary to implement alternative Edman degradation reaction conditions that are compatible with DNA. Suitable alternative Edman degradation reaction conditions identified by the inventors include, but are not limited to, Lewis acids in aprotic solvents. In some cases, the Lewis acid is BF3 etherate, BCl3, BBr3, scandium (III) trifluoromethanesulfonate, or any combination thereof. For example, the Lewis acid can comprise BF3 etherate or consist of BF3 etherate. According to some embodiments, the aprotic solvent is acetonitrile, N,N-dimethylformamide (DMF), dimethyl sulfoxide (DMSO), or any combination thereof. In a non-limiting example, alternative Edman degradation reaction conditions include BF3 etherate in an aprotic solvent, such as 40mM BF3 etherate in anhydrous acetonitrile. Suitable alternatives further include triethylamine acetate in dimethylformamide (DMF), for example at 70°C.
[0059] In order to further enhance the stability of the nucleic acids implemented in the method, one or more non-natural nucleotides that enhance stability may be employed in any DNA utilized in the method. According to some embodiments, one or more DNAs in the DNA utilized in the method comprise one or more nucleotides that increase thermal stability. Non-limiting examples of nucleotides that increase thermal stability include 7-deaza-8-aza-purine-triphosphate, 2-amino-2'-deoxyadenosine-5'-triphosphate (2-amino-dATP), 5-methyl-2'-deoxycytidine-5'-triphosphate (5-Me-dCTP), 5-propynyl-2'-deoxycytidine-5'-triphosphate (5-Pr-dCTP), 5-propynyl-2'-deoxyuridine-5'-triphosphate (5-Pr-dUTP) and or halogenated deoxyuridine (XdU), such as 5-chloro-2'-deoxyuridine-5'-triphosphate (5-Cl-dUTP), 5-bromo-2'-deoxyuridine-5'-triphosphate (5-Br-dUTP) and any combination thereof. In certain embodiments, one or more nucleic acids comprise non-natural nucleotides that stabilize the nucleic acid during the degradation step, a non-limiting example of which is a 7-deaza purine nucleotide. For example, the present disclosure surprisingly demonstrates that polymerases (e.g., Sequenase version 2.0, Klenow (exo-), and Bst 3.0) are able to accept 7-deazapurine nucleotide substituted template-primer duplexes and 7-deazapurine nucleoside triphosphates as substrates. See, e.g., Examples 3 and 4 in the Experimental section below. Figure 5 D.
[0060] In certain embodiments, the method includes conjugating a degradation moiety to the N-terminal amino acid. A "degradation moiety" refers to a moiety that, when conjugated to the N-terminal amino acid of a polypeptide, promotes the cleavage of the N-terminal amino acid from the polypeptide under conditions compatible with the degradation moiety. In a non-limiting example, the degradation moiety employed is an isothiocyanate (ITC). When practicing the methods of the present disclosure, non-limiting examples of ITCs that can be used as degradation moieties include phenyl isothiocyanate (PITC), substituted phenyl isothiocyanate (e.g., para-substituted phenyl isothiocyanate, ortho-substituted phenyl isothiocyanate, meta-substituted phenyl isothiocyanate, pentafluorophenyl isothiocyanate, etc.), alkyl isothiocyanate, naphthyl isothiocyanate, etc. The term "conjugation" or "conjugating" generally refers to a covalent or non-covalent (usually covalent) chemical connection that associates a molecule of interest with the proximal end of a second molecule of interest.
[0061] According to some embodiments, the method includes conjugating a primer to a degradation moiety, wherein the primer conjugated to the degradation moiety comprises a cycle number barcode and a sequence at 3' of the cycle number barcode that is complementary to the primer binding site of the nucleic acid that fixes the polypeptide to the solid support. A variety of methods can be used to conjugate the primer to the degradation moiety. In a non-limiting example, the degradation moiety includes an isothiocyanate (e.g., PITC, etc.) with a reactive group for conjugating to a nucleic acid comprising a cycle number barcode. In certain embodiments, the reactive group is a click chemistry reactive group.
[0062] Click chemistry reactions that can be employed include (i) nucleophilic substitutions; (ii) additions to C-C multiple bonds (e.g., Michael additions, epoxidations, dihydroxylations, aziridinations); (iii) non-aldol chemistry (e.g., N-hydroxysuccinimide active ester couplings); and (iv) cycloadditions (e.g., Diels-Adler reactions, Huisgen cycloadditions). Huisgen cycloadditions have been applied to various branches of chemistry. It consists of the condensation of an organic azide with an alkyne group to form a 1,2,3-triazole bond. Azide and alkyne functional groups can be readily incorporated into large organic scaffolds with biological relevance. The reaction can be catalyzed by the introduction of copper(I). The Cu(I) core has a dual role in that it activates the sluggish alkyne group, thereby accelerating the kinetics of the azide-alkyne condensation by approximately 10. 7 –10 8 times, and it organizes the reactive groups by "templation" so that only regiospecific 1,4-disubstituted adducts are formed. This reaction is called copper-catalyzed azide-alkyne cycloaddition (CuAAC), and its compatibility with a variety of biological substrates and synthetic conditions has made CuAAC a benchmark in click conjugation. Since its discovery, Cu(I)-catalyzed azide-alkyne cycloaddition has been widely used in biology, biochemistry and biotechnology. Click chemistry reactions that can be used include but are not limited to Huisgen azide-alkyne 1,3-dipolar cycloaddition, copper-catalyzed azide-alkyne cycloaddition (CuAAC), ruthenium-catalyzed azide-alkyne cycloaddition (RuAAC), etc. Details on nucleic acid click chemistry can be found in Fantoni et al. (2021) Chem. Rev. 121(12):7122–7154.
[0063] The method of the present disclosure includes one or more annealing steps, for example, annealing a primer conjugated to a degradation portion with a primer binding site of a nucleic acid that fixes the polypeptide to a solid support, annealing a primer with a nucleic acid marker of the N-terminal amino acid being degraded, and / or similar steps. Those of ordinary skill in the art can design the various nucleic acids used in the method of the present disclosure so that they can anneal to each other as needed, for example, by appropriately designing the nucleic acid to have a complementary region. As used herein, the term "complementary" or "complementarity" refers to the nucleotide sequence of a first nucleic acid that is base-paired with a region of a second nucleic acid by a non-covalent bond, or the nucleotide sequence of a first region of a nucleic acid that is base-paired with a second region (for example, a stem region) of a nucleic acid by a non-covalent bond. In classical Watson-Crick base pairing, adenine (A) forms a base pair with thymine (T), just as guanine (G) forms a base pair with cytosine (C) in DNA. In RNA, thymine is replaced by uracil (U). Therefore, A is complementary to T, and G is complementary to C. In RNA, A is complementary to U, and vice versa. Generally, "complementary" or "complementarity" refers to a nucleotide sequence that is at least partially complementary. These terms can also encompass fully complementary duplexes so that each nucleotide in a chain is complementary to each nucleotide in another chain in a corresponding position. In some cases, a nucleotide sequence can be partially complementary to a target, wherein not all nucleotides are complementary to each nucleotide in a target nucleic acid in all corresponding positions. For example, the region of a first nucleic acid can be fully (i.e., 100%) complementary to the region of a second nucleic acid, or the region of the first nucleic acid can share a certain degree of incomplete (e.g., 70%, 75%, 85%, 90%, 95%, 99%) complementarity. The percent identity of two nucleotide sequences can be determined by comparing sequences for optimal comparison purposes (e.g., for optimal comparison, a gap (gap) can be introduced into the sequence of the first sequence). The nucleotides at the corresponding positions are then compared, and the percent identity between the two sequences is a function of the number of identical positions shared by the sequences (i.e., the total number of the number / position of identity %=identical positions × 100). When a position in a sequence is occupied by the same nucleotide as the corresponding position in another sequence, the molecules are identical at that position. Non-limiting examples of such mathematical algorithms are described in Karlin et al., Proc. Natl. Acad. Sci. USA 90:5873-5877 (1993). Such algorithms are incorporated into NBLAST and XBLAST programs (version 2.0), as described in Altschul et al., Nucleic Acids Res. 25:389-3402 (1997). When utilizing BLAST and gapped BLAST programs, the default parameters of the corresponding programs (e.g., NBLAST) can be used.In some embodiments, the parameters for sequence comparison can be set to score = 100, wordlength = 12, or can vary (eg, wordlength = 5 or wordlength = 20).
[0064] The conditions during the annealing step can be those under which the first nucleic acid (e.g., primer) specifically hybridizes to the second nucleic acid (e.g., template nucleic acid). Whether specific hybridization occurs is determined by factors such as the degree of complementarity between the relevant portions of the nucleic acids, their lengths, and the temperature at which hybridization occurs, which may be determined by the melting temperatures (T M The melting temperature is the temperature at which half of the nucleic acids remain hybridized and half dissociate into single strands. The Tm of a duplex can be determined experimentally or using the following formula: Tm = 81.5 + 16.6 (log10 [Na + ])+0.41(fraction G+C)–(600 / N), where N is the chain length, and [Na + ] is less than 1 M. See Sambrook and Russell (2001; Molecular Cloning: A Laboratory Manual, 3rd ed., Cold Spring Harbor Press, Cold Spring Harbor, NY, Chapter 10). Other more advanced models that rely on various parameters can also be used to predict the Tm of nucleic acid duplexes based on various hybridization conditions. Methods for achieving specific nucleic acid hybridizations can be found, for example, in Tijssen, Laboratory Techniques in Biochemistry and Molecular Biology—Hybridization with Nucleic Acid Probes, Part I, Chapter 2, "Overview of principles of hybridization and the strategy of nucleic acid probe assays," Elsevier (1993).
[0065] In certain embodiments, the polypeptide sequencing methods of the present disclosure comprise annealing a primer to a nucleic acid marker of the degraded N-terminal amino acid, wherein the primer comprises a barcode corresponding to the identity of the degraded N-terminal amino acid. In a non-limiting embodiment, during the annealing step, the primer comprising a barcode corresponding to the identity of the degraded N-terminal amino acid is conjugated to a binding moiety that specifically binds to the degraded N-terminal amino acid, and wherein annealing is dependent on binding of the binding moiety to the degraded N-terminal amino acid.
[0066] A variety of suitable binding moieties can be employed, non-limiting examples of which include polypeptide binding moieties (eg, antibodies), small molecules, aptamers, and the like.
[0067] According to some embodiments, binding moiety is an antibody. The term "antibody" can include antibodies or immunoglobulins (e.g., IgG (e.g., IgG1, IgG2, IgG3 or IgG4), IgE, IgD, IgA, IgM, etc.) of any isotype, a full antibody (e.g., an antibody consisting of a tetramer, which in turn consists of two dimers of heavy and light chain polypeptides); a single-chain antibody (e.g., scFv); an antibody fragment (e.g., a full chain or single-chain antibody fragment) that retains specific binding to the cell surface molecules of a target cell, including but not limited to single-chain Fv (scFv), Fab, (Fab')2, (scFv')2, and diabody; a chimeric antibody; a monoclonal antibody, a human antibody, a humanized antibody (e.g., a humanized full antibody, a humanized half antibody, or a humanized antibody fragment, such as a humanized scFv); and a fusion protein comprising the antigen-binding portion of an antibody and a non-antibody protein. According to some embodiments, the antibody is selected from IgG, Fv, single-chain antibody, scFv, Fab, F(ab')2 or Fab'. In certain embodiments, the antibody is a nanobody (an antibody fragment consisting of a single monomeric variable antibody domain - also known as a single domain antibody (sdAb)), a monobody (a synthetic binding protein constructed using a fibronectin type III domain (FN3) as a molecular scaffold), or a bispecific T cell engager (BiTE).
[0068] Immunoglobulin light or heavy chain variable region (V L and V H) consists of "framework" regions (FRs) interrupted by three hypervariable regions (also called "complementarity determining regions" or "CDRs"). The extent of the framework regions and CDRs has been defined (see, E. Kabat et al., Sequences of proteins of immunological interest, 4th ed., U.S. Department of Health and Human Services, Public Health Services, Bethesda, MD (1987); and Lefranc et al., "IMGT, International Immunogenetics Information (IMGT,the internationalImMunoGeneTics information The sequences of the framework regions of different light or heavy chains are relatively conserved across species. The framework region of an antibody, i.e., the combined framework regions of the light and heavy chains, serves to position and align the CDRs. The CDRs are primarily responsible for binding to the epitope of the antigen.
[0069] Thus, an "antibody" encompasses a protein having one or more polypeptides that can be genetically encoded by, for example, immunoglobulin genes or fragments of immunoglobulin genes. Recognized immunoglobulin genes include kappa, lambda, alpha, gamma, delta, epsilon, and mu constant region genes, as well as numerous immunoglobulin variable region genes. Light chains are classified as kappa or lambda. Heavy chains are classified as gamma, mu, alpha, delta, or epsilon, which in turn define the immunoglobulin classes, i.e., IgG, IgM, IgA, IgD, and IgE, respectively.
[0070] As used herein, the term "monoclonal antibody" refers to an antibody obtained from a colony of substantially homogeneous antibodies, that is, except for the possible naturally occurring mutation that may exist in a small amount, the independent antibodies constituting the colony are identical. For example, a monoclonal antibody can be an antibody derived from a single clone (including any eukaryotic, prokaryotic, yeast or phage clone), or produced via a cell-free expression system, rather than a method for producing the same. Monoclonal antibody compositions demonstrate single binding specificity and affinity for a specific epitope. Monoclonal antibodies have a high degree of specificity for a single antigenic site. In addition, in contrast to conventional (polyclonal) antibody preparations that typically include different antibodies for different determinants (epitopes), each monoclonal antibody is for a single determinant on the antigen. The modifier "monoclonal" represents the characteristic of an antibody obtained from a colony of substantially homogeneous antibodies, and should not be construed as needing to produce antibodies by any specific method. Monoclonal antibodies can be prepared using various techniques known in the art, including but not limited to hybridoma, recombinant, yeast display technology, phage display technology, ribosome display technology, DNA display technology, etc. For example, monoclonal antibodies can be prepared by the hybridoma method first described by Kohler et al., Nature 256:495 (1975), or can be prepared by recombinant DNA methods (see, e.g., U.S. Patent No. 4,816,567). "Monoclonal antibodies" can also be isolated from phage antibody libraries using, for example, Clackson et al., Nature 352:624-628 (1991) and Marks et al., J. Mol. Biol. 222:581-597 (1991).
[0071] According to some embodiments, binding moiety is small molecule." small molecule " compound means that molecular weight is 1000 atomic mass units (amu) or smaller compound.In some embodiments, small molecule is 900amu or smaller, 750amu or smaller, 500amu or smaller, 400amu or smaller, 300amu or smaller or 200amu or smaller.In some cases, small molecule is not made up of the repeating molecular unit such as existing in polymer.
[0072] In certain embodiments, the binding moiety is an aptamer. "Aptamer" refers to a nucleic acid (e.g., an oligonucleotide) that has specific binding affinity for a target cell surface molecule. Aptamers exhibit certain desirable properties, such as ease of selection and synthesis, high binding affinity and specificity, and versatile synthetic accessibility.
[0073] The phrases "specifically binds," "specific for," "immunoreactive," "immunoreactivity," and "antigen binding specificity" when referring to a binding moiety (e.g., an antibody, a small molecule, an aptamer, etc.) refer to a binding reaction with an antigen (e.g., a particular amino acid or a post-translationally modified form thereof) that is highly preferential for that antigen, thereby determining the presence of that antigen in the presence of a heterogeneous population of antigens (e.g., a mixture of different amino acids) and / or being selective for that antigen. Thus, under specified conditions, a specified binding moiety binds to a specific antigen and does not bind to other antigens present in the sample in significant amounts. Specific binding to an antigen under such conditions may require a binding moiety that is selected for its specificity for a particular antigen. For example, a binding moiety (e.g., an antibody) can specifically bind to a specific amino acid and not exhibit comparable binding (e.g., not exhibit detectable binding) to other amino acids present in the sample.
[0074] In some embodiments, if the binding moiety is present in an amount greater than or equal to about 10 5 M -1 Affinity or K a A binding moiety "specifically binds" to a particular amino acid if it binds or associates with it at a specific binding constant (i.e., the equilibrium association constant for the specific binding interaction in units of 1 / M). In certain embodiments, the binding moiety binds or associates with a specific amino acid at a specific binding constant of greater than or equal to about 10 6 M -1 , 10 7 M -1 , 10 8 M -1 , 10 9 M -1 , 10 10 M -1 , 10 11 M -1 , 10 12 M -1 or 10 13 M -1 K a Binds to a specific amino acid. "High affinity" binding refers to K a For at least 10 7 M -1 , at least 10 8 M -1 , at least 10 9 M -1 , at least 10 10 M -1 , at least 10 11 M -1 , at least 10 12 M -1 , at least 10 13 M-1 Alternatively, affinity can be defined as having units of M (e.g., 10 -5 M to 10 -13 The equilibrium dissociation constant (K) of a specific binding interaction of D ). In some embodiments, specific binding means that the binding moiety is bound to less than or equal to about 10 -5 M, less than or equal to about 10 -6 M, less than or equal to about 10 -7 M, less than or equal to about 10 -8 M, or less than or equal to about 10 -9 M, 10 -10 M, 10 -11 M, or 10 -12 M or smaller K D Binding to a specific amino acid. The binding affinity of a binding moiety for a specific amino acid can be readily determined using conventional techniques, for example, by competitive ELISA (enzyme-linked immunosorbent assay), equilibrium dialysis, by using surface plasmon resonance (SPR) technology (e.g., a BIAcore 2000 instrument using the general procedures outlined by the manufacturer), by radioimmunoassay, or similar techniques.
[0075] In certain embodiments, the binding moiety specifically binds to a degraded N-terminal amino acid that comprises a post-translational modification (PTM), and wherein the barcode indicates the identity of the degraded N-terminal amino acid and the PTM. PTMs are chemical modifications that play a key role in functional proteomics because they regulate activity, localization, and interactions with other cellular molecules such as proteins, nucleic acids, lipids, and cofactors. PTMs of interest include, but are not limited to, phosphorylation, glycosylation, ubiquitination, nitrosylation, methylation, acetylation, or lipidation.
[0076] Protein phosphorylation, mainly phosphorylation on serine, threonine or tyrosine residues, is one of the most important and well-studied post-translational modifications. Phosphorylation plays a vital role in the regulation of many cellular processes (including cell cycle, growth, apoptosis and signal transduction pathways). Protein glycosylation is considered to be one of the main post-translational modifications, which has a significant impact on the folding, conformation, distribution, stability and activity of proteins. Glycosylation covers a variety of options for adding sugar moieties to proteins, ranging from simple monosaccharide modifications of nuclear transcription factors to highly complex branched polysaccharide changes of cell surface receptors. Carbohydrates in the form of asparagine-linked (N-linked) or serine / threonine-linked (O-linked) oligosaccharides are the main structural components of many cell surface and secretory proteins. Ubiquitin is an 8-kDa polypeptide consisting of 76 amino acids that is linked to the ε-NH2 of lysine in target proteins via the C-terminal glycine of ubiquitin. After the initial monoubiquitination event, ubiquitin polymers may be formed, and then the polyubiquitinated protein is recognized by the 26S proteasome, which catalyzes the degradation of ubiquitinated proteins and the recycling of ubiquitin. S-nitrosylation is a key PTM used by cells to stabilize proteins, regulate gene expression, and provide NO donors, and the production, positioning, activation, and catabolism of SNO are strictly regulated. S-nitrosylation is a reversible reaction, and SNO has a short half-life in the cytoplasm because there are many reductases that denitrosylate proteins, including glutathione (GSH) and thioredoxin. Therefore, SNO is usually stored in membranes, vesicles, interstitial spaces, and lipophilic protein folds to protect them from denitrosylation. Transferring a one-carbon methyl group to the nitrogen or oxygen of the amino acid side chain (N- and O-methylation, respectively) increases the hydrophobicity of the protein and can neutralize the negative charge of the amino acid when combined with a carboxylic acid. Methylation is mediated by methyltransferases, and S-adenosylmethionine (SAM) is the main methyl group donor. N-acetylation, or the transfer of an acetyl group to nitrogen, occurs in almost all eukaryotic proteins by both irreversible and reversible mechanisms. N-terminal acetylation requires cleavage of the N-terminal methionine by methionine aminopeptidase (MAP) before replacing the amino acid with an acetyl group from acetyl-CoA by N-acetyltransferase (NAT). This type of acetylation is co-translational because the N-terminus is acetylated on the growing polypeptide chain still attached to the ribosome. 80% to 90% of eukaryotic proteins are acetylated in this way. Lipidation is a method for targeting proteins to membranes in organelles (endoplasmic reticulum [ER], Golgi apparatus, mitochondria), vesicles (endosomes, lysosomes), and plasma membranes. The four types of lipidation are: C-terminal glycosylphosphatidylinositol (GPI) anchor; N-terminal myristoylation; S-myristoylation; and S-prenylation.Each type of modification confers a different membrane affinity on the protein, although all types of lipidation increase the hydrophobicity of the protein and, therefore, its affinity for membranes. Different types of lipidation are also not mutually exclusive, as two or more lipids can be attached to a given protein.
[0077] As summarized above, the labeling, degradation, annealing, and extension steps are performed in successive cycles to generate a plurality of extension products, each of which comprises a UMI, a corresponding cycle number barcode, and a barcode corresponding to the identity of the corresponding degraded N-terminal amino acid. Once the plurality of extension products has been generated, the method further comprises sequencing the plurality of extension products. As will be understood after reading this disclosure, the sequence of the polypeptide can be determined based on the sequences of the plurality of extension products.
[0078] Sequencing of the plurality of extension products can be performed using any of a variety of available high-throughput nucleic acid sequencers and systems. Illustrative sequencing systems include the Illumina iSeq 100, Miniseq, MiSeq series, NextSeq series (e.g., NextSeq 500 series, NextSeq 1000, NextSeq 2000), and NovaSeq sequencing systems (Illumina, Inc., San Diego, Calif.), Pacific Biosciences Sequel (e.g., Sequel II) sequencing systems (Pacific Biosciences, Menlo Park, Calif.), Oxford Nanopore Technologies MinION TM GridIONx5 TM 、PromethION TM or SmidgION TM Nanopore-based sequencing systems (Oxford Nanopore Technologies, Oxford, UK), as well as other systems with similar capabilities.
[0079] On the Illumina platform, the sequencing process involves cloning amplification adapter-connected DNA fragments on the surface of a glass slide. A cyclic reversible termination strategy is used to read the bases, which sequences the template strand one nucleotide at a time through a progressive cycle of base incorporation, washing, imaging, and cutting. In this strategy, fluorescently labeled 3'-O-azidomethyl-dNTPs are used to pause the polymerization reaction, enabling the removal of unincorporated bases and enabling fluorescence imaging to determine the added nucleotides. After scanning the flow cell with a coupled charge device (CCD) camera, the fluorescent moiety and 3' block are removed and the process is repeated.
[0080] In sequence analysis based on zero-mode waveguides (ZMWs), ZMWs are nanometer-sized holes that serve as optical confinement to allow observation of individual polymerase molecules. As a result, nucleotide incorporation events provide observation of incorporated nucleotide analogs that are easily distinguished from unincorporated nucleotide analogs. For a description of ZMWs and their use in nucleic acid sequencing, see, for example, U.S. Patent Application Publication No. 2003 / 0044781 and U.S. Patent No. 6,917,726, each of which is incorporated herein by reference in its entirety for all purposes. See also Levene et al. (2003) "Zero-mode waveguides for single-molecule analysis at high concentrations," Science 299:682-686; Eid et al. (2009) "Real-time DNA sequencing from single polymerase molecules," Science 323:133-138; and U.S. Pat. Nos. 7,056,676, 7,056,661, 7,052,847, 7,033,764, and 7,907,800, the entire disclosures of which are incorporated herein by reference in their entirety for all purposes.
[0081] In nanopore sequencing, nanopore is used as a biosensor and provides a unique passage through which the ionic solution on the cis side of the membrane contacts the ionic solution on the trans side. A constant bias (positive on the trans side) generates an ionic current through the nanopore and drives the ssDNA or ssRNA in the cis chamber to reach the trans chamber through the hole. Processive enzymes (e.g., helicases, polymerases, nucleases, etc.) can be incorporated into polynucleotides so that they gradually control motion and allow nucleotides to pass through small-diameter nanopores one core base at a time. Because the ionic conductivity through the nanopore is sensitive to the presence of the quality of the core base and its associated electric field, the ionic current level through the nanopore reveals the sequence of the core base in the translocation chain. Patch clamps, voltage clamps, etc. can be used.
[0082] Details for obtaining raw sequencing reads of nucleic acid molecules using nanopores are described in, for example, Feng et al. (2015) Genomics, Proteomics & Bioinformatics 13(1): 4-16. Nanopore-based sequencing systems are available and include the SmidgION, MinION, GridION, and PromethION nanopore-based sequencing systems available from Oxford Nanopore Technologies Limited. Detailed design considerations and protocols for performing nucleic acid sequencing are provided by such systems.
[0083] The method of the present disclosure can be carried out in any suitable container / confinement. One or more steps of the method can be carried out in a first container, and one or more other steps can be carried out in a second container. Non-limiting examples of containers in which one or more steps of the method can be carried out include tubes, bottles, plates, holes of multi-well plates (e.g., 6-well plates, 12-well plates, 24-well plates, 48-well plates, 96-well plates or 384-well plates), confinements in microfluidic devices, etc.
[0084] Compositions and kits
[0085] Aspects of the present disclosure further include compositions. In some embodiments, a composition is provided that includes one or more of any polypeptide and / or reagents for performing the polypeptide sequencing methods of the present disclosure described elsewhere herein, or any combination thereof. Non-limiting examples of reagents and combinations thereof that may be present in the compositions of the present disclosure include UMI-functionalized solid supports, degradation moieties with reactive groups for conjugation to nucleic acids, primers containing cycle number barcodes, Edman degradation reagents (including those for providing alternative DNA-compatible Edman degradation conditions described elsewhere herein), conjugates comprising primers conjugated to binding moieties that specifically bind to amino acids, conjugates comprising primers conjugated to binding moieties that specifically bind to amino acids bearing post-translational modifications, nucleic acid sequencing aptamers, and any combination thereof.
[0086] According to some embodiments, the compositions of the present disclosure include one or any combination of any polypeptide and / or reagent present in a liquid medium. The liquid medium can be an aqueous liquid medium, such as water, a buffer solution, etc. One or more additives, such as salts (e.g., NaCl, MgCl2, KCl, MgSO4), buffers (Tris buffer, N-(2-hydroxyethyl)-piperazine-N'-(2-ethanesulfonic acid) (HEPES), 2-(N-morpholino)-ethanesulfonic acid (MES), 2-(N-morpholino)-ethanesulfonic acid sodium salt (MES), 3-(N-morpholino) propanesulfonic acid (MOPS), N-tris[hydroxymethyl]methyl-3-aminopropanesulfonic acid (TAPS) etc.), solubilizers, detergents (e.g., nonionic detergents such as Tween 20 etc.), nuclease inhibitors, glycerol, chelating agents etc. can be present in such compositions.
[0087] Theme composition can be present in any suitable environment.According to one embodiment, composition is present in reaction tubes (for example, 0.2mL pipe, 0.6mL pipe, 1.5mL pipe etc.) or hole.In some aspects, composition is present in two or more (for example, a plurality of) reaction tubes or hole (for example, plate, such as 6 orifice plates, 12 orifice plates, 24 orifice plates, 48 orifice plates, 96 orifice plates or 384 orifice plates).Pipe and / or plate can be made by any suitable material (for example, polypropylene etc.).In some aspects, wherein there is the pipe and / or plate of composition and provide to the effective heat transfer (for example, when being placed in heating block, water-bath, thermal cycler etc.) of composition, make it possible to change the temperature of composition in short period of time, for example, according to the needs of specific degraded or enzymatic reaction that occurs.According to some embodiments, composition is present in thin-walled polypropylene tube or the plate with thin-walled polypropylene hole.
[0088] Other suitable environments for the subject compositions include, for example, microfluidic chips (e.g., "lab-on-a-chip devices"). The composition can be present in an instrument configured to bring the composition to a desired temperature (e.g., a temperature-controlled water bath, a heat block, etc.). The instrument configured to bring the composition to a desired temperature can be configured to bring the composition to a series of different desired temperatures, each for a suitable period of time (e.g., the instrument can be a thermal cycler).
[0089] Aspects of the present disclosure also include kits. The kits may include, for example, one or any combination of reagents for performing the polypeptide sequencing methods of the present disclosure described elsewhere herein. Non-limiting examples of reagents and combinations thereof that may be present in the compositions of the present disclosure include UMI-functionalized solid supports, degradation moieties with reactive groups for conjugation to nucleic acids, primers comprising cycle number barcodes, Edman degradation reagents (including those for providing alternative DNA-compatible Edman degradation conditions described elsewhere herein), conjugates comprising primers conjugated to binding moieties that specifically bind to amino acids, conjugates comprising primers conjugated to binding moieties that specifically bind to amino acids bearing post-translational modifications, nucleic acid sequencing aptamers, and any combination thereof.
[0090] According to some embodiments, a subject kit includes one or any combination of the following reagents: (i) a UMI-functionalized solid support; (ii) a degradation moiety with a reactive group for conjugating to a nucleic acid; (iii) a primer comprising a cycle number barcode; (iv) an Edman degradation reagent; (v) a conjugate comprising a primer conjugated to a binding moiety that specifically binds an amino acid; (vi) a conjugate comprising a primer conjugated to a binding moiety that specifically binds to an amino acid bearing a post-translational modification; and (vii) a nucleic acid sequencing aptamer. According to some embodiments, the degradation moiety includes a PITC with a reactive group for conjugating to the primer comprising the cycle number barcode. In some cases, the reactive group is a click chemistry reactive group. In certain embodiments, the Edman degradation reagent includes BF3 etherate and an aprotic solvent. According to some embodiments, the Edman degradation reagent includes triethylamine acetate and N,N-dimethylformamide (DMF). In certain embodiments, one or more of the nucleic acid-based reagents includes a non-natural nucleotide (e.g., a 7-deazapurine nucleotide) that stabilizes nucleic acids under Edman degradation conditions. In some cases, the binding moiety is a polypeptide (e.g., an antibody). In other cases, the binding moiety is a small molecule or an aptamer.
[0091] The components of the kit may be present in separate containers, or multiple components may be present in a single container. For example, two or more components of the kit may be provided in a single tube, or may be provided in different tubes.
[0092] In addition to the components mentioned above, the kit of the present invention may further include instructions for use of one or any combination of reagents, for example, to perform any of the polypeptide sequencing methods of the present invention. Instructions are generally recorded on a suitable recording medium. For example, instructions can be printed on a substrate (such as paper or plastic, etc.). Therefore, instructions can be present in the kit as a package insert, present in the label of the container of the kit or its component (that is, associated with packaging or sub-packaging), etc. In other embodiments, instructions exist as an electronic storage data file present in a suitable computer-readable storage medium, and the computer-readable storage medium is, for example, a CD-ROM, a disk, a hard disk drive (HDD), etc. In other embodiments, there are no actual instructions in the kit, but there is provided a means for obtaining instructions from a remote source (for example, via the Internet). The example of this embodiment is a kit comprising a website, in which instructions can be viewed and / or instructions can be downloaded from the website. As with instructions, this means for obtaining instructions is recorded on a suitable substrate.
[0093] For purposes of completeness, the present disclosure is further defined in the following numbered clauses.
[0094] 1. A method for sequencing a polypeptide, the method comprising:
[0095] (a) labeling the N-terminal amino acid of a polypeptide with a nucleic acid tag comprising a unique molecular identifier (UMI);
[0096] (b) degrading the N-terminal amino acid from the polypeptide;
[0097] (c) annealing a primer to the nucleic acid marker of the degraded N-terminal amino acid, wherein the primer comprises a barcode corresponding to the identity of the degraded N-terminal amino acid;
[0098] (d) extending the primer annealed to the nucleic acid marker to generate an extension product, wherein the extension product comprises the UMI and the barcode corresponding to the identity of the degraded N-terminal amino acid;
[0099] (e) performing steps (a)-(d) in consecutive cycles to generate a plurality of extension products, each extension product in the plurality of extension products comprising the UMI, a corresponding cycle number barcode, and a barcode corresponding to the identity of a corresponding degraded N-terminal amino acid;
[0100] (f) sequencing the plurality of extension products; and
[0101] (g) determining the sequence of the polypeptide based on the sequences of the plurality of extension products.
[0102] 2. The method of clause 1, comprising indexing the extension products produced in step (d) with the cycle number barcode.
[0103] 3. The method of clause 1, wherein step (a) comprises labeling the N-terminal amino acid of the polypeptide with a nucleic acid marker comprising the UMI and the cycle number barcode.
[0104] 4. The method of any one of clauses 1-3, wherein prior to step (a), the polypeptide is immobilized on the solid support via a nucleic acid attached to the C-terminus of the polypeptide and to the surface of the solid support, wherein the nucleic acid comprises the UMI and a primer binding site 3' or 5' to the UMI.
[0105] 5. The method of clause 4, wherein the marking step (a) comprises:
[0106] (i) conjugating a degradation moiety to the N-terminal amino acid;
[0107] (ii) conjugating a primer to the degradation moiety, wherein the primer conjugated to the degradation moiety comprises a sequence complementary to the primer binding site of the nucleic acid that immobilizes the polypeptide to the solid support;
[0108] (iii) annealing the primer conjugated to the degradation moiety to the primer binding site; and
[0109] (iv) using the nucleic acid that immobilizes the polypeptide to the solid support as a template to extend the primer conjugated to the degradation moiety, thereby labeling the N-terminal amino acid with the nucleic acid label comprising the UMI.
[0110] 6. The method of clause 4, wherein the marking step (a) comprises:
[0111] (i) conjugating a degradation moiety conjugated to a primer to the N-terminal amino acid, the primer comprising a sequence complementary to the primer binding site of the nucleic acid that immobilizes the polypeptide to the solid support;
[0112] (ii) annealing the primer conjugated to the degradation moiety to the primer binding site; and
[0113] (iii) using the nucleic acid that immobilizes the polypeptide to the solid support as a template to extend the primer conjugated to the degradation moiety, thereby labeling the N-terminal amino acid with the nucleic acid label comprising the UMI.
[0114] 7. The method of clause 5 or clause 6, wherein the degradation moiety comprises phenyl isothiocyanate (PITC) carrying a reactive group for conjugation to the nucleic acid comprising the cycle number barcode.
[0115] 8. The method according to clause 7, wherein the reactive group is a click chemistry reactive group.
[0116] 9. The method according to any one of clauses 1 to 8, wherein the degradation step (b) is carried out under conditions comprising a Lewis acid in an aprotic solvent.
[0117] 10. The method of clause 9, wherein the Lewis acid is BF3 etherate, BCl3, BBr3, scandium(III) trifluoromethanesulfonate, or any combination thereof.
[0118] 11. The method according to clause 9, wherein the Lewis acid is BF3 etherate.
[0119] 12. The method according to any one of clauses 9-11, wherein the aprotic solvent is acetonitrile, N,N-dimethylformamide (DMF), dimethyl sulfoxide (DMSO), or any combination thereof.
[0120] 13. The method according to clause 12, wherein the degradation step (b) is carried out under conditions comprising triethylamine acetate in N,N-dimethylformamide (DMF).
[0121] 14. The method according to any one of clauses 1 to 13, wherein the one or more nucleic acids employed in step (a) and / or step (b) comprise non-natural nucleotides that stabilize the nucleic acid during degradation step (b).
[0122] 15. The method of clause 14, wherein the non-natural nucleotide comprises a 7-deaza purine nucleotide.
[0123] 16. A method according to any one of clauses 1 to 15, wherein in step (c), the primer comprising the barcode corresponding to the identity of the degraded N-terminal amino acid is conjugated to a binding portion that specifically binds to the degraded N-terminal amino acid, and wherein the annealing is dependent on binding of the binding portion to the degraded N-terminal amino acid.
[0124] 17. A method according to clause 16, wherein the binding moiety specifically binds to a degraded N-terminal amino acid comprising a post-translational modification, and wherein the barcode indicates the identity of the degraded N-terminal amino acid and the post-translational modification.
[0125] 18. The method according to clause 16 or clause 17, wherein the post-translational modification is phosphorylation, glycosylation, ubiquitination, nitrosylation, methylation, acetylation or lipidation.
[0126] 19. A method according to any one of clauses 16 to 18, wherein the binding moiety is a polypeptide.
[0127] 20. The method according to clause 19, wherein the polypeptide is an antibody.
[0128] 21. The method according to any one of clauses 16 to 18, wherein the binding moiety is a small molecule or an aptamer.
[0129] 22. The method according to any one of clauses 1 to 21, wherein the polypeptide to be sequenced is present in a protein sample isolated from a single cell.
[0130] 23. The method according to clause 22, wherein the method is a single-cell protein sequencing method for a plurality of polypeptides present in the protein sample.
[0131] 24. The method according to any one of clauses 1 to 23, wherein the polypeptide to be sequenced is present in a protein sample isolated from a tissue sample.
[0132] 25. The method according to clause 24, wherein the tissue sample is a biopsy sample.
[0133] 26. The method according to clause 25, wherein the biopsy sample is a tumor biopsy sample.
[0134] 27. A method according to any one of clauses 1 to 23, wherein the polypeptide to be sequenced is present in a protein sample isolated from a biological fluid.
[0135] 28. A composition comprising one or any combination of the following:
[0136] (a) UMI-functionalized solid support,
[0137] (b) a degradation moiety carrying a reactive group for conjugation to a nucleic acid,
[0138] (c) primers comprising a cycle number barcode,
[0139] (d) Edman degradation reagent,
[0140] (e) a conjugate comprising a primer conjugated to a binding moiety that specifically binds to an amino acid,
[0141] (f) a conjugate comprising a primer conjugated to a binding moiety that specifically binds to an amino acid bearing a post-translational modification, and
[0142] (g) Nucleic acid sequencing aptamers.
[0143] 29. A kit comprising:
[0144] (a) One or any combination of the following:
[0145] (i) UMI-functionalized solid support,
[0146] (ii) a degradation moiety carrying a reactive group for conjugation to a nucleic acid,
[0147] (iii) primers containing cycle number barcodes,
[0148] (iv) Edman degradation reagent,
[0149] (v) a conjugate comprising a primer conjugated to a binding moiety that specifically binds to an amino acid,
[0150] (vi) a conjugate comprising a primer conjugated to a binding moiety that specifically binds to an amino acid bearing a post-translational modification, and
[0151] (vii) nucleic acid sequencing aptamers; and
[0152] (b) instructions for using one or any combination of the reagents to perform a method according to any one of clauses 1 to 27.
[0153] 30. The kit according to clause 29, wherein the degradation moiety comprises PITC carrying a reactive group for conjugation to the primer comprising the cycle number barcode.
[0154] 31. The kit according to clause 30, wherein the reactive group is a click chemistry reactive group.
[0155] 32. The kit according to any one of clauses 29 to 31, wherein the Edman degradation reagent comprises a Lewis acid in an aprotic solvent.
[0156] 33. The kit according to clause 32, wherein the Lewis acid is BF3 etherate, BCl3, BBr3, scandium(III) trifluoromethanesulfonate, or any combination thereof.
[0157] 34. The kit according to clause 32, wherein the Lewis acid is BF3 etherate.
[0158] 35. The kit according to any one of clauses 32 to 34, wherein the aprotic solvent is acetonitrile, N,N-dimethylformamide (DMF), dimethyl sulfoxide (DMSO), or any combination thereof.
[0159] 36. The kit according to any one of clauses 29 to 35, wherein one or more of the nucleic acid-based reagents comprises non-natural nucleotides that stabilize the nucleic acid under Edman degradation conditions.
[0160] 37. The kit according to clause 36, wherein the non-natural nucleotide comprises a 7-deazapurine nucleotide.
[0161] 38. A kit according to any one of clauses 29 to 37, wherein the binding moiety is a polypeptide.
[0162] 39. The kit according to clause 38, wherein the polypeptide is an antibody.
[0163] 40. The kit according to any one of clauses 29 to 37, wherein the binding moiety is a small molecule or an aptamer.
[0164] The following examples are offered by way of illustration only and not by way of limitation.
[0165] Experimental
[0166] Example 1 - Development of an alternative Edman degradation reaction compatible with DNA
[0167] Conventional Edman degradation conditions are incompatible with DNA. The cleavage and conversion reactions are performed with pure trifluoroacetic acid (TFA) and a TFA solution at elevated temperatures, respectively. The use of strong proton acids poses the greatest challenge to DNA, which is easily depurinated under acidic conditions. Here, the cleavage reaction conditions were observed to degrade poly(dT) oligonucleotides. Furthermore, PITC and its derivatives, used to modify the N-terminus of peptides, may react with the exocyclic amines of the nucleobases.
[0168] In this example, the development of an alternative Edman degradation reaction compatible with DNA is described. It is assumed that the degradation of DNA is mainly caused by the protonation of the nucleobase under strongly acidic conditions. Therefore, an Edman degradation procedure using BF3 etherate in an aprotic solvent for the cleavage step was adopted. Consistent with this hypothesis, polypyrimidine sequences are stable under these conditions for a long period of time. However, under such conditions, natural purine nucleotides still undergo depurination, although at a significantly slower rate. In order to further improve the stability of DNA, chemically modified purine nucleotides were studied. 7-deaza purine nucleotides lack a nitrogen atom at the 7-position and are reported to be resistant to depurination.
[0169] Oligonucleotides containing 7-deazapurine nucleotides were stable under degradation conditions for a duration of 4 hours, as confirmed by LC and MS ( Figure 3 A and Figure 3 C). Considering the rapid N-terminal amino acid cleavage in the presence of BF3 etherate (see below), the stability of 7-deazapurine-modified DNA is sufficient for Edman degradation. Finally, 7-deazapurine-modified DNA was subjected to PITC in water / pyridine (1:1) at 50°C for 16 hours, and no DNA modification was detected. This is consistent with the low nucleophilicity of the exocyclic amine of the nucleobase.
[0170] In the presence of 40 mM BF3 etherate, cleavage of the N-terminal amino acid was completed within 5 minutes. The cleavage induced by BF3 etherate was significantly faster than the TFA cleavage reaction, which required 30 minutes. Anilinothiazolinone (ATZ) amino acids were the major fragments produced under these conditions, and ATZ amino acids were converted to stable PTC amino acids ( Figure 3 B) PTC amino acids can be readily synthesized by reacting PITC or its derivatives with amino acids, and thus allow easy access to target molecules for binding agent (eg, antibody) production.
[0171] Example 2— Development of conditions for solid-phase Edman degradation of DNA-peptide conjugates
[0172] This embodiment relates to converting DNA-compatible Edman degradation into a solid phase format. Solid phase reactions bring two benefits. First, solid phase reactions allow the use of large excesses of reagents that can be easily removed by filtration. This greatly simplifies the design of iterative Edman degradation cycles. In addition, due to the insolubility of oligonucleotides in organic solvents, DNA-conjugated peptides have different reactivities in these organic solvents compared to unconjugated peptides, which has been well documented in the field of DNA-encoded libraries. The initial experiments performed here are consistent with the literature and show that no Edman degradation occurs on DNA-conjugated PTC-peptides in anhydrous acetonitrile. It has been reported that fixing DNA on a solid phase makes chemical transformations in non-aqueous solvents easy to perform DNA-encoded synthesis. In view of this, Edman degradation of peptide-DNA conjugates on solid supports was studied.
[0173] As demonstrated herein, Edman degradation can be performed on immobilized DNA-peptide conjugates. To test Edman degradation on a solid support, it was first determined that the stability of the DNA under degradation was not affected by immobilization. Next, DNA-peptide conjugates were synthesized on a solid support to test the feasibility of Edman degradation. This was achieved by synthesizing dibenzocyclooctyne (DBCO)-modified DNA sequences on controlled pore glass (CPG) or polystyrene-coated carboxylic acid magnetic beads. CPG is commonly used for solid phase DNA synthesis, and magnetic beads are generally compatible with enzymes and allow for easy separation. PTC-peptides containing a C-terminal azidolysine were conjugated to DBCO-modified DNA via strain-promoted alkyne-azide cycloaddition (SPAAC) to form model DNA-peptide conjugates ( Figure 4 A). Cleavage reactions were performed using 40 mM BF3 etherate in anhydrous acetonitrile. The supernatants were collected and the release of the PTC amino acid was confirmed by LC-MS on two solid supports ( Figure 4 B). DNA-peptide conjugates were cleaved from CPG after degradation and analyzed by HPLC. The results showed that degradation was complete within 10 minutes ( Figure 4 C).
[0174] Example 3 - Barcoding of degradation fragments of the N-terminal amino acid of a model peptide
[0175] In this example, the DNA-compatible Edman degradation reaction described above will be used to implement the first cycle of INDEED. To achieve this goal, a 7-deazapurine-modified DNA (UMI) with 3'-amino and 5'-DBCO groups will be synthesized. The UMI is immobilized on 1 μm carboxylic acid Dynabeads using carbodiimide chemistry. TMAbove. A peptide containing a C-terminal azidolysine is conjugated to DNA via SPAAC. Secondly, the N-terminus of the peptide is modified with a PITC derivative (2) bearing an alkyne. Thirdly, methyltetrazine azide (3) is conjugated to the alkyne via copper (I)-catalyzed azide-alkyne cycloaddition (CuAAC), and a trans-cyclooctene (TCO)-modified primer is installed via an inverse electron demand Diels-Alder (IEDDA) reaction. In the fourth step, the DNA UMI is transcribed via a primer extension reaction. In the final step, the N-terminal amino acid is cleaved from the peptide by treatment with BF3 etherate ( Figure 5 B) This step is achieved through a tandem cleavage-hydrolysis reaction that generates PTC amino acid fragments barcoded with DNA UMIs.
[0176] Experiments conducted to investigate several key steps in DNA barcoding of degradation fragments established that Edman degradation could be performed on intact UMI-peptide-primer constructs. Peptide-UMI conjugates were immobilized on beads using the conditions established above. Alkyne-modified PITC (2) was synthesized by first treating 4-(2-aminoethyl)aniline with an alkyne NHS ester (1) at 0°C to achieve selective acylation of the aliphatic amine. The aniline moiety was then converted to an isothiocyanate ( Figure 5 A). Modification of the on-bead peptide-UMI conjugate by (2) was quantitative. The primers were successfully installed by the aforementioned CuAAC-IEDDA cascade to generate the complete UMI-peptide-primer construct ( Figure 5 B). The relative amount of primer sequence on the beads can be quantified by flow cytometry via annealing to fluorescently labeled complementary strands. The yield of Edman degradation can be determined by comparing the fluorescence intensity before and after the reaction. The degradation yield is approximately 85% after 10 minutes ( Figure 5 C).
[0177] Next, we determined that 7-deazapurine-modified DNA was accepted by DNA polymerases. Barcoding of PTC amino acids requires that the polymerase accept the 7-deazapurine nucleotide-substituted template-primer duplex and 7-deazapurine nucleoside triphosphate as substrate. We examined primer extension using Sequenase version 2.0, Klenow (exo-), and Bst 3.0. Figure 5 D) All of these polymerases are able to efficiently incorporate 7-deazapurine nucleotide triphosphates, and Sequenase version 2.0 will be used in future experiments.
[0178] Example 4 - Alternative Methods for Barcoding Degraded Fragments of the N-Terminal Amino Acids of Model Peptides
[0179] In this example, a 3'-dibenzocyclooctene (DBCO) modified deazapurine substituted DNA template was immobilized on magnetic beads ( Figure 5 F). Subsequently, a 7-amino acid long model peptide with a C-terminal azidolysine was conjugated to DNA via strain-promoted alkyne-azide cycloaddition (SPAAC). Azide-modified PITC (4, Figure 5 E) Reaction with the N-terminus of a model peptide.
[0180] To introduce primers for barcode transfer, conjugation of an alkyne-modified primer to the azide group on 4 was employed. SPAAC was chosen because of the sensitivity of the phenylthiocarbamoyl group in 2 to oxidation, which could lead to undesirable reactions with reactive oxygen species generated during copper(I)-catalyzed azide-alkyne cycloaddition (CuAAC). One disadvantage of SPAAC is its relatively slow reaction rate. However, it has been determined that the reaction rate can be increased by hybridizing the template and primer, which increases the effective concentration of the reactants. For example, in the presence of a 7-amino acid long peptide, primer conjugation was significantly complete within 10 minutes.
[0181] After incorporation of the primers, they were extended using Klenow ( Figure 5 F), and the beads were subjected to Lewis acid-catalyzed Edman degradation. The insolubility of DNA in acetonitrile was exploited during this step, which allowed the DNA-barcoded ATZ amino acids to remain hybridized to the template DNA after the initial cleavage reaction. Subsequently, the basic conversion reaction served two purposes: (i) it converted the ATZ amino acids into stable PTC amino acids, and (ii) it denatured the double-stranded DNA, releasing the DNA-barcoded PTC amino acids. In addition, DTT was introduced to inhibit the oxidative degradation of PTC amino acids under alkaline conditions. This cleavage and conversion reaction produced an overall yield of 96%, and the resulting structure was confirmed by mass spectrometry analysis ( Figure 5 G).
[0182] Example 5 - Characterization of binders that recognize PTC amino acids
[0183] BD-PEX has the function of converting the binding events between DNA-barcoded PTC amino acids and binders (e.g., antibodies) into DNA output. PTC amino acids retain the structural characteristics of the original amino acids and differ only in the PTC modification on the amino group. Because these amino groups are often modified to conjugate with carrier proteins during the generation of antibodies against amino acids, it is expected that antibodies generated against amino acids will also recognize PTC amino acids. Furthermore, it is expected that commercially available antibodies can be used to detect these amino acids. The fingerprint of four amino acids is sufficient to identify most proteins in the human proteome.
[0184] In this example, it is first demonstrated that antibodies directed against amino acids recognize the PTC amino acid. As a proof of principle, PTC-tryptophan was synthesized by reacting a PITC derivative (1) with tryptophan and conjugated to azide-modified DNA. This conjugate is structurally identical to the conjugate formed by INDEED, except for the difference in the linker at the distal end of the PTC-tryptophan. The binding affinity of a commercially available tryptophan mAb to DNA-conjugated PTC-tryptophan was measured by biolayer interferometry (BLI) and surface plasmon resonance (SPR). The tryptophan mAb had a Kd of 280 nM for PTC-tryptophan, and the binding was highly specific, as no binding was observed to PTC-tyrosine and PTC-phenylalanine ( Figure 6 A). In addition, an antibody against phosphotyrosine (PY20) was also tested and a Kd of 20 nM was obtained ( Figure 6 B). This suggests that PTC amino acids with post-translational modifications (PTMs) may also be recognized by their corresponding anti-PTM antibodies. Additional anti-PTM antibodies were tested, leading to the identification of antibodies that recognize PTCs of asymmetric dimethylarginine (ADMA), acetyl-lysine, and phosphoserine ( Figure 6 E).
[0185] As demonstrated above, PTC amino acids with click handles can be easily synthesized and conjugated to azide-modified carrier proteins such as BSA. In this example, it was next demonstrated that antibodies against PTC amino acids can be generated. In order to obtain new PTC amino acid-specific antibodies, an antibody discovery activity was initiated. PTC-tyrosine and PTC-phenylalanine were synthesized and conjugated to azide-modified BSA. Mice were challenged with modified BSA, and a strong immune response was observed against both antigens. Subsequently, antibody-producing B cells were harvested and fused to form hybridomas. In subsequent subcloning and screening, several candidates with high specificity were identified ( Figure 6 C and Figure 6 D). Using this strategy, antibodies against PTC-modified phenylalanine, tyrosine, tryptophan, arginine, and aspartic acid were identified ( Figure 6 F)
[0186] Example 6—Conversion of DNA barcoded degradation fragments into DNA sequences by binding-dependent primer extension (BD- Development of PEX
[0187] Highly specific binding agents (e.g., antibodies) will be used to recognize PTC amino acids released by Edman degradation of DNA-encoded proteins, and the binding events will be converted into DNA output. Conversion to DNA output can reveal all antibody-antigen interactions simultaneously, and the resulting DNA can be further amplified to increase detection sensitivity. To achieve this, and referring to Figure 7A , will use SiteClick TM The kit performs site-specific modification of the Fc region of an antibody to introduce an azide functionality. Subsequently, a DBCO-modified primer carrying an antibody-specific barcode is conjugated to the antibody. The primer sequence is designed to be short and, therefore, not conducive to intermolecular primer extension. Binding of the antibody to its PTC amino acid target increases the effective molar concentration of the primer, promoting duplex formation between the primer and the complementary sequence on the template. The resulting complex can be used as a substrate for primer extension, which converts the binding event into a sequenceable DNA output. The primer extension products can be analyzed by qPCR and / or DNA sequencing to determine the reaction yield and detection limit.
[0188] Example 7—Conversion of DNA barcoded degradation fragments into DNA sequences using biotinylated primers DNA-dependent primer extension (BD-PEX)
[0189] This example describes the use of biotinylated primers during the INDEED process, enabling barcoded DNA to be pulled down by streptavidin beads and BD-PEX to be performed on the beads ( Figure 7B ). To achieve this, and refer to Figure 7B , use a tool like SiteClick TM The kit or oYo-Link kit performs site-specific modification on the Fc region of the antibody to introduce a click handle, such as an azide or tetrazine. Subsequently, a DBCO- or TCO-modified primer carrying an antibody-specific barcode is conjugated to the antibody. The stability of the primer-template complex plays a crucial role in determining the efficiency and specificity of intramolecular primer extension. Although increasing the length of the primer can improve the efficiency of primer extension, longer primers also tend to promote extension in the absence of an antibody-antigen recognition event. It was found that primers with a length of 6 nucleotides allowed efficient primer extension and simultaneously minimized nonspecific primer extension (Figure 7C). In addition, intermolecular primer extension was used as another source of nonspecific primer extension on the beads. This form of nonspecific primer extension can be effectively inhibited by reducing the surface density of the PTC amino acids barcoded by the DNA. At a density of 1 pmol of DNA per mg of magnetic beads, intermolecular primer extension was almost eliminated (Figure 7E).
[0190] Example 8 - Sequencing / fingerprinting of short peptides
[0191] The peptide with the sequence RGFDWGK{N3} was subjected to five cycles of the INDEED process. The resulting DNA-barcoded PTC amino acids from each cycle were pulled onto streptavidin beads in separate containers. Proximity primer extension was performed using a mixture of DNA-barcoded PTC amino acid-specific antibodies (100 nM each of anti-PTC-Arg antibody, anti-PTC-Phe antibody, anti-PTC-Asp antibody, and anti-PTC-Trp antibody). After primer extension, adapter PCR was performed in each container using adapter primers with cycle number barcodes. Finally, all DNA was pooled, indexed, and sequenced on a MiSeq sequencer ( Figure 8A ). Cycle number barcodes and antibody barcodes were extracted from the sequencing results. The read counts of all possible combinations of barcodes were plotted on a heat map ( Figure 8B ), and the results were consistent with the sequence of the peptide.
[0192] Example 9 - Sequencing / fingerprinting of short peptides and quantification of single amino acid substitutions
[0193] Single amino acid replacement is caused by non-synonymous single nucleotide polymorphism (nsSNP), and usually destroys the function of protein by changing protein structure. DNA sequencing allows sensitive detection of SNP, but detection of single amino acid replacement by MS is usually subject to the limitation of sensitivity. It is expected that the method described herein will be able to realize the detection of single amino acid replacement by sequencing peptide with single amino acid resolution. In addition, DNA sequencing is read out and will allow signal amplification, and therefore improves sensitivity.
[0194] To display peptide sequencing via INDEED, a mixture of peptides with single amino acid substitutions will be immobilized on UMI-coated magnetic beads using the chemical method established above. INDEED will be performed iteratively, and DNA-barcoded PTC amino acids will be collected. Prior to the enzymatic reaction, the organic solvent from the degradation fragment mixture will be removed by solid phase extraction or buffer exchange. The combined DNA-barcoded PTC amino acids are then converted into DNA sequences by BD-PEX using a mixture of DNA-barcoded antibodies against PTC amino acids. The resulting DNA will be sequenced. It is expected that single amino acid substitutions can be identified regardless of their position in the polypeptide. In addition to identifying single amino acid substitutions in a highly parallel manner, the sequencing method of the present invention will also allow these variations to be quantified using read counts. If each polypeptide molecule is uniquely barcoded, the occurrence of single amino acid substitutions can be digitally quantified, and this quantification will not be affected by PCR bias.
[0195] Example 10—Mapping and Quantification of Post-Translational Modifications (PTMs) in Model Peptides
[0196] The identification and quantification of PTMs are crucial for understanding protein function. Currently, PTMs are typically studied using antibody-based techniques and mass spectrometry. However, antibody-based techniques are generally not site-specific. Although PTM analysis by MS can provide information about the site of modification, accurate quantification of PTMs generally requires the use of chemically synthesized isotope-labeled peptide standards. It is expected that most PTM-specific antibodies can be used in the sequencing method disclosed herein. Because PTC amino acid fragments are barcoded with DNA containing information about the order of amino acids within the peptide, the method of the present invention can specifically map PTM sites. In addition, quantification of PTMs can be achieved by DNA sequencing without the need to synthesize isotope-labeled peptide standards that are specific to the protein of interest.
[0197] To demonstrate this capability, a peptide containing two tyrosine amino acids and all possible phosphotyrosine derivatives thereof will be synthesized. A mixture of these peptides will be INDEED. Tyrosine will be recognized by an antibody specific for PTC-tyrosine, and phosphotyrosine will be recognized by an anti-phosphotyrosine antibody (such as PY20 shown above). Recognition events will be recorded by BD-PEX, and the resulting DNA will be sequenced. It is expected that the site of phosphorylation will be encoded in the DNA sequence, and the relative abundance of phosphorylation will be quantified using read counts. Depending on the stability of the PTM during INDEED, this approach can be extended to other PTMs, such as phosphorylation, methylation, acylation, and glycosylation on serine and threonine.
[0198] Example 11: Fingerprint analysis of full-length proteins
[0199] The average protein length in eukaryotes is 400 amino acids. Fingerprinting full-length proteins by Edman degradation may not be optimal due to limitations in degradation efficiency. Therefore, to fingerprint full-length proteins, proteins can be digested with endopeptidases, such as trypsin, to produce short peptides, which can then be subjected to INDEED. This capability will be demonstrated by fingerprinting tryptic digests of full-length proteins.
[0200] Due to the wide range of cysteine-specific reactions, such as α-halocarbonyl and maleimide, peptide immobilization via cysteine can be adopted. By controlling the pH, selective modification of cysteine to other nucleophilic residues such as lysine, histidine and the N-terminus can be achieved. However, the low abundance of cysteine (2%) may lead to incomplete capture of tryptic peptides. The C-terminal carboxylic acid is a more versatile conjugation handle for peptide immobilization. It has been reported that the C-terminal carboxylic acid can be selectively labeled by carboxypeptidases, the proteolytic activity of which is suppressed at high pH, while the transpeptidase activity catalyzes the attachment of nucleophilic molecules to the C-terminal carboxylic acid. Recently, a photoredox-catalyzed decarboxylation reaction of the C-terminal carboxylic acid has been described. Although this method has only been demonstrated for short peptides less than 10 amino acids in length, it can be used as a more versatile and efficient method for C-terminal immobilization. Conjugation of click handles such as alkynes has been demonstrated, and therefore these methods can be easily implemented into the INDEED workflow.
[0201] The side chains of cysteine and lysine can be capped before degradation. It has been well-studied that cysteine can be capped by alkylation. Lysine capping can be achieved by first masking the N-terminus with a reversible modification, and then irreversibly capping the lysine with a reagent such as an NHS ester. After lysine capping, the N-terminal amino group is released by removing the reversible modification. In addition, these capping reactions can be used to introduce affinity tags that are recognized by existing affinity reagents, thereby further expanding the range of sequenceable amino acids.
[0202] Example 12—Mapping of protein variants from single cells
[0203] Mapping protein variants at the single cell level can reveal cellular heterogeneity beyond the gene or even protein level and can greatly advance our understanding of cell function, organism development, and disease mechanisms. The polypeptide sequencing methods disclosed herein can be used for single molecule analysis of protein variants, such as single amino acid substitutions and post-translational modifications. Here, a workflow for mapping these protein variants at the single cell level will be developed ( Figure 9 ). First, to isolate and enrich the protein of interest, single cells are isolated via FACS in a multiwell plate containing cleavage buffer and beads coated with antibodies against the protein of interest. Second, the protein of interest is eluted from the antibody-coated beads and digested with trypsin. Finally, the resulting peptides are conjugated to DNA UMIs and sequenced using our single-molecule peptide sequencing technology. Importantly, the UMIs used in this workflow can also include barcodes specific to each well, thereby allowing identification and quantification of protein variants in each cell.
[0204] Therefore, the above only illustrates the principle of the present disclosure. It should be understood that those skilled in the art will be able to design various arrangements, which, although not explicitly described or shown in this article, embody the principle of the present invention and are included in its spirit and scope. In addition, all examples and conditional language narrated in this article are mainly intended to help readers understand the principle of the present invention and the concept contributed by the inventor to promote this area, and should be interpreted as not being limited to the examples and conditions of such specific narration. In addition, all statements of the principles, aspects and embodiments of the present invention and its specific examples are narrated herein and are intended to cover both its structural equivalents and functional equivalents. In addition, it is intended that such equivalents include both currently known equivalents and equivalents to be developed in the future, i.e., any element of the performance of the same function of development, regardless of structure. Therefore, the scope of the present invention is not intended to be limited to the exemplary embodiments shown and described herein.
Claims
1. A method for sequencing a polypeptide, the method comprising: (a) labeling the N-terminal amino acid of a polypeptide with a nucleic acid tag comprising a unique molecular identifier (UMI); (b) degrading the N-terminal amino acid from the polypeptide; (c) annealing a primer to the nucleic acid marker of the degraded N-terminal amino acid, wherein the primer comprises a barcode corresponding to the identity of the degraded N-terminal amino acid; (d) extending the primer annealed to the nucleic acid marker to generate an extension product, wherein the extension product comprises the UMI and the barcode corresponding to the identity of the degraded N-terminal amino acid; (e) performing steps (a)-(d) in consecutive cycles to generate a plurality of extension products, each extension product in the plurality of extension products comprising the UMI, a corresponding cycle number barcode, and a barcode corresponding to the identity of a corresponding degraded N-terminal amino acid; (f) sequencing the plurality of extension products; as well as (g) determining the sequence of the polypeptide based on the sequences of the plurality of extension products.
2. The method of claim 1, comprising indexing the extension products produced in step (d) with the cycle number barcode.
3. The method of claim 1 , wherein step (a) comprises labeling the N-terminal amino acid of the polypeptide with a nucleic acid marker comprising the UMI and the cycle number barcode.
4. The method of any one of claims 1 to 3, wherein prior to step (a), the polypeptide is immobilized on the solid support via a nucleic acid attached to the C-terminus of the polypeptide and to the surface of the solid support, wherein the nucleic acid comprises the UMI and a primer binding site 3' or 5' to the UMI.
5. The method of claim 4, wherein the marking step (a) comprises: (i) conjugating a degradation moiety to the N-terminal amino acid; (ii) conjugating a primer to the degradation moiety, wherein the primer conjugated to the degradation moiety comprises a sequence complementary to the primer binding site of the nucleic acid that immobilizes the polypeptide to the solid support; (iii) annealing the primer conjugated to the degradation moiety to the primer binding site; as well as (iv) using the nucleic acid that immobilizes the polypeptide to the solid support as a template to extend the primer conjugated to the degradation moiety, thereby labeling the N-terminal amino acid with the nucleic acid label comprising the UMI.
6. The method of claim 4, wherein the marking step (a) comprises: (i) conjugating a degradation moiety conjugated to a primer to the N-terminal amino acid, the primer comprising a sequence complementary to the primer binding site of the nucleic acid that immobilizes the polypeptide to the solid support; (ii) annealing the primer conjugated to the degradation moiety to the primer binding site; as well as (iii) using the nucleic acid that immobilizes the polypeptide to the solid support as a template to extend the primer conjugated to the degradation moiety, thereby labeling the N-terminal amino acid with the nucleic acid label comprising the UMI.
7. The method of claim 5 or claim 6, wherein the degradation moiety comprises phenyl isothiocyanate (PITC) carrying a reactive group for conjugation to the nucleic acid comprising the cycle number barcode. The method of claim 7 , wherein the reactive group is a click chemistry reactive group.
9. The method according to any one of claims 1 to 8, wherein the degradation step (b) is carried out under conditions comprising a Lewis acid in an aprotic solvent.
10. The method of claim 9, wherein the Lewis acid is BF3 etherate, BCl3, BBr3, scandium(III) trifluoromethanesulfonate, or any combination thereof.
11. The method of claim 9, wherein the Lewis acid is BF3 etherate.
12. The method according to any one of claims 9 to 11, wherein the aprotic solvent is acetonitrile, N,N-dimethylformamide (DMF), dimethyl sulfoxide (DMSO), or any combination thereof.
13. The method of claim 12, wherein the degradation step (b) is performed under conditions comprising triethylamine acetate in N,N-dimethylformamide (DMF).
14. The method according to any one of claims 1 to 13, wherein the one or more nucleic acids employed in step (a) and / or step (b) comprise non-natural nucleotides that stabilize the nucleic acid during degradation step (b).
15. The method of claim 14, wherein the non-natural nucleotide comprises a 7-deaza purine nucleotide.
16. The method according to any one of claims 1 to 15, wherein in step (c), the primer comprising the barcode corresponding to the identity of the degraded N-terminal amino acid is conjugated to a binding moiety that specifically binds to the degraded N-terminal amino acid, and wherein the annealing is dependent on binding of the binding moiety to the degraded N-terminal amino acid.
17. The method of claim 16, wherein the binding moiety specifically binds to a degraded N-terminal amino acid comprising a post-translational modification, and wherein the barcode indicates the identity of the degraded N-terminal amino acid and the post-translational modification.
18. The method of claim 16 or claim 17, wherein the post-translational modification is phosphorylation, glycosylation, ubiquitination, nitrosylation, methylation, acetylation or lipidation.
19. The method of any one of claims 16 to 18, wherein the binding moiety is a polypeptide.
20. The method of claim 19, wherein the polypeptide is an antibody.
21. The method of any one of claims 16 to 18, wherein the binding moiety is a small molecule or an aptamer.
22. The method according to any one of claims 1 to 21, wherein the polypeptide to be sequenced is present in a protein sample isolated from a single cell.
23. The method according to claim 22, wherein the method is a single-cell protein sequencing method performed on a plurality of polypeptides present in the protein sample.
24. The method according to any one of claims 1 to 23, wherein the polypeptide to be sequenced is present in a protein sample isolated from a tissue sample.
25. The method of claim 24, wherein the tissue sample is a biopsy sample.
26. The method of claim 25, wherein the biopsy sample is a tumor biopsy sample.
27. The method of any one of claims 1 to 23, wherein the polypeptide to be sequenced is present in a protein sample isolated from a biological fluid.
28. A composition comprising one or any combination of the following: (a) UMI-functionalized solid support, (b) a degradation moiety carrying a reactive group for conjugation to a nucleic acid, (c) primers comprising a cycle number barcode, (d) Edman degradation reagent, (e) a conjugate comprising a primer conjugated to a binding moiety that specifically binds to an amino acid, (f) a conjugate comprising a primer conjugated to a binding moiety that specifically binds to an amino acid bearing a post-translational modification, and (g) Nucleic acid sequencing aptamers 29. A kit comprising: (a) One or any combination of the following: (i) UMI-functionalized solid support, (ii) a degradation moiety carrying a reactive group for conjugation to a nucleic acid, (iii) primers containing cycle number barcodes, (iv) Edman degradation reagent, (v) a conjugate comprising a primer conjugated to a binding moiety that specifically binds to an amino acid, (vi) a conjugate comprising a primer conjugated to a binding moiety that specifically binds to an amino acid bearing a post-translational modification, and (vii) nucleic acid sequencing aptamers; as well as (b) Instructions for using one or any combination of the reagents to perform a method according to any one of claims 1 to 27.
30. The kit of claim 29, wherein the degradation moiety comprises PITC carrying a reactive group for conjugation to the primer comprising the cycle number barcode.
31. The kit of claim 30, wherein the reactive group is a click chemistry reactive group.
32. The kit of any one of claims 29 to 31 , wherein the Edman degradation reagent comprises a Lewis acid in an aprotic solvent.
33. The kit of claim 32, wherein the Lewis acid is BF3 etherate, BCl3, BBr3, scandium(III) trifluoromethanesulfonate, or any combination thereof.
34. The kit of claim 32, wherein the Lewis acid is BF3 etherate.
35. The kit according to any one of claims 32 to 34, wherein the aprotic solvent is acetonitrile, N,N-dimethylformamide (DMF), dimethyl sulfoxide (DMSO), or any combination thereof.
36. The kit of any one of claims 29 to 35, wherein one or more of the nucleic acid-based reagents comprises non-natural nucleotides that stabilize the nucleic acid under Edman degradation conditions.
37. The kit of claim 36, wherein the non-natural nucleotide comprises a 7-deaza purine nucleotide.
38. A kit according to any one of claims 29 to 37, wherein the binding moiety is a polypeptide.
39. The kit of claim 38, wherein the polypeptide is an antibody.
40. The kit of any one of claims 29 to 37, wherein the binding moiety is a small molecule or an aptamer.
Citation Information
Patent Citations
Method for sequencing nucleic acid molecules
US20030044781A1
Recombinant immunoglobin preparations
US4816567A
Zero-mode clad waveguides for performing spectroscopy with confined effective observation volumes
US6917726B2
Method for sequencing nucleic acid molecules
US7033764B2
Method for sequencing nucleic acid molecules
US7052847B2