Macromolecular analysis using nucleic acid encoding
The method of conjugating macromolecules to solid supports with recording tags and using coding agents for nucleic acid sequencing addresses proteomics challenges, achieving high-throughput and accurate analysis of proteins and peptides.
Patent Information
- Application Number
- JP2024006005
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2016-08-18
- Filing Date
- 2024-01-18
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2037-05-02
AI Technical Summary
Current proteomics technologies face challenges in achieving high-throughput, multiplexed, and accurate analysis of proteins and peptides due to limitations in multiplexing, cross-reactivity, and throughput, especially in characterizing post-translational modifications and dynamic range of protein concentrations.
A method involving macromolecules conjugated to solid supports with recording tags, using sequential or simultaneous binding with agents having coding tags to generate extended recording tags, followed by nucleic acid sequencing for high-throughput analysis.
Enables highly parallel, sensitive, and accurate analysis of proteins and peptides, overcoming limitations of existing methods by enhancing multiplexing and throughput, and providing detailed information on post-translational modifications.
Smart Images

Figure 0007796432000012 
Figure 0007796432000013 
Figure 0007796432000014
Abstract
Description
[Technical Field]
[0001] Sequence Listing Description The sequence listing accompanying this application is provided in text format in lieu of a paper copy and is hereby incorporated by reference. The name of the text file containing the sequence listing is 760229_401WO_SEQUENCE_LISTING.txt. The text file is 38.7 KB, was created on May 2, 2017, and is submitted electronically via EFS-Web.
[0002] The present disclosure relates generally to the analysis of macromolecules, including peptides, polypeptides, and proteins, using barcoding and nucleic acid encoding of molecular recognition events. [Background technology]
[0003] Proteins perform and facilitate many different biological functions and play essential roles in cell biology and physiology. The repertoire of different protein molecules is extensive and much more complex than the transcriptome due to the additional diversity introduced by post-translational modifications (PTMs). Furthermore, proteins within cells dynamically change (expression levels and modification states) in response to the environment, physiological conditions, and pathological conditions. Therefore, proteins contain a vast amount of largely uncovered relevant information, especially compared to genomic information. In general, proteomics analysis has lagged behind genomics analysis. While next-generation sequencing (NGS) has transformed the field of genomics by enabling the analysis of billions of DNA sequences in a single instrument run, throughput in protein and peptide analysis remains limited.
[0004] Nevertheless, this protein information is directly needed to better understand proteome dynamics in health and disease, and to help enable precision medicine. As such, there is great interest in developing "next-generation" tools to miniaturize and highly parallelize the collection of this proteomic information.
[0005] Highly parallel macromolecular characterization and recognition of proteins is challenging for several reasons. The use of affinity-based assays is often hampered by several key challenges. One challenge is multiplexing the readout of a set of affinity agents to a set of related macromolecules; another challenge is minimizing cross-reactivity between affinity agents and off-target macromolecules; and a third challenge is developing an efficient high-throughput readout platform. An example of this problem arises in proteomics, where one goal is to identify and quantify the majority or all of the proteins in a sample. Furthermore, it is desirable to characterize various post-translational modifications (PTMs) of proteins at the single-molecule level. Currently, this is a formidable challenge to achieve in a high-throughput manner.
[0006] Molecular recognition and characterization of protein or peptide macromolecules are commonly performed using immunoassays. Many different immunoassay formats exist, including ELISA, multiplex ELISA (e.g., spotted antibody array, liquid particle ELISA array), digital ELISA (e.g., Quanterix, Singulex), reverse phase protein array (RPPA), and many others. All of these different immunoassay platforms face similar challenges, including the development of high-affinity and highly specific (or selective) antibodies (binders), limited multiplexing capabilities at both the sample and analyte levels, limited sensitivity and dynamic range, and cross-reactivity and background signals. Binder-agnostic approaches, such as direct protein characterization by peptide sequencing (Edman degradation or mass spectrometry), offer useful alternatives. However, none of these approaches are highly parallel or high-throughput.
[0007] Peptide sequencing based on Edman degradation, first proposed by Pehr Edman in 1950, involves the stepwise degradation of the N-terminal amino acid of a peptide through a series of chemical modifications and downstream HPLC analysis (later replaced by mass spectrometry). In the first step, the N-terminal amino acid is modified with phenylisothiocyanate (PITC) under mildly basic conditions (NMP / methanol / HO) to form the phenylthiocarbamoyl (PTC) derivative. In the second step, the PTC-modified amino group is treated with acid (anhydrous TFA) to create a cleaved cyclic ATZ (2-anilino-5(4)-thiazolinone)-modified amino acid, leaving a new N-terminus on the peptide. The cleaved cyclic ATZ amino acid is converted to a PTH amino acid derivative and analyzed by reverse-phase HPLC. This process is continued iteratively until all or a partial number of the amino acids comprising the peptide sequence have been removed from the N-terminus and identified. Generally, Edman degradation peptide sequencing is time-consuming and has limited throughput, with only a few peptides per day.
[0008] Over the past 10–15 years, peptide analysis using MALDI, electrospray mass spectrometry (MS), and LC-MS / MS has largely replaced Edman degradation. Despite recent advances in MS instrumentation (Riley et al., 2016, Cell Syst, 2:142–143), MS still suffers from several drawbacks, including high instrumentation costs, sophisticated user requirements, inadequate quantification capabilities, and a limited ability to perform measurements across the dynamic range of the proteome. For example, proteins ionize with different levels of efficiency, making absolute and even relative quantification between samples difficult. The implementation of mass tags has helped improve relative quantification, but requires labeling of the proteome. An additional complication is the dynamic range, where the concentration of proteins in a sample can vary over a very large range (over 10 orders of magnitude for plasma). Generally, only more abundant species are analyzed with MS, making characterization of less abundant proteins difficult. Finally, sample throughput is generally limited to a few thousand peptides per run, and for data-independent analysis (DIA), this throughput is insufficient for true bottom-up, high-throughput proteomic analysis. Furthermore, there are significant computational requirements for deconvoluting the thousands of complex MS spectra recorded for each sample. [Prior art documents] [Non-patent literature]
[0009] [Non-Patent Document 1] Riley et al., 2016, Cell Syst, 2:142-143 Summary of the Invention [Means for solving the problem]
[0010] Thus, there remains a need in the art for improved techniques for macromolecular sequencing and / or analysis applied to protein sequencing and / or analysis, and for products, methods, and kits for achieving the same. There is a need for highly parallel, accurate, sensitive, and high-throughput proteomics technologies. The present disclosure meets these and other needs.
[0011] These and other aspects of the present invention will become apparent upon reference to the following detailed description, and to this end, various references are set forth herein in which certain background information, procedures, compounds and / or compositions are described in more detail, each of which is hereby incorporated by reference in its entirety.
[0012] Embodiments of the present disclosure generally relate to methods for highly parallel, high-throughput digital macromolecular analysis, and peptide analysis in particular.
[0013] A first embodiment is a method for analyzing a macromolecule, comprising: (a) providing a macromolecule conjugated to a solid support and an associated recording tag; (b) contacting the macromolecule with a first binding agent capable of binding to the macromolecule, the first binding agent including a first coding tag having identifying information related to the first binding agent; (c) transferring information of the first coding tag to the recording tag to generate a primary decompressed recording tag; (d) contacting the macromolecule with a second binding agent capable of binding to the macromolecule, the second binding agent comprising a second coding tag having identifying information for the second binding agent; (e) transferring information of the second coding tag to the primary decompression recording tag to generate a secondary decompression recording tag; (f) analyzing the secondary decompression record tag; The method includes:
[0014] A second embodiment is the method of embodiment 1, wherein the contacting steps (b) and (d) are performed sequentially.
[0015] A third embodiment is the method of embodiment 1, wherein the contacting steps (b) and (d) are performed simultaneously.
[0016] A fourth embodiment further comprises, between steps (e) and (f): (x) repeating steps (d) and (e) one or more times by replacing the second binding agent with a third (or higher order) binding agent capable of binding to the macromolecule, the third (or higher order) binding agent comprising a third (or higher order) coding tag having identifying information for the third (or higher order) binding agent; (y) transferring information of the third (or higher) coding tag to the second (or higher) decompressed recording tag to generate a third (or higher) decompressed recording tag; further comprising 2. The method of embodiment 1, wherein in step (f) the third (or higher) extended recording tag is analyzed.
[0017] A fifth embodiment is a method for analyzing a macromolecule, comprising: (a) providing a macromolecule conjugated to a solid support, an associated first recording tag, and an associated second recording tag; (b) contacting the macromolecule with a first binding agent capable of binding to the macromolecule, the first binding agent including a first coding tag having identifying information related to the first binding agent; (c) transferring information of the first coding tag to the first recording tag to generate a first decompressed recording tag; (d) contacting the macromolecule with a second binding agent capable of binding to the macromolecule, the second binding agent comprising a second coding tag having identifying information for the second binding agent; (e) transferring information of the second coding tag to the second recording tag to generate a second decompressed recording tag; (f) analyzing the first and second decompressed recording tags; The method includes:
[0018] A sixth embodiment is the method of embodiment 5, wherein the contacting steps (b) and (d) are performed sequentially.
[0019] A seventh embodiment is the method of embodiment 5, wherein the contacting steps (b) and (d) are performed simultaneously.
[0020] An eighth embodiment is the method of embodiment 5, wherein step (a) further comprises providing an associated third (or higher order) recording tag attached to said solid support.
[0021] A ninth embodiment further comprises, between steps (e) and (f): (x) repeating steps (d) and (e) one or more times by replacing the second binding agent with a third (or higher order) binding agent capable of binding to the macromolecule, the third (or higher order) binding agent comprising a third (or higher order) coding tag having identifying information for the third (or higher order) binding agent; (y) transferring information of the third (or higher) coding tag to the third (or higher) recording tag to generate a third (or higher) decompressed recording tag; further comprising 9. The method of embodiment 8, wherein in step (f) the first extended recording tag, the second extended recording tag and the third (or higher order) extended recording tag are analyzed.
[0022] A tenth embodiment is the method of any one of embodiments 5 to 9, wherein the first coding tag, the second coding tag, and any higher order coding tag comprise a binding cycle-specific spacer sequence.
[0023] An eleventh embodiment is a method for analyzing a peptide, comprising: (a) providing a peptide conjugated to a solid support and an associated recording tag; (b) modifying the N-terminal amino acid (NTAA) of the peptide with a chemical agent; (c) contacting the peptide with a first binding agent capable of binding to the modified NTAA, the first binding agent comprising a first coding tag having identifying information for the first binding agent; (d) transferring information of the first coding tag to the recording tag to generate an expanded recording tag; (e) analyzing the decompressed record tag; The method includes:
[0024] A twelfth embodiment is the method described in embodiment 11, wherein step (c) further comprises contacting the peptide with a second (or higher order) binding substance comprising a second (or higher order) coding tag having identifying information for the second (or higher order) binding substance, the second (or higher order) binding substance being capable of binding to a modified NTAA other than the modified NTAA of step (b).
[0025] A thirteenth embodiment is the method of the twelfth embodiment, wherein the contacting of the peptide with the second (or higher order) binding substance is performed sequentially after the contacting of the peptide with the first binding substance.
[0026] A 14th embodiment is the method of the 12th embodiment, wherein the contacting of the peptide with the second (or higher order) binding substance is performed simultaneously with the contacting of the peptide with the first binding substance.
[0027] A 15th embodiment is the method of any one of embodiments 11 to 14, wherein the chemical agent is an isothiocyanate derivative, 2,4-dinitrobenzenesulfonic acid (DNBS), 4-sulfonyl-2-nitrofluorobenzene (SNFB) 1-fluoro-2,4-dinitrobenzene, dansyl chloride, 7-methoxycoumarin acetic acid, a thioacylating reagent, a thioacetylating reagent, or a thiobenzylating reagent.
[0028] A sixteenth embodiment is a method for analyzing a peptide, comprising: (a) providing a peptide conjugated to a solid support and an associated recording tag; (b) modifying the N-terminal amino acid (NTAA) of the peptide with a chemical agent to obtain a modified NTAA; (c) contacting the peptide with a first binding agent capable of binding to the modified NTAA, the first binding agent comprising a first coding tag having identifying information for the first binding agent; (d) transferring information of the first coding tag to the recording tag to generate a first decompressed recording tag; (e) removing the modified NTAA to expose a new NTAA; (f) modifying the new NTAA of the peptide with a chemical agent to obtain a new modified NTAA; (g) contacting the peptide with a second binding agent capable of binding to the newly modified NTAA, the second binding agent comprising a second coding tag having identifying information for the second binding agent; (h) transferring information of the second coding tag to the first decompressed recording tag to generate a second decompressed recording tag; (i) analyzing the second decompressed record tag; The method includes:
[0029] A seventeenth embodiment is a method for analyzing a peptide, comprising: (a) providing a peptide conjugated to a solid support and an associated recording tag; (b) contacting the peptide with a first binding agent capable of binding to an N-terminal amino acid (NTAA) of the peptide, the first binding agent comprising a first coding tag having identifying information for the first binding agent; (c) transferring information of the first coding tag to the recording tag to generate an expanded recording tag; (d) analyzing the decompressed record tag; The method includes:
[0030] In an 18th embodiment, the method of the 17th embodiment is characterized in that step (b) further comprises contacting the peptide with a second (or higher order) binding substance comprising a second (or higher order) coding tag having identifying information for the second (or higher order) binding substance, the second (or higher order) binding substance being capable of binding to an NTAA other than the NTAA of the peptide.
[0031] A 19th embodiment is the method of the 18th embodiment, wherein the contacting of the peptide with the second (or higher order) binding substance is performed sequentially after the contacting of the peptide with the first binding substance.
[0032] A 20th embodiment is the method of the 18th embodiment, wherein the contacting of the peptide with the second (or higher order) binding substance is carried out simultaneously with the contacting of the peptide with the first binding substance.
[0033] A twenty-first embodiment is a method for analyzing a peptide, comprising: (a) providing a peptide conjugated to a solid support and an associated recording tag; (b) contacting the peptide with a first binding agent capable of binding to an N-terminal amino acid (NTAA) of the peptide, the first binding agent comprising a first coding tag having identifying information for the first binding agent; (c) transferring information of the first coding tag to the recording tag to generate a first decompressed recording tag; (d) removing the NTAA to expose a new NTAA of the peptide; (e) contacting the peptide with a second binding agent capable of binding to the new NTAA, the second binding agent comprising a second coding tag having identifying information for the second binding agent; (h) transferring information of the second coding tag to the first decompressed recording tag to generate a second decompressed recording tag; (i) analyzing the second decompressed record tag; The method includes:
[0034] A 22nd embodiment is the method of any one of embodiments 1 to 10, wherein the macromolecule is a protein, polypeptide, or peptide.
[0035] A 23rd embodiment is the method of any one of embodiments 1 to 10, wherein the macromolecule is a peptide.
[0036] A 24th embodiment is the method of any one of embodiments 11 to 23, wherein the peptide is obtained by fragmenting a protein from a biological sample.
[0037] A 25th embodiment is the method of any one of embodiments 1 to 10, wherein the macromolecule is a lipid, carbohydrate, or macrocycle.
[0038] A 26th embodiment is a method described in any one of embodiments 1 to 25, wherein the recording tag is a DNA molecule, a DNA with pseudo-complementary bases, an RNA molecule, a BNA molecule, an XNA molecule, an LNA molecule, a PNA molecule, a γPNA molecule, or a combination thereof.
[0039] A 27th embodiment is the method of any one of embodiments 1 to 26, wherein the recording tag comprises a universal priming site.
[0040] A 28th embodiment is the method of embodiment 27, wherein the universal priming sites comprise priming sites for amplification, sequencing, or both.
[0041] A 29th embodiment is the method of any one of embodiments 1 to 28, wherein the recording tag comprises a unique molecular identifier (UMI).
[0042] A thirtieth embodiment is the method according to any one of the first to twenty-ninth embodiments, wherein the recording tag includes a barcode.
[0043] A thirty-first embodiment is the method of any one of embodiments 1 to 30, wherein the recording tag comprises a spacer at its 3' end.
[0044] A 32nd embodiment is the method of any one of embodiments 1 to 31, wherein the macromolecule and the associated recording tag are covalently attached to the solid support.
[0045] A 33rd embodiment is the method of any one of embodiments 1 to 32, wherein the solid support is a bead, a porous bead, a porous matrix, an array, a glass surface, a silicon surface, a plastic surface, a filter, a membrane, nylon, a silicon wafer chip, a flow-through chip, a biochip including signal transduction electronics, a microtiter well, an ELISA plate, a spin interference disk, a nitrocellulose membrane, a nitrocellulose-based polymer surface, a nanoparticle, or a microsphere.
[0046] A 34th embodiment is the method of embodiment 33, wherein the solid support is a polystyrene bead, a polymer bead, an agarose bead, an acrylamide bead, a solid core bead, a porous bead, a paramagnetic bead, a glass bead, or a controlled pore bead.
[0047] A thirty-fifth embodiment is the method of any one of embodiments 1 to 34, wherein a plurality of macromolecules and associated recording tags are attached to a solid support.
[0048] A 36th embodiment is the method of embodiment 35, wherein the plurality of macromolecules are spaced on the solid support by an average distance of >50 nm.
[0049] A 37th embodiment is the method of any one of embodiments 1 to 36, wherein the binding agent is a polypeptide or protein.
[0050] A 38th embodiment is the method of embodiment 37, wherein the binding agent is a modified aminopeptidase, a modified aminoacyl-tRNA synthetase, a modified anticalin, or a modified ClpS.
[0051] A 39th embodiment is the method of any one of embodiments 1 to 38, wherein the binding agent is capable of selectively binding to a macromolecule.
[0052] A 40th embodiment is a method according to any one of embodiments 1 to 39, wherein the coding tag is a DNA molecule, an RNA molecule, a BNA molecule, an XNA molecule, an LNA molecule, a PNA molecule, a γPNA molecule, or a combination thereof.
[0053] A 41st embodiment is the method of any one of embodiments 1 to 40, wherein the coding tag comprises an encoder sequence.
[0054] A 42nd embodiment is the method of any one of embodiments 1 to 41, wherein the coding tag further comprises a spacer, a binding cycle-specific sequence, a unique molecular identifier, a universal priming site, or any combination thereof.
[0055] A 43rd embodiment is the method of any one of embodiments 1 to 42, wherein the binding agent and the coding tag are joined by a linker.
[0056] A 44th embodiment is the method according to embodiments 1 to 42, wherein the binding agent and the coding tag are joined by a SpyTag / SpyCatcher or SnoopTag / SnoopCatcher peptide-protein pair.
[0057] A 45th embodiment is the method of any one of embodiments 1 to 44, wherein the transfer of information from the coding tag to the recording tag is mediated by DNA ligase.
[0058] A 46th embodiment is the method according to any one of embodiments 1 to 44, wherein the transfer of information from the coding tag to the recording tag is mediated by a DNA polymerase.
[0059] A 47th embodiment is the method of any one of embodiments 1 to 44, wherein the transfer of information from the coding tag to the recording tag is mediated by chemical ligation.
[0060] A 48th embodiment is the method of any one of embodiments 1 to 47, wherein analysis of the extended recording tag comprises nucleic acid sequencing.
[0061] A 49th embodiment is the method of embodiment 48, wherein the nucleic acid sequencing method is sequencing by synthesis, sequencing by ligation, sequencing by hybridization, polony sequencing, ion semiconductor sequencing, or pyrosequencing.
[0062] A 50th embodiment is the method of embodiment 48, wherein the nucleic acid sequencing method is single molecule real-time sequencing, nanopore-based sequencing, or direct imaging of DNA using advanced microscopy.
[0063] A 51st embodiment is the method of any one of embodiments 1 to 50, wherein the extended recording tag is amplified before analysis.
[0064] A 52nd embodiment is the method of embodiments 1 to 51, wherein the order of coding tag information contained in the extended recording tags provides information about the order of binding of the binding agents to the macromolecule.
[0065] A 53rd embodiment is the method of embodiments 1 to 52, wherein the frequency of the coding tag information contained in the extended recording tag provides information about the frequency of binding of the binding substance to the macromolecule.
[0066] A 54th embodiment is the method of any one of embodiments 1 to 53, wherein multiple extended recording tags representing multiple macromolecules are analyzed in parallel.
[0067] A 55th embodiment is the method of embodiment 54, wherein a plurality of extended recording tags representing said plurality of macromolecules are analyzed in a multiplexed assay.
[0068] A 56th embodiment is the method of any one of embodiments 1 to 55, wherein the plurality of extended recording tags are subjected to a target enrichment assay prior to analysis.
[0069] A 57th embodiment is the method of any one of embodiments 1 to 56, wherein the plurality of extended recording tags is subjected to a subtraction assay prior to analysis.
[0070] A 58th embodiment is the method of any one of embodiments 1 to 57, wherein the plurality of extended recording tags undergoes a normalization assay prior to analysis to reduce highly abundant species.
[0071] A 59th embodiment is the method of any one of embodiments 1 to 58, wherein the NTAA is removed by modified aminopeptidase, modified amino acid tRNA synthetase, mild Edman degradation, Edmanase enzyme, or anhydrous TFA.
[0072] A 60th embodiment is the method of any one of embodiments 1 to 59, wherein at least one binding agent is attached to a terminal amino acid residue.
[0073] A 61st embodiment is the method of any one of embodiments 1 to 60, wherein at least one binding agent binds to a post-translationally modified amino acid.
[0074] A 62nd embodiment is a method for analyzing one or more peptides from a sample containing multiple protein complexes, proteins, or polypeptides, comprising:
[0075] (a) distributing the plurality of protein complexes, proteins, or polypeptides in the sample into a plurality of compartments, each compartment comprising a plurality of compartment tags optionally attached to a solid support, the compartment tags being the same within an individual compartment and different from compartment tags in other compartments; (b) fragmenting said plurality of protein complexes, proteins, and / or polypeptides into a plurality of peptides; (c) contacting the plurality of peptides with the plurality of compartment tags under conditions sufficient to allow annealing or joining of the plurality of peptides with the plurality of compartment tags within the plurality of compartments, thereby producing a plurality of compartment-tagged peptides; (d) collecting said compartment-tagged peptides from said plurality of compartments; (e) analyzing one or more compartment-tagged peptides according to a method according to any one of embodiments 1 to 21 and embodiments 26 to 61; The method includes:
[0076] A 63rd embodiment is the method of embodiment 62, wherein the compartments are microfluidic droplets.
[0077] A 64th embodiment is the method of embodiment 62, wherein the compartment is a microwell.
[0078] A 65th embodiment is the method of embodiment 62, wherein the compartments are separate regions on a surface.
[0079] A 66th embodiment is the method of any one of embodiments 62 to 65, wherein each compartment contains, on average, a single cell.
[0080] A 67th embodiment is a method for analyzing one or more peptides from a sample containing multiple protein complexes, proteins, or polypeptides, comprising: (a) labeling said plurality of protein complexes, proteins, or polypeptides with a plurality of universal DNA tags; (b) distributing the plurality of labeled protein complexes, proteins, or polypeptides in the sample into a plurality of compartments, each compartment comprising a plurality of compartment tags, the compartment tags being the same within an individual compartment and different from compartment tags in other compartments; (c) contacting said plurality of protein complexes, proteins, or polypeptides with said plurality of compartment tags under conditions sufficient to allow annealing or joining of said plurality of protein complexes, proteins, or polypeptides with said plurality of compartment tags within said plurality of compartments, thereby producing a plurality of compartment-tagged protein complexes, proteins, or polypeptides; (d) collecting said compartment-tagged protein complexes, proteins, or polypeptides from said plurality of compartments; (e) optionally fragmenting said compartment-tagged protein complex, protein, or polypeptide into compartment-tagged peptides; (f) analyzing one or more compartment-tagged peptides according to a method according to any one of embodiments 1 to 21 and embodiments 26 to 61; The method includes:
[0081] A 68th embodiment is the method of any one of embodiments 62 to 67, wherein the compartment tag information is transferred to the recording tag associated with the peptide by primer extension or ligation.
[0082] A 69th embodiment is the method of any one of embodiments 62 to 68, wherein the solid support comprises beads.
[0083] A 70th embodiment is the method of embodiment 69, wherein the beads are polystyrene beads, polymer beads, agarose beads, acrylamide beads, solid core beads, porous beads, paramagnetic beads, glass beads, or controlled pore beads.
[0084] A 71st embodiment is the method of any one of embodiments 62 to 70, wherein the compartment tag comprises a single-stranded or double-stranded nucleic acid molecule.
[0085] A 72nd embodiment is the method of any one of embodiments 62 to 71, wherein the compartment tag comprises a barcode and optionally a UMI.
[0086] A 73rd embodiment is the method of embodiment 72, wherein the solid support is a bead, the compartment tag comprises a barcode, and further wherein the beads having the multiple compartment tags attached thereto are formed by split-and-pool synthesis.
[0087] A 74th embodiment is the method of embodiment 72, wherein the solid support is a bead, the compartment tag comprises a barcode, and further wherein beads having multiple compartment tags attached thereto are formed by individual synthesis or immobilization.
[0088] A 75th embodiment is the method of any one of embodiments 62 to 74, wherein the compartment tag is a component within a recording tag, and the recording tag optionally further comprises a spacer, a unique molecular identifier, a universal priming site, or any combination thereof.
[0089] A 76th embodiment is the method of any one of embodiments 62 to 75, wherein the compartment tag further comprises a functional moiety capable of reacting with an internal or N-terminal amino acid of the plurality of protein complexes, proteins, or polypeptides.
[0090] A 77th embodiment is the method of embodiment 76, wherein the functional moiety is an NHS group.
[0091] A 78th embodiment is the method of embodiment 76, wherein the functional moiety is an aldehyde group.
[0092] A 79th embodiment is a method described in any one of embodiments 62 to 78, wherein the multiple compartment tags are formed by printing, spotting, ink-jetting the compartment tags into the compartments, or a combination thereof.
[0093] An 80th embodiment is the method of any one of embodiments 62 to 79, wherein the compartment tag further comprises a peptide.
[0094] An 81st embodiment is the method of embodiment 80, wherein the compartment tag peptide comprises a protein ligase recognition sequence.
[0095] An 82nd embodiment is the method described in embodiment 81, wherein the protein ligase is butelase I or a homolog thereof.
[0096] An 83rd embodiment is the method of any one of embodiments 62 to 82, wherein the plurality of polypeptides is fragmented with a protease.
[0097] An 84th embodiment is the method of embodiment 83, wherein the protease is a metalloprotease.
[0098] An 85th embodiment is the method of embodiment 84, wherein the activity of the metalloprotease is modulated by light-activated release of a metal cation.
[0099] An 86th embodiment is the method of any one of embodiments 62 to 85, further comprising subtracting one or more abundant proteins from the sample before distributing the plurality of polypeptides into the plurality of compartments.
[0100] An 87th embodiment is the method of any one of embodiments 62 to 86, further comprising releasing the compartment tag from the solid support prior to conjugating the plurality of peptides and the compartment tag.
[0101] An 88th embodiment is the method of embodiment 62, further comprising, after step (d), attaching said compartment-tagged peptide to a solid support with a recording tag.
[0102] An 89th embodiment is the method of embodiment 88, further comprising transferring information of the compartment tag on the compartment-tagged peptide to the associated recording tag.
[0103] A 90th embodiment is the method of embodiment 89, further comprising removing the compartment tag from the compartment-tagged peptide prior to step (e).
[0104] A 91st embodiment is a method according to any one of embodiments 62 to 90, further comprising determining the identity of the single cell from which the analyzed peptide originates based on the compartment tag sequence of the analyzed peptide.
[0105] A 92nd embodiment is the method of any one of embodiments 62 to 90, further comprising determining the identity of the protein or protein complex from which the analyzed peptide is derived based on the compartment tag sequence of the analyzed peptide.
[0106] A 93rd embodiment is a method for analyzing a plurality of macromolecules, comprising: (a) providing a plurality of macromolecules conjugated to a solid support and associated recording tags; (b) contacting the plurality of macromolecules with a plurality of binding agents capable of binding to the plurality of macromolecules, each binding agent comprising a coding tag having identifying information related to the binding agent; (c) (i) transferring information from a recording tag associated with the macromolecule to the coding tag of the binding substance bound to the macromolecule to generate an extended coding tag; or (ii) transferring information from a recording tag associated with the macromolecule and the coding tag of the binding substance bound to the macromolecule to a ditag construct; (d) collecting the extended coding tag or ditag construct; (e) optionally repeating steps (b) through (d) for one or more binding cycles; (f) analyzing the collection of extended coding tag or ditag constructs; The method includes:
[0107] A 94th embodiment is the method of embodiment 93, wherein the macromolecule is a protein.
[0108] A 95th embodiment is the method of embodiment 93, wherein the macromolecule is a peptide.
[0109] A 96th embodiment is the method of embodiment 95, wherein the peptide is obtained by fragmenting a protein from a biological sample.
[0110] A 97th embodiment is a method described in any one of embodiments 93 to 96, wherein the recording tag is a DNA molecule, an RNA molecule, a PNA molecule, a BNA molecule, an XNA molecule, an LNA molecule, a γPNA molecule, or a combination thereof.
[0111] A 98th embodiment is a method according to any one of embodiments 93 to 97, wherein the recording tag comprises a unique molecular identifier (UMI).
[0112] A 99th embodiment is a method according to any of embodiments 93 to 98, wherein the recording tag comprises a compartment tag.
[0113] A 100th embodiment is the method of any one of embodiments 93 to 99, wherein the recording tag comprises a universal priming site.
[0114] A 101st embodiment is the method of any one of embodiments 93 to 100, wherein the recording tag comprises a spacer at its 3' end.
[0115] A 102nd embodiment is a method according to any one of embodiments 93 to 101, wherein the 3' end of the recording tag is blocked to prevent extension of the recording tag by a polymerase, and information from the recording tag associated with the macromolecule and the coding tag of the binding substance bound to the macromolecule is transferred to a ditag construct.
[0116] A 103rd embodiment is a method according to any one of embodiments 93 to 102, wherein the coding tag comprises an encoder sequence.
[0117] A 104th embodiment is a method described in any one of embodiments 93 to 103, wherein the coding tag includes a UMI.
[0118] A 105th embodiment is the method of any one of embodiments 93 to 104, wherein the coding tag comprises a universal priming site.
[0119] A 106th embodiment is the method of any one of embodiments 93 to 105, wherein the coding tag comprises a spacer at its 3' end.
[0120] A 107th embodiment is the method of any one of embodiments 93 to 106, wherein the coding tag comprises a binding cycle-specific sequence.
[0121] A 108th embodiment is the method of any one of embodiments 93 to 107, wherein the binding agent and the coding tag are joined by a linker.
[0122] A 109th embodiment is the method according to any one of embodiments 93 to 108, wherein the transfer of information from the recording tag to the coding tag is effected by primer extension.
[0123] A 110th embodiment is the method according to any one of embodiments 93 to 108, wherein the transfer of information from the recording tag to the coding tag is effected by ligation.
[0124] A 111th embodiment is the method of any one of embodiments 93 to 108, wherein the ditag construct is generated by gap filling, primer extension, or both.
[0125] A 112th embodiment is the method of any one of embodiments 93 to 97, 107, 108, and 111, wherein the ditag molecule comprises a universal priming site derived from the recording tag, a compartment tag derived from the recording tag, a unique molecular identifier derived from the recording tag, an optional spacer derived from the recording tag, an encoder sequence derived from the coding tag, a unique molecular identifier derived from the coding tag, an optional spacer derived from the coding tag, and a universal priming site derived from the coding tag.
[0126] A 113th embodiment is the method of any one of embodiments 93 to 112, wherein the macromolecule and the associated recording tag are covalently attached to the solid support.
[0127] A 114th embodiment is the method of embodiment 113, wherein the solid support is a bead, a porous bead, a porous matrix, an array, a glass surface, a silicon surface, a plastic surface, a filter, a membrane, nylon, a silicon wafer chip, a flow-through chip, a biochip containing signal transduction electronics, a microtiter well, an ELISA plate, a spin interference disk, a nitrocellulose membrane, a nitrocellulose-based polymer surface, a nanoparticle, or a microsphere.
[0128] A 115th embodiment is the method of embodiment 114, wherein the solid support is a polystyrene bead, a polymer bead, an agarose bead, an acrylamide bead, a solid core bead, a porous bead, a paramagnetic bead, a glass bead, or a controlled pore bead.
[0129] A 116th embodiment is the method of any one of embodiments 93 to 115, wherein the binding agent is a polypeptide or protein.
[0130] A 117th embodiment is the method of embodiment 116, wherein the binding agent is a modified aminopeptidase, a modified aminoacyl-tRNA synthetase, a modified anticalin, or an antibody or binding fragment thereof.
[0131] A 118th embodiment is the method of any one of embodiments 95 to 117, wherein the binding agent binds to a single amino acid residue, a dipeptide, a tripeptide, or a post-translational modification of the peptide.
[0132] A 119th embodiment is the method of embodiment 118, wherein the binding agent binds to an N-terminal amino acid residue, a C-terminal amino acid residue, or an internal amino acid residue.
[0133] A 120th embodiment is the method of embodiment 118, wherein the binding agent binds to an N-terminal peptide, a C-terminal peptide, or an internal peptide.
[0134] A 121st embodiment is the method of embodiment 119, wherein the binding agent is bound to an N-terminal amino acid residue, and the N-terminal amino acid residue is cleaved after each binding cycle.
[0135] A 122nd embodiment is the method of embodiment 119, wherein the binding agent is bound to a C-terminal amino acid residue, and the C-terminal amino acid residue is cleaved after each binding cycle.
[0136] Embodiment 123. The method of embodiment 121, wherein the N-terminal amino acid residue is cleaved by Edman degradation.
[0137] Embodiment 124 The method of embodiment 93, wherein the binding agent is a site-specific covalent label of an amino acid or post-translational modification.
[0138] Embodiment 125. The method of any one of embodiments 93 to 124, wherein after step (b), the complex comprising the macromolecule and the associated binding agent is dissociated from the solid support and distributed into a droplet or emulsion of microfluidic droplets.
[0139] Embodiment 126. The method of embodiment 125, wherein each microfluidic droplet contains, on average, one complex comprising the macromolecule and the binding substance.
[0140] Embodiment 127. The method of embodiment 125 or 126, wherein the recording tag is amplified before generating the extended coding tag or ditag construct.
[0141] Embodiment 128. The method of any one of embodiments 125 to 127, wherein emulsion fusion PCR is used to transfer the recording tag information to the coding tag or to create a population of ditag constructs.
[0142] Embodiment 129. The method of any one of embodiments 93 to 128, wherein the collection of extended coding tags or ditag constructs is amplified prior to analysis.
[0143] Embodiment 130. The method of any one of embodiments 93 to 129, wherein the analysis of the collection of extended coding tags or ditag constructs comprises nucleic acid sequencing.
[0144] Embodiment 131. The method of embodiment 130, wherein the nucleic acid sequencing method is sequencing by synthesis, sequencing by ligation, sequencing by hybridization, polony sequencing, ion semiconductor sequencing, or pyrosequencing.
[0145] Embodiment 132. The method of embodiment 130, wherein the nucleic acid sequencing method is single-molecule real-time sequencing, nanopore-based sequencing, or direct imaging of DNA using advanced microscopy.
[0146] Embodiment 133. The method of embodiment 130, wherein the partial composition of the macromolecule is determined by analysis of multiple extended coding tag or ditag constructs using unique compartment tags and optionally UMIs.
[0147] Embodiment 134. The method of any one of embodiments 1 to 133, wherein the analyzing step is performed using a sequencing method having an error rate per base of >5%, >10%, >15%, >20%, >25%, or >30%.
[0148] Embodiment 135. The method of any one of embodiments 1 to 134, wherein the identification component of the coding tag, the recording tag, or both, includes an error correction code.
[0149] Embodiment 136. The method of embodiment 135, wherein the identification component is selected from an encoder sequence, a barcode, a UMI, a compartment tag, a cycle-specific sequence, or any combination thereof.
[0150] Embodiment 137. The method of embodiment 135 or 136, wherein the error correction code is selected from a Hamming code, a Lee distance code, an asymmetric Lee distance code, a Reed-Solomon code, and a Levenshtein-Tenengolts code.
[0151] Embodiment 138. A method according to any one of embodiments 1 to 134, wherein the identification component of the coding tag, the recording tag, or both, is capable of generating a unique current or ion flux or optical signature, and the analyzing step includes detecting the unique current or ion flux or optical signature to identify the identification component.
[0152] Embodiment 139. The method of embodiment 138, wherein the identification component is selected from an encoder sequence, a barcode, a UMI, a compartment tag, a cycle-specific sequence, or any combination thereof.
[0153] Embodiment 140. A method for analyzing a plurality of macromolecules, comprising: (a) providing a plurality of macromolecules conjugated to a solid support and associated recording tags; (b) contacting the plurality of macromolecules with a plurality of binding agents capable of binding to cognate macromolecules, each binding agent comprising a coding tag having identifying information related to the binding agent; (c) transferring information from a first coding tag of a first binding substance to a first recording tag associated with a first macromolecule to generate a primary extended recording tag, wherein the first binding substance binds to the first macromolecule; (d) contacting the plurality of macromolecules with a plurality of binding agents capable of binding to cognate macromolecules; (e) transferring information of a second coding tag of a second binding substance to the first extension recording tag to generate a second extension recording tag, wherein the second binding substance binds to the first macromolecule; (f) optionally repeating steps (d)-(e) for "n" binding cycles, transferring information from each coding tag of each binding substance that binds to the first macromolecule to an extended recording tag generated in a previous binding cycle to generate an nth extended recording tag representing the first macromolecule; (g) analyzing the n-th order decompressed record tag; A method comprising:
[0154] Embodiment 141. The method of embodiment 140, wherein a plurality of n-th order elongation recording tags representing a plurality of macromolecules are generated and analyzed.
[0155] Embodiment 142. The method of embodiment 140 or 141, wherein the macromolecule is a protein.
[0156] Embodiment 143. The method of embodiment 142, wherein the macromolecule is a peptide.
[0157] Embodiment 144. The method of embodiment 143, wherein the peptide is obtained by fragmenting a protein from a biological sample.
[0158] Embodiment 145. The method of any one of embodiments 140 to 144, wherein the plurality of macromolecules comprises macromolecules from multiple pooled samples.
[0159] Embodiment 146. The method of any one of embodiments 140 to 145, wherein the recording tag is a DNA molecule, an RNA molecule, a BNA molecule, an XNA molecule, an LNA molecule, a PNA molecule, a γPNA molecule, or a combination thereof.
[0160] Embodiment 147. The method of any one of embodiments 140 to 146, wherein the recording tag comprises a unique molecular identifier (UMI).
[0161] Embodiment 148. The method of any one of embodiments 140 to 147, wherein the recording tag comprises a compartment tag.
[0162] Embodiment 149. The method of any one of embodiments 140 to 148, wherein the recording tag comprises a universal priming site.
[0163] Embodiment 150. The method of any one of embodiments 140 to 149, wherein the recording tag comprises a spacer at its 3' end.
[0164] Embodiment 151. A method according to any one of embodiments 140 to 150, wherein the coding tag comprises an encoder sequence.
[0165] Embodiment 152. A method according to any one of embodiments 140 to 151, wherein the coding tag includes a UMI.
[0166] Embodiment 153. The method of any one of embodiments 140 to 152, wherein the coding tag comprises a universal priming site.
[0167] Embodiment 154. The method of any one of embodiments 140 to 153, wherein the coding tag comprises a spacer at its 3' end.
[0168] Embodiment 155. The method of any one of embodiments 140 to 154, wherein the coding tag comprises a binding cycle-specific sequence.
[0169] Embodiment 156. The method of any one of embodiments 140 to 155, wherein the coding tag comprises a unique molecular identifier.
[0170] Embodiment 157. The method of any one of embodiments 140 to 156, wherein the binding agent and the coding tag are joined by a linker.
[0171] Embodiment 158. The method of any one of embodiments 140 to 157, wherein the transfer of information from the recording tag to the coding tag is mediated by primer extension.
[0172] Embodiment 159. A method according to any one of embodiments 140 to 158, wherein the transfer of information from the recording tag to the coding tag is mediated by ligation.
[0173] Embodiment 160. The method of any one of embodiments 140 to 159, wherein the plurality of macromolecules, the associated recording tags, or both, are covalently attached to the solid support.
[0174] Embodiment 161. The method of any one of embodiments 140 to 160, wherein the solid support is a bead, a porous bead, a porous matrix, an array, a glass surface, a silicon surface, a plastic surface, a filter, a membrane, nylon, a silicon wafer chip, a flow-through chip, a biochip containing signal transduction electronics, a microtiter well, an ELISA plate, a spin interference disk, a nitrocellulose membrane, a nitrocellulose-based polymer surface, a nanoparticle, or a microsphere.
[0175] Embodiment 162. The method of embodiment 161, wherein the solid support is a polystyrene bead, a polymer bead, an agarose bead, an acrylamide bead, a solid core bead, a porous bead, a paramagnetic bead, a glass bead, or a controlled pore bead.
[0176] Embodiment 163. The method of any one of embodiments 140 to 162, wherein the binding agent is a polypeptide or protein.
[0177] Embodiment 164. The method of embodiment 163, wherein the binding agent is a modified aminopeptidase, a modified aminoacyl-tRNA synthetase, a modified anticalin, or an antibody or binding fragment thereof.
[0178] Embodiment 165. The method of any one of embodiments 142 to 164, wherein the binding agent binds to a single amino acid residue, a dipeptide, a tripeptide, or a post-translational modification of the peptide.
[0179] Embodiment 166 The method of embodiment 165, wherein the binding agent binds to an N-terminal amino acid residue, a C-terminal amino acid residue, or an internal amino acid residue.
[0180] Embodiment 167. The method of embodiment 165, wherein the binding agent binds to an N-terminal peptide, a C-terminal peptide, or an internal peptide.
[0181] Embodiment 168. The method of any one of embodiments 142 to 164, wherein the binding agent binds to a chemical label at a modified N-terminal amino acid residue, a modified C-terminal amino acid residue, or a modified internal amino acid residue.
[0182] Embodiment 169. The method of embodiment 166 or 168, wherein the binding agent is bound to an N-terminal amino acid residue or a chemical label of the modified N-terminal amino acid residue, and the N-terminal amino acid residue is cleaved after each binding cycle.
[0183] Embodiment 170. The method of embodiment 166 or 168, wherein the binding agent is attached to a chemical label at the C-terminal amino acid residue or the modified C-terminal amino acid residue, and the C-terminal amino acid residue is cleaved after each binding cycle.
[0184] Embodiment 171. The method of embodiment 169, wherein the N-terminal amino acid residue is cleaved by Edman degradation, Edmanase, a modified aminopeptidase, or a modified acylpeptide hydrolase.
[0185] Embodiment 172. The method of embodiment 163, wherein the binding agent is a site-specific covalent label of an amino acid or post-translational modification.
[0186] Embodiment 173. A method according to any one of embodiments 140 to 172, wherein the plurality of n-th order elongation recording tags are amplified before analysis.
[0187] Embodiment 174. The method of any one of embodiments 140 to 173, wherein the analysis of the n-th order extension recording tag comprises nucleic acid sequencing.
[0188] Embodiment 175. The method of embodiment 174, wherein multiple n-th order elongation recording tags representing multiple macromolecules are analyzed in parallel.
[0189] Embodiment 176. The method of embodiment 174 or 175, wherein the nucleic acid sequencing method is sequencing by synthesis, sequencing by ligation, sequencing by hybridization, polony sequencing, ion semiconductor sequencing, or pyrosequencing.
[0190] Embodiment 177. The method of embodiment 174 or 175, wherein the nucleic acid sequencing method is single-molecule real-time sequencing, nanopore-based sequencing, or direct imaging of DNA using advanced microscopy.
[0191] Non-limiting embodiments of the present invention will now be described by way of example with reference to the accompanying drawings, which are schematic and not intended to be drawn to scale. For illustrative purposes, not every component will be labeled in the drawings, nor will every component of each embodiment of the present invention be shown unless illustration is necessary for an understanding of the invention by those skilled in the art. [Brief explanation of the drawings]
[0192] [Figure 1-1] FIG. 1A shows a legend for the functional elements shown in the figure. FIG. 1B shows a basic outline for converting a protein code to a DNA code. In this conversion, proteins or polypeptides are fragmented into peptides, which are then converted into a library of extended record tags representing the peptides. The extended record tags constitute a DNA-encoded library representing the peptide sequences. The library can be appropriately modified and sequenced on any next-generation sequencing (NGS) platform. [Figure 1-2] Same as above.
[0193] [Figure 2-1]Figures 2A-2D show binding agents (e.g., antibodies, anticalins, N-recognins) containing coding tags that interact with immobilized proteins that are colocalized or colabeled with single or multiple recording tags. An example of protein macromolecule analysis by the methods disclosed herein is shown, using multiple cycles of a specific protein (e.g., ATP-dependent Clp protease adaptor protein (ClpS), aptamers, etc., and their mutants / homologs). The recording tag is composed of a universal priming site, a barcode (e.g., partition barcode, compartment barcode, fraction barcode), an optional unique molecular identifier (UMI) sequence, and a spacer sequence (Sp) used for information transfer from the coding tag. The spacer sequence (Sp) may be constant across all binding cycles, may be specific to the binder, or may be specific to the number of binding cycles. The coding tag hybridizes to an encoder sequence that provides identification information for the binder, an optional UMI, and a complementary spacer sequence of the recording tag, allowing for transfer of the coding tag information to the recording tag (e.g., primer extension, herein referred to as polymerase extension). The extended recording tag is composed of a spacer sequence that facilitates the binding of cognate binders to proteins (also referred to as "binding agent coding tag"). Figure 2A illustrates the process of creating an extended recording tag by cyclic binding of a cognate binder to a protein and the corresponding information transfer from the binder's coding tag to the protein's recording tag. After a series of successive binding and coding tag information transfer steps, a final extended recording tag is produced that includes binder coding tag information, including the encoder sequence from "n" binding cycles, providing the identity of the binder (e.g., antibody 1 (Ab1), antibody 2 (Ab2), antibody 3 (Ab3), ... antibody "n" (Abn)), a barcode / optional UMI derived from the recording tag, an optional UMI sequence derived from the binder's coding tag, and flanking universal priming sequences at each end of the library construction to facilitate amplification and analysis by digital next-generation sequencing. Figure 2B illustrates an example scheme for labeling proteins with DNA barcoded recording tags.In the top panel, N-hydroxysuccinimide (NHS) is an amine-reactive coupling agent, and dibenzocyclooctyl (DBCO) is a strained alkyne useful for "click" coupling to the surface of a solid substrate. In this scheme, a recording tag is coupled to the epsilon amine of a lysine (K) residue (and optionally the N-terminal amino acid) of a protein via the NHS moiety. In the bottom panel, a heterobifunctional linker, NHS-alkyne, is used to label the epsilon amine of the lysine (K) residue, creating an alkyne "click" moiety. Azide-labeled DNA recording tags can then be easily attached to these reactive alkyne groups via standard click chemistry. Furthermore, DNA recording tags can be designed with an orthogonal methyltetrazine (mTet) moiety for downstream coupling to a TCO-derivatized sequencing substrate via a reverse iEDDA reaction. Figure 2C shows two examples of protein analysis methods using recording tags. In the top panel, a protein macromolecule is immobilized to a solid support by a capture agent and optionally crosslinked. Either the protein or the capture agent may be labeled with a recording tag. In the bottom panel, a protein with an associated recording tag is directly immobilized on a solid support. Figure 2D shows an example of the overall workflow of a simple protein immunoassay using DNA encoding of cognate binders and sequencing of the resulting extended recording tag. Proteins can be barcoded (i.e., indexed) with the recording tag and pooled prior to cycle binding analysis, significantly increasing sample throughput and binding reagent savings. This approach represents an effective, digitally enabled, simpler, and more scalable method for performing reverse-phase protein assays (RPPA). [Figure 2-2] Same as above. [Figure 2-3] Same as above. [Figure 2-4] Same as above.
[0194] [Figure 3]3A-3D show the process of degradation-based peptide sequencing by constructing a DNA-extended recording tag that represents the peptide sequence. This is achieved using a cyclic process of N-terminal amino acid (NTAA) binding, transfer of coding tag information to the recording tag attached to the peptide, and NTAA cleavage, all of which are repeated in a cyclical manner on a solid support. An outline of an exemplary construction of an extended recording tag derived from N-terminal degradation of a peptide is provided: (A) the N-terminal amino acid of the peptide is labeled (e.g., with phenylthiocarbamoyl (PTC), dinitrophenyl (DNP), sulfonylnitrophenyl (SNP), acetyl, or guanidindyl moiety); (B) shows a binding agent and associated coding tag attached to the labeled NTAA; (C) shows a peptide attached to a solid support (e.g., a bead) and attached to the recording tag (e.g., via a trifunctional linker); upon binding of the binding agent to the NTAA of the peptide, information from the coding tag is transferred to the recording tag (e.g., by primer extension), generating the extended recording tag; (D) the labeled NTAA is cleaved by chemical or enzymatic means to expose a new NTAA. This cycle is repeated "n" times, as indicated by the arrows, to generate the final extended recording tag. The final extended recording tag is optionally flanked by universal priming sites to facilitate downstream amplification and DNA sequencing. The forward universal priming site (e.g., Illumina's P5-S1 sequence) may be part of the original recording tag design, and the reverse universal priming site (e.g., Illumina's P7-S2' sequence) may be added as a final step in extending the recording tag, which may be performed independently of the binding agent.
[0195] [Figure 4-1]4A-4B show an exemplary protein sequencing workflow according to the methods disclosed herein. FIG. 4A shows an exemplary workflow, with alternative modes outlined in light gray dashed lines and specific embodiments shown in boxes associated with arrows. Alternative modes for each step of the workflow are shown in boxes below the arrows. FIG. 4B shows options for improving the efficiency of information transfer when performing the cycle binding and coding tag information transfer steps. Multiple recording tags can be used per molecule. Furthermore, for a given binding event, multiple transfers of coding tag information to recording tags can be performed, or alternatively, a surface amplification step can be used to create copies, such as an extended recording tag library. [Figure 4-2] Same as above.
[0196] [Figure 5]Figures 5A-5B outline an exemplary construction of an extended recording tag for transferring identifying information from a coding tag of a binding agent to a recording tag associated with a macromolecule (e.g., a peptide) using primer extension to generate an extended recording tag. The coding tag includes a unique encoder sequence with identifying information for the binding agent, optionally flanked at each end by a common spacer sequence (Sp'). Figure 5A shows an NTAA-binding agent including a coding tag that binds to the NTAA of a recording tag-labeled peptide linked to a bead. The recording tag anneals to the coding tag via a complementary spacer sequence (Sp), and a primer extension reaction mediates the transfer of the coding tag information to the recording tag using the spacer (Sp) as a priming site. The coding tag is shown as a duplex with a single-stranded spacer (Sp') sequence at the end distal to the binding agent. This configuration minimizes hybridization of the coding tag to internal portions of the recording tag and favors hybridization of the recording tag's terminal spacer (Sp) sequence with the coding tag's single-stranded spacer overhang (Sp'). Additionally, the extended recording tag may be pre-annealed with an oligonucleotide (encoder, complementary to the spacer sequence) to prevent hybridization of the coding tag with internal recording tag sequence elements. Figure 5B shows the final extended recording tag produced after "n" cycles of binding, transfer of coding tag information, and addition of a universal priming site to the 3' end ("***" represents intervening binding cycles not shown in the extended recording tag).
[0197] [Figure 6]Figure 6 shows that coding tag information is transferred to an extended recording tag by enzymatic ligation. Two different macromolecules with respective recording tags are shown, and recording tag extension proceeds in parallel. Ligation can be facilitated by designing the double-stranded coding tag so that the spacer sequence (Sp) has a "sticky end" overhang that anneals to the complementary spacer (Sp') of the recording tag. The complementary strand of the double-stranded coding tag transfers the information to the recording tag. When ligation is used to extend the recording tag, the direction of extension can be 5' to 3' as shown, or optionally 3' to 5'.
[0198] [Figure 7] Figure 7 illustrates a "spacerless" approach to transferring coding tag information to a recording tag by chemical ligation, linking the recording tag or the 3' nucleotide of the extended recording tag to the 5' nucleotide of the coding tag (or its complement) without inserting a spacer sequence into the extended recording tag. Alternatively, the orientation of the extended recording tag and coding tag may be reversed, with the 5' end of the recording tag ligated to the 3' end of the coding tag (or its complement). In the example shown, hybridization of the recording tag's complementary "helper" oligonucleotide sequence ("recording helper") to the coding tag is used to stabilize the complex and allow specific chemical ligation of the recording tag to the coding tag's complementary strand. The resulting extended recording tag lacks a spacer sequence. Also illustrated is "click chemistry"-type chemical ligation, e.g., using azide and alkyne moieties (shown as triple-line symbols) that can be used with DNA, PNA, or similar nucleic acid polymers.
[0199] [Figure 8]8A-8B show an exemplary method for writing post-translational modification (PTM) information of a peptide into an extended recording tag prior to N-terminal amino acid degradation. Figure 8A: A binding agent (e.g., a phosphotyrosine antibody) containing a coding tag with identifying information for the binding agent can be bound to the peptide. If a phosphotyrosine is present in the recording tag-labeled peptide as shown, upon binding of the phosphotyrosine antibody to the phosphotyrosine, the coding tag and recording tag anneal via the complementary spacer sequence, transferring the coding tag information to the recording tag and generating an extended recording tag. Figure 8B: The extended recording tag can contain coding tag information for both the primary amino acid sequence (e.g., "aa1," "aa2," "aa3," ..., "aaN") and the post-translational modifications of the peptide (e.g., "PTM1," "PTM2").
[0200] [Figure 9]Figures 9A-9B illustrate a multi-cycle process of binding a binding agent to a macromolecule and transferring information from a coding tag attached to the binding agent to an individual recording tag among multiple recording tags co-localized at the site of a single macromolecule attached to a solid support (e.g., a bead), thereby generating multiple extended recording tags that collectively represent the macromolecule. In this figure, for illustrative purposes only, the macromolecule is a peptide, and each cycle involves binding a binding agent to the N-terminal amino acid (NTAA), recording the binding event by transferring the coding tag information to the recording tag, and then removing the NTAA to expose a new NTAA. Figure 9A shows multiple recording tags (including a universal forward priming sequence and a UMI) co-localized with the macromolecule on a solid support. Each recording tag has a common spacer sequence (Sp) complementary to the common spacer sequence in the coding tag of the binding agent, which can be used to prime an extension reaction and transfer the coding tag information to the recording tag. FIG. 9B shows different pools of cycle-specific NTAA binding agents used for binding in each successive cycle, each pool having a cycle-specific spacer sequence.
[0201] [Figure 10]Figures 10A-10C show an exemplary mode involving multiple cycles of transferring information from a coding tag attached to a binding substance to one of multiple recording tags co-localized at the site of a single macromolecule attached to a solid support (e.g., a bead), thereby generating multiple extended recording tags that collectively represent the macromolecule. In this figure, for illustrative purposes only, the macromolecule is a peptide, and each round of the process involves binding to an NTAA, recording the binding event, and then removing the NTAA to expose a new NTAA. Figure 10A shows multiple recording tags (including a universal forward priming sequence and a UMI) co-localized on a solid support with a macromolecule, preferably a single molecule per bead. Each recording tag has a different spacer sequence at its 3' end, with a different "cycle-specific" sequence (e.g., C1, C2, C3, ... Cn). Preferably, the recording tags on each bead share the same UMI sequence. In the first cycle of binding (cycle 1), multiple NTAA-binding substances are contacted with the macromolecule. The binding agent used in cycle 1 has a common 5' spacer sequence (C'1) complementary to the cycle 1 C1 spacer sequence of the recording tag. The binding agent used in cycle 1 also has a 3' spacer sequence (C'2) complementary to the cycle 2 spacer C2. During binding cycle 1, a first NTAA binding agent binds to the free N-terminus of the macromolecule, and the information of the first coding tag is transferred to the cognate recording tag by primer extension from the C1 sequence hybridized to the complementary C'1 spacer sequence. After removing the NTAA to expose a new NTAA, binding cycle 2 contacts the macromolecule with multiple NTAA binding agents that have a cycle 2 5' spacer sequence (C'2) identical to the 3' spacer sequence of the cycle 1 binding agent and a common cycle 3 3' spacer sequence (C'3). The second NTAA-binding substance binds to the NTAA of the macromolecule, and the information of the second coding tag is transferred from the complementary C2 and C'2 spacer sequences to the cognate coding tag by primer extension.These cycles are repeated for up to "n" binding cycles, capping the final extension recording tag with a universal reverse priming sequence to generate multiple extension recording tags colocalized with a single macromolecule, each with coding tag information from one binding cycle. Each set of binding agents used in each successive binding cycle has a cycle-specific spacer sequence in its coding tag, allowing the binding cycle information to be associated with the binding agent information of the resulting extension recording tag. Figure 10B shows various pools of cycle-specific binding agents used in each successive binding cycle, each pool with a cycle-specific spacer sequence. Figure 10C illustrates how cycle-specific spacer sequences can be used to sequentially assemble a collection of extension recording tags colocalized with sites on a macromolecule based on PCR assembly of the extension recording tags, thereby providing an ordered sequence of the macromolecule. In a preferred mode, multiple copies of each extension recording tag are generated by amplification prior to concatemer formation.
[0202] [Figure 11] 11A-11B show information transfer from recording tags to coding tag or ditag constructs. Two methods for recording binding information are shown (A) and (B). The binding substance can be any type of binding substance described herein; while an anti-phosphotyrosine binding substance is shown, it is for illustrative purposes only. In the case of extended coding tag or ditag constructs, rather than transferring binding information from the coding tag to the recording tag, the information is either transferred from the recording tag to the coding tag to generate an extended coding tag (A), or information is transferred from both the recording tag and the coding tag to a third ditag-forming construct (B). Ditags and extended coding tags contain the recording tag (barcode, optional UMI sequence, and optional compartment tag (CT) sequence (not shown)) and the coding tag information. Ditags and extended coding tags can be eluted from the recording tag, collected, optionally amplified, and read by a next-generation sequencer.
[0203] [Figure 12] Figures 12A-12D show the design of PNA combinatorial barcode / UMI recording tags and ditag detection of binding events. Figure 12A shows the construction of combinatorial PNA barcodes / UMIs of four basic PNA word sequences (A, A'-B, B'-C, and C') by chemical ligation. Hybridization of DNA arms is included to create a spacerless combinatorial template for the combinatorial assembly of PNA barcodes / UMIs. Chemical ligation is used to stitch together annealed PNA "words." Figure 12B shows a method for transferring the PNA information of recording tags to DNA intermediates. The DNA intermediates can transfer information to coding tags. That is, complementary DNA word sequences are annealed to the PNA and chemically ligated (optionally, enzymatically ligated if a ligase that uses the PNA template is discovered). In Figure 12C, the DNA intermediate is designed to interact with the coding tag via the spacer sequence Sp. A strand displacement primer extension step displaces the ligated DNA and transfers the recording tag information from the DNA intermediate to the coding tag, generating an extended coding tag. Terminator nucleotides may be incorporated at the end of the DNA intermediate to prevent transfer of the coding tag information to the DNA intermediate by primer extension. Figure 12D: Alternatively, information may be transferred from the coding tag to the DNA intermediate to generate a ditag construct. Terminator nucleotides may be incorporated at the end of the coding tag to prevent transfer of the recording tag information from the DNA intermediate to the coding tag.
[0204] [Figure 13]Figures 13A-13E show the generation of a library of elements representing peptide sequence composition by partitioning a proteome onto compartment bar-coated beads followed by ditag assembly via emulsion fusion PCR. The amino acid content of the peptides can then be characterized by N-terminal sequencing or, alternatively, by attaching (covalently or noncovalently) binding entities carrying amino acid-specific chemical labels or coding tags. The coding tags consist of a universal priming sequence, an encoder sequence for identifying the amino acid, a compartment tag, and the amino acid UMI. After information transfer, the ditags are mapped back to the original molecules via the recording tag UMI. In Figure 13A, the proteome is compartmentalized into droplets containing bar-coded beads. Peptides with associated recording tags (containing compartment barcode information) are attached to the bead surface. The droplet emulsion is then broken to release the bar-coded beads onto which the peptides are distributed. In Figure 13B, specific amino acid residues of a peptide are chemically labeled with a DNA coding tag conjugated to a site-specific labeling moiety. The DNA coding tag contains amino acid barcode information and optionally an amino acid UMI. Figure 13C: The labeled peptide-recording tag complex is released from the bead. Figure 13D: The labeled peptide-recording tag complex is emulsified into a nanoemulsion or microemulsion such that there is an average of less than one peptide-recording tag complex per compartment. Figure 13E: Emulsion fusion PCR transfers the recording tag information (e.g., compartment barcode) to all of the DNA coding tags attached to amino acid residues.
[0205] [Figure 14]Figure 14 shows the generation of extended coding tags from emulsified peptide recording tag-coding tag complexes. The peptide complexes of Figure 13C are co-emulsified with PCR reagents into droplets, resulting in an average of a single peptide complex per droplet. A three-primer fusion PCR approach is used to amplify the recording tag associated with the peptide, fuse the amplified recording tag with multiple binding substance coding tags or covalently labeled amino acid coding tags, extend the coding tag by primer extension to transfer peptide UMI and compartment tag information from the recording tag to the coding tag, and amplify the resulting extended coding tag. Multiple extended coding tag species are present per droplet, with a different species for each amino acid encoder sequence-UMI coding tag present. In this way, both the identity and count of amino acids within a peptide can be determined. The U1 universal primer and Sp primer are designed to have a higher melting Tm than the U2tr universal primer. This allows for two-stage PCR, in which the first few cycles are performed at a higher annealing temperature to amplify the recording tag, followed by a step with a lower Tm so that the recording tag and coding tag can prime each other during PCR to produce an extended coding tag, and the U1 and U2tr universal primers are used to prime the amplification of the resulting extended coding tag product. In certain embodiments, premature polymerase extension from the U2tr primer can be prevented by using a photolabile 3' blocking group (Young et al., 2008, Chem. Commun. (Camb) 4:462-464). After the first round of PCR to amplify the recording tag and the second round of fusion PCR step in which the coding tag Sptr primes the extension of the coding tag from the amplified Sp' sequence of the recording tag, the 3' blocking group of U2tr is removed, and PCR is initiated at a higher temperature to amplify the extended coding tag with the U1 and U2tr primers.
[0206] [Figure 15] Figure 15 illustrates the use of proteome partitioning and barcoding to facilitate enhanced protein mappability and phasing. Peptide sequencing typically involves digesting proteins into peptides. This process loses information about the relationships between individual peptides derived from the parent protein molecule and their relationship to the parent protein molecule. To reconstruct this information, individual peptide sequences are mapped back to the collection of protein sequences from which they may have been derived. The task of finding unique matches within such a set becomes more difficult with increasing collection size and complexity (e.g., proteome sequence complexity) for short and / or partial peptide sequences. Partitioning the proteome into barcoded (e.g., compartment-tagged) compartments or sections, followed by digesting proteins into peptides and splicing the compartment tags onto the peptides, reduces the "protein" space over which peptide sequences must be mapped, significantly simplifying the task for complex protein samples. Labeling proteins with unique molecular identifiers (UMIs) before digesting them into peptides facilitates mapping peptides back to the original protein molecule, allowing annotation of phasing information between post-translational modification (PTM) variants derived from the same protein molecule and identification of individual proteoforms. Figure 15A shows an example of proteome partitioning, involving labeling proteins with recording tags containing partitioning barcodes and subsequent fragmentation into recording tag-labeled peptides. Figure 15B: Even with partial peptide sequence or composition information alone, this mapping is highly degenerate. However, partial peptide sequence or composition information combined with information from multiple peptides derived from the same protein allows for unique identification of the original protein molecule.
[0207] [Figure 16]Figure 16 shows an exemplary mode of compartment-tagged bead array design. The compartment tag includes a barcode of X5-20 to identify an individual compartment and a unique molecular identifier (UMI) of N5-10 to identify the peptide to which the compartment tag is attached, where X and N represent degenerate nucleobases or nucleobase words. The compartment tag can be single-stranded (shown in the top row) or double-stranded (shown in the bottom row). Optionally, the compartment tag can be a chimeric molecule (shown on the left) containing a peptide sequence with a recognition sequence for a protein ligase (e.g., butelase I) for ligation with a peptide of interest. Alternatively, a chemical moiety can be included in the compartment tag for coupling with a peptide of interest (e.g., an azide, as shown in the right diagram).
[0208] [Figure 17]17A-17B show an exemplary method for target peptide enrichment using (A) multiple extended recording tags representing multiple peptides and (B) standard hybrid capture techniques. For example, hybrid capture enrichment may use one or more biotinylated "bait" oligonucleotides that hybridize with extended recording tags representing one or more peptides of interest ("target peptides") from a library of extended recording tags representing a library of peptides. The bait oligonucleotide:targeted extended recording tag hybridization pair is pulled down from solution by the biotin tag after hybridization to generate an enriched fraction of extended recording tags representing one or more peptides of interest. Separation ("pull-down") of the extended recording tags can be achieved, for example, using streptavidin-coated magnetic beads. The biotin moiety is bound to the streptavidin on the beads, and separation is achieved by localizing the beads using a magnet and removing or exchanging the solution. Optionally, non-biotinylated competitor enrichment oligonucleotides that competitively hybridize to the extended recording tags representing undesired or over-abundant peptides can be included in the hybridization step of the hybrid capture assay to modulate the amount of enriched target peptides. The non-biotinylated competitor oligonucleotides compete for hybridization with the target peptide, but the hybridization duplexes are not captured during the capture step due to the absence of biotin moieties. Therefore, by adjusting the ratio of competitor oligonucleotides to biotinylated "bait" oligonucleotides, the enriched extended recording tag fraction can be modulated over a wide dynamic range. This step will be important for addressing the dynamic range problem of protein abundance within a sample.
[0209] [Figure 18]18A-18B show exemplary methods for distributing single cells and bulk proteomes into individual droplets, each containing beads with multiple compartment tags attached to correlate peptides with their original protein complexes or proteins derived from a single cell. The compartment tags include barcodes. Manipulation of droplet components after droplet formation: (A) distributing single cells into individual droplets, followed by cell lysis to release the cellular proteome, proteolytically digesting the cellular proteome into peptides, and inactivating the proteases after sufficient proteolysis; (B) distributing the bulk proteome into multiple droplets, each containing a protein complex, proteolytically digesting the protein complex into peptides, and inactivating the proteases after sufficient proteolysis. Thermolabile metallo-proteases can be used to digest encapsulated proteins into peptides after photoactivation of photocaged divalent cations. The protease may be heat-inactivated after sufficient proteolysis or may chelate divalent cations. The droplets contain a hybridized or releasable compartment tag that contains a nucleic acid barcode (separate from the recording tag) that can be ligated to either the N- or C-terminal amino acid of the peptide.
[0210] [Figure 19]Figures 19A-19B show an exemplary method for distributing single cells and bulk proteomes into individual droplets, each containing beads with multiple bifunctional recording tags attached to compartment tags for correlating peptides with their original proteins or protein complexes, or proteins with their original single cells. Manipulation of droplet components after droplet formation: (A) Distributing single cells into individual droplets, followed by cell lysis to release the cellular proteome, proteolytically digesting the cellular proteome into peptides, and inactivating the proteases after sufficient proteolysis; (B) Distributing the bulk proteome into multiple droplets, each containing a protein complex, followed by proteolytically digesting the protein complex into peptides, and inactivating the proteases after sufficient proteolysis. Thermolabile metallo-proteases can be used to digest encapsulated proteins into peptides after photorelease of a photocaged divalent cation (e.g., Zn2+). The protease may be heat-inactivated after sufficient proteolysis or may chelate divalent cations. The droplets contain a hybridized or releasable compartment tag that contains a nucleic acid barcode (separate from the recording tag) that can be ligated to either the N- or C-terminal amino acid of the peptide.
[0211] [Figure 20-1]Figures 20A-20L show the generation of compartment-barcoded recording tags attached to peptides. Compartment barcoding techniques (e.g., barcoded beads in microfluidic droplets) can be used to transfer compartment-specific barcodes to molecular contents encapsulated within specific compartments. (A) In certain embodiments, protein molecules are denatured, and the ε-amine groups of lysine residues (K) are chemically conjugated to activated universal DNA tag molecules (containing a universal priming sequence (U1) shown with an NHS moiety at the 5' end). After conjugation of the universal DNA tag to the polypeptide, excess universal DNA tag is removed. (B) The universal DNA-tagged polypeptide is hybridized with nucleic acid molecules bound to beads, and each bead-bound nucleic acid molecule contains a unique population of compartment tag (barcode) sequences. Compartmentalization can occur by separating a sample into distinct physical compartments (indicated by dashed ellipses), such as droplets. Alternatively, compartmentalization can be achieved directly, without the need for additional physical separation, by immobilizing the labeled polypeptide on the bead surface, for example, by annealing the polypeptide's universal DNA tag to the bead's compartment DNA tag. A single polypeptide molecule interacts only with a single bead (e.g., a single polypeptide does not span multiple beads). However, multiple polypeptides may interact with the same bead. In addition to the compartment barcode sequence (BC), the nucleic acid molecule bound to the bead may consist of a common Sp (spacer) sequence, a unique molecular identifier (UMI), and a sequence complementary to the polypeptide DNA tag U1'. (C) After annealing the universal DNA-tagged polypeptide to the bead-bound compartment tag, the compartment tag is released from the bead by cleaving the attached linker. (D) The annealed U1 DNA tag primer is extended by polymerase-based primer extension using the bead-derived compartment tag nucleic acid molecule as a template.The primer extension step can be performed after releasing the compartment tag from the bead, as shown in (C), or optionally while the compartment tag is still attached to the bead (not shown). This effectively writes the barcode sequence from the bead's compartment tag into the polypeptide's U1 DNA tag sequence. This new sequence constitutes the recording tag. After primer extension, the polypeptide is cleaved into peptide fragments using a protease, e.g., Lys-C (cleaves C-terminal to lysine residues), Glu-C (cleaves C-terminal to glutamic acid residues, and to a lesser extent, glutamic acid residues), or a random protease such as proteinase K. (E) Each peptide fragment is labeled at its C-terminal lysine with the extended DNA tag sequence that constitutes the recording tag for downstream peptide sequencing as disclosed herein. (F) The recording-tagged peptides are coupled to azide beads via strained alkyne-labeled DBCO. The azide beads also optionally contain a capture sequence complementary to the recording tag to facilitate the efficiency of DBCO-azide immobilization. It should be noted that removing the peptide from the original bead and re-immobilizing it on a new solid support (e.g., beads) allows optimal intermolecular spacing between the peptides, facilitating the peptide sequencing methods disclosed herein. Figures 20G-20L illustrate a similar concept to that shown in Figures 20A-20F, except that they use click chemistry conjugation of a DNA tag to an alkyne-prelabeled polypeptide (as described in Figure 2B). Azide and mTet chemistries are orthogonal, allowing click conjugation to DNA tags and click iEDDA conjugation (mTet and TCO) to sequencing substrates. [Figure 20-2] Same as above. [Figure 20-3] Same as above. [Figure 20-4] Same as above.
[0212] [Figure 21]Figure 21 shows an exemplary method for compartmentalizing single cells into compartment-tagged (e.g., barcoded) beads using a flow-focusing T-junction. Using two water streams, cell lysis and protease activation (Zn2+ mixing) can be easily initiated upon droplet formation.
[0213] [Figure 22] 22A-22B show exemplary tagging details. (A) Compartment tags (DNA-peptide chimeras) are attached to peptides using peptide ligation with Butelase I. (B) Compartment tag information is transferred to the associated recording tag before peptide sequencing begins. Optionally, the compartment tag can be cleaved after information transfer to the recording tag using the endopeptidase AspN, which selectively cleaves peptide bonds at the N-terminus of aspartic acid residues.
[0214] [Figure 23]23A-23C show array-based barcoding for spatial proteomic analysis of tissue sections. (A) An array of spatially encoded DNA barcodes (featuring barcodes designated BCij) is combined with tissue sections (FFPE or frozen). In one embodiment, the tissue sections are fixed and permeabilized. In a preferred embodiment, the array feature size is smaller than the cell size (approximately 10 μm for human cells). (B) Array-mounted tissue sections are treated with reagents to reverse crosslinking (e.g., citraconic anhydride-based antigen retrieval protocol (Namimatsu, Ghazizadeh, et al., 2005)), and then proteins within are labeled with site-reactive DNA labels (e.g., labeling of lysines released after antigen retrieval), which effectively labels every protein molecule with a DNA recording tag. After labeling and washing, the array-bound DNA barcode sequences are cleaved and allowed to diffuse into the mounted tissue section, where they hybridize with the DNA recording tags attached to proteins within it. (C) The array-mounted tissue is then subjected to polymerase extension to transfer the information from the hybridized barcodes to the DNA recording tags that label the proteins. After barcode transfer, the array-mounted tissue is scraped off the slide and, optionally, digested with a protease to extract proteins or peptides into solution.
[0215] [Figure 24-1]Figures 24A-24B show two different exemplary DNA target macromolecules (AB and CD) immobilized on beads and assayed with binding entities attached to coding tags. This model system serves to illustrate the single-molecule behavior of coding tag transfer from bound entities to proximal recording tags. In a preferred embodiment, the coding tag is incorporated into an extended recording coding tag by primer extension. Figure 24A shows an AB macromolecule interacting with an A-specific binding entity ("A'", an oligonucleotide sequence complementary to the "A" component of the AB macromolecule), with the associated coding tag information being transferred to the recording tag by primer extension, and with a B-specific binding entity ("B'", an oligonucleotide sequence complementary to the "B" component of the AB macromolecule), with the associated coding tag information being transferred to the recording tag by primer extension. Coding tags A and B differ in sequence and, in this illustration, in length for easy identification. The different lengths facilitate analysis of the coding tag transfer by gel electrophoresis, but are not required for analysis by next-generation sequencing. Binding of A' and B' binding substances is shown as an alternative possibility for a single binding cycle. Adding a second cycle would further extend the extended recording tag. Depending on whether an A' or B' binding substance is added in the first and second cycles, the extended recording tag can contain coding tag information in the form AA, AB, BA, and BB. Thus, the extended recording tag contains information not only about the identity of the binding substance but also about the order of the binding events. Similarly, Figure 24B shows a CD macromolecule interacting with a C-specific binding substance ("C'," an oligonucleotide sequence complementary to the "C" component of the CD macromolecule), transferring the associated coding tag information to the recording tag by primer extension, and with a D-specific binding substance ("D'," an oligonucleotide sequence complementary to the "D" component of the CD macromolecule), transferring the associated coding tag information to the recording tag by primer extension. Coding tags C and D differ in sequence and, in this illustration, in length for easy identification.The different lengths facilitate analysis of coding tag transitions by gel electrophoresis, but are not required for analysis by next-generation sequencing. Binding of C' and D' binders is shown as an alternative possibility for a single binding cycle. Adding a second cycle would further extend the extended recording tag. Depending on whether a C' or D' binder is added in the first or second cycle, the extended recording tag can contain coding tag information in the form CC, CD, DC, and DD. The coding tag may optionally include a UMI. The inclusion of a UMI in the coding tag allows additional information about the binding event to be recorded, thereby enabling binding events to be distinguished at the level of individual binders. This can be useful when an individual binder can participate in more than one binding event (e.g., when its binding affinity is such that it can dissociate and reassociate frequently enough to participate in more than one event). It can also be useful for error correction. For example, under some circumstances, a coding tag may transfer information to a recording tag twice or more times in the same binding cycle, and the use of UMI will reveal that these are likely to be repeated information transfer events all associated with a single binding event. [Figure 24-2] Same as above.
[0216] [Figure 25]Figure 25 shows an exemplary DNA target macromolecule (AB) immobilized on a bead and assayed by a binding substance attached to a coding tag. The A-specific binding substance ("A'", an oligonucleotide complementary to the A component of the AB macromolecule) interacts with the AB macromolecule, and the information in the associated coding tag is transferred to the recording tag by ligation. The B-specific binding substance ("B'", an oligonucleotide complementary to the B component of the AB macromolecule) interacts with the AB macromolecule, and the information in the associated coding tag is transferred to the recording tag by ligation. Coding tags A and B differ in sequence and, in this illustration, in length for easy identification. The different lengths facilitate analysis of the coding tag transfer by gel electrophoresis, but the different lengths are not required for analysis by next-generation sequencing.
[0217] [Figure 26]Figures 26A-26B show exemplary DNA-peptide macromolecules for binding / coding tag transfer by primer extension. Figure 26A shows an exemplary oligonucleotide-peptide target macromolecule ("A" oligonucleotide-cMyc peptide) immobilized on a bead. A cMyc-specific binding substance (e.g., an antibody) interacts with the cMyc peptide portion of the macromolecule, transferring the associated coding tag information to the recording tag. Transfer of the cMyc coding tag information to the recording tag can be analyzed by gel electrophoresis. Figure 26B shows an exemplary oligonucleotide-peptide target macromolecule ("C" oligonucleotide-hemagglutinin (HA) peptide) immobilized on a bead. An HA-specific binding substance (e.g., an antibody) interacts with the HA peptide portion of the macromolecule, transferring the associated coding tag information to the recording tag. Transfer of the coding tag information to the recording tag can be analyzed by gel electrophoresis. Binding of a cMyc antibody-coding tag and an HA antibody-coding tag is shown as alternative possibilities for a single binding cycle. If a second cycle is performed, the extended recording tag will be further extended. Depending on whether a cMyc antibody-coding tag or an HA antibody-coding tag is added in the first and second binding cycles, the extended recording tag can contain coding tag information in the form cMyc-HA, HA-cMyc, cMyc-cMyc, and HA-HA. Also, although not shown, additional binding entities can be introduced to enable detection of the A and C oligonucleotide components of the macromolecule. Thus, hybrid macromolecules containing different types of backbones can be analyzed by transferring information to the recording tag and reading out the extended recording tag, which contains information about the order of binding events and the identity of the binding entities.
[0218] [Figure 27-1]Figures 27A-27D show the generation of error-correction barcodes. (A) A subset of 65 error-correction barcodes (SEQ ID NOs: 1-65) was selected from a set of 77 barcodes derived from the R software package "DNABarcodes" (https: / / bioconductor.riken.jp / packages / 3.3 / bioc / manuals / DNABarcodes / man / DNABarcodes.pdf) using the command parameters [create.dnabarcodes(n=15, dist=10)]. This algorithm generates 15-mer "Hamming" barcodes capable of correcting substitution errors up to a distance of four substitutions and detecting errors up to nine substitutions. The subset of 65 barcodes was generated by filtering out barcodes that did not exhibit various nanopore current levels (as in nanopore-based sequencing) or that were correlated with other members of this set. (B) Plot of predicted nanopore current levels for 15-mer barcodes passing through the pore. Predicted currents were calculated by dividing each 15-mer barcode word into a composite set of 11 overlapping 5-mer words and using the 5mer R9 nanopore current level lookup table (template_median68pA. 5mers.model (https: / / github.com / jts / nanopolish / tree / master / etc / r9-models) to predict the corresponding current levels as the barcodes thread through the nanopore one base at a time. As can be seen in (B), this set of 65 barcodes exhibits a unique current signature for each of its members. (C) The generation of a PCR product as a model extended recording tag for nanopore sequencing using an overlapping set of DTR and DTR primers is shown. The PCR amplicons are then ligated to form a chained extended recording tag model. (D) An exemplary "extended recording tag" model nanopore sequencing read (read length 734 bases) generated as shown in Figure 27C. MinIon The R9.4 read has a quality score of 7.2 (poor read quality).However, barcode sequences can be easily identified using lalign, even when read quality is low (Qscore = 7.2). The 15mer spacer element is underlined. Barcodes can be aligned in either the forward or reverse direction, denoted by the BC or BC' symbols. [Figure 27-2] Same as above. [Figure 27-3] Same as above.
[0219] [Figure 28]Figures 28A-28D show analyte-specific labeling of proteins with recording tags. (A) A binding agent that targets a protein analyte of interest in its native conformation contains an analyte-specific barcode (BCA') that hybridizes to the complementary analyte-specific barcode (BCA) of a DNA recording tag. Alternatively, the DNA recording tag can be attached to the binding agent via a cleavable linker, which is "clicked" directly onto the protein and then cleaved from the binding agent (via the cleavable linker). The DNA recording tag contains a reactive coupling moiety (a click chemistry reagent (e.g., azide, mTet, etc.) for coupling to a protein of interest) and other functional components (e.g., a universal priming sequence (P1), a sample barcode (BCS), an analyte-specific barcode (BCA), and a spacer sequence (Sp)). The sample barcode (BCS) can be used to label and distinguish proteins from different samples. The DNA recording tag may also contain an orthogonal coupling moiety (e.g., mTet) for subsequent coupling to a substrate surface. When the recording tag is click-chemistry coupled to a protein of interest, the protein is coupled to a click chemistry coupling moiety of the same type as the click chemistry coupling moiety of the DNA recording tag. (B) After the binding entity binds to the proximal target protein, the reactive coupling moiety (e.g., azide) of the recording tag is covalently attached to the cognate click chemistry coupling moiety (shown as a triple-line symbol) of the proximal protein. (C) After the target protein analyte is labeled with the recording tag, the attached binding entity is removed by digesting uracil (U) using a uracil-specific cleavage reagent (e.g., USER™).(D) Target protein analytes labeled with DNA recording tags are immobilized on a substrate surface using a suitable bioconjugate chemistry, such as click chemistry (e.g., alkyne-azide binding pair, methyltetrazine (mTET)-trans-cyclooctene (TCO) binding pair, etc.). In certain embodiments, the entire target protein-recording tag labeling assay is performed in a single tube containing multiple different target protein analytes using a pool of binding agents and a pool of recording tags. After target labeling of protein analytes within a sample with recording tags containing sample barcodes (BCSs), multiple protein analyte samples may be pooled prior to the immobilization step in (D). Thus, in certain embodiments, up to thousands of protein analytes across hundreds of samples can be labeled and immobilized in a single-tube next-generation protein assay (NGPA), significantly saving expensive affinity reagents (e.g., antibodies).
[0220] [Figure 29]Figures 29A-29E show the conjugation of DNA recording tags to polypeptides. (A) A modified polypeptide is labeled with a bifunctional click chemistry reagent, such as an alkyne-NHS ester (acetylene-PEG-NHS ester) reagent or an alkyne-benzophenone, to generate an alkyne-labeled (three-line symbol) polypeptide. The alkyne may also be a strained alkyne, such as a cyclooctyne, including dibenzocyclooctyl (DBCO). (B) An example of a DNA recording tag design is shown, which is chemically coupled to an alkyne-labeled polypeptide. The recording tag contains a universal priming sequence (P1), a barcode (BC), and a spacer sequence (Sp). The recording tag is labeled with an mTet moiety for coupling to the substrate surface and an azide moiety for coupling to the alkyne moiety of the labeled polypeptide. (C) A modified alkyne-labeled protein or polypeptide is labeled with the recording tag via the alkyne and azide moieties. Optionally, the recording tag-labeled polypeptide can be further labeled with a compartment barcode, for example, by annealing to a complementary sequence attached to a compartment bead and primer extension (also called polymerase extension), or as shown in Figures 20H-20J. (D) A population of recording tag-labeled peptides is created by protease digestion of the recording tag-labeled polypeptide. In some embodiments, some peptides will not be labeled with any recording tag. In other embodiments, some peptides may have one or more recording tags attached. (E) The recording tag-labeled peptides are immobilized on a substrate surface using inverse electron demand Diels-Alder (iEDDA) click chemistry between a substrate surface functionalized with TCO groups and the mTet portion of the recording tag attached to the peptide. In certain embodiments, cleanup steps may be used between the different steps shown.The use of orthogonal click chemistry (e.g., azide-alkyne and mTet-TCO) allows for both click chemistry labeling of polypeptides with recording tags and click chemistry immobilization of recording tag-labeled peptides to substrate surfaces (see McKay et al., 2014, Chem. Biol. 21:1075-1101, incorporated by reference in its entirety).
[0221] [Figure 30]Figures 30A-30E show the writing of a sample barcode to a recording tag after initial DNA labeling of the polypeptide. (A) A denatured polypeptide is labeled with a bifunctional click chemistry reagent, such as an alkyne-NHS reagent or an alkyne-benzophenone, to generate an alkyne-labeled polypeptide. (B) After labeling the polypeptide with an alkyne (or alternatively, a click chemistry moiety), a DNA tag containing a universal priming sequence (P1) and labeled with an azide and mTet moiety is coupled to the polypeptide via an azide-alkyne interaction. It is understood that other click chemistry interactions may also be used. (C) A recording tag DNA construct containing sample barcode information (BCS') and other recording tag functional components (e.g., a universal priming sequence (P1'), a spacer sequence (Sp')) anneals to a DNA tag-labeled polypeptide via the complementary universal priming sequence (P1-P1'). The recording tag information is transferred to the DNA tag by polymerase extension. (D) A population of recording tag-labeled peptides is created by protease digestion of the recording tag-labeled polypeptide. (E) The recording tag-labeled peptide is immobilized on a substrate surface using inverse electron demand Diels-Alder (iEDDA) click chemistry between a surface functionalized with TCO groups and the mTet portion of the recording tag attached to the peptide. In certain embodiments, cleanup steps may be used between the different steps shown. The use of orthogonal click chemistry (e.g., azide-alkyne and mTet-TCO) allows both click chemistry labeling of polypeptides with recording tags and click chemistry immobilization of recording tag-labeled peptides to substrate surfaces (see McKay et al., 2014, Chem. Biol., 21:1075-1101, incorporated by reference in its entirety).
[0222] [Figure 31]Figures 31A-31E show bead compartmentalization for barcoding polypeptides. (A) A polypeptide is labeled in solution using standard bioconjugation or photoaffinity labeling techniques with heterobifunctional click chemistry reagents. Possible labeling sites include the ε-amine of a lysine residue (e.g., with an NHS-alkyne as shown) or the carbon backbone of a peptide (e.g., with a benzophenone-alkyne). (B) An azide-labeled DNA tag containing a universal priming sequence (P1) is coupled to the alkyne moiety of the labeled polypeptide. (C) The DNA-tagged polypeptide is annealed to a DNA recording tag-labeled bead via complementary DNA sequences (P1 and P1'). The DNA recording tag on the bead contains a spacer sequence (Sp'), a compartment barcode sequence (BCP'), an optional unique molecular identifier (UMI), and a universal sequence (P1'). The DNA recording tag information is transferred to the DNA tag of the polypeptide by polymerase extension (alternatively, ligation can be used). After information transfer, the resulting polypeptide contains multiple recording tags containing several functional elements, including compartment barcodes. (D) Protease digestion of the recording tag-labeled polypeptides creates a population of recording tag-labeled peptides. The recording tag-labeled peptides are released from the beads and (E) re-immobilized to a sequencing substrate (e.g., using iEDDA click chemistry between the mTet and TCO moieties, as shown).
[0223] [Figure 32-1]Figures 32A-32H show an example workflow for next-generation protein assays (NGPAs). Protein samples are labeled with DNA recording tags, which consist of several functional units, such as a universal priming sequence (P1), a barcode sequence (BC), an optional UMI sequence, and a spacer sequence (Sp) (which allows information transfer with the binding substance coding tag). (A) The labeled protein is immobilized (passively or covalently) on a substrate (e.g., beads, porous beads, or a porous matrix). (B) The substrate is blocked with the protein, and optionally, a competitor oligonucleotide (Sp') complementary to the spacer sequence is added to minimize nonspecific interactions of the analyte recording tag sequence. (C) An analyte-specific antibody (with an associated coding tag) is incubated with the substrate-bound protein. The coding tag may contain a uracil base for subsequent uracil-specific cleavage. (D) After antibody binding, excess competitor oligonucleotide (Sp'), if applicable, is washed away. The coding tag is transiently annealed to the recording tag via a complementary spacer sequence, and the coding tag information is transferred to the recording tag in a primer extension reaction to generate an extended recording tag. If the immobilized protein is denatured, the bound antibody and annealed coding tag can be removed under alkaline wash conditions, such as 0.1 N NaOH. If the immobilized protein is in its native conformation, milder conditions may be required to remove the bound antibody and coding tag. Examples of milder antibody removal conditions are outlined in panels E–H. (E) After information transfer from the coding tag to the recording tag, the coding tag is nicked (cut) at its uracil site using a uracil-specific excision reagent (e.g., USER™) enzyme mix. (F) The bound antibody is removed from the protein using high salt and low / high pH washes. The cleaved DNA coding tag, which remains attached to the antibody, is short and similarly rapidly eluted. The longer DNA coding tag fragment may or may not remain annealed to the recording tag.(G) A second binding cycle begins similarly to steps (B)-(D), with a second primer extension step in which primer extension transfers the coding tag information from the second antibody to the extended recording tag. (H) The result of the two binding cycles is a concatenation of the binding information from the first and second antibodies attached to the recording tag. [Figure 32-2] Same as above.
[0224] [Figure 33]Figures 33A-33D show a one-step next-generation protein assay (NGPA) using multiple binding agents and enzyme-mediated sequential information transfer. The NGPA assay uses an immobilized protein molecule to which two cognate binding agents (e.g., antibodies) are simultaneously bound. After multiple cognate antibody binding events, a combined primer extension and DNA nicking step is used to transfer information from the coding tag of the bound antibody to a recording tag. The caret symbol (^) on the coding tag represents a double-stranded DNA nicking endonuclease site. (A) In the illustrated example, the coding tag of an antibody bound to epitope 1 (Epi#1) of a protein transfers the coding tag information (e.g., encoder sequence) to the recording tag in a primer extension step after hybridization of a complementary spacer sequence. (B) Once double-stranded DNA is formed between the extended recording tag and the coding tag, the coding tag is cleaved using a nicking endonuclease that cleaves only one strand of DNA in a double-stranded DNA substrate, such as Nt.BsmAI, which is active at 37°C. After the nicking step, the duplex formed by the cleavable coding tag-binding agent and the extended recording tag becomes thermodynamically unstable and dissociates. The longer coding tag fragment may or may not remain annealed to the recording tag. (C) This allows the coding tag of an antibody bound to epitope #2 (Epi#2) of the protein to anneal to the extended recording tag via a complementary spacer sequence, allowing the extended recording tag to be further extended by primer extension, transferring information from the coding tag of the Epi#2 antibody to the extended recording tag. (D) Again, after double-stranded DNA is formed between the extended recording tag and the coding tag of the Epi#2 antibody, the coding tag is nicked with a nicking endonuclease such as Nb.BssSI. In certain embodiments, the use of a non-strand-displacing polymerase during primer extension (also called polymerase extension) is preferred. A non-strand-displacing polymerase prevents extension of the cleaved coding tag remainder, where more than a single base remains annealed to the recording tag.The processes (A)-(D) can be repeated spontaneously until all coding tags of proximal bound binding entities are "consumed" by hybridization, information transfer to extended recording tags, and nicking steps. The coding tags may contain an encoder sequence that is identical for all binding entities (e.g., antibodies) specific for a given analyte (e.g., cognate protein), may contain an epitope-specific encoder sequence, or may contain a unique molecular identifier (UMI) to distinguish between different molecular events.
[0225] [Figure 34] Figures 34A-34C show density control of recording tag-peptide immobilization using titration of reactive moieties on the substrate surface. (A) Peptide density on the substrate surface can be titrated by controlling the density of functional coupling moieties on the substrate surface. This can be achieved by derivatizing the substrate surface with an appropriate ratio of active coupling molecules to "dummy" coupling molecules. In the example shown, NHS-PEG-TCO reagent (active coupling molecule) is combined with NHS-mPEG (dummy molecule) in a defined ratio to derivatize the amine surface with TCO. Functionalized PEGs are available in a variety of molecular weights, from 300 to over 40,000. (B) A bifunctional 5' amine DNA recording tag (mTet is the other functional moiety) is coupled to the N-terminal Cys residue of the peptide using succinimidyl 4-(N-maleimidomethyl)cyclohexane-1 (SMCC) bifunctional crosslinker. The internal mTet-dT group of the recording tag is created from the azide-dT group using mtetrazine-azide. (C) The recording tag-labeled peptide is immobilized on the activated substrate surface of (A) using iEDDA click chemistry between mTet and TCO. The mTet-TCO iEDDA coupling reaction is very rapid, efficient, and stable (mTet-TCO is more stable than Tet-TCO).
[0226] [Figure 35]Figures 35A-35C show next-generation protein sequencing (NGPS) binding cycle-specific coding tags. (A) Design of an NGPS assay using cycle-specific N-terminal amino acid (NTAA) binding substance coding tags. An NTAA binding substance (e.g., an antibody specific for an N-terminal DNP-labeled tyrosine) binds to the DNP-labeled NTAA of a peptide attached to a recording tag containing a universal priming sequence (P1), a barcode (BC), and a spacer sequence (Sp). When the binding substance binds to the cognate NTAA of the peptide, the coding tag attached to the NTAA binding substance approaches the recording tag and anneals to it via the complementary spacer sequence. The coding tag information is transferred to the recording tag by polymerase extension. The coding tag may include a cycle-specific barcode to record which binding cycle the coding tag represents. In certain embodiments, the coding tags of binding substances that bind to analytes have the same encoder barcode, independent of the cycle number, combined with a unique binding cycle-specific barcode. In other embodiments, the coding tag of the binding substance for the analyte includes a unique encoder barcode for the analyte-binding cycle combination information. In either approach, a common spacer sequence can be used in the coding tag of the binding substance for each binding cycle. (B) In this example, the binding substance for each binding cycle has a short binding cycle-specific barcode to identify the binding cycle, thereby providing a unique combination barcode that identifies the specific binding substance binding cycle combination along with the encoder barcode that identifies the binding substance. (C) After the binding cycle is completed, the extended recording tag can be converted into an amplifiable library using a capping cycle step, in which, for example, a cap comprising a universal priming sequence P1' linked to a universal priming sequence P2 and a spacer sequence Sp' first anneals to the extended recording tag via the complementary P1 and P1' sequences, bringing the cap into close proximity with the extended recording tag.The complementary Sp and Sp' sequences of the extended recording tag and cap anneal, and primer extension adds a second universal primer sequence (P2) to the extended recording tag.
[0227] [Figure 36-1]Figures 36A-36E show a DNA-based model system for demonstrating information transfer from coding tags to recording tags. Exemplary binding and intramolecular writing were demonstrated using an oligonucleotide model system. The targeting molecules A' and B' of the coding tags were designed to hybridize with the target-binding regions A and B of the recording tags. A recording tag (RT) mix was prepared by pooling equal concentrations of two recording tags, saRT_Abc_v2 (target A) and saRT_Bbc_V2 (target B). The recording tags are biotinylated at the 5' end and contain a unique target-binding region, a universal forward primer sequence, a unique DNA barcode, and an 8-base common spacer sequence (Sp). The coding tags contain a unique encoder barcode base flanked by an 8-base common spacer sequence (Sp'), one of which is covalently linked to the A or B target via a polyethylene glycol linker. (A) Biotinylated recording tag oligonucleotides (saRT_Abc_v2 and saRT_Bbc_V2) were immobilized on streptavidin beads along with biotinylated dummy T10 oligonucleotides. Recording tags were designed with either the A or B capture sequence (recognized by cognate binders A' and B', respectively) and corresponding barcodes (rtA_BC and rtB_BC) to identify the binding target. All barcodes in this model system were selected from a set of 65 15-mer barcodes (SEQ ID NOs: 1-65). In some cases, 15-mer barcodes were combined to construct longer barcodes to facilitate gel analysis. Specifically, rtA_BC = BC_1 + BC_2; rtB_BC = BC_3. Two coding tags for the cognate binders, namely CT_A'-bc (encoder barcode = BC_5) and CT_B'-bc (encoder barcode = BC_5 + BC_6), were also synthesized, based on the A and B sequences of the recording tag. Optionally, blocking oligos (DupCT_A'BC and DupCT_AB'BC) complementary to portions of the coding tag sequence (leaving behind a single-stranded Sp' sequence) were pre-annealed to the coding tag before it was annealed to the recording tag immobilized on the bead.Strand-displacing polymerases remove the blocking oligos during polymerase extension. The barcode legend (inset) shows the assignment of 15-mer barcodes to functional barcodes for recording tags and coding tags. (B) The recording tag barcode design and coding tag encoder barcode design provide easy gel analysis of "intra- versus intermolecular" interactions between recording tags and coding tags. With this design, undesired "intermolecular" interactions (between the A recording tag and the B' coding tag, and between the B recording tag and the A' coding tag) generate gel products that are either 15 bases longer or shorter than the desired "intra-molecular" (between the A recording tag and the A' coding tag; between the B recording tag and the B' coding tag) interaction products. In the primer extension step, the A' and B' coding tag barcodes (ctA'_BC, ctB'_BC) are modified to their reverse complement barcodes (ctA'_BC and ctB'_BC). (C) Primer extension assays demonstrated that information was transferred from the coding tag to the recording tag, and that the adapter sequence was attached to the annealed EndCap oligo by primer extension for PCR analysis. (D) Optimization of "intramolecular" information transfer by titrating the surface density of the recording tag using dummy T20 oligos. Biotinylated recording tag oligos were mixed with biotinylated dummy T20 oligos at various ratios ranging from 1:0 to 1:10 to 1:10,000. At reduced recording tag densities (1:103 and 1:104), "intramolecular" interactions prevail over "intermolecular" interactions. (F) As a simple extension of the DNA model system, a simple protein binding system involving the Nano-Tag15 peptide-streptavidin binding pair has been demonstrated (KD ∼4 nM) (Perbandt et al., 2007, Proteins 67:1147-1153), although any number of peptide-binding agent model systems can be used. The Nano-Tag15 peptide sequence is (fM)DVEAWLGARVPLVET (SEQ ID NO: 131) (fM = formyl-Met). The Nano-Tag15 peptide further contains a short, flexible linker peptide (GGGGS) and a cysteine residue for coupling to the DNA recording tag.Other exemplary peptide tag-cognate binder pairs include calmodulin-binding peptide (CBP)-calmodulin (KD ∼2 pM) (Mukherjee et al., 2015, J. Mol. Biol., 427:2707-2725), amyloid beta (Aβ16-27) peptide-US7 / Lcn2 anticalin (0.2 nM) (Rauth et al., 2016, Biochem. J., 473:1563-1578), PA tag / NZ-1 antibody (KD ∼400 pM), FLAG-M2 Ab (28 nM), HA-4B2 Ab (1.6 nM), and Myc-9E10 Ab (2.2 nM) (Fujii et al., 2014, Protein Expr. Purif., 95:240-247). (E) To test intramolecular information transfer from the binding agent coding tag to the recording tag by primer extension, an oligonucleotide "binding agent" that binds to the complementary DNA sequence "A" can be used for testing and development. This hybridization event inherently exhibits an affinity greater than fM. Streptavidin may be used as a test binding agent for the Nano-tag15 peptide epitope. The peptide tag-binding agent interaction is high affinity but can be easily disrupted by acidic and / or high salt washes (Perbandt et al., supra). [Figure 36-2] Same as above. [Figure 36-3] Same as above.
[0228] [Figure 37]Figures 37A-37B illustrate the use of nano- or micro-emulsion PCR to transfer information from a UMI-labeled N- or C-terminus of a peptide to a DNA-tagged counterpart. (A) The N- or C-terminus of a polypeptide is labeled with a nucleic acid molecule containing a unique molecular identifier (UMI). The UMI may be flanked by sequences used to prime subsequent PCR. Internal portions of the polypeptide are then "body-tagged" with separate DNA tags containing sequences complementary to the priming sequences flanking the UMI. (B) The resulting labeled polypeptides are emulsified, and emulsion PCR (ePCR) (alternatively, emulsion in vitro transcription-RT-PCR (IVT-RT-PCR) reactions or other suitable amplification reactions can be performed) is performed to amplify the N- or C-terminal UMI. Microemulsions or nanoemulsions are formed with an average droplet diameter of 50-1000 nm, with an average of less than one polypeptide per droplet. Pre- and post-PCR droplet contents are shown in the left and right panels, respectively. The UMI amplicon is hybridized to the internal polypeptide body DNA via a complementary priming sequence, and the UMI information is transferred from the amplicon to the internal polypeptide body DNA tag by primer extension.
[0229] [Figure 38]Figure 38 illustrates single-cell proteomics. Cells are encapsulated in droplets containing polymer-forming subunits (e.g., acrylamide) and lysed. The polymer-forming subunits are polymerized (e.g., polyacrylamide), and the proteins are crosslinked to the polymer matrix. The emulsion droplets are disrupted, releasing polymerized gel beads containing single-cell protein lysate attached to a permeable polymer matrix. Proteins are crosslinked to the polymer matrix either in their native conformation or in a denatured state by including a denaturant such as urea in the lysis and encapsulation buffer. Recording tags containing compartment barcodes and other recording tag components (e.g., universal priming sequence (P1), spacer sequence (Sp), optional unique molecular identifier (UMI)) are attached to the proteins using several methods known in the art and described herein, including emulsification or combinatorial indexing with barcoded beads. Alternatively, polymerized gel beads containing single-cell proteins can be subjected to proteinase digestion after recording tag addition to generate recording tag-labeled peptides suitable for peptide sequencing. In certain embodiments, the polymer matrix can be designed to dissolve in a suitable additive, such as a disulfide-crosslinked polymer that is destroyed upon exposure to a reducing agent such as tris(2-carboxyethyl)phosphine (TCEP) or dithiothreitol (DTT).
[0230] [Figure 39]Figures 39A-39E show the enhancement of amino acid cleavage reactions using bifunctional N-terminal amino acid (NTAA) modifiers and chimeric cleavage reagents. (A) and (B) Peptides attached to a solid substrate are modified with a bifunctional NTAA modifier, such as biotin-phenylisothiocyanate (PITC). (C) Low-affinity Edmanase (>μM Kd) is mobilized to biotin-PITC-labeled NTAA using streptavidin-Edmanase chimeric proteins. (D) The efficiency of Edmanase cleavage is significantly improved due to the increased effective local concentration as a result of the biotin-streptavidin interaction. (E) The cleaved biotin-PITC-labeled NTAA and accompanying streptavidin-Edmanase chimeric proteins diffuse far away after cleavage. Several other bioconjugation mobilization strategies can also be used. Azide-modified PITC is commercially available (4-azidophenylisothiocyanate, Sigma), allowing several simple transformations of azide-PITC into other bioconjugates of PITC, such as biotin-PITC via click chemistry reaction with alkyne-biotin.
[0231] [Figure 40-1]Figures 40A-40I show the generation of C-terminal recording tag-labeled peptides from protein lysates (optionally encapsulated in gel beads). (A) Denatured polypeptides are reacted with an acid anhydride to label lysine residues. In one embodiment, a mix of alkyne (mTet)-substituted citraconic anhydride and propionic anhydride is used to label lysines with mTet (shown as striped rectangles). (B) The result is an alkyne (mTet)-labeled polypeptide in which some lysines are blocked with propionic groups (shown as squares in the polypeptide chain). The alkyne (mTet) moiety is useful for click chemistry-based DNA labeling. (C) DNA tags (shown as solid rectangles) are attached to the alkyne or mTet moiety using azide or trans-cyclooctene (TCO) labels, respectively, via click chemistry. (D) Barcodes and functional elements, such as spacer (Sp) and universal priming sequences, are added to the DNA tags using a primer extension step as shown in Figure 31 to produce recording tag-labeled polypeptides. The barcode may be a sample barcode, a partition barcode, a compartment barcode, a spatial location barcode, or any combination thereof. (E) The resulting recording tag-labeled polypeptide is protease- or chemically fragmented into recording tag-labeled peptides. (F) For illustrative purposes, a peptide fragment labeled with two recording tags is shown. (G) A DNA tag containing a universal priming sequence complementary to the universal priming sequence of the recording tag is ligated to the C-terminus of the peptide. The C-terminal DNA tag also contains a moiety for conjugating the peptide to a surface. (H) The complementary universal priming sequence of the C-terminal DNA tag and the stochastically selected recording tag are annealed. An intramolecular primer extension reaction is used to transfer information from the recording tag to the C-terminal DNA tag. (I) The internal recording tag of the peptide is coupled to a lysine residue via maleic anhydride. This coupling is reversible at acidic pH. The internal recording tag is cleaved from the lysine residue of the peptide at acidic pH, leaving the C-terminal recording tag intact.Optionally, the newly exposed lysine residues may be bridged with a non-hydrolyzable anhydride such as propionic anhydride. [Figure 40-2] Same as above.
[0232] [Figure 41] Figure 41 shows the workflow of a preferred embodiment of the NGPS assay.
[0233] [Figure 42]Figures 42A-42D show exemplary steps in an NGPS sequencing assay. The N-terminal amino acid (NTAA) acetylation or amidination step of a peptide bound to a surface labeled with a recording tag can occur before or after binding by an NTAA-binding substance, depending on whether the NTAA-binding substance is engineered to bind to acetylated NTAA or native NTAA. In the first case, (A) the NTAA of the peptide is first acetylated chemically using acetic anhydride or enzymatically using an N-terminal acetyltransferase (NAT). (B) The NTAA is recognized by an NTAA-binding substance, such as an engineered anticalin, aminoacyl-tRNA synthetase (aaRS), or ClpS. A DNA coding tag is attached to the binding substance and contains a barcode encoder sequence that identifies the specific NTAA-binding substance. (C) After the acetylated NTAA binds to the NTAA-binding substance, the DNA coding tag temporarily anneals to the recording tag via complementary sequences, and the coding tag information is transferred to the recording tag by polymerase extension. In an alternative embodiment, the recording tag information is transferred to the coding tag by polymerase extension. (D) The acetylated NTAA is cleaved from the peptide by an engineered acylpeptide hydrolase (APH), which catalyzes the hydrolysis of the terminal acetylated amino acid of the acetylated peptide. After cleavage of the acetylated NTAA, this cycle spontaneously repeats, starting with acetylation of the newly exposed NTAA. While N-terminal acetylation is used as an exemplary mode of NTAA modification / cleavage, other N-terminal moieties, such as guanyl moieties, may alternatively be substituted by modifying the cleavage chemistry accordingly. When guanidination is used, guanylated NTAA can be cleaved under mild conditions using a 0.5-2% NaOH solution (see Hamada, 2016, incorporated by reference in its entirety). APH is a serine peptidase that can catalyze the removal of Nα-acetylated amino acids of blocked peptides and belongs to the prolyl oligopeptidase (POP) family (clan SC, family S9).APH is a key regulator of N-terminally acetylated proteins in eukaryotic, bacterial, and archaeal cells.
[0234] [Figure 43] Figures 43A-43B show exemplary recording tag-coding tag design features. (A) Structure of an exemplary recording tag-associated protein (or peptide) and a bound binding entity (e.g., anticalin) with an associated coding tag. A thymidine (T) base is inserted between the spacer (Sp') and barcode (BC') sequences of the coding tag to accommodate stochastic non-templated 3'-terminal adenosine (A) addition during primer extension reactions. (B) The DNA coding tag is attached to the binding entity (e.g., anticalin) via a SpyCatcher-SpyTag protein-peptide interaction.
[0235] [Figure 44] Figures 44A-44E show (A) and (B) the enhancement of the NTAA cleavage reaction using hybridization of a cleaving agent to a recording tag. The NTAA of a recording tag-labeled peptide attached to a solid-phase substrate (e.g., beads) is modified or labeled (Mod) with, for example, PITC, DNP, SNP, acetyl modifiers, or guanidination. (C) A cleavage enzyme (e.g., acylpeptide hydrolase (APH), aminopeptidase (AP), edmanase, etc.) is attached to a DNA tag containing a universal priming sequence complementary to the universal priming sequence of the recording tag. The cleavage enzyme is recruited to the modified NTAA by hybridization between the DNA tag of the cleavage enzyme and the complementary universal priming sequence of the recording tag. (D) This hybridization step significantly increases the effective affinity of the cleavage enzyme for NTAA. (E) The cleaved NTAA diffuses away, and the accompanying cleavage enzyme can be removed by peeling off the hybridized DNA tag.
[0236] [Figure 45]Figure 45 shows cyclic resolution peptide sequencing using peptide ligase, protease, and diaminopeptidase. Butelase I ligates the TEV-Butelase I peptide substrate (TENLYFQNHV, SEQ ID NO: 132) to the NTAA of the query peptide. Butelase requires an NHV motif at the C-terminus of the peptide substrate. After ligation, tobacco etch virus (TEV) protease is used to cleave the chimeric peptide substrate after the glutamine (Q) residue, resulting in a chimeric peptide with an asparagine (N) residue attached to the N-terminus of the query peptide. Diaminopeptidase (DAP) or dipeptidyl peptidase cleaves two amino acid residues from the N-terminus, shortening the N-tagged query peptide by two amino acids, effectively removing the asparagine (N) residue and the original NTAA of the query peptide. The newly exposed NTAAs are read using a binding agent such as those provided herein, and the entire cycle is then repeated "n" times to sequence "n" amino acids. Control of DAP processivity can be achieved by using a streptavidin-DAP metalloenzyme chimeric protein and by tethering a biotin moiety to the N-terminal asparagine residue. DETAILED DESCRIPTION OF THE INVENTION
[0237] Terms not specifically defined herein should be given the meaning that they would be given by one of ordinary skill in the art in light of the present disclosure and the context, but as used herein, unless specified to the contrary, the terms have the indicated meaning.
[0238] I. Introduction The present disclosure provides, in part, a highly parallel, high-throughput digital macromolecule characterization and quantification method directly applicable to protein and peptide characterization and sequencing (see Figures 1B and 2A). The methods described herein use binding agents comprising coding tags carrying identifying information in the form of nucleic acid molecules or sequenceable polymers, where the binding agents interact with a macromolecule of interest. Multiple successive binding cycles are performed, each cycle involving exposing multiple macromolecules immobilized on a solid support, preferably representing a pooled sample, to multiple binding agents. During each binding cycle, the identity of each binding agent that binds to the macromolecule, and optionally the number of binding cycles, are recorded by transferring the information from the binding agent coding tag to a recording tag co-localized with the macromolecule. In an alternative embodiment, information from the recording tag, including identifying information about the associated macromolecule, can be transferred to the binding agent's coding tag (e.g., to form an extended coding tag) or to a third "ditag" construct. Multiple cycles of binding events build historical binding information about the recording tags that co-localize with the macromolecule, resulting in an extended recording tag containing multiple coding tags in a collinear order that represents the temporal binding history for a given macromolecule. Furthermore, cycle-specific coding tags can be used to track information from each cycle, so that if a cycle is skipped for some reason, the extended recording tag continues to collect information in subsequent cycles, allowing the cycle that lacked information to be identified.
[0239] Alternatively, instead of writing or transferring information from the coding tag to the recording tag, information can be transferred from the recording tag containing identifying information about the associated macromolecule to a coding tag or third ditag construct that forms an extended coding tag. The resulting extended coding tag or ditag can be collected after each binding cycle for subsequent sequence analysis. Using the identifying information on the recording tag, including barcodes (e.g., partition tags, compartment tags, sample tags, fraction tags, UMIs, or any combination thereof), the extended coding tag or ditag sequence reads can be mapped back to the original macromolecule. In this way, a nucleic acid code library representation of the macromolecule's binding history is generated. This nucleic acid code library can be amplified and analyzed using very high-throughput next-generation digital sequencing methods, thereby analyzing millions to billions of molecules per run. The creation of a nucleic acid code library of binding information is uniquely useful in that it allows for enrichment, subtraction, and normalization using DNA-based techniques that use hybridization. These DNA-based methods are easily and rapidly scalable and customizable, and are more cost-effective than those available for direct manipulation of other types of macromolecular libraries, such as protein libraries. Thus, nucleic acid-encoded libraries of binding information can be processed by one or more techniques prior to sequencing to enrich and / or subtract and / or normalize the sequence representation. This allows for much more efficient, rapid, and cost-effective extraction of the most desired information from very large libraries, where the abundance of individual members can initially vary over many orders of magnitude. Importantly, these nucleic acid-based techniques for manipulating library representation are orthogonal to, and can be used in conjunction with, more conventional methods. For example, common, highly abundant proteins, such as albumin, can be subtracted using protein-based methods that can remove most, but not all, of the undesired proteins.Subsequently, albumin-specific members of the extended record tag library can also be subtracted, thus achieving a more thorough global subtraction.
[0240] In one aspect, the present disclosure provides a highly parallelized method for peptide sequencing using the Edman-like degradation technique, which enables sequencing from large populations (e.g., millions to billions) of peptides labeled with DNA recording tags. These recording tag-labeled peptides are derived from proteolytic digestion or limited hydrolysis of a protein sample, and the recording tag-labeled peptides are randomly immobilized on a sequencing substrate (e.g., porous beads) with appropriate intermolecular spacing on the substrate. Modification of the N-terminal amino acid (NTAA) residue of a peptide with small chemical moieties, such as phenylthiocarbamoyl (PTC), dinitrophenol (DNP), sulfonylnitrophenol (SNP), dansyl, 7-methoxycoumarin, acetyl, or guanidinyl, that catalyze or mobilize the NTAA cleavage reaction allows for periodic control of the Edman-like degradation process. The modifying chemical moiety can also result in enhanced binding affinity for cognate NTAA-binding substances. The modified NTAA of each immobilized peptide is identified by binding of a cognate NTAA-binding substance containing a coding tag and transfer (e.g., primer extension or ligation) of the coding tag information (e.g., an encoder sequence that provides identifying information about the binding substance) from the coding tag to the peptide's recording tag. The modified NTAA is then removed by chemical or enzymatic means. In certain embodiments, an enzyme (e.g., edmanase) is engineered to catalyze the removal of the modified NTAA. In other embodiments, naturally occurring exopeptidases, such as aminopeptidases or acylpeptide hydrolases, can be engineered to cleave the terminal amino acid only in the presence of the appropriate chemical modification.
[0241] II. Definition In the following description, certain specific details are set forth to provide a thorough understanding of various embodiments. However, it will be understood by those skilled in the art that the compounds can be made and used without these details. In other instances, well-known structures have not been shown or described in detail to avoid unnecessarily obscuring the description of the embodiments. Unless the context requires otherwise, throughout this specification and the claims which follow, the word "comprise" and variations thereof, such as "comprises" and "comprising," should be interpreted in an open-ended, inclusive sense, i.e., "including, but not limited to." Furthermore, the term "comprising" (and related terms such as "comprise" or "comprises" or "having" or "including") does not exclude that in certain other embodiments, an embodiment, such as, for example, any composition of matter, composition, method, or process described herein, may "consist of" or "consist essentially of" the described features. The headings provided herein are merely for convenience and are not to be interpreted as dictating the scope or meaning of the claimed embodiments.
[0242] Throughout this specification, a reference to "one embodiment" or "an embodiment" means that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment. Thus, the appearances of the phrase "in one embodiment" or "in an embodiment" in various places throughout this specification are not necessarily all referring to the same embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0243] As used herein, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a peptide" includes one or more peptides, or mixtures of peptides. Also, unless otherwise specified or apparent from the context, as used herein, the term "or" is understood to be inclusive and encompasses both "or" and "and."
[0244] As used herein, the term "macromolecule" encompasses large molecules composed of smaller subunits. Examples of macromolecules include, but are not limited to, peptides, polypeptides, proteins, nucleic acids, carbohydrates, lipids, and macrocycles. Macromolecules also include chimeric macromolecules (e.g., peptides linked to nucleic acids) that are composed of combinations of two or more types of macromolecules linked by covalent bonds. Macromolecules can also include "macromolecular assemblies" that are composed of noncovalent complexes of two or more macromolecules. Macromolecular assemblies can be composed of the same type of macromolecule (e.g., protein-protein) or two or more different types of macromolecules (e.g., protein-DNA).
[0245] As used herein, the term "peptide" encompasses peptides, polypeptides, and proteins and refers to a molecule comprising a chain of two or more amino acids joined by peptide bonds. Generally speaking, peptides with more than 20-30 amino acids are commonly referred to as polypeptides, and peptides with more than 50 amino acids are commonly referred to as proteins. The amino acids of a peptide are most typically L-amino acids, but may also be D-amino acids, modified amino acids, amino acid analogs, amino acid mimetics, or any combination thereof. Peptides may be naturally occurring, synthetically produced, or recombinantly expressed. Peptides may also contain additional groups that modify the amino acid chain, such as functional groups added by post-translational modification.
[0246] As used herein, the term "amino acid" refers to an organic compound having an amine group, a carboxylic acid group, and a side chain specific to each amino acid that functions as a monomeric subunit of a peptide. Amino acids include the 20 standard naturally occurring or canonical amino acids as well as non-standard amino acids. Standard naturally occurring amino acids include alanine (A or Ala), cysteine (C or Cys), aspartic acid (D or Asp), glutamic acid (E or Glu), phenylalanine (F or Phe), glycine (G or Gly), histidine (H or His), isoleucine (I or Ile), lysine (K or Lys), leucine (L or Leu), methionine (M or Met), asparagine (N or Asn), proline (P or Pro), glutamine (Q or Gln), arginine (R or Arg), serine (S or Ser), threonine (T or Thr), valine (V or Val), tryptophan (W or Trp), and tyrosine (Y or Tyr). Amino acids may be L- or D-amino acids. Non-standard amino acids can be naturally occurring or chemically synthesized modified amino acids, amino acid analogs, amino acid mimetics, non-standard proteinogenic amino acids, or non-proteinogenic amino acids. Examples of non-standard amino acids include, but are not limited to, selenocysteine, pyrrolysine, and N-formylmethionine, β-amino acids, homoamino acids, proline and pyruvate derivatives, 3-substituted alanine derivatives, glycine derivatives, ring-substituted phenylalanine and tyrosine derivatives, linear core amino acids, and N-methyl amino acids.
[0247] As used herein, the term "post-translational modification" refers to a modification that occurs on a peptide after translation of the peptide by the ribosome is complete. A post-translational modification can be a covalent modification or an enzymatic modification. Examples of post-translational modifications include, but are not limited to, acylation, acetylation, alkylation (including methylation), biotinylation, butyrylation, carbamylation, carbonylation, deamidation, deimination, diphthamide formation, disulfide bridge formation, eliminylation, flavin attachment, formylation, gamma-carboxylation, glutamylation, glycylation, glycosylation, glycosylphosphatidylinositol glypiation, heme Modifications include C-attachment, hydroxylation, hypusine formation, iodination, isoprenylation, lipidation, lipoylation, malonylation, methylation, myristoylation, oxidation, palmitoylation, pegylation, phosphopantetheinylation, phosphorylation, prenylation, propionylation, retinylidene Schiff base formation, S-glutathionylation, S-nitrosylation, S-sulfenylation, selenation, succinylation, sulfination, ubiquitination, and C-terminal amidation. Post-translational modifications include modifications of the amino and / or carboxyl termini of peptides. Modifications of the terminal amino group include, but are not limited to, desamino, N-lower alkyl, N-di-lower alkyl, and N-acyl modifications. Modifications of the terminal carboxy group include, but are not limited to, amide, lower alkyl amide, dialkyl amide, and lower alkyl ester modifications (e.g., lower alkyl is C1-C4 alkyl). Post-translational modifications also include, for example, but are not limited to, modifications of amino acids between the amino and carboxy termini, such as those described above. The term post-translational modifications can also include peptide modifications that include one or more detectable labels.
[0248] As used herein, the term "binding agent" refers to a nucleic acid molecule, peptide, polypeptide, protein, carbohydrate, or small molecule that binds to, associates with, associates with, recognizes, or combines with a macromolecule or a component or feature of a macromolecule. A binding agent can form a covalent or noncovalent bond with a macromolecule or a component or feature of a macromolecule. A binding agent can also be a chimeric binding agent composed of two or more types of molecules, such as a nucleic acid molecule-peptide chimeric binding agent or a carbohydrate-peptide chimeric binding agent. A binding agent can be a naturally occurring molecule, a synthetically produced molecule, or a recombinantly expressed molecule. A binding agent can bind to a single monomer or subunit of a macromolecule (e.g., a single amino acid of a peptide) or to multiple linked subunits of a macromolecule (e.g., a di-peptide, tri-peptide, or higher-order peptide of a longer peptide, polypeptide, or protein molecule). A binding agent may bind to a linear molecule or a molecule having a three-dimensional structure (also referred to as a conformation). For example, an antibody binding agent may bind to a linear peptide, polypeptide, or protein, or to a conformational peptide, polypeptide, or protein. A binding agent may bind to the N-terminal peptide, C-terminal peptide, or intervening peptide of a peptide, polypeptide, or protein molecule. A binding agent may bind to the N-terminal amino acid, C-terminal amino acid, or intervening amino acid of a peptide molecule. A binding agent may preferably bind preferentially to chemically modified or labeled amino acids over unmodified or unlabeled amino acids. For example, a binding agent may preferably bind preferentially to amino acids modified with acetyl moieties, guanyl moieties, dansyl moieties, PTC moieties, DNP moieties, SNP moieties, etc. over amino acids without such moieties. A binding agent may bind to post-translational modifications of a peptide molecule.A binding agent may exhibit selective binding to a macromolecular component or feature (e.g., a binding agent may selectively bind to one of the 20 possible naturally occurring amino acid residues and bind with very low affinity or not at all to the other 19 naturally occurring amino acid residues). A binding agent may exhibit less selective binding if it is capable of binding to multiple macromolecular components or features (e.g., a binding agent may bind with similar affinity to two or more different amino acid residues). A binding agent comprises a coding tag, which is joined to the binding agent by a linker.
[0249] As used herein, the term "linker" refers to one or more of a nucleotide, nucleotide analog, amino acid, peptide, polypeptide, or non-nucleotide chemical moiety used to join two molecules. Linkers can be used to join a binding agent and a coding tag, a recording tag and a macromolecule (e.g., a peptide), a macromolecule and a solid support, a recording tag and a solid support, etc. In certain embodiments, the linker joins two molecules via an enzymatic or chemical reaction (e.g., click chemistry).
[0250] As used herein, the term "proteomics" refers to the quantitative analysis of the proteome within cells, tissues, and body fluids, as well as the spatial distribution of the proteome within the corresponding cells and tissues. Furthermore, proteomic studies encompass the dynamic state of the proteome, which changes continuously over time in response to biology and defined biological or chemical stimuli.
[0251] As used herein, the term "non-cognate binder" refers to a binding substance that is unable to bind or binds with low affinity to a macromolecular feature, component, or subunit being interrogated in a particular binding cycle reaction, compared to a "cognate binder" that binds with high affinity to the corresponding macromolecular feature, component, or subunit. For example, when tyrosine residues of peptide molecules are interrogated in a binding reaction, a non-cognate binder is one that binds to tyrosine residues with low affinity or not at all, and therefore does not efficiently transfer coding tag information to a recording tag under conditions suitable for transferring coding tag information from a cognate binder to a recording tag. Alternatively, when tyrosine residues of peptide molecules are interrogated in a binding reaction, a non-cognate binder is one that binds to tyrosine residues with low affinity or not at all, and therefore does not efficiently transfer recording tag information to a coding tag under conditions suitable for embodiments involving extended coding tags rather than extended recording tags.
[0252] The terminal amino acid at one end of a peptide chain, having a free amino group, is referred to herein as the "N-terminal amino acid" (NTAA). The terminal amino acid at the other end of the chain, having a free carboxyl group, is referred to herein as the "C-terminal amino acid" (CTAA). The amino acids that make up a peptide can be numbered sequentially, making the peptide "n" amino acids long. As used herein, an NTAA is considered to be the nth amino acid (also referred to herein as "n NTAA"). Using this nomenclature, the next amino acid down the length of the peptide from the N-terminus to the C-terminus is the n-1 amino acid, then the n-2 amino acid, and so on. In certain embodiments, the NTAA, the CTAA, or both can be modified or labeled with a chemical moiety.
[0253] As used herein, the term "barcode" refers to a nucleic acid molecule of about 2 to about 30 bases (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 bases) that provides a unique identifier tag or provenance information for a macromolecule (e.g., a protein, polypeptide, peptide), a binding agent, a set of binding agents from a binding cycle, a sample macromolecule, a set of samples, a macromolecule in a compartment (e.g., a droplet, bead, or separate location), a macromolecule in a set of compartments, a fraction of a macromolecule, a set of macromolecule fractions, a spatial region or set of spatial regions, a library of macromolecules, or a library of binding agents. Barcodes can be artificial or naturally occurring sequences. In certain embodiments, each barcode in a population of barcodes is different. In other embodiments, a portion of the barcodes in a population of barcodes are different, e.g., at least about 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, or 99% of the barcodes in the population of barcodes are different. The population of barcodes can be randomly or non-randomly generated. In certain embodiments, the population of barcodes is an error-correcting barcode. Barcodes can be used to computationally deconvolute multiplexed sequencing data and identify sequence reads originating from individual macromolecules, samples, libraries, etc. Barcodes can also be used to deconvolute collections of macromolecules distributed in small compartments to enhance mapping. For example, rather than mapping peptides back to a proteome, peptides can be mapped back to the protein molecules or protein complexes from which they originated.
[0254] A "sample barcode," also called a "sample tag," identifies the sample from which a macromolecule originates.
[0255] "Spatial barcode" identifies which region of a 2D or 3D tissue section a macromolecule originates from. Spatial barcode can be used for molecular pathology on tissue sections. Spatial barcode allows multiplex sequencing of multiple samples or libraries from tissue sections (multiple).
[0256] As used herein, the term "coding tag" refers to a nucleic acid molecule of about 2 to about 100 bases, including 2 and 100, and all integers therebetween, that contains identifying information about an associated binding entity. A "coding tag" may be made of a "sequenceable polymer" (see, e.g., Niu et al., 2013, Nat. Chem., 5:282-292; Roy et al., 2015, Nat. Commun., 6:7237; Lutz, 2015, Macromolecules, 48:4759-4767, each of which is incorporated by reference in its entirety). A coding tag optionally includes an encoder sequence flanked by a spacer on one side or by spacers on both sides. A coding tag may also be comprised of an optional UMI and / or an optional binding cycle-specific barcode. A coding tag may be single-stranded or double-stranded. The double-stranded coding tag may comprise blunt ends, overhanging ends, or both. The coding tag may refer to a binding substance, a complementary sequence hybridized with a coding tag directly attached to the binding substance (e.g., for double-stranded coding tags), or a coding tag directly attached to the coding tag information present in the extended recording tag. In certain embodiments, the coding tag may further comprise a binding cycle-specific spacer or barcode, a unique molecular identifier, a universal priming site, or any combination thereof.
[0257] As used herein, the term "encoder sequence" or "encoder barcode" refers to a nucleic acid molecule about 2 bases to about 30 bases in length (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 bases) that provides identifying information about the binding agent to which it is associated. An encoder sequence is capable of uniquely identifying the binding agent to which it is associated. In certain embodiments, the encoder sequence provides identifying information about the binding agent to which it is associated and the binding cycle in which the binding agent is used. In other embodiments, the encoder sequence is combined with another binding cycle-specific barcode in the coding tag. Alternatively, the encoder sequence can distinguish the binding agent to which it is associated as belonging to a set of two or more different binding agents. In some embodiments, this level of identification is sufficient for analysis. For example, in some embodiments involving binding agents that bind to amino acids, rather than definitively identifying the amino acid residue at a particular position in the peptide, it may be sufficient to know that the peptide contains one of two possible amino acids at that position. Another example uses a common encoder sequence for polyclonal antibodies that recognize more than one epitope of a protein target and include a mixture of antibodies with different specificities. In other embodiments, where an encoder sequence identifies a set of potential binding agents, a sequential decoding approach can be used to provide a unique identification of each binding agent. This is achieved by varying the encoder sequence for a given binding agent over repeated cycles of binding (see Gunderson et al., 2004, Genome Res., 14:870-7).The partially distinguishing coding tag information from each binding cycle, when combined with coding information from other cycles, provides a unique identifier for the binding agent, e.g., it is the particular combination of coding tags, rather than the individual coding tags (or encoder sequences), that provides unique identifying information for the binding agent. Preferably, the encoder sequences within a library of binding agents have the same or similar number of bases.
[0258] As used herein, the terms "binding cycle-specific tag," "binding cycle-specific barcode," or "binding cycle-specific sequence" refer to a unique sequence used to identify a library of binding agents used within a particular binding cycle. A binding cycle-specific tag can comprise a length of about 2 bases to about 8 bases (e.g., 2, 3, 4, 5, 6, 7, or 8 bases). A binding cycle-specific tag can be incorporated into a binding agent's coding tag, as part of a spacer sequence, as part of an encoder sequence, as part of a UMI, or as a separate component within the coding tag.
[0259] As used herein, the term "spacer" (Sp) refers to a nucleic acid molecule of about 1 base to about 20 bases (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 bases) in length present on the end of a recording tag or coding tag. In certain embodiments, the spacer sequence flanks the encoder sequence of the coding tag at one or both ends. After binding of the binding agent to the macromolecule, annealing between the complementary spacer sequences on the associated coding tag and recording tag allows the transfer of binding information to the recording tag, coding tag, or ditag construct by primer extension or ligation, respectively. Sp' refers to a spacer sequence complementary to Sp. Spacer sequences in a library of binding agents preferably have the same number of bases. Common (shared or identical) spacers can be used in the library of binding agents. The spacer sequence may have a "cycle-specific" sequence to track the binding substance used in a particular binding cycle. The spacer sequence (Sp) may be constant across all binding cycles, specific to a particular class of macromolecule, or specific to the number of binding cycles. The macromolecule class-specific spacer allows the coding tag information present in the extension recording tag of a cognate binding substance from a completed binding / extension cycle to anneal via the class-specific spacer to the coding tag of another binding substance that recognizes the same class of macromolecule in a subsequent binding cycle. Only sequential binding of the correct cognate pair results in interacting spacer elements and effective primer extension. The spacer sequence may contain a sufficient number of bases to anneal with a complementary spacer sequence in the recording tag to initiate a primer extension (also called polymerase extension) reaction, or to provide a "splint" for a ligation reaction, or to mediate a "sticky end" ligation reaction. The spacer sequence may contain fewer bases than the encoder sequence in the coding tag.
[0260] As used herein, the term "recording tag" refers to a nucleic acid molecule or sequenceable polymer molecule that contains identifying information about the macromolecule to which it is associated (see, e.g., Niu et al., 2013, Nat. Chem., 5:282-292; Roy et al., 2015, Nat. Commun., 6:7237; Lutz, 2015, Macromolecules, 48:4759-4767, each of which is incorporated by reference in its entirety). In certain embodiments, after a binding agent binds to a macromolecule, information from a coding tag linked to the binding agent can be transferred to a recording tag associated with the macromolecule while the binding agent is bound to the macromolecule. In other embodiments, after a binding agent binds to a macromolecule, information from a recording tag associated with the macromolecule can be transferred to a coding tag linked to the binding agent while the binding agent is bound to the macromolecule. The recording tag may be directly linked to the macromolecule, linked to the macromolecule via a multifunctional linker, or associated with the macromolecule by being in proximity (or co-localized) with the macromolecule on a solid support. The recording tag may be linked to the 5' or 3' end, or to an internal site, so long as the linkage is compatible with the method used to transfer the coding tag information to the recording tag or vice versa. The recording tag may further comprise other functional components, such as a universal priming site, a unique molecular identifier, a barcode (e.g., a sample barcode, a fraction barcode, a spatial barcode, a compartment tag, etc.), a spacer sequence complementary to the spacer sequence of the coding tag, or any combination thereof. The spacer sequence of the recording tag is preferably at the 3' end of the recording tag in embodiments that use polymerase extension to transfer the coding tag information to the recording tag.
[0261] As used herein, the term "primer extension," also referred to as "polymerase extension," refers to a reaction catalyzed by a nucleic acid polymerase (e.g., a DNA polymerase) in which a nucleic acid molecule (e.g., an oligonucleotide primer, a spacer sequence) that is annealed to a complementary strand is extended by the polymerase using the complementary strand as a template.
[0262] As used herein, the term "unique molecular identifier" or "UMI" refers to a nucleic acid molecule about 3 to about 40 bases (3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 bases) in length that provides a unique identifier tag for each macromolecule (e.g., peptide) or binding entity to which the UMI is attached. A macromolecular UMI is generated by computationally deconvoluting sequencing data from multiple extended sequence tags to identify individual macromolecules. The UMI can be used to identify the originating elongation record tag. The binding agent UMI can be used to identify each individual binding agent that binds to a particular macromolecule. For example, the UMI can be used to identify the number of individual binding events for binding agents specific to a single amino acid present in a particular peptide molecule. When both a UMI and a barcode are mentioned in relation to a binding agent or macromolecule, it is understood that the barcode refers to identifying information other than the UMI for the individual binding agent or macromolecule (e.g., sample barcode, compartment barcode, binding cycle barcode).
[0263] As used herein, the terms "universal priming site" or "universal primer" or "universal priming sequence" refer to a nucleic acid molecule that can be used for library amplification and / or sequencing reactions. Universal priming sites include, but are not limited to, priming sites (primer sequences) for PCR amplification, flow cell adapter sequences that anneal to complementary oligonucleotides on the flow cell surface to enable bridge amplification in some next-generation sequencing platforms, sequencing priming sites, or combinations thereof. Universal priming sites can also be used for other types of amplification, including those commonly used in conjunction with next-generation digital sequencing. For example, extended recording tag molecules can be circularized to form DNA nanoballs that can be used as sequencing templates using universal priming sites in rolling circle amplification (Drmanac et al., 2009, Science 327:78-81). Alternatively, the recording tag molecule can be circularized and sequenced directly by polymerase extension from the universal priming site (Korlach et al., 2008, Proc. Natl. Acad. Sci., 105:1176-1181). The term "forward" when used in reference to a "universal priming site" or "universal primer" can also be referred to as "5'" or "sense." The term "reverse" when used in reference to a "universal priming site" or "universal primer" can also be referred to as "3'" or "antisense."
[0264] As used herein, the term "extended recording tag" refers to a recording tag to which the information of at least one coding tag (or its complementary sequence) of a binding agent has been transferred after the binding agent has bound to a macromolecule. The information of the coding tag can be transferred to the recording tag directly (e.g., by ligation) or indirectly (e.g., by primer extension). The information of the coding tag can be transferred to the recording tag enzymatically or chemically. The extended recording tag may comprise binding agent information of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200 or more coding tags. The base sequence of the extended recording tag may reflect the temporal and sequential order of binding of the binding entities identified by the coding tags, may reflect a partial sequential order of binding of the binding entities identified by the coding tags, or may not reflect any order of binding of the binding entities identified by the coding tags. In certain embodiments, the coding tag information present in the extended recording tag represents the analyzed macromolecular sequence with at least 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity. In certain embodiments where the analyzed macromolecule sequence is not represented with 100% identity by the extension recording tag, errors may result from off-target binding by the binding agent, or from "skipped" binding cycles (e.g., due to a failure of the primer extension reaction due to the inability of the binding agent to bind to the macromolecule during the binding cycle), or both.
[0265] As used herein, the term "extended coding tag" refers to a coding tag to which at least one recording tag (or its complementary sequence) information has been transferred after the binding substance to which the coding tag is attached binds to the macromolecule to which the recording tag is attached. The recording tag information can be transferred directly to the coding tag (e.g., ligation) or indirectly (e.g., primer extension). The recording tag information can be transferred enzymatically or chemically. In certain embodiments, the extended coding tag contains one recording tag information reflecting one binding event. As used herein, the term "ditag" or "ditag construct" or "ditag molecule" refers to a nucleic acid molecule to which at least one recording tag (or its complementary sequence) information and at least one coding tag (or its complementary sequence) have been transferred after the binding substance to which the coding tag is attached binds to the macromolecule to which the recording tag is attached (see Figure 11B). The recording tag information and the coding tag can be transferred indirectly to the ditag (e.g., primer extension). The information in the recording tag can be transferred enzymatically or chemically. In certain embodiments, the ditag comprises a UMI of the recording tag, a compartment tag of the recording tag, a universal priming site of the recording tag, a UMI of the coding tag, an encoder sequence of the coding tag, a binding cycle-specific barcode, a universal priming site of the coding tag, or any combination thereof.
[0266] As used herein, the terms "solid support," "solid surface," or "solid substrate" or "substrate" refer to any solid material, including porous and non-porous materials, to which macromolecules (e.g., peptides) can be directly or indirectly attached by any means known in the art, including covalent and non-covalent interactions, or any combination thereof. A solid support can be two-dimensional (e.g., planar) or three-dimensional (e.g., gel matrix or beads). A solid support can be any support surface, including, but not limited to, beads, microbeads, arrays, glass surfaces, silicon surfaces, plastic surfaces, filters, membranes, nylon, silicon wafer chips, flow-through chips, flow cells, biochips including signal transduction electronics, channels, microtiter wells, ELISA plates, spin interference disks, nitrocellulose membranes, nitrocellulose-based polymer surfaces, polymer matrices, nanoparticles, or microspheres. Materials for solid supports include, but are not limited to, acrylamide, agarose, cellulose, nitrocellulose, glass, gold, quartz, polystyrene, polyethylene vinyl acetate, polypropylene, polymethacrylate, polyethylene, polyethylene oxide, polysilicate, polycarbonate, Teflon, fluorocarbon, nylon, silicone rubber, polyanhydrides, polyglycolic acid, polylactic acid, polyorthoesters, functionalized silanes, polypropyl fumerate, collagen, glycosaminoglycans, polyamino acids, dextran, or any combination thereof. Solid supports further include thin films, membranes, bottles, dishes, shaped polymers such as fibers, woven fibers, tubes, particles, beads, microspheres, microparticles, or any combination thereof. For example, when the solid surface is a bead, the bead may include, but is not limited to, ceramic beads, polystyrene beads, polymer beads, methylstyrene beads, agarose beads, acrylamide beads, solid core beads, porous beads, paramagnetic beads, glass beads, or controlled pore beads.The beads may be spherical or irregularly shaped. The size of the beads may range from nanometers, e.g., 100 nm, to millimeters, e.g., 1 mm. In certain embodiments, the size of the beads ranges from about 0.2 microns to about 200 microns, or from about 0.5 microns to about 5 microns. In some embodiments, the beads may have a diameter of about 1 μm, about 1.5 μm, about 2 μm, about 2.5 μm, about 2.8 μm, about 3 μm, about 3.5 μm, about 4 μm, about 4.5 μm, about 5 μm, about 5.5 μm, about 6 μm, about 6.5 μm, about 7 μm, about 7.5 μm, about 8 μm, about 8.5 μm, about 9 μm, about 9.5 μm, about 10 μm, about 10.5 μm, about 15 μm, or about 20 μm. In certain embodiments, a "bead" solid support can refer to an individual bead or to multiple beads.
[0267] As used herein, the terms "nucleic acid molecule" or "polynucleotide" refer to single- or double-stranded polynucleotides containing deoxyribonucleotides or ribonucleotides linked by 3'-5' phosphodiester linkages, as well as polynucleotide analogs. Nucleic acid molecules include, but are not limited to, DNA, RNA, and cDNA. Polynucleotide analogs may have backbones other than the standard phosphodiester linkages found in naturally occurring polynucleotides, and optionally modified sugar moieties or moieties other than ribose or deoxyribose. Polynucleotide analogs contain bases capable of hydrogen bonding with standard polynucleotide bases by Watson-Crick base pairing, where the analog backbone presents the bases in a manner that allows such hydrogen bonding in a sequence-specific manner between the oligonucleotide analog molecule and the bases in the standard polynucleotide. Examples of polynucleotide analogs include, but are not limited to, xeno nucleic acids (XNAs), bridged nucleic acids (BNAs), glycol nucleic acids (GNAs), peptide nucleic acids (PNAs), gamma PNAs, morpholino polynucleotides, locked nucleic acids (LNAs), threose nucleic acids (TNAs), 2'-O-methyl polynucleotides, 2'-O-alkylribosyl-substituted polynucleotides, phosphorothioate polynucleotides, and boronophosphate polynucleotides. Polynucleotide analogs may have purine or pyrimidine analogs, including universal base analogs capable of pairing with any base, including, for example, 7-deazapurine analogs, 8-halopurine analogs, 5-halopyrimidine analogs, or hypoxanthine, nitroazole, isocarbostyril analogs, azolecarboxamide, and aromatic triazole analogs, or base analogs with additional functionality, such as a biotin moiety for affinity binding.
[0268] As used herein, "nucleic acid sequencing" means determining the order of nucleotides in a nucleic acid molecule or in a sample of nucleic acid molecules.
[0269] As used herein, "next-generation sequencing" refers to a high-throughput sequencing method that allows millions to billions of molecules to be sequenced in parallel. Examples of next-generation sequencing methods include sequencing by synthesis, sequencing by ligation, sequencing by hybridization, polony sequencing, ion semiconductor sequencing, and pyrosequencing. By attaching a primer to a solid substrate and a complementary sequence to a nucleic acid molecule, the nucleic acid molecule can be hybridized to the solid substrate via the primer, and then amplified by using a polymerase in separate regions on the solid substrate to generate multiple copies (these groups are sometimes referred to as polymerase colonies or polony). Therefore, during the sequencing process, a nucleotide at a specific position can be sequenced multiple times (e.g., hundreds or thousands of times) - this depth of coverage is referred to as "deep sequencing". Examples of high-throughput nucleic acid sequencing technologies include platforms offered by Illumina, BGI, Qiagen, Thermo-Fisher, and Roche, including formats such as parallel bead arrays, sequencing-by-synthesis, sequencing-by-ligation, capillary electrophoresis, electronic microchips, "biochips," microarrays, parallel microchips, and single molecule arrays, as reviewed by Service (Science 311:1544-1546, 2006).
[0270] As used herein, "single molecule sequencing" or "third-generation sequencing" refers to next-generation sequencing methods in which reads from a single molecule sequencing device are generated by sequencing a single molecule of DNA. Unlike next-generation sequencing methods that rely on amplification to clone many DNA molecules in parallel for stepwise sequencing, single molecule sequencing interrogates a single molecule of DNA and does not require amplification or synchronization. Single molecule sequencing includes methods that require pausing the sequencing reaction after each base incorporation (a "wash-and-scan" cycle) and methods that do not require pausing between read steps. Examples of single molecule sequencing methods include single molecule real-time sequencing (Pacific Biosciences), nanopore-based sequencing (Oxford Nanopore), duplex interrupted nanopore sequencing (DI), and direct imaging of DNA using advanced microscopy.
[0271] As used herein, "analysis" of a macromolecule refers to quantifying, characterizing, distinguishing, or a combination thereof, all or part of the components of the macromolecule. For example, analysis of a peptide, polypeptide, or protein includes determining all or part of the peptide's amino acid sequence (contiguous or non-contiguous). Analysis of a macromolecule also includes partial identification of the components of the macromolecule. For example, partial identification of amino acids within a macromolecular protein sequence can identify the protein's amino acids as belonging to a subset of possible amino acids. Analysis typically begins with analysis of n NTAAs and then proceeds to the next peptide amino acid (i.e., n-1, n-2, n-3, etc.). This is achieved by cleaving n NTAAs, thereby converting the n-1 peptide amino acid into the N-terminal amino acid (referred to herein as "n-1 NTAA"). Analysis of a peptide can also include determining the presence and frequency of post-translational modifications on the peptide, which may or may not include information about the sequential order of post-translational modifications on the peptide. Analysis of a peptide may also include determining the presence and frequency of epitopes within the peptide, which may or may not include information regarding the sequential order or location of the epitopes within the peptide. Analysis of a peptide may include combining different types of analysis, for example, obtaining epitope information, amino acid sequence information, post-translational modification information, or any combination thereof.
[0272] As used herein, the term "compartment" refers to a physical region or volume that separates or isolates a subset of macromolecules from a sample of macromolecules. For example, a compartment can separate individual cells from other cells or separate a subset of a sample's proteome from the rest of the sample's proteome. A compartment can be an aqueous compartment (e.g., a microfluidic droplet), a solid compartment (e.g., a picotiter or microtiter well on a plate, a tube, a vial, a gel bead), or a separate area on a surface. A compartment can contain one or more beads to which macromolecules can be immobilized.
[0273] As used herein, the term "compartment tag" or "compartment barcode" refers to a single- or double-stranded nucleic acid molecule of about 4 bases to about 100 bases (including 4 bases, 100 bases, and any integer in between) that contains identifying information about a constituent (e.g., the proteome of a single cell) within one or more compartments (e.g., a microfluidic droplet). Compartment barcodes identify a subset of macromolecules in a sample, such as a subset of a protein sample separated into the same physical compartment or group of compartments from multiple (e.g., millions to billions) compartments. Thus, compartment tags can be used to distinguish constituents from one or more compartments with the same compartment tag from constituents in another compartment with a different compartment tag, even after these constituents have been pooled together. Labeling proteins and / or peptides within each compartment or within a group of two or more compartments with a unique compartment tag allows for identification of peptides originating from the same protein, protein complex, or cell within an individual compartment or group of compartments. A compartment tag includes a barcode, optionally flanked on one or both sides by a spacer sequence, and an optional universal primer. The spacer sequence may be complementary to the spacer sequence of a recording tag, thereby enabling transfer of the compartment tag information to the recording tag. A compartment tag may also include a universal priming site, a unique molecular identifier (providing identifying information about the attached peptide), or both, particularly for embodiments in which the compartment tag includes a recording tag used in downstream peptide analysis methods described herein. A compartment tag may include a functional moiety for coupling to a peptide (e.g., an aldehyde, NHS, mTet, alkyne, etc.). Alternatively, a compartment tag may include a peptide containing a recognition sequence for a protein ligase to enable ligation of the compartment tag to a peptide of interest.A compartment may contain a single compartment tag, multiple identical compartment tags storing an optional UMI sequence, or two or more different compartment tags. In certain embodiments, each compartment contains a unique compartment tag (one-to-one mapping). In other embodiments, multiple compartments from a larger population of compartments contain the same compartment tag (many-to-one mapping). The compartment tag may be attached to a solid support (e.g., a bead) within the compartment or to the surface of the compartment itself (e.g., the surface of a picotiter well). Alternatively, the compartment tag may be free in solution within the compartment.
[0274] As used herein, the term "partitioning" refers to the random assignment of unique barcodes to subpopulations of macromolecules from a population of macromolecules in a sample. In certain embodiments, partitioning can be achieved by dividing the macromolecules into compartments. Partitioning can consist of macromolecules in a single compartment or can consist of macromolecules in multiple compartments from a population of compartments.
[0275] As used herein, "partition tag" or "partition barcode" refers to a single- or double-stranded nucleic acid molecule of about 4 bases to about 100 bases (including 4 bases, 100 bases, and any integer in between) that contains identifying information for partitioning. In certain embodiments, a partition tag for a macromolecule refers to the same compartment tag resulting from partitioning the macromolecule into compartment(s) labeled with the same barcode.
[0276] As used herein, the term "fraction" refers to a subset of macromolecules (e.g., proteins) within a sample that have been sorted out from the remainder of the sample or organelles using physical or chemical separation methods, such as fractionation by size, hydrophobicity, isoelectric point, affinity, etc. Separation methods include HPLC separation, gel separation, affinity separation, cell fractionation, organelle fractionation, tissue fractionation, etc. Physical properties such as fluid flow, magnetism, electric current, mass, density, etc. can also be used for separation.
[0277] As used herein, the term "fraction barcode" refers to a single- or double-stranded nucleic acid molecule of about 4 bases to about 100 bases (including 4 bases, 100 bases, and any integer in between) that contains identifying information about a macromolecule within a fraction.
[0278] III. Methods for analyzing macromolecules The method described herein provides a highly parallelized approach for macromolecule analysis. Highly multiplexed macromolecule binding assays are converted into nucleic acid molecule libraries for next-generation sequencing. The method presented herein is particularly useful for protein or peptide sequencing.
[0279] In a preferred embodiment, a protein sample is labeled at the single-molecule level with at least one nucleic acid recording tag that includes a barcode (e.g., sample barcode, compartment barcode) and an optional unique molecular identifier. The protein sample is subjected to proteolytic digestion to generate a population (e.g., millions to billions) of peptides labeled with recording tags. These recording tag-labeled peptides are pooled and randomly immobilized on a solid support (e.g., porous beads). The pooled, immobilized recording tag-labeled peptides are subjected to multiple successive binding cycles, each involving exposure to multiple binding agents (e.g., binding agents for all 20 naturally occurring amino acids) labeled with coding tags that contain encoder sequences that identify the associated binding agents. During each binding cycle, information regarding binding of the binding agents to the peptides is captured by transferring the binding agent's coding tag information to the recording tag (or by transferring the recording tag information to the coding tag or both the recording tag information and the coding tag information to a separate ditag construct). Once the binding cycles are complete, a library of extended recording tags (or extended coding tags or ditag constructs) representing the binding history of the assayed peptides is generated, which can be analyzed using very high-throughput next-generation digital sequencing methods. The use of nucleic acid barcodes for the recording tags allows for the deconvolution of large amounts of peptide sequencing data to identify, for example, the sample, cell, proteome subset, or protein from which the peptide sequences originated.
[0280] In one aspect, a method for analyzing a macromolecule is provided, comprising: (a) providing a macromolecule and an associated or co-localized recording tag attached to a solid support; (b) contacting the macromolecule with a first binding substance capable of binding to the macromolecule, the first binding substance comprising a first coding tag having identifying information related to the first binding substance; (c) transferring information from the first coding tag to the recording tag to generate a primary extension recording tag; (d) contacting the macromolecule with a second binding substance capable of binding to the macromolecule, the second binding substance comprising a second coding tag having identifying information related to the second binding substance; (e) transferring information from the second coding tag to the primary extension recording tag to generate a secondary extension recording tag; and (f) analyzing the secondary extension tag (see, e.g., Figures 2A-D).
[0281] In certain embodiments, contacting steps (b) and (d) are performed sequentially, e.g., the first binding agent and the second binding agent are contacted with the macromolecule in separate binding cycle reactions. In other embodiments, contacting steps (b) and (d) are performed simultaneously, e.g., in a single binding cycle reaction including the first binding agent, the second binding agent, and optionally additional binding agents. In preferred embodiments, contacting steps (b) and (d) each include contacting the macromolecule with multiple binding agents.
[0282] In certain embodiments, the method further includes, between steps (e) and (f), (x) repeating steps (d) and (e) one or more times by replacing the second binding substance with a third (or higher order) binding substance capable of binding to the macromolecule, the third (or higher order) binding substance comprising a third (or higher order) coding tag having identifying information for the third (or higher order) binding substance; (y) transferring information from the third (or higher order) coding tag to the second (or higher order) extended recording tag to generate a third (or higher order) extended recording tag; and (z) analyzing the third (or higher order) extended recording tag.
[0283] The third (or higher order) binding substance can be contacted with the macromolecule in a separate binding cycle reaction from the first and second binding substances, or the third (or higher order) binding substance can be contacted with the macromolecule together with the first and second binding substances in a single binding cycle reaction.
[0284] In a second aspect, a method for analyzing a macromolecule is provided, the method comprising the steps of: (a) providing a macromolecule attached to a solid support, an associated first recording tag, and an associated second recording tag; (b) contacting the macromolecule with a first binding substance capable of binding to the macromolecule, the first binding substance comprising a first coding tag having identifying information related to the first binding substance; (c) transferring information from the first coding tag to the first recording tag to generate a first extended recording tag; (d) contacting the macromolecule with a second binding substance capable of binding to the macromolecule, the second binding substance comprising a second coding tag having identifying information related to the second binding substance; (e) transferring information from the second coding tag to the second recording tag to generate a second extended recording tag; and (f) analyzing the first extended recording tag and the second extended recording tag.
[0285] In certain embodiments, contacting steps (b) and (d) are performed sequentially, e.g., the first binding agent and the second binding agent are contacted with the macromolecule in separate binding cycle reactions. In other embodiments, contacting steps (b) and (d) are performed simultaneously, e.g., in a single binding cycle reaction including the first binding agent, the second binding agent, and optionally additional binding agents.
[0286] In certain embodiments, step (a) further comprises providing an associated third (or higher order) recording tag attached to the solid support. In further embodiments, the method further comprises, between steps (e) and (f), (x) repeating steps (d) and (e) one or more times by replacing the second binding substance with a third (or higher order) binding substance capable of binding to the macromolecule, the third (or higher order) binding substance comprising a third (or higher order) coding tag having identifying information for the third (or higher order) binding substance; (y) transferring information from the third (or higher order) coding tag to the third (or higher order) recording tag to generate a third (or higher order) extended recording tag; and (z) analyzing the first extended recording tag, the second extended recording tag, and the third (or higher order) extended recording tag.
[0287] The third (or higher order) binding substance can be contacted with the macromolecule in a separate binding cycle reaction from the first and second binding substances, or the third (or higher order) binding substance can be contacted with the macromolecule together with the first and second binding substances in a single binding cycle reaction.
[0288] In certain embodiments, the first coding tag, the second coding tag, and any higher order coding tag each have a binding cycle-specific sequence.
[0289] In a third aspect, a method for analyzing a peptide is provided, comprising: (a) providing a peptide and an associated recording tag attached to a solid support; (b) modifying the N-terminal amino acid (NTAA) of the peptide with a chemical moiety to produce a modified NTAA; (c) contacting the peptide with a first binding substance capable of binding to the modified NTAA, the first binding substance comprising a first coding tag having identifying information related to the first binding substance; (d) transferring information from the first coding tag to the recording tag to produce an extended recording tag; and (e) analyzing the extended recording tag (see, e.g., Figure 3).
[0290] In certain embodiments, step (c) further comprises contacting the peptide with a second (or higher order) binding substance comprising a second (or higher order) coding tag having identifying information for the second (or higher order) binding substance, the second (or higher order) binding substance being capable of binding to a modified NTAA other than the modified NTAA of step (b). In further embodiments, contacting the peptide with the second (or higher order) binding substance is performed sequentially after contacting the peptide with the first binding substance, e.g., the first binding substance and the second (or higher order) binding substance are contacted with the peptide in separate binding cycle reactions. In other embodiments, contacting the peptide with the second (or higher order) binding substance is performed simultaneously with contacting the peptide with the first binding substance, e.g., in a single binding cycle reaction including the first binding substance and the second (or higher order) binding substance.
[0291] In certain embodiments, the chemical moiety is attached to the NTAA via a chemical or enzymatic reaction.
[0292] In certain embodiments, the chemical moiety used to modify NTAA is a phenylthiocarbamoyl (PTC), a dinitrophenol (DNP) moiety; a sulfonyloxynitrophenyl (SNP) moiety, a dansyl moiety; a 7-methoxycoumarin moiety; a thioacyl moiety; a thioacetyl moiety; an acetyl moiety; a guanidinyl moiety; or a thiobenzyl moiety.
[0293] Chemical moieties can be added to NTAA using chemical agents. In certain embodiments, the chemical agent for modifying NTAA with a PTC moiety is phenylisothiocyanate or a derivative thereof; the chemical agent for modifying NTAA with a DNP moiety is an aryl halide such as 2,4-dinitrobenzenesulfonic acid (DNBS) or 1-fluoro-2,4-dinitrobenzene (DNFB); the chemical agent for modifying NTAA with a sulfonyloxynitrophenyl (SNP) moiety is 4-sulfonyl-2-nitrofluorobenzene (SNFB); the chemical agent for modifying NTAA with a dansyl group is a sulfonyl chloride such as dansyl chloride. a chemical agent for modifying NTAA with a 7-methoxycoumarin moiety is 7-methoxycoumarin acetic acid (MCA); a chemical agent for modifying NTAA with a thioacyl moiety is a thioacylation reagent; a chemical agent for modifying NTAA with a thioacetyl moiety is a thioacetylation reagent; a chemical agent for modifying NTAA with an acetyl moiety is an acetylation reagent (e.g., acetic anhydride); a chemical agent for modifying NTAA with a guanidinyl(amidinyl) moiety is a guanidinylation reagent; or a chemical agent for modifying NTAA with a thiobenzyl moiety is a thiobenzylation reagent.
[0294] In a fourth aspect, the disclosure provides a method for analyzing a peptide, the method comprising: (a) providing a peptide conjugated to a solid support and an associated recording tag; (b) modifying an N-terminal amino acid (NTAA) of the peptide with a chemical moiety to produce a modified NTAA; (c) contacting the peptide with a first binding agent capable of binding to the modified NTAA, the first binding agent including a first coding tag having identifying information related to the first binding agent; (d) transferring information from the first coding tag to the recording tag to produce a first extended recording tag; (e) transferring information from the first coding tag to the recording tag to produce a first extended recording tag; (f) removing the modified NTAA to expose a new NTAA; (f) modifying the new NTAA of the peptide with a chemical moiety to generate a newly modified NTAA; (g) contacting the peptide with a second binding substance capable of binding to the newly modified NTAA, the second binding substance including a second coding tag having identifying information about the second binding substance; (h) transferring information of the second coding tag to a first extended recording tag to generate a second extended recording tag; and (i) analyzing the second extended recording tag.
[0295] In certain embodiments, the contacting steps (c) and (g) are performed sequentially, eg, the first binding agent and the second binding agent are contacted with the peptide in separate binding cycle reactions.
[0296] In certain embodiments, the method further includes, between steps (h) and (i), (x) repeating steps (e), (f) and (g) one or more times by replacing the second binding substance with a third (or higher order) binding substance capable of binding to the modified NTAA, the third (or higher order) binding substance including a third (or higher order) coding tag having identifying information for the third (or higher order) binding substance; (y) transferring information from the third (or higher order) coding tag to the second (or higher order) extended recording tag to generate a third (or higher order) extended recording tag; and (z) analyzing the third (or higher order) extended recording tag.
[0297] In certain embodiments, the chemical moiety is attached to the NTAA via a chemical or enzymatic reaction.
[0298] In certain embodiments, the chemical moiety is a phenylthiocarbamoyl (PTC), a dinitrophenol (DNP) moiety; a sulfonyloxynitrophenyl (SNP) moiety, a dansyl moiety; a 7-methoxycoumarin moiety; a thioacyl moiety; a thioacetyl moiety; an acetyl moiety; a guanyl moiety; or a thiobenzyl moiety.
[0299] Chemical moieties can be added to NTAA using chemical agents. In certain embodiments, the chemical agent for modifying NTAA with a PTC moiety is phenylisothiocyanate or a derivative thereof; the chemical agent for modifying NTAA with a DNP moiety is an aryl halide such as 2,4-dinitrobenzenesulfonic acid (DNBS) or 1-fluoro-2,4-dinitrobenzene (DNFB); the chemical agent for modifying NTAA with a sulfonyloxynitrophenyl (SNP) moiety is 4-sulfonyl-2-nitrofluorobenzene (SNFB); the chemical agent for modifying NTAA with a dansyl group is a sulfonyl group such as dansyl chloride. the chemical agent for modifying NTAA with a 7-methoxycoumarin moiety is 7-methoxycoumarin acetic acid (MCA); the chemical agent for modifying NTAA with a thioacyl moiety is a thioacylation reagent; the chemical agent for modifying NTAA with a thioacetyl moiety is a thioacetylation reagent; the chemical agent for modifying NTAA with an acetyl moiety is an acetylating agent (e.g., acetic anhydride); the chemical agent for modifying NTAA with a guanyl moiety is a guanidinylation reagent, or the chemical agent for modifying NTAA with a thiobenzyl moiety is a thiobenzylation reagent.
[0300] In a fifth aspect, there is provided a method for analyzing a peptide, the method comprising the steps of: (a) providing a peptide and an associated recording tag attached to a solid support; (b) contacting the peptide with a first binding substance capable of binding to the N-terminal amino acid (NTAA) of the peptide, the first binding substance comprising a first coding tag having identifying information related to the first binding substance; (c) transferring information from the first coding tag to the recording tag to generate an extended recording tag; and (d) analyzing the extended recording tag.
[0301] In certain embodiments, step (b) further comprises contacting the peptide with a second (or higher order) binding substance comprising a second (or higher order) coding tag having identifying information for the second (or higher order) binding substance, the second (or higher order) binding substance being capable of binding to an NTAA other than the NTAA of the peptide. In further embodiments, the contacting of the peptide with the second (or higher order) binding substance is performed sequentially after contacting the peptide with the first binding substance, e.g., the first binding substance and the second (or higher order) binding substance are contacted with the peptide in separate binding cycle reactions. In other embodiments, the contacting of the peptide with the second (or higher order) binding substance is performed simultaneously with contacting the peptide with the first binding substance, e.g., in a single binding cycle reaction including the first binding substance and the second (or higher order) binding substance.
[0302] In a sixth aspect, a method for analyzing a peptide is provided, the method comprising the steps of: (a) providing a peptide and an associated recording tag attached to a solid support; (b) contacting the peptide with a first binding substance capable of binding to an N-terminal amino acid (NTAA) of the peptide, the first binding substance comprising a first coding tag having identifying information related to the first binding substance; (c) transferring information from the first coding tag to the recording tag to generate a first extended recording tag; (d) removing the NTAA to expose a new NTAA of the peptide; (e) contacting the peptide with a second binding substance capable of binding to the new NTAA, the second binding substance comprising a second coding tag having identifying information related to the second binding substance; (f) transferring information from the second coding tag to the first extended recording tag to generate a second extended recording tag; and (g) analyzing the second extended recording tag.
[0303] In certain embodiments, the method further includes, between steps (f) and (g), (x) repeating steps (d), (e) and (f) one or more times by replacing the second binding substance with a third (or higher order) binding substance capable of binding to the macromolecule, the third (or higher order) binding substance comprising a third (or higher order) coding tag having identifying information for the third (or higher order) binding substance; and (y) transferring information from the third (or higher order) coding tag to a second (or higher order) extended recording tag to generate a third (or higher order) extended recording tag, and analyzing the third (or higher order) extended recording tag in step (g).
[0304] In certain embodiments, the contacting steps (b) and (e) are performed sequentially, eg, the first binding agent and the second binding agent are contacted with the peptide in separate binding cycle reactions.
[0305] In any of the embodiments presented herein, the method comprises analyzing multiple macromolecules in parallel. In a preferred embodiment, the method comprises analyzing multiple peptides in parallel.
[0306] In any of the embodiments presented herein, the step of contacting the macromolecule (or peptide) with a binding agent comprises contacting the macromolecule (or peptide) with a plurality of binding agents.
[0307] In any of the embodiments presented herein, the macromolecule may be a protein, polypeptide, or peptide. In further embodiments, the peptide may be obtained by fragmenting a protein or polypeptide from a biological sample.
[0308] In any of the embodiments presented herein, the macromolecule may be or include a carbohydrate, lipid, nucleic acid, or macrocycle.
[0309] In any of the embodiments presented herein, the recording tag may be a DNA molecule, a DNA molecule with modified bases, an RNA molecule, a BNA molecule, an XNA molecule, an LNA molecule, a PNA molecule, a γPNA molecule (Dragulescu-Andrasi et al., 2006, J. Am. Chem. Soc., 128:10258-10267), a GNA molecule, or any combination thereof.
[0310] In any of the embodiments presented herein, the recording tag may include a universal priming site. In further embodiments, the universal priming site includes a priming site for amplification, ligation, sequencing, or a combination thereof.
[0311] In any of the embodiments presented herein, the recording tag may comprise a unique molecular identifier, a compartment tag, a partition barcode, a sample barcode, a fraction barcode, a spacer sequence, or any combination thereof.
[0312] In any of the embodiments presented herein, the coding tag may comprise a unique molecular identifier (UMI), an encoder sequence, a binding cycle-specific sequence, a spacer sequence, or any combination thereof.
[0313] In any of the embodiments presented herein, the binding cycle-specific sequence in the coding tag may be a binding cycle-specific spacer sequence.
[0314] In certain embodiments, the binding cycle-specific sequence is encoded as a separate barcode from the encoder sequence, hi other embodiments, the encoder sequence and the binding cycle-specific sequence are written in a single barcode that is unique to the binding agent and to each binding cycle.
[0315] In certain embodiments, the spacer sequence comprises a common binding cycle sequence shared among binding substances from multiple binding cycles, while in other embodiments, the spacer sequence comprises a unique binding cycle sequence shared among binding substances from the same binding cycle.
[0316] In any of the embodiments presented herein, the recording tag may include a barcode.
[0317] In any of the embodiments presented herein, the macromolecule and associated recording tag(s) may be covalently attached to a solid support.
[0318] In any of the embodiments presented herein, the solid support may be a bead, a porous bead, a porous matrix, an expandable gel bead or matrix, an array, a glass surface, a silicon surface, a plastic surface, a filter, a membrane, nylon, a silicon wafer chip, a flow-through chip, a biochip containing signal transduction electronics, a microtiter well, an ELISA plate, a spin interference disk, a nitrocellulose membrane, a nitrocellulose-based polymer surface, a nanoparticle, or a microsphere.
[0319] In any of the embodiments presented herein, the solid support may be a polystyrene bead, a polymer bead, an agarose bead, an acrylamide bead, a solid core bead, a porous bead, a paramagnetic bead, a glass bead, or a controlled pore bead.
[0320] In any of the embodiments provided herein, a plurality of macromolecules and associated recording tags can be attached to a solid support. In further embodiments, the plurality of macromolecules are spaced apart on the solid support by an average distance of ≧50 nm, ≧100 nm, or ≧200 nm.
[0321] In any of the embodiments presented herein, the binding agent can be a polypeptide or protein. In further embodiments, the binding agent is a modified or variant aminopeptidase, a modified or variant aminoacyl-tRNA synthetase, a modified or variant anticalin, or a modified or variant ClpS.
[0322] In any of the embodiments presented herein, the binding agent may be capable of selectively binding to a macromolecule.
[0323] In any of the embodiments presented herein, the coding tag may be a DNA molecule, a DNA molecule with modified bases, an RNA molecule, a BNA molecule, an XNA molecule, an LNA molecule, a GNA molecule, a PNA molecule, a γPNA molecule, or a combination thereof.
[0324] In any of the embodiments presented herein, the binding agent and the coding tag may be joined by a linker.
[0325] In any of the embodiments presented herein, the binding agent and coding tag may be joined by a SpyTag / SpyCatcher or SnoopTag / SnoopCatcher peptide-protein pair (Zakeri et al., 2012, Proc Natl Acad Sci USA 109(12):E690-697; Veggiani et al., 2016, Proc. Natl. Acad. Sci. USA 113:1202-1207, respectively, which are incorporated by reference in their entireties).
[0326] In any of the embodiments presented herein, the transfer of information from the coding tag to the recording tag is mediated by DNA ligase, or alternatively, the transfer of information from the coding tag to the recording tag is mediated by DNA polymerase or chemical ligation.
[0327] In any of the embodiments presented herein, analyzing the elongated recording tag may include nucleic acid sequencing. In further embodiments, the nucleic acid sequencing is sequencing by synthesis, sequencing by ligation, sequencing by hybridization, polony sequencing, ion semiconductor sequencing, or pyrosequencing. In other embodiments, the nucleic acid sequencing is single-molecule real-time sequencing, nanopore-based sequencing, nanogap tunneling sequencing, or direct imaging of DNA using advanced microscopy.
[0328] In any of the embodiments presented herein, the extended recording tag can be amplified before analysis.
[0329] In any of the embodiments presented herein, the order of the coding tag information contained in the extended recording tag can provide information regarding the order of binding to the macromolecule by the binding substance, and therefore the sequence of the analyte detected by the binding substance.
[0330] In any of the embodiments presented herein, the frequency of particular coding tag information (e.g., encoder sequence) contained in the extended recording tag can provide information regarding the frequency of binding of a particular binding substance to a macromolecule, and therefore the frequency of the analyte within the macromolecule that is detected by the binding substance.
[0331] In any of the embodiments disclosed herein, multiple macromolecule (e.g., protein) samples can be pooled, where a population of macromolecules within each sample is labeled with a recording tag that includes a sample-specific barcode. Such a pool of macromolecule samples can be subjected to binding cycles in a single reaction tube.
[0332] In any of the embodiments presented herein, multiple extended recording tags representing multiple macromolecules can be analyzed in parallel.
[0333] In any of the embodiments presented herein, multiple extended recording tags representing multiple macromolecules can be analyzed in a multiplexed assay.
[0334] In any of the embodiments presented herein, a target enrichment assay can be performed on multiple extended recording tags prior to analysis.
[0335] In any of the embodiments presented herein, multiple extended recording tags can be subjected to a subtraction assay prior to analysis.
[0336] In any of the embodiments presented herein, multiple extended recording tags can be subjected to a normalization assay prior to analysis to reduce highly abundant species.
[0337] In any of the embodiments presented herein, NTAA can be removed by modified aminopeptidases, modified amino acid tRNA synthetases, mild Edman degradation, Edmanase enzyme, or anhydrous TFA.
[0338] In any of the embodiments presented herein, at least one binding agent may be attached to a terminal amino acid residue. In certain embodiments, the terminal amino acid residue is the N-terminal amino acid or the C-terminal amino acid.
[0339] In any of the embodiments described herein, at least one binding agent may bind to a post-translationally modified amino acid.
[0340] Features of the above-described embodiments are presented in further detail in the following sections.
[0341] IV. Macromolecules In one aspect, the present disclosure relates to the analysis of macromolecules. Macromolecules are large molecules composed of smaller subunits. In certain embodiments, the macromolecule is a protein, a protein complex, a polypeptide, a peptide, a nucleic acid molecule, a carbohydrate, a lipid, a macrocycle, or a chimeric macromolecule.
[0342] Macromolecules (e.g., proteins, polypeptides, peptides) analyzed according to the methods disclosed herein can be obtained from biological samples such as, but not limited to, cells (both primary cells and cultured cell lines), cell lysates or extracts, organelles or vesicles, including exosomes, tissues and tissue extracts; biopsies; feces; bodily fluids of virtually any organism (e.g., blood, whole blood, serum, plasma, urine, lymph, bile, cerebrospinal fluid, interstitial fluid, aqueous or vitreous humor, colostrum, sputum, amniotic fluid, saliva, anal and vaginal secretions, sweat and semen, transudates, exudates (e.g., fluid obtained from an abscess or any other site of infection or inflammation), or joints (normal joints or joint lymphocytes). The samples may be obtained from any suitable source or sample, including fluids obtained from joints affected by diseases such as horse arthritis, osteoarthritis, gout or septic arthritis (preferably mammalian-derived samples including samples containing the microbiome, and particularly preferably human-derived samples including samples containing the microbiome); environmental samples (e.g., air, agricultural, water and soil samples, etc.); microbial samples, including microbial biofilms and / or microbial communities, and samples derived from microbial spores; and research samples, including extracellular fluids, extracellular supernatants from cell cultures, inclusion bodies within bacteria, cellular compartments including mitochondrial compartments, and cellular periplasm.
[0343] In certain embodiments, the macromolecule is a protein, protein complex, polypeptide, or peptide. The amino acid sequence information and post-translational modifications of the peptide, polypeptide, or protein are converted into a nucleic acid-encoded library that can be analyzed by next-generation sequencing. The peptide may contain L-amino acids, D-amino acids, or both. The peptide, polypeptide, protein, or protein complex may contain standard, naturally occurring amino acids, modified amino acids (e.g., post-translationally modified), amino acid analogs, amino acid mimetics, or any combination thereof. In some embodiments, the peptide, polypeptide, or protein is naturally occurring, synthetically produced, or recombinantly expressed. In any of the above-described peptide embodiments, the peptide, polypeptide, protein, or protein complex may further contain post-translational modifications.
[0344] Standard, naturally occurring amino acids include alanine (A or Ala), cysteine (C or Cys), aspartic acid (D or Asp), glutamic acid (E or Glu), phenylalanine (F or Phe), glycine (G or Gly), histidine (H or His), isoleucine (I or Ile), lysine (K or Lys), leucine (L or Leu), methionine (M or Met), asparagine (N or Asn), proline (P or Pro), glutamine (Q or Gln), arginine (R or Arg), serine (S or Ser), threonine (T or Thr), valine (V or Val), tryptophan (W or Trp), and tyrosine (Y or Tyr). Non-standard amino acids include selenocysteine, pyrrolysine, and N-formylmethionine, β-amino acids, homoamino acids, proline and pyruvate derivatives, 3-substituted alanine derivatives, glycine derivatives, ring-substituted phenylalanine and tyrosine derivatives, straight-chain core amino acids, and N-methyl amino acids.
[0345] Post-translational modifications (PTMs) of peptides, polypeptides, or proteins can be covalent or enzymatic. Examples of post-translational modifications include, but are not limited to, acylation, acetylation, alkylation (including methylation), biotinylation, butyrylation, carbamylation, carbonylation, deamidation, deimination, diphthamide formation, disulfide bridge formation, eliminylation, flavin attachment, formylation, gamma-carboxylation, glutamylation, glycylation, glycosylation (e.g., N-linked, O-linked, C-linked, phosphoglycosylation), glycosylphosphatidylinositol addition ( Post-translational modifications include modifications of the amino and / or carboxyl termini of peptides, polypeptides, or proteins. Modifications of the terminal amino group include, but are not limited to, desamino, N-lower alkyl, N-di-lower alkyl, and N-acyl modifications. Modifications of the terminal carboxyl group include, but are not limited to, amide, lower alkyl amide, dialkyl amide, and lower alkyl ester modifications (e.g., lower alkyl is a C1-C4 alkyl). Post-translational modifications also include, but are not limited to, modifications of amino acids between the amino and carboxy termini of a peptide, polypeptide, or protein, such as those described above. Post-translational modifications can regulate the "biological properties" of proteins within cells, such as their activity, structure, stability, or localization. Phosphorylation is the most common post-translational modification and plays an important role in protein regulation, particularly in cell signaling (Prabakaran et al., 2012, Wiley Interdiscip Rev Syst Biol Med 4:565-583).The addition of sugar to protein, such as glycosylation, has been shown to promote protein folding, improve stability and modify regulatory function.The attachment of lipid to protein allows targeting to cell membrane.Post-translational modification can also include the modification of peptide, polypeptide or protein to include one or more detectable labels.
[0346] In certain embodiments, peptides, polypeptides, or proteins can be fragmented. For example, fragmented peptides can be obtained by fragmenting proteins derived from a sample, such as a biological sample. Peptides, polypeptides, or proteins can be fragmented by any means known in the art, including fragmentation with a protease or endopeptidase. In some embodiments, fragmentation of peptides, polypeptides, or proteins is targeted by using specific proteases or endopeptidases that bind to and cleave specific consensus sequences (e.g., TEV protease specific for the ENLYFQ\S consensus sequence). In other embodiments, fragmentation of peptides, polypeptides, or proteins is non-targeted or random by using non-specific proteases or endopeptidases that can bind to and cleave specific amino acid residues rather than consensus sequences (e.g., proteinase K is a non-specific serine protease). Proteinases and endopeptidases are well known in the art, and examples of those that can be used to cleave proteins or polypeptides into smaller peptide fragments include proteinase K, trypsin, chymotrypsin, pepsin, thermolysin, thrombin, factor Xa, furin, endopeptidase, papain, pepsin, subtilisin, elastase, enterokinase, Genenase™ I, endoprotease LysC, endoprotease AspN, endoprotease GluC, and the like (Granvogl et al., 2007, Anal Bioanal Chem, 389:991-1002). In certain embodiments, peptides, polypeptides, or proteins are fragmented with proteinase K, or optionally, a heat-labile form of proteinase K to allow for rapid inactivation. Proteinase K is highly stable in denaturing reagents such as urea and SDS, allowing for the digestion of fully denatured proteins.Fragmentation of proteins and polypeptides into peptides can be performed before or after the attachment of DNA tags or DNA recording tags.
[0347] Chemical reagents can also be used to digest proteins into peptide fragments. These can cleave at specific amino acid residues (e.g., cyanogen bromide hydrolyzes peptide bonds at the C-terminus of methionine residues). Chemical reagents used to fragment polypeptides or proteins into smaller peptides include cyanogen bromide (CNBr), hydroxylamine, hydrazine, formic acid, BNPS-skatole [2-(2-nitrophenylsulfenyl)-3-methylindole], iodosobenzoic acid, and NTCB+Ni (2-nitro-5-thiocyanobenzoic acid).
[0348] In certain embodiments, after enzymatic or chemical cleavage, the resulting peptide fragments are of approximately the same desired length, e.g., from about 10 to about 70 amino acids, from about 10 to about 60 amino acids, from about 10 to about 50 amino acids, from about 10 to about 40 amino acids, from about 10 to about 30 amino acids, from about 20 to about 70 amino acids, from about 20 to about 60 amino acids, from about 20 to about 50 amino acids, from about 20 to about 40 amino acids, from about 20 to about 30 amino acids, from about 30 to about 70 amino acids, from about 30 to about 60 amino acids, from about 30 to about 50 amino acids, or from about 30 to about 40 amino acids. The cleavage reaction can be monitored, preferably in real time, by spiking the protein or polypeptide sample with a short test FRET (fluorescence resonance energy transfer) peptide comprising a peptide sequence containing a proteinase or endopeptidase cleavage site. In intact FRET peptides, fluorescent and quencher groups are attached to either end of the peptide sequence containing the cleavage site, and low fluorescence occurs due to fluorescence resonance energy transfer between the quencher and the fluorophore. When the test peptide is cleaved by a protease or endopeptidase, the quencher and the fluorophore are separated, resulting in a large increase in fluorescence. The cleavage reaction can be stopped when a certain fluorescence intensity is achieved, allowing for reproducible cleavage endpoints.
[0349] A sample of macromolecules (e.g., peptides, polypeptides, or proteins) can be subjected to protein fractionation before attachment to a solid support, in which proteins or peptides are separated by one or more properties, such as subcellular location, molecular weight, hydrophobicity, or isoelectric point, or by protein enrichment methods. Alternatively, or in addition, protein enrichment methods can be used to select specific proteins or peptides (see, e.g., Whiteaker et al., 2007, Anal. Biochem., 362:44-54, incorporated by reference in its entirety) or to select specific post-translational modifications (see, e.g., Huang et al., 2014, J. Chromatogr. A, 1372:1-17, incorporated by reference in its entirety). Alternatively, specific classes of proteins, such as immunoglobulins, or immunoglobulin (Ig) isotypes, such as IgG, can be enriched or selected for analysis by affinity. In the case of immunoglobulin molecules, it is particularly interesting to analyze the sequence and abundance or frequency of hypervariable sequences involved in affinity binding, especially because they fluctuate in response to disease progression or correlate with health, immune, and / or disease phenotypes.Excessively abundant proteins can also be subtracted from samples using standard immunoaffinity methods.Depletion of abundant proteins can be useful for plasma samples, where more than 80% of the protein composition is albumin and immunoglobulin.Several commercial products are available for depleting excessively abundant proteins from plasma samples, such as PROTIA and PROT20 (Sigma-Aldrich).
[0350] In certain embodiments, the macromolecule is composed of a protein or polypeptide. In one embodiment, the protein or polypeptide is labeled with a DNA recording tag using standard amine coupling chemistry (see, e.g., Figures 2B, 2C, 28, 29, 31, 40). ε-amino groups (e.g., of lysine residues) and N-terminal amino groups are particularly amenable to labeling with amine-reactive coupling agents, depending on the pH of the reaction (Mendoza and Vachet, 2009). In certain embodiments (see, e.g., Figures 2B and 29), the recording tag is composed of a reactive moiety (e.g., a moiety for conjugation to a solid surface, a multifunctional linker, or a macromolecule), a linker, a universal priming sequence, a barcode (e.g., a compartment tag, partition barcode, sample barcode, fraction barcode, or any combination thereof), an optional UMI, and a spacer (Sp) sequence to facilitate information transfer to / from the coding tag. In another embodiment, proteins can be first labeled with a universal DNA tag, followed by an enzymatic or chemical coupling step to attach a barcode-Sp sequence (representing a sample, compartment, physical location on a slide, etc.) to the protein (see, e.g., Figures 20, 30, 31, and 40). Universal DNA tags are used to label protein or polypeptide macromolecules and contain a short sequence of nucleotides that can be used as an attachment point for a barcode (e.g., a compartment tag, a recording tag, etc.). For example, the recording tag may contain a sequence at its end that is complementary to the universal DNA tag. In certain embodiments, the universal DNA tag is a universal priming sequence. Once the universal DNA tag on the labeled protein hybridizes to a complementary sequence in the recording tag (e.g., attached to a bead), the annealed universal DNA tag can be extended by primer extension, thereby transferring the recording tag information to the DNA-tagged protein. In certain embodiments, proteins are labeled with a universal DNA tag before being digested into peptides with a proteinase.The universal DNA tags on the labeled peptides derived from the digest can then be converted into informative and useful recording tags.
[0351] In certain embodiments, protein macromolecules can be immobilized (and optionally covalently cross-linked) to a solid support by an affinity capture reagent, where the recording tag is directly associated with the affinity capture reagent, or the protein can be directly immobilized to the solid support along with the recording tag (see, e.g., Figure 2C).
[0352] V. Solid support The macromolecules of the present disclosure are attached to the surface of a solid support (also referred to as a "substrate surface"). The solid support may be any porous or non-porous support surface, including, but not limited to, beads, microbeads, arrays, glass surfaces, silicon surfaces, plastic surfaces, filters, membranes, nylon, silicon wafer chips, flow cells, flow-through chips, biochips containing signal transduction electronics, microtiter wells, ELISA plates, spin interference disks, nitrocellulose membranes, nitrocellulose-based polymer surfaces, nanoparticles, or microspheres. Materials for solid supports include, but are not limited to, acrylamide, agarose, cellulose, nitrocellulose, glass, gold, quartz, polystyrene, polyethylene vinyl acetate, polypropylene, polymethacrylate, polyethylene, polyethylene oxide, polysilicate, polycarbonate, Teflon, fluorocarbon, nylon, silicone rubber, polyanhydrides, polyglycolic acid, polylactic acid, polyorthoesters, functionalized silanes, polypropyl fumerate, collagen, glycosaminoglycans, polyamino acids, or any combination thereof. Solid supports further include thin films, membranes, bottles, dishes, shaped polymers such as fibers, woven fibers, tubes, particles, beads, microparticles, or any combination thereof. For example, when the solid surface is a bead, the bead can include, but is not limited to, polystyrene beads, polymer beads, agarose beads, acrylamide beads, solid core beads, porous beads, paramagnetic beads, glass beads, or controlled pore beads.
[0353] In certain embodiments, the solid support is a flow cell. The form of the flow cell can vary among different next-generation sequencing platforms. For example, an Illumina flow cell is a flat, optically transparent surface, similar to a microscope slide, with a series of oligonucleotide anchors attached to its surface. Template DNA is ligated at its ends with adapters complementary to the oligonucleotides on the flow cell surface. The adapter-tagged single-stranded DNA is bound to the flow cell, amplified by solid-phase "bridge" PCR, and then sequenced. The 454 flow cell (454 Life Sciences) supports a "picotiter" plate, a fiber-optic slide with approximately 1.6 million 75-picoliter wells. Individual molecules of sheared template DNA are captured on separate beads, and each bead is compartmentalized into an individual droplet of aqueous PCR reaction mixture in an oil emulsion. The template is clonally amplified on the surface of beads by PCR, and then the template-loaded beads are distributed into the wells of a picotiter plate for sequencing reaction, ideally one or less beads per well. In the SOLiD (Supported Oligonucleotide Ligation and Detection) instrument from Applied Biosystems, template molecules are amplified by emulsion PCR, similar to the 454 system. After a step of selecting beads that do not contain amplified template, the template bound to the beads accumulates on a flow cell. The flow cell can be a simple filter frit, such as the TWIST™ DNA synthesis column (Glen Research).
[0354] In certain embodiments, the solid support is a bead, which may refer to an individual bead or multiple beads. In some embodiments, the beads are compatible with the selected next-generation sequencing platform (e.g., SOLiD or 454) used for downstream analysis. In some embodiments, the solid support is an agarose bead, a paramagnetic bead, a polystyrene bead, a polymer bead, an acrylamide bead, a solid-core bead, a porous bead, a glass bead, or a controlled-pore bead. In further embodiments, the beads can be coated with binding functionalities (e.g., amine groups, biotin-labeled macromolecules, affinity ligands such as streptavidin for binding to antibodies) to facilitate binding to macromolecules.
[0355] Proteins, polypeptides, or peptides can be attached to a solid support directly or indirectly by any means known in the art, including covalent and non-covalent interactions or any combination thereof (see, for example, Chan et al., 2007, PLoS One, 2:e1164; Cazalis et al., Bioconj. Chem., 15:1005-1009; Soellner et al., 2003, J. Am. Chem. Soc., 125:11790-11791; Sun et al., 2006, Bioconjug. Chem., 17:52-57; Decreau et al., 2007, J. Org. Chem., 72:2794-2802; Camarero et al., 2004, J. Am. Chem., 1999, pp. 141-142, each of which is hereby incorporated by reference in its entirety). Soc., 126:14730-14731; Girish et al., 2005, Bioorg. Med. Chem. Lett., 15:2447-2451; Kalia et al., 2007, Bioconjug. Chem., 18:1064-1069; Watzke et al., 2006, Angew Chem. Int. Ed. Engl., 45:1408-1412; Parthasarathy et al., 2007, Bioconjugate Chem., 18:469-476; and Bioconjugate Techniques, GT (See Hermanson, Academic Press (2013)). For example, peptides can be attached to the solid support by a ligation reaction. Alternatively, the solid support may include an agent or coating to facilitate direct or indirect attachment of the peptide to the solid support. Any suitable molecule or material can be used for this purpose, including proteins, nucleic acids, carbohydrates, and small molecules. For example, in one embodiment, the agent is an affinity molecule. In another example, the agent is an azide group that can react with an alkynyl group of another molecule to facilitate association or bonding between the solid support and the other molecule.
[0356] Proteins, polypeptides, or peptides can be conjugated to solid supports using a method called "click chemistry." To this end, any reaction that is rapid and substantially irreversible can be used to attach the protein, polypeptide, or peptide to the solid support. Exemplary reactions include the copper-catalyzed reaction of azides with alkynes to form triazoles (Husgen 1,3-dipolar cycloaddition), strain-promoted azide-alkyne cycloaddition (SPAAC), the reaction of dienes with dienophiles (Diels-Alder), strain-promoted alkyne-nitrone cycloaddition, the reaction of strained alkenes with azides, tetrazines, or tetrazoles, the alkene-azide [3 + 2] cycloaddition, the alkene-tetrazine inverse electron demand Diels-Alder (IEDDA) reaction (e.g., m-tetrazine (mTet)-trans-cyclooctene (TCO)), the alkene-tetrazole photoreaction, the Staudinger ligation of azides and phosphines, and various substitution reactions such as the displacement of leaving groups by nucleophilic attack on an electrophilic atom (Horisawa 2014; Knall, Hollauf et al. 2014). Exemplary substitution reactions include the reaction of an amine with an activated ester; the reaction of an amine with an N-hydroxysuccinimide ester; the reaction of an amine with an isocyanate; the reaction of an amine with an isothiocyanate; and the like.
[0357] In some embodiments, the macromolecule and the solid support are joined by a functional group that can be formed by the reaction of two complementary reactive groups, such as a functional group that is the product of one of the aforementioned "click" reactions. In various embodiments, the functional group can be formed by the reaction of an aldehyde, oxime, hydrazone, hydrazide, alkyne, amine, azide, acyl azide, acyl halide, nitrile, nitrone, sulfhydryl, disulfide, sulfonyl halide, isothiocyanate, imide ester, activated ester (e.g., N-hydroxysuccinimide ester, pentanoic acid STP ester), ketone, α,β-unsaturated carbonyl, alkene, maleimide, α-haloimide, epoxide, aziridine, tetrazine, tetrazole, phosphine, biotin, or thiirane functional group with a complementary reactive group. An exemplary reaction is the reaction of an amine (e.g., a primary amine) with an N-hydroxysuccinimide ester or isothiocyanate.
[0358] In still other embodiments, the functional group comprises an alkene, ester, amide, thioester, disulfide, carbocyclic, heterocyclic, or heteroaryl group. In further embodiments, the functional group comprises an alkene, ester, amide, thioester, thiourea, disulfide, carbocyclic, heterocyclic, or heteroaryl group. In other embodiments, the functional group comprises an amide or thiourea. In some more specific embodiments, the functional group is a triazolyl functional group, an amide, or a thiourea functional group.
[0359] In a preferred embodiment, iEDDA click chemistry is used to immobilize macromolecules (e.g., proteins, polypeptides, peptides) to solid supports because it is rapid and provides high yields at low input concentrations. In another preferred embodiment, m-tetrazine is used in the iEDDA click chemistry reaction rather than tetrazine because m-tetrazine has improved binding stability.
[0360] In a preferred embodiment, the substrate surface is functionalized with TCO, and proteins, polypeptides, or peptides labeled with recording tags are immobilized on the TCO-coated substrate surface via the attached m-tetrazine moieties (Figure 34).
[0361] Proteins, polypeptides, or peptides can be immobilized to solid support surfaces via their C-terminus, N-terminus, or internal amino acids, for example, via amine, carboxyl, or sulfhydryl groups. Standard activated supports used for coupling to amine groups include CNBr-activated, NHS-activated, aldehyde-activated, azlactone-activated, and CDI-activated supports. Standard activated supports used for carboxyl coupling include carbodiimide-activated carboxyl moieties coupled to amine supports. For cysteine coupling, maleimide-, iodoacetyl-, and pyridyl disulfide-activated supports can be used. An alternative method for peptide carboxy-terminal immobilization uses anhydrotrypsin, a catalytically inactive derivative of trypsin, which couples to peptides containing C-terminal lysine or arginine residues without cleaving them.
[0362] In certain embodiments, the protein, polypeptide, or peptide is immobilized on a solid support by covalently attaching a linker bound to the solid surface to a lysine group on the protein, polypeptide, or peptide.
[0363] A recording tag can be attached to a protein, polypeptide, or peptide before or after immobilization on a solid support. For example, a protein, polypeptide, or peptide can be first labeled with a recording tag and then immobilized on a solid surface via the recording tag, which contains two functional moieties for coupling (see Figure 28). One functional moiety of the recording tag is coupled to the protein, and the other functional moiety immobilizes the protein labeled with the recording tag on a solid support.
[0364] Alternatively, the protein, polypeptide, or peptide can be immobilized on a solid support before being labeled with a recording tag. For example, the protein can first be derivatized with a reactive group, such as a click chemistry moiety. The activated protein molecule can then be attached to a suitable solid support and then labeled with a recording tag using a complementary click chemistry moiety. For example, a protein derivatized with an alkyne and mTet moiety can be immobilized on beads derivatized with azide and TCO, and then attached to a recording tag labeled with azide and TCO.
[0365] It is understood that the methods presented herein for attaching a macromolecule (e.g., a protein, polypeptide, or peptide) to a solid support can also be used to attach a recording tag to a solid support or to attach a recording tag to a macromolecule (e.g., a protein, polypeptide, or peptide).
[0366] In certain embodiments, the surface of the solid support is passivated (blocked) to minimize nonspecific absorption to binding substances. A "passivated" surface refers to a surface that has been treated with an outer layer of material to minimize nonspecific binding of binding substances. Surface passivation methods include coating the surface with polyethylene glycol (PEG) (Pan et al., 2015, Phys. Biol., 12:045006), polysiloxane (e.g., Pluronic F-127), star polymers (e.g., star PEG) (Groll et al., 2010, Methods Enzymol., 472:1-18), hydrophobic dichlorodimethylsilane (DDS) plus self-assembled Tween-20 (Hua et al., 2014, Nat. Methods, 11:1233-1236), and diamond-like carbon (DLC), DLC plus PEG (Stavis et al., 2011, Proc. Natl. Standard methods from the literature for fluorescent single-molecule analysis include passivation with polymers such as PEG (Acad. Sci. USA, 108:983-988). In addition to covalent surface modification, several passivation agents can be used as well, including surfactants such as Tween-20, solution-based polysiloxanes (Pluronic series), polyvinyl alcohol (PVA), and proteins such as BSA and casein. Alternatively, the density of proteins, polypeptides, or peptides on the surface or within the volume of a solid substrate can be adjusted by spiking competitors or "dummy" reactive molecules when immobilizing the proteins, polypeptides, or peptides to the solid substrate (see Figure 36A).
[0367] In certain embodiments, when multiple macromolecules are immobilized on the same solid support, appropriate spacing between the macromolecules can be used to reduce or prevent cross-linking or intermolecular events, such as when a binding substance binds to a first macromolecule and transfers its coding tag information to a recording tag associated with a neighboring macromolecule rather than the recording tag associated with the first macromolecule. To control the spacing of macromolecules (e.g., protein, polypeptide, or peptide spacing) on a solid support, the density of functional coupling groups (e.g., TCO) on the substrate surface can be adjusted (see Figure 34). In some embodiments, multiple macromolecules are spaced apart on the surface or within the volume of a solid support (e.g., a porous support) by a distance of about 50 nm to about 500 nm, or about 50 nm to about 400 nm, or about 50 nm to about 300 nm, or about 50 nm to about 200 nm, or about 50 nm to about 100 nm. In some embodiments, a plurality of macromolecules are spaced apart on the surface of a solid support by an average distance of at least 50 nm, at least 60 nm, at least 70 nm, at least 80 nm, at least 90 nm, at least 100 nm, at least 150 nm, at least 200 nm, at least 250 nm, at least 300 nm, at least 350 nm, at least 400 nm, at least 450 nm, or at least 500 nm. In some embodiments, a plurality of macromolecules are spaced apart on the surface of a solid support by an average distance of at least 50 nm. In some embodiments, macromolecules are spaced apart on the surface or within the volume of a solid support so that, empirically, the relative frequency of intermolecular to intramolecular events is <1:10; <1:100; <1:1,000; or <1:10,000. Appropriate spacing frequencies can be determined empirically using functional assays (see Example 23) and can be achieved by dilution and / or by spiking in "dummy" spacer molecules that compete for attachment to sites on the substrate surface.
[0368] For example, as shown in Figure 34, PEG-5000 (MW approximately 5000) is used to block the interstitial space between peptides on a substrate surface (e.g., a bead surface). The peptides are then coupled to functional moieties also attached to PEG-5000 molecules. In a preferred embodiment, this is achieved by coupling a mixture of NHS-PEG-5000-TCO and NHS-PEG-5000-methyl to amine-derivatized beads (see Figure 34). The stoichiometric ratio between the two PEGs (TCO to methyl) is adjusted to produce an appropriate density of functional coupling moieties (TCO groups) on the substrate surface; methyl-PEG is inert to coupling. The effective spacing between TCO groups can be calculated by measuring the density of TCO groups on the surface. In certain embodiments, the average spacing between coupling moieties (e.g., TCO) on the solid surface is at least 50 nm, at least 100 nm, at least 250 nm, or at least 500 nm. After the beads are derivatized with PEG5000-TCO / methyl, excess NH2 groups on the surface are quenched with a reactive anhydride (e.g., acetic anhydride or succinic anhydride).
[0369] VI. Recording Tags At least one recording tag is directly or indirectly associated with or co-localized with the macromolecule and attached to a solid support (see, e.g., FIG. 5). The recording tag may comprise DNA, RNA, PNA, γPNA, GNA, BNA, XNA, TNA, a polynucleotide analog, or a combination thereof. The recording tag may be single-stranded or partially or completely double-stranded. The recording tag may have blunt ends or overhanging ends. In certain embodiments, once the binding agent binds to the macromolecule, the identity of the binding agent's coding tag is transferred to the recording tag to generate an extended recording tag. The extended recording tag can be further extended in subsequent binding cycles.
[0370] The recording tag can be attached to the solid support directly or indirectly (e.g., via a linker) by any means known in the art, including covalent and non-covalent interactions, or any combination thereof. For example, the recording tag can be attached to the solid support by a ligation reaction. Alternatively, the solid support may include an agent or coating to facilitate direct or indirect attachment of the recording tag to the solid support. Strategies for immobilizing nucleic acid molecules to solid supports (e.g., beads) are described in U.S. Pat. No. 5,900,481; Steinberg et al. (2004, Biopolymers, 73:597-605); Lund et al., 1988 (Nucleic Acids, 1988), each of which is incorporated herein by reference in its entirety. Res., 16:10861-10880); and Steinberg et al. (2004, Biopolymers, 73:597-605).
[0371] In certain embodiments, colocalization of a macromolecule (e.g., a peptide) and an associated recording tag is achieved by conjugating the macromolecule and recording tag with a bifunctional linker attached directly to the solid support surface. Steinberg et al. (2004, Biopolymer 73:597-605). In a further embodiment, a solid support (e.g., a bead) is derivatized with a trifunctional moiety, and the resulting bifunctional moiety is coupled to both the macromolecule and the recording tag.
[0372] Methods and reagents such as those described for the attachment of macromolecules and solid supports (eg, click chemistry reagents and photoaffinity labeling reagents) can also be used for the attachment of recording tags.
[0373] In certain embodiments, a single recording tag is attached to a macromolecule (e.g., a peptide), preferably by attachment to a deblocked N-terminal or C-terminal amino acid. In other embodiments, multiple recording tags are attached to a macromolecule (e.g., a protein, polypeptide, or peptide), preferably to a lysine residue or peptide backbone. In some embodiments, a macromolecule (e.g., a protein or polypeptide) labeled with multiple recording tags is fragmented or digested into smaller peptides, each labeled with, on average, one recording tag.
[0374] In certain embodiments, the recording tag optionally includes a unique molecular identifier (UMI) that provides a unique identifier tag for each macromolecule (e.g., protein, polypeptide, peptide) to which the UMI is associated. The UMI can be about 3 to about 40 bases, about 3 to about 30 bases, about 3 to about 20 bases, or about 3 to about 10 bases, or about 3 to about 8 bases. In some embodiments, the UMI is about 3 bases, 4 bases, 5 bases, 6 bases, 7 bases, 8 bases, 9 bases, 10 bases, 11 bases, 12 bases, 13 bases, 14 bases, 15 bases, 16 bases, 17 bases, 18 bases, 19 bases, 20 bases, 25 bases, 30 bases, 35 bases, or 40 bases in length. The UMI can be used to deconvolute sequencing data from multiple extended recording tags to identify sequence reads from individual macromolecules. In some embodiments, within a library of macromolecules, each macromolecule is associated with a single recording tag, and each recording tag includes a unique UMI. In other embodiments, multiple copies of the recording tag are associated with a single macromolecule, with each copy of the recording tag containing the same UMI. In some embodiments, the UMI has a different base sequence from the spacer or encoder sequence in the binding agent coding tag to facilitate distinguishing these components during sequence analysis.
[0375] In certain embodiments, the recording tag includes a barcode other than the UMI, for example, if present. The barcode is a nucleic acid molecule about 3 to about 30 bases, about 3 to about 25 bases, about 3 to about 20 bases, about 3 to about 10 bases, about 3 to about 10 bases, or about 3 to about 8 bases in length. In some embodiments, the barcode is about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, or 30 bases in length. In one embodiment, the barcode enables multiplexed sequencing of multiple samples or libraries. Barcodes can be used to identify the partition, fraction, compartment, sample, spatial location, or library from which a macromolecule (e.g., a peptide) originates. Barcodes are used to deconvolute multiplexed sequence data and identify sequence reads from individual samples or libraries. For example, barcoded beads are useful in methods involving emulsion and partitioning of samples, eg, for partitioning proteomes.
[0376] A barcode may represent a compartment tag, in which a unique barcode is assigned to a compartment, such as a droplet, a microwell, or a physical region on a solid support. The association of a compartment with a specific barcode can be achieved in any number of ways, such as by enclosing a single barcoded bead in the compartment, by directly mixing or adding a barcoded droplet to the compartment, or by directly printing or injecting a barcode reagent into the compartment. The barcode reagent in the compartment is used to attach a compartment-specific barcode to a macromolecule or fragment thereof in the compartment. When applied to the partitioning of proteins into compartments, barcodes can be used to map analyzed peptides back to their original protein molecules in the compartment, greatly facilitating protein identification. Compartment barcodes can also be used to identify protein complexes.
[0377] In other embodiments, multiple compartments representing a subset of a population of compartments can be assigned unique barcodes representing that subset.
[0378] Alternatively, the barcode may be a sample-identifying barcode. Sample barcodes are useful for multiplexed analysis of a set of samples in a single reaction vessel or immobilized on a single solid substrate or collection of solid substrates (e.g., a flat slide, a collection of beads contained in a single tube or vessel, etc.). Macromolecules from many different samples can be labeled with recording tags bearing sample-specific barcodes, and all samples can then be pooled together prior to immobilization to the solid support, cyclic binding, and recording tag analysis. Alternatively, samples can be kept separate until after DNA-encoded library creation, with sample barcodes attached during PCR amplification of the DNA-encoded library, and then mixed prior to sequencing. This approach can be useful when assaying analytes (e.g., proteins) of different abundance classes. For example, a sample can be split, barcoded, and one portion treated with a binding agent for a low-abundance analyte and the other portion treated with a binding agent for a higher-abundance analyte. In certain embodiments, this approach serves to adjust the dynamic range of a particular protein analyte assay to fall within the "sweet spot" of normal expression levels for the protein analyte.
[0379] In certain embodiments, peptides, polypeptides, or proteins from multiple different samples are labeled with recording tags containing sample-specific barcodes. Peptides, polypeptides, or proteins labeled with multiple sample barcodes can be mixed before cyclic binding reactions. This effectively creates a highly multiplexed alternative to digital reverse-phase protein arrays (RPPAs) (Guo, Liu, et al., 2012; Assadi, Lamerz, et al., 2013; Akbani, Becker, et al., 2014; Creighton and Huang, 2015). The creation of digital RPPA-like assays has numerous applications in translational research, biomarker validation, drug discovery, clinical practice, and precision medicine.
[0380] In certain embodiments, the recording tag includes a universal priming site, e.g., a forward or 5' universal priming site. A universal priming site is a nucleic acid sequence that can be used to prime a library amplification reaction and / or for sequencing. Universal priming sites can include, but are not limited to, priming sites for PCR amplification, flow cell adapter sequences that anneal to complementary oligonucleotides on the flow cell surface (e.g., Illumina next-generation sequencing), sequencing priming sites, or combinations thereof. The universal priming site can be from about 10 bases to about 60 bases. In some embodiments, the universal priming site includes an Illumina P5 primer (5'-AATGATACGGCGACCACCGA-3'-SEQ ID NO: 133) or an Illumina P7 primer (5'-CAAGCAGAAGACGGCATACGAGAT-3'-SEQ ID NO: 134).
[0381] In certain embodiments, the recording tag comprises a spacer at its end, e.g., the 3' end. As used herein, a reference to a spacer sequence in relation to a recording tag includes a spacer sequence identical to the spacer sequence associated with the cognate binder or a spacer sequence complementary to the spacer sequence associated with the cognate binder. The end, e.g., the 3' spacer, on the recording tag allows the identity of the cognate binder to be transferred from its coding tag to the recording tag during the first binding cycle (e.g., by annealing a complementary spacer sequence for primer extension or sticky end ligation).
[0382] In one embodiment, the spacer sequence is about 1 to 20 bases in length, about 2 to 12 bases in length, or about 5 to 10 bases in length. The length of the spacer may depend on factors such as the temperature and reaction conditions of the primer extension reaction for transferring the coding tag information to the recording tag.
[0383] In a preferred embodiment, the spacer sequence in the recording tag is designed to have minimal complementarity to other regions in the recording tag; similarly, the spacer sequence in the coding tag should have minimal complementarity to other regions in the coding tag. In other words, the spacer sequences of the recording tag and coding tag should have minimal sequence complementarity to components present in the recording tag or coding tag, such as unique molecular identifiers, barcodes (e.g., compartments, partitions, samples, spatial locations), universal primer sequences, encoder sequences, cycle-specific sequences, etc.
[0384] As described for binder spacers, in some embodiments, the recording tags associated with a library of macromolecules share a common spacer sequence. In other embodiments, the recording tags associated with a library of macromolecules have binding cycle-specific spacer sequences that are complementary to the binding cycle-specific spacer sequences of their cognate binders, which can be useful when using non-concatenated extension recording tags (see Figure 10).
[0385] The collection of extension recording tags can be concatenated later (see, for example, Figure 10). After the binding cycles are completed, a beaded solid support, where each bead has, on average, one or less macromolecules per bead, and each macromolecule has a collection of extension recording tags co-localized at the site of the macromolecule, is placed in an emulsion. The emulsion is formed so that each droplet is occupied by, on average, at most one bead. An optional assembly PCR reaction is performed in the emulsion to amplify the extension recording tags co-localized with the macromolecules on the beads and assemble them in a colinear order by priming between different cycle-specific sequences of different extension recording tags (Xiong, Peng et al., 2008). The emulsion is then broken, and the assembled extension recording tags are sequenced.
[0386] In another embodiment, the DNA recording tag is composed of a universal priming sequence (U1), one or more barcode sequences (BCs), and a spacer sequence (Sp1) specific to the first binding cycle. In the first binding cycle, the binding agent uses a DNA coding tag composed of an Sp1 complementary spacer, an encoder barcode, an optional cycle barcode, and a second spacer element (Sp2). The utility of using at least two different spacer elements is that the first binding cycle potentially selects one of multiple DNA recording tags and extends a single DNA recording tag, resulting in a new Sp2 spacer element at the end of the extended DNA recording tag. In the second and subsequent binding cycles, the binding agent contains only the Sp2' spacer, not the Sp1'. In this way, only the single extended recording tag from the first cycle is extended in subsequent cycles. In another embodiment, a binding agent-specific spacer can be used in the second and subsequent cycles.
[0387] In some embodiments, the recording tag comprises, from 5' to 3', a universal forward (or 5') priming sequence, a UMI, and a spacer sequence. In some embodiments, the recording tag comprises, from 5' to 3', a universal forward (or 5') priming sequence, an optional UMI, a barcode (e.g., a sample barcode, a partition barcode, a compartment barcode, a spatial barcode, or any combination thereof), and a spacer sequence. In some other embodiments, the recording tag comprises, from 5' to 3', a universal forward (or 5') priming sequence, a barcode (e.g., a sample barcode, a partition barcode, a compartment barcode, a spatial barcode, or any combination thereof), an optional UMI, and a spacer sequence.
[0388] Combinatorial approaches can be used to generate UMIs from modified DNA and PNA. In one example, UMIs can be constructed by "chemical ligation" of a set of short word sequences (4-15 mers) designed to be orthogonal to one another (Spiropoulos and Heemstra, 2012). A DNA template is used to guide the chemical ligation of the "word" polymer. The DNA template is constructed with hybridizing arms that allow the combinatorial template structure to be assembled simply by mixing the subcomponents together in solution (see Figure 12C). In certain embodiments, there are no "spacer" sequences in this design. The size of the word space can vary from 10 words to 10,000 or more words. In certain embodiments, the words are selected to be different from one another so as not to cross-hybridize, yet still maintain relatively uniform hybridization conditions. In one embodiment, the word length is approximately 10 bases, with about 1000 words in the subset (which is only 0.1% of the total 10-mer word space, approximately 4 10 = million words). These sets of words (1000 in a subset) are concatenated together to get a complexity of 1000. nFor the four words concatenated, this produces a final combinatorial UMI of power: 12 This results in UMI diversity for different elements of a species. These UMI sequences are added to macromolecules (peptides, proteins, etc.) at the single-molecule level. In one embodiment, the diversity of UMIs exceeds the number of macromolecules to which they are attached. In this way, the UMI uniquely identifies the macromolecule of interest. The use of combinatorial word UMIs facilitates reading on sequencers with high error rates (e.g., nanopore sequencers, nanogap tunneling sequencing, etc.) because single-base resolution is not required to read multi-base-long words. The combinatorial word approach can also be used to generate other identity-informative recording or coding tag components, such as compartment tags, partition barcodes, spatial barcodes, sample barcodes, encoder sequences, cycle-specific sequences, and barcodes. Methods for nanopore sequencing and DNA encoding information with error-tolerant words (codes) are known in the art (e.g., Kiah et al., 2015, Codes for DNA sequence profiles. IEEE International Symposium on Information Theory (ISIT); Gabrys et al., 2015, Asymmetric Lee distance codes for DNA-based storage. IEEE Symposium on Information Theory (ISIT); Laure et al., 2016, Coding in 2D: Using Intentional Dispersity to Enhance the Information Capacity of Sequence-Coded Polymer Barcodes. Angew. Chem. Int. Ed. doi:10.1002 / anie.201605279; Yazdi et al., 2015, IEEE Transactions on Molecular, Biological and Multi-Scale Communications, vol. 1:230-248; and Yazdi et al., 2015, Sci Rep, vol. 5:14138). Thus, in certain embodiments, the extended recording tag, extended coding tag, or ditag construct in any of the embodiments described herein is comprised of an identification component (e.g., a UMI, an encoder sequence, a barcode, a compartment tag, a cycle-specific sequence, etc.) that is an error-correcting code. In some embodiments, the error-correcting code is selected from a Hamming code, a Lee distance code, an asymmetric Lee distance code, a Reed-Solomon code, and a Levenshtein-Tenengolts code. With regard to nanopore sequencing, current or ion flux profiles and asymmetric base calling errors are intrinsic to the type and biochemistry of the nanopore used, and this information can be used to design more robust DNA codes using the error correction techniques described above. As an alternative to using robust DNA nanopore sequencing barcodes, the current or ion flux signature of the barcode sequence can be used directly (U.S. Patent No. 7,060,507, incorporated by reference in its entirety), thereby avoiding DNA base calling entirely and allowing barcode sequences to be readily identified by mapping them back to the predicted current / flux signature, as described in Laszlo et al. (2014, Nat. Biotechnol., 32:829-833, incorporated by reference in its entirety).In this paper, Laszlo et al. describe the ability to map and identify DNA strands from a biological nanopore, MspA, by examining the current signatures generated when different word strings are passed through the nanopore and mapping the resulting current signatures back to in silico predictions of possible current signatures from the universe of sequences (2014, Nat. Biotechnol. 32:829-833). A similar concept can be applied to the electrical signals generated by DNA sequencing based on the DNA code and nanogap tunneling currents (Ohshiro et al., 2012, Sci Rep 2:501).
[0389] Thus, in certain embodiments, the identification component of the coding tag, the recording tag, or both, is capable of generating a unique current or ion flux or optical signature, and the analyzing step of any of the methods presented herein comprises detecting the unique current or ion flux or optical signature to identify the identification component. In some embodiments, the identification component is selected from an encoder sequence, a barcode, a UMI, a compartment tag, a cycle-specific sequence, or any combination thereof.
[0390] In certain embodiments, all or a substantial amount (e.g., at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%) of the macromolecules (e.g., proteins, polypeptides, or peptides) in a sample are labeled with a recording tag. Labeling of the macromolecules can be performed before or after immobilization of the macromolecules to a solid support.
[0391] In other embodiments, a subset of macromolecules (e.g., proteins, polypeptides, or peptides) in a sample are labeled with a recording tag. In certain embodiments, a subset of macromolecules from a sample is labeled with a recording tag for targeted (analyte-specific) labeling. Targeted recording tag labeling of proteins can be achieved using a target protein-specific binding agent (e.g., an antibody, aptamer, etc.) linked to a complementary target-specific bait sequence, e.g., a short target-specific DNA capture probe that anneals to the analyte-specific barcode in the recording tag (see Figure 28A). The recording tag contains a reactive moiety for a cognate reactive moiety (e.g., click chemistry label, photoaffinity label) present on the target protein. For example, the recording tag can contain an azide moiety for interaction with an alkyne-derivatized protein, or the recording tag can contain a benzophenone for interaction with a native protein, etc. (see Figures 28A-B). Once the target protein-specific binding substance is attached to the target protein, the recording tag and the target protein are coupled via their corresponding reactive moieties (see Figures 28B-C). After labeling the target protein with the recording tag, the target protein-specific binding substance can be removed by digestion of a DNA capture probe linked to the target protein-specific binding substance. For example, the DNA capture probe can be designed to contain uracil bases, which can then be targeted for digestion with a uracil-specific cleavage reagent (e.g., USER™) to dissociate the target protein-specific binding substance from the target protein.
[0392] In one example, antibodies specific to a set of target proteins are coupled to complementary bait sequences (e.g., analyte barcodes BC1, BC2, BC3, BC4, BC5, BC6, BC7, BC8, BC9, BC10, BC11, BC12, BC13, BC14, BC15, BC16, BC17, BC18, BC19, BC20, BC2 A A DNA capture probe (e.g., analyte barcode BC in Figure 28) hybridizes to a recording tag designed using the ASample-specific labeling of proteins can be achieved by using antibodies labeled with DNA capture probes that hybridize to complementary bait sequences on recording tags containing sample-specific barcodes.
[0393] In another example, a target protein-specific aptamer is used to target a subset of proteins in a sample with a recording tag. The target-specific aptamer is linked to a DNA capture probe that anneals to a complementary bait sequence in the recording tag. The recording tag contains a reactive or photoreactive chemical probe (e.g., benzophenone (BP)) for coupling with a target protein having a corresponding reactive moiety. The aptamer binds to the target protein molecule, bringing the recording tag into close proximity with the target protein, thereby coupling the recording tag and the target protein.
[0394] Photoaffinity (PA) protein labeling using photoreactive chemical probes attached to small molecule protein affinity ligands has been previously described (Park, Koh, et al., 2016). Exemplary photoreactive chemical probes include benzophenone-based probes (reactive diradical, 365 nm), phenyldiazirine-based probes (reactive carbon, 365 nm), and phenylazide-based probes (reactive nitrene free radical, 260 nm) activated under previously described irradiation wavelengths (Smith and Collins, 2015). In a preferred embodiment, target proteins in a protein sample are labeled with recording tags containing sample barcodes using the method described by Li et al. (Li, Liu, et al., 2013), in which the bait sequence in the benzophenone-labeled recording tag is hybridized with a DNA capture probe attached to a cognate binding entity (e.g., a nucleic acid aptamer (see Figure 28)). For photoaffinity-labeled protein targets, DNA / RNA aptamers are preferred over antibodies as target protein-specific binding agents because the photoaffinity moiety can self-label the antibody rather than the target protein. In contrast, photoaffinity labeling is less efficient for nucleic acids than for proteins, making aptamers a better vehicle for DNA-directed chemical or photolabeling. Similar to photoaffinity labeling, DNA-directed chemical labeling of reactive lysines (or other moieties) near the aptamer binding site can also be used, as described by Rosen et al. (Rosen, Kodal et al., 2014; Kodal, Rosen et al., 2016).
[0395] In the above-described embodiments, other types of linkages besides hybridization can be used to link the target-specific binding substance and the recording tag (see Figure 28A). For example, as shown in Figure 28B, the two moieties can be covalently linked using a linker designed to be cleaved and release the binding substance once the captured target protein (or other macromolecule) is covalently bound to the recording tag. A suitable linker can be attached to various positions on the recording tag, such as at the 3' end or within a linker attached to the 5' end of the recording tag.
[0396] VII. Binding Substances and Coding Tags The methods described herein use binding agents capable of binding to macromolecules. A binding agent can be any molecule capable of binding to a component or feature of a macromolecule (e.g., a peptide, polypeptide, protein, nucleic acid, carbohydrate, small molecule, etc.). A binding agent can be a naturally occurring molecule, a synthetically produced molecule, or a recombinantly expressed molecule. A binding agent can bind to a single monomer or subunit of a macromolecule (e.g., a single amino acid of a peptide), or can bind to multiple linked subunits of a macromolecule (e.g., dipeptides, tripeptides, or higher peptides of a longer peptide molecule).
[0397] In certain embodiments, the binding substance can be designed to bind by a covalent bond. The covalent bond can be designed so that binding to a precise moiety is conditional or advantageous. For example, an NTAA and its cognate NTAA-specific binding substance can each be modified with a reactive group so that, once the NTAA-specific binding substance binds to the cognate NTAA, a coupling reaction occurs, resulting in a covalent linkage between the two. Nonspecific binding of the binding substance to other positions lacking the cognate reactive group does not result in covalent attachment. The covalent bond between the binding substance and its target allows for the use of more stringent washes to remove nonspecifically bound binding substances, thus increasing the specificity of the assay.
[0398] In certain embodiments, a binding agent may be a selective binding agent. As used herein, selective binding refers to the ability of a binding agent to preferentially bind to a specific ligand (e.g., an amino acid or class of amino acids) compared to binding to a different ligand (e.g., an amino acid or class of amino acids). Selectivity is generally taken as the equilibrium constant for a reaction in which one ligand is replaced by another in a complex with the binding agent. Generally, such selectivity is related to the spatial geometry of the ligand and / or the manner and degree to which the ligand binds to the binding agent, such as attachment to the binding agent by hydrogen bonding or van der Waals forces (non-covalent interactions) or by reversible or irreversible covalent bonds. It should also be understood that selectivity may be relative as opposed to absolute, and that various factors, including ligand concentration, can affect selectivity. Thus, in one example, a binding agent selectively binds to one of the 20 common amino acids. In an example of non-selective binding, the binding agent may bind to two or more of the 20 common amino acids.
[0399] In practicing the methods disclosed herein, the ability of a binding agent to selectively bind to a feature or component of a macromolecule need only be sufficient to allow the transfer of the coding tag information of the binding agent to a recording tag associated with the macromolecule, the transfer of the recording tag information to the coding tag, or the transfer of the coding tag information and the recording tag information to a ditag molecule. Thus, selectivity need only be relative to other binding agents to which the macromolecule is exposed. It should also be understood that the selectivity of a binding agent need not be absolute for a particular amino acid, but may be selective for a class of amino acids, such as amino acids with nonpolar or nonpolar side chains, or amino acids with electrically (positively or negatively) charged side chains, or amino acids with aromatic side chains, or for certain specific classes or sizes of side chains.
[0400] In certain embodiments, the binding agent has high affinity and high selectivity for the target macromolecule. In particular, a high binding affinity with a low dissociation rate is effective for information transfer between the coding tag and the recording tag. In certain embodiments, the Kd of the binding agent is <10 nM, <5 nM, <1 nM, <0.5 nM, or <0.1 nM. In certain embodiments, the binding agent is added to the macromolecule at a concentration >10x, >100x, or >1000x the Kd of the binding agent to drive binding to completion. A detailed discussion of the binding kinetics of antibodies to single protein molecules is provided in Chang et al. (Chang, Rissin et al., 2012).
[0401] To increase the affinity of a binder for small N-terminal amino acids (NTAA) of peptides, the NTAA can be modified with an "immunogenic" hapten such as dinitrophenol (DNP). This can be achieved by cyclic sequencing using dinitrofluorobenzene (DNFB), a Sanger reagent that attaches the DNP group to the amine group of the NTAA. Commercial anti-DNP antibodies have affinities in the low nM range (approximately 8 nM, LO-DNP-2) (Bilgicer, Thomas, et al., 2009); it should therefore be possible to engineer NTAA binders with high affinity for several NTAAs modified with DNP (via DNFB) while simultaneously achieving good binding selectivity for specific NTAAs. In another example, NTAAs can be modified with sulfonylnitrophenol (SNFB) using 4-sulfonyl-2-nitrofluorobenzene (SNFB). Similar affinity enhancements can also be achieved using alternative NTAA modifiers, such as acetyl or amidinyl (guanidinyl) groups.
[0402] In certain embodiments, a binding agent may bind to the NTAA, CTAA, intervening amino acids, dipeptides (sequences of two amino acids), tripeptides (sequences of three amino acids), or higher order peptides of a peptide molecule. In some embodiments, each binding agent in a library of binding agents selectively binds to a particular amino acid, for example, one of the 20 standard naturally occurring amino acids. Standard, naturally occurring amino acids include alanine (A or Ala), cysteine (C or Cys), aspartic acid (D or Asp), glutamic acid (E or Glu), phenylalanine (F or Phe), glycine (G or Gly), histidine (H or His), isoleucine (I or Ile), lysine (K or Lys), leucine (L or Leu), methionine (M or Met), asparagine (N or Asn), proline (P or Pro), glutamine (Q or Gln), arginine (R or Arg), serine (S or Ser), threonine (T or Thr), valine (V or Val), tryptophan (W or Trp), and tyrosine (Y or Tyr).
[0403] In certain embodiments, the binding agent may bind to post-translational modifications of amino acids. In some embodiments, the peptide comprises one or more post-translational modifications, which may be the same or different. The NTAA, CTAA, intervening amino acids, or a combination thereof of the peptide may be post-translationally modified. Post-translational modifications of amino acids include acylation, acetylation, alkylation (including methylation), biotinylation, butyrylation, carbamylation, carbonylation, deamidation, deimination, diphthamide formation, disulfide bridge formation, eliminylation, flavin attachment, formylation, gamma-carboxylation, glutamylation, glycylation, glycosylation, glycosylphosphatidylinositol glypiation, heme C attachment, hydroxylation, These include hypusine formation, iodination, isoprenylation, lipidation, lipoylation, malonylation, methylation, myristoylation, oxidation, palmitoylation, pegylation, phosphopantetheinylation, phosphorylation, prenylation, propionylation, retinylidene Schiff base formation, S-glutathionylation, S-nitrosylation, S-sulfenylation, selenation, succinylation, sulfination, ubiquitination, and C-terminal amidation (see also Seo and Lee, 2004, J. Biochem. Mol. Biol., 37:35-44).
[0404] In certain embodiments, lectins are used as binding agents for detecting the glycosylation status of proteins, polypeptides, or peptides. Lectins are carbohydrate-binding proteins that can selectively recognize free carbohydrates or glycan epitopes on glycoproteins. The list of lectins that recognize various glycosylation states (e.g., core-fucose, sialic acid, N-acetyl-D-lactosamine, mannose, and N-acetyl-glucosamine) includes A, AAA, AAL, ABA, ACA, ACG, ACL, AOL, ASA, BanLec, BC2L-A, BC2LCN, BPA, BPL, Calsepa, CGL2, CNL, Con, ConA, DBA, Discoidin, DSA, ECA, EEL, F17AG, Gal1, Gal1-S, Gal2, Gal3, Gal3C-S, Gal7-S, Gal9, GNA, GRFT, GS-I, GS-II, GSL-I, GSL-II, HHL, HIHA, HPA, I, II, Jacalin, and LBA. , LCA, LEA, LEL, Lentil, Lotus, LSL-N, LTL, MAA, MAH, MAL_I, Malectin, MOA, MPA, MPL, NPA, Orysata, PA-IIL, PA-IL, PALa, PHA-E, PHA-L, PHA-P, PHAE, PHAL, PNA, PPL, PSA, PSL1a, PTL, PT L-I, PWM, RCA120, RS-Fuc, SAMB, SBA, SJA, SNA, SNA-I, SNA-II, SSA, STL, TJA-I, TJA-II, TxL Including CI, UDA, UEA-I, UEA-II, VFA, VVA, WFA, WGA (see Zhang et al., 2016, MABS, 8:524-535).
[0405] In certain embodiments, the binding agent may bind to a modified or labeled NTAA, which may be labeled with PITC, 1-fluoro-2,4-dinitrobenzene (Sanger's reagent, DNFB), dansyl chloride (DNS-Cl, or 1-dimethylaminonaphthalene-5-sulfonyl chloride), 4-sulfonyl-2-nitrofluorobenzene (SNFB), an acetylating reagent, a guanidinating reagent, a thioacylation reagent, a thioacetylation reagent, or a thiobenzylation reagent.
[0406] In certain embodiments, the binding agent may be an aptamer (e.g., a peptide aptamer, a DNA aptamer, or an RNA aptamer), an antibody, anticalin, an ATP-dependent Clp protease adaptor protein (ClpS), an antibody-binding fragment, an antibody mimetic, a peptide, a peptidomimetic, a protein, or a polynucleotide (e.g., DNA, RNA, peptide nucleic acid (PNA), γPNA, bridged nucleic acid (BNA), heterologous nucleic acid (XNA), glycerol nucleic acid (GNA), or threose nucleic acid (TNA), or variants thereof).
[0407] As used herein, the terms antibody and antibodyy are used broadly to include intact antibody molecules, such as, but not limited to, Immunoglobulin A, Immunoglobulin G, Immunoglobulin D, Immunoglobulin E, and Immunoglobulin M, as well as any immunoreactive component(s) of an antibody molecule that immunospecifically binds to at least one epitope. Antibodies may be naturally occurring, synthetically produced, or recombinantly expressed. Antibodies may be fusion proteins. Antibodies may be antibody mimetics. Examples of antibodies include, but are not limited to, Fab fragments, Fab' fragments, F(ab')2 fragments, single-chain antibody fragments (scFv), miniantibodies, diabodies, cross-linked antibody fragments, Affibodies™, nanobodies, single-domain antibodies, DVD-Ig molecules, alphabodies, affimers, affitins, cyclotides, molecules, and the like. Immunoreactive products obtained using antibody or protein engineering techniques also expressly fall within the meaning of the term antibody. Detailed descriptions of antibody and / or protein engineering, including associated protocols, can be found, among other places, in J. Maynard and G. Georgiou, 2000, Ann. Rev. Biomed. Eng., 2:339-76; Antibody Engineering, R. Kontermann and S. Dubel (eds.), Springer Lab Manual, Springer Verlag (2001); U.S. Patent No. 5,831,012; and S. Paul, Antibody Engineering Protocols, Humana Press (1995).
[0408] Similar to antibodies, nucleic acid and peptide aptamers that specifically recognize peptides can be produced using known methods.Aptamers bind to target molecules in a highly specific, conformation-dependent manner, and generally have very high affinity, but aptamers with lower binding affinity can also be selected if desired.Aptamers have been shown to distinguish between targets based on very small structural differences, such as the presence or absence of methyl or hydroxyl groups, and some aptamers can distinguish between D-enantiomers and L-enantiomers.Aptamers have been obtained that bind to small molecule targets, including drugs, metal ions, and organic dyes, peptides, biotin, and proteins, including but not limited to streptavidin, VEGF, and viral proteins. Aptamers have been shown to retain functional activity after biotinylation, fluorescein labeling, and when attached to glass surfaces and microspheres (see Jayasena, 1999, Clin Chem, 45:1628-50; Kusser, 2000, J. Biotechnol., 74:27-39; Colas, 2000, Curr Opin Chem Biol, 4:54-9). Aptamers that specifically bind to arginine and AMP have also been described (see Patel and Suri, 2000, J. Biotech., 74:39-60). Oligonucleotide aptamers that bind to specific amino acids have been disclosed by Gold et al. (1995, Ann. Rev. Biochem., 64:763-97). RNA aptamers that bind to amino acids have also been described (Ames and Breaker, 2011, RNA Biol. 8:82-89; Mannironi et al., 2000, RNA 6:520-27; Famulok, 1994, J. Am. Chem. Soc. 116:1698-1706).
[0409] Binding agents can be generated by genetically modifying naturally occurring or synthetically produced proteins to introduce one or more mutations into their amino acid sequences, resulting in engineered proteins that bind to specific macromolecular components or features (e.g., NTAAs, CTAAs, or post-translationally modified amino acids or peptides). For example, exopeptidases (e.g., aminopeptidases, carboxypeptidases), exoproteases, mutant exoproteases, mutant anticalins, mutant ClpSs, antibodies, or tRNA synthetases can be modified to create binding agents that selectively bind to specific NTAAs. In another example, carboxypeptidases can be modified to create binding agents that selectively bind to specific CTAAs. Binding agents can also be designed or engineered and utilized to specifically bind modified NTAA or CTAA, such as those bearing post-translational modifications (e.g., phosphorylated NTAA or phosphorylated CTAA) or those modified with a label (e.g., using PTC, 1-fluoro-2,4-dinitrobenzene (using Sanger's reagent, DNFB), dansyl chloride (DNS-Cl, or 1-dimethylaminonaphthalene-5-sulfonyl chloride), or a thioacylation reagent, thioacetylation reagent, acetylation reagent, amidination (guanidination) reagent, or thiobenzylation reagent). Strategies for directed evolution of proteins are known in the art (e.g., reviewed by Yuan et al., 2005, Microbiol. Mol. Biol. Rev. 69:373-392) and include phage display, ribosome display, mRNA display, CIS display, CAD display, emulsion, cell surface display, yeast surface display, bacterial surface display, and the like.
[0410] In some embodiments, binding agents that selectively bind modified NTAA can be utilized. For example, NTAA can be reacted with phenylisothiocyanate (PITC) to form a phenylthiocarbamoyl-NTAA derivative. In this manner, binding agents can be adapted to selectively bind both the phenyl group of the phenylthiocarbamoyl moiety and the alpha-carbon R group of NTAA. The use of PITC in this manner allows for subsequent cleavage of NTAA by Edman degradation, as discussed below. In another embodiment, NTAA can be reacted with Sanger reagent (DNFB) to generate DNP-labeled NTAA (see Figure 3). Optionally, DNFB is used in conjunction with an ionic liquid, such as 1-ethyl-3-methylimidazolium bis[(trifluoromethyl)sulfonyl]imide ([emine][TfN]), in which DNFB is highly soluble. In this manner, binding agents can be engineered to selectively bind to a combination of DNP and the R group of NTAA. The addition of the DNP moiety provides a larger "handle" for the interaction between the binding agent and the NTAA, which should lead to a higher affinity interaction. In yet another embodiment, the binding agent may be an aminopeptidase engineered to recognize the DNP-labeled NTAA, thereby providing cyclical control of the peptide's aminopeptidase degradation. Once the DNP-labeled NTAA is cleaved, another cycle of DNFB derivatization is performed for binding to and cleavage of the newly exposed NTAA. In certain preferred embodiments, the aminopeptidase is a monomeric metallo-protease, and such aminopeptidases are activated by zinc (Calcagno and Klein, 2016). In another example, the binding agent may selectively bind to NTAAs modified with sulfonylnitrophenol (SNFB), for example, by using 4-sulfonyl-2-nitrofluorobenzene (SNFB). In yet another embodiment, the binding agent may selectively bind to acetylated or amidinated NTAAs.
[0411] Other reagents that can be used to modify NTAA include trifluoroethyl isothiocyanate, allyl isothiocyanate, and dimethylaminoazobenzene isothiocyanate.
[0412] Binding agents can be engineered for high affinity for the modified NTAA, high specificity for the modified NTAA, or both. In some embodiments, binding agents can be developed by directed evolution of promising affinity scaffolds using phage display.
[0413] Engineered aminopeptidase variants have been described that bind to and cleave individual or small groups of labeled (biotinylated) NTAA (see PCT Publication No. WO2010 / 065322, incorporated by reference in its entirety). Aminopeptidases are enzymes that cleave amino acids from the N-terminus of proteins or peptides. Natural aminopeptidases have very limited specificity and generally cleave the N-terminal amino acid in a processive manner, cleaving amino acids one after the other (Kishor et al., 2015, Anal. Biochem., 488:6-8). However, residue-specific aminopeptidases have been identified (Eriquez et al., J. Clin. Microbiol., 1980, 12:667-71; Wilce et al., 1998, Proc. Natl. Acad. Sci. USA, 95:3472-3477; Liao et al., 2004, Prot. Sci., 13:1802-10). Aminopeptidases can be engineered to specifically bind to 20 different NTAAs, representing standard amino acids, labeled with specific moieties (e.g., PTC, DNP, SNP, etc.). Control of the stepwise degradation of the N-terminus of peptides is achieved by using engineered aminopeptidases that are active (e.g., binding or catalytic) only in the presence of the label. In another example, Havranak et al. (U.S. Patent Publication No. 2014 / 0273004) describe the engineering of aminoacyl-tRNA synthetases (aaRSs) as specific NTAA binders. The amino acid binding pockets of aaRSs have an intrinsic ability to bind cognate amino acids, but generally exhibit poor binding affinity and specificity. Furthermore, these natural amino acid binders do not recognize N-terminal labels. Directed evolution of aaRS scaffolds can be used to generate higher affinity and more specific binders that recognize the N-terminal amino acid for the N-terminal tag.
[0414] In another example, highly selective engineered ClpS has also been described in the literature: Emili et al. describe that directed evolution of the E. coli ClpS protein by phage display resulted in four different variants capable of selectively binding NTAA to aspartic acid, arginine, tryptophan, and leucine residues (U.S. Patent No. 9,566,335, incorporated by reference in its entirety).
[0415] In certain embodiments, anticalins are engineered for both high affinity and high specificity for labeled NTAA (e.g., DNP, SNP, acetylation, etc.). Certain variants of the anticalin scaffold have a shape suitable for binding to a single amino acid due to their beta-barrel structure. N-terminal amino acids (with or without modifications) can potentially fit into this "beta-barrel" bucket and be recognized. High-affinity anticalins with engineered novel binding activities have been described (reviewed by Skerra, 2008, FEBS J., 275:2677-2683). For example, anticalins with high affinity (low nM) binding to fluorescein and digoxigenin have been engineered (Gebauer and Skerra, 2012). Engineering alternative scaffolds for new binding functions has also been reviewed by Banta et al. (2013, Annu. Rev. Biomed. Eng. 15:93-113).
[0416] The functional affinity (avidity) of a given monovalent binding agent can be increased by at least an order of magnitude by using bivalent or higher-order multimers of the monovalent binding agent (Vauquelin and Charlton, 2013). Avidity refers to the cumulative strength of multiple simultaneous noncovalent interactions. Individual binding interactions can easily dissociate. However, when multiple binding interactions exist simultaneously, transient dissociation of a single binding interaction does not dissociate the binding protein, and the binding interactions may be restored. An alternative method for increasing the avidity of a binding agent is to include complementary sequences in the coding tag attached to the binding agent and the recording tag associated with the macromolecule.
[0417] In some embodiments, binding agents that selectively bind to modified C-terminal amino acids (CTAAs) can be utilized. Carboxypeptidases are proteases that cleave terminal amino acids containing free carboxyl groups. Several carboxypeptidases exhibit amino acid preferences; for example, carboxypeptidase B preferentially cleaves at basic amino acids such as arginine and lysine. Carboxypeptidases can be engineered to create binding agents that selectively bind to specific amino acids. In some embodiments, carboxypeptidases can be engineered to selectively bind both the modifying moiety and the alpha carbon R group of CTAA. Thus, engineered carboxypeptidases can specifically recognize 20 different CTAAs, representing standard amino acids in conjunction with a C-terminal label. Control of stepwise degradation from the C-terminus of a peptide is achieved by using engineered carboxypeptidases that are active (e.g., binding or catalytic) only in the presence of a label. In one example, CTAA can be modified with a para-nitroanilide group or a 7-amino-4-methylcoumarinyl group.
[0418] Other potential scaffolds that can be engineered to generate binding agents for use in the methods described herein include anticalins, amino acid tRNA synthetases (aaRS), ClpS, Affilin®, Adnectin™, T cell receptors, zinc finger proteins, thioredoxin, GST A1-1, DARPins, affimers, affitins, alphabodies, avimers, Kunitz domain peptides, monobodies, single domain antibodies, EETI-II, HPST1, intrabodies, lipocalins, PHD-finger, V(NAR)LDTI, evibodies, Ig(NAR), knottins, maxibodies, neocarzinostatin, pVIII, tendamistat, VLR, protein A scaffolds, MTI-II, ecotin, GCN4, Im9, Kunitz domains, Examples of suitable antibodies include in, microbody, PBP, trans-body, tetranectin, WW domain, CBM4-2, DX-88, GFP, iMab, LDL receptor domain A, Min-23, PDZ-domain, avian pancreatic polypeptide, charybdotoxin / 10Fn3, domain antibody (Dab), a2p8 ankyrin repeat, insect defense A peptide, designed AR protein, C-type lectin domain, staphylococcal nuclease, Src homology domain 3 (SH3), or Src homology domain 2 (SH2).
[0419] Binding agents can be engineered to withstand higher temperatures and milder denaturing conditions (e.g., the presence of urea, guanidine thiocyanate, ionic solutions, etc.). The use of denaturants helps reduce secondary structure of surface-bound peptides, such as α-helices, β-hairpins, β-strands, and other such structures, which may interfere with binding of the binding agent to linear peptide epitopes. In one embodiment, an ionic liquid such as 1-ethyl-3-methylimidazolium acetate ([EMIM] + [ACE]) is used to reduce peptide secondary structure during binding cycles (Lesch, Heuer, et al., 2015).
[0420] Any of the described binding agents also include coding tags containing identifying information for the binding agent. A coding tag is a nucleic acid molecule of about 3 to about 100 bases that provides unique identifying information for the binding agent to which it is associated. A coding tag can contain about 3 to about 90 bases, about 3 to about 80 bases, about 3 to about 70 bases, about 3 to about 60 bases, about 3 to about 50 bases, about 3 to about 40 bases, about 3 to about 30 bases, about 3 to about 20 bases, about 3 to about 10 bases, or about 3 to about 8 bases. In some embodiments, the coding tag is approximately 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 bases in length. The coding tag may be composed of DNA, RNA, polynucleotide analogs, or combinations thereof. Polynucleotide analogs include PNA, γPNA, BNA, GNA, TNA, LNA, morpholino polynucleotides, 2'-O-methyl polynucleotides, alkylribosyl-substituted polynucleotides, phosphorothioate polynucleotides, and 7-deazapurine analogs.
[0421] The coding tag includes an encoder sequence that provides identifying information about the associated binding agent. The encoder sequence is about 3 to about 30 bases, about 3 to about 20 bases, about 3 to about 10 bases, or about 3 to about 8 bases. In some embodiments, the encoder sequence is about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, or 30 bases in length. The length of the encoder sequence determines the number of unique encoder sequences that can be generated. Shorter coding sequences generate fewer unique coding sequences, which can be useful when using a small number of binding agents. Longer encoder sequences may be desirable when analyzing populations of large molecules. For example, a 5-base encoder sequence has the formula 5'-NNNNN-3' (SEQ ID NO: 135), where N can be any naturally occurring nucleotide or analog. Using the four naturally occurring nucleotides A, T, C, and G, the total number of unique encoder sequences having a length of 5 bases is 1,024. In some embodiments, the total number of unique encoder sequences can be reduced by, for example, removing encoder sequences in which all bases are identical, at least three consecutive bases are identical, or both. In certain embodiments, a set of ≧50 unique encoder sequences is used in the binding agent library.
[0422] In some embodiments, the identifying component of a coding tag or recording tag, such as an encoder sequence, a barcode, a UMI, a compartment tag, a partition barcode, a sample barcode, a spatial region barcode, a cycle-specific sequence, or any combination thereof, is subjected to Hamming distance, Lee distance, asymmetric Lee distance, Reed-Solomon, Levenshtein-Tenengolts, or similar error correction methods. Hamming distance refers to the number of different positions between two links of equal length. Hamming distance measures the minimum number of substitutions required to change one link to another. Hamming distance can be used to correct errors by selecting encoder sequences that are reasonably spaced apart. Thus, in the example where the encoder sequence is 5 bases, the number of usable encoder sequences is reduced to 256 unique encoder sequences (1→4 4 (Hamming distance of the encoder sequence = 256 encoder sequences). In another embodiment, the encoder sequence, barcode, UMI, compartment tag, cycle-specific sequence, or any combination thereof, is designed to be easily read by a cyclic decoding process (Gunderson, 2004, Genome Res. 14:870-7). In another embodiment, the encoder sequence, barcode, UMI, compartment tag, partition barcode, spatial barcode, sample barcode, cycle-specific sequence, or any combination thereof, is designed to be read by nanopore sequencing, which has low accuracy because it requires reading many-base words (approximately 5-20 bases in length) rather than single-base resolution. A subset of 15-mer error-correcting Hamming barcodes that can be used in the methods of the present disclosure are set forth in SEQ ID NOS: 1-65, and their corresponding reverse-complementary sequences are set forth in SEQ ID NOS: 66-130.
[0423] In some embodiments, each unique binding agent in a library of binding agents has a unique encoder sequence. For example, 20 unique encoder sequences can be used for a library of 20 binding agents that bind to 20 standard amino acids. Additional coding tag sequences can be used to identify modified amino acids (e.g., post-translationally modified amino acids). In another example, 30 unique encoder sequences can be used for a library of 30 binding agents that bind to 20 standard amino acids and 10 post-translationally modified amino acids (e.g., phosphorylated amino acids, acetylated amino acids, methylated amino acids). In other embodiments, two or more different binding agents can share the same encoder sequence. For example, two binding agents that each bind to a different standard amino acid can share the same encoder sequence.
[0424] In certain embodiments, the coding tag further comprises a spacer sequence at one or both ends. The spacer sequence is about 1 base to about 20 bases, about 1 base to about 10 bases, about 5 bases to about 9 bases, or about 4 bases to about 8 bases. In some embodiments, the spacer is about 1 base, 2 bases, 3 bases, 4 bases, 5 bases, 6 bases, 7 bases, 8 bases, 9 bases, 10 bases, 11 bases, 12 bases, 13 bases, 14 bases, 15 bases, or 20 bases in length. In some embodiments, the spacer in the coding tag is shorter than the encoder sequence, e.g., at least 1 base, 2 bases, 3 bases, 4 bases, 5 bases, 6 bases, 7 bases, 8 bases, 9 bases, 10 bases, 11 bases, 12 bases, 13 bases, 14 bases, 15 bases, 20 bases, or 25 bases shorter than the encoder sequence. In other embodiments, the spacer in the coding tag is the same length as the encoder sequence. In certain embodiments, the spacer is binder-specific, such that a spacer from a previous binding cycle only interacts with a spacer from the appropriate binder in the current binding cycle. An example is a pair of related antibodies containing a spacer sequence that allows information transfer only if both antibodies sequentially bind to the macromolecule. The spacer sequence can be used as a primer annealing site for primer extension reactions or as a splint or sticky end in ligation reactions. The 5' spacer on the coding tag (see Figure 5A, "*Sp'") can be a T m To increase the yield, one may optionally include pseudo-complementary bases to the 3' spacer on the recording tag (Lehoud et al., 2008, Nucleic Acids Res. 36:3409-3419).
[0425] In some embodiments, coding tags within a collection of binding agents share a common spacer sequence used in the assay (e.g., an entire library of binding agents used in a multiple binding cycle method has a common spacer in their coding tags). In another embodiment, the coding tags are comprised of binding cycle tags that identify a specific binding cycle. In other embodiments, coding tags within a library of binding agents have binding cycle-specific spacer sequences. In some embodiments, coding tags comprise one binding cycle-specific spacer sequence. For example, the coding tag for a binding agent used in the first binding cycle comprises a "cycle 1"-specific spacer sequence, the coding tag for a binding agent used in the second binding cycle comprises a "cycle 2"-specific spacer sequence, and so on for "n" binding cycles. In a further embodiment, the coding tag for a binding agent used in the first binding cycle comprises a "cycle 1"-specific spacer sequence and a "cycle 2"-specific spacer sequence, and the coding tag for a binding agent used in the second binding cycle comprises a "cycle 2"-specific spacer sequence and a "cycle 3"-specific spacer sequence, and so on for "n" binding cycles. This embodiment is useful for PCR assembly of unconcatenated extended recording tags after binding cycles are completed (see Figure 10). In some embodiments, the spacer sequence comprises a sufficient number of bases to anneal with a complementary spacer sequence within the recording tag or extended recording tag to initiate a primer extension or sticky end ligation reaction.
[0426] When a population of recording tags is associated with a macromolecule, cycle-specific spacer sequences can be used to concatenate the information from the coding tags onto a single recording tag. Information can be transferred from the coding tag to a randomly selected recording tag in the first binding cycle, and then cycle-dependent spacer sequences can be used to prime only the extended recording tag in subsequent binding cycles. More specifically, the coding tag for the binding substance used in the first binding cycle includes a "cycle 1"-specific spacer sequence and a "cycle 2"-specific spacer sequence, while the coding tag for the binding substance used in the second binding cycle includes a "cycle 2"-specific spacer sequence and a "cycle 3"-specific spacer sequence, and so on up to "n" binding cycles. The coding tag of the binding substance from the first binding cycle can anneal to the recording tag via the complementary cycle 1-specific spacer sequence. Once the coding tag information has been transferred to the recording tag, at the end of binding cycle 1, a cycle 2-specific spacer sequence is located at the 3' end of the extended recording tag. The coding tag of the binding agent from the second binding cycle can anneal to the extended recording tag via the complementary cycle 2-specific spacer sequence. Once the coding tag information is transferred to the extended recording tag, at the end of binding cycle 2, a cycle 3-specific spacer sequence is located at the 3' end of the extended recording tag, and so on up to "n" binding cycles. This embodiment assumes that the transfer of binding information in a particular binding cycle among multiple binding cycles occurs only on the (extended) recording tag that has undergone the previous binding cycle. However, sometimes binding agents fail to bind to cognate macromolecules. An oligonucleotide containing a binding cycle-specific spacer can be used after each binding cycle as a "tracking" step to keep the binding cycles synchronized even if binding cycle events fail. For example, if a cognate binding agent fails to bind to the macromolecules during binding cycle 1, a tracking step can be added after binding cycle 1 using an oligonucleotide containing both a cycle 1-specific spacer, a cycle 2-specific spacer, and a "null" encoder sequence.A "null" encoder sequence can be the absence of an encoder sequence, or preferably, a specific barcode that positively identifies a "null" binding cycle. The "null" oligonucleotide can anneal to the recording tag via the cycle 1-specific spacer, and the cycle 2-specific spacer is transferred to the recording tag. Thus, despite a failed binding cycle 1 event, binding entities from binding cycle 2 can anneal to the extended recording tag via the cycle 2-specific spacer. The "null" oligonucleotide marks binding cycle 1 as a failed binding event within the extended recording tag.
[0427] In a preferred embodiment, a binding cycle-specific encoder sequence is used in the coding tag. The binding cycle-specific encoder sequence can be achieved either by using a completely unique analyte (e.g., NTAA)-binding cycle encoder barcode or by the combined use of an analyte (e.g., NTAA) encoder sequence and a cycle-specific barcode (see Figure 35). The advantage of using a combinatorial approach is that fewer total barcodes need to be designed. For a set of 20 analyte-binding substances used over 10 cycles, only 20 analyte encoder sequence barcodes and 10 binding cycle-specific barcodes need to be designed. In contrast, if the binding cycle is directly embedded in the binding substance encoder sequence, a total of 200 independent encoder barcodes may need to be designed. The advantage of directly embedding the binding cycle information in the encoder sequence is that the total length of the coding tag can be minimized when using error-correcting barcodes for nanopore reading. The use of error-tolerant barcodes allows for highly accurate barcode identification using sequencing platforms and more error-prone techniques, but offers other advantages such as rapid analysis speed, low cost, and / or more portable instrumentation. One such example is nanopore-based sequencing readout.
[0428] In some embodiments, the coding tag comprises a cleavable or nickable DNA strand within a second (3') spacer sequence proximal to the binding agent (see Figure 32). For example, the 3' spacer may have one or more uracil bases that can be nicked by a uracil-specific excision reagent (USER), which creates a one-nucleotide gap at the uracil position. In another example, the 3' spacer may contain a recognition sequence for a nicking endonuclease that hydrolyzes only one strand of the duplex. Preferably, the enzyme used to cleave or nick the 3' spacer sequence acts on only one DNA strand (the 3' spacer of the coding tag), thus leaving the other strand of the duplex belonging to the (extended) recording tag intact. These embodiments are particularly useful in assays in which proteins are analyzed in their native conformation, as they allow the binding agent to be removed from the (extended) recording tag without denaturation after primer extension has occurred, leaving a single-stranded DNA spacer sequence on the extended recording tag available for subsequent binding cycles.
[0429] Coding tags can also be designed to contain palindromic sequences. The inclusion of palindromic sequences in the coding tag allows the nascent, growing extended recording tag to fold back on itself as the coding tag information is transferred. The extended recording tag folds into a more compact structure, effectively reducing undesired intermolecular binding and primer extension events.
[0430] In some embodiments, the coding tag contains an analyte-specific spacer that allows priming extension only on recording tags previously extended using a binding substance that recognizes the same analyte. Extended recording tags can be assembled from a series of binding events using a coding tag comprising an analyte-specific spacer and an encoder sequence. In one embodiment, the first binding event uses a binding substance with a coding tag consisting of a general 3' spacer primer sequence and an analyte-specific spacer sequence at the 5' end for use in the next binding cycle; then, subsequent binding cycles use binding substances with the encoded analyte-specific 3' spacer sequence. This design results in amplifiable library elements created only from the correct series of cognate binding events. Off-target and cross-reactive binding interactions lead to non-amplifiable extended recording tags. In one example, a pair of cognate binding substances and specific macromolecular analytes is used in ...
Claims
1. (a) a solid support comprising at least 1,000,000 peptide molecules covalently attached to the solid support; (i) each of the at least 1,000,000 peptide molecules is associated with a nucleic acid tag; and (ii) a solid support, wherein adjacent peptide molecules on the solid support are separated from each other by an average distance of 50 nm or more on the surface or within the volume of the solid support; and (b) a plurality of binding agents, each binding to one of the peptide molecules, the plurality of binding agents comprising at least three different binding agents, each of the binding agents comprising a nucleic acid code tag comprising an encoder sequence containing identifying information for the binding agent; A composition comprising: The composition, wherein the nucleic acid coded tag of at least one binding substance that binds to one peptide molecule among the at least 1,000,000 peptide molecules binds to the nucleic acid tag associated with the peptide molecule by hybridizing to a complementary sequence of the nucleic acid tag or by forming a covalent bond with the nucleic acid tag by enzymatic or chemical ligation.
2. The composition of claim 1 , wherein each peptide molecule is covalently bound to the nucleic acid tag associated with each peptide molecule.
3. The composition of claim 1 , wherein the at least 1,000,000 peptide molecules are randomly bound to the solid support.
4. The composition of claim 1 , wherein each peptide molecule is covalently attached to the solid support via the nucleic acid tag.
5. The composition of claim 1 , wherein the nucleic acid tag is covalently attached to the C-terminal amino acid residue of the peptide molecule.
6. The composition of claim 1 , wherein the solid support is a bead, microsphere, or microparticle.
7. The composition of claim 1 , wherein the nucleic acid tag associated with each peptide molecule comprises a barcode or unique molecular identifier (UMI).
8. 10. The composition of claim 1, wherein each of the at least three different binding agents is selective for a different peptide feature or component.
9. 9. The composition of claim 8, wherein the peptide feature or component is selected from the group consisting of an N-terminal amino acid residue, a C-terminal amino acid residue, a dipeptide sequence, a tripeptide sequence, a modified N-terminal amino acid residue, a modified C-terminal amino acid residue, and a post-translationally modified amino acid residue of the peptide molecule.
10. The composition of claim 1, wherein the peptide molecules covalently attached to the solid support each contain an N-terminal amino acid residue modified by a chemical agent, and the modified N-terminal amino acid is bound by a cognate NTAA binding substance.
11. The composition of claim 1 , wherein each of the at least three different binding agents comprises an aptamer, a polypeptide, or a protein.
12. The composition of claim 1 , wherein each of the at least three different binding substances binds to the N-terminal amino acid residue or the C-terminal amino acid residue of the peptide molecule.
13. 2. The composition of claim 1, wherein the at least 1,000,000 peptide molecules covalently attached to the solid support are obtained from a plurality of different biological samples, and each of the associated nucleic acid tags comprises a sample-specific barcode.
Citation Information
Patent Citations
Enzymatic encoding methods for efficient synthesis of large libraries
WO2007062664A2
Systems and methods for barcoding nucleic acids
WO2015164212A1
Method for analyzing nucleic acid derived from single cell
WO2015166768A1