Nanopore-based sequencing of peptides
Patent Information
- Authority / Receiving Office
- AU · AU
- Patent Type
- Applications
- Current Assignee / Owner
- GLYPHIC BIOTECHNOLOGIES INC
- Filing Date
- 2025-01-30
- Publication Date
- 2026-07-30
AI Technical Summary
Current technologies for studying proteins face challenges in selectivity, sensitivity, and throughput, particularly due to the folded nature of proteins and the simultaneous entry of amino acids into nanopores, leading to superimposed signals that are difficult to resolve.
The method involves attaching polymerizable molecules to amino acids, generating modified amino acids through an intramolecular expansion process, and sequencing these modified amino acids using nanopores or nanogaps to achieve high-throughput, high-accuracy protein sequencing.
This approach enables high-throughput single-molecule protein sequencing with an average read accuracy greater than 80% for at least three different modified amino acid types, allowing for precise identification of individual amino acids in peptides.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
NANOPORE-BASED SEQUENCING OF PEPTIDESCROSS REFERENCE
[0001] This application claims benefit of U.S. Provisional Patent App. No. 63 / 694,299, filed September 13, 2024, U.S. Provisional Patent App. No. 63 / 683,941, filed August 16, 2024, U.S. Provisional Patent App. No. 63 / 558,344, filed February 27, 2024, and U.S. Provisional Patent App. No. 63 / 627,214, filed January 31, 2024, each of which applications is incorporated by reference herein in its entirety.BACKGROUND
[0002] Technological improvements in the analysis and characterization of biological molecules have proven to be critical in understanding biological and pathological mechanisms, which has implications in disease diagnosis and modeling, development of therapeutics and treatment, and improving health outcomes. Among these technological improvements, nucleic acid sequencing has emerged as an important tool for genomic and transcriptomic analysis of biological samples.
[0003] Protein signaling underpins a variety of cellular processes and serve important functions in viruses, cells, and living organisms. However, current technologies for studying proteins are limited in selectivity, sensitivity, throughput, or require a priori knowledge. As such, new approaches for characterizing and analyzing proteins is needed.SUMMARY
[0004] Recognized herein is a need for technologies for studying proteins de novo, with improved accuracy and in a high-throughput format. The use of nanopores in protein or peptide sequencing through conventional methods faces challenges. Peptide or protein translocation through a nanopore and subsequent readout is hindered by protein characteristics: proteins are folded and are not uniformly charged, and the tight intramolecular spacing of amino acids means many amino acids enter the nanopore simultaneously, leading to superimposed signals that are difficult to resolve from one another. Provided herein are systems, compositions, kits, and methods for analyzing proteins that address the abovementioned needs. A method of the present disclosure may comprise attaching a plurality of polymerizable molecules to amino acids of a peptide, thereby generating modified amino acids. The modified amino acids may be generated via an intramolecular expansion process. One or more processes described herein may involvesequencingvia a nanopore or nanogap sequencer, allowing for identification of individual amino acids of the peptide in the order in which they appear or occur in the peptide. The methods, systems, compositions, and kits provided herein enable high-throughput single-molecule protein sequencing with high accuracy.
[0005] In an aspect, provided herein is a method for sequencing a peptide comprising a plurality of amino acids, comprising: (a) providing a plurality of modified amino acids generated from at least a sub set of the plurality of amino acids, wherein the plurality of modified amino acids comprises a plurality of polymerizable molecules; and (b) sequencing the plurality of modified amino acids, thereby determining an amino acid identity of each modified amino acid of the plurality of modified amino acids; wherein the sequencing has an average read accuracy that is greater than 80% for at least 3 different modified amino acid types.
[0006] In an aspect, provided herein is a method for sequencing a peptide at sub-attomole resolution comprising a plurality of amino acids, comprising: (a) providing a plurality of modified amino acids generated from atleast a subset of the plurality of amino acids, wherein the plurality of modified amino acids comprises a plurality of polymerizable molecules; and (b) sequencing the plurality of modified amino acids, thereby determining an amino acid identity of each modified amino acid of the plurality of modified amino acids; wherein the sequencing has an individual identification accuracy that is greater than 80% for at least 3 different modified amino acid types.
[0007] In some embodiments, the plurality of polymerizable molecules is covalently linked.
[0008] In some embodiments, the plurality of polymerizable molecules comprises a plurality of nucleic acid molecules. In some embodiments, the plurality of nucleic acid molecules comprises partially double-stranded DNA molecules. In some embodiments, the plurality of nucleic acid molecules comprises branched nucleic acid molecules.
[0009] In some embodiments, the plurality of modified amino acids comprises a modified amino acid that comprises a non-naturally occurring chemical modification. In some embodiments, the non-naturally occurring chemical modification is a protecting group. In some embodiments, the non-naturally occurring chemical modification comprises phenylisothiocyanate, a xanthate, a guanidinylating agent, a dithioester, or a thiocarbamoyl. In some embodiments, the non-naturally occurring chemical modification comprises a click chemistry moiety.
[0010] In some embodiments, the sequencing is performed using a nanopore or a nanogap. In some embodiments, the nanopore comprises a transmembrane protein. In some embodiments, the nanogap comprises an inorganic material. In some embodiments, (b) comprises translocating the plurality of modified amino acids through or adjacent to the nanopore or the nanogap; measuringa plurality of signals from the plurality of modified amino acids; and using the plurality of signals to determine the amino acid identity of each modified amino acid. In some embodiments, the sequencing is perf ormed under conditions sufficient to reduce a translocation speed of the plurality of modified amino acids through the nanopore as compared to a control condition. In some embodiments, the conditions sufficient to reduce the translocation speed comprise an increased viscosity, an addition of a radical agent, an addition of a DNA-binding peptide, an addition of ATP inhibitors, an addition of metal ions, an addition of an intercalating dye, a decreased temperature, an altered pH, an altered voltage, an addition of a strand displacing enzyme, replacement or addition of a helicase, addition of replication protein A, addition of a DNA repair or replication protein, addition of a chromatin remodeling protein, or addition of a denaturant, or a combination thereof, as compared to the control condition. In some embodiments, the conditions sufficient to reduce said translocation speed comprise an increased viscosity, an addition of a radical agent, an addition of a DNA-binding peptide, an addition of ATP inhibitors, an addition of metal ions, an addition of an intercalating dye, a decreased temperature, an altered pH, an altered voltage, an addition of a strand displacing enzyme, replacement or addition of a helicase, addition of replication protein A, addition of a denaturant, or a combination thereof as compared to the control condition.
[0011] In some embodiments, the method further comprises, prior to (a), generating the plurality of modified amino acids. In some embodiments, the generating comprises (I) providing a linker and a polymerizable molecule, (II) coupling the linker and the polymerizable molecule to an amino acid of a peptide, thereby generating an amino acid-linker complex, and (III) cleaving the amino acid, thereby generating the modified amino acid, wherein the modified amino acid comprises a cleaved amino acid, the linker, and the polymerizable molecule. In some embodiments, (II) comprises coupling the linker to (i) the amino acid of the peptide and (ii) the polymerizable molecule. In some embodiments, the linker is pre-coupled to the polymerizable molecule. In some embodiments, the method further comprises contacting the modified amino acid with a helper molecule, wherein the helper molecule comprises a charged moiety, a chelator, or a hydrophobic or hydrophilic moiety. In some embodiments, the method further comprises repeating (I)-(III). In some embodiments, the method further comprises derivatizingthe modified amino acid. In some embodiments, the method further comprises (IV) coupling the amino acidlinker complex to a capture moiety. In some embodiments, the capture moiety is coupled to a substrate. In some embodiments, (I), (II), (III), (IV) ora combination thereofare performed on the substrate. In some embodiments, (I), (II), (III), (IV) or a combination thereof are performed apart from the substrate. In some embodiments, the capture moiety comprises a cleavable moiety andfurther comprising, cleaving the cleavable moiety. In some embodiments, the cleaving occurs subsequent to (II) and prior to (III). In some embodiments, the method further comprises, (V) removing the amino acid-linker complex from the substrate. In some embodiments, the removing comprises an enzymatic digestion, heat denaturation, or toehold-mediated strand displacement. In some embodiments, the capture moiety is coupled to the substrate via an anchor molecule. In some embodiments, the anchor molecule or the capture moiety comprises a PEG linker. In some embodiments, the substrate comprises a plurality of anchor molecules configured to couple to the capture moiety, wherein an average distance between the plurality of anchor molecules is greater than 100 nanometers. In some embodiments, the anchor molecule comprises a peptide nucleic acid (PNA). In some embodiments, the capture moiety is coupled to the peptide. In some embodiments, the capture moiety is coupledto a C-terminus of the peptide. In some embodiments, the capture moiety comprises a nucleic acid barcode molecule. In some embodiments, the nucleic acid barcode molecule identifies the peptide. In some embodiments, the nucleic acid barcode molecule comprises temporal information. In some embodiments, (IV) is performed priorto (III). In some embodiments, the method further comprises repeating (I)-(IV) on the peptide. In some embodiments, the repeating yields a stacked plurality of modified amino acids. In some embodiments, the repeating yields a plurality of detectable products, wherein the plurality of detectable products comprises a plurality of modified amino acids that are not coupled to one another. In some embodiments, the method further comprises concatemerizing at least a portion of the detectable products, thereby generating a stacked plurality of modified amino acids. In some embodiments, the method further comprises cleaving the modified amino acid from the capture moiety. In some embodiments, the capture moiety comprises a cleavable moiety. In some embodiments, the polymerizable molecule and the capture moiety comprise nucleic acid molecules. In some embodiments, the coupling of (IV) comprises hybridization. In some embodiments, the hybridization is performed using a splint oligonucleotide. In some embodiments, splint oligonucleotide comprises a hairpin nucleic acid molecule. In some embodiments, the capture moiety or the polymerizable molecule comprises a hairpin nucleic acid molecule. In some embodiments, the coupling of (IV) comprises ligation. In some embodiments, the capture moiety or the polymerizable molecule comprises a hairpin nucleic acid molecule. In some embodiments, the ligation is performed using a splint oligonucleotide. In some embodiments, the splint oligonucleotide comprises a hairpin nucleic acid molecule. In some embodiments, the linker comprises an isothiocyanate, a xanthate, a guanidinylating agent, a dithioester, or a thiocarbamoyl. In some embodiments, the linker is coupled to the polymerizable molecule. In some embodiments, the linker is coupled to the polymerizable molecule using clickchemistry. In some embodiments, the polymerizable molecule is a DNA molecule comprising a click chemistry moiety. In some embodiments, the click chemistry moiety is coupled to a nucleobase of the DNA molecule. In some embodiments, the click chemistry moiety is coupled to a backbone of the DNA molecule. In some embodiments, the amino acid is a terminal amino acid. In some embodiments, the linker comprises a charged moiety.
[0012] In another aspect, disclosed herein is a method for sequencing a peptide comprising a plurality of amino acids, comprising: (a) providing the peptide; and (b) sequencing the peptide, thereby identifying at least 2 contiguous amino acids of the peptide; wherein the sequencing has an average read accuracy that is greater than 80% for at least 3 different amino acid types.
[0013] In some embodiments, (b) comprises generating a plurality of modified amino acids from the peptide and identifying the plurality of modified amino acids. In some embodiments, the generating comprises (I) providing, a linker and a polymerizable molecule, (II) couplingthe linker to (i) an amino acid of a peptide and (ii) the polymerizable molecule to generate an amino acidlinker complex, and (III) cleaving the amino acid, thereby generating the modified amino acid, wherein the modified amino acid comprises a cleaved amino acid, the linker, and the polymerizable molecule. In some embodiments, the method further comprises contacting the modified amino acid with a helper molecule, wherein the helper molecule comprises a charged moiety, a chelator, or a hydrophobic or hydrophilic moiety. In some embodiments, the linker comprises a charged moiety. In some embodiments, the method further comprises, repeating (I)- (III). In some embodiments, the method further comprises, derivatizingthe modified amino acid. In some embodiments, the method further comprises (IV) couplingthe amino acid-linker complex to a capture moiety. In some embodiments, the capture moiety is coupled to a substrate. In some embodiments, (I), (II), (III), (IV) or a combination thereof are performed on the substrate. In some embodiments, (I), (II), (III), (IV) or a combination thereof are performed apartfromthe substrate. In some embodiments, the capture moiety comprises a cleavable moiety and further comprising cleaving the cleavable moiety. In some embodiments, the cleaving occurs subsequent to (II) and prior to (III). In some embodiments, the method further comprises, (V) removing the amino acidlinker complex from the substrate. In some embodiments, the removing comprises an enzymatic digestion, heat denaturation, ortoehold-mediated strand displacement. In some embodiments, the capture moiety is coupled to the substrate via an anchor molecule. In some embodiments, the anchor molecule or the capture moiety comprises a PEG linker. In some embodiments, the substrate comprises a plurality of anchor molecules configured to couple to the capture moiety, wherein an average distance between the plurality of anchor molecules is greater than 100 nanometers. In some embodiments, the anchor molecule comprises a peptide nucleic acid (PNA).In some embodiments, the capture moiety is coupled to the peptide. In some embodiments, the capture moiety is coupled to a C-terminus of the peptide. In some embodiments, the capture moiety comprises a nucleic acid barcode molecule. In some embodiments, the nucleic acid barcode molecule identifiesthe peptide. In some embodiments, the nucleic acid barcode molecule comprises temporal information. In some embodiments, (IV) is performed prior to (III). In some embodiments, the method further comprises repeating (I)-(IV) on the peptide. In some embodiments, the repeating yields a stacked plurality of modified amino acids. In some embodiments, the repeating yields a plurality of detectable products, wherein the plurality of detectable products comprises a plurality of modified amino acids that are not coupled to one another. In some embodiments, the method further comprises concatemerizing at least a portion of the detectable products, thereby generating a stacked plurality of modified amino acids. In some embodiments, the method further comprises cleaving the modified amino acid from the capture moiety. In some embodiments, the capture moiety comprises a cleavable moiety. In some embodiments, the polymerizable molecule and the capture moiety comprise nucleic acid molecules. In some embodiments, the coupling of (IV) comprises hybridization. In some embodiments, the hybridization is performed using a splint oligonucleotide. In some embodiments, splint oligonucleotide comprises a hairpin nucleic acid molecule. In some embodiments, the capture moiety or the polymerizable molecule comprises a hairpin nucleic acid molecule. In some embodiments, the coupling of (IV) comprises ligation. In some embodiments, the capture moiety or the polymerizable molecule comprises a hairpin nucleic acid molecule. In some embodiments, the ligation is performed using a splint oligonucleotide. In some embodiments, the splint oligonucleotide comprises a hairpin nucleic acid molecule. In some embodiments, the linker comprises an isothiocyanate, a xanthate, a guanidinylating agent, a dithioester, or a thiocarbamoyl. In some embodiments, the linker is coupled to the polymerizable molecule. In some embodiments, the linker is coupled to the polymerizable molecule using click chemistry. In some embodiments, the polymerizable molecule is a DNA molecule comprising a click chemistry moiety. In some embodiments, the click chemistry moiety is coupled to a nucleobase of the DNA molecule. In some embodiments, the click chemistry moiety is coupled to a backbone of the DNA molecule. In some embodiments, the amino acid is a terminal amino acid. In some embodiments, the linker comprises a charged moiety.
[0014] In another aspect, provided herein is a method for sequencing a peptide, comprising: (a) providing a modified amino acid generated from the peptide, wherein the modified amino acid comprises a polymerizable molecule; (b) translocating the modified amino acid through or adjacent to a nanopore, wherein (b) is performed under conditions sufficient to reduce atranslocation speed of the modified amino acid through the nanopore as compared to a control condition; (c) measuring a signal generated from the modified amino acid during (b); and (d) using the signal generated from the modified amino acid to determine an identity of the modified amino acid.
[0015] In some embodiments, the polymerizable molecule comprises a nucleic acid molecule. In some embodiments, the nucleic acid molecule is a partially double-stranded DNA molecule. In some embodiments, the nucleic acid molecule is a branched nucleic acid molecule. In some embodiments, the nucleic acid molecule comprises a modified base or modified nucleic acid backbone. In some embodiments, the modified base or modified nucleic acid backbone is selected from the group consisting of a locked nucleic acid (LNA), a phosphoramidite, a click chemistry- conjugated base or click chemistry-conjugated sugar backbone, a spacer moiety, and a combination thereof.
[0016] In some embodiments, (b) occurs at ambient temperature.
[0017] In some embodiments, the conditions sufficient to reduce the translocation speed comprise an increased viscosity, an addition of a radical agent, an addition of a DNA-binding peptide, an addition of ATP inhibitors, an addition of metal ions, an addition of an intercalating dye, a decreased temperature, an altered pH, an altered voltage, an addition of a strand displacing enzyme, replacement or addition of a helicase, addition of replication protein A, or addition of a denaturant, as compared to the control condition.
[0018] In some embodiments, the modified amino acid comprises a non-naturally occurring chemical modification. In some embodiments, the non-naturally occurring chemical modification is a protecting group. In some embodiments, the non-naturally occurring chemical modification comprises phenylisothiocyanate, a xanthate, a guanidinylating agent, a dithioester, or a thiocarbamoyl. In some embodiments, the non-naturally occurring chemical modification comprises a click chemistry moiety.
[0019] In some embodiments, the nanopore comprises a transmembrane protein.
[0020] In some embodiments, the method further comprises, prior to (a), generating the modified amino acid. In some embodiments, the generating comprises (I) providing, a linker and a polymerizable molecule, (II) coupling the linker to (i) an amino acid of the peptide and (ii) the polymerizable molecule to generate an amino acid-linker complex, and (III) cleaving the amino acid, thereby generating the modified amino acid, wherein the modified amino acid comprises a cleaved amino acid, the linker, and the polymerizable molecule. In some embodiments, the method further comprises contactingthe modified amino acid with a helper molecule, wherein the helper molecule comprises a charged moiety, a chelator, or a hydrophobic or hydrophilic moiety.In some embodiments, the linker comprises a charged moiety. In some embodiments, the method further comprises repeating (I)-(III). In some embodiments, the method further comprises derivatizing the modified amino acid. In some embodiments, the method further comprises (IV) coupling the amino acid-linker complex to a capture moiety. In some embodiments, the capture moiety is coupled to a substrate. In some embodiments, (I), (II), (III), (IV) or a combination thereof are performed on the substrate. In some embodiments, (I), (II), (III), (IV) or a combination thereof are performed apart from the substrate. In some embodiments, the capture moiety comprises a cleavable moiety and further comprising, cleaving the cleavable moiety. In some embodiments, the cleavingoccurs subsequentto (II) andpriorto (III). In some embodiments, the method further comprises, (V) removing the amino acid-linker complex from the substrate. In some embodiments, the removing comprises an enzymatic digestion, heat denaturation, or toehold- mediated strand displacement. In some embodiments, the capture moiety is coupled to the substrate via an anchor molecule. In some embodiments, the anchor molecule or the capture moiety comprises a PEG linker. In some embodiments, the substrate comprises a plurality of anchor molecules configured to couple to the capture moiety, wherein an average distance between the plurality of anchor moleculesis greater than 100 nanometers. In some embodiments, the anchor molecule comprises a peptide nucleic acid (PNA). In some embodiments, the capture moiety is coupled to the peptide. In some embodiments, the capture moiety is coupled to a C- terminus of the peptide. In some embodiments, the capture moiety comprises a nucleic acid barcode molecule. In some embodiments, the nucleic acid barcode molecule identifies the peptide. In some embodiments, the nucleic acid barcode molecule comprises temporal information. In some embodiments, (IV) is performed prior to (III). In some embodiments, the method further comprises repeating (I)-(IV) on the peptide. In some embodiments, the repeating yields a stacked plurality of modified amino acids. In some embodiments, the repeating yields a plurality of detectable products, wherein the plurality of detectable products comprises a plurality of modified amino acids that are not coupled to one another. In some embodiments, the method further comprises concatemerizing at least a portion of the detectable products, thereby generating a stacked plurality of modified amino acids. In some embodiments, the method further comprises cleaving the modified amino acid from the capture moiety. In some embodiments, the capture moiety comprises a cleavable moiety. In some embodiments, the polymerizable molecule and the capture moiety comprise nucleic acid molecules. In some embodiments, the coupling of (IV) comprises hybridization. In some embodiments, the hybridization is performed using a splint oligonucleotide. In some embodiments, splint oligonucleotide comprises a hairpin nucleic acid molecule. In some embodiments, the capture moiety or the polymerizable molecule comprises ahairpin nucleic acid molecule. In some embodiments, the coupling of (IV) comprises ligation. In some embodiments, the capture moiety or the polymerizable molecule comprises a hairpin nucleic acid molecule. In some embodiments, the ligation is performed using a splint oligonucleotide. In some embodiments, the splint oligonucleotide comprisesa hairpin nucleic acid molecule. In some embodiments, the linker comprises an isothiocyanate, a xanthate, a guanidinylating agent, a dithioester, or a thiocarbamoyl. In some embodiments, the linker is coupled to the polymerizable molecule. In some embodiments, the linker is coupled to the polymerizable molecule using click chemistry. In some embodiments, the polymerizable molecule is a DNA molecule comprising a click chemistry moiety. In some embodiments, the click chemistry moiety is coupled to a nucleobase of the DNA molecule. In some embodiments, the click chemistry moiety is coupled to a backbone of the DNA molecule. In some embodiments, the amino acid is a terminal amino acid. In some embodiments, the linker comprises a charged moiety.
[0021] In some embodiments, (d) comprises determining a probability that a modified amino acid is an amino acid type or a subset of amino acid types.
[0022] In yet another aspect, disclosedhereinisamethod of processing a peptide, comprising: (a) providing the peptide and a linker, wherein the linker is capable of coupling to an amino acid of the peptide and wherein the linker is coupled to a nucleic acid molecule; (b) coupling the linker to the amino acid of the peptide to generate an amino acid-linker complex; (c) couplingthe nucleic acid molecule to a capture moiety; (d) cleavingthe amino acid from the peptide to yield a modified amino acid comprising a cleaved amino acid, the linker, and the nucleic acid molecule; and (e) performing a nucleic acid extension reaction of the nucleic acid molecule, thereby generating a detectable product comprising the modified amino acid.
[0023] In some embodiments, the method further comprises contacting the modified amino acid or the detectable product with a helper molecule, wherein the helper molecule comprises a charged moiety, a chelator, or a hydrophobic or hydrophilic moiety. In some embodiments, the linker comprises a charged moiety.
[0024] In some embodiments, the modified amino acid comprises a non-naturally occurring chemical modification. In some embodiments, the non-naturally occurring chemical modification is a protecting group. In some embodiments, the non-naturally occurring chemical modification comprises phenylisothiocyanate, a xanthate, a guanidinylating agent, a dithioester, or a thiocarbamoyl. In some embodiments, the non-naturally occurring chemical modification comprises a click chemistry moiety. In some embodiments, the nucleic acid molecule is a branched nucleic acid molecule.
[0025] In some embodiments, the method further comprises sequencing the detectable product. In some embodiments, the sequencing is performed using a nanopore or a nanogap. In some embodiments, the nanopore comprises a transmembrane protein. In some embodiments, the nanogap comprises an inorganic material. In some embodiments, the sequencing comprises translocating the detectable product through or adjacent to the nanopore or the nanogap; measuring a signal from the detectable product; and using the signal to determine an amino acid identity of the detectable product. In some embodiments, the sequencing is performed under conditions sufficient to reduce a translocation speed of the detectable product through the nanopore as compared to a control condition. In some embodiments, the conditions sufficient to reduce the translocation speed comprise an increased viscosity, an addition of a radical agent, an addition of a DNA-binding peptide, an addition of ATP inhibitors, an addition of metal ions, an addition of an intercalating dye, a decreased temperature, an altered pH, an altered voltage, an addition of a strand displacing enzyme, replacement or addition of a helicase, addition of replication protein A, or addition of a denaturant, as compared to the control condition.
[0026] In some embodiments, the capture moiety is coupled to a substrate. In some embodiments, the capture moiety is coupled to the substrate via a PEG linker. In some embodiments, the sub strate comprises a plurality of capture moieties, wherein an average distance between the plurality of capture moieties is greater than 100 nanometers.
[0027] In some embodiments, the capture moiety is coupled to the peptide. In some embodiments, the capture moiety is coupled to a C-terminus of the peptide.
[0028] In some embodiments, the capture moiety comprises a nucleic acid barcode molecule. In some embodiments, the nucleic acid barcode molecule identifies the peptide. In some embodiments, the nucleic acid barcode molecule comprises temporal information.
[0029] In some embodiments, the method further comprises repeating (a)-(e) on the peptide. In some embodiments, the repeating yields a plurality of detectable products that are not coupled to one another.
[0030] In some embodiments, the method further comprises releasingthe modified amino acid from the capture moiety. In some embodiments, the capture moiety comprises a cleavable moiety. In some embodiments, the releasing comprises dehybridization.
[0031] In some embodiments, the capture moiety comprises an additional nucleic acid molecule. In some embodiments, the coupling of (c) comprises hybridization. In some embodiments, the hybridization is performed using a splint oligonucleotide. In some embodiments, the coupling of (c) comprises ligation.
[0032] In some embodiments, the linker comprises an isothiocyanate, a xanthate, a guanidinylating agent, a dithioester, or a thiocarbamoyl. In some embodiments, the linker is coupled to the nucleic acid molecule using click chemistry.
[0033] In some embodiments, the amino acid is a terminal amino acid.
[0034] In yet another aspect, provided herein is a method of generating increased reads of a modified amino acid, comprising (a) providing the modified amino acid, wherein the modified amino acid comprises a polymerizable molecule; (b) translocating the modified amino acid through or adjacent to a nanopore or a nanogap; (c) circularizing the polymerizable molecule thereby generating a circularized, modified amino acid; and (d) translocating the circularized, modified amino acid through or adjacent to the nanopore or the nanogap.
[0035] In some embodiments, the method further comprises, during (b) and (d), measuring a signal generated from the nanopore or nanogap while the modified amino acid translocates through the nanopore or the nanogap. In some embodiments, the method further comprises using the signal to determine an identity of the modified amino acid.
[0036] In some embodiments, the polymerizable molecule comprises a nucleic acid molecule.
[0037] In some embodiments, the modified amino acid comprises a non-naturally occurring chemical modification. In some embodiments, the non-naturally occurring chemical modification is a protecting group. In some embodiments, the non-naturally occurring chemical modification comprises phenylisothiocyanate, a xanthate, a guanidinylating agent, a dithioester, or a thiocarbamoyl. In some embodiments, the non-naturally occurring chemical modification comprises a click chemistry moiety.
[0038] In some embodiments, the nanopore comprises a transmembrane protein. In some embodiments, the nanogap comprises an inorganic material.
[0039] In some embodiments, (b) is performed under conditions sufficient to reduce a translocation speed of the modified amino acid through the nanopore as compared to a control condition. In some embodiments, the conditions sufficient to reduce the translocation speed comprise an increased viscosity, an addition of a radical agent, an addition of a DNA-binding peptide, an addition of ATP inhibitors, an addition of metal ions, an addition of an intercalating dye, a decreased temperature, an altered pH, an altered voltage, an addition of a strand displacing enzyme, replacement or addition of a helicase, addition of replication protein A, or addition of a denaturant, as compared to the control condition.
[0040] In some embodiments, the method further comprises, generating the modified amino acid wherein the generating comprises (I) providing, a linker and a polymerizable molecule, (II) coupling the linker to (i) an amino acid of a peptide and (ii) the polymerizable molecule to generatean amino acid-linker complex, and (III) cleaving the amino acid, thereby generating the modified amino acid, wherein the modified amino acid comprises a cleaved amino acid, the linker, and the polymerizable molecule. In some embodiments, the method further comprises contacting the modified amino acid with a helper molecule, wherein the helper molecule comprises a charged moiety, a chelator, or a hydrophobic or hydrophilic moiety. In some embodiments, the method further comprises repeating (I)-(III) on the peptide.
[0041] In another aspect of the present disclosure, provided herein is a method for sequencing a peptide, comprising: (a) providing a modified amino acid generated from the peptide, wherein the modified amino acid comprises a polymerizable molecule compri sing M onomers, wherein M is a positive integer; (b) translocating the modified amino acid through or adjacent to a nanopore, wherein the translocating comprises ratcheting of a first monomer of the M monomers through or adjacent to the nanopore; (c) during (b), measuring a first state of a first set of N monomers of the M monomers, wherein N < M; wherein the first state is associated with the ratcheting of the first monomer; (d) ratcheting a second monomer of theM monomers through die nanopore; (e) measuring a second state of a second set of N monomers of the M monomers; wherein the second state is associated with the ratcheting of the second monomer; (f) repeating (d)-(e) N-2 times, thereby obtaining N measured states; and (g) using the N measured states to determine an identity' of the modified amino acid.
[0042] In some embodiments, the N measured states correspond to N contiguous monomers of the M monomers. In some embodiments, the N measured states correspond to N monomers of the M monomers, wherein a subset of the N monomers is non-contiguous. In some embodiments, the N measured states correspond to fewer than N monomers. In some embodiments, the N measured states correspond to a non-integer value of monomers. In some embodiments, the measuring comprises measuring an ionic current blockade. In some embodiments, the N measured states represent a conformational or orientational state of theN monomers. In some embodiments, (g) comprises comparing the N measured states to a reference measurement of known modified amino acid types. In some embodiments, the method further comprises, repeating (a)-(g) for a plurality of modified amino acids generated from the peptide, thereby sequencing the peptide.
[0043] In another aspect of the present disclosure, provided herein is a method for sequencing at sub-attomole resolution, a peptide comprising a plurality of amino acids, comprising: (a) providing a plurality of modified amino acids generated from at least a subset of the plurality of amino acids, wherein the plurality of modified amino acids comprises a plurality of polymerizable molecules; and (b) sequencing the plurality of modified amino acids, thereby determining an amino acid identity of each modified amino acid of the plurality of modified amino acids; whereinthe sequencing has an average identification accuracy that is greater than 80% for at least 3 different amino acid types of the 20 canonical proteinogenic amino acid types.
[0044] Another aspect of the present disclosure provides a method for peptide sequencing comprising sequencing a peptide or plurality of peptides at sub-attomole resolution, wherein the sequencing is capable of discriminating all 20 proteinogenic amino acids. In another aspect, provided herein is a method for peptide sequencing, comprising sequencing a peptide or plurality of using a nanopore, wherein the sequencing is capable of discriminating all 20 proteinogenic amino acids.
[0045] In some embodiments, the sequencing is capable of discriminating post-translationally modified amino acids, which may be synthetic or naturally occurring. In some embodiments, (a) comprises generating a plurality of modified amino acids from the peptide and identifying the plurality of modified amino acids. In some embodiments, the generating comprises (I) providing a linker and a polymerizable molecule, (II) coupling the linker to (i) an amino acid of the peptide and (ii) the polymerizable molecule to generate an amino acid-linker complex, and (III) cleaving the amino acid, thereby generating a modified amino acid of the plurality of modified amino acids, wherein the modified amino acid comprises a cleaved amino acid, the linker, and the polymerizable molecule. In some embodiments, the method further comprises contacting the modified amino acid with a helper molecule, wherein the helper molecule comprises a charged moiety, a chelator, or a hydrophobic hydrophilic moiety. In some embodiments, the linker comprises a charged moiety. In some embodiments, the method further comprises repeating (I)- (III). In some embodiments, the method further comprises derivatizing the modified amino acid. In some embodiments, the method further comprises (IV) couplingthe amino acid-linker complex to a capture moiety. In some embodiments, the capture moiety is coupled to a substrate. In some embodiments, the capture moiety is coupled to the substrate via a PEG linker. In some embodiments, (I), (II), (III), (IV) or a combination thereof are performed on the substrate. In some embodiments, (I), (II), (III), (IV) or a combination thereof are performed apartfromthe substrate. In some embodiments, the capture moiety comprises a cleavable moiety and further comprising cleaving the cleavable moiety. In some embodiments, the cleaving occurs subsequent to (II) and prior to (III). In some embodiments, the method further comprises, (V) removing the amino acidlinker complex from the substrate. In some embodiments, the removing comprises an enzymatic digestion, heat denaturation, or toehold-mediated strand displacement. In some embodiments, the capture moiety is coupled to the substrate via an anchor molecule. In some embodiments, the anchor molecule or the capture moiety comprises a PEG linker. In some embodiments, the substrate comprises a plurality of anchor molecules configured to couple to the capture moiety,wherein an average distance between the plurality of anchor molecules is greater than 100 nanometers. In some embodiments, the anchor molecule comprises a peptide nucleic acid (PNA). In some embodiments, the capture moiety is coupled to the peptide. In some embodiments, the capture moiety is coupled to a C-terminus of the peptide. In some embodiments, the capture moiety comprises a nucleic acid barcode molecule. In some embodiments, the nucleic acid barcode molecule identifiesthe peptide. In some embodiments, the nucleic acid barcode molecule comprises temporal information. In some embodiments, (IV) is performed prior to (III). In some embodiments, the method further comprises repeating (I)-(IV) on the peptide. In some embodiments, the repeating yields a stacked plurality of modified amino acids. In some embodiments, the repeating yields a plurality of detectable products, wherein the plurality of detectable products comprises a plurality of modified amino acids that are not coupled to one another. In some embodiments, the method further comprises concatemerizing at least a portion of the detectable products, thereby generating a stacked plurality of modified amino acids. In some embodiments, the method further comprises cleaving the modified amino acid from the capture moiety. In some embodiments, the capture moiety comprises a cleavable moiety. In some embodiments, the polymerizable molecule and the capture moiety comprise nucleic acid molecules. In some embodiments, the coupling of (IV) comprises hybridization. In some embodiments, the hybridization is performed using a splint oligonucleotide. In some embodiments, splint oligonucleotide comprises a hairpin nucleic acid molecule. In some embodiments, the capture moiety or the polymerizable molecule comprises a hairpin nucleic acid molecule. In some embodiments, the coupling of (IV) comprises ligation. In some embodiments, the capture moiety or the polymerizable molecule comprises a hairpin nucleic acid molecule. In some embodiments, the ligation is performed using a splint oligonucleotide. In some embodiments, the splint oligonucleotide comprises a hairpin nucleic acid molecule. In some embodiments, the linker comprises an isothiocyanate, a xanthate, a guanidinylating agent, a dithioester, or a thiocarbamoyl. In some embodiments, the linker is coupled to the polymerizable molecule. In some embodiments, the linker is coupled to the polymerizable molecule using click chemistry. In some embodiments, the polymerizable molecule is a DNA molecule comprising a click chemistry moiety. In some embodiments, the click chemistry moiety is coupled to a nucleobase of the DNA molecule. In some embodiments, the click chemistry moiety is coupled to a backbone of the DNA molecule. In some embodiments, the amino acid is a terminal amino acid. In some embodiments, the linker comprises a charged moiety.
[0046] Another aspect of the present disclosure provides a non-transitory computer readable medium comprising machine executable code that, upon execution by one or more computer processors, implements any of the methods above or elsewhere herein.
[0047] Another aspect of the present disclosure provides a system comprising one or more computer processors and computer memory coupled thereto. The computer memory comprises machine executable code that, upon execution by the one or more computer processors, implements any of the methods above or elsewhere herein.
[0048] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in this art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.INCORPORATION BY REFERENCE
[0049] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also “Figure” and “FIG.” herein), of which:
[0051] FIG. 1 A schematically showsan example workflowforprocessingpolymericanalytes molecules (e.g., peptides) described herein. FIG. IB schematically shows another example workflow for processing polymeric analytes in solution or on a substrate. FIG. 1C schematically shows another example workflow for processing polymeric analytes in solution or on a substrate. FIG. ID schematically shows yet another example workflow for processing polymeric analytesin solution or on a substrate. FIG. IE schematically shows a workflow for analyzing a polymeric analyte. FIG. IF schematically shows contacting a helper molecule with a modified amino acid or stacked plurality of modified amino acids. FIG. 1G schematically shows another workflow for generating a stacked plurality of modified monomers. FIG. 1H schematically shows a workflow for reducing intermolecular crosstalk as described herein.
[0052] FIG. 2A schematically shows an exemplary linker for coupling polymerizable molecules to polymeric analytes. FIG. 2B shows another exemplary linker for coupling polymerizable molecules to polymeric analytes.
[0053] FIG. 3 schematically shows an example modified amino acid, as described herein.
[0054] FIG. 4A schematically shows a workflow for increasing the number of reads of a modified amino acid or stacked plurality of modified amino acids. FIG. 4B schematically shows a workflow for sequencing a stacked plurality of modified amino acids in which the modified amino acids are spatially separated. FIG. 4C schematically shows a workflow for generating a plurality of stacked plurality of modified amino acids. FIG. 4D schematically shows a workflow for increasing the number of reads from a stacked plurality of modified amino acids.
[0055] FIG. 5 schematically shows a computer system described herein.
[0056] FIG. 6A shows example data of a current profile generated stacked pluralities of modified amino acids. FIG. 6B shows example current traces of individual stacked pluralities of modified amino acids.
[0057] FIG. 7 shows example current traces of a stacked plurality of modified amino acids.
[0058] FIG.8 shows example data of classification accuracy of a stacked plurality of modified amino acids.
[0059] FIG. 9 shows example data of classification accuracy of different modified amino acid types.
[0060] FIG. 10 shows example data of classification accuracy as a function of an applied confidence threshold.
[0061] FIG. 11 shows example data of classification accuracy of different modified amino acids comprising post-translationally modifications or other modifications.
[0062] FIG. 12 shows example data of a stacked plurality of five modified amino acids from a peptide.
[0063] FIG. 13 shows example current traces of a stacked plurality of five modified amino acids from a peptide.
[0064] FIG. 14 shows example current trace data from a tripeptide and a modified amino acid.
[0065] FIG. 15 shows an exampleworkflowforprocessingproteins orpeptidesin preparation for peptide sequencing.
[0066] FIG. 16 shows example data of a sample processing approach to conjugate one or more linkers to a processed peptide.
[0067] FIG. 17 shows example data of conjugationof a capture moietyto a processed peptide.
[0068] FIG. 18 shows example data of cleavage of a processed peptide comprising a linker coupled thereto.
[0069] FIG. 19 shows example data of current traces of individual modified amino acids comprising a spacer moiety.
[0070] FIG. 20 shows example data of alternative polymerizable molecules for processing polymeric analytes described herein.
[0071] FIG. 21 shows an example scheme for chemical expansion of a peptide described herein.
[0072] FIG. 22A shows example alternative workflows for processing polymeric analytes. FIG. 22B shows example data from one example alternative workflow. FIG.22C shows example data from another example alternative workflow. FIG. 22D shows example data of yet another example alternative workflow. FIG. 22E shows example data of yet another example alternative workflow.
[0073] FIG. 23 shows an example scheme of reducing intermolecular crosstalk as described herein.DETAILED DESCRIPTION
[0074] While various embodiments of the invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only . Numerous variations, changes, and substitutionsmay occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed.Definitions
[0075] Whenever the term “at least,” “greater than,” or “greater than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “at least,” “greater than” or “greater than or equal to” applies to each of the numerical values in that series ofnumerical values. For example, greater than or equal to 1 , 2, or 3 is equivalent to greater than or equal to 1, greater than or equal to 2, or greater than or equal to 3.
[0076] Whenever the term “no more than,” “less than,” or “less than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “no more than,” “less than,” or “less than or equal to” applies to each of the numerical values in that series of numerical values. For example, less than or equal to 3, 2, or 1 is equivalent to less than or equal to 3, less than or equal to 2, or less than or equal to 1.
[0077] References to “one embodiment,” “an embodiment,” “example embodiment,” “some embodiments,” “certain embodiments,” “various embodiments,” etc., indicate that the embodiment(s) of the disclosed technology so described may include a particular feature, structure, or characteristic, but not every embodiment necessarily includes the particular feature, structure, or characteristic. Further, repeated use of the phrase “in one embodiment” does not necessarily refer to the same embodiment, although it may.
[0078] Ranges may be expressed hereinas from “about” or “approximately” or “substantially” one particular value and / or to “about” or “approximately” or “substantially” another particular value. When such a range is expressed, other exemplary embodiments include from the one particular value and / or to the other particular value. Further, the term “about” means within an acceptable error range for the particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, “about” can mean within an acceptable standard deviation, per the practice in the art. Alternatively, “about” can mean a range of up to ±20%, preferably up to ±10%, more preferably up to ±5%, and more preferably still up to ±1% of a given value. Alternatively, particularly with respect to biological systems or processes, the term can mean within an order of magnitude, preferably within 2-fold, of a value. Where particular values are described in the application and claims, unless otherwise stated, the term “about” is implicit and in this context means within an acceptable error range for the particular value.
[0079] By “comprising” or “containing” or “including” is meant that at least the named compound, element, particle, or method step is present in the composition or article or method, but does not exclude the presence of other compounds, materials, particles, method steps, even if the other such compounds, material, particles, method steps have the same function as what is named.
[0080] Throughout this description, various components may be identified having specific values or parameters, however, these items are provided as exemplary embodiments. Indeed, the exemplary embodiments do not limit the various aspects and concepts of the present disclosure asmany comparable parameters, sizes, ranges, and / orvalues may be implemented. The terms “first,” “second,” and the like, “primary,” “secondary,” and the like, do not denote any order, quantity, or importance, but rather are used to distinguish one element from another.
[0081] As used herein, the term “protein” generally refers to a molecule comprising two or more amino acids joined by a peptidebond. A protein may also be referred to as a “polypeptide”, “oligopeptide”, or “peptide”. A protein can be a naturally occurring molecule, or a synthetic molecule. A protein may include one or more non-natural amino acids, modified amino acids, or non-amino acid linkers. A protein may contain D-amino acid enantiomers, L- amino acid enantiomers or both. Amino acids of a protein may be modified naturally or synthetically, such as by post-translational modifications or by chemical modification. In some circumstances, different proteins may be distinguished from each other based on different genes from which they are expressed in an organism, different primary sequence length or different primary sequence composition. Proteins expressed from the same gene may nonetheless be different proteoforms, for example, being distinguished based on non-identical length, non-identical amino acid sequence or non-identical post-translational modifications. Different proteins can be distinguished based on one or both of gene of origin and proteoform state.
[0082] As used herein, the term “peptide” may refer to any short, single peptide chain. A peptide may be no more than about 100, 95, 90, 85, 80, 75, 70, 65, 60, 55, 50, 45, 40, 35, 30, 25, 20, 15, 10, 5, or less than about5 amino acidsinlength. A peptide may have a known or unknown biological function or activity. Peptides can include natural, synthetic, modified, or degraded proteins or peptides, or a combination thereof. Peptides can include proteinogenic, natural, synthetic, or modified amino acids or amino acid residues, or a combination thereof.
[0083] As used herein, the term “single analyte” may refer to an analyte that is individually manipulated or distinguished from other analytes. A single analyte may comprise a biomolecule or a synthetic molecule. A single analyte may comprise a small molecule. A single analyte can be a single molecule (e.g., a single biomolecule such as a single protein, nucleic acid molecule, affinity reagent, lipid, carbohydrate, metabolite, hapten, small molecule, pharmaceutical compound, nanoparticle, amino acid derivative, synthetic amino acid, etc.), a single complex of two or more molecules (e.g., a multimeric protein having two or more separable subunits, a single protein attached to a nucleic acid molecule or a single protein attached to an affinity reagent), a single particle, or the like. Reference herein to a “single analyte” in the context of a composition, system or method herein does not necessarily exclude application of the composition, system or method to multiple single analytes that are manipulated or distinguished individually, unless indicated contextually or explicitly to the contrary.
[0084] As used herein, “polypeptide” refers to two or more amino acids linked together by a peptide bond. The term “polypeptide” includes proteins that have a C-terminal end and an N- terminal end as generally known in the art and may be synthetic in origin or naturally occurring As used herein “at least a portion of the polypeptide” refers to 2 or more amino acids of the polypeptide. A polypeptide may comprise one or more peptides. Optionally, a portion of the polypeptide includes at least: 1, 5, 10, 20, 30 or 50 amino acids, either consecutive or with gaps, of the complete amino acid sequence of the polypeptide, or the full amino acid sequence of the polypeptide.
[0085] As used herein, “affixed” refers to a connection between a polypeptide and a substrate such that at least a portion of the polypeptide and the substrate are held in physical proximity. The term “affixed” encompasses both an indirect or direct connection and may be reversible or irreversible, for example the connection is optionally a covalent bond or a non-covalent bond.
[0086] As used herein, the term “sample” refers to a collected substance or material that comprises or is suspected to comprise one or more analytes of interest (e.g., biomolecules, e.g, polypeptides). A sample may be modified for purposes such as storage or stability. A sample may be naturally occurring or synthetic. A sample may be processed to separate or remove unwanted fractions or impurities from the analyte(s) of interest. A sample may be enriched or purified. For example, a sample may comprise a fraction of a separation process (e.g., chromatography, fractionation, electrophoresis, etc.). Alternatively, a sample may not be subjected to processing that separates or removes any unwanted fractions or impurities from the analyte(s) of interest. A sample may be obtained from any suitable source or location, including from organisms, cells, tissues, cell preparations, cell-free compositions, the environment (e.g., air, water, dirt, soil, agriculture, soil, dust). A sample may be obtained from an organism or part of an organism, such as from a fluid, tissue, or cell. A sample may include biological and / or non- biological components. As used herein, the terms “biological sample” or “biological source” refer to a sample that is derived from a predominantly biological system or organism, such as one or more viral particles, cells (e.g. individualized cells), organelles (e.g. individualized organelles), tissues, organs, bodily fluids, bone, cartilage, and exoskeleton. Abiological sample may comprise a majority of biological material on a mass basis, excluding the weight of fluid within the sample. Biological samples may comprise one or more proteins, referred to herein as protein samples. Biological samples can be acquired from various sources, e.g. , from a clinical patient sample, such as blood, serum, plasma, Cerebral Spinal Fluid (CSF), saliva, mucosal secretions, urine, lymph, perspiration, vaginal fluid, semen, etc. A biological sample may be processed to purify and retain one or more biomolecules (e.g., proteins, nucleic acids, carbohydrates, lipids, glycoproteins,lipoproteins, metabolites, etc.) from the biological sample. A biological sample (e.g., a protein sample) may be derived from cultured cells, which may be treated or untreated. A biological sample (e.g., a protein sample) can also result from tissue specimens, such as biopsy samples, which may optionally be processed to liberate biomolecules (e.g., proteins) contained therein. Tissue samples may also be derived from in vivo specimens, including fresh, frozen, acute, and fixed tissues. A sample or a biological sample may comprise non-biological molecules, including but not limited to nanoparticles, polymers, haptens, small molecules, chemicals, fluorescent reagents, inert materials, pharmaceuticals, food additives, environmental contaminants, solvents, industrial chemicals, nanomaterials, radioisotopes, by-products from non-biological molecules.
[0087] As used herein, the terms “antibody” and “immunoglobulin” may generally refer to proteins that can recognize and bind to a specific antigen. An antibody or immunoglobulin may refer to an antibody isotype, fragments of antibodies including, but not limited to, Fab, Fv, scFv, vHH, and Fd fragments, chimeric antibodies, humanized antibodies, single-chain antibodies, and fusion proteins including an antigen-binding portion of an antibody and a non-antibody protein. The antibodies may be detectably labeled, e.g., with a fluorophore, radioisotope, enzyme (e.g, a peroxidase), epitope tag, which generates a detectable product, fluorescent protein, nucleic acid barcode sequence, and the like. The antibodies may be further conjugated to other moieties, such as members of specific binding pairs, e.g., biotin (member of biotin-avidin specific binding pair), and the like. Also encompassed by the terms are Fab', Fv, F(ab')2, and other antibody fragments that retain specific binding to antigen. Antibodies may exist in a variety of other forms including for example, Fv, Fab, and (Fab)2, as well asbi-functional (i.e., bi-specific) hybrid antibodies (e.g, Lanzavecchiaet al., Eur. J. Immunol. 17, 105 (1987)) and in single chains (e.g., Huston et al., Proc. Natl. Acad. Sci. U.S.A., 85, 5879-5883 (1988)andBird etal., Science, 242, 423-426(1988), which are incorporated herein by reference). (See, generally, Hood etal., Immunology, Benjamin, N.Y., 2nded. (1984), and Hunkapiller and Hood, Nature, 323, 15-16 (1986), which are herein incorporated by reference).
[0088] “Binding” as used herein generally refers to a covalent or non-covalent interaction between two molecules (referred to herein as “binding partners”, e.g., a substrate and an enzyme or an antibody and an epitope). Bindingbetween binding partners may be specific or non-specific. Binding between binding partners may involve one or more additional molecules (e.g., biomolecules) or enhancer molecules or substrates.
[0089] As used herein, “specifically binds” or “binds specifically” generally refers to an interaction between bindingpartners (e.g., abindingpartnerand a cognate molecule) such thatthe binding partners bind to one another, but do not bind to other molecules that may be present in theenvironment (e.g., in a biological sample, in tissue, in an in vitro assay) under a set of conditions. A specific binding interaction may entail a binding partner that binds to a cognate molecule. The specific binding interaction may entail the binding of the binding partner to its cognate molecule at a significantly or substantially higher level or with greater affinity as compared to the binding of the binding partner to a non-cognate molecule. A specific binding interaction may entail a first binding partner that has greater selectivity of binding to the cognate molecule as compared to a non-cognate molecule.
[0090] The terms “nucleic acid”, “nucleic acid molecule”, “oligonucleotide” and “polynucleotide” may be used interchangeably herein and generally refer to a polymeric form of naturally occurring or synthetic nucleotides, or analogs thereof, of any length. A nucleic acid molecule may comprise one or more deoxyribonucleotides, deoxynucleotide triphosphates, dideoxynucleotide triphosphates, deoxynucleotide hexaphosphates, dideoxynucleotide hexaphosphates, ribonucleotides, hexitol nucleotides, cyclohexane nucleotides, or analogs or combinations thereof. A nucleic acid molecule may comprise, e.g., DNA, RNA, HNA, CeNA, and modified forms thereof. A nucleic acid molecule may comprise nucleotides that are linked by phosphodiester bonds. Anucleic acid molecule may have any two- orthree-dimensional structure, and may perform any function, known or unknown. A nucleic acid molecule may be single stranded, double stranded, or partially double stranded. Non-limiting examples of polynucleotides include a gene, a gene fragment, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, noncoding RNA, small interfering RNA, short hairpin RNA, micro RNA, scaRNA, ribozymes, riboswitches, viral RNA, complementary DNA (cDNA), cosmid DNA, mitochondrial DNA, chromosomal or genomic DNA, viral DNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, control regions, isolated RNA of any sequence, nucleic acid probes, nucleic acid adapters, and primers. The nucleic acid molecule may be linear, circular, or any other geometry. Examples of polynucleotide analogs include but are not limited to xeno nucleic acid (XNA), bridged nucleic acid (BNA), glycol nucleic acid (GNA), hexitol nucleic acid (HNA), cyclohexane nucleic acid (CeNA), peptide nucleic acids (PNAs), yPNAs, morpholino polynucleotides, locked nucleic acids (LNAs), threose nucleic acid (TNA), 2 '-O-Methyl polynucleotides, 2'-O-alkyl ribosyl substituted polynucleotides, phosphorothioate polynucleotides, and boronophosphate polynucleotides. A polynucleotide analog may possess purine or pyrimidine analogs, including for example, 7-deaza purine analogs, 8-halopurine analogs, 5-halopyrimidine analogs, inverted base, or universal base analogs that can pair with any base, including hypoxanthine, nitroazoles, isocarbostyril analogues, azolecarboxamides, and aromatic triazole analogues, orbase analogs with additional functionality, such as a biotin moiety for affinity binding.
[0091] As used herein, the term “amino acid” generally refers to an organic compound that combines to form a protein or peptide. An amino acid generally comprises an amine group, a carboxylic acid group, and a side-chain specific to each amino acid, which serve as a monomeric subunit of a peptide. An amino acid may include the 20 standard, naturally occurring or canonical amino acids as well as non-standard or non-canonical amino acids. The standard, naturally- occurring or canonical amino acids include Alanine (A or Ala), Cysteine (C or Cys), Aspartic Acid (D or Asp), Glutamic Acid (E or Glu), Phenylalanine (F or Phe), Glycine (G or Gly), Histidine (H or His), Isoleucine (I or He), Lysine (K or Lys), Leucine (L or Leu), Methionine (M or Met), Asparagine (N or Asn), Proline (P or Pro), Glutamine (Q or Gin), Arginine (R or Arg), Serine (S or Ser), Threonine (T or Thr), Valine (V or Vai), Tryptophan (W or Trp), and Tyrosine (Y or Tyr). An amino acid may be an L-amino acid or a D-amino acid. Non-standard amino acids may be modified amino acids, amino acid analogs, amino acid mimetics, non-standard proteinogenic amino acids, or non-proteinogenic amino acids that occur naturally or are chemically synthesized. Examples of non-standard amino acids include, but are not limited to, selenocysteine, pyrro lysine, and N-formylmethionine, (3 -amino acids, Homo-amino acids, Proline and Pyruvic acid derivatives, 3 -substituted alanine derivatives, glycine derivatives, ring- substituted phenylalanine and tyrosine derivatives, linear core amino acids, and N-methyl amino acids.
[0092] As used herein, the term “amino acid type” generally refers to one of the standard, naturally-occurring or canonical amino acids, e.g., one member of the group consisting of Alanine (A or Ala), Cysteine (C or Cys), Aspartic Acid (D or Asp), Glutamic Acid (E or Glu), Phenylalanine (F or Phe), Glycine (G or Gly), Histidine (H or His), Isoleucine (I or He), Lysine (K or Lys), Leucine (L or Leu), Methionine (M or Met), Asparagine (N or Asn), Proline (P or Pro), Glutamine (Q or Gin), Arginine (R or Arg), Serine (S or Ser), Threonine (T or Thr), Valine (V or Vai), Tryptophan (W or Trp), Tyrosine (Y or Tyr), derivatives thereof, and modified forms of any of the aforementioned amino acids. The term “amino acid type” may be used herein to distinguish a plurality of amino acids that comprise different side chain groups, rather than a plurality of amino acids that are identical (e.g., different positional amino acids of a single peptide that have the same side chain). An amino acid type may comprise a modified version of one of the standard, naturally-occurring or canonical amino acids e.g., post translational modifications, an epigenetic modification, or chemical or enzymatic modifications. In some instances, an amino acid type can include non-canonical amino acids.
[0093] As used herein, the term “post-translational modification” refers to modifications that occur on a peptide subsequentto translation. A post-translational modification maybe a covalent modification or enzymatic modification. Examples of post-translation modifications include, but are not limited to, acylation, acetylation, alkylation (including methylation), benzoylation, biotinylation, butyrylation, carbamylation, carbonylation, carboxylation, crotonylation, deamidation, deiminiation, dimethylation, diphthamide formation, disulfide bridge formation, eliminylation, flavin attachment, formylation, gamma-carboxylation, glutamylation, glutarylation, glycylation, glycosylation, glypiation, heme C attachment, hydroxylation, hypusine formation, iodination, isoprenylation, lipidation, lipoylation, malonylation, methylation, myristolylation, nitration, oxidation, palmitoylation, pegylation, phosphopantetheinylation, phosphorylation, prenylation, propionylation, pyroglutamate formation, retinylidene Schiff base formation, S-glutathionylation, S-nitrosylation, S-sulfenylation, sulfation, selenation, stearoylation, succinylation, sulfination, trimethylation, ubiquitination, and C-terminal amidation. A post-translational modification includes modifications of the amino terminus and / or the carboxyl terminus of a peptide. Modifications (both naturally occurring and synthetic) of the terminal amino group include, but are not limited to, des-amino, N-lower alkyl, N-di-lower alkyl, and N-acyl modifications, N-terminal cyclization, deamination, oxidation, ubiquitination, SUMOylation, Neddylation, ISGylation, pupylation, eliminylation, biotinylation, lipidation, N- terminal methylation, N-terminal acetylation, N-terminal propionylation, N-terminal butyrylation, N-terminal crotonylation, N-terminal myristoylation, N-terminal palmitoylation, N-terminal stearoylation, andN-terminal benzoylation. Modifications of the terminal carboxy group include, but are not limited to, amide, lower alkyl amide, dialkyl amide, and lower alkyl ester modifications (e.g., wherein lower alkyl is C1-C4 alkyl). A post-translational modification also includes modifications, such as but not limited to those described above, of amino acids falling between the amino and carboxy termini. The term post-translational modification can also include peptide modifications that include one or more detectable labels. A post-translational modification may be naturally occurring or synthetic.
[0094] As used herein, the term “binding agent” refers to a molecule, e.g., a nucleic acid molecule, a peptide, a polypeptide, a protein, carbohydrate, a synthetic molecule, or a small molecule that binds to, associates with, unites with, recognizes, or combines with another molecule. The binding agent may bind to a macromolecule or a component or feature of a macromolecule. A binding agent may form a covalent association or non-covalent association with a molecule, a macromolecule, or a component or feature of a macromolecule. A binding agent may also be a chimeric binding agent, composed of two or more types of molecules, suchas a nucleic acid molecule-peptide chimeric binding agent, a carbohydrate-peptide chimeric binding agent, or a lipid-peptide chimeric binding agent. A binding agent may be a naturally occurring, synthetically produced, or recombinantly expressed molecule. A binding agent may bind to a single monomer or subunit of a macromolecule (e.g., a single amino acid of a peptide) or bind to a plurality of linked subunits of a macromolecule (e.g., a di-peptide, tri-peptide, or higher order peptide of a longer peptide, polypeptide, or protein molecule). A binding agent may bind to a linear molecule or a molecule having a three-dimensional structure (also referred to as conformation). For example, an antibody binding agent may bind to linear peptide, polypeptide, or protein, or bind to a conformational peptide, polypeptide, or protein. A binding agent may bind to anN-terminal peptide, a C-terminal peptide, oraninterveningpeptideof apeptide, polypeptide, or protein molecule. A binding agent may bind to an N-terminal amino acid, C-terminal amino acid, or an intervening amino acid of a peptide molecule. A binding agent may preferably bind to a chemically modified or labeled amino acid over a non-modified or unlabeled amino acid. For example, a binding agent may preferably bind to an amino acid that has been modified with an acetyl moiety, guanyl moiety, dansyl moiety, PTC moiety, DNP moiety, SNP moiety, etc., over an amino acid that does not possess such a moiety. A binding agent may bind to a post- translational modification of a peptide molecule. A binding agent may exhibit selective binding to a component or feature of a macromolecule (e.g., a binding agent may selectively bind to one of the 20 possible natural amino acid residues and bind with very low affinity or not at all to the other 19 natural amino acid residues). A binding agent may exhibit less selective binding, where the binding agent is capable of binding a plurality of components or features of a macromolecule (e.g., a binding agent may bind with similar affinity to two or more different amino acid residues). A binding agent may comprise a tag, which may be coupled to the binding agent via a linker.
[0095] As used herein, the term “linker” generally refers to a molecule or moiety that is involved in joining two or more molecules. A linker may facilitate a covalent or noncovalent interaction of two or more molecules. A linker may be a crosslinker. The linker can be unifunctional, bifunctional, trifunctional, quadrifunctional, or poly functional. A linker can be or comprise a nucleotide, a nucleotide analog, an amino acid, a peptide, a polypeptide, or a nonnucleotide chemical moiety, such as an organic or inorganic compound. A linker may comprise a polymer, such as a polyethylene glycol (PEG), polyethylene, polypropylene, polyvinyl chloride, polystyrene or other organic or inorganic polymer. A linker may comprise one or more reactive ends, e.g., an amine-reactive group, a carboxyl-reactive group, a sulfhydryl-reactive group, a hydroxyl-reactive group, etc. Alternatively, a linker may not comprise a reactive end. In some examples, a linkermay be usedtojoin different molecule types, e.g., different biomolecule typessuch as a peptide with a nucleic acid molecule, a lipid with a peptide, a carbohydrate with a peptide, etc.; non-biomolecule types; or a biomolecule to anon-biomolecule. For example, a linker may be used to join a binding agent with a tag, a tag with a macromolecule (e.g., peptide, nucleic acid molecule), a macromolecule with a solid support, a tag with a solid support, etc. Alinkermay join two molecules via enzymatic reaction or chemistry reaction (e.g., click chemistry). A linker may join more than two molecules, e.g., via enzymatic or chemical reactions. A linker can be relatively linear or non-linear, e.g., cyclic or circularized, branched, polygonal, etc.
[0096] The term “conjugated” asused herein generally refers to a covalent or ionic interaction between two entities, e.g., molecules, compounds, or combinations thereof.
[0097] As used herein, the term “tag” generally refers to a molecule or moiety that is conjugated to a molecule. Atagmay comprise a detectable label, e.g., a fluorophore or fluorescent protein, a radioactive isotope, an enzyme (e.g., a chromogenic or fluorescent protein, proteins that can catalyze chromogenic substrates), a mass tag, a hapten (e.g., biotin, digoxigenin, urushiol, fluorescein), a vibrational or FTIR tag (e.g., alkyne group). A tag may comprise a biomolecule, such as a nucleic acid molecule, a protein, a lipid, a carbohydrate, or a combination thereof. A tag may comprise one or more nucleic acid molecules, which may optionally encode information regarding the tag or the molecule onto which a tag is conjugated (e.g., a binding agent, such as an antibody). For example, a tag may comprise a nucleic acid barcode molecule. A tag may comprise an organic compound or an inorganic compound.
[0098] As used herein, the term “barcode” generally refers to an identifying feature that may be used to distinguish similar items. A barcode may comprise a nucleic acid molecule of about 2 to about 150 bases. A barcode may comprise a nucleic acid molecule of about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150 or more bases, which may provide a unique identifier tag or origin information for a molecule (e.g., protein, polypeptide, peptide), a binding agent, a set of binding agents from a binding cycle, a sample molecule, a set of samples, molecules within a compartment (e.g., droplet, bead, partition or separated location), macromolecules within a set of compartments, a fraction of macromolecules, a set of macromolecule fractions, a spatial region or set of spatial regions, a library of macromolecules, or a library of binding agents. A barcode canbe an artificial sequence or a naturally occurring sequence including peptides, proteins, protein complexes, carbohydrates, and synthetic polymeric materials such as peptoids, polysaccharides, polymers, fluorescent tags,chemical tags, magnetic tags, isobaric tags, Raman spectroscopic tags, quantum dots, etc. In certain embodiments, each barcode within a population of barcodes is different. In other embodiments, a portion of barcodes in a population of barcodes is different, e.g., at least about 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, or 99% of the barcodes in a population of barcodes is different. A population of barcodes may be randomly generated or non-randomly generated. A population of barcodes may comprise error correcting barcodes. Barcodes can be used to computationally deconvolute sequence reads derived from an individual molecule, sample, library, etc. Barcodes may comprise multiplexed information, e.g., arising from different samples, compartments, individual molecules, etc. A barcode can also be used for deconvolution of a collection of molecules that have been distributed into small compartments for enhanced mapping. For example, rather than mapping a peptide back to the proteome, the peptide can be mapped back to its originating protein molecule or protein complex, a sample or partition from which it originated, etc. A barcode may comprise any useful sequence, including repeat sequences (e.g., a poly -A, poly-T, poly-C, poly- G region) or the barcode may comprise non-repeat sequences. A barcode may encode for information including, but not limited to time, lineage, sample types, cell number, beads, single molecule information, meta data, space / location (e.g., a slide, well, tissue), proximity (e.g., to other molecules, cells, metabolites, DNA, RNA), patient info, biological sample information, library information computer data, weather, physical parameters such as temperature, humidity, precipitation.
[0099] As used herein, a “sample barcode”, also referred to as “sample tag” generally refers to a barcode molecule comprising identifying information of a sample from which a barcoded molecule derives.
[0100] As used herein, a “spatial barcode” generally refers to a barcode molecule comprising identifying information of a region of a 2-D or 3-D sample (e.g., a tissue section) from which a molecule originates or is derived. Spatial barcodes may be used for molecular pathology on tissue sections. A spatial barcode may allow for multiplex sequencing of a plurality of samples or libraries from tissue section(s).
[0101] As used herein, a “temporal barcode” generally refers to a barcode molecule comprising time-based information relating to the barcoded molecule. The types of time-based data encoded in a temporal barcode can include information such as a lifetime of a barcoded molecule, a time of collection of a sample, a time or duration since the beginning of an experiment or induction with a stimulus, information on the age of a cell or tissue, a sequence of interactions between molecules, a time or cycle or round (e.g., of an iterative process) in which the barcodemolecule is provided, among others. It is possible for different types of barcodes (e.g., spatial, temporal, cell-specific) to be combined in one multiplexed barcode.
[0102] As used herein, the term “nucleic acid sequence” or “oligonucleotide sequence” generally refers to a contiguous string of nucleotide bases and may refer to the particular placement of nucleotide bases in relation to each other as they appear in an oligonucleotide. Similarly, the term “polypeptide sequence” or “amino acid sequence” refers to a contiguous string of amino acids and may refer to the particular placement of amino acids in relation to each other as they appear in a polypeptide.
[0103] A “nucleotide sequence” according to the present invention may include any polymer or oligomer of nucleotides such as pyrimidine and purine bases, such as cytosine, thymine, and uracil, and adenine and guanine, respectively and combinations thereof. The nucleotide sequence may comprise any deoxyribonucleotide, ribonucleotide, hexitol-nucleotide, cyclohexanenucleotide, peptide nucleic acid component, and any chemical variants thereof, such as methylated, 7-deaza purine analogs, 8-halopurine analogs, hydroxymethylated or glycosylated forms of these bases, and the like. The polymers or oligomers may be heterogeneous or homogenous in composition and may be isolated from naturally occurring sources or may be artificially or synthetically produced. In addition, a nucleotide sequence may be DNA, RNA, HNA, CeNA or a mixture thereof, and may exist permanently or transitionally in single-stranded or double-stranded form, including homoduplex, heteroduplex, and hybrid states.
[0104] The terms “complementary” or “complementarity” refer to polynucleotides (i.e., a sequence of nucleotides) related by base-pairing rules. For example, the sequence “5'-AGT-3',” is complementary to the sequence “5'- ACT-3'”. Complementarity may be “partial,” in which only some of the nucleic acids’ bases are matched according to the base pairing rules, or there may be “complete” or “total” complementarity between the nucleic acids. The degree of complementarity between nucleic acid strands can have significant effects on the efficiency and strength of hybridization between nucleic acid strands under defined conditions.
[0105] As used herein, the term “hybridization” is used in reference to the pairing of complementary nucleic acids. Hybridization and the strength of hybridization (e.g., the strength of the association between the nucleic acids) is influenced by such factors as the degree of complementary between the nucleic acids, stringency of the conditions involved, and the melting temperature of the formed hybrid. Hybridization methods involve the annealing of one nucleic acid to another, complementary nucleic acid, e.g., based on Watson-Crick base pairing.
[0106] As used herein, the term “proteomics” generally refers to quantitative and / or qualitative analysis of the proteome within a sample, such as biological sample, e.g., from cells,tissues, or bodily fluids. Proteomics may include the analysis of spatial distributions of proteins within a sample (e.g., cell and / or tissues). Proteomics may include studies of the dynamic state of the proteome, e.g., how one or more proteins change in time.
[0107] The terminal amino acid at one end of the peptide chain that has a free amino group may be referred to herein as the “N-terminal amino acid” (NTAA). The terminal amino acid at the other end of the chain that has a free carboxyl group may be referred to herein as the “C-terminal amino acid” (CTAA). The amino acids making up a peptide may be numbered in order, with the peptide being “n” amino acids in length. As used herein, in some instances, NTAA may be considered the nth amino acid (also referred to herein as the “n NTAA”). In such cases, the next amino acid is the n- 1 amino acid, then the n-2 amino acid, and so on down the length of the peptide from the N-terminal end to C-terminal end. Alternatively, CTAA may be consideredthe nth amino acid (also referred to herein as the “n CTAA”). In such cases, the next amino acidis the n-1, then the n-2 amino acid, and so on down the length of the peptidefromthe C-terminal end toN-terminal end. An NTAA, CTAA, or both may be modified or labeled with a chemical moiety.
[0108] As used herein, the terms “determining,” “measuring,” “assessing,” and“assaying” are used interchangeably and include both quantitative and qualitative determinations.
[0109] As used herein, the term “unique molecular identifier” or “UMI” generally refers to a molecule barcodecomprisingindexinginformation. AUMImay comprise a nucleic acid molecule of about 3 to about 150 bases (3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22,23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48,49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74,75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91 , 92, 93, 94, 95, 96, 97, 98, 99, 100,105, 110, 115, 120, 125, 130, 135, 140, 145, or 150 bases) in length. A UMI may provide a unique identifier tag for each molecule (e.g., peptide, binding agent, a nucleic acid molecule) that comprises or is coupled to a UMI. A UMI may comprise a random sequence (e.g., a random N- mer).
[0110] As used herein, a “derivative” of a nucleic acid molecule generally refers to a nucleic acid molecule that is derived from an originating nucleic acid molecule. The derivative may have the same or substantially the same nucleotide sequence as the originating nucleic acid molecule, or the derivative may comprise a complement or partial complement as the originating nucleic acid molecule. A derivative may be the same type of nucleic acid (e.g., DNA or RNA) as the originating nucleic acid molecule, or the derivative may be a different type of nucleic acid (e.g, cDNA generated from an RNA molecule). A nucleic acid molecule derivative may display sequence identity as the originating nucleic acid molecule. The derivative nucleic acid moleculemay also be subjected to additional processing from the originating nucleic acid molecule, e.g., chemical or enzymatic modification, splicing, ligation, polymerization, fragmentation, tagmentation (e.g., using a transposase), digestion, etc.
[0111] A derivative polypeptide or peptide may be derived from an originating polypeptide (or peptide). A derivative may comprise the same amino acid sequence as the originating polypeptide, or the sequence may be different. The derivative polypeptide may result from or be subjected to additional processing from the originating polypeptide, e.g., chemical or enzymatic modification. The derivative polypeptide may comprise one or more tags, nucleic acid molecules, barcode molecules, labels (e.g., detectablelabels), fluorophores, probes, linkers, post-translational modifications, chemical protecting groups, or other chemical moieties.
[0112] As used herein, the term “compartment” or “partition” generally refers to a physical area or volume that separates or isolates a subset of molecules from a sample of molecules. For example, a compartment or partition may separate an individual cell from other cells, or a subset of a sample’s proteome from the rest of the sample’s proteome. A compartment or partition may be an aqueous compartment (e.g., microfluidic droplet), a solid compartment (e.g., picotiter well or microtiter well on a plate, tube, vial, gel bead), a liquid-liquid phase separation, a liquid condensate, a sub cellular region, or a separatedregion on a surface. A compartmentmay comprise one or more beads to which macromolecules may be immobilized. A compartment may be transient.
[0113] As used herein, the term “solid support”, “solid surface”, or “solid substrate” or “substrate” refers to any solid material, including porous and non-porous materials, to which a molecule can be associated directly or indirectly. The molecule may be associated with the substrate by covalent or non-covalent interactions, or a combination thereof. A substrate may be two-dimensional (e.g., planar surface) or three-dimensional (e.g., gel matrix or bead). A solid support may comprise, in non-limiting examples, a bead, a microbead, an array, a glass surface, a silicon surface, a plastic surface, a filter, a membrane, nylon or other polymer, a silicon wafer chip, a flow through chip, a flow cell, a microfluidic device or chip or a surface thereof, a biochip including signal transducing electronics, a channel, a microtiter well, an ELISA plate, a spinning interferometry disc, a nitrocellulosemembrane, a nitrocellulose-based polymer surface, a polymer matrix, a nanoparticle, or a microsphere. Materials for a solid support include but are not limited to acrylamide, agarose, cellulose, nitrocellulose, glass, gold, quartz, polystyrene, polyethylene vinyl acetate, polypropylene, polymethacrylate, polyethylene, polyethylene oxide, poly silicates, polycarbonates, Teflon, fluorocarbons, nylon, silicon rubber, polyanhydrides, polyglycolic acid, polylactic acid, polyorthoesters, functionalized silane, polypropylfumerate, collagen,glycosaminoglycans, poly amino acids, dextran, or any combination thereof. Solid supports further include thin film, membrane, bottles, dishes, fibers, woven fibers, shaped polymers such as tubes, (e.g., nanotubes), particles, beads, DNA origami, microspheres, microparticles, or any combination thereof. For example, when solid surface is a bead, the bead can include, but is not limited to, a ceramic bead, polystyrene bead, a polymer bead, a methylstyrene bead, an agarose bead, an acrylamide bead, a solid core bead, a porous bead, a paramagnetic bead, a glass bead, or a controlled pore bead. Ahead may be spherical or an irregularly shaped. Ahead’ s size may range from nanometers, e.g., 1 nm, lO nm, lOO nm, to millimeters, e.g., 1 mm. In certain embodiments, beads range in size from about 0.2 micron to about200microns, or from about 0.5 micron to about 5 microns. In some embodiments, beads can be about 1, 1.5, 2, 2.5, 2.8, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81 , 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 pm in diameter. In certain embodiments, “a bead” solid support may refer to an individual bead or a plurality of beads.
[0114] As used herein, “sequencing” generally refers to determining the order and identity of: (A) nucleotides (base sequences) in a nucleic acid sample, e.g., DNA orRNA; or determining the order and identity of (B) amino acids in all or part of a polymer, such as a protein, peptide, or other multimeric molecule. Many techniques are available for nucleic acid sequencing, such as Sanger sequencing or High Throughput Sequencing technologies (HTS). Sanger sequencing may involve sequencing via detection through (capillary) electrophoresis, in which up to 384 capillaries may be sequence analyzed in one run. High throughput sequencing involves the parallel sequencing of thousands or millions or more sequences at once. HTS can be defined as Next Generation sequencing (NGS), i.e. techniques based on solid phase pyrosequencing or as Next-Next Generation sequencingbased on single nucleotide real time sequencing (SMRT). HTS technologies are available such as offered by Roche, Illumina and Applied Biosystems (Life Technologies). Further high throughput sequencing technologies are described by and / or available from Helicos, Pacific Biosciences, Complete Genomics, Ion Torrent Systems, Oxford Nanopore Technologies, Nabsys, ZS Genetics, GnuBio.
[0115] As used herein, “next generation sequencing” refers to high-throughput sequencing methods that allow the sequencing of millions to billions of molecules in parallel. Examples of next generation sequencing methods include sequencing by synthesis, sequencing by ligation, sequencing by hybridization, polony sequencing, ion semiconductor sequencing, nanoporesequencing, and pyro sequencing. By attaching primers to a solid substrate and a complementary sequence to a nucleic acid molecule, a nucleic acid molecule can be hybridized to the solid substrate via the primer and then multiple copies can be generated in a discrete area on the solid substrate by using polymerase to amplify (these groupings are sometimes referred to as polymerase colonies or polonies). Consequently, during the sequencing process, a nucleotide at a particular position can be sequenced multiple times (e.g., hundreds or thousands of times) — this depth of coverage is referred to as “deep sequencing.” Examples of high throughput nucleic acid sequencing technology include platforms provided by Illumina, BGI, Qiagen, ThermoFisher, and Roche, including formats such as parallel bead arrays, sequencing by synthesis, sequencing by ligation, capillary electrophoresis, electronic microchips, “biochips,” microarrays, parallel microchips, and single-molecule arrays, as reviewedby Service (Science s 11 : 1544-1546, 2006).
[0116] As used herein, “analyzing” the macromolecule means to quantify, characterize, distinguish, or a combination thereof, all or a portion of the components of a molecule (e.g., a macromolecule, a biological molecule such as a protein, amino acid, nucleic acid molecule, etc.). For example, analyzing a peptide, polypeptide, or protein may comprise determining all or a portion of the amino acid sequence (contiguous or non-continuous) of the peptide. Analyzing a macromolecule may include partial identification of a component of the macromolecule. For example, partial identification of amino acids in a protein sequence can identify an amino acid in the protein as belonging to a subset of possible amino acids. Analysis may be performed sequentially, e.g., beginning with analysis of the n NTAA, and then proceeding to the next amino acid of the peptide (i.e., n-1, n-2, n-3, and so forth). In such instances, sequencing may be performed by cleavage ofthen NTAA, thereby converting the n-1 amino acid of the peptide to an N-terminal amino acid (referred to herein as the “n-1 NTAA”). Similarly, analysis of a peptide may begin from C-terminus towards the N-terminus with each round of cleavage from the C- terminus creating a new CTAA. Cleavage of the n CTAA converts the n-1 amino acid of the peptide to a C-terminal amino acid, referred to herein as an “n-1 CTAA”. Analyzing the peptide may also include determining a presence and frequency of post-translational modifications on the peptide, which may or may not include information regarding the sequential order of the post- translational modifications on the peptide. Analyzing the peptide may also include determining the presence and frequency of epitopes in the peptide, which may or may not include information regarding the sequential order or location of the epitopes within the peptide. Analyzingthe peptide may include combining different types of analysis, for example obtaining epitope information, amino acid sequence information, post-translational modificationinformation, or any combination thereof.
[0117] As used herein, the term “array” generally refers to a population of molecules that is attached to one or more solid supports such that the molecules at one address can be distinguished from molecules at other addresses. An array can include different molecules that are each located at different addresses on a solid support. Alternatively, an array can include separate solid supports each functioning as an address that bears a different molecule, wherein the different molecules can be identified according to the locations of the solid supports on a surface to which the solid supports are attached, or according to the locations of the solid supports in a liquid such as a fluid stream. The molecules of the array can be, for example, nucleic acids such as SNAPs, polypeptides, proteins, peptides, oligopeptides, enzymes, ligands, or receptors such as antibodies, functional fragments of antibodies or aptamers. The addresses of an array can optionally be optically observable, and, in some configurations, adjacent addresses can be optically distinguishable when detected using a method or apparatus set forth herein.
[0118] As used herein, the term “functionalized” refers to any material or substance that has been modified to include a functional group. A functionalized material or substance may be naturally or synthetically functionalized. For example, a polypeptide can be naturally functionalized with a phosphate group, oligosaccharide (e.g., glycosyl, glycosylphosphatidylinositol or phosphoglycosyl), nitrosyl, methyl, acetyl, lipid (e.g., glycosyl phosphatidylinositol, myristoyl orprenyl), ubiquitin or other naturally occurringpost-translational modification. A functionalized material or substance may be functionalized for any given purpose, including altering chemical properties (e.g., altering hydrophobicity or changing surface charge density) or altering reactivity (e.g., capable of reactingwith a moiety or reagentto form a covalent bond to the moiety or reagent).
[0119] As used herein, the term “click reaction,” “click chemistry,” or “bioorthogonal reaction” refers to single-step, thermodynamically favorable conjugation reaction utilizing biocompatible reagents. A click reaction may utilize no toxic or biologically incompatible reagents (e.g., acids, bases, heavy metals) or generate no toxic or biologically incompatible byproducts. A click reaction may utilize an aqueous solvent or buffer (e.g., phosphate buffer solution, Tris buffer, saline buffer, MOPS, etc.). A click reaction may be thermodynamically favorable if it has a negative Gibbs free energy of reaction, for example a Gibbs free energy of reaction of less than about -5 kiloJoules / mole (kJ / mol), -10 kJ / mol, -25 kJ / mol, -50 kJ / mol, -100 kJ / mol, -200 kJ / mol, -300 kJ / mol, -400 kJ / mol, or less than -500 kJ / mol. Exemplary bioorthogonal and click reactions are describedin detail in WO 2019 / 195633A1, which is herein incorporated by reference in its entirety. Exemplary click reactions may include metal-catalyzed azide-alkyne cycloaddition, strain-promoted azide-alkyne cycloaddition, strain-promoted azide-nitrone cycloaddition, strained alkene reactions, thiolene reaction, Diels-Alder reaction, inverse electron demand Diels-Alder reaction, [3+2] cycloaddition, [4+1] cycloaddition, nucleophilic substitution, dihydroxylation, thiolyne reaction, photoclick, nitrone dipole cycloaddition, norbomene cycloaddition, oxanobornadiene cycloaddition, tetrazine ligation, and tetrazole photoclick reactions. Exemplary functional groups or reactive handles utilized to perform click reactions (also referred to herein as “click chemistry moieties”) may include alkenes (e.g., linear alkenes or cyclic alkenes such as trans-cyclooctene (TCO)), alkynes (e.g., linear alkynes or cycloalkynes (e.g., cyclooctynes or derivatives thereof, e.g., aza-dimethoxycyclooctyne (DIMAC), symmetrical pyrrolocyclooctyne (SYPCO), pyrrolocyclooctyne (PYRROC), difluorocyclooctyne (DIFO), a,a-bis(trifluoromethyl)pyrrolocyclooctyne(TRIPCO), bicyclo[6.1.0]nonyne (BCN), dibenzocyclooctyne (DIBO), difluorinated cyclooctyne (DIFO), difluorobenzocyclooctyne (DIFBO), dibenzoazacyclo-octyne (DBCO), difluoro-aza- dibenzocyclooctyne (F2-DIBAC), biaryl-azacyclooctynone (BARAC), difluorodimeth oxydib enzocyclooctynol (FMDIBO), difluorodimeth oxydibenzocyclooctynone (keto-FMDIBO), and 3,3,6,6-tetramethylthiacycloheptyne (TMTH)), TMTH-sulf oximine (TMTHSI), azides, epoxides, amines, thiols, nitrones, isonitriles, isocyanides, aziridines, activated esters, and tetrazines, triazoles, and combinations, variations, or derivatives thereof. The click chemistry moieties may be subjected to conditions sufficient to react the first click chemistry moiety to the second click chemistry moiety, e.g., provision of metal catalysts, appropriate solvents, pH, temperature, ionic concentration, or light / energy, for any useful duration of time.
[0120] As used herein, the terms “group” and “moiety” are intended to be synonymous when used in reference to the structure of a molecule. The terms refer to a component or part of the molecule. The terms do not necessarily denote the relative size of the component or part compared to the molecule, unless indicated otherwise. The terms do not necessarily denote the relative size of the component or part compared to any other component or part of the molecule, unless indicated otherwise. A group or moiety can contain one or more atoms.
[0121] As used herein, “primers” generally refer to nucleic acid molecules which can prime the synthesis of a nucleic acid molecule (e.g., DNA or RNA). A primer may be single stranded. A primer may comprise one or more recognition sites for a protein (e.g., a polymerizing enzyme, a restriction enzyme, a cleaving enzyme, etc.) to bind to the primer or a primer hybridized to a template strand. A primer may comprise DNA, RNA, or other nucleic acid analogs or noncanonical bases (e.g., spacer moieties, uracils, abasic sites). A primer may optionally comprise any number of functional sequences such as sequencing primer sequences (e.g., P5 or P7sequences), sequencing primer-binding sequences, read sequences (e.g., R1 or R2 sequences), restriction sites, transposition sites (e.g., mosaic end sequences), etc.
[0122] “Amplification” or “amplifying” generally refers to a polynucleotide amplification reaction, namely, a population of polynucleotides that are replicated from one or more starting sequences. Amplifying may refer to a variety of amplification reactions, including but not limited to polymerase chain reaction (PCR), linear polymerase reactions, nucleic acid sequence- based amplification, rolling circle amplification and similar reactions. An amplification reaction may generate an amplicon. Amplification or amplifying may refer to an increase in quantity of a measurable output, for example, signal amplification.
[0123] An “adapter” as referred to herein, generally refers to a short nucleic acid molecule (e.g., about 10 to about 100 base pairs in length). An adapter may comprise a short double-stranded DNA molecule. An adapter may be attached, e.g., via polymerization or ligation, to an end of a DNA fragments or amplicons. Adapters may comprise synthetic oligonucleotides, e.g., oligonucleotides that have nucleotide sequences which are at least partially complementary to each other. An adapter may have blunt ends, may have staggered ends (also referred to herein as a 3 ’ or 5’ “overhang sequence” or “sticky end”, or a blunt end and a staggered end. Adapters may be attached (e.g., via ligation) to fragments to provide an adapter-ligated fragment; the adapter- ligated fragment may serve as a starting point for subsequent manipulation e.g., for amplification or sequencing. An adapter may be functionalized, e.g., conjugated with a tag, probe, detectable label, affinity capture reagent (e.g., biotin or streptavidin).
[0124] The term “translocating” and “translocation,” as used herein, generally refers to the movement of a molecule through a medium (e.g., a gas, a liquid, a solid, ora multiphase medium). Translocation of a molecule may occur spontaneously (e.g., through diffusion, Brownian motion, etc.). Alternatively, or in addition to, translocation of a molecule may occur with an application of force or pressure, e.g., using frictional force, tension force, a normal force, air resistance force, spring force, a temperature gradient, gravitational force, electrical force, magnetic force, acoustic force (e.g., acoustophoresis) etc. In some examples, translocation of a molecule may be achieved by application of pressure-driven flow or electrophoretic forces. Translocation may occur through a liquid or through a solid or semi-solid substrate (e.g., through a pore or gap) or adjacent or in proximity to the solid or semi-solid substrate.
[0125] As used herein, the abbreviations for the natural 1 -enantiomeric amino acids are conventional and can be as follows: alanine (A, Ala); arginine (R, Arg); asparagine (N, Asn); aspartic acid (D, Asp); cysteine (C, Cys); glutamic acid (E, Glu); glutamine (Q, Gin); glycine (G, Gly); histidine (H, His); isoleucine (I, He); leucine (L, Leu); lysine (K, Lys); methionine (M, Met);phenylalanine (F, Phe); proline (P, Pro); serine (S, Ser); threonine (T, Thr); tryptophan (W, Trp); tyrosine (Y, Tyr); valine (V, Vai). Unless otherwise specified, X can indicate any amino acid. In some aspects, X can be asparagine (N), glutamine (Q), histidine (H), lysine (K), or arginine (R). References to these amino acids are also in the form of “[amino acid] [residues / residues]” (e.g, lysine residue, lysine residues, leucine residue, leucine residues, etc.).
[0126] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure, suitable methods and materials are described below.Protein Sequencing Using Intramolecular Expansion and Nanopore Analysis
[0127] Provided herein are methods, systems, compositions, and kits for characterizing polymeric analytes, such as proteins. The methods, systems, compositions, and kits of the present disclosure provide for the analysis of individual or clusters of monomers of a polymeric analyte, e.g., amino acids of a protein, thereby providing information on the composition or sequence of monomers (e.g., amino acids in a protein, also referred to herein as “protein sequencing”). In some aspects, a method of the present disclosure comprises providing a polymeric analyte, such as a peptide, and determining the identity of the individual (or clusters of) monomers comprised by the polymeric analyte. Alternatively, or in addition to, the present disclosure may provide for methods of identifying a peptide or protein without determining the identity of the individual monomers. In some aspects, the methods, systems, compositions, and kits provided herein entail coupling polymerizable molecules to monomers (e.g., amino acids), thereby generating modified monomers (e.g., modified amino acids), and detecting the modified monomers (e.g., modified amino acids). In some embodiments, the methods provided herein entail repeating one or more operations on a single or plurality of peptides to generate a plurality of modifiedmonomers, which optionally may be tethered together (e.g., in a stacked plurality of modified monomers), and analyzing or detecting the plurality of modified monomers. In some embodiments, the detecting is performed using nanoscale objects (e.g., nanopores or nanogaps). Beneficially, the polymerizable molecule may alter or modulate a property of the monomer, suchthatthe monomer (e.g., amino acid) is more accurately identifiable, e.g., while translocating adjacent to or through the nanoscale object. Accordingly, the methods, systems, and compositions disclosed herein provide a more facile and highly parallelized approach to protein sequencing with increased accuracy. The methods, systems, and compositions disclosed herein may enable high -resolutionsequencing of polymeric analytes, e.g., at single-molecule, sub -zeptomole, sub-attomole, sub- femtomole, or sub -picomole resolution.
[0128] In one aspect of the present disclosure, provided herein is a method for sequencing a peptide comprising a plurality of amino acids, comprising (a) providing a plurality of modified amino acids generated from at least a subset of the plurality of amino acids, and(b) sequencing the plurality of modified amino acids, thereby determining an amino acid identity of each modified amino acid of the plurality of modified amino acids. The methods provided herein may advantageously provide for highly -accurate identification and sequencing of individual amino acids from a peptide; for instance, sequencing of the plurality of modified amino acids may have an average read accuracy that is greater than 80% for at least 2 different modified amino acid types. In some embodiments, the plurality of modified amino acids is derived from two contiguous amino acids of a peptide. In some embodiments, the modified amino acids herein comprise polymerizable molecules.
[0129] In some embodiments, the methods provided herein comprise providing a modified amino acid, in which the modified amino acid comprises a polymerizable molecule and translocating the modified amino acid or derivative thereof through or adjacent to a nanopore or a nanogap. The method may further comprise sensing the modified amino acid as it translocates through or adjacent to the nanopore or nanogap. In some instances, the sensing comprises measuring a signal from the nanopore or the nanogap as the modified amino acid or derivative thereof translocates through the nanopore or nanogap. In some instances, the method further comprises using the signal to determine the identity of the modified amino acid or derivative thereof. In some instances, the polymerizable molecule facilitates the translocation or modifies an interaction between the modified amino acid and the nanopore or nanogap. The interaction may be a covalent or noncovalent interaction and can, in some instances, induce a biophysical change within the nanopore or nanogap. In some instances, the polymerizable molecule modulates the translocation speed of the modified amino acid through the nanopore or nanogap, which may render the modified amino acid more detectable as compared to an unmodified amino acid. The modulation of the interaction of the modified amino acid may control or cause a change in the translocation speed of the modified amino acid through the nanopore or nanogap. In some embodiments, the translocation of the modified amino acid through the nanopore is performed under conditions sufficient to reduce the translocation speed of the modified amino acid through the nanopore as compared to a control condition. In some instances, the reduced translocation speed of the modified amino acid through the nanopore or nanogap may improve the measured signal (e.g., current blockade) or increase the signal-to-noise ratio of the measured signal ascompared to an unmodified amino acid or the control condition. In some instances, the reduced translocation speed of the modified amino acid through the nanopore or nanogap improves the accuracy of identification of the modified amino acids.
[0130] In some embodiments, the methods provided herein comprise performing an intramolecular expansion of a polymeric analyte, e.g., a peptide, to generate a modified monomer (e.g., modified amino acid). Such an intramolecular expansion process may comprise providing a linker, coupling the linker to a monomer of the polymeric analyte to generate a monomer-linker complex, coupling the linker or the monomer-linker complex to a capture moiety, and cleaving the monomer from the polymeric analyte, thereby yielding a modified monomer. In some instances, the method may further comprise repeatingthe intramolecular expansion or one or more operations of the intramolecular expansion on the next monomer of the polymeric analyte to generate another modified monomer. In some instances, the monomer-linker complex or the modified monomer resulting from a round or cycle of intramolecular expansion may be coupled to that of the previous round or cycle of intramolecular expansion to generate a stacked plurality of modified monomers, e.g., linked by a polymerizable molecule backbone. In some instances, the modified monomer or stacked plurality of modified monomers is detected using a nanopore sequencer, thereby outputting the identity of the individual monomers of the polymeric analyte.
[0131] Modified amino acids: The modified amino acid or derivative thereof may originate from or be part of a protein or peptide; for example, the modified amino acid may comprise or be derived from an amino acid located at a terminus (N-terminus or C-terminus) of a peptide, or the modified amino acid may comprise or be derived from an amino acid located within the peptide. The modified amino acid may comprise a proteinogenic amino acid or derivative thereof with any number of modifications. Examples of modifications include, in non-limiting examples, chemical modifications (e.g., protecting groups), biological modifications (e.g., post-translational modifications, modifications introduced by enzymatic treatment or digestion), physical modifications^. g., mutations introduced by irradiation, heat, etc.), andthelike. In some instances, the modified amino acid or derivative thereof comprises or is coupled to a binding agent, such as an antibody, antibody fragment, nanobody, aptamer, peptide, a small molecule, an inorganic compound, a polymer, or any variations or combinations thereof. The modified amino acid may comprise a covalent or noncovalent modification. In some instances, the modified amino acid comprises a non-naturally occurring chemical modification. For example, the modified amino acid may comprise a protecting group, such as, in non-limiting examples, a methyl, formyl, ethyl, acetyl, t-butyl, anisyl, benzyl, tifluoroacetyl, N-hydroxysuccinimide, t-butyloxycarbonyl, benzoyl, 4-methyl benzyl, thioanizyl, thiocresyl, benzyloxymethyl, 4 -nitrophenyl,benzyloxycarbonyl, 2-nitrobenzoyl, 2-nitrophenylsulphenyl, 4-toluenesulphonyl, pentafluorophenyl, diphenylmethyl, 2-chlorobenzyloxy carbonyl, 2,4, 5 -trichlorophenyl, 2- bromobenzyloxycarbonyl, 9-fluorenylmethyloxycarbonyl, triphenylmethyl, or 2, 2, 5,7,8- pentamethyl-chroman-6-sulphonyl group. Additional non-limiting examples of modifications include fluorescence labeling, Isotope labeling, PEGylation, alkylation, arylation, acylation, benzylation, carbamoylation, guanidination, eliminylation, iodination, seleniation, thiolation, acetoacetylation, aminoethylation, carbonylation, cyanylation, glucuronidation, glucosylation, hydroxymethylation, lipoylation, mannosylation, naphthoquinone addition, nucleotide addition, phosphopantetheinylation, polysialylation, andtyramine addition. In some instances, the modified amino acid comprises a proteinogenic amino acid or derivative thereof that is coupled to a polymerizable molecule. In some instances, the modified amino acid comprises a proteinogenic amino acid or derivative thereof, a linker, and a polymerizable molecule.
[0132] The modified amino acid may comprise any useful modification. Modifications may be naturally-occurring (e.g., post translational modifications) or non-naturally occurring, such as by labeling or tagging, e.g., with an amino acid- or amine-reactive agent or linker comprising the amino acid- or amine-reactive agent. Examples of amino acid- or amine-reactive agents include isothiocyanate (e.g., PITC, NITC), l-fluoro-2,-4-dinitrobenzene (DNFB), dansyl chloride, 4- sulfonyl-2-nitrobfluorobenzene (SNFB), an acetylating agent, an acylating agent, an alkylating agent, a guanidination agent, a thioacetylation agent, a thioacylation agent, a thiobenzoylation agent, or a derivative or combination thereof. Alternatively, or in addition to, the one or more modified amino acids may comprise an adduct (e.g., a polymer such as PEG, a polymerizable molecule such as a nucleic acid molecule, a nanoparticle or nanotube, a peptide or protein), a lipid, a carbohydrate, a metabolite, a fluorophore, a hapten, a peptide, a synthetic peptide, a peptoid, a quencher, a tag (e.g., a fluorescent tag, a magnetic tag, a radioactive tag), a barcode, or other moiety. In some instances, a modified amino acid may comprise a modification that facilitates recruitment of an enzyme (or ribozyme or DNAzyme) to recognize or cleave a terminal amino acid, e.g., a NTAA or CTAA of a peptide. For example, a terminal amino acid of a peptide may be modified with a saccharide in order to recruit a lectin or lectin-bound protease. In another example, one or more modified amino acids may comprise or be coupled to a nucleic acid molecule having a first sequence that is complementary to a second sequence comprised by an oligo-bound protease. Hybridization of the first sequence to the second sequence may facilitate local recruitment of the protease to the amino acid to be cleaved. In yet another example, a peptide may be modified with phenylisothiocyanate (PITC), which may allow for recruitment and cleavage of the modified amino acid by an Edmanase. In another example, a peptide may bemodified with a peptide or protein, which may allow for recruitment and cleavage of the modified amino acid by a protease. In yet another example, a peptide may be modified with a tag that may allow for recruitment and cleavage of the modified amino acid by an enzyme. In other examples, a peptide may be modified with functional moiety that is recognized by a specific cleaving enzyme; recognition and binding of the cleaving enzyme to the functional moiety may result in cleavage of the modified amino acid. In some examples, modificationsto amino acids may include epitope tags, which can facilitate binding of a binding agent to the modified amino acid. Examples of such epitope tags include fluorophores, nucleic acid molecules, peptides, haptens, polymers, chemical moieties, or other adduct molecule.
[0133] The methods described herein may further comprise generating the modified amino acid. In some instances, the modified amino acid comprises a proteinogenic amino acid or derivative thereof, a linker, and a polymerizable molecule. In one example, an amino acid of a peptide may be contacted with a linker that comprises (i) first reactive moiety capable of reacting with the amino acid and (ii) a second reactive moiety. Prior to, during, or subsequent to the reaction of the first reactive moiety with the amino acid, a polymerizable molecule comprising a third reactive moiety that is capable of reacting with the second reactive moiety may be provided. The second and third reactive moieties may comprise click chemistry moieties that can react with one another (e.g., azide and DBCO, azide andBCN, alkyne and DBCO, TCO and tetrazine, etc.). The reaction of the amino acid with the linker and the linker to the polymerizable molecule may thus yield a modified amino acid comprising the amino acid, the linker, and the polymerizable molecule. Alternatively, or in addition to, the polymerizable molecule may comprise the amino acid- reactive moiety (optionally via a linker) and may react directly with the amino acid. In some instances, cleavage of the amino acid from the peptide may be performed, and the modified amino acid may comprise the cleaved product comprising the cleaved, and optionally derivatized, amino acid, the linker (if present), and the polymerizable molecule. In some instances, the amino acid is a terminal amino acid.
[0134] Alternatively, or in addition to, generating the modified amino acid may comprise (i) contacting an amino acid of a peptide with a polymerizable molecule comprising an amino acid reactive group and (ii) polymerizing the polymerizable molecule, thereby generating the modified amino acid. In one such example, the polymerizable molecule may comprise a modified nucleotide comprising an amino acid reactive moiety (e.g., isothiocyanate, guanidinylating group, dithioester, xanthate, etc.). The amino acid reactive group may react with the amino acid (e.g., a terminal amino acid). The modified nucleotide may then be subject to a nucleic acid reaction such as ligation (e.g., usingligase or chemical ligation such as click chemistry) or an extension reaction,e.g., using polymerase or terminal deoxynucleotidyl transferase (TdT), thereby generating the modified amino acid. See, e.g., FIG. IB (inset).
[0135] A modified amino acid may comprise a single amino acid or a plurality of amino acids (e.g., dipeptide, tripeptide, etc.). For example, the modified amino acid may comprise, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or greater amino acids. In instances where the modified amino acid comprises more than one amino acid, the plurality of amino acids comprised by the modified amino acid may be of the same amino acid type (e.g., leu-leu, val-val, ile-ile) or different amino acid types (e.g., val-pro, arg-his, gly-leu). The plurality of amino acids comprised by the modified amino acid may comprise any number of modifications, e.g., as described elsewhere herein.
[0136] Polymerizable Molecule: The polymerizable molecules described herein may comprise any useful polymerizable moiety. The polymerizable moiety may comprise a naturally occurring or synthetic polymer (organic or inorganic) or biopolymer. The polymer may comprise one or more monomer types, which may assemble covalently or non covalently . The polymer may comprise one or more repeating monomers. The polymer may comprise non-repeating units. In instances where more than one monomer type is used, the polymer may form an alternating copolymer structure, a periodic copolymer structure, a random copolymer structure, a block copolymer structure, a chained or grafted copolymer, or any other useful structure. The polymer may be linear or non-linear. The polymer may comprise polyethylene glycol (PEG), PEG- diacrylate, PEG-acrylate, PEG-thiol, PEG-azide, PEG-alkyne, polyacrylamide, agarose, collagen, fibrin, gelatin, chitosan, hyaluronic acid, alginate, polyvinyl alcohol, or another polymer. A polymerizable molecule may comprise one or more repeating units (e.g., monomers), or a polymerizable molecule may not comprise repeatable units. A polymerizable may be expandable or extendible (e.g., capable of being polymerized) or the polymerizable molecule may not be expandable or extendible (e.g., comprises a terminal monomer or a terminal end that cannot be expanded). The polymerizable molecule may comprise a scaffold.
[0137] In some examples, the polymerizable molecule comprises a biomolecule, e.g., a protein or peptide, a nucleic acid molecule, e.g., DNA, RNA, XNA, PNA, LNA, a carbohydrate or lipid chain, or a combination thereof. In some instances, the polymerizable molecule comprises a nucleic acid molecule, which can comprise any useful number and type of nucleotides, e.g., including canonical and noncanonical bases, and the number of nucleotides may be modulated based on the intended purpose. For example, the length or sequence of the nucleic acid molecules may be modulated to alter a property of the amino acid (e.g., volume, aspect ratio, charge, etc.) to which the polymerized molecule is tethered. The nucleic acid molecule may additionally or alternatively comprise any useful functional sequences or moieties, including but not limited tobarcode sequences or other identifying sequences, UMI sequences, template or recognition sequences for DNA repair or replication, 5’ blocking groups, 3’ blocking groups, protecting groups, enzyme recognition sites (e.g., transposition sites, restriction sites), spacer sequences or spacer moieties (e.g., dSpacer, 3 ’ C3 Spacer, 2-aminopurine, hexanediol, Spacer 9, Spacer 18, etc.) sequencing primer sequences, read sequences, or primer sequences. The nucleic acid molecule may comprise canonical bases, noncanonical or modified bases, naturally occurring bases, synthetic bases, abasic sites, nucleotide analogs, or a combination thereof. The nucleic acid molecule can be single stranded, double stranded, or partially double stranded. The nucleic acid molecule may comprise RNA, DNA, or a hybrid of RNA and DNA. In some instances, the modified amino acid or derivative thereof is generated using an iterative process, as described elsewhere herein; accordingly, the nucleic acid molecule may comprise information on the round or cycle number of the iterative process. In some instances, the nucleic acid molecule comprises a nucleic acid barcode molecule and may comprise useful information on the identity of the amino acid, temporal information, spatial information, etc.
[0138] The polymerizable molecules described herein may be any useful type of polymerizable molecule. The polymerizable molecules may be naturally occurring, such as biological polymers (e.g., nucleic acid molecules, peptides, polysaccharides, fatty acids), or other naturally occurring polymers, e.g., rubber, cellulose, starches, polyhydroxyalkanoates, chitosan, dextran, structural proteins (e.g., collagen, hyaluronic acid, glycosaminoglycans), agarose, carrageenan, isphagula, acacia, agar, gelatin, shellac, xanthan gum, guar gum, alginate, etc. The polymerizable molecules may be synthetic, e.g., acrylics, nylons, silicones, viscose, rayon, polyesters, poly carboxylic acids, polyvinyl acetate, polyacrylamide, polyacrylate, polyethylene glycol, polyurethane, polylactic acid, silica, polystyrene, polyacrylonitrile, polybutadiene, polycarbonate, polyethylene terephthalate, poly(chlorotrifluoroethylene), polyethylene oxide), poly(ethylene terephthalate), polyethylene, polyisobutylene, poly(methyl methacrylate), poly(oxymethylene), poly formaldehyde, polypropylene, polystyrene, polytetrafluoroethylene), poly(vinyl acetate), poly(vinyl alcohol), poly(vinyl chloride), poly(vinylidene dichloride), poly(vinylidene difluoride), poly(vinyl fluoride) and combinations thereof. The polymerizable molecule may be charged (positively charged, negatively charged) uniformly or nonuniformly. The polymerizable molecule may comprise both a positively charged region and a negatively charged region. The polymerizable molecules may comprise one or more reactive moieties (e.g, radical groups) to initiate polymerization or to add a functional group, or the polymerizable molecules may be polymerized via contacting of an initiating agent (e.g., ammonium persulfate, peroxide, or other radicalizing agent). The polymerizable molecules may be polymerizable via anenzymatic reaction or contacting of an enzyme (e.g., polymerizing enzyme such as polymerases), ribozyme or DNAzyme. Alternatively, or in addition to, the polymerizable molecules may be polymerizable via self-assembly. The polymerizable molecules may comprise a single polymer type (e.g., a homopolymer) or more than one polymer type (e.g., a copolymer) and may comprise random or arranged monomers. The polymerizable molecules may be a block polymer, alternating copolymer, periodic copolymer, statistical copolymer, stereoblock copolymer, gradient copolymer, branched copolymer, graft copolymer, etc. The polymerizable molecule may comprise one or more non-repeat monomers.
[0139] The polymerizable molecules may encompass any useful geometry or shape. For instance, the polymerizable molecule may comprise a linear, circular, branched or other polymer. In some instances, the polymerizable molecules comprise one or more nucleic acid molecules, which may have any useful shape or geometry, e.g., single-stranded, double-stranded, partially- double stranded, hairpin, 2D or 3D structure (e.g., DNA origami).
[0140] The same or different types of polymerizable molecules may be used in the methods described herein. For example, a first polymerizable molecule may be a nucleic acid molecule, and a second polymerizable molecule may be a peptide. In another example, both the first polymerizable molecule and the second polymerizable molecule are nucleic acid molecules. In such an example, the first polymerizable molecule may be coupled to the second polymerizable molecule via ligation, hybridization, an extension reaction, or a combination thereof. For instance, the first polymerizable molecule may comprise a first nucleic acid sequence and the second polymerizable molecule may comprise a second nucleic acid sequence. The first nucleic acid sequence may be complementary or partially complementary to the second nucleic acid sequence, and the coupling may comprise hybridizing the first nucleic acid sequence or portion thereof to the second nucleic acid sequence or portion thereof. Alternatively, the first nucleic acid sequence and the nucleic acid sequence may be complementary to two sequences of a splint or bridge oligonucleotide, and coupling may be mediated via hybridization to the splint oligo. The first nucleic acid sequence may be ligated to the second nucleic acid sequence, either chemically (e.g, via click chemistry approaches in which the first polymerizable molecule and the second polymerizable molecule each comprise one member of a click chemistry pair) or enzymatically (e.g., using a ligase).
[0141] The polymerizable molecules may comprise functional portions or functional groups. For example, the polymerizable molecules may comprise a nucleic acid molecule comprising a functional sequence, such as a primer sequence (e.g., universal priming site), a sequencing sequence, a read sequence, a unique molecular identifier (UMI), a barcode sequence, a cleavagesequence (e.g., a restriction site, a Cas-binding sequence), a transposition sequence (e.g., a mosaic end sequence), a spacer moiety, a primer-binding sequence, or a combination thereof. The polymerizable molecule may comprise a linker or linking moiety, e.g., a reactive or cross-linking moiety such as click chemistry moieties (e.g., alkyne, azide, DBCO, BCN, tetrazine, TCO), photoreactive groups (e.g., benzophenone), l-ethyl-3-(3-dimethylaminopropyl)carbodiimide hydrochloride (EDC), N-hydroxysulfosuccinimide (NHS), Sulfo-NHS, or NHS-esters, sulfhydryls, amine groups, maleimides, hydrazines, hydroxyl amines, thiols, biotin or streptavidin, cystamine, glutaraldehyde, formaldehyde, succinimidyl 4-(N-maleimidomethyl)cyclohexame-l- carboxylate (SMCC), Sulfo-SMCC, 4-(4,6-Dimeth oxy-1 ,3,5-triazin-2-yl)^l- methylmorpholinium chloride (DMTMM), silane (e.g., amino silanes), combinations thereof, etc. The polymerizable molecule may comprise an amino acid reactive group, e.g., an isothiocyanate, an aldehyde, a guanidinylating agent, a xanthate, a dithioester or thiocarbamoyl, or other amino acid reactive group, as described elsewhere herein.
[0142] The polymerizable molecules may comprise a tag, which may be useful in identifying enriching, or purifying the polymerizable molecule. A tag may be a barcode, e.g., a nucleic acid barcode molecule, fluorophore or fluorescent protein, bioluminescenttag, chemiluminescent tag isobaric tag, radioisotope, mass tag, or other detectable moiety, or a combination thereof. The tag may be used in enriching or purifyingthe polymerizablemolecule and may comprise, for example, an affinity tag (e.g., an antibody, an aptamer, a biotin molecule, a streptavidin molecule, etc.). In some examples, a polymerizable molecule may comprise a biotin tag, which may enable pulldown or purification using streptavidin (e.g., streptavidin beads).
[0143] The polymerizable molecule may be cleavable or comprise a cleavable moiety. The cleavage of the polymerizable molecule or cleavable moiety may be achieved using a stimulus. The stimulus can be, for example, a chemical stimulus (e.g., application of an acid or base, a reducing agent), a biological stimulus (e.g., a cleaving enzyme, a protease, a nuclease), a thermal stimulus (e.g., application of heat), a photo-stimulus, a physical or mechanical stimulus, or other type of stimulus or a combination of stimuli, as described elsewhere herein. In some instances, the polymerizable molecule may comprise a cleavable tag, e.g., a photocleavable biotin moiety, which can be useful for purifying or enriching the polymerizable molecule from a mixture (e.g, during an operation of intramolecular expansion of a polymeric analyte). In some instances, the polymerizable molecule comprises a nucleic acid sequence that is cleavable using a chemical stimulus; for example, the nucleic acid sequence may comprise only purine nucleobases and may be cleavable using acid-catalyzed depurination and cleavage (e.g., using trifluoracetic acid or boron triflate etherate).
[0144] The polymerizable molecules may be any useful size. The polymerizable molecules may be about 1 angstrom, about 2 angstrom, about 3 angstrom, about 4 angstrom, about 5 angstrom, about 6 angstrom, about ? angstrom, about 8 angstrom, about 9 angstrom, about 10 angstrom, about 20 angstrom, about 30 angstrom, about 40 angstrom, about 50 angstrom, about 60 angstrom, about 70 angstrom, about 80 angstrom, about 90 angstrom, about 100 angstrom, about 200 angstrom, about 300 angstrom, about 400 angstrom, bout 500 angstrom, about 600 angstrom, about 700 angstrom, about 800 angstrom, about 900 angstrom, about 1000 angstrom, about 10,000 angstrom, about 100,000 angstrom or greater in size, length, or another dimension. In some instances, the polymerizable molecule (e.g., the first polymerizable molecule or the second polymerizable molecule) comprises a nucleic acid molecule comprising one or more nucleotide bases. The polymerizable molecule may comprise any useful number of nucleotide bases, e.g., about 1 base, about 2 bases, about 3 bases, about 4 bases, about 5 bases, about 6 bases, about ? bases, about 8 bases, about 9 bases, about 10 bases, about 20 bases, about 30 bases, about 40 bases, about 50 bases, about 60 bases, about 70 bases, about 80 bases, about 90 bases, about 100 bases, about 200 bases, about 300 bases, about 400 bases, about 500 bases, about 600 bases, about 700 bases, about 800 bases, about 900 bases, about 1000 bases, or a greater numb er of bases.
[0145] In some instances, the polymerizable molecules comprises a nucleic acid molecule or analog thereof. The nucleic acid molecule can be single stranded, double stranded, or partially double-stranded. The nucleic acid molecule may comprise a crosslinked nucleic acid molecule. The nucleic acid molecule may comprise a modified nucleotide or non-canonical base. For instance, the polymerizable molecules may comprise a pseudo-complementary base, a bridged nucleic acid (BNA), a xenonucleic acid (XNA), a locked nucleic acid (LNA), a peptide nucleic acid (PNA), a gamma-PNA molecule, a morpholino, or a combination thereof. In some instances, the polymerizable molecule comprises a PNA, which may be advantageous in increasing the stability of the molecule under reaction conditions (e.g., acidic cleavage). In some instances, a polymerizable molecule may comprise a hexitol nucleic acid (HNA) or a cyclohexyl nucleic acid (CeNA), which may be useful in rendering the polymerizable molecule more resistant to acid degradation (e.g., as used in conventional Edman degradation). Alternatively, or in addition to, a polymerizable molecule may comprise naturally occurring bases that are more resistant to acid degradation, e.g., be composed of primarily thymine or cytosine or analogs thereof. For example, a nucleic acid molecule may comprise at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% thymines or cytosines, which can render the nucleic acid molecule more acid resistant as compared to a nucleic acid molecule comprising adenines or guanines.
[0146] The polymerizable molecules may comprise more than one molecule type. For example, the polymerizable molecule may comprise a conjugate of two or more polymer types, such as two or more biopolymers, a biopolymer and a synthetic polymer, two or more synthetic polymers, etc. In some instances, the polymerizable molecule may comprise a peptide-oligo conjugate, a peptide-polymer conjugate, an oligo-polymer conjugate, etc.
[0147] The modified amino acid may comprise a polymerizable molecule coupled covalently or non -covalently thereto. The coupling may be performed using any suitable chemistries and reaction conditions and may comprise the use of a linker. In one such example, a polymerizable molecule such as a nucleic acid molecule may comprise a first reactive group, e.g., a first click chemistry moiety, as described elsewhere herein, and may be contacted with a linker comprising a second reactive group, e.g., a second click chemistry moiety that is able to react with the first reactive group. The linker may also comprise an additional reactive group that is able to tether to an amino acid (e.g., a terminal amino acid) and optionally, cleave the amino acid from a peptide. For example, the additional reactive group may be a thiocyanate conjugate, e.g., an isothiocyanate (ITC) such as phenyl isothiocyanate (PITC) or naphthylisothiocyanate (NITC), or an aldehyde group, e.g., ortho-phthalaldehyde (OPA), 2,3-naphthalenedicarboxyaldehyde (NDA), a guanidinylating agent, dinitrofluorobenzene (DNFB), dansyl chloride, a dithioester, a thiobenzoyl, a thioacetyl, a xanthate, or other amino acid-reactive group. The linker may be reacted with an amino acid of a peptide, e.g., the NTAA or CTAA. Use of such a linker comprising at least two reactive groups may allow for (i) tethering of the amino acid to the linker and (ii) tethering of the linker to the polymerizable molecule (such as a nucleic acid molecule; see, e.g., FIG. 2A-2B) In some instances, the linker may be provided pre-tethered to the polymerizable molecule (e.g., nucleic acid molecule) prior to contacting with the amino acid. In some instances, the conjugation of the polymerizable molecule to the amino acid, either via a linker or without a linker, may change the chemical structure of the amino acid. For example, if using a linker comprising an isothiocyanate moiety, the amino acid may be derivatized to a thiocarbamyl group (e.g., under alkaline conditions), a thiazolone group (e.g., under acid conditions), a thiohydantoin group, or other chemical moiety, thereby generating a modified amino acid comprising a polymerizable molecule coupled thereto.
[0148] In some instances, the polymerizable molecule comprises a linker that can couple to an amino acid-linker complex. For instance, the polymerizable molecule may comprise a nucleic acid molecule that comprises a linker that comprises a click chemistry moiety. The linker of the polymerizable molecule may be coupled via a synthetic nucleobase or to the backbone. For example, an octadiynyl deoxy nucleotide or ribonucleotide or an ethynyl deoxy nucleotide orribonucleotide may be incorporated into a DNA molecule, thereby yielding a linker-conjugated DNA molecule. The linker-conjugated DNA molecule may then react, e.g., via a click chemistry reaction, with an amino acid-linker complex or another linker (e.g., a linker comprising an amino acid reactive moiety or that is already reacted to an amino acid and additionally comprises a click chemistry moiety).
[0149] The polymerizable molecule of the modified amino acid may facilitate translocation of the modified amino acid through the nanopore or nanogap. For example, the polymerizable molecule may be able to couple to a processive enzyme that allows for ratcheting of the polymerizable molecule through the nanopore or nanogap. In some instances, the polymerizable molecule comprises a nucleic acid molecule, and translocation of the modified amino acid through the nanopore or nanogap is facilitated using a DNA or RNA-processing enzyme, such as a helicase, polymerase, topoisomerase, or other enzyme. The ratcheting enzyme may be naturally derived (e.g., from a virus such as poxvirus helicase-primase D5, El (papillomaviridae), or herpesviridae (UL5), poxvirus DNA polymerase holoenzyme, E9, A20, D4, bacteria, mammal), or the enzyme may be an engineered variant.
[0150] A modified amino acid having the same amino acid type (e.g., two modified amino acids both comprising leucine or derivatized leucine) may be coupled to the same or different linkers and / or polymerizable molecules. For instance, first modified amino acid and a second modified amino acid may comprise the same polymerizable molecule but different linkers (e.g., different click chemistry moieties), or the same linker but different polymerizable molecules (e.g., a protein and a nucleic acid molecule, two different nucleic acid sequences, etc.).
[0151] Linkers: One or more linkers maybe used to couple molecules, e.g., the polymerizable molecule to an amino acid or derivative thereof to thereby generate the modified amino acid or a precursor or derivative thereof. In some instances, the linker comprises a click chemistry moiety. The click chemistry moiety may comprise any suitable clickable or bioorthogonal moieties, as described elsewhere herein, e.g., alkenes, alkynes (e.g., alkyne, cycloalkynes such as DBCO and BCN), azides, epoxides, amines, thiols, nitrones, isonitriles, isocyanides, aziridines, activated esters, and tetrazines, and combinations, variations, or derivatives thereof. The linker may be subjected to conditions sufficient to react the click chemistry moiety of the linker to an additional click chemistry moiety (e.g., comprised by the polymerizable molecule), e.g., provision of metal catalysts, appropriate solvents, pH, temperature, ionic concentration, orlight / energy for any useful duration of time.
[0152] The coupling of the linker to a monomer of a polymeric analyte e.g., an amino acid of a peptide, and / or to the polymerizable molecule or to a capture moiety may be covalent ornoncovalent. In an example, a linker may comprise a first reactive group that is able to couple to a monomer of the polymeric analyte (e.g., an amino acid of a peptide) and optionally, cleave the amino acid from a peptide. For example, the first reactive group may be an amino-acid reactive group, e.g., an isothiocyanate (ITC) such as isothiocyanate, phenyl isothiocyanate (PITC), 3- pyridyl isothiocyanate (PYITC), 2-piperidinoethyl isothiocyanate (PEITC), 3-(4-morpholino) propyl isothiocyanate (MPITC), 3- (diethylamino)propyl isothiocyanate (DEPTIC) or naphthylisothiocyanate (NITC), fluorescein isothiocyanate (FITC), ammonium thiocyanate, potassium thiocyanate, trimethylsilyl isothiocyanate (TMS-ITC), phenyl phosphoroisothiocyanatidate, acetyl isothiocyanate (AITC), or an aldehyde group, e.g., orthophthalaldehyde (OP A), 2,3 -naphthalenedicarboxy aldehyde (ND A), 2 -pyridinecarboxy aldehyde, a dithioester, a xanthate, a guanidinylating agent, or other amino acid reactive moiety which can react with an N-terminal amino acid (NTAA). The linker may additionally comprise a second reactive group that is capable of coupling, either directly or indirectly, to the capture moiety. In an example of direct coupling, the capture moiety may comprise a click chemistry moiety (e.g., alkyne), and the second reactive group of the linker may comprise an additional click chemistry moiety (e.g., azide) that can react with the click chemistry moiety of the capture moiety. Alternatively, the linker may be coupled indirectly to the capture moiety, e.g., via noncovalent interaction or via an intermediate linking molecule. In some examples, the intermediate linking molecule may comprise a polymerizable molecule (e.g., a polymer or nucleic acid molecule) that can couple the linker to the capture moiety . In one such example, the polymerizable molecule may comprise (i) a third reactive group that is capable of coupling to the second reactive group (e.g, via alkyne-azide click chemistry) of the linker and (ii) a moiety that can couple to the capture moiety (e.g., another orthogonal click chemistry reaction, avidin-biotin interaction, nucleic acid coupling or hybridization). In some instances, the linking polymerizable molecule comprises a nucleic acid molecule that comprises (i) a click chemistry moiety (e.g., alkyne) that can conjugate to a reactive group (e.g., azide) of the linker and (ii) a nucleic acid sequence that can couple to the capture moiety, e.g., via ligation, splint ligation, or hybridization. Alternatively, or in addition to, the linking polymerizable molecule may comprise the linker comprising an amino acid-reactive moiety. In some instances, the linker comprises a linking nucleic acid molecule that comprises a self-splinting moiety.
[0153] In some instances, the linker may be directly coupled to the polymerizable molecule (e.g., the polymerizable molecule comprises the linker that can couple to a monomer of the polymeric analyte). For example, the polymerizable molecule may comprise an amino acid reactive moiety, which may enable direct coupling of the polymerizable molecule to the monomerof the polymeric analyte. In some instances, the polymerizable molecule comprises a nucleic acid molecule that comprises an amino acid-reactive moiety, e.g., a guanidinylating group, a dithioester, an isothiocyanate (e.g., PITC), etc., which can be used to conjugate the polymerizable molecule to an amino acid, e.g., an N-terminal amino acid of a peptide analyte. In some instances, the peptide analyte may comprise or be coupled to an additional nucleic acid molecule; in such instances, the additional nucleic acid molecule may be used to localize (e.g., via hybridization) the polymerizable molecule to enable more efficient reaction of the amino acid-reactive group with the amino acid.
[0154] When applicable, the click chemistry moieties of the linker and capture moiety or intermediate linking molecule may comprise useful clickable moieties, as described elsewhere herein, e.g., alkenes, alkynes, azides, epoxides, amines, thiols, nitrones, isonitriles, isocyanides, aziridines, activated esters, and tetrazines, and combinations, variations, or derivatives thereof. The linker may be subjected to conditions sufficient to react the first click chemistry moiety to the second click chemistry moiety, e.g., provision of metal catalysts, appropriate solvents, pH, temperature, ionic concentration, or light / energy for any useful duration of time.
[0155] The linker may comprise an amino acid-reactive moiety. The amino acid- reactive moiety of the linker may be any useful moiety that enables the reactive moiety to conjugate to and optionally cleave an amino acid. In some examples, the first reactive moiety can react with a terminal amino acid (e.g., NTAA or CTAA). In such examples, the first reactive moiety may comprise any primary amine or carboxylic group reactive group, including but not limited to isocyanates, acyl azides, NHS esters, sulfonyl chlorides, aldehydes, glyoxals, epoxides, oxiranes, carbonates, aryl halides, acyl halides, aldehydes, imidoesters, carbodiimides, anhydrides, phenyl esters, isothiocyanates (e.g., phenyl isothiocyanate, sodium isothiocyanate, ammonium isothiocyanates (e.g., tetrabutylammonium isothiocyanate, tetrabutylammonium isothiocyanate), diphenylphosphoryl isothiocyanate), acetyl chloride, cyanogen bromide, carboxypeptidases, azide, alkyne, DBCO, maleimide, succinimide, thiol-thiol disulfide bonds, tetrazine, TCO, vinyl, methylcyclopropene, acryloyl, allyl, ketones, Michael acceptors, alkyl halides, carboxylic acids, carbon dioxide, diazomethane and other diazo compounds, among others. Additional examples of amino acid reactive groups are linkers or sequencing reagents comprising amino acid reactive groups are provided in U.S. Pat. Pub. No. 2020 / 0217853, International Patent App. Nos. PCT / US2023 / 079684, filed November 14, 2023, and PCT / US2024 / 013211, filed January 26, 2024, and U.S. Provisional Patent App. No. 63 / 601,389, filed November 21, 2023, each of which is incorporated by reference herein in its entirety.
[0156] The linker may comprise any additional useful moieties. For example, the linker may comprise a releasable or cleavable moiety, which may facilitate removal of the monomer from the polymeric analyte, or portion thereof, or from the substrate. Such a releasable or cleavable moiety may comprise, for example, a disulfide bond, which may be releasable by contacting with a reducing agent (e.g., DTT, TCEP). In some examples, the linker may couple to the third polymerizable molecule via the releasable or cleavable moiety, alternatively or in addition to the coupling via click chemistry moieties. As such, the coupling between the polymerizable molecule and the linker may be reversible. The linker may additionally or alternatively comprise any number of spacing moieties, e.g., polymers (e.g., PEG, PVA, polyacrylamide), peptide nucleic acids (PNAs), aminohexanoic acid, nucleic acids, alkyl chains, etc. Such spacing moieties may increase the distance between any other moieties of the linker, e.g., the amino acid-reactive group and the polymerizable molecule-reactive group. The spacing moiety may comprise any useful properties, e.g., comprise one or more moieties that are charged, polar, nonpolar, hydrophobic, hydrophilic, or a combination thereof. Additional moieties, such as those to change a property of the linker, such as charge, hydrophobicity, hydrophilicity, flexibility, etc. may be employed. The linker may comprise a carbon, ethylene glycol, acetylene, isothianapthene, dimethyl silane, urethane, glycolic acid, lactic acid, dioxanone, methyl methacrylate, hydroxy ethyl methacrylate, vinyl chloride, tetrafluoroethylene, propylene, ethylene, ether ketone, ether suylfonefluorene, aniline, phenylene, polypyrrole, phenylenevinylene, fluorene, thiophene, or 3,4- ethylenedioxythiophene.
[0157] The linker may comprise or be coupled to a detectable moiety, e.g., a fluorophore, radioisotope, mass tag, nucleic acid molecule (which can also act as a releasable or cleavable moiety), or other detectable moiety. In some examples, the linker comprises a fluorophore, which can enable localization visualization of the linker using single- molecule imaging. In another example, the monomer may be labeled with a first fluorophore and the linker may comprise a second fluorophore to enable localization visualization of the linker and the monomer (e.g., using two-channel imaging or FRET). In some instances, the linker comprises a barcode, e.g., a nucleic acid barcode or a peptide barcode.
[0158] Use of a linker comprising two reactive groups may allow for coupling of the linker to (i) the monomer of the polymeric analyte and (ii) a linking polymerizable molecule (e.g., linking nucleic acid molecule) or (iii) the capture moiety . In some instances, when a linking polymerizable molecule is used, the linker may be pre-coupled to the linking polymerizable molecule. For example, a precursor linker may comprise a monomer-binding group (e.g., PITC) and a click chemistry moiety (e.g., azide), which may be reacted with a polymerizable molecule (e.g.,oligonucleotide) comprising complementary click chemistry moiety (e.g., alkyne) to generate a linker that is capable of coupling to the monomer and the capture moiety (e.g., another oligonucleotide). In another example, the linkingpolymerizable molecule may comprisean amino acid reactive group (e.g., an isothiocyanate such as PITC, guanidinylating agent, xanthate, dithioester, etc.). In some instances, the linker may be provided pre-coupled to the linking polymerizable molecule.
[0159] FIG. 2A schematically shows an example linker that may be used in sequencing polymeric analytes such as peptides. FIG. 2A Panel A shows a bifunctional linker 203 (e.g., 1- (but-3-yn-l-yl)-4-isothiocyanatobenzene) comprising an amino acid reactive moiety (e.g., PITC) and an alkyne click chemistry moiety, which may be reacted with a polymerizable molecule 201 (e.g., a linking nucleic acid molecule) comprising a complementary azide click chemistry moiety. The bifunctional linker may also comprise a spacer moiety, e.g., an alkyl chain (an ethyl group is depicted) of any length, a polymer (e.g., PEG) of any length, etc. The spacer moiety may be located between the amino acid reactive moiety and the click chemistry moiety. FIG. 2A Panel B shows the product of a click chemistry cycloaddition reaction between the azide and alkyne groups to generate a linker molecule comprising the polymerizable molecule and the amino acid reactive moiety . The conjugation of the polymerizable molecule 201 to the bifunctional linker 203 may occur at any useful or convenient step. In alternative examples (not shown), the bifunctional linker 203 may comprise an azide group, e.g., l-(2-azidoethyl)-4- isothiocyanatobenzene, which can be reacted to a polymerizable molecule 201 comprising an alkyne moiety.
[0160] FIG. 2B illustrates another example linker that can be used in sequencing polymeric analytes such as peptides. The linker comprises a guanidinylating group, which can react with and conjugate to an N-terminal amino acid with high efficiency under relatively mild conditions. The linker also comprises a spacer moiety that comprises a charged linker (ammonium) moiety that connects the guanidinylating group with a click chemistry moiety (modified tetrazine comprising dihydro-2H-pyran). The charged linker moiety may be useful in increasing the attraction (decreasing electrostatic repulsion) of the linker to another charged polymerizable molecule, e.g, a linking DNA molecule, and / or decreasing the translocation speed of the linker (or the modified amino acid comprisingthe linker) through a nanopore sequencer. The modified tetrazine can react with another molecule, e.g., a polymerizable molecule or capture moiety that comprises a transcyclooctene (TCO) moiety, in a highly efficient click chemistry reaction. Beneficially, the use of a linker such as shown in FIG. 2B can allow for efficient conjugation to an N-terminal amino acid and cleavage of the N-terminal amino acid under relatively mild conditions, as compared to traditional Edman degradation (e.g., using trifluoroacetic acid).
[0161] In some instances where the polymeric analyte comprises a peptide that comprises amino acid monomers, the coupling of the linker to an amino acid (e.g., NTAA or CTAA) changes the chemical structure of the amino acid. For example, if using a linker comprising an isothiocyanate moiety, the amino acid may be derivatized to a thiocarbamyl group (e.g., under mildly alkaline conditions) during or subsequent to contact with the isothiocyanate moiety. One or more further derivatizations may be performed. For instance, the amino acid or amino acid derivative (e.g., thiocarb amyl-derivatized amino acid) may be further derivatized to a thiazolone group (e.g., under acid conditions), a thiohydantoin group, or other chemical moiety. Similarly, a thiazolone group or thiohydantoin group may be further derivatized to a thiocarbamyl group.
[0162] In some instances, the polymerizable molecule comprises a linker. The linker may be used, for instance, for coupling a reactive moiety to the polymerizable molecule, which reactive moiety can react with that of another linker or with a monomer (e.g., terminal amino acid). In one such example, a polymerizable molecule, e.g., nucleic acid molecule, may comprise a linker that comprises a click chemistry moiety. The linker comprising the click chemistry moiety may be coupled to the polymerizable molecule using any useful approach, e.g., by incorporation of a linker-conjugated nucleotide or nucleoside and may be located at any useful position (e.g., at a 5’ end, at a 3’ end, in the center or other position of the polymerizable molecule). For example, a click-functionalized nucleotide or nucleoside, e.g., ethynyl deoxyuridine, octadiynyl deoxyuridine, can be incorporated into a DNA or RNA molecule. As such, the click chemistry moiety of the polymerizable molecule may then couple to another linker that comprises a complementary click chemistry moiety and also an amino acid reactive group (e.g., isothiocyanate, dansyl chloride, DNFB, xanthate, dithioester, guanidinylating agent, NHS ester, etc.). In another example, the polymerizable molecule may comprise afree amine (e.g., a modified nucleobase orbackbone with a free amine). In such an example, the free amine may be able to react with another linker that comprises, for example, an NHS ester group and also an amino acid reactive group. Alternatively, or in addition to, the polymerizable molecule may comprise the amino acid reactive group.
[0163] In some instances, the linker may be coupled to a backbone of a polymerizable molecule. For instance, the linker may form a phosphodiester or analogous bond with a DNA polymerizable molecule, or the linker may form an amide bond with a peptide polymerizable molecule.
[0164] The polymerizable molecule may comprise any useful number of linkers; for example, the polymerizable molecule may comprise a first linker that can couple to a second linker that can couple to a third linker that is coupled to or configure to couple to a monomer of a polymericanalyte, e.g., an amino acid or a peptide. Accordingly, the modified amino acid (e.g., comprising the cleaved amino and the polymerizable molecule), may comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more linkers.
[0165] The linker may comprise any number of spacing moieties, e.g., alkyl chains, polymer spacers (e.g., PEG), nucleic acid or oligo spacers, or other useful spacing moieties which may be useful in modulating the size or molecular weight of the linker. For example, the linker may comprise atleast 1, at least2, at least 3, at least4, at least 5, at least 6, at least 7, atleast 8, at least 9, at least 10, or a greater number of spacing moieties (e.g., hydrocarbon units, PEG units, nucleotides, spacer sequences etc.). The linker may comprise at most about 100, at most about 10, at most about 9, at most about 8, at most about 7, at most about 6, at most about 5, at most about 4, at most about 3, at most about 2, or at most 1 spacing moiety. The linker may comprise any useful number of functional groups, e.g., for attachment to multiple molecules. The linker may comprise atleast 1, at least2, at least 3, at least4, at least 5, at least 6, at least 7, atleast 8, at least 9, at least 10, or a greater number of functional groups.
[0166] The linker may be modulated to achieve a size range or relative size as compared to another molecule in the system. For instance, it may be advantageous to use a linker that has a ratiometric length or radius as compared to a nanopore or nanogap channel, e.g., to increase dwell time in a nanopore. Accordingly, the linker size or molecular weight may be adjusted to achieve an optical length or molecular weight, e.g., via addition of spacer moieties, monomers, or by linking multiple linkers together.
[0167] The linker may comprise a charged moiety (e.g., cation, anion, polyatomic ion, a nucleic acid molecule) or any other useful functional groups such as hydrophilic moieties, hydrophobic moieties, chelators, or lipophilic moieties. The linker may comprise a carboxylic moiety, a primary amine, a secondary amine, a tertiary amine, or a quaternary amine.
[0168] Intramolecular Expansion: The modified amino acid or derivative thereof may be produced from a peptide using an intramolecular expansion process, e.g., using one or more linkers and polymerizable molecules. In intramolecular expansion, individual amino acids, clusters of amino acids, or small peptides (e.g., a dipeptide, tripeptide, or quadripeptide) of a protein (e.g., a peptide or protein analyte) may be sequentially removed and re-tethered together, suchthatthe distance (e.g., a chemical distance suchasthe number of atomsor a physical distance) between the individual amino acids or clusters of amino acids is increased. Beneficially, performing an intramolecular expansion process of one or more amino acids of a peptide may obviate or overcome several issues with nanopore sequencing of peptides. For instance, by increasing the spacing between amino acids, fewer amino acids may enter the nanopore at a giveninstance, thereby reducing the quantity of superimposed signals arising from the number of amino acids within the nanopore. Further, increasing the spacing between amino acids may disrupt intramolecularinteractionsthatconvolvethe currentblockade signal, thereby allowing for higher- accuracy signal output from the nanopore. In some instances, the intramolecular expansion process may sufficiently separate the amino acids, such that only one amino acid or derivative thereof is present in a sensing region of a nanopore or nanogap in a given moment.
[0169] In an example, a method for intramolecular expansion may comprise providing a peptide comprising a plurality of amino acids, a linker (e.g., as described elsewhere herein), a polymerizable molecule (e.g., as described elsewhere herein), and a capture moiety. The linker may be configured to couple to (i) an amino acid (e.g., NTAA, CTAA, or an internal amino acid) of the peptide and (ii) the polymerizable molecule. The method may further comprise contacting the linker with the amino acid and the polymerizable molecule. Alternatively, or in addition, the linker may be provided pre-tethered to the polymerizable molecule and subsequently reacted with the amino acid. The linker may couple to the amino acid of the peptide to generate an amino acidlinker complex, which may or may not comprise the polymerizable molecule. In the instances that the amino acid-linker complex comprises the polymerizable molecule, the amino acid-linker complex may then couple to the capture moiety via the polymerizable molecule. For example, the polymerizable molecule and the capture moieties may both comprise nucleic acid molecules, which may be coupled via hybridization, ligation, an extension reaction, or combination thereof. In some instances, the method may further comprise, cleaving the amino acid from the peptide to yield a modified amino acid that comprises the amino acid-linker complex, and optionally repeating the process. In an example in which the process is repeated, an additional linker may be provided which is configured to couple to (i) an additional amino acid of the peptide (e.g., the n- 1 NTAA or n-1 CTAA) and (ii) an additional polymerizable molecule. Alternatively, the additional linker may be provided pre-coupled to the additional polymerizable molecule. The method may further comprise contacting the additional linker with the additional amino acid to generate an additional amino acid-linker complex. The additional polymerizable molecule may be coupled to the additional linker prior to, during, or subsequent to the coupling of the additional linker to the additional amino acid. The additional polymerizable molecule may be configured to couple to the (first) modified amino acid, e.g., via coupling of the polymerizable molecule and the additional polymerizable molecule. As such, in some examples, sub sequent to generation of the additional linker- additional amino acid complex, the additional linker-additional amino acid complex may couple to the modified amino acid, thereby generating a stacked plurality of modified amino acids (see, e.g., FIGs. 1 A-1B and ID). The additional amino acid may be cleavedfrom the peptide prior to, during, or sub sequent to generation of the stacked plurality of modified amino acids.
[0170] In some instances, intramolecular expansion of the peptide or protein may occur across a plurality of capture moieties. For example, a substrate may be provided that comprises a plurality of capture moieties, and, in some instances, the capture moieties are located adj acent to the peptide or protein. A first amino acid (e.g., n NTAA or n CTAA) of the peptide or protein may be coupled to a first capture moiety (e.g., via a first linker and a first polymerizable molecule), a second amino acid (e.g., n-1 NTAA or n-1 CTAA) may be coupled to a second capture moiety (e.g., via a second linker and a second polymerizable molecule), and a third amino acid (e.g., n-2 NTAA or n-2 CTAA) may be coupled to a third capture moiety (e.g., via a third linker and third polymerizable molecule). In another example, a first amino acid (e.g., n NTAA) may be coupled to a first capture moiety, a second amino acid (e.g., n-1 NTAA) may be coupled to the modified amino acid (e.g, from the n NTAA), thereby generating a stacked plurality of modified amino acids, and a third amino acid (e.g., n-2 NTAA) may be coupled to a second capture moiety. In some instances, the polymerizable molecule may comprise temporal information, e.g., abarcode on the round or cycle numberthatthe polymerizable molecule is provided. As will be appreciated, any numberof amino acids (or modified amino acids) may be coupled to any number of capture moieties.
[0171] In some instances, intramolecular expansion is performed without the use of a polymerizable molecule. For example, intramolecular expansion may be performed using a chemical expansion process, e.g., using amide skeletal elongation via amino acid insertion, as described by Z. Liu et al. 2023. Chemistry a European Journal. Vol 29, Issue 46, which is incorporated by reference herein in its entirety. In such instances, rather than coupling individual amino acids together along a polymerizable molecule backbone, a peptide may be intramolecularly expanded by chemical processing (see, e.g., FIG. 21).
[0172] A “modified amino acid” as used herein may refer to the amino acid-linker complex, the amino acid-linker-polymerizable molecule complex, or derivatives thereof. The modified amino acid may be used to refer to the amino acid-linker complex or the amino acid-linker- polymerizable molecule complex before or after cleavage. In some instances, the modified amino acid may refer to a portion of the amino acid-linker complex or the amino acid-linker- polymerizable molecule complex (e.g., justthe comprised aminoacid portion, justthe amino acidlinker complex portion, etc.).
[0173] Capture Moieties: A capture moiety may couple to the amino acid, the linker, the polymerizable molecule, or a combination thereof. The coupling of the amino acid, the linker, or the polymerizable molecule to the capture moiety may comprise a covalent interaction or anoncovalent interaction. The coupling may occur by interaction of binding pairs, e.g., biotin and avidin (or streptavidin), antigen or epitope and antibody or antibody fragment, cyclodextrins and small hydrophobic molecules (e.g., alkanes, benzene, polycyclics), cucurbiturils and adamantaneammonium or trimethylammoniomethyl ferrocene, cyclophane (e.g., calixarenes, cavitands, pillararenes, tetralactams), etc. In some embodiments, the coupling of the amino acid, the linker, or the polymerizable molecule to the capture moiety occurs through coupling of nucleic acid molecules (e.g., hybridization to one another or to a splint molecule, ligation, or a nucleic acid extension reaction).
[0174] In some instances, the capture moiety comprises an additional polymerizable molecule (e.g., a nucleic acid molecule or peptide). In one such example, both the polymerizable molecule of the modified amino acid and the capture moiety may comprise nucleic acid molecules. The nucleic acid molecules may be coupled to one another, e.g., via complementary base pairing directly or via a splint or bridge molecule and optional ligation (e.g., enzymatic or chemical ligation). In some instances, the splint molecule may comprise a hairpin nucleic acid molecule. Alternatively, or in addition to, the nucleic acid molecules may be coupled via a nucleic acid extension or amplification reaction. In some instances, the capture moiety and the polymerizable molecule comprise click chemistry moieties or reactive moieties which can allow for chemical ligation of the capture moiety to the polymerizable molecule. Non-exhaustive examples of chemical attachment of oligonucleotides can be found in M. Greenberg. Current Protocols in Nucleic Acid Chemistry. (2000). 1.4.1-4.5.19, which is incorporated by reference in its entirety.
[0175] The capture moiety or the polymerizable molecule, or, if applicable, a splint molecule, can comprise any naturally occurring, non-naturally occurring or engineered nucleotide base. For example, the capture moiety or the polymerizable molecule may comprise a nucleic acid molecule or analog thereof, which may comprise a pseudo-complementary base, a bridged nucleic acid, a xenonucleic acid, a locked nucleic acid, a peptide nucleic acid (PNA), a gamma-PNA, a morpholino, etc., as is described elsewhere herein. The capture moiety may comprise one or more functional sequences, including, but not limited to a priming sequence, sequencing sequence (e.g, P5 or P7 sequence), sequencing read sequence (e.g., R1 or R2 sequence), a protein binding site such as a mosaic end sequence, a transposase recognition sequence, a transcription factor binding site, a cleavage site (e.g., restriction site), a UMI, a blocking group, a spacer sequence, a barcode sequence, or other functional sequence. In some instances, the capture moiety comprises a cleavable or releasable moiety, e.g., a restriction enzyme recognition site, an abasic site, a uracil which can be cleaved using USER® or uracil DNA glycosylase, a disulfide bond that can be releasable upon addition of a reducing agent, a photocleavable moiety that can be cleaved using aphotostimulus, a thermolabile bond that is cleaved using a thermal stimulus, etc. In some instances, the capture moiety comprises a partial restriction site; e.g., the capture moiety may comprise a first partial restriction site and the polymerizable molecule may comprise a second partial restriction site; upon coupling or ligation of the polymerizable molecule to the capture moiety, the two partial restriction sites may generate a complete restriction site, such that the individual molecules (capture moiety and polymerizable molecule) are not cleavableby restriction digest individually but the ligated or coupled product is. In some instances, the capture moiety comprises a barcode sequence that comprises any useful information, e.g., the identity of the peptide that is to be analyzed, temporal information, spatial information, the origin of the peptide (e.g., from a sample, partition, protein, cell, experiment, etc.). The barcode may encode for information on a protein, cell type, demographic, patient, organism, cell state, phenotype, disease, age, sex, ancestral lineage, fitness or athletic performance, nutrient or metabolic state, paternity or kinship, drug-association, behavioral or cognitive state, pharmacokinetics, a synthetic molecule, etc.
[0176] In some instances, the capture moiety is provided coupled to a substrate or is configured to couple to a substrate (e.g., couple to an anchor moiety or molecule on the substrate). In one example, the substrate comprises one or more identical capture nucleic acid molecules; these identical capture nucleic acid molecules may act as a capture moiety for coupling to a polymerizable molecule of a modified amino acid, e.g., for a modified amino acid comprising a terminal amino acid, an n-1 amino acid, etc. In some instances, commercially available substrates, e.g., beads (e.g., DNA beads or barcoded beads), flow cells, or chips, e.g., Illumina® HiSeq, iSeq, MiniSeq, NextSeq, NovaSeq, etc. may be used as the substrates described herein. In some instances, the capture moietiesmay comprise additional useful sequences, e.g., primer sequences (e.g., P5 or P7 sequences) or read sequences (e.g., R1 or R2).
[0177] Alternatively, or in addition to, the capture moiety may be configured to couple to a substrate. The substrate may comprise one or more anchor molecules to which the capture moiety binds. Non-limiting examples of capture-anchor molecule binding pairs include biotinstreptavidin, nucleic acid coupling, and click chemistry pairs (e.g., azide-alkyne, TCO-tetrazine, etc.).
[0178] The capture moiety may be coupled to a substrate using any useful approach. In some instances, the capture moiety comprises a substrate-tethering group or linker or additional functional group. In some examples, the capture moiety comprises a nucleic acid molecule that comprises a substrate-tethering group, e.g., biotin, a click chemistry moiety such as an azide, that can couple to a substrate, e.g., a substrate comprising streptavidin or a complementary clickchemistry moiety that can react with that of the substrate-tethering group. Alternatively, or in addition to, the capture moiety may comprise a nucleic acid molecule, which may be coupled to a substrate-tethered nucleic acid molecule (e.g., an anchor nucleic acid molecule). The capture moiety may additionally comprise a binding sequence, to which another nucleic acid molecule (e.g., a polymerizable molecule that is part of or coupled to the modified amino acid). In some instances, the capture moiety comprises a single-stranded oligonucleotide or a single-stranded region in which a complementary oligonucleotide can hybridize. The complementary oligonucleotide may comprise a detectable label (e.g., fluorophore) that allows for detection of the capture moiety.
[0179] In some instances, the capture moiety maybe directly coupledto the polymeric analyte (e.g., peptide) that is to be analyzed or is undergoing intramolecular expansion. In such examples, the capture moiety may additionally comprise a nucleic acid barcode molecule that encodes the identity of the peptide or the originating sample or partition from which the peptide originated. The capture moiety may be coupled to any useful segment of the peptide, e.g., at a terminus (e.g, C-terminus) or at an internal residue (e. g. , at a side chain) or a derivatized or functionalized portion of the peptide, e.g., a functionalized C-terminus, a functionalized side chain, etc. Alternatively, or in addition to, the capture moiety may be provided in a solution and may not be coupled to the substrate or the peptide.
[0180] In some embodiments, the capture moiety is coupled to or configured to couple to both the polymeric analyte and the substrate. For example, the capture moiety may be coupled to the polymeric analyte and optionally comprise information (e.g., a barcode molecule) about the polymeric analyte or thatidentifiesthe polymeric analyte. The capturemoiety may also be coupled to or configured to couple to a substrate, e.g., via a linker, coupling of affinity or interactive pairs (e.g., biotin-streptavidin, antibody-antigen), nucleic acid coupling, etc. In one example, the capture moiety comprises a nucleic acid molecule which can be linked to a peptide analyte, e.g, at a terminus or a side chain of the peptide analyte using, for example click chemistry, linkers, or reaction of amine groups of lysine side chains or carboxyl group of aspartic acid or glutamic acid residues. The capture moiety may couple to a substrate comprising one or more anchor nucleic acid molecules, e.g., via hybridization or ligation. See, e.g., FIG. IB.
[0181] The capture moiety may be coupled to the polymeric analyte using any useful attachment approach, e.g., click chemistry moieties (e.g., alkyne-azide coupling), photoreactive groups (e.g., benzophenone), l-ethyl-3-(3-dimethylaminopropyl)carbodiimide hydrochloride (EDC) (e.g., to couple amino-oligosorpeptides),N-hydroxysulfosuccinimide(NHS), Sulfo-NHS, orNHS-esters (e.g., to couple sulfhydryl oligos), maleimides, hydrazines, hydroxyl amines, thiols,biotin-streptavidin interactions, cystamine, glutaraldehyde, formaldehyde, succinimidyl 4-(N- maleimidomethyl)cyclohexame-l-carboxylate (SMCC), Sulfo-SMCC, 4-(4, 6-Dimeth oxy- 1,3,5- triazin-2-yl)-4-methylmorpholinium chloride (DMTMM), silane (e.g., amino silanes), combinations thereof, etc. In some instances, the peptide or the capture moiety may be functionalized to comprise a coupling chemistry to couple the polymeric analyte to the capture moiety. In one non-limiting example, a polymeric analyte may comprise an alkyne such as dibenzocyclooctyne (DBCO), which may be configuredto react to an amine (e.g., DBCO-alcohol, DBCO-Boc, DBCO-NHS), a carboyxl or carbonyl (e.g., DBCO, DBCO-silane), a sulfhydryl, etc. An azide-functionalized capture moiety, e.g., a capture nucleic acid molecule may react with DBCO to link the polymeric analyte and the capture moiety. In other examples, linkers such as bifunctional linkers may be used to attach a polymeric analyte to a capture moiety; such bifunctional linkers may comprise the same reactive moiety on both ends or a different moiety at each end (e.g., heterobifunctional linker). Additional examples of linkers are described elsewhere herein.
[0182] In some examples, the capture moiety and the polymeric analyte (e.g., peptide) are coupled using a linker. In an example, terminal amino acid residue attachment can be achieved by reacting the peptide with a linker comprising (i) an amine-reactive group (e.g., isothiocyanates such as PITC, guanidinylating agents, dithioesters, xanthates, NHS esters, etc.) and (ii) a reactive group (e.g., click chemistry group). The linker can be, for example, PITC-conjugated click chemistry moieties. The linker reacts with and “blocks” the primary amines (e.g., modifies lysines), including the N-terminus. Subsequent cleavage of the N-terminal amino acid (e.g., using an Edman reagent, such as acid), can be performed, and one of the remaining modified lysines may be attached to the capture moiety (e.g. , usingthe click chemistry moiety coupled to the aminereactive group). Optionally, the peptide may be treated with a protease, e.g., LysC, which cleaves peptides such that a remaining peptide has a C-terminal lysine and suchthatthe remaining peptide comprises a primary amine only at the C-terminal lysine residue and the N-terminus. See also, Example 6 and¥G. 15. In some instances, the capture moiety may comprise the amine-reactive group, which can then couple to the N-terminal amine or to the amine of a lysine side chain. Additional examples of reactions for coupling the capture moiety and the polymeric peptide include, but are not limited to: thiol-based reactions (e.g., disulfide bonds with cysteine residues, Michael addition or retro-Michael reaction with maleimides or electrophilic alkenes), imine-based reactions (e.g., lysine residues formingreversible imines with aldehydes or ketones), oxime-based reactions (e.g., lysine residues forming oxime bonds with hydroxylamine derivatives or aminooxy compounds), hydrazone / oxime-based reactions (e.g., ketone or aldehyde moieties from N-terminal residues or modified side chains forming reversible hydrazone or oxime bonds with hydrazides or aminooxy compounds), boronic acid-based reactions (e.g., 1,2- or 1,3 -diol containing residues such as serine, threonine or tyrosine may react with boronic acids), and metalbased reactions (e.g., histidine or cysteine forming reversible coordination complexes with transition metals such as nickel, copper, zinc, etc.).
[0183] Cleaving: In some instances, generating the modified amino acid further comprises cleaving the amino acid or the amino acid-linker complex from the peptide. The cleaving of the amino acid or amino acid-linker complex may be achieved using any suitable mechanism, such as via application of a stimulus. The stimulus can be, for example, a chemical stimulus, a biological stimulus, a thermal stimulus (e.g., application of heat), a photo-stimulus, a physical or mechanical stimulus, or other type of stimulus or a combination of stimuli. In some instances, the stimulus comprises a chemical stimulus, e.g., a change in pH, application of an acid (e.g., trifluoroacetic acid, heptafluorobutyric acid, formic acid, phosphoric acid, acetic acid) or base, addition of a lytic agent, initiating agent, radical-generating agent, reducing agent, etc. In some instances, the chemical stimulus comprises application of a Lewis acid (e.g., boron triflate, boron trifluoride etherate, boron trichloride, boron tribromide, boron triiodide, or scandium triflate). In some instances, the stimulus comprises a biological stimulus, e.g., enzyme (e.g., Edmanase, protease, nuclease such as endonuclease or exonuclease) or ribozyme or DNAzyme that can cleave or catalyze cleavage of the amino acid or amino acid-linker complex. In some instances, the stimulus may be neutralized subsequent to the cleaving using any useful technique, e.g., neutralization of an acid, heat-killing of enzyme, nucleic acid removal or degradation, buffer exchange, or other approach.
[0184] In some examples, the methods provided herein may comprise using a linker comprising an amino acid reactive group (e.g., PITC, a xanthate, a guanidinylating agent, a dithioester or thiocarbamoyl) and coupling the amino acid reactive group of the linker with the amino acid and cleavingthe amino acid from the peptide using a stimulus (e.g., change in pH, temperature). In an example, the linker may comprise a PITC moiety that couples to an NTAA under mildly alkaline conditions to generate a phenylthiocarbamoyl (PTC) derivative of the NTAA, and cleavage of theNTAA from the peptide may be achievedusingan Edman degradation reaction (e.g., application of an acid such as trifluoroacetic acid or boron triflate, optionally with heat), to generate a thiazolinone (ATZ) derivative or a phenylthiohydantoin (PTH) derivative. As described elsewhere herein, the linker may comprise a moiety or molecule (e.g., polymerizable molecule such as a nucleic acid molecule) that can also couple to the capture moiety such that theamino acid-linker complex may be coupled to the capture moiety, thereby generating an amino acid-linker-capture moiety complex.
[0185] Given the harsh reaction conditions of standard Edman degradation, the polymerizable molecules described herein (e.g., nucleic acid molecules, peptides, lipids) etc. may comprise alterations or modifications to render them more resistant to the reaction conditions. For example, nucleic acid molecules may comprise predominantly pyrimidines (e.g., thymines, cytosines, uracils) which are more resistantto acid degradation and heatas compared to purines(e.g., adenine and guanine). For example, a nucleic acid molecule may comprise at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% thymines or cytosines. Alternatively, or in addition to, canonical nucleotides may be substituted or may comprise acid-resistant nucleotide analogs, e.g., hexitol nucleic acids.
[0186] Alternative degradation chemistries may also be employed. Milder degradation under basic conditions forN-terminal amino acid removal can include the use of triethylamine acetate in acetonitrile or other solvent such as water, N, N-dimethylformamide (DMF), or a mixture of solvents. Alternatively, degradation may be achieved using a thioacylation approach, the use of milder acid reagents, e.g., trichloroacetic acid (pKa of 0.66) or dichloroacetic acid (pKa of 1 .35), acetic acid, or alternative basic reaction conditions, e.g., using acid-base pairs such as N, N- Diisopropylethylamine (DIPEA), pyridine, acetic acid derivatives, etc. In some instances, use of a base (alkaline conditions) may be sufficient for cleavage. For instance, for linkers comprising guanidinylating agents, use of a mild base can be used to cleave amino acid-linker complex.
[0187] C-terminal degradation strategies are also provided herein. C-terminal degradation may comprise Edman-like degradation approaches. C-terminal degradation may employ the use of activatingreagents that react with the C-terminal carboxyl group of a peptide, and a derivatizing agent (e.g., a thiocyanate to generate a peptide-thiocyanate or peptide-thiohydantoin). Nonlimiting examples of activating reagents include acetyl chloride and acetic anhydride. Alternatively, or in addition to, single-step C-terminal derivatization of a peptide to a peptidyl- thiohydantoin may be performed, e.g., using Schlack-Kumpf approach, in which a peptide is reacted with thiocyanic acid (e.g., in acetone) to generate a peptidyl-thiohydantoin. The peptide- thiohydantoin may be cleaved, e.g., using basic conditions, to generate an amino acid thiohydantoin and remaining peptide.
[0188] Cleavage of amino acids may also be achieved using enzymatic or enzyme-analog (e.g., ribozyme or DNAzyme) approaches. Example enzymatic cleavage may include the use of Edmanases (e.g., modified cruzain), aminopeptidases (e.g., Pfu aminopeptidase I, PhTET aminopeptidases, P. horikoshii aminopeptidases), metalloenzymatic aminopeptidases,acylpeptide hydrolases, tRNA synthetases, endopeptidases, carb oxy peptidases, and the like. The enzymes or ribozymes or DNAzymes may be modified or engineered to recognize a modified amino acid, e.g., an amino acid that has a chemical moiety attached thereto (e.g., PITC, NITC, dansyl chloride, SNFB, DNP, SNP, guanidinyl group, biotin, streptavidin, nucleic acid molecules, lipids, carbohydrates, acetyl groups, acyl groups, guandinylation agents, etc.).
[0189] One or more reactions may be accelerated by application of energy or radiation, e.g, electromagnetic radiation. For example, degradation or cleavage of the terminal amino acid of a peptide may be facilitated by applying microwave energy to accelerate the reaction kinetics. For example, hydrolysis of proteins may be facilitated by application of microwave energy, e.g., as described in Margolis et al., 1991, Journal of Automatic Chemistry. Vol 13, No. 3, pp 93-95, which is incorporated by reference herein.
[0190] Additional processingofthe cleavedamino acids may also be performed. For instance, in some instances, cleavage of the amino acid may result in generation of stereoisomers. The cleaved products maybe treated to remove the stereocenter during the cleavage reaction. In some instances, stereoisomers may be enzymatically converted to a single isomer, e.g., using an isomerase.
[0191] In some instances, more than one amino acid may be cleaved from the peptide per cleavage event. The cleaving may comprise cleaving 2 amino acids, 3 amino acids, 4 amino acids, 5 amino acids, 6 amino acids, 7 amino acids, 8 amino acids, 9 amino acids, 10 amino acids, or more. For example, the polymeric analyte may comprise a peptide comprising a plurality of amino acids, and single amino acids, di-peptides, tri-peptides, quadri-peptides, or larger may be cleaved in the methods described herein. In some instances, at most about 10 amino acids, at most about 9 amino acids, at most about 8 amino acids, at most about 7 amino acids, at most about 6 amino acids, at most about 5 amino acids, at most about 4 amino acids, at most about 3 amino acids, or fewer amino acids may be cleaved in a given cleavage event. In some instances, cleavage of greater than one amino acid may be mediated using an enzyme (e.g., Edmanase, protease) or ribozyme or DNAzyme that is capable of recognizing or cleaving more than a single amino acid.
[0192] Cleavage of the amino acid from the peptide may be conducted using a biological stimulus, such as an enzyme or ribozyme or DNAzyme. The enzyme can be any useful cleaving enzyme, e.g., a protease, such as an Edmanase, hydrolase, lyase, transferase, cruzain, a cleaving protein (e.g., ClpS, ClpX), Proteinase K, exopeptidase, aminopeptidase, diaminopeptidase, serine protease, cysteine protease, threonine protease, aspartic protease, aspartic protease, glutamic protease, metalloprotease, asparagine peptide lyase, pepsin, trypsin, pancreatin, Lys-C, Arg-C, Glu-C, Asp-N, chymotrypsin, carboxypeptidase (e.g., carboxypeptidase A, carboxypeptidase B,carb oxy peptidase Y), SUMO protease, elastase, papain, endoproteinase, proteinase, TrypZean®, bromelain, collagenase, hyaluronase, therm oly sin, ficin, keratinase, tryptase, fibroblast activation, enterokinase, chymotrypsinogen, chymase, clostripain, calpain, alpha-lytic protease, proline specific endopeptidase, furin, thrombin, subtilisin, genenase, PCSK9, cathepsin, prolidase, methionine aminopeptidase, cathepsin C, 1-cyclohexen-l-yl-boronic acid pinacol ester, pyroglutamate aminopeptidase, renin, kininogen, kallikrein,DPPIV / CD26, thimet oligopeptidase, prolyl oligopeptidase, leucine aminopeptidase, dipeptidylpeptidase, or other enzyme or protease, or a combination or variation (e.g., engineered mutant or variant) thereof. In some instances, the cleaving enzyme or ribozyme orDNAzymemay be configured or engineered to cleave a terminal amino acid or plurality of amino acids; alternatively, the cleaving enzyme or ribozyme or DNAzyme may be configured or engineered to cleave off-site at a non-terminal location of the peptide, e.g., at an internal amino acid at an n-1, n-2, n-3, n-4, n-5, n-6, n-7, n-8, n-9, n-10, etc. position, where n is the number of amino acids in the peptide.
[0193] In the instances of enzymatic cleavage, additional reagents may be providedto catalyze or induce the cleavage. For instance, metalloproteases, aminopeptidases, or exopeptidases may facilitate cleavage of an amino acid or plurality of amino acids in the presence of a catalyst, e.g, metal or metal ion (e.g., cobalt). Accordingly, a catalyst may be provided in order to facilitate the binding of the enzyme to an amino acid or the subsequent cleavage of the amino acid from the peptide. In some examples, cleavage may be mediated by an apo-enzyme, which is inactive in the absence of a metal catalyst of cofactor, and cleavage may be controlled by addition of metal or metal ions.
[0194] Other examples of cleaving stimuli include: a photo stimulus (e.g., application of UV, X-rays, gamma rays, or other wavelength of light), mechanical stimulus (e.g., sonication, high pressure, electromagnetic energy), thermal stimulus (e.g., application of heat), or chemical stimulus. In some instances, the peptide may comprise or be altered to comprise a cleavable or labile bond that can be cleaved upon application of the appropriate stimulus, e.g., disulfide bonds (e.g., cleavable upon application of a chemical stimulus such as a reducing agent), ester linkages (e.g., cleavable with a change of pH), a vicinal-diol linkage (e.g., cleavable with sodium periodate), a Diels-Alder linkage (e.g., cleavable upon application of heat), a sulfonelinkage (e.g, cleavable via a base), a silyl ether linkage (e.g., cleavable via an acid), a glycosidic linkage (e.g, cleavable via an amylase), a peptide linkage (e.g., cleavable via a protease), or a phosphodiester linkage (e.g., cleavable via a nuclease (e.g., DNase)).
[0195] In some instances, the capturemoiety may be cleaved from the peptide orthe substrate. The cleaving may occur at any useful or convenient step, e.g., after generation of the modifiedamino acid or stacked plurality of modified amino acids. In some instances, cleavage of the capture moiety may occur subsequent to the formation of a stacked plurality of modified amino acids, and the cleaved product may be sequenced, e.g., using nanopore sequencing or imaging approaches described elsewhere herein.
[0196] Nanopore s / Nanogaps: Characterization of the modified amino acid may be performed using nanoscale technologies such as nanopores, nanogaps, or nanochannels. In some instances, a nanopore, nanogap, or nanochannel may be provided on a membrane in an ionic solution. In commercially available nanopore sequencers, individual analytes enter the nanopore under an applied electric potential, thereby altering the flow of ions through the nanopore in a timedependent manner. Measurement of the modulation of the ionic current as the individual analytes translocate through the nanopore can be performed, and the measured signal can be computationally decoded, e.g., to yield a DNA sequence of a DNA analyte. Similarly, a nanopore sequencer may be used to characterize or analyze the detectable complexes described herein, e.g, modified amino acids or stacked pluralities of modified amino acids. Additional example methods and systems for nanopore sequencing of intramolecularly expanded peptides can be found in International Patent App. No. PCT / US2023 / 071456, which is incorporated by reference herein.
[0197] A signal may be measured from the nanopore, membrane, or surrounding solution. For example, a conductance, current, current blockage, current density, current change, voltage, impedance, resistance, inductance, capacitance, frequency, phase, power, electric field, magnetic field, or other parameter within the nanopore may be monitored as a function of time. In some instances, the current signal may be an ionic current signal, a cross-pore or transverse-to-pore drain current, or a source current. As a molecule enters the nanopore, nanogap, or nanochannel, a change in the conductance, current, impedance, or other parameter may occur and provide information (e.g., size, charge, aspect ratio, volume, hydrophobicity, chemical structure) on the molecule. Each amino acid or modified amino acid, or a subset of amino acids or modified amino acids, may generate a unique signal that is distinguishable from other amino acids or modified amino acids. As such, the unique signal signatures may be assigned to the amino acids or modified amino acids in order to determine the identity of the amino acids, polymerizable molecules, or both.
[0198] In some instances, a modified amino acid (e.g., an individual modified amino acid or a modified amino acid comprised by a stacked plurality of modified amino acids) may b e analyzed numerous times, e.g., via translocation and measuring of a current or conductance of the modified amino acid, through the same or different nanopore or nanogap. For example, iterative reading of a modified amino acid may be beneficial in improving the accuracy of the reads or identificationof the modified amino acid (or polymerizable molecule). In such cases, the modified amino acid may be translocated through one or more nanopores at least 2 times, at least 3 times, at least 4 times, at least 5 times, at least 6 times, at least 7 times, at least 8 times, at least 9 times, at least 10 times, at least 20 times, at least 50 times, at least 100 times, at least 200 times, or even greater. Alternatively, or additionally, the modified amino acid may be ratcheted back and forth within the nanopore, e.g., using a processive enzyme such as phi29 polymerase such as described in Cherf et al. 2012. Nat. Biotechnol. 30(4):344-348, which is incorporated by reference herein in its entirety. Other enzymes or motor proteins such as helicases, unfoldases, or polymerases can also be modulated or engineered to perform iterative reading of the modified amino acid or polymerizable molecule. For instance, a motor protein may be altered or modified to induce slipping of the motor protein as it translocates along the modified amino acid (or polymerizable molecule). Alternatively, or in addition to, “flossing” of the modified amino acid may be performed, e.g., using a dual nanopore device and competing voltage forces, e.g., as described in Liu et al. Small. 2020. 16(3):el 905379. Flossing of the modified amino acid may also be performed by reversing the electric field, modulating the electric current flow (e.g., switching from alternating current to direct current or from direct current to alternating current), applying a magnetic field, modulating impedance, by inclusion of structural elements on the modified amino acids (e.g., comprised by or attached to the polymerizable molecules) such as DNA hairpins, hybridizing primers (e.g., for generating double-stranded regions), dumbbell shaped complexes, e.g., as described by Kasianowicz. 2004 Nature Materials. 3, 355-356., use of magnetic particles (e.g., attached to a portion or an end of the modified amino acid or stacked plurality of modified amino acids), and the like. Alternatively, or in addition to, slowing or stuttering of the modified amino acid during translocation through the nanopore may be performed, for instance, via introducing noncan onical or non-incorporable nucleotides (e.g., sulfur modified nucleotides). In such instances, a polymerase maybe used as a molecular motor protein to ratchet the modified amino acid through the nanopore, and the addition of the non-incorporable nucleotides may change the translocation speed (e.g., slow down) or introduce stutter, skipping, backstepping, or other useful ratcheting manner of the modified amino acid.
[0199] In some instances, a peptide or portion thereof may be analyzed using a nanopore or nanogap. For example, a peptide may be provided, and a modified amino acid or stacked plurality of modified amino acids may be generated from the peptide using the methods described herein. The remaining peptide, e.g., after any number of cycles of generating modified amino acids or a stacked plurality of modified amino acids, may be analyzed using the nanopore or nanogap (see, e.g., FIG. IE). The signal (e.g., from the current or other parameter such as impedance) may bemeasured, which may output information such as size (e.g., molecular weight, number of amino acids, hydrodynamic radius, length), conformation, charge, polarity, hydrophobicity, hydrophilicity, composition (e.g., amino acid sequence) or other parameter about the remaining peptide. In some instances, the peptide may notbe subjected to intramolecular expansion and may be directly translocated through a nanopore or nanogap for signal measurement and analysis.
[0200] In another example, a peptide may be provided, digested or fragmented, and then the digested or fragmented peptides or individual amino acids can be analyzed using a nanopore or nanogap. The peptide may be digested or fragmented using any suitable approach, e.g., enzymatic approach (e.g., using a protease or proteasome, LysC), mechanical approach (e.g., sonication, mechanical shearing), chemical approach (e.g., using an Edman degradation reaction), or other useful approach. The peptide may be digested or fragmented into smaller length peptides or individual amino acids. The fragmented or digested product may then translocate through a nanopore or nanogap and be analyzed, e.g., via measurement of a signal (e.g., current) and using the measured signal to identify or characterize the fragmented or digested product. In some instances, the fragmented or digested product may be contacted with a helper molecule, which may assist with translocatingthe fragmented or digested product through the nanopore. The helper molecule may comprise, for example, a charged molecule (e.g., a DNA molecule) and a linker that comprises an amino acid reactive moiety (e.g., an amine-reactive moiety such as isothiocyanate, guanidinyl group, dithioester, xanthate, etc. or a carboxyl-reactive moiety such as carbodiimide, EDC, etc.). The amino acid reactive group may be used to couple the helper molecule to the fragmented or digested product and facilitate translocation (e.g., via electrophoresis) of the helper molecule-digested product complex through the nanopore.
[0201] The nanopore, nanogap, or nanochannel may be generated from an organic or inorganic material and may comprise a protein, e.g., a pore-forming protein or a transmembrane protein. Such a protein may be naturally occurring, synthetic, or engineered. Examples of naturally occurring organic nanopores include wild-type aerolysin, alpha-hemolysin, mycobacterial porins (e.g., MspA porin), Phi29 connector channels, Fragaceatoxin C, Cytolysin A, Ferric hydroxyamate uptake component A, Curb specific gene G, outer membrane porin G, viral DNA packaging motors, etc. Alternatively, or in addition to, the nanopore, nanochannel, or nanogap may comprise an engineered variant of a naturally occurring nanopore. Such engineered variants can be generated, for example, using protein engineering approaches (e.g., directed evolution approaches such as yeast-surface, bacterial-surface, or phage display, ribosomal display, CIS display, droplet directed evolution, etc.), genetic engineering, or chemical modification (e.g, using maleimide reaction with cysteine residues, isothiocyanate or NHS ester reaction with lysineresidues). In some instances, the engineered variant may comprise a chemical modification such as addition of a chelator molecule (e.g., nickel), e.g., as described by Wang et al. 2023. Nature Methods. 21, 92-101 or may be otherwise functionalized. The modification may be designed at any useful region of the nanopore, e.g., the lumen or pore constriction area. In some instances, a nanopore may be engineered to change the pore or lumen size.
[0202] In some embodiments, the nanopore, nanogap, or nanochannel is comprised of an inorganic material. For example, solid-state nanoporesmay be madefrom dielectric materials such as a silicon compound (e.g., silicon nitride, silicon dioxide), an aluminum compound (e.g., aluminum oxide), a titanium compound (e.g., titanium oxide), a molybdenum compound (e.g, molybdenum sulfate), hafnium, graphene, etc. The nanopore, nanochannel, or nanogap may assume any useful form factor or geometry, e.g., gaps or channels within membranes, capillaries, etc., and may be generated using any suitable process, e.g., ion-beam sculpting, electron beam exposure. The nanopore, nanochannel, or nanogap may comprise an elastomeric material.
[0203] In some instances, the nanopore, nanogap, or nanochannel is coupled to or configured to couple to a protein. The protein may be a molecular motor, which may facilitate movement of the modified amino acid or portion thereof (e.g., the polymerizable material) through the nanopore. In a non-limiting example, the modified amino acid may comprise an amino acid coupled, optionally via a linker, to a nucleic acid molecule, and the nanopore may be coupled to a ratcheting enzyme such as a helicase or a molecule comprising helicase activity. The helicase may be used to translocate (ratchet) the nucleic acid molecule and the coupled amino acid through the nanopore. Non-limiting examples of helicases include PcrAl , Rep A, UvrD, Dda, HSV UL5, HSV UL9, DnaB, PriA, T7gp4A and 4B, T4gp41, SV40, TAG, Polyoma TAG, BPV El, MCM 4 / 6 / 7, Dna2, FFA-1, RecD, Tral, NS3, RecQL4, UvrD, UvrAB, PcrA, Rad3, helicase E, XPD, XPB, Dna2, RecD2, BACH1, HDH II, RecQ, WRN, Rtell, BLM, RuvB, Mph 1, CHD4, CMG helicase, RecBCD, RecG, RecQ, RuvAB, PriA, UvrD, T4 UvsW, HDH II, HDH IV, WRN, Tra I, Rho, PDH65, BLM, Srs2, Sgsl, Rtell, SWI2, SNF2, TFIIH, Rho, Factor 2, TRCF, RecQL5, ERCC6 / RAD26, HSVUL5 , eIF4 A, RHA, Ded 1 p, vasa, Rad54, ATRX, BLM, CHD4, Pif 1 , Dna2, Rtel 1, WRN, BLM, FANC, Dna2, Pifl, WRN, and Hel308.
[0204] In some instances, the molecular motor may increase, decrease, or otherwise change the translocation speed of the modified amino acid through the nanopore as compared to an unmodified amino acid. In other examples, the molecular motor may comprise a topoisomerase, a polymerase, a nuclease (e.g., endonucleases such as restriction endonucleases or Cas proteins, or exonuclease) an unfoldase, mitotic spindle protein (e.g., nuclear mitotic apparatus protein, kinesin, dynein), or other motor protein (e.g., myosin), a G-protein coupled receptor or signalingprotein, or a combination or a variant thereof, or other poly mer-processing protein. In some instances, the protein comprises a protease (e.g., exopeptidase, Edmanase) or proteosome, which can enable “chop-n-drop” or cleaving of the modified amino acid or portion thereof prior to translocation in the nanopore. In some instances, the protein may not be coupled to the nanopore, nanogap, or nanochannel but may be adjacent or in proximity to the nanopore, nanogap, or nanochannel. For instance, the protein and the nanopore may be provided separately in an array (e.g., a planar substrate, a microwell or nanowell array), such that the protein may contact the modified amino acid (or stacked plurality of modified amino acids) in proximity to the nanopore, nanogap, or nanochannel.
[0205] In some instances, the modified amino acid or the stacked plurality of modified amino acids may comprise a nucleic acid molecule comprising non-canonical nucleotides, which can be processed by a motor protein adjacent to the nanopore, nanogap, or nanochannel. In one such example, the modified amino acid or stacked plurality of modified amino acids may comprise one or more nucleotide hexaphosphate groups. The modified amino acid or stacked plurality of modified amino acids may be contacted with a molecular motor protein that accepts hexaphosphates (e.g., an exonuclease, a 5’ exo polymerase such as Bst or engineered variant thereof, a DNA polymerase from G. kaustophillus or engineered variant thereof) and which, in some instances, may cleave the nucleotide comprising the amino acid portion, thereby generated a liberated portion of a modified amino acid. The liberated portion of the modified amino acid may then translocate through the nanopore, nanogap or nanochannel and be detected.
[0206] Alternatively, or in addition to, translocation of the molecules described herein (e.g, modified amino acids, polymerizable molecules, etc.) through a medium or through a nanopore or nanogap may be facilitated by application of a force. For example, a molecule may be translocated by application of pressure (e.g., pressure-driven flow), an electric field (e.g., via electrophoresis, electroosmotic flow, isoelectric focusing), a magnetic field (e.g., using a ferromagnetic fluid, magnetic particles), light (e.g., optoelectronics, optical tweezers), hydrodynamic force, centrifugal force, or gravitational force.
[0207] A commercially available nanopore system may be used in the methods described herein. For example, a nanopore system from Oxford Nanopore Technologies (ONT), such as the MinlON, VolTRAX, GridlON, PromethlON, MinIT, Flongle, or Q-Line products maybe used to identify and characterize the modified amino acids described herein. A solid-state nanopore system may be used, e.g., Northern Nanopore Instruments, Norcada, or IMEC.
[0208] In one example, a method provided herein may comprise providing a peptide, and generating a modified amino acid from the peptide. The generating the modified amino acid maycomprise providing a linker and a polymerizable molecule, coupling the linker to the amino acid of the peptide and to the polymerizable molecule, thereby generating an amino acid-linker complex, and cleaving the amino acid from the peptide, thereby generating the modified amino acid comprising the linker and the polymerizable molecule. In some instances, the amino acidlinker complex or the optionally-cleaved modified amino acid is coupled to a capture moiety. One or more operations of the process may be repeated, thereby generating a stacked plurality of modified amino acids comprising a plurality of stacked polymerizable molecules. The modified amino acids or stacked plurality of modified amino acids may then be analyzed.
[0209] Imaging Detection: In some instances, the modified amino acid is characterized or identified using a binding agent comprising a detectable label and detecting the detectable label, thereby identifying an amino acid type of the modified amino acid. In some embodiments, the modified amino acid is analyzed using super-resolution imaging. In some embodiments, the method may comprise attaching a portion of the modified amino acid to a substrate, linearizing the modified amino acid or portion thereof, and attaching an additional portion of the modified amino acid to the substrate. In some instances, the method comprises attaching a first sequence of a DNA molecule or modified DNA molecule (e.g., a DNA molecule comprised by a modified amino acid) to a substrate; linearizing the DNA molecule and attaching a second sequence of the DNA molecule or the modified DNA molecule to the substrate.
[0210] Additional methods and systems of processing polymeric analytes such as peptides are also described in U.S. Pat. No. 11,499,979, International Pat. App. Nos. PCT / US2022 / 081392, PCT / US2023 / 017954 and PCT / US2023 / 071456, and U.S. Pat. App. No. 18 / 740,088, filed June 11, 2024, each of which is incorporated by reference herein in its entirety.
[0211] FIGs. 1A-1D schematically show example workflows for generating a modified monomer from a polymeric analyte (e.g., a modified amino acid from a peptide) either on a substrate (FIG. 1A) or with or without a substrate (FIG. IB and FIG. 1C). In workflow 100a of FIG. 1A, a polymeric analyte 103 (e.g., a peptide) and a capture moiety 105 are provided, which optionally are coupled to a substrate 101 . The capture moiety 105 may comprise a first nucleic acid molecule (e.g., DNA). In process 106, a linker 109 and a polymerizable molecule, e.g., a linking nucleic acid molecule 111, are provided. In some instances, the linker 109 is pre-tethered to the polymerizable molecule (depicted as a linking nucleic acid molecule 111); alternatively, the linker 109 and the polymerizable molecule maybe provided separately. In process 106, the linker 109 may couple to a monomer, e.g., an amino acid (e.g., NTAA) of the polymeric analyte 103 (e.g., a peptide) to generate a monomer-linker complex. In process 112, the monomer-linker complex may couple to the capture moiety 105, thereby generating a monomer-capture moietycomplex. Coupling of the monomer-linker complex to the capture moiety 105 may be mediated by the polymerizable molecule, e.g., the linking nucleic acid molecule 111. Optionally, the monomer-linker complex and the capture moiety 105 may be covalently linked together (e.g., the linking nucleic acid molecule 111 may be covalently linked to the capture moiety 105), using chemical (e.g., click chemistry) or enzymatic (e.g., a ligase) approaches. Alternatively, or in addition to, the polymerizable molecule may comprise a first sequence that is complementary and may hybridize to a second sequence of the capture moiety 105 (not shown), or the polymerizable molecule may be linked to the capture moiety 105 via a splint or bridge molecule, which may comprise sequences that are complementary to the first sequence of the polymerizable molecule and the second sequence of the capture moiety 105 (not shown). In process 113, the monomer may be cleaved from the polymeric analyte 103, thereby providing a modified monomer that comprises the cleaved monomer-capture moiety complex; the modified monomer may comprise the cleaved monomer coupled to the linker 109, the polymerizable molecule (shown as a linking nucleic acid molecule 111), the capture moiety 105, or a combination thereof (e.g., the modified monomer may comprise the cleaved monomer, the linker, and the polymerizable molecule).
[0212] Any of the processes, e.g., 106, 112, and 113 may be iterated and repeated any number of times (“rounds”) using additional linkers 109 and polymerizable molecules (e.g., linking nucleic acid molecules optionally comprising cycle / round information), and tethering the additional polymerizable molecules together (e.g., tethering an additional polymerizable molecule to the polymerizable molecule of the monomer-capture moiety complex). Multiple rounds may be performed until all or a subset of the monomers in the polymeric analyte 103 are cleaved and tethered together. In some instances, processes 106, 112, and 113 may be iterated to generate a stacked plurality of modified amino acids 123 comprising a set of cleaved monomers, e.g., a concatenated set of modified monomers that each comprise a polymerizable molecule coupled thereto. The stacked plurality of modified amino acids 123 may comprise a stacked set of polymerizable molecules (e.g., via coupling of the polymerizable molecule from a second round to the polymerizable molecule from a first round and coupling of the polymerizable molecule from the third round to that of the second round, and so on) from the individual modified monomers. The stacked plurality of modified amino acids 123 may comprise polymerizable molecules that are coupled or concatenated in a linear fashion (e.g., polymerizable molecules that form a linear polymerizable molecule “backbone”), or in a non-linear fashion (e.g., random or semi-random arrangement, branched coupling). The stacked plurality of modified amino acids 123 may be circularized; for example, the stacked plurality of modified amino acids 123 forming a linear polymerizable backbone may be joined at the ends to form a circularized product. In someinstances, the stacked plurality of amino acids 123 may be cleaved from the substrate. For example, the capture moiety 105 or the polymerizable molecule (e.g., linking nucleic acid molecule 111) may comprise a restriction or cleavable site that can be cleaved upon addition of the proper cleaving reagent, e.g., a restriction enzyme, a reducing agent (for disulfide bonds), etc. Alternatively, or in addition to, the stacked set of polymerizable molecules may be subject to amplification to generate amplicons of the polymerizable molecules coupled to or comprised by the stacked plurality of modified amino acids. The stacked plurality of modified amino acids 123 may be subjected to analysis or characterization, e.g., by contacting the stacked plurality of modified amino acids 123 with a library of binding agents, which can bind to their respective monomer targets (e.g., an amino acid type), and detecting the binding agents, or by using a nanopore sequencer.
[0213] FIG. IB schematically illustrates another example workflow of generating a modified monomer, e.g., modified amino acid, in presence or absence of a substrate. In such an example workflow 100b, a polymeric analyte 103, e.g., a peptide and a capture moiety 105 are provided. The capture moiety 105 may comprise a first nucleic acid molecule (e.g., DNA molecule) and may comprise identifying information of the polymeric analyte 103, e.g., an identifying barcode sequence. The capture moiety 105 may additionally comprise a releasable or cleavable moiety. The polymeric analyte and the capture moiety 105 may, in some instances, be coupled to a substrate (FIG. IB inset), e.g., via an anchor molecule (e.g., anchor nucleic acid molecule). The polymeric analyte may be coupled to the capture moiety, either directly or indirectly (e.g., both the polymeric analyte and the capture moiety may be coupled to a substrate). For instance, the capture moiety 105 may comprise a nucleic acid sequence that is complementary to a sequence on a substrate (e.g., bead, flat surface), or the capture moiety may be coupled to the substrate via a linker (not shown). Alternatively, the capture moiety 105 may comprise a nucleic acid sequence that is partially complementary to a splint oligonucleotide that is also partially complementary to an anchor oligonucleotide on the substrate, and proximity ligation (e.g., using ligase or chemical ligation) may be performed to couple the capture moiety 105 to the substrate (not shown). In process 106, a linker 109 and polymerizable molecule, such as a linking nucleic acid molecule 111 , are provided. In some instances, the linker 109 is pre-tethered to the polymerizable molecule (linking nucleic acid molecule 111); alternatively, the linker 109 and the polymerizable molecule (linking nucleic acid molecule 111) may be provided separately. The polymerizable molecule may comprise identifying temporal information, e.g., the cycle or round in which it is provided. In process 106, the linker 109 may couple to a monomer, e.g., an amino acid (e.g., NTAA) of the polymeric analyte 103 (e.g., peptide) to generate a monomer-linker complex. In process 112, themonomer-linker complex may couple to the capture moiety 105. Coupling of the monomer-linker complex to the capture moiety 105 may be mediatedby the polymerizable molecule and optionally an additional polymerizable molecule 116. In some instances, the additional polymerizable molecule 116 may comprise a cleavable tag moiety, which can enable rapid and precise purification. For example, the additional polymerizable molecule 116 may comprise a photocleavable biotin moiety, which may allow for purification of the monomer-linker complex using streptavidin pulldown. Subsequently, the photocleavable biotin moiety maybe cleaved and removed at any convenient or useful step. Alternatively, the capture moiety 105 may be directly hybridized or ligated to the polymerizable molecule (linking nucleic acid molecule 111). Optionally, the monomer-linker complex and the capture moiety may be covalently linked together (e.g., via ligation). Alternatively, or in addition to, the polymerizable molecule may comprise a first sequence that is complementary and may hybridize to a second sequence of the capture moiety 105 (not shown), or the polymerizable molecule may be linked to the capture moiety 105 via a splint or bridge molecule, which may comprise sequences that are complementary to the first sequence of the polymerizable molecule and the second sequence of the capture moiety 105 (not shown). In process 113, the monomer may be cleaved from the polymeric analyte 103 to generate the modified monomer that comprises the monomer-capture moiety complex. The modified monomer may comprise the cleaved monomer, the linker 109, the polymerizable molecule (e.g., linking nucleic acid molecule 111), the capture moiety 105, or a combination thereof (e.g., the cleaved monomer, the linker, and the polymerizable molecule, just the cleaved monomer, or just the cleaved monomer-linker complex). Processes 106, 112, and 113 may be iterated and repeated any number of times (“rounds”) using additional linkers 109 and polymerizable molecules, and tethering the additional polymerizable molecules together (e.g., tethering an additional (e.g., second) polymerizable molecule to the monomer-capture moiety complex, tethering a polymerizable molecule from the third round to that of the second round, and so on). Multiple rounds may continue until all or a subset of the monomers of the polymeric analyte 103 are tethered together. For example, the process may be iterated to generate a stacked plurality of modified monomers (e.g., stacked plurality of amino acids 123) comprising a set of concatenated modified monomers, e.g., concatenated monomer-linker-polymerizable complexes. The polymerizable molecules (e.g., linking nucleic acid molecules 111) of the stacked plurality of modified monomers (e.g., stacked plurality of amino acids 123) may beidentical molecules (e.g, same nucleic acid sequence), or they may be different. The stacked plurality of modified amino acids 123 may comprise polymerizable molecules that are coupled or concatenated in a linear fashion (e.g., polymerizable molecules that form a linear polymerizable molecule “backbone”),or in a non-linear fashion (e.g., random or semi-random arrangement, branched coupling). The stacked plurality of modified amino acids 123 may be circularized; for example, the stacked plurality of modified amino acids 123 forming a linear polymerizable backbone may be joined at the ends to form a circularized product. After any useful number of rounds, the stacked plurality of modified monomers (e.g., stacked plurality of modified amino acids 123) may be cleaved from or at the capture moiety 105, e.g., using the cleavable moiety. Alternatively, or in addition to, the stacked set of polymerizable molecules may be subject to amplification to generate amplicons of the polymerizable molecules coupled to or comprised by the stacked plurality of modified amino acids. In instances where a substrate is used, the stacked plurality of modified amino acids may be removed from the substrate, e.g., via cleavage of the capture moiety or dehybridization (e.g, of a nucleic acid capture moiety that is annealed to an anchor nucleic acid molecule of the substrate). The stacked plurality of modified amino acids 123 may be subjected to analysis or characterization, e.g., by contacting the stacked plurality of modified amino acids 123 with a library of binding agents, which can bind to their respective monomer targets (e.g., an amino acid type), and detecting the binding agents or using a nanopore sequencer.
[0214] Alternatively, or in addition to, generating the modified monomer (e.g., modified amino acid) may comprise providing a polymerizable molecule comprising a reactive moiety, reacting the reactive moiety with the monomer, and polymerizing the polymerizable molecule. For example, referring to FIG. IB (inset), in process 106’, a polymerizable molecule comprising an incorporable monomer, e.g., a modified nucleotide, that has an amino acid reactive moiety (e.g., isothiocyanate, PITC) is provided and reacted with the modified amino acid. In process 107, a linking polymerizable molecule 111 (e.g., linking nucleic acid molecule) is provided. The linking polymerizable molecule 111 may act as a splint molecule to hybridize the capture moiety 105 or portion thereof. The linking polymerizable molecule 111 may additionally comprise a terminated end such that it is not extendable. In process 112’, an extension reaction is performed, e.g., using polymerase or a TdT enzyme to elongate the capture moiety 105 to comprise a complementary sequence to a portion of the linking polymerizable molecule 111. The extension reaction may incorporate canonical or noncanonical nucleotides (e.g., hexaphosphate nucleotides, deaza-nucleotides), ribonucleotides, or a combination thereof. In some instances, the reaction may be catalyzed by addition of cations (e.g., manganese ions, magnesium ions, calcium ions, etc.). The linking polymerizable molecule 111 may, in some instances, comprise identifying temporal information, e.g., the cycle or round in which it is provided, which may be copied to the extended nucleic acid molecule. In some embodiments, the modified nucleotide may be ligated to thecapture moiety 105. The remaining operations of workflow 100b, e.g., process 113, iteration, etc. may be performed.
[0215] FIG. 1C schematically illustrates another example workflow of generating a modified monomer, e.g., modified amino acid, or a detectable product comprising the modified monomer, in presence or absence of a substrate. In such an example workflow 100c, similar to that of 100b, a polymeric analyte 103, e.g., a peptide, and a capture moiety 105 are provided. The capture moiety 105 may comprise a first nucleic acid molecule (e.g., DNA molecule) and may comprise identifying information of the polymeric analyte 103, e.g., an identifying barcode sequence such as a peptide-identifying barcode. The capture moiety 105 may additionally comprise a releasable or cleavable moiety. The polymeric analyte and the capture moiety 105 may, in some instances, be coupled to a substrate (FIG. 1C inset). The polymeric analyte may be coupled to the capture moiety, either directly or indirectly (e.g., both the polymeric analyte and the capture moiety may be coupled to a substrate). For instance, the capture moiety 105 may comprise a nucleic acid sequence that is complementary to a sequence on a substrate (e.g., bead, flat surface), or the capture moiety may be coupled to the substrate via a linker (not shown). In process 106, a linker 109 and polymerizable molecule, such as a linking nucleic acid molecule 111, are provided. In some instances, the linker 109 is pre-tethered to the polymerizable molecule (linking nucleic acid molecule 111); alternatively, the linker 109 and the polymerizable molecule (linking nucleic acid molecule 11 l)maybeprovidedseparately. Thepolymerizablemoleculemay comprise identifying temporal information, e.g., the cycle or round in which it is provided. In process 106, the linker 109 may couple to a monomer, e.g., an amino acid (e.g., NTAA) of the polymeric analyte 103 (e.g., peptide) to generate a monomer-linker complex. In process 112, the monomer-linker complex may couple to the capture moiety 105, e.g., a portion of the capture moiety 105 may hybridize to the polymerizable molecule (linking nucleic acid molecule 111). In some instances, a nucleic acid extension reaction, e.g., using a polymerase, may be performed, e.g., to copy a sequence (e.g., a barcode sequence) of the capture moiety 105 to the polymerizable molecule (linking nucleic acid molecule 111). In process 113, the monomer may be cleaved from the polymeric analyte 103 to generate the modified monomer that comprises the cleaved monomer, the linker 109, the polymerizable molecule (e.g., linking nucleic acid molecule 111), the capture moiety 105, or a combination thereof (e.g., the cleaved monomer, the linker, and the polymerizable molecule, justthe cleaved monomer, orjustthe cleaved monomer-linker complex). In some instances, the cleavage of the monomer results in formation of a detectable product 131, which comprises a complementary sequence 105’ to a portion of the capture moiety 105, the linking nucleic acid molecule 111, the linker 109, and the cleaved monomer (e.g., cleaved aminoacid). The detectable product 131 may subsequently be dehybridized or removed (e.g., via cleavage) from the capture moiety 105, thereby regenerating the capture moiety 105, which can be used in subsequent iterations.
[0216] Processes 106, 112, and 113 may be iterated and repeated any number of times (“rounds”) using additional linkers 109 and polymerizable molecules and tethering the additional polymerizable molecules to the capture moiety 105 (e.g., via ligation, hybridization, extension), cleaving the additional monomers, removing the detectable product, and repeating. Multiple rounds may continue until all or a subset of the monomers of the polymeric analyte 103 are generated into discrete detectable products. For example, the process may be iterated to generate a plurality of detectable products 131 that each comprise a modified monomer. After any useful number of rounds, the detectable products may be collected, optionally processed, and analyzed, e.g., using a nanopore sequencer. In some instances, the discrete detectable products may be combined or concatemerized, e.g., hybridized or ligated to one another, to generate a stacked plurality of modified amino acids. The stacked plurality of modified amino acids 123 may comprise polymerizable molecules that are coupled or concatenated in a linear fashion (e.g., polymerizable molecules that form a linear polymerizable molecule “backbone”), or in a nonlinear fashion (e.g., random or semi-random arrangement, branched coupling, circularized molecule). The stacked plurality of modified amino acids 123 may be circularized; for example, the stacked plurality of modified amino acids 123 forming a linear polymerizable backbone may be joined at the ends to form a circularized product.
[0217] In some instances, a branched polymerizable molecule may be used or constructed to generate the modified monomer. FIG. ID schematically illustrates another example workflow of generating a modified monomer, e.g., modified amino acid using a branched linking polymerizable molecule. In such an example workflow 1 OOd, similar to that of 100b and 100c, a polymeric analyte 103, e.g., a peptide, and a capture moiety 105 are provided. The capture moiety 105 may comprise a first nucleic acid molecule (e.g., DNA molecule) and may comprise identifying information of the polymeric analyte 103, e.g., an identifying barcode sequence such as a peptide-identifying barcode. The capture moiety 105 may additionally comprise a releasable or cleavable moiety. The polymeric analyte and the capture moiety 105 may, in some instances, be coupled to a substrate (see, e.g., FIG. 1C inset) The polymeric analyte may be coupled to the capture moiety, either directly or indirectly (e.g., both the polymeric analyte and the capture moiety may be coupled to a substrate). For instance, the capture moiety 105 may comprise a nucleic acid sequence that is complementary to a sequence on a substrate (e.g., bead, flat surface). In process 106, a linker 109 and polymerizable molecule, such as a linking nucleic acid molecule111, are provided. In some instances, the polymerizable molecule is a branched polymerizable molecule (e.g., a branched DNA molecule). In one such example, the branched DNA molecule may comprise a first nucleic acid molecule that comprises a first click chemistry moiety (e.g, an ethynyl or octadiynyl nucleobase or nucleotide analogue) that can conjugate to a second nucleic acid molecule comprising a second click chemistry moiety (e.g., azide), thereby generating the branched DNA molecule.
[0218] In some instances, the linker 109 is pre-tethered to the polymerizable molecule; alternatively, the linker 109 and the polymerizable molecule (e.g., linking nucleic acid molecule 111) may be provided separately . The polymerizable molecule may comprise identifyingtemporal information, e.g., the cycle or round in which it is provided. In process 106, the linker 109 may couple to a monomer, e.g., an amino acid (e.g., NTAA) of the polymeric analyte 103 (e.g., peptide) to generate a monomer-linker complex. In process 112, the monomer-linker complex may couple to the capture moiety 105, e.g., a portion of the capture moiety 105 may hybridize to the polymerizable molecule (linking nucleic acid molecule 111, hybridization not shown) or alternatively, the polymerizable molecule may be coupled to the capture moiety 105 using a splint molecule (not shown) and ligated. In another example, the coupling of the capture moiety 105 and the polymerizable molecule (linking nucleic acid molecule 111) may be performed using click chemistry. In some instances, a nucleic acid extension reaction, e.g., using a polymerase, may be performed (not shown). In some instances, a different order of operations may be performed; for instance, the coupling of the polymerizable molecule to the capture moiety 105 may occur first, followed by coupling of the linker 109 to the polymerizable molecule.
[0219] In process 113, the monomer is cleaved from the polymeric analyte 103 to generate the modified monomer that comprises the cleaved monomer, the linker 109, the polymerizable molecule (e.g., linkingnucleic acid molecule 111), which altogether may be coupledto the capture moiety 105. Processes 106, 112, and 113 may be iterated and repeated any number of times (“rounds”) using additional linkers 109 and polymerizable molecules, and tethering the additional polymerizable molecules together (e.g., tethering a second polymerizable molecule provided in a second round to polymerizable molecule of the first round). Multiple rounds may continue until all or a subset of the monomers of the polymeric analyte 103 are tethered together. For example, the process may be iterated to generate a stacked plurality of modified monomers (e.g., stacked plurality of amino acids 123) comprising a set of concatenated modified monomers, e.g., concatenated monomer-linker-polymerizable complexes. For example, the polymerizable molecule from a second round may be coupledto the polymerizablemolecule from the firstround, and the polymerizable molecule from a third round may couple that of the second round, and soon. Accordingly, in some embodiments, the polymerizable molecules from the multiple rounds may thus form a polymerizable molecule backbone that comprises pendant, individual modified monomers. The polymerizable molecules (e.g., linking nucleic acid molecules 111) of the stacked plurality of modified monomers (e.g., stacked plurality of amino acids 123) may be identical molecules (e.g., same nucleic acid sequence), orthey may be different. The stacked plurality of modified amino acids 123 may comprise polymerizable molecules that are coupled or concatenated in a linear fashion (e.g., polymerizable molecules that form a linear polymerizable molecule “backbone”), or in a non-linear fashion (e.g., random arrangement, branched coupling). The stacked plurality of modified amino acids 123 may be circularized; for example, the stacked plurality of modified amino acids 123 forming a linear polymerizable backbone may be joined at the ends to form a circularized product. After any useful number of rounds, the stacked plurality of modified monomers (e.g., stacked plurality of amino acids 123) maybe cleaved at or from the capture moiety 105, e.g., using the cleavable moiety. Alternatively, or in addition to, the stacked set of polymerizable molecules may be subject to amplification to generate amplicons of the polymerizable molecules coupled to or comprised by the stacked plurality of modified amino acids. The stacked plurality of modified amino acids 123 may be subjected to analysis or characterization, e.g., by contacting the stacked plurality of modified amino acids 123 with a library of binding agents, which can bind to their respective monomer targets (e.g., an amino acid type), and detecting the binding agents, or using a nanopore sequencer.
[0220] In some instances, the polymerizable molecule, e.g., linkingnucleic acid molecule 111 comprises temporal information on the cycle in which it is provided; as such, the temporal information may be used for reconstructing the sequence of the peptide or for quality control. For example, referring to FIG. IB, if a particular cycle number is missing, then it can be inferred that an amino acid is missing or was not present in the peptide, that cleavage of the amino acid did not occur, or other error. For reconstruction purposes, e.g., referring to FIG. 1C, the presence of a temporal (e.g., round or cycle number) barcode and a peptide-identifying barcode can be used to attribute a particular identified modified amino acid to the order or position (using the temporal barcode) in which it occurs in a specific peptide (using the peptide-identifying barcode).
[0221] Accordingly, in some aspects of the present disclosure, a method for processing a peptide may comprise (a) providing the peptide and a linker, wherein the linker is coupled to a polymerizable molecule (e.g., nucleic acid molecule); (b) couplingthe linker to the amino acid of the peptide to generate an amino acid-linker complex; (c) couplingthe polymerizable molecule to a capture moiety, e.g., via hybridization; (d) cleaving the amino acid from the peptide to yield a modified amino acid comprising the cleaved amino acid, the linker, and the polymerizable-limolecule; and performing an extension reaction (e.g., nucleic acid extension reaction) of the polymerizable molecule, thereby generating a detectable product comprising the modified amino acid.
[0222] It will be appreciated that the modified amino acid may comprise other additional modifications that are not depicted, such as posttranslational modifications or chemical modifications (e.g., protecting groups), described elsewhere herein. In some instances, the modified amino acid comprises a derivatized amino acid. For instance, the linker may comprise a PITC moiety as the amino acid reactive group, and upon conjugation of PITC to an amino acid (e.g., NTAA) of a peptide under mildly basic conditions, a phenylthiocarbamoyl (PTC) derivative of the amino acid is generated. The PTC-derivatized amino acid may be treated with acid (e.g., TFA or a Lewis acid) to generate a cleaved cyclic 2-anilino-5(4)- thiazolinone (ATZ)-derivatized amino acid, leaving a new N-terminus on the remaining peptide. The ATZ-derivatized amino acid may be converted to a phenylthiohydantoin (PTH) derivative or PTC derivative.
[0223] It will also be appreciated that the polymerizable molecules and capture moieties described herein may comprise other molecule types, e.g., peptides, lipids, carbohydrates, polymers (both naturally occurring and synthetic), or a combination thereof. For instance, referring again to FIGS. 1A-1D, the capture moiety 105 and the polymerizable molecule (e.g, shown as a linking nucleic acid molecule 111) may each comprise a peptide, optionally comprising a peptide barcode sequence. In such examples, process 112 (coupling of the monomer-linker complex to the capture moiety) may be mediated using a peptide enzyme such as sortase A. In one such example, the capture moiety (or the polymerizable molecule) may comprise an oligoglycine or poly -glycine peptide sequence at the C-terminus and the polymerizable molecule (or capture moiety) may comprise a sortase recognition sequence (e.g. , LPXTG, where Xis any amino acid) at the N-terminus and optionally, an oligo-glycine or poly-glycine peptide sequence at the C-terminus (to facilitate further attachment). Sortase A may then be used to catalyze the formation of a peptide bond between the capture moiety and the polymerizable molecule. Accordingly, subsequent to cleavage of the terminal monomer (e.g., terminal amino acid of a peptide analyte), and iteration of the workflow, a stacked plurality of modified monomers 123 that are connected via a peptide backbone may be generated. In some instances, the polymerizable molecule and / or capture moiety may comprise a protecting group that can be deprotected at any useful or convenient step. Beneficially, the use of a peptide backbone may allow for performing harsher reaction conditions (e.g., traditional Edman degradation using strong acids) and for assisting in nanopore readout of the individual monomers spaced along the peptide backbone, e.g., asdescribed by K. Motone et al. Multi-pass, single-molecule nanopore reading of long protein strands. Nature (2024), which is incorporated by reference herein in its entirety.
[0224] A “derivative” of the modified amino acid may generally refer to a molecule that is derived from the modified amino acid . A derivative may b e a product of a reaction (e . g. , chemical, enzymatic) or interaction of the modified amino acid with another molecule. Further processing of the modified amino acid may optionally be performed to arrive at the “derivative thereof.” For example, further chemical or enzymatic treatment, cleavage of the amino acid, an extension or amplification reaction, cleavage or removal from a substrate, physical processes such as mechanical shearing or fragmentation may be performed on the modified amino acid to obtain a derivative of the amino acid or derivative of the modified amino acid. In an example, a modified amino acid may be derivatized by chemical reaction, e.g., using maleimide to react with cysteine residues, NHS esters or isothiocyanates to react with lysine residues, etc. The derivativemay result from performing a nucleic acid reaction (e.g., nucleic acid extension reaction, amplification, ligation, transposition, hybridization, dehybridization, etc.).
[0225] Iteration: In some instances, one or more of the operations described herein may be iterated or repeated. Iteration of the operations may allow for sequential processing, analysis, or identification of the individual monomers of the polymeric analyte, which can allow for reconstruction of the entire polymeric analyte. For example, referring to FIGs. 1A-1D, the operations of the workflow 100a, 100b, 100c, and 1 OOd may be conducted to generate a modified amino acid sequentially for each terminal monomer (e.g., NTAA) of the polymeric analyte (e.g, peptide). In some instances, for analyzing peptides, the individual modified amino acids may then be immobilized to a substrate (e.g., the same or separate substrate as shown in FIG. 1A), linearized, optionally immobilized (e.g., at another end), and detected. Alternatively, orin addition to, the individual modified amino acids may be combined, e.g., via hybridization or ligation of the polymerizable molecules to generate the stacked plurality of modified amino acids. As such, the stacked plurality of modified amino acids may comprise a plurality of polymerizable molecules from multiple rounds or cycles. In some instances, the polymerizable molecule of the second (or third, fourth, fifth, . . . nth) cycle may be configured to only couple to the first (or second, third, fourth, ...n-lth) polymerizable molecule. For example, the first cycle polymerizable molecule may comprise a unique binding sequence that is absent on the capture moiety of the substrate, and to which the second cycle polymerizable molecule can bind. Accordingly, the second cycle polymerizable molecule may only bind to the first cycle polymerizable molecule and not to any of the capture moieties. In the event that an amino acid is missed during a cycle (a “null” event, e.g., not cleaved, not coupled to the linker or polymerizable molecule, etc.), a bridgingpolymerizable molecule may be provided that encodes for a null event but comprises the unique binding sequence, such that subsequent rounds may continue, even if a null event occurs. Alternatively, the polymerizable molecules across cycles or rounds may comprise the same binding sequence (e.g., an adapter sequence).
[0226] In some instances, a plurality of bioorthogonal click chemistry moieties may be used in the polymerizable molecules for different cycles. For instance, during a first cycle, the polymerizable molecule (e.g., alinkingnucleic acid molecule) may comprise tetrazine, which may react with a linker provided in the first cycle that comprises TCO. During the second cycle, the polymerizable molecule and the linker may comprise moieties for an orthogonal click chemistry to that of the first cycle, e.g., sulfur fluoride exchange (SuFEx) click chemistry. Accordingly, temporal information may be provided by the different click chemistry moieties, in addition to or alternatively to using cycle barcodes.
[0227] In instances where one or multiple additional capture moieties are used to couple to the modified monomers, the polymerizable molecules (e.g., linking nucleic acid molecules 111) may additionally encode temporal information, e.g., the cycle or iteration number, such that the order of the individual monomers may be determined. For example, for a given peptide, the terminal amino acid may be coupled to a polymerizable molecule that comprises a barcode sequence that identifies the cycle number (e.g., cycle 1 ) (not shown). The information encoded by the barcode sequence may be coupled to an adjacent (additional) capture moiety (not shown). Following cleavage of the monomer from the polymeric analyte (e.g., as shown in process 113 of FIG. IB and FIG. 1C), the workflow may be repeated for the n-1 terminal amino acid, which may again be coupled to a capture moiety via a linker and barcoded polymerizable molecule and cleaved. The barcoded polymerizable molecule may comprise the cycle number (e.g., cycle 2).
[0228] In some instances, temporal information may be provided separately. For example, prior to, during, or subsequent to coupling of a polymerizable molecule 111 to the capture moiety, a temporal barcode may be provided that can couple to the polymerizable molecule 111 or to the capture moiety 105, or a combination thereof. The temporal barcode may comprise any useful agent, including a nucleic acid molecule, a peptide, a lipid, a carbohydrate, an enzyme (e.g., a chromogenic or fluorogenic enzyme) or a ribozyme or DNAzyme, a fluorophore, a dye, an intercalating agent, a dideoxynucleotide, a fluorescent nucleic acid molecule or nucleotide, a radioisotope, a mass tag, or other detectable label that can indicate the time or cycle (or iteration) number in which it is provided. In some instances, the temporal barcode comprises a cycle- specific nucleic acid barcode molecule, which can couple to the polymerizable molecule 111 or to the capture moiety 105. The temporal barcode may comprise any additional useful functionalsequences, e.g., primer sites, sequencing sites, restriction sites, abasic or cleavable sites, etc. In some instances, the temporal barcode may comprise an amplification site that allows forbridge amplification of the temporal barcode and optionally, the coupled polymerizable molecules, to other capture or polymerizable molecules.
[0229] Analysis of the remaining peptide subsequent to any number of cycles of intramolecular expansion may be performed, which may be useful in providing additional information or identification or “fingerprinting” of the peptide, as described elsewhere herein. FIG. IE schematically illustrates an example workflow for analyzing a remaining peptide after performing intramolecular expansion, e.g., as shown in FIG. 1A-1D. The polymeric analyte 103 (e.g., a peptide) and the capture moiety 105 may optionally be or remain coupled to a substrate, such as a bead. The capture moiety 105 may comprise a cleavable moiety, e.g., a restriction site, a uracil, an abasic site, a photocleavable or photolabile moiety, a disulfide bond, etc.
[0230] In process 126 of FIG. IE, similar to that of 106, a linker and a polymerizable molecule, e.g., a linking nucleic acid molecule are provided. In some instances, the linker pretethered to the polymerizable molecule (e.g., a linking nucleic acid molecule); alternatively, the linker and the polymerizable molecule may be provided separately. In process 126, the linker may couple to a monomer, e.g., an amino acid (e.g., NTAA) of the polymeric analyte 103 (e.g., the remaining peptide subsequent to one or more cycles of intramolecular expansion) to generate a monomer-linker complex. In process 128, the monomer-linker complex may couple to the stacked plurality of modified amino acids 123. Coupling of the monomer-linker complex to the stacked plurality of modified amino acids 123 may be mediated by the polymerizable molecule (e.g, linking nucleic acid molecule). Optionally, the monomer-linker complex and the stacked plurality of modified amino acids 123 may be covalently linked together using chemical (e.g., click chemistry) or enzymatic (e.g., a ligase, polymerase) approaches, or they may be noncovalently linked (e.g., via hybridization). In process 130, cleavage of the capture moiety 105 may be performedby application of a stimulus, e.g., abiological stimulus (e.g., restriction enzyme, UDG), photo-stimulus, chemical stimulus (e.g., reducing agent), thereby providing a liberated stacked plurality of modified amino acids 123 coupled to the remainder of the polymeric analyte 103. In some instances, a portion 105 ’ of the capture moiety may remain tetheredto the polymeric analyte 103. The portion 105’ of the capture moiety may serve as an attachment site for an additional adapter 125 (e.g., a sequencing adapter), e.g., via ligation, hybridization, or an extension reaction. In process 132, the remaining polymeric analyte 103 and the stacked plurality of modified amino acids 123 may be translocated through a nanopore for readout or analysis.
[0231] Beneficially, readout of the remaining polymeric analyte may provide additional multiplexed information on the polymeric analyte. For instance, readout of the stacked plurality of modified amino acids 123 may yield information on the primary structure (amino acid sequence) of a portion of the peptide, and readout of the remaining polymeric analyte (e.g., peptide) may yield additional information on the peptide, e.g., the length or size of the peptide (or the remaining portion of the peptide). The combination of information may be used for peptide identification or fingerprinting in a rapid and accurate manner.
[0232] In some instances, the remainingpolymeric analyte may be further processedor treated prior to or duringtranslocation through the nanopore. The polymeric analyte may be digested with an enzyme (e.g., fragmenting enzyme, aminopeptidase, trypsin, peptidase, or other enzyme), degraded or cleaved (e.g., using Edman degradation).
[0233] Improving Nanopore Sequencing Accuracy of Modified Monomers: The methods, systems, kits, and compositions described herein may advantageously improve nanopore sequencing readout and identification of the modified monomers (e.g., modified amino acids) described herein. In someinstances, the methods, systems, kits, and compositionsdescribed herein may improve the accuracy of identification of an amino acid type of a modified amino acid. In some embodiments, the improvement to identification accuracy may be achieved by modulating the translocation of the modified monomer adjacent to or through the nanopore. In some embodiments, the translocation of the modified amino acid through the nanopore can be modulated by alteringthe conditions in which the translocation occurs. For instance, a commercial nanopore sequencer system may conventionally use a particular set of conditions (which may be referred to herein as “control condition”) for translocation of a molecule through the nanopore. By altering or changing the conditions under which the nanopore sequencer is run, the translocation speed of the analyte of interest (e.g., a modified amino acid) may be modulated. Advantageously, reducingthe translocation speed of the analyte of interest (e.g., modified amino acid) through the nanopore may improve the accuracy of the readout.
[0234] In some instances, the conditions sufficient to reduce a translocation speed of the analyte of interest (e.g., a modified amino acid) include alterations to a sequencing run buffer (a “control” solution). Such alterations may include an increased viscosity (e.g., addition of PEG, glycerol, glucose, dextran, bovine serum albumin, methylcellulose, dextrose, etc.), an increased concentration of ATP inhibitors or addition of ATP competition agents (e.g., ATP analogs such as adenylyl-imidodiphosphate), which may aid in slowing down the ratcheting enzyme (e.g., helicase), addition of cofactors, an increased ionic strength, e.g., addition of cations, metal ions, metal salts, halide salts, including but not limited to lithium chloride (LiCl), potassium chloride(KC1), magnesium chloride (MgCl), sodium chloride (NaCl), caesium chloride (CsCl), tetramethyl ammonium chloride, trimethylphenyl ammonium chloride, phenyltrimethyl ammonium chloride, or l-ethyl-3 -methyl imidazolium chloride, or other salts, an altered pH, an addition of an enzyme, e.g., strand displacing enzyme such as Phi29, replication protein A, DNA replication or repair enzymes, single-stranded binding proteins, DNA-binding proteins (e.g., dsDNA binding proteins, ssDNA binding proteins, Ku70 / 80) or peptides, addition of a peptide or peptoid, addition of helicase inhibitors, addition of polymerase inhibitors, replacement or addition of a helicase, or addition of a denaturant or agent that can modulate DNA structure (e.g., betaine, DMSO). In some instances, the conditions sufficient to reduce a translocation speed of the analyte of interest includes use of a stimulus, as described elsewhere herein, such as a photostimulus, thermal stimulus, chemical stimulus, or biological stimulus.
[0235] Translocation speed of the analyte of interest may also be modulated using molecular sieving mechanisms, e.g., incorporation of gel matrices (e.g., polyacrylamide, agarose, photo- crosslinkable gels, thermosensitive gels such as SOL gels). In one such example, the nanopore or nanogap may be embedded within a gel matrix, which may help to control translocation of the modified amino acids through the nanopore or nanogap.
[0236] In some instances, the modified amino acid or stacked plurality of modified amino acids is contacted with one or more helper molecules, which can aid in reducing the translocation speed, modulate or alter the interaction with the nanopore, or otherwise improve the accuracy of the readout obtained from the nanopore or nanogap sequencing. The one or more helper molecules may bind to the modified amino acid or stacked plurality of modified amino acids, or a portion thereof, and may comprise one or more moieties that can interact with the amino acid portion of the modified amino acid or the amino acid portions of the stacked plurality of modified amino acids. Alternatively, or in addition to, the helper molecule may comprise one or more moieties that are configured to interact with a portion of the nanopore or nanogap sequencer. For example, the helper molecule may be configured to interact with a portion of the sensing region (e.g., a sensing region in the lumen) of a biological nanopore. The interaction of the helper molecule with the modified amino acids or the nanopore or nanogap sequencer may comprise a covalent or non- covalent (e.g., ionic, dipole-dipole interaction, electrostatic interaction, hydrophobic or van der Waals) bond. For example, the helper molecule may comprise a hydrophobic moiety (e.g., a hydrocarbon, a lipophilic moiety, a cholesterol moiety) that can interact with nonpolar or hydrophobic residues, or a polar or charged moiety (e.g., a cation, an anion, a polyatomic ion, ionic compounds that can interact with polar or oppositely-charged residues), an electronwithdrawing group (e.g., fluorine, chlorine, or other halogen), an amine group, an oxygen group,or other moiety . In some instances, the helper molecule comprises an amphiphilic moiety. In some instances, the helper molecule comprises a chelator. The one or more helper molecules may comprise any useful molecule type, e.g., a protein or peptide, a nucleic acid, a lipid, a carbohydrate, a synthetic polymer (e.g., PEG, PVA, polyacrylamide), etc. In some instances, the one or more helper molecules comprises a DNA molecule that can hybridize to the polymerizable molecule (e.g., another DNA molecule) of the modified amino acid. In some instances, the linker comprises the one or more helper molecules.
[0237] Alternatively, or in addition to, the helper molecule may compriseone or more reactive groups that can couple or link to additional moieties that may interact with the modified amino acids. In one example, the helper molecule may comprise a DNA molecule that comprises one or more reactive groups. The DNA molecule may hybridize to the polymerizable molecule (e.g, another DNA molecule) of a modified amino acid or stacked plurality of amino acids. Subsequently, a library of molecules with different properties (e.g., charged, polar, hydrophobic, etc.) may be contacted with the helper molecule; the library of molecules may selectively bind to different amino acid types based on the amino acid property (e.g., a charged molecule may bind to a modified amino acid comprising an oppositely charged residue, a hydrophobic molecule may bind to a hydrophobic residue) and may comprise another reactive group which can couple or link to the reactive group of the helper molecule (e.g., via click chemistry, affinity binding, etc.). In some instances, the helper molecule, andnotthe stacked plurality of aminoacids, may be analyzed using the nanopore sequencer, obviating the need to analyze the modified amino acid or stacked plurality of amino acids.
[0238] Beneficially, the helper molecule may be useful in nanopore sequencing identification of the modified amino acids (e.g., the amino acid type comprised by the modified amino acids). For instance, the helper molecule may aid in further distinction of the signals generated from a nanopore sequencer from the different amino acid types. In one such example, the helper molecules comprising hydrophobic regions may bind to hydrophobic aminoacid residues, thereby increasingthe detectability (e.g., signal amplitude, signal intensity, current deflection or blockade, etc.) of one or subsets of the modified amino acids comprising hydrophobic residues. In some instances, the helper molecule may change a property, e.g., hydrophobicity, hydrophilicity, flexibility or rigidity, size, shape, conformation, charge, etc., of the modified amino acid, which may render the modified amino acid more detectable, e.g., by increasing the distinction of the signal generated from the different amino acid types.
[0239] FIG. IF schematically illustrates binding of one or more helper molecules with a stacked plurality of modified amino acids. The stacked plurality of modified amino acids 123 maybe generated using any useful technique or combination of techniques, e.g., as shown in FIGs. 1A-1D. The stacked plurality of modified amino acids 123 is contacted with one or more helper molecules 131. The one or more helper molecules 131 may be identical or different. For instance, a plurality of helper molecules comprising different functional groups or moieties may be provided. In one such example, a first helper molecule may comprise a hydrophobic moiety disposed at one end of the molecule and a polar or charged moiety disposed at the other end. A second helper molecule may comprise two ends, each having a polar or charged moiety. A third helper molecule may comprise two ends, each having a hydrophobic or nonpolar moiety. The hydrophobic moieties may be configured to form hydrogen bonds or van der Waals interactions with nonpolar side chains (e.g., Gly, Ala, Vai, Leu, He, Met, Phe, Trp, etc.). Similarly, the charged or polar moieties may be configured to form ionic or electrostatic interactions with polar or charged side chains (e.g., Asn, Ser, Gin, Asp, Glu, Lys, Arg, His). In some instances, the helper molecules comprise a polymerizable molecule, e.g., nucleic acid molecule, which can hybridize to a complementary sequence of the polymerizable molecules (e.g., DNA molecules) of the stacked plurality of modified amino acids. The helper molecules bound to the stacked plurality of modified amino acids may then be translocated through a nanopore (e.g., a biological or solid- state nanopore) or nanogap for detection and can aid in decreasing the translocation time of the individual modified amino acids through the nanopore or nanogap. In some instances, the helper molecules may comprise single nucleotides, di-nucleotides, or other small nucleic acid molecule, which may be configured to hybridize with a polymerizable molecule of the modified amino acid or stacked plurality of modifiedamino acids atany useful position or location, e.g., in contactwith the amino acid, adjacent to the amino acid, upstream or downstream along the polymerizable molecule backbone, etc.
[0240] Alternatively, or in addition to, the helper molecules may be configured to interactwith the nanopore or nanogap sequencer. For example, the helper molecule may comprise a moiety that interacts with a moiety or binding pocket within the lumen or sensing region of a biological nanopore. In some instances, the nanopore or nanogap cancompriseamodification,e.g., abinding moiety, chelator molecule (e.g., nickel modification), etc. which can interactwith the moiety of the helper molecule. In some instances, the helper molecule may induce a structural or chemical change in the nanopore, which may improve sequencing or identification accuracy, e.g., by decreasing the translocation speed of the modified amino acid through the nanopore.
[0241] Alternatively, or in addition to, the linkers described herein may comprise different moieties that can facilitate or modulate the interaction of the modified amino acids with the nanopore. For example, a library of linkers with different properties (e.g., charged, polar,hydrophobic, etc.) and comprising amino acid reactive groups may be contacted with one or more peptides (e.g., from a sample or peptide library); the library of linkers may have a variety of moieties (e.g., a charged moiety, a hydrophobic moiety, etc.), which may couple preferentially to particular amino acid types (e.g., hydrophobic moieties may interact preferentially with hydrophobic residues, negatively charged groups may interact with positively charged residues, etc.) and may, in some instances, facilitate the conjugation of the linker to the amino acid. The library of linkers may additionally comprise the same or different reactive groups (e.g., click chemistry moieties, see, e.g., FIGs. 2A-2B) for facilitating attachment to a polymerizable molecule.
[0242] The nanopore, nanogap, or nanochannel may be modified or engineered to comprise a moiety that can interact with the modified amino acid (or portion thereof) or a helper molecule. For instance, a biological nanopore may be engineered using protein engineering techniques to improve pi-pi interactions between the linker comprised by the modified amino acid and the nanopore (e.g., in the lumen or sensingregion of the nanopore). In one such example, the nanopore lumen may comprise one or more phenylalanine residues, which can then interact with a phenyl group of the modified amino acid (e.g., added to a portion of the linker or the polymerizable molecule). In another example, a coordination interaction between the modified amino acid and the nanopore lumen may occur. In one such example, the modified amino acid (e.g., the linker) may comprise a carboxylate or amine group, and the nanopore lumen may comprisea modification to facilitate coordination chemistry, e.g., a lumen-exposed charged residue such as histidine, aspartic acid, or glutamic acid or a polar side chain. The coordination interaction maybe facilitated by a metal ion, e.g., Cu(II), Ca(II), Co(II), Mg (II), Zn(II), Fe(II), or other metal ion. In some instances, the metal ion may coordinate 6 bonds, any number of which may arise from the modified nanopore and the modified amino acid; for example, the ratio of bonds arising from the modified nanopore to the modified amino acid may be 5 :1, 4:2, 3 :3, 2:4, or 1 :5.
[0243] Interaction of the modified amino acid and the nanopore, nanogap, or nanochannel may be facilitated via hydrogen bonding, van der Waals interaction, or entropic forces. In such instances, either or both the nanopore and the modified amino acid or portion thereof (e.g., the linker, the polymerizable molecule) may comprise oxygen or hydrogen groups to facilitate hydrogen bonding, hydrophobic regions (e.g., hydrocarbons) to facilitate hydrophobic interactions.
[0244] In yet other examples, the nanopore, nanogap, or nanochannel may comprise one or more binding agents or moieties, which can facilitate the interaction of the nanopore, nanogap, or nanochannel with the modified amino acid. For instance, the nanopore, nanogap, or nanochannelmay comprise an antibody, nanobody, antibody fragment, aptamer, biotin, streptavidin, integrin or binding moiety (e.g., an RGD tripeptide), DNA, RNA, polymer, nanoparticle, fluorophore. Additional examples of binding agents are described elsewhere herein. In some instances, low- affinity binders may be attached to the nanopore or nanogap to enable weak interactions of the modified amino acids during translocation and / or measurement.
[0245] Similarly, the nanopore (e.g., a solid-state nanopore) may comprise a surface coating that can facilitate interaction with the modified amino acid or portion thereof. For instance, the nanopore may comprise any number of useful moieties such as charged moieties, hydrophobic moieties, polar moieties, or nonpolar moieties.
[0246] In some instances, the polymerizable molecule may be treated with a radical or radicalgenerating agent, e.g., UV rays, hydrogen peroxide, ozone, or other reactive oxygen species. In some instances, the radical agent may be provided prior to or during translocation of the modified amino acid through the nanopore or nanogap, e.g., to induce damage on the polymerizable molecule, which may be useful improving the detection of the nanopore or nanogap sequencer, e.g., increasingthe measured current signal or blockade, producing a more distinct signal, etc. In some instances, the damage or change caused to the polymerizable molecule decreases the translocation speed of the polymerizable molecule.
[0247] Modifications to the polymerizable molecules may also be useful in improving the sequencing accuracy. For example, a polymerizable molecule that comprises DNA may comprise or be treated with, an intercalating dye (e.g., SYBR) or a polyamine (e.g., spermine, spermidine), which may change the charge or another property of the DNA molecule, thereby rendering the polymerizable molecule and / orthe modified monomer (e.g., modified amino acid) comprisingthe polymerizable molecule more detectable. In some instances, the change in charge or other property may modulate the translocation speed of the modified monomer. Modifications to the polymerizable molecules may be performed using a chemical or enzymatic reaction or by changing the buffer or surrounding solution conditions. In some instances, the modification of the polymerizable molecule comprises inducingDNA damage, e.g., irradiation such as treatment with UV, gamma, or X-rays, treatment with a mutagen (e.g., ethidium bromide), an oxidizing agent or reactive oxygen species, heat, a radical agent (e.g., peroxide), etc.
[0248] The polymerizable molecule orthe modified monomer, e.g., modified aminoacid, may comprise or be coupled to other structural or chemical elements (e.g., provided separately or comprised by a helper molecule); such structural or chemical elements may beneficially assist in enhancing sequencing or molecular identification, e.g., by decreasing the translocation speed of the polymerizable molecule as compared to a control molecule (e.g., a polymerizable moleculethatdoes not have the structural or chemical element). In some instances, the structural or chemical element may be useful in analysis or determining the identity of the modified amino acid. For instance, the structural or chemical element may incur a change in the measured signal of a portion of the modified amino acid or stacked plurality of modified amino acids during translocation through the nanopore; such a signal change can be used as a distinctive signal or signal fiducial marker, e.g., for signal processing or alignment purposes, an indicator of a discrete modified amino acid or a stacked plurality of modified amino acids, or other spatiotemporal marker.
[0249] The structural or chemical element may include, for example, a locked nucleic acid (LNA), peptide nucleic acid (PNA), a spacer, a loop sequence, a hairpin sequence, a coiled or super-coiled DNA molecule, a charged linker or adduct, a modified base or backbone, a lesion (e.g., UV-induced DNA lesion), a biotin moiety (e.g., a biotinylated nucleobase), an avidin or streptavidin moiety (e.g., a streptavidin-conjugated nucleobase), an inverted DNA base, a modified DNA moiety such as an Int 5 -Nitroindole, Int 5-TAMRA™ (Azide), Int Super G®, 2'- Deoxy-P-nucleoside-5 '-Triphosphate - (N-2037), Pyrene-dU-CEPhosphoramidite, Perylene-dU- CE Phosphoramidite, Pyrrolo-dC-CEPhosphoramidite, TMP-F-dU-CEPhosphoramidite, dP-CE Phosphoramidite, O4-Triazolyl-dU-CE Phosphoramidite, 5-Hydroxymethyl-dC II-CE Phosphoramidite, tC°-CE Phosphoramidite. The structural or chemical element may comprise a DNA adduct (e.g., an adduct introduced chemically or using a mutant polymerase), such as Pt- (GpG), N2-benzo[a]pyrene diolepoxide-2’-deoxyguanosine, 8-oxo-7,8-dihydro-2’- deoxy guanosine, abasic sites, 5-guanidinohydantoin, 2 ’-deoxy inosine, DNA mismatch sites, 2’- deoxy cytidine derivatives, O6-carboxymethyl-2’ -deoxyguanosine, an epigenetic mark (e.g., 5’- methyl-2’ -deoxy citidine and 5-hydroxymethyl-2’-deoxy cytidine), or other DNA adduct, e.g., as described by Nookaewet al 2020. Chemical Research, in Toxicology. 33 (12), 2944-2952, which is incorporated by reference herein in its entirety. In some embodiments, bulkier DNA adducts may be used or attached to the polymerizable molecule, e.g., a DNA binding protein (e.g., recombinase or other single stranded binding proteins, strand displacing enzymes, streptavidin, traptavidin, biotin, etc.). The structural or chemical element may comprise an electronic tag, e.g, a molecule comprising saturated double bonds, triple bonds, dye molecules, etc. The structural or chemical element may comprise a single- or double-stranded break, which optionally may be introduced at any convenient step or process, e.g., using non-homologous end joining homologous recombination, polymerase repair, etc.
[0250] In some instances, the structural or chemical element may change an interaction of the polymerizable molecule or modified monomer (e.g., modified amino acid) with the nanopore. In some instances, the structural or chemical element may induce a structural change in the nanopore(e.g., biological nanopore), such that the measured signal from the nanopore (e.g., current) is altered.
[0251] The polymerizable molecule may comprise a modification site, e.g., in proximity to the attachmentpointofthe linker orthe amino acid-linker complex, which may enable conjugation of another molecule (e.g., a helper molecule, as described elsewhere herein) that can interact with the amino acid portion of a modified amino acid. The polymerizable molecule may comprise single-stranded DNA, double-stranded DNA, or partially double-stranded DNA with any useful number or variety of structural or chemical elements, which may assist in increasing the nanopore sequencing accuracy, for example by decreasing the translocation speed of the polymerizable molecule. In some instances, the polymerizable molecule comprises a DNA molecule that comprises a particular sequence of nucleotides or nucleotide analogs that decreases the translocation speed of the DNA molecule through the nanopore.
[0252] In some instances, the accuracy of detection and identification of the modified amino acid through the nanopore sequencer includes alteration of an operatingparameter of the nanopore sequencer. For instance, a reduced or decreased temperature, different voltage (e.g., change in voltage, altering voltage, oscillatingvoltage, etc.), or different samplingrate may be used. In some instances, the translocation may be performed at freezing or sub-freezing temperatures. In such instances, additional reagents, such as glycerol or cryoprotective agents, may be utilized to facilitate the translocation of the analytes and preservation of the nanopore (e.g., biological nanopore).
[0253] The linker of the polymerizable molecule may have any useful properties that may aid in distinction, discrimination, or identification of the modified amino acids. For example, the polymerizable molecule may have a designated size, molecular weight, hydrodynamic radius, hydrodynamic volume, hydrophobicity, hydrophilicity, hygroscopicity, charge, polarity, flexibility, rigidity, etc.
[0254] FIG. 3 schematically shows an example of a modified amino acid, which can be generated by coupling a polymerizable molecule having a linker to an amino acid-linker complex. The polymerizable molecule comprises an alkyne-linkedDNA molecule; the alkyne is provided as a synthetic nucleobase, octadiynyl deoxyuridine, thatisincorporatedinto aDNAmolecule. The amino acid-linker complex is generated using a bifunctional linker, l-(2-azidoethyl)4- isothiocyanatobenzene, which comprises (1) a PITC moiety (amino acid reactive group) and (2) an azide moiety that is capable of reacting with the alkyne of the polymerizable molecule. Three different example amino acid-linker complexes are shown comprising Asp, Trp, and Tyrthat have each been reacted with the bifunctional linker. The amino acid-linker complexes may react withthe alkyne-linked DNA molecule, thereby generating a modified amino acid. Alternatively, the alkyne-linkedDNA molecule may first be reacted with the bifunctional linker, which can then subsequently react with an amino acid (e.g., an NTAA of a peptide).
[0255] In other examples, the modified amino acid may comprise a plurality of linkers that are linked together. For instance, the polymerizable molecule may comprise a backbone linker (e.g., a phosphodiesterbackbone)comprisinga functional moiety that can couple to another linker which may be capable of coupling to an amino acid-linker complex, to generate a modified amino acid comprising multiple linkers. In an example, the polymerizable molecule comprises a backbone linker (e.g., on a phosphate or sugar of a nucleic acid molecule) that comprises a free amine group. A second linker comprising an amine-reactive group, such as NHS ester and a click chemistry moiety (e.g., BCN, DBCO, azide, alkyne, etc.) may couple to the free amine group of the polymerizable molecule. Finally, an amino acid-linker complex, e.g., generated by reacting an N-terminal amino acid of a peptide with a third linker comprising an amino acid reactive group (e.g., PITC, a xanthate, a guanidinylating agent, a dithioester, etc.) may be provided. The amino acid-linker complex may comprise an additional click chemistry moiety that can react to the click chemistry moiety of the second linker. As such, reaction of the amino acid-linker complex (the third linker) with the second linker may yield a modified amino acid the comprises the backbone linker, the second linker (coupled to the backbone linker via NHS ester and amine reaction), and the amino acid-linker complex (comprising the third linker coupled to the second linker via click chemistry).
[0256] In some instances, the translocation of the modified amino acid may be modulated or controlled using a mutant helicase or other ratcheting enzyme, e.g., polymerase, exonuclease, single stranded and double stranded binding protein, or topoisomerase, such as gyrases. Alternatively, or in addition, translocation of the modified amino acid may be modulated by the applied voltage. For instance, voltage-gated pausing may be implemented to pause a modified amino acid as it translocates through the nanopore.
[0257] In some instances, improvements to nanopore sequencing accuracy may be obtained by modulating the translocation speed of an analyte (e.g., a modified amino acid) through the nanopore. The translocation speed of the modified amino acid may be higher or lower than the translocation speed of the unmodified amino acid or a modified amino acid that does not comprise a polymerizable molecule. In some instances, the translocation speed of the modified amino acid is lower than that of the unmodified amino acid or the modified amino acid that does not comprise the polymerizable molecule. For example, the translocation speed of the modified amino acid may be decreased by at least about 0.1%, at least about 1%, at least about 2%, at least about 3%, atleast about 4%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, at least about 10%, at least about 20%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, atleast about 90%, at least about 100%, atleast about200%, at least about 500%, at least about 1000%, at least about 5000% or greater. The translocation speed of the modified amino acid may be decreased by a range of percentages, e.g., from about 5%- 20%, from about 500%-3000%, etc. In some instances, the translocation speed of the modified amino acid is non-uniform, e.g., may fall within a range of percentages by which the translocation speed changes, and may depend on a property (e.g., size, charge, polarity, etc.) of the polymerizable molecule, the amino acid type encompassed by the modified amino acid, the linker, or any other component of the molecule.
[0258] The change in translocation speed of the modified amino acid, as compared to an unmodified amino acid (e.g., an amino acid without a coupled polymerizable molecule) may render the modified amino acid more detectable. For example, the polymerizable molecule coupled to an amino acid may decrease or alter the translocation speed as it translocates through the nanopore or nanogap, such that a more accurate current reading may be obtained. Alternatively, or in addition to, the polymerizable molecule may alter the charge, size, masscharge ratio, aspect ratio of the amino acid to which it is coupled, which can alter the current signature to improve detectability of the amino acid, the polymerizable molecule, or both.
[0259] Alternatively, or in addition to, the media (e.g., liquid, buffer, solution) in which the nanopore is in contact may comprise one or more agents that can alter or modulate the translocation speed of the modified amino acid. For example, the media may comprise ions, salts, or other molecules which may selectively or preferentially alter the interaction between the modified amino acid and the nanopore. For instance, a change in the buffer composition may result in increased retention or dwell times of the modified amino acid within the sensing region of the nanopore, thereby producing a more detectable signal that can be used to identify or detect the modified amino acid.
[0260] In some instances, a measured signal (e.g., current blockade) of the modified amino acid as it translocates through the nanopore or nanogap is substantially different than that of an unmodified amino acid or a modified amino acid that does not comprise a polymerizable molecule. In some instances, a measured signal of the modified amino acid as it translocates through the nanopore or nanogap is substantially different than that of the polymerizable molecule alone. The difference of the measured signal may be measured by any useful metric, e.g., fold-change, percentage change, absolute current measurement, signal-to-noise ratio, signal amplitude difference, signal signature, frequency, etc.
[0261] Identification Accuracy: The methods, systems, compositions, and kits provided herein may allow for high-accuracy sequencing of polymeric analytes, such as peptides. For instance, for a given amino acid type, the methods provided herein may provide an average read accuracy of greater than 50%, greater than 60%, greater than 70%, greater than 80%, greater than 90% or higher. For a plurality of amino acid types, the methods provided herein may provide an average read accuracy of greater than 50%, greater than 60%, greater than 70%, greater than 80%, greater than 90% or higher for each of the amino acid types of the plurality of amino acid types. The plurality of amino acid types may include 2 amino acid types, 3 amino acid types, 4 amino acid types, 5 amino acid types, 6 amino acid types, 7 amino acid types, 8 amino acid types, 9 amino acid types, 10 amino acid types, 11 amino acid types, 12 amino acid types, 13 amino acid types, 14 amino acid types, 15 amino acid types, 16 amino acid types, 17 amino acid types, 18 amino acid types, 19 amino acid types, or 20 amino acid types of the 20 proteinogenic amino acids. In some instances, the plurality of amino acid types may include greater than 20 amino acid types, e.g., including post-translationally modified amino acids or noncanonical amino acids. The average read accuracy may fall within a range of accuracies based on the number of amino acid types being identified. For instance, the average read accuracy for identifying a smaller set of amino acid types (e.g., between about 3 and about 5 amino acid types, between about 5 and about 10 amino acid types) may be greater than the average read accuracy for identifying a larger set of amino acid types (e.g., 10 or more amino acid types, all 20 proteinogenic amino acid types, more than 20 proteinogenic amino acid types and post-translational modifications). Accordingly, it will be appreciated that the average read accuracy may fall within a range, depending on the number of amino acid types that are identified.
[0262] The methods, systems, compositions, and kits provided herein may allow for high individual identification accuracies of amino acids (or an amino acid type comprised by a modified amino acid). The individual identification accuracy of a modified amino acid type maybe greater than 50%, greater than 60%, greater than 70%, greater than 80%, greater than 90% or higher. For a plurality of amino acid types, the methods provided herein may provide an individual identification accuracies of greater than 50%, greater than 60%, greater than 70%, greater than 80%, greater than 90% or higher for each of the amino acid types of the plurality of amino acid types. The plurality of amino acid types may include 2 amino acid types, 3 amino acid types, 4 amino acid types, 5 amino acid types, 6 amino acid types, 7 amino acid types, 8 amino acid types, 9 amino acid types, 10 amino acid types, 11 amino acid types, 12 amino acid types, 13 amino acid types, 14 amino acid types, 15 amino acid types, 16 amino acid types, 17 amino acid types, 18 amino acid types, 19 amino acid types, or 20 amino acid types. In some instances, the plurality ofamino acid types may include greater than 20 amino acid types, e.g., post-translationally modified amino acids or noncanonical amino acids. The individual identification accuracy may fall within a range of accuracies based on the number of amino acid types being identified. For instance, the individual identification accuracy for identifying a smaller set of amino acid types (e.g., between about 3 and about 5 amino acid types, between about 5 and about 10 amino acid types) may be greater than the individual identification accuracy for identifying a larger set of amino acid types (e.g., 10 or more amino acid types, all 20 proteinogenic amino acid types, more than 20 proteinogenic amino acid types and post-translational modifications). Accordingly, it will be appreciated that the individual identification accuracy may fall within a range, depending on the number of amino acid types that are identified.
[0263] The methods, systems, compositions, and kits provided herein may also allow for high average identification accuracies of amino acids (or an amino acid type comprised by a modified amino acid). The average identification accuracy of a given modified amino acid type may be greater than 50%, greater than 60%, greater than 70%, greater than 80%, greater than 90% or higher. For a plurality of amino acid types, the methods provided herein may provide an average identification accuracy of greater than 50%, greater than 60%, greater than 70%, greaterthan 80%, greater than 90% or higher for each of the amino acid types of the plurality of amino acid types. The plurality of amino acid types may include 2 amino acid types, 3 amino acid types, 4 amino acid types, 5 amino acid types, 6 amino acid types, 7 amino acid types, 8 amino acid types, 9 amino acid types, 10 amino acid types, 11 amino acid types, 12 amino acid types, 13 amino acid types, 14 amino acid types, 15 amino acid types, 16 amino acid types, 17 amino acid types, 18 amino acid types, 19 amino acid types, or 20 amino acid types. In some instances, the plurality of amino acid types may include greater than 20 amino acid types, e.g., post-translationally modified amino acids or noncanonical amino acids. The average identification accuracy may fall within a range of accuracies based on the number of amino acid types being identified. For instance, the average identification accuracy for identifying a smaller set of amino acid types (e.g., between about 3 and about 5 amino acid types, between about 5 and about 10 amino acid types) may be greater than the average identification accuracy for identifying a larger set of amino acid types (e.g., 10 or more amino acid types, all 20 proteinogenic amino acid types, more than 20 proteinogenic amino acid types and post-translational modifications). Accordingly, it will be appreciated that the average identification accuracy may fall within a range, depending on the number of amino acid types that are identified.
[0264] The individual or average identification accuracy may be determined probabilistically. For example, an individual identification accuracy or an average identification accuracy may becharacterized by a probabilistic determination of a modified amino acid belonging to an amino acid type or subset of amino acid types. The probabilistic determination may comprise determining a probabilistic distribution of a range of classes that correspond to amino acid types (e.g., leucine, isoleucine, valine, arginine, etc.) or determining the probability that a modified amino acid is derived from an amino acid type or a subset of amino acid types.
[0265] Individual read accuracy may also be improved by use of calibration molecules. For instance, reference calibration molecules, e.g., modified amino acids of known identity, may be used to calibrate a nanopore sequencer or account for batch-to-batch or run-to-run variability. Alternatively, or in addition to, an adapter sequence comprising one or more known modified amino acids (e.g., a cleaved amino acid that is linked to a known DNA sequence) may be attached to each or a subset of polymeric analytes. In such examples, the calibrating molecule may serve as a reference standard for the individual monomers of the polymeric analytes. The calibration molecules may comprise one or more small molecules, which may vary in one or more properties such as hydrophobicity, size, charge, flexibility, polarity, etc. A calibration molecule may comprise a polymerizable molecule and optionally, a linker, without an amino acid attached thereto. In some instances, the calibration molecules may comprise similar, but not identical, modified amino acids. For instance, a calibration molecule may comprise an amino acid, a polymerizable molecule (e.g., nucleic acid molecule) that is used in the intramolecular expansion process of a peptide analyte, but which may comprise a different linker type (e.g., isothiocyanate instead of PITC, isothiocyanate instead of a guanidinylating agent), etc. In some instances, a library of calibration molecules may be used to enable calibration across different runs, devices, flow cells, etc.
[0266] In some instances, intramolecular calibration may be performed. For instance, a modified amino acid or stacked plurality of modified amino acids may comprise or be appended to a polymerizable molecule that is identical to that of the modified amino acid (or the polymerizable molecules comprised by the stacked plurality of modified amino acids), but without the amino acid or linker. Byway of example, and referring to FIG. 3, the polymerizable molecule (“DNA backbone”) may comprise a nucleic acid sequence which comprises a substituted, alkyne- containingbase (e.g., octadiynyl dU) to which the amino acid-linker complex may bind. The DNA backbone may also comprise or be coupled to a calibration molecule, which comprises the same sequence as the DNA backbone but with a conventional nucleotide (e.g., A, C, T, G, U) instead of the substituted base. Accordingly, the calibration molecule (sequence) may serve as a baseline sequence for alignment or comparison of the read arising from the modified amino acid portion.
[0267] Read accuracy can also be improved computationally. For instance, accuracy can be improved by imposing confidence thresholds of individual reads, such that higher-confidence reads are output, and lower-confidence reads are removed. Confidence can be, for example, a probability that is assigned by a computational algorithm for the predicted amino acid type.
[0268] In some instances, the identification accuracy of a modified amino acid may be improved by increasing the read count of a particular modified amino acid. For instance, the modified amino acid may be ratcheted back and forth through a particular nanopore, thereby obtaining a plurality of reads of the modified amino acid. Additional approaches for ratcheting a modified amino acid back and forth (“flossing”) through a particular nanopore are described elsewhere herein.
[0269] In some instances, the read count may be increased by circularizing the polymerizable molecule or plurality of polymerizable molecules of a modified amino acid or stacked plurality of amino acids. In one such example, a method of generating increased reads of a modified amino acid may comprise: providing the modified amino acid comprising a polymerizable molecule; translocating the modified amino through or adjacent to a nanopore or a nanogap; circularizing the polymerizable molecule, thereby generating a circularized modified amino acid; and translocating the circularized, modified amino acid through or adjacent to the nanopore or the nanogap.
[0270] FIG. 4A schematically shows an example workflow for generating increased reads (repeat reads) for a single modified amino acid or stacked plurality of modified amino acids. In FIG. 4A Panel A, a stacked plurality of modified amino acids 401 is provided, and can be generated using the methods described herein, see, e.g., FIGs. 1A-1D. The stacked plurality of modified amino acids may comprise a plurality of modified amino acids that are linked together by a polymerizable molecule backbone (e.g., a nucleic acid backbone) that comprises linked polymerizable molecules 111. Alternatively, or in addition to, single modified amino acids comprising a polymerizable molecule 111 may be provided. The modified amino acid or stacked plurality of modified amino acids may be translocated through a nanopore 405, optionally using a ratcheting enzyme 403 such as a polymerase, helicase, or other processive enzyme. As the modified amino acid or stacked plurality of modified amino acids translocates through the nanopore, one or more measurements is made (e.g., current blockade, current signal, impedance, current amplitude, etc.). The one or more measurements may be deconvolved computationally to output the identity of the modified amino acid. In FIG. 4A Panel B, multiple reads are generated from a single modified amino acid or stacked plurality of modified amino acids by circularizing the polymerizable molecule or plurality of stacked polymerizable molecules. In one example, thepolymerizable molecule or plurality of stacked polymerizable molecules may translocate counter- directionally (e.g., in the trans to cis direction) through an adjacent nanopore. One portion of the polymerizable molecule or plurality of stacked polymerizable molecules may then self-ligate to another portion of the polymerizable molecule or plurality of stacked polymerizable molecules, thereby generating a circularized molecule (a circularized modified amino acid or circularized stacked plurality of modified amino acids). Alternatively, or in addition to, additional polymerizable molecules (e.g., splint or bridge oligos) may be provided to enable circularization. The circularized molecule may then iteratively translocate through both pores in a cyclic manner, and the signal may be measured from one or both nanopores. The measured signal may thus generate multiple reads of the same circularized molecule, which reads can be output and used to identify the modified amino acid or plurality of modified amino acids.
[0271] Repeat reads may notbe necessary to increaseread accuracy of a given modified amino acid. In some instances, the stacked plurality of modified amino acids may be positioned along a polymerizable molecule backbone such that individual reads are obtained for a given modified amino acid without overlapping or interfering signal from another modified amino acid of the stacked plurality of modified amino acids. For instance, FIG. 4B schematically shows nanopore sequencing of a stacked plurality of modified amino acids that are spaced along a polymerizable molecule backbone, e.g., as generated using workflow lOOd of FIG. ID. The stacked plurality of modified amino acids may comprise individual modified amino acids 402 that are spatially resolved from one another and spaced along the polymerizable molecule backbone such that only approximately one modified amino acid translocates through a given nanopore at a given instant. A molecular motor protein such as a ratcheting enzyme, e.g., a polymerase or helicase, or engineered variantthereof, may be used to translocate the polymerizable molecule backbone (e.g, a single or double-stranded DNA molecule) adjacent to the nanopore. The modified amino acids (comprising a cleaved amino acid, linker, and at least a portion of a polymerizable molecule) may translocate through the nanopore (e.g., using electrophoretic force) as the polymerizable molecule backbone is ratcheted through the ratcheting enzyme (shown in FIG. 4B as occurring in an orthogonal direction to the ratcheting). The ionic current signal can be measured through the nanopore as the translocation event occurs. Continuous ratcheting of the polymerizable molecule backbone can result in movement of the stacked plurality of modified amino acids adjacent to the nanopore, and discrete measurements of the individual modified amino acids can be made as a function of time or ratcheting of the polymerizable molecule backbone. In some instances, the ratcheting enzyme may have cleaving activity (e.g., an engineered Bst polymerase that has been modified to cleave hexaphosphates or other modified nucleotides or nucleosides), such that aportion of the individual modified amino acids 402 may be released from the polymerizable molecule backbone and may translocate (e.g., via diffusion or active transport such as electrophoresis) through the nanopore. In some embodiments, the ratcheting enzyme is engineered at the ATP binding site.
[0272] Alternatively, or in addition to, an increase in read countmay be obtainedby increasing the numb er of engagement sites of the modified amino acid or stacked plurality of modified amino acids with one or more nanopores. For instance, introducing more structures that can engage with a motor or ratcheting enzyme (e.g., helicase) or to facilitate entry to the nanopore may be performed. FIG. 4C schematically shows a workflowto introduce additional 5’ double-stranded or partially double-stranded ends that can engage with or enter a nanopore using hybridization chain reaction. In such an example, a stacked plurality of modified amino acids 423 may be coupled (e.g., using ligation) to different hairpin molecules at each end of the stacked plurality of modified amino acids 423. A plurality of the stacked plurality of modified amino acids 423 may then be concatenated or joined together by introduction of helper molecules, such as additional hairpin molecules that can promote the hybridization chain reaction. Accordingly, a plurality of the stacked plurality of modified amino acids may be formed, with each stacked plurality of modified amino acids comprising a 5 ’ end double-stranded or partially double-stranded end that can facilitate entry into the nanopore. Hybridization chain reaction may additionally be used to generate or introduce additional structural elements, e.g., loop orhairpin regions, A-tailed regions, etc. to the plurality of stacked plurality of modified amino acids (not shown).
[0273] FIG. 4D schematically shows nanopore sequencing of a circularized stacked plurality of modified monomers, e.g., modified amino acids. The left-hand schematic illustrates a circularized stacked plurality of modified monomers, e.g., individual modified amino acids 402, which may be connected to one another along a polymerizable molecule backbone. The circularized stacked plurality of modified monomers may be generated using an intramolecular expansion process to generate a linear stacked plurality of modified monomers and then circularizing the linear stacked plurality of modified monomers (e.g., via ligation of one end to another). Alternatively, during intramolecular expansion, the polymerizable molecule may be provided as a circularized molecule comprising a plurality of individual linkers (e.g., comprising amino acid reactive groups) to which the monomers may be coupled individually and then cleaved. In some instances, a portion of the polymeric analyte (e.g., peptide) may notbe expanded and may remain coupled to the polymerizable molecule backbone. The circularized stacked plurality of modified monomers may be contacted with a nanopore 405 and a locally coupled molecular motor protein, such as a ratcheting enzyme 403. The ratcheting enzyme 403 may facilitate translocationof the circularized stacked plurality of modified monomers adjacent to the nanopore 405, such that the individual modified amino acids 402 may translocate through the nanopore 405 for detection. In some instances, the individual modified amino acids 402 may comprise a charged moiety (e.g., along the polymerizable molecule or via a linker), which can aid in translocation of the individual modified amino acids 402 through the nanopore 405. In some instances, if a remaining portion of the polymeric analyte is present, that may also be translocated through the nanopore 405 for detection, or alternatively, detected using other approaches such as imaging (e.g., super-resolution imaging). Detection of the remainingportion of the polymeric analyte, e.g, using the nanopore may yield useful information such as the size or length, identity, charge, etc. of the remaining portion of the polymeric analyte, which, in combination with identification of the individual modified monomers (modified amino acids), may identify the originating analyte (e.g., identify the starting protein or peptide). In some embodiments, imaging may be used to detect the remaining polymeric analyte. For instance, a dye or fluoroph ore may be coupled to the N-terminus of a remaining peptide after intramolecular expansion and to the capture moiety coupled to the remaining peptide, and imaging (e.g., confocal microscopy or super resolution imaging) may be used, e.g., to determine a distance between the capture moiety and the N- terminus, thereby yieldinginformation on the number of amino acids in the remaining polymeric analyte. In another example, the N-terminus of the remaining peptide may be labeled with a fluorophore and the nanopore may also comprise a fluorophore, and FRET can be used to determine the number of amino acids in the remaining peptide.
[0274] Multi-state measurements'. The methods provided herein may also employ measuring multiple states of polymeric analytes. A measured state may correspond to any useful measured parameter, e.g., a measured ionic current or current blockade, that can be indicative of a characteristic of a portion of the polymeric analyte that is translocating through or adjacent to the nanopore or nanogap. In some aspects, a method for sequencing a peptide may comprise (a) providing a modified amino acid generated from the peptide, wherein the modified amino acid comprises a polymerizable molecule; (b) translocating at least a portion of the modified amino acid through or adjacent to a nanopore or nanogap; (c) measuring one or more states of the modified amino acid or portion thereof. In some embodiments, the method may comprise: (a) providing a modified amino acid generated from the peptide, wherein the modified amino acid comprises a polymerizable molecule comprising M monomers, wherein M is a positive integer;(b) translocating the modified amino acid through or adjacent to a nanopore, wherein the translocating comprises ratcheting of a first monomer of the M monomers through the nanopore;(c) during (b), measuring a first state of a first set of N monomers of the M monomers, wherein N< M; wherein the first state is associated with the ratcheting of the first monomer; (d) ratcheting a second monomer of the M monomers through the nanopore; (e) measuring a second state of a second set of N monomers of the M monomers; wherein the second state is associated with the ratcheting of the second monomer; (f) repeating (d)-(e) N-2 times, thereby obtaining N measured states; and (g) using the N measured states to determine an identity of the modified amino acid. In some instances, (a)-(g) are repeated in order to sequence the entire peptide or a subset of amino acids of the peptide.
[0275] N may be any useful integer or non-integer number. In some instances, N is a characteristic number that represents the number of monomers that is present in the nanopore or in a sensing or interacting portion of the nanopore in a given instant. For example, N may represent the number or average number of monomers of a polymerizable molecule that traverse through the nanopore in a given instant. As such, N measured states may correspond to the N contiguous monomers in the nanopore in a given instant (e.g., the ratio of the measured states to the number of monomers measured is 1 : 1). In another example, the N measured states may correspond to N measurements taken from N discrete, non-contiguous monomers, e.g., corresponding from a slip in the ratcheting of the monomers (e.g., the ratio of the measured states to the number of discrete monomers measured is 1 : 1 for non-contiguous monomers). Alternatively, the N measured states may correspond to fewer than N monomers, e.g., corresponding from a repeat read of one or more monomers (e.g., the ratio of the measured states to the number of discrete monomers measured is <1 ). In some instances, the N measured states may correspond to a non-integer value of monomers. For instance, if a non-integer number of monomers (e.g., 1.2, 1.5, etc.) is ratcheted through the nanopore, the N measured states may correspond to a measurement obtained from a non-integer value of monomers.
[0276] The measured state may be derived from a measurement of the current signal, current b...
Claims
CLAIMSWHAT IS CLAIMED IS:
1. A method for sequencing, at sub-attomole resolution, a peptide comprising a plurality of amino acids, comprising:(a) providing a plurality of modified amino acids generated from at least a subset of said plurality of amino acids, wherein said plurality of modified amino acids comprises a plurality of polymerizable molecules; and(b) sequencing said plurality of modified amino acids, thereby determining an amino acid identity of each modified amino acid of said plurality of modified amino acids; wherein said sequencing has an individual identification accuracy that is greater than 80% for at least 3 different modified amino acid types.
2. The method of claim 1, wherein said plurality of polymerizable molecules is covalently linked.
3. The method of claim 1 or 2, wherein said sequencing is performed using a nanopore or a nanogap.
4. The method of claim 3, wherein said sequencing is performed under conditions sufficient to reduce a translocation speed of said plurality of modified amino acids through said nanopore as compared to a control condition.
5. The method of claim 4, wherein said conditions sufficient to reduce said translocation speed comprise an increased viscosity, an addition of a radical agent, an addition of a DNA-binding peptide, an addition of ATP inhibitors, an addition of metal ions, an addition of an intercalating dye, a decreased temperature, an altered pH, an altered voltage, an addition of a strand displacing enzyme, replacement or addition of a helicase, addition of replication protein A, addition of a denaturant, or a combination thereof as compared to said control condition.
6. The method of any one of claims 1-5, further comprising, prior to (a), generating said plurality of modified amino acids.
7. The method of claim 6, wherein said generating comprises (I) providing a linker and a polymerizable molecule, (II) coupling said linker and said polymerizable molecule to an amino acid of a peptide, thereby generating an amino acid-linker complex, and (III) cleaving said amino acid, thereby generating said modified amino acid, wherein said modified amino acid comprises a cleaved amino acid, said linker, and said polymerizable molecule.
8. The method of claim 7, further comprising, contacting said modified amino acid with a helper molecule, wherein said helper molecule comprises a charged moiety, a chelator, or a hydrophobic or hydrophilic moiety.
9. The method claim 7 or 8, further comprising (IV) coupling said amino acid-linker complex to a capture moiety.
10. The method of claim 9, wherein said capture moiety is coupled to a substrate.
11. The method of claim 10, further comprising, (V) removing said amino acid-linker complex from said substrate.
12. The method of claim 11, wherein said removing comprises an enzymatic digestion, heat denaturation, or toehold-mediated strand displacement.
13. The method of claim 10, wherein said capture moiety is coupled to said substrate via an anchor molecule.
14. The method of claim 13, wherein said anchor molecule or said capture moiety comprises a PEG linker.
15. The method of claim 10, wherein said substrate comprises a plurality of anchor molecules configured to couple to said capture moiety, wherein an average distance between said plurality of anchor molecules is greater than 100 nanometers.
16. The method of any one of claims 13-15, wherein said anchor molecule comprises a peptide nucleic acid (PNA).
17. The method of any one of claims 9-16, wherein said capture moiety is coupled to said peptide.
18. The method of claim 9, wherein said coupling of (IV) comprises ligation.
19. The method of claim 18, wherein said capture moiety or said polymerizable molecule comprises a hairpin nucleic acid molecule.
20. The method of claim 18, wherein said ligation is performed using a splint oligonucleotide.
21. The method of claim 20, wherein said splint oligonucleotide comprises a hairpin nucleic acid molecule.
22. A method for sequencing a peptide comprising a plurality of amino acids, comprising:(a) providing said peptide; and(b) sequencing said peptide, thereby identifying at least 2 contiguous amino acids of said peptide;wherein said sequencing has an average read accuracy that is greater than 80% for at least 3 different amino acid types.
23. A method for sequencing a peptide, comprising:(a) providing a modified amino acid generated from said peptide, wherein said modified amino acid comprises a polymerizable molecule;(b) translocating said modified amino acid through or adjacent to a nanopore; wherein (b) is performed under conditions sufficient to reduce a translocation speed of said modified amino acid through said nanopore as compared to a control condition;(c) measuring a signal generated from said modified amino acid during (b); and(d) using said signal generated from said modified amino acid to determine an identity of the modified amino acid.
24. The method of claim 23, wherein said polymerizable molecule comprises a nucleic acid molecule.
25. The method of claim 24, wherein said nucleic acid molecule is a branched nucleic acid molecule.
26. The method of claim 24, wherein said nucleic acid molecule comprises a modified base or modified nucleic acid backbone.
27. The method of claim 26, wherein said modified base or modified nucleic acid backbone is selected from the group consisting of a locked nucleic acid (LNA), a phosphoramidite, a click chemistry-conjugated base or click chemistry-conjugated sugar backbone, a spacer moiety, and a combination thereof.
28. The method of any one of claims 23-27, wherein (b) occurs at ambient temperature.
29. The method of any one of claims 23-27, wherein said conditions sufficient to reduce said translocation speed comprise an increased viscosity, an addition of a radical agent, an addition of a DNA-binding peptide, an addition of ATP inhibitors, an addition of metal ions, an addition of an intercalating dye, a decreased temperature, an altered pH, an altered voltage, an addition of a strand displacing enzyme, replacement or addition of a helicase, addition of replication protein A, or addition of a denaturant, as compared to said control condition.
30. The method of any one of claims 23-29, further comprising, prior to (a), generating said modified amino acid.31 . The method of claim 30, wherein said generating comprises (I) providing a linker and a polymerizable molecule, (II) coupling said linker and said polymerizable molecule to an amino acid of said peptide, thereby generating an amino acid-linker complex, and (III)cleaving said amino acid, thereby generating said modified amino acid, wherein said modified amino acid comprises a cleaved amino acid, said linker, and said polymerizable molecule.
32. The method of claim 31, further comprising, contacting said modified amino acid with a helper molecule, wherein said helper molecule comprises a charged moiety, a chelator, or a hydrophobic or hydrophilic moiety.
33. A method of processing a peptide, comprising:(a) providing said peptide and a linker, wherein said linker is capable of coupling to an amino acid of said peptide and wherein said linker is coupled to a nucleic acid molecule;(b) coupling said linker to said amino acid of said peptide to generate an amino acidlinker complex;(c) coupling said nucleic acid molecule to a capture moiety;(d) cleaving said amino acid from said peptide to yield a modified amino acid comprising a cleaved amino acid, said linker, and said nucleic acid molecule; and(e) performing a nucleic acid extension reaction of said nucleic acid molecule, thereby generating a detectable product comprising said modified amino acid.
34. The method of claim 33, further comprising, contacting said modified amino acid or said detectable product with a helper molecule, wherein said helper molecule comprises a charged moiety, a chelator, or a hydrophobic or hydrophilic moiety.
35. The method of claim 33, wherein said linker comprises a charged moiety.
36. The method of any one of claims 33-35, further comprising sequencing said detectable product.
37. The method of claim 36, wherein said sequencing is performed using a nanopore or a nanogap.
38. The method of claim 37, wherein said sequencing is performed under conditions sufficient to reduce a translocation speed of said detectable product through said nanopore as compared to a control condition.
39. The method of claim 38, wherein said conditions sufficient to reduce said translocation speed comprise an increased viscosity, an addition of a radical agent, an addition of a DNA-binding peptide, an addition of ATP inhibitors, an addition of metal ions, an addition of an intercalating dye, a decreased temperature, an altered pH, an altered voltage, an addition of a strand displacing enzyme, replacement or addition of a helicase, addition of replication protein A, or addition of a denaturant, as compared to said control condition.
40. The method of any one of claims 33-39, wherein said capture moiety is coupled to a substrate.41 . The method of claim 40, wherein said capture moiety is coupled to said substrate via a PEG linker.
42. The method of claim 40, wherein said substrate comprises a plurality of capture moieties, wherein an average distance between said plurality of capture moieties is greater than 100 nanometers.
43. The method of any one of claims 33-42, wherein said capture moiety is coupled to said peptide.
44. The method of claim 43, wherein said capture moiety is coupled to a C-terminus of said peptide.
45. The method of any one of claims 33-44, further comprising repeating (a)-(e) on said peptide.
46. The method of claim 45, wherein said repeating yields a plurality of detectable products that are not coupled to one another.
47. The method of any one of claims 33-46, further comprising releasing said modified amino acid from said capture moiety.
48. A method for sequencing, at sub -attomole resolution, a peptide comprising a plurality of amino acids, comprising:(a) providing a plurality of modified amino acids generated from at least a subset of said plurality of amino acids, wherein said plurality of modified amino acids comprises a plurality of polymerizable molecules; and(b) sequencing said plurality of modified amino acids, thereby determining an amino acid identity of each modified amino acid of said plurality of modified amino acids; wherein said sequencing has an average identification accuracy that is greater than 80% for at least 3 different amino acid types of the 20 canonical proteinogenic amino acid types.
49. A method for peptide sequencing, comprising(a) sequencing a peptide or plurality of peptides at sub-attomole resolution, wherein said sequencing is capable of discriminating all 20 proteinogenic amino acids.
50. A method for peptide sequencing, comprising(a) sequencing a peptide or plurality of using a nanopore, wherein said sequencing is capable of discriminating all 20 proteinogenic amino acids.