Methods of n-terminal tagging of proteins
The method addresses the challenge of N-terminal tagging by denaturing proteins, attaching specific linkers and tags, and blocking residues, enhancing conjugation efficiency and compatibility with nanopore experiments.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- NORTHEASTERN UNIV (US)
- Filing Date
- 2025-10-31
- Publication Date
- 2026-05-07
AI Technical Summary
Existing methods for N-terminal tagging of proteins face challenges in conjugation efficiency due to steric hindrance and protein conformation, particularly when using high concentration denaturants which interfere with subsequent processes like nanopore experiments.
A method involving denaturation, linker attachment, and specific tagging of proteins using linkers and tags like PEG-D20, E76, and Cu(II) catalysis to enhance conjugation efficiency, followed by steps to block lysine and glutamate residues, ensuring compatibility with nanopore translocation.
The method achieves effective N-terminal tagging of proteins, enhancing their suitability for nanopore experiments by improving conjugation efficiency and maintaining protein integrity for analysis.
Smart Images

Figure US2025053583_07052026_PF_FP_ABST
Abstract
Description
Docket No. 5200.2432001 (INV-25057)Methods of N-Terminal Tagging of ProteinsRELATED APPLICATION
[0001] This application claims the benefit of U.S. Provisional Application No.63 / 714,831, filed on October 31, 2024. The entire teachings of the above application are incorporated herein by reference.GOVERNMENT SUPPORT
[0002] This invention was made with government support under R01HG012553 awarded by the National Institutes of Health. The government has certain rights in the invention.INCORPORATION BY REFERENCE OF MATERIAL IN XML
[0003] This application incorporates by reference the Sequence Listing contained in the following extensible Markup Language (XML) file being submitted concurrently herewith: a) File name: 5200-2432-001_Sequence_Listing.xml; created October 31, 2025, 1,955 Bytes in size.BACKGROUND
[0004] In recent years, multiple chemical moieties have been used to target the N- terminus of native peptides and proteins [De Rosa L, Di Stasi R, Romanelli A, D'Andrea LD. Exploiting Protein N-Terminus for Site-Specific Bioconjugation. Molecules. 2021 Jun 9;26(12):3521.]. However, compared with model peptides, proteins with more complex conformations generally demonstrate fluctuating or low conjugation efficiencies due to steric hinderance that limits access to the N-terminus. MacDonald, J., Munch, H., Moore, T. et al. One-step site-specific modification of native proteins with 2- pyridinecarboxyaldehydes. Nat Chem Biol 11, 326-331 (2015). Studies have not yet demonstrated N-terminal tagging proteins for nanopore experiments.SUMMARY
[0005] Recent advances in single-file protein translocation through nanopores foresee a new paradigm of proteomics, including characterization, fingerprinting and potentially sequencing. In some promising architectures of nanopore setup, a tag at one end of the protein is needed. For example, in [Yu, L., Kang, X., Li, F. et al. Unidirectional single-file- 1 -4242682. vlDocketNo. 5200.2432001 (INV-25057) transport of full-length proteins through a nanopore. Nat Biotechnol 41, 1130-1139 (2023)], a poly(aspartic acid) tag is needed to incorporate high charge density in favor of electrophoretic force. In another example, [Nivala, J., Marks, D. & Akeson, M. Unfoldase-mediated protein translocation through an a-hemolysin nanopore. Nat Biotechnol 31, 247-250 (2013).,] uses an SsrA tag is needed to enable recognition of the peptide chain by an unfoldase. WO 2024 / 094986 discloses polynucleotides that are conjugated to both ends to attach the peptide chain to the carrier strand. These references used recombinant proteins with a built-in tag sequence or certain wild type proteins bearing desired features. To use nanopore technologies to study protein samples of interest from specimens, a process for terminal tagging native protein molecules is needed.
[0006] The methods described herein relate to a process for library preparation of protein samples for proteomic assays, such as nanopore translocation experiments. In several embodiments, the methods include: (i) An optional preprocessing step in which the protein sample is denatured and unfolded; trace amount of transition metal ions are removed to avoid interference with the conjugation chemistry; and reduction and alkylation are performed on proteins to cleave and block disulfide bonds; (ii) a step in which an N-terminal specific linker is attached to the protein sample; (iii) a step in which covalent blocking of lysine and / or glutamate and aspartate is performed to avoid carbamylation by degradation products of ureaderived denaturants and to adjust the electric charge of the protein sample; (iv) a step in which tag that enables effective nanopore capture and / or translocation is added to the protein sample via the linker; (v) a step in which tagged proteins are extracted; (vi) an optional postprocessing step in which reduction is performed on proteins to cleave disulfide bonds, for testing on proteomic assays such as nanopore translocation experiments. The utility of the methods described herein are exemplified in, but not limited to, nanopore translocation experiments.
[0007] The methods described herein involve protein library preparation by N-terminal tagging for nanopore translocation. In some embodiments, high concentration of denaturant to unfold the proteins was utilized, but it was observed that using high concentration denaturants interferes with various processes including the conjugation chemistry, purification of proteins, and nanopore experiments, depending on the type of denaturant. The methods described herein address such conflicts by design of the tagging process.
[0008] In some embodiments, a method of tagging a protein is described. The method can include: a) optionally denaturing a protein; b) reacting a linker and the protein to form a- 2 -4242682. vlDocket No. 5200.2432001 (INV-25057) linker-functionalized protein, wherein the linker is selected from the group consisting of:from 1 to 12, and m is an integer from 1 to 6; c) optionally denaturing the linker- functionalized protein, wherein at least step a) or step c) is performed prior to step d); andN3^ X reacting the linker-functionalized protein with a tag, wherein the tag includes AN3^X^ , N3^ XA J , or A J, , wherein a is an integer from 1 to 100, X is a canonical poly amino acid repeating in x residues, wherein the canonical poly amino acid residue includes negatively charged amino acids or a combination of negatively charged amino acid and neutral amino acids, wherein x is an integer from 1 to 200, J is an amino acid including SEQ ID NO: 1, and E is a solid phase bead. In some embodiments, the method further includes reacting the linker-functionalized protein with an amine specific protecting compound of a molecular weight less than 1 kDa after step c) and before step d). In further embodiments, the method includes reducing the linker- functionalized protein and / or reacting the linker-functionalized protein with an alkylating agent before reacting the linker functionalized protein with the amine specific protecting compound. In further embodiments, step d) further includes reacting a Cu(II) salt and- 3 -4242682. vlDocket No. 5200.2432001 (INV-25057) quenching. In some other embodiments, A isL J aand has a molecular weight greater than IkDa. In some other embodiments, the amine specific protecting compound includes an N-hydroxysuccinimide (NHS) ester, a pyrocarbonate, or a carboxylic anhydride. In some other embodiments, steps a) and b) are performed at the same time. In some embodiments,polyethylene glycol (PEG)5K-poly(aspartic acid)3K, N3-poly(glutamic acid)10K, or N3-PEG10K. In some further embodiments, the tag further includes SEQ ID NO: 1. In some other embodiments, the concentration of the denaturant is about 0.1% to about 96% wt / wt%.
[0009] In some other embodiments, an N-terminally modified protein includes the4242682. vlDocket No. 5200.2432001 (INV-25057)from 1 to 6, R is a canonical amino acid side chain of the protein, wherein Z is4242682. vlDocket No. 5200.2432001 (INV-25057), wherein a is an integer from 1 to 100, X is a canonical poly amino acid repeating in x residues, wherein x is an integer from 1 to 200, J is an amino acid sequence including SEQ ID NO: 1, and E is a solid phase bead. In some further embodiments, A includeshas a molecular weight greater than IkDa.
[0010] In some other embodiments, a kit for chemically modifying proteins to make a tagged protein is described. The kit includes: i) a denaturant for a protein; ii) a linker, wherein4242682. vlDocket No. 5200.2432001 (INV-25057)N3_X„, N3_X^, EA J , or A J , wherein A is, wherein a is an integer from 1 to 100, X is a canonical poly amino acid repeating in x residues, wherein x is an integer from 1 to 200, J is an amino acid sequence including SEQ ID NO: 1, and E is a solid phase bead. In some embodiments, the kit further includes a Cu(II) salt, sodium ascorbate, and a chelator. In some other embodiments, the kit includes an N-Hydroxysuccinimide (NHS) ester of a molecular weight less than 1 kDa or a pyrocarbonate of a molecular weight less than 1 kDa. In some other embodiments, the linker includessome other embodiments, the tag includes Ns-polyethylene glycol (PEG)5K-poly(aspartic acid)3K (PEG- 020 tag), N3-poly(glutamic acid)10K (E76 tag), or N3-PEGIOK (PEG10K tag). In some further embodiments, the tag includes SEQ ID NO: 1. In some other embodiments, the denaturant includes guanidinium salts, perchlorate salts, urea, urea derivatives, sodium dodecyl sulfate (SDS), polyethylene glycol tert-octyl phenyl ether, polyoxyethylene (10) lauryl ether, polyoxyethylene (20) cetyl ether, dodecyulguanidine acetate (dodin), cetyltrimethylammonium bromide (CTAB), Nonyl phenoxypolyethoxylethanol (NP-40), dimethyl sulfoxide (DMSO), dimethyl formamide (DMF), or N-methylpyrrolidone (NMP), or any combination thereof.BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The foregoing will be apparent from the following more particular description of example embodiments, as illustrated in the accompanying drawings in which like reference characters refer to the same parts throughout the different views. The drawings are not necessarily to scale, emphasis instead being placed upon illustrating embodiments.
[0012] FIG. 1 shows a sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS- PAGE) gel of tagging concanavalin A (ConA) with Ns-polyethylene glycol 5K (PEG5K)-4242682. vlDocket No. 5200.2432001 (INV-25057) poly(aspartic acid)3K (PEG-D20; the 5K and 3K represent the molecular weights of the polymers) experiments. Lanes are numbered 2-7 (from right to left) and are from one experiment and lanes 8-13 are from another experiment with the lanes labelled as follows: 1) pristine ConA control; 2) reaction mixture after step 4 showing ConA-PEG-D20 product; 3) strong cation exchange (SCX) binding flow-through; 4) SCX wash flow-through; 5) strong anion exchange (SAX) elution peak 1; 6) SAX elution peak 2; 7) SAX elution peak 3; 8) reaction mixture after step 4 showing ConA-PEG-D20 product; 9) SCX binding flow- through; 10) SCX wash flow-through; 11) SAX elution peak 1; 12) SAX elution peak 2; 13) SAX elution peak 3.
[0013] FIG. 2 shows an SDS-PAGE gel of a tagging ConA with PEG-D20 experiment. Lanes (from right to left) are numbered as follows: 1) pristine ConA control; 2) SCX elution;3) SAX elution peak 1; 4) SAX elution peak 2; 5) SAX elution peak 3.
[0014] FIG. 3 shows an SDS-PAGE gel of tagging ConA with PEG-D20 and N3- poly(glutamic acid) 1 OK (E76 or E76 tag) experiments. Lanes 1-4 and 8 (from right to left) are results of PEG-D20 tag and Lanes 5-7 and 9 are results of E76 tag with the lanes labelled as follows: 1) pristine ConA control; 2) reaction mixture after step 4 showing ConA-PEG- D20 product; 3) SCX binding flow-through; 4) SCX elution; 5) reaction mixture after step 4; 6) SCX binding flow through; 7) SCX elution; 8) SCX elution with heating at 95 °C for 5 min for the SDS-PAGE; 9) SCX elution with heating at 95 °C for 5 min for the SDS-PAGE.
[0015] FIG. 4 shows an SDS-PAGE gel of tagging ConA with an E76 tag. Lanes 2 and 3 (from right to left) are the results of pyridine-2-carbaldehyde (PCA)-PEG4 (4 PEG polymers)-dibenzocyclooctyne amine (DBCO), and lanes 4-6 are results of linker 5-ethynyl- pyridine-2-carbaldehyde (5-E-2-PCA) and lanes are labelled as follows: 1) pristine ConA control; 2) SAX peak 1; 3) SAX peak 2; 4) SAX peak 1; 5) SAX peak 2; 6) SAX peak 3. N- hydroxysuccinimide (NHS) acetate was used in step 3 to block the lysine sidechains of the protein.
[0016] FIG. 5 shows an SDS-PAGE gel of tagging ConA with PEG-D20 tag. Lanes (2-4), (5, 6), (7, 8), (9,10) (from right to left) are results of separate experiments. The linker for Lanes (2-4), (5, 6) and (9, 10) is 5-E-2-PCA. The linker for lanes (7, 8) is PCA-PEG4-DBCO. The sample of Lanes (5, 6) was reacted with NHS acetate to block the lysine sidechains during Step 3. The SDS-PAGE was heated at 95 °C for 5 min. The lanes are labelled as follows: 1) pristine ConA control; 2) SAX Peak 1; 3) SAX Peak 2; 4) SAX Peak 3; 5) SAX Peak 1; 6) SAX Peak 2; 7) SAX Peak 2; 8) SAX Peak 3; 9) SAX Peak 2; 10) SAX Peak 3.- 8 -4242682. vlDocketNo. 5200.2432001 (INV-25057)
[0017] FIG. 6 shows an SDS PAGE gel of tagging streptavidin and ovalbumin. The SDS- PAGE was heated at 95 °C for 5 min. Lanes are labelled as follows (from right to left): 1) ovalbumin control; 2) ovalbumin after step 4, using PCA-PEG4-DBCO and PEG-D20; 3) ovalbumin after step 4, using 5-E-2-PCA and PEG-D20; 4) streptavidin control; 5) streptavidin after step 4, using PCA-PEG4-DBCO and PEG-D20; 6) streptavidin after step 4, using 5-E-2-PCA and PEG-D20.
[0018] FIG. 7 shows SDS-PAGE gel experiments of tagging streptavidin and ovalbumin. Lanes are labeled as follows (from right to left): 1) streptavidin control; 2) streptavidin after step 2 using PCA-PEG4-DBCO; 3) streptavidin after step 4 using PCA-PEG4-DBCO; 4) streptavidin after step 2 using 5-E-2-PCA; 5) streptavidin after step 4 using 5-E-2-PCA; 6) ovalbumin control; 7) ovalbumin after step 2 using PCA-PEG4-DBCO; 8) ovalbumin after step 4 using PCA-PEG4-DBCO; 9) ovalbumin after step 2 using 5-E-2-PCA; 10) ovalbumin after step 4 using 5-E-2-PCA; 11-15) irrelevant samples.
[0019] FIG. 8 shows SDS-PAGE gel experiments of tagging ConA with PEG-D20. Lanes (2-5) and (7-10) (from right to left) are separate experiments. The lanes are labelled as follows: 1) pristine ConA control; 2) ConA after step 2 using PCA-PEG4-DBCO; 3) SAX thereof, peak 1; 4) SAX thereof, peak 2; 5) SAX thereof, peak 3; 6) blank; 7) ConA after step 2 using 5-E-2-PCA; 8) SAX thereof, peak 1; 9) SAX thereof, peak 2; 10) SAX thereof, peak 3.
[0020] FIG. 9 shows SDS-PAGE gel experiments of tagging ConA and lysozyme using PEG-D20. Lanes (1-4) and (6-9) (from right to left) are separate experiments. The lanes are labelled as follows: 1) Pristine lysozyme control; 2) lysozyme SAX peak 1; 3) SAX thereof, peak 2; 4) SAX thereof, peak 3; 5) blank; 6) pristine ConA control; 7) SAX thereof, peak 1;8) SAX thereof, peak 2; 9) SAX thereof, peak 3.
[0021] FIG. 10 shows SDS-PAGE gel experiments of various incubation times in step 2 using 0.2 mM Cu(II) catalyzed reaction. The SDS-PAGE was heated at 37 °C for 1 hr. The lanes are numbered from right to left. The lanes are labelled as follows: 1) pristine lysozyme control; 2) lysozyme + PCA-PEG4-DBCO, 6 hrs; 3) lysozyme + PCA-PEG4-DBCO, 24 hrs;4) lysozyme + PCA-PEG4-DBCO, 29 hrs; 5) lysozyme + 5-E-2-PCA, 6 hrs; 6) lysozyme + 5- E-2-PCA, 24 hrs; 7) lysozyme + 5-E-2-PCA, 29 hrs; 8) ConA control; 9) ConA + PCA- PEG4-DBCO + PEG-D20, 5 hrs; 10) ConA + 5-E-2-PCA + PEG-D20, 5 hrs.
[0022] FIG. 11 shows SDS-PAGE gel experiments of various incubation times in step 2 using 0.2 mM Cu(II) catalyzed reaction. The SDS-PAGE was heated at 37 °C for 1 hr. The- 9 -4242682. vlDocket No. 5200.2432001 (INV-25057) lanes are numbered from right to left. The lanes are labelled as follows: 1) pristine lysozyme control; 2) lysozyme + 5-E-2-PCA + PEG-D20, 6 hrs, size exclusion chromatography (SEC) fraction 6; 3) SEC thereof, fraction 7; 4) SEC thereof, fraction 11; 5) lysozyme + PCA- PEG4-DBC0 + PEG-D20, 6 hrs; 6) lysozyme + PCA-PEG4-DBC0 + PEG20, 24 hrs; 7) lysozyme + PCA-PEG4-DBCO + PEG20 29 hrs; 8) lysozyme + 5-E-2-PCA + PEG-D20, 6 hrs; 9) lysozyme + 5-E-2-PCA + PEG-D20, 24 hrs; 10) lysozyme + 5-E-2-PCA + PEG-D20, 29 hrs; 11-15) irrelevant samples.
[0023] FIG. 12 shows SDS-PAGE gel experiments of comparison of various tagging methods of lysozyme. Lanes are numbered form right to left. The lanes are labelled as follows: 1) pristine lysozyme control; 2) lysozyme + PCA-PEG4-DBCO, catalyzed by 0.2 mM Cu(II), 6 hrs; 3) lysozyme + PCA-PEG4-DBCO + PEG-D20, catalyzed by 0.2 mM Cu(II), 6 hrs; 4) Lysozyme + PCA-PEG4-DBCO + PEG-D20, catalyzed by 0.2 mM Cu(II) 6 hrs, treated with 200 mM citric acid, 100 mM (Gly)s tripeptide, 200 mM MeNFE-HCl; 5) SEC thereof, fraction 6; 6) SEC thereof, fraction 7; 7) SEC thereof, fraction 11; 8) lysozyme + PCA-PEG4-DBCO + E76; 9) lysozyme + PCA-PEG4-DBCO + E76, treated with 200 mM citric acid, 100 mM (Gly)s tripeptide, 200 mM MeNEh-HCl; 10) SEC thereof, fraction 9; 11) SEC thereof, fraction 11; 12) lysozyme + PCA-PEG4-DBCO + PEG-10K; 13) lysozyme + PCA-PEG4-DBCO + PEG-D20; 14) lysozyme + 5-E-2-PCA catalyzed by 0.2 mM Cu(II), 6 hrs; 15) Lysozyme + 5-E-2-PCA + PEG-D20 catalyzed by 0.2 mM Cu(II), 6 hrs.
[0024] FIG. 13 shows SDS-PAGE gel experiments at varying concentrations of Cu(II) and varying temperatures. Lanes are numbered from right to left and are labelled as follows: 1) pristine lysozyme control; 2) lysozyme + PCA-PEG4-DBCO + E76, no catalyst, 37 °C, 27 hrs; 3) lysozyme + PCA-PEG4-DBCO + E76, catalyzed by 1 mM Cu(II), 37 °C, 6 hrs; 4) lysozyme + PCA-PEG4-DBCO + E76, catalyzed by 1 mM Cu(II), 4 °C, 6 hrs; 5) lysozyme + PCA-PEG4-DBCO + E76, catalyzed by 0.2 mM Cu(II), 37 °C, 6 hrs; 6) lysozyme + PCA- PEG4-DBCO + E76, no catalyst, 60 °C, 27 hrs; 7) lysozyme + 5-E-2-PCA + E76, catalyzed by 1 mM Cu(II), 37 °C, 6 hrs; 8) lysozyme + PCA-PEG4-DBCO + E76, no catalyst, 37 °C, 27 hrs; 9) SEC thereof, fraction 8; 10) SEC thereof, fraction 12; 11) pristine ConA control; 12) ConA + PCA-PEG4-DBCO + E76, 37°C, 26 days; 13) E76 tag control.
[0025] FIG. 14 shows SDS-PAGE gel experiments of lysozyme with varying Cu(II) concentrations and varying temperatures. Lanes 1-10 are irrelevant. Lanes are numbered form right to left and are labelled as follows: 11) pristine lysozyme control; 12) lysozyme + PCA-- 10 -4242682. vlDocket No. 5200.2432001 (INV-25057)PEG4-DBC0 + PEG-D20, 50°C, 2 days; 13) lysozyme + PCA-PEG4-DBCO + PEG-D20, 37°C, 2 days, 1 mM Cu(II).
[0026] FIG. 15 shows SDS-PAGE gel experiments. Lanes 1-6 are experiments of IgG- binding protein A and ConA with various methods and proteins. Experiments of lysozyme show extraction using SEC for Lanes 7-12. Lanes 7 and 8 are one experiment and lanes 9-11 are one experiment. Lanes are numbered from right to left and labelled as follows: 1) pristine protein A control; 2) protein A + PCA-PEG4-DBCO + PEG-D20; 3) pristine ConA control;4) ConA + PCA-PEG4-DBCO + PEG-D20; 5) pristine lysozyme control; 6) lysozyme + PCA-PEG4-DBCO + PEG-D20; 7) SEC fraction 7; 8) SEC fraction 6; 9) SEC fraction 8; 10) SEC fraction 7; 11) SEC fraction 6; 12) lysozyme + 5-E-2-PCA + PEG-D20.
[0027] FIG. 16 shows SDS-PAGE gel experiments. Lanes 1-4 are experiments comparing various denaturants for step 4. Lanes 5-12 are experiments studying cleavage effects of treatment of 200 mM citric acid, 100 mM (Gly)s tripeptide, 200 mM MeNFE-HCl. The model is lysozyme + PCA-PEG4-DBCO + PEG10K. The lanes are numbered from right to left and are labelled as follows: 1) Pristine lysozyme control; 2) lysozyme + PCA-PEG4-DBCO + PEG-D20; 3) lysozyme + 5-E-2-PCA + PEG-D20, Step 4 with 8M urea; 4) lysozyme + 5-E- 2-PCA + PEG-D20, Step 2 buffer exchanged to 2% 3-((3-cholamidopropyl) dimethylammonio)-l -propanesulfonate (CHAPS), then adding 8M tetramethylurea (TMU);5) SEC fraction 6; 6) SEC fraction 8; 7) SEC fraction 9; 8) SEC fraction 11; 9) SEC fraction 13; 10) SEC fraction 15; SEC fraction 18; SEC fraction 8 of non-treated sample.
[0028] FIG. 17 shows SDS-PAGE gel experimenters with validation of gel excision for lanes 1-5, validation of one-pot lysozyme + PCA-PEG4-DBCO + PEG10K extraction using SEC in (0.5% SDS, phosphate-buffered saline (PBS)) for lanes 6-12. The lanes are numbered from right to left and are labelled as follows: 1) pristine lysozyme control; 2) irrelevant sample: 3) lysozyme + PCA-PEG4-DBCO + PEG-D20 extraction by gel excision; 4) pristine lysozyme (6 M guanidinium chloride (GdmCl), PBS) buffer exchanged to (8 M urea, PBS) by ultrafiltration; 5) irrelevant sample; 6) SEC fraction 6; 7) SEC fraction 8; 8) SEC fraction 9; 9) SEC fraction 11; 10) SEC fraction 13; 11) SEC fraction 15; 12) SEC fraction 18; 13) pristine protein A control; 14) pristine ConA control; 15) pristine BSA control.
[0029] FIG. 18 shows SDS-PAGE gel experiments comparing PEG-D20 tags and PEG10K tags. Lanes are numbered from right to left and are labelled as follows: 1) pristine lysozyme control; 2) lysozyme + PCA-PEG4-DBCO + PEG10K; 3) lysozyme + PCA-PEG4- DBCO + PEG10K; 4-6) Irrelevant samples; 7) lysozyme + PCA-PEG4-DBCO + PEG-D20-- 11 -4242682. vlDocket No. 5200.2432001 (INV-25057) acetate; 8) lysozyme + PCA-PEG4-DBCO + PEG-D20-Boc; 9) lysozyme + PCA-PEG4- DBCO + PEG-D20; 10-14) irrelevant samples; 15) pristine lysozyme control.
[0030] FIG. 19 shows SDS-PAGE gel validating lysozyme purification by various methods. Lanes are numbered right to left and are labelled as follows: 1) pristine lysozyme control; 2) pristine lysozyme precipitated from (6 M GdmCl, PBS) by 10:1 volume ratio isopropyl alcohol, overnight; 3) pristine lysozyme precipitated from (6 M GdmCl, PBS) by 10:1 volume ratio ethyl alcohol, overnight; 4) pristine lysozyme precipitated from (6 M GdmCl, PBS) by 40:1 volume ratio isopropyl alcohol, overnight, 4 °C; 5) pristine lysozyme precipitated from (6 M GdmCl, PBS) by 40:1 volume ratio ethyl alcohol, overnight, 4 °C; 6) pristine lysozyme control; 7) pristine lysozyme desalted by Zeba 7kDa spin column, in (2% Triton™ X-100 (polyethylene glycol / c / 7-octyl phenyl ether), PBS); 8) pristine lysozyme desalted by Zeba 7kDa spin column, in (2% 3-([3-cholamidopropyl]dimethylammonio)-2- hydroxy-1 -propanesulfonate (CHAPSO), PBS) 9) pristine lysozyme ultrafiltered by Amicon Ultra 4 molecular weight cut off (MWCO) lOkDa, from (6M GdmCl, PBS) to (8M Urea, PBS); 10) pristine lysozyme ultrafiltered by Amicon Ultra 4 MWCO lOkDa, in (8M TMU, PBS); 11) pristine lysozyme ultrafiltered by Amicon Ultra 4 MWCO lOkDa, in (2% CHAPSO, PBS); 12) pristine lysozyme control in (6M GdmCl, PBS) directly mixed with gel loading buffer and loaded into gel.
[0031] FIG. 20 shows SDS-PAGE gel validation of step 2 reaction conditions using lysozyme. Reaction was carried out as a one-pot mixture with PCA-PEG4-DBCO linker and 1 : 1 molar ratio of PEG10K tag. The lanes are numbered right to left and are labelled as follows: 1) pristine lysozyme control; 2-4) irrelevant samples; 5) reaction with 2% SDS, 1 mM linker; 6) reaction with 2% Triton™ X-100 (polyethylene glycol / cvV-octyl phenyl ether), 1 mM linker; 7) reaction with no denaturant, 1 mM linker; 8) reaction with 2% SDS, 5 mM linker; 9) reaction with no denaturant, 5 mM linker; 10) pristine lysozyme control; 11) pristine lysozyme control with reducing loading buffer; 12) pristine lysozyme control in (6 M GdmCl, PBS) directly mixed with gel loading buffer and loaded into gel.
[0032] FIG. 21 shows an SDS-PAGE gel of BSA using various denaturants in step 2. Reaction with 1 mM PCA-PEG4-DBCO linker and 1 : 1 molar ratio of PEG10K tag, then free linker removed with desalting (Zeba 7kDa) or ultrafiltration (Pierce Concentrator lOkDa). Lanes are numbered from right to left with lanes 2-5 using desalting and lanes 7-11 using ultrafiltration and are labelled as follows: 1) ladder; 2) non-denatured in PBS; 3) 8 M urea, PBS; 4) 2% Triton™ X-100 (polyethylene glycol / cvV-octyl phenyl ether), PBS; 5) 2% SDS,- 12 -4242682. vlDocket No. 5200.2432001 (INV-25057)PBS; 6) irrelevant sample; 7) non-denatured in PBS; 8) 8 M urea, PBS; 9) 2% Triton™ X- 100 (polyethylene glycol tert-octyl phenyl ether), PBS; 10) 2% SDS, PBS; 11) irrelevant sample; 12) irrelevant sample.
[0033] FIG. 22 is a chromatogram of lysozyme + PCA-PEG4-DBCO + PEG-D20 extraction by SEC. The chromatography column was Superdex 75 10 / 300 GL.
[0034] FIG. 23 is a chromatogram of ConA + 5-E-2-PCA + PEG-D20 extraction by SAX. The chromatography column was EconoFit Macro-Prep High Q. The pH of buffer was around 5, adjusted by N,N’ -dimethylpiperazine.
[0035] FIG. 24 shows a trace of tagged product of nanopore translocation. The nanopore was alpha hemolysin (aHL), and buffer was 1 M KC1, 2 M GdmCl, 10 mM Tris, pH 7.5. From to bottom, the traces come from species of: 1) ConA - 5-E-2-PCA - PEG-D20; 2) lysozyme - PCA-PEG4-DBCO - E76; 3) lysozyme - 5-E-2-PCA - E76; 4) lysozyme - PCA- PEG4-DBCO - PEG-D20; 5) lysozyme - PCA-PEG4-DBCO - E76; 6) lysozyme - 5-E-2- PCA - E76.
[0036] FIGs. 25A-J show a scheme of a process for making tagged protein. FIG. 25A is a scheme showing the two chemical reactions of the process; FIG. 25B shows two examples of a tagged protein; FIG. 25C is a chromatogram showing extraction of tagged ConA; FIG. 25D is a trace of translocation of a tagged ConA; FIG. 25E is a detailed view of a translocation event of a tagged ConA; FIG. 25F is a scatter plot showing a blockade - dwell time distribution of two stages of the translocation events of tagged ConA; FIG. 25G is a chromatogram showing extraction of a tagged lysozyme; FIG. 25H is a trace of a translocation of a tagged lysozyme; FIG. 251 is a detailed view of a translocation event of a tagged lysozyme; FIG. 25J is a scatter plot showing a blockade - dwell time distribution of a tagged lysozyme.
[0037] FIG. 26 shows a structure of tags that can be bound to a protein using methods described herein.
[0038] FIG. 27 is a graph showing reduced reactivity of a CuAAC reaction in the presence of increased guanidinium salt concentration.
[0039] FIG. 28 shows SDS-PAGE gel experiments of extracted tagged proteins. The lanes are numbered from right to left and are labelled as follows: 1) ConA + 5-E-2-PCA control; 2) ConA + 5-E-2-PCA + PEG-D20 reaction mixture; 3) ConA-PEG-D20 extracted by gel excision and electroelution; 4) blank; 5) beta Casein + 5-E-2-PCA control; 6) beta Casein + 5-E-2-PCA + PEG-D20 reaction mixture; 7) beta Casein-PEG-D20 extracted by gel- 13 -4242682. vlDocket No. 5200.2432001 (INV-25057) excision and electroelution; 8) blank; 9) 5) lactalbumin + 5-E-2-PCA control; 10) lactalbumin + 5-E-2-PCA + PEG-D20 reaction mixture; 11) lactalbumin-PEG-D20 extracted by gel excision and electroelution; 12) blank; 13) blank; 14) ladder.
[0040] FIG. 29 is a graph of raw translocation data of tagged ConA through CytK. Details of sample preparation are listed in Example 3.
[0041] FIGs. 30A-B are graphs of segmentation of translocation events into steps. The data was a translocation of PEG-D20 tagged ConA through CytK. The triangle symbol in FIG. 30A shows a segmentation start and end point of each segment. The triangle symbol in FIG. 30B shows the dwell time and blockade of each segment in a scatter plot. A well- defined two-step translocation was observed.
[0042] FIGs. 31A-F are scatter plots of parts of translocation events of PEG-D20 tagged ConA through CytK. The segmentation of the data was performed as shown in FIGs. 30A-B. FIGs. 31A-B, FIGs. 31C-D, and FIGs. 31E-F were events of translocation under 100 mV, 120m V, and 140m V, respectively. Each part of the events showed clear aggregated distribution, indicating well-defined dynamics of translocation. The first part corresponds to the tag, and the second part to the analyte protein.
[0043] FIGs. 32A-B are graphs of the cumulative distribution function of the dwell time of the two parts of the translocation events of PEG-D20 tagged ConA through CytK. Clearly the dwell time of the second part (FIG. 32B) showed a decrease as the voltage increased, consistent with faster translocation of the protein with higher voltage. The dwell time of the first part (FIG. 32A) did not show a clear trend as the voltage increased, indicating that the time of this part is not dominated by the translocation of the tag per se, but by the initiation of the translocation of the protein, which may require entropy-dominated conformation adjustments, not strongly dependent on the driving voltage.
[0044] FIGs. 33 A-D are graphs and scatter plots identifying that there are two different directions of translocation of the PEG-D20 tagged ConA through MspA. Note the reversed order of the shallow tag level and the protein. The two types of translocation events each form a clear population on the scatter plot. FIG. 33B is the scatter plot corresponding to FIG. 33A and FIG. 33D corresponds to FIG. 33C. The triangles in each scatter plot indicate the corresponding steps shown as lines their respective graphs (FIGs. 33A and 33C).- 14 -4242682. vlDocket No. 5200.2432001 (INV-25057)DETAILED DESCRIPTION
[0045] A description of example embodiments follows.
[0046] A challenge of N-terminal tagging of proteins is conjugation efficiency, or reactivity challenges, which can be due to the conformation of the protein. Denaturing proteins unfolds the protein structure and may increase conjugation efficiency in some circumstances. For example, [Sahu, S., Emenike, B., Beusch, C.M. et al. Copper(I)-nitrene platform for chemoproteomic profiling of methionine. Nat Commun 15, 4243 (2024)] used a low concentration of SDS to improve conjugation of methionine side chain residues, but not the amino acid backbone. However, denaturant selection can affect the ability to use the tagged protein in further analysis. As discussed in Example 2, it was observed that SDS was incompatible with nanopores due to destruction of lipid membranes and was difficult to remove. To thoroughly unfold the proteins, a high concentration of denaturant should be used, which has not been previously explored with these conjugation chemistries.Table 1: Glossary- 15 -4242682. vlDocket No. 5200.2432001 (INV-25057)
[0047] For organics in Table 1, salts formed with simple inorganic counterions are included within the scope of the abbreviations.Definitions
[0048] As used herein, an “alkylating agent” refers to a compound capable of reacting with a primary amine, secondary amine, and / or a sulfur group that binds an alkyl group to the primary amine, secondary amine, and / or sulfur group. Some non-limiting examples of alkylating agents include: iodoacetamide, N-ethylmaleimide (NEM), haloacetamides, and maleimides.
[0049] As used herein, a “canonical amino acid” refers to the 20 naturally occurring amino acids.
[0050] As used herein, “poly amino acid” refers to a repeating polymer of amino acids as monomeric units. A non-limiting example is poly(glutamic acid) which is a polymer of- 16 -4242682. vlDocket No. 5200.2432001 (INV-25057) glutamic acids, which can be defined as having an average number residues, molecular weight, or a combination thereof.
[0051] As used herein, “deep eutectic solvents’ refers to a solution of Lewis or Bronsted acids and bases that form a homogenous mixture that has a melting point lower than the acid or base individually.
[0052] As used herein, “chaotrope” refers to a molecule in an aqueous solution that disrupts the hydrogen bonding network of between the molecules of water.Protein Samples
[0053] Protein samples can be crude protein extract of solid or solution form, from biological material of interest, e.g., cell lysate, tissue biopsy, secretions, etc. They can also be any further purified protein material by means such as chromatography, electrophoresis or precipitation. Noteworthily, antibodies can be used to enrich certain proteins of interest. The sample can contain one or more types of proteins. In addition to proteins mentioned in the Exemplification, this method can be applied to a variety of proteins (FIGs. 6, 7, 15, 21, and 28)Use of Denaturants
[0054] In some embodiments, the whole process (i) through (vi) is conducted in common biologically compatible buffers e.g. PBS without denaturants. Alternatively, in some embodiments, denaturants are added to keep the proteins in the sample denatured and unfolded which provides advantages including: (a) The quaternary, tertiary and secondary structures of the proteins are broken so that the protein molecules assume a looser conformation, which enhances the access of the N-terminus of the molecule for the conjugation chemistry (FIG. 20); (b) Alterations of proteins due to activity of enzymes contained in the sample, such as degradation due to proteases and modifications due to PTM enzymes e.g. kinases, may compromise the integrity of the analyte proteins. Most enzymes are non-functional in denatured conditions, so the integrity of the analyte proteins is improved with denaturants; (c) Samples containing toxic proteins may pose safety concerns during handling. The toxicity of toxic proteins may be reduced in denatured conditions, so the safety of the process may be improved; (d) Some proteins, especially membrane proteins, are prone to aggregation in aqueous buffer. Denaturants are typically used to solubilize such proteins before further processing; (e) Proteins of interest can be of low abundance, and- 17 -4242682. vlDocket No. 5200.2432001 (INV-25057) immunoadsorption can be used to pool the protein. Denaturants can separate the antigenantibody complex to make the proteins of interest available for tagging.
[0055] Denaturants may be used alone or in any combination thereof. Denaturants used include: a) Chaotropes, including in non-limiting embodiments, guanidinium salts, perchlorate salts, urea, urea derivatives, thiourea, thiourea derivatives, imidazolidinone, bromide salts, hexafluorophosphate salts, tetrafluoroborate salts, thiocyanate salts, tetraalkylammonium salts, lithium salts; b) detergents, including in non-limiting embodiments, SDS, Triton™ X-100 (polyethylene glycol tert-octyl phenyl ether), Brij™-35 (polyoxyethylene (10) lauryl ether), Brij™-58 (polyoxyethylene (20) cetyl ether), dodecyulguanidine acetate (dodin), CTAB, NP-40 (Nonyl phenoxypolyethoxylethanol), cetyltrimethylammonium salts, Nonyl phenoxypolyethoxylethanol (NP-40), etc.; c) organic solvents, including in non-limiting embodiments, DMSO, DMF, NMP, N-methylpyrrolidone (NMP), formamide, N-methylformamide, acetamide, N-methylacetamide, dimethylacetamide (DMAc), hexafluoro-2-propanol (HFIP), trifluoroethanol (TFE), acetonitrile, phenol etc.
[0056] Denaturants should be used at high enough concentrations to favor thorough unfolding of proteins. Table 2 lists typical minimum concentrations for denaturants when used alone. It is further noted that percentages used in Table 2 refer to wt / wt% for solids and vol / vol% for liquids.Table 2: Typical minimum concentrations of denaturants4242682. vlDocket No. 5200.2432001 (INV-25057)
[0057] To be noted is that SDS and guanidinium salts cannot be combined because precipitation will happen. Similarly, CTAB and perchlorate salts cannot be combined. Urea and DMU degrade to give isocynates, which will further react with lysine on proteins. To retain protein sequence integrity due to such reactions, generally, the buffer containing urea or DMU should be prepared afresh and refrigerated for storage. It is observed that when urea is used for the 2-PCA chemistry, smeared bands appear on the SDS-PAGE gel, indicating urea or its degradation product interferes with the 2-PCA chemistry (FIG.16). It is also observed that when urea is used for the CuAAC chemistry, an unexpected band appears on the SDS-PAGE gel. Thus, for certain steps, the usage of urea is avoided. Using DMU for CuAAC does not appear to give unexpected bands. In some embodiments, the concentration of the denaturant is preferred to be as high as possible without deactivating the protein.
[0058] Guanidinium salts is a class of denaturant that functions by the guanidinium onium. Guanidinium chloride is most commonly known in protein research, while other salts,- 19 -4242682. vlDocket No. 5200.2432001 (INV-25057) e.g. guanidinium thiocyanate, guanidinium bromide, can also be used. Guanidinium salts have the advantage of strong denaturing ability, chemical stability (not degrading to cyanates, not reacting with the linkers), easy to remove (as a small molecule, easy to remove by standard desalting methods). Thus, in some embodiments, guanidinium salts are preferred for 2-PCA chemistry. However, it was observed that high concentration of GdmCl strongly inhibits the reactivity of the CuAAC (FIG. 27). Without being bound to one single theory, one possible factor of such inhibition is the sequestering effect of a halide (e.g. chloride) on Cu(I) species needed by CuAAC. The use of GdmCl is preferably less than 2 M in concentration if CuAAC is to be performed within its presence. It is also noted that guanidinium salts are incompatible with ion exchange chromatography (IEX) methods.
[0059] Detergents of various types can be used. Most detergents are difficult to remove by desalting, ultrafiltration, or dialysis. The CHAPS or CHAPSO detergent is an exception because of its small micelle size, making it removable with desalting or ultrafiltration, but it is less potent in unfolding proteins. Notably, ionic detergents are incompatible with IEX methods. Thus, if a detergent is to be used alone or together with other denaturants, non-ionic detergents are preferred. Detergents should be removed during the process for biological nanopore experiments, because the presence of detergents can compromise the stability of bilayer membrane of the nanopore.
[0060] Organic solvents can be used to facilitate unfolding of proteins and increase the solubility of the linkers. Organic solvents that are soluble in water, and tend not to precipitate proteins should be used, such as DMSO, DMF, NMP. TMU can also be considered as an organic solvent. When organic solvent is added into the sample, ultrafiltration should be preceded by diluting the sample, so the concentration of the organic solvent does not damage the ultrafilter, especially the ultrafiltration membrane.Linkers
[0061] FIG. 25A is a schematic showing an embodiment of a linker. The linker includes a pyridine-2-carbaldehyde moiety as shown in FIG. 25A.- 20 -4242682. vlDocket No. 5200.2432001 (INV-25057)
[0062] In some embodiments, the linker is selected from:from 1 to 12, and m is an integer from 1 to 6. More specific examples are described herein.
[0063] In some embodiments, the linker is ethynyl-PCA, where an ethynyl is attached to a 2-PCA ring.
[0064] In some embodiments,preferred due to lower steric hinderance and plenteous commercial availability.
[0065] In some embodiments, the linker is an alkynyl-PCA with a short spacer:- 21 -4242682. vlDocket No. 5200.2432001 (INV-25057)where m is an integer from 1 to 6
[0066] In some embodiments, the linker is PCA-spacer-DBCO or PCA-spacer-propargyl.
[0067] The spacer is a biochemically stable linkage, e.g., PEGs, amides, and hydrocarbon chains. In some embodiments, the linker is PCA-PEG-DBCO or PCA-PEG-propargyl
[0069] where n is an integer from 1 to 12
[0070] 4242682. vlDocket No. 5200.2432001 (INV-25057)
[0071] where m is an integer from 1 to 6
[0072] where n is an integer from 1 to 12
[0073]
[0074] An example of a PCA-PEG linker is PCA-PEG4-DBCO, which has the following structure:
[0075] Tags
[0076] Tags can be a negatively charged oligomer for electrophoretic dragging of the analyte protein towards and through a nanopore. Tags can also be peptide motifs that enable recognition of the analyte protein by a complex of nanopore and motor protein. One example of such motif is SEQ ID NO: 1, which is a small stable RNA A fragment (ssRa) tag. Tags can- 23 -4242682. vlDocket No. 5200.2432001 (INV-25057) also be neutral polymers that assume a loose conformation, so the tag enters the nanopore by electroosmotic force and drags the appended protein payload through the pore. One example of such motif is PEG>lkDa. In the scheme of FIG. 26, X, Y and Z parts show possible components of the tag. All the three components are optional, as long as at least one component is present. The direction of the monomers can be either side for parts X and Y. The curved lines denote common biochemical linkages such as amide, ether, hydrocarbon, etc. to couple the components together. While ssrA is a C-terminal degron, the molecular motor ClpX that recognizes it is capable of traversing a protein chain in either N-C or C-N directions. Thus, it is possible for an N-terminal tagged ssrA degron to deliver the protein through ClpX complex. According to the nanopore setup to be used, if the desired direction of the electric field is opposite, positively charged oligomers will be used in place of negatively charged oligomers. Some examples of tag composition are shown in FIG. 26.
[0077] Tags can serve for more than enhancement of capture or translocation. Segments can be chosen to have characteristic translocation trace when translocating through the nanopore, serving as markers to facilitate interpretation of the translocation signal, including but not limited to calibration of conductivity, probing the status of the pore, barcoding for sample multiplexing, or identification of translocation direction. A characteristic translocation trace includes, but is not limited to, a specific blockade level, certain patterns of spikes or steps, or a sequence of current that matches a certain filter or neural network feature. When barcoding is used for sample multiplexing, the composition of the tag can include block-copolymers to encode the barcode.
[0078] Examples of the tag include N3-PEG5K-poly(aspartic acid)3K (referred to herein as “PEG-D20 tag”), N3-poly(glutamic acid)10K (referred to herein as “E76 tag”), N3- PEG10K (referred to herein as “PEG10K tag”). In N3-PEG5K-poly(aspartic acid)3K, the 5K is referring to the PEG polymer having a molecular weight of about 5 kDa and the 3K is referring to the poly(aspartic acid) of having a molecular weight of about 3 kDa, which is on average 20 aspartic acid residues. This is consistent with the other descriptions of the tags. In some embodiments, the poly(amino acid) may include other sequences with alternating neutral and negatively charged amino acids (i.e. (GD)20, (GSD)30). The negatively charged amino acid residues are advantageous for capture by a nanopore.
[0079] In some embodiments, the tag is tethered to solid phase beads via a cleavable linkage, such as disulfide. After the proteins are conjugated with the tag, the solid phase substrate enables a simple extraction by washing off all other materials. Thereafter, the- 24 -4242682. vlDocket No. 5200.2432001 (INV-25057) tethered protein-tag complex can be eluted with breaking the cleavable linkage. The eluent only contains tagged proteins and free tags and is easy to separate.Purification of Proteins
[0080] Examples of protein purification include desalting, ultrafiltration, dialysis, precipitation. Use of denaturants may render certain combinations of purification nonfunctional. An example of functional and nonfunctional combinations with GdmCl as the denaturant can be seen in FIG. 19.
[0081] Considering the MW of common proteins, desalting resins with a MWCO of 5~10kDa is preferrable. Higher MWCO can be used if the user knows the protein of interest has a high MW. Spin desalting columns or chromatographic (including variants using gravity flow, syringe, or FPLC system, etc.) desalting columns can be used, and spin columns are preferred because it does not significantly dilute the sample. Examples of such spin desalting columns include Zeba™ Spin Desalting Columns, 7K MWCO (Thermo Scientific™ 89882), Micro Bio-Spin™ P-6 Gel Columns (Bio-Rad, 7326200), etc. It was observed that for removal of free linkers, a single desalting using such columns is not sufficient, so desalting is repeated several (3-5) times.
[0082] Ultrafiltration is conducted with spin ultrafilters (also known as protein concentrators) or lateral flow ultrafilters. A MWCO of 5~10kDa is preferrable. Higher MWCO can be used if the user knows the protein of interest has a high MW. The purification process involves repeated cycles of diluting the sample with the target buffer, concentrating it using the ultrafilter, and refilling it with the target buffer again. The initial dilution is recommended, so that the concentration of proteins after concentrating does not exceed the solubility, and the concentration of organic solvents, if used, does not exceed the tolerance of the ultrafilter to cause damage. For example, an Amicon-4 Ultra lOkDa unit is used to handle samples of 100 ~ 500 pL. The sample is diluted with target buffer to a volume of 4mL and then spun down to 250 pL and filled with the target buffer again to 4mL. The cycle is repeated 4 times to ensure complete buffer exchange. If the target buffer contains detergent other than CHAPS or CHAPSO, the detergent will be retained in the ultrafilter and accumulated as the cycles repeat, yielding concentrated detergent solution. Such usage should be avoided.
[0083] Dialysis involves using dialysis membrane soaked in the target buffer and equilibration for overnight. It minimizes protein loss, at the cost of time.- 25 -4242682. vlDocket No. 5200.2432001 (INV-25057)
[0084] Precipitation involves using precipitant such as ethanol, isopropanol, PEG etc. It provides facile and economic separation of proteins from the solution. Chilling the precipitant solution and elongating the precipitation time helps increase the yield of protein that is precipitated; however, it is noted that not all protein will precipitate, as some will remain in solution and is not carried forward. This is suitable for processing large amounts of proteins.Nanopore Detection, Fingerprinting, or Sequencing of Proteins
[0085] In an embodiment, a method for nanopore detection, fingerprinting, or sequencing of proteins includes: containing a fluid in a structure, the structure defining an interior cavity and a barrier dividing the interior cavity into a first chamber and a second chamber, the barrier containing a channel structure defining a nanopore therethrough fluidically coupling the first and second chambers, the nanopore having a flow path bounded by an interior surface of the channel structure; wherein the fluid includes (i) an N-terminally modified protein described herein; applying a voltage gradient between the first chamber and the second chamber via the flow path, the voltage gradient inducing an ionic flow through the flow path of the nanopore; observing an electronic signature produced by interactions between ions comprising the ionic flow and the linear series of amino acid residues of the N- terminally modified protein passing through the flow path of the nanopore altering the ionic current through the flow path of the nanopore; and determining a property of the N-terminally modified protein based on the observed electronic signature.
[0086] In some embodiments, the N-terminally modified protein passing through the flow path of the nanopore is subject to an electroosmotic force in the flow path of the nanopore induced by the nanopore
[0087] In some other embodiments, the property of the N-terminally modified protein is: an identity of the N-terminally modified protein, detection of the N-terminally modified protein, an amino acid residue sequence of the N-terminally modified protein, direction of movement through the nanopore of the N-terminally modified protein, or any combination thereof. In some other embodiments, the nanopore is alpha-hemolysin (aHL), Bacillus cereus Cytotoxin K, Mycobacterium smegmatis porin A (MspA) or mutants thereof.
[0088] In some instances, when the N-terminally modified protein has the tag of the negatively charged X segment, there is an increase in the driving force by the electric field which can enhance the capture rate or facilitate translocation of proteins that otherwise have lower capture rate or hindered translocation through a nanopore. In some other instances,- 26 -4242682. vlDocket No. 5200.2432001 (INV-25057) when the N-terminally modified protein has the tag of the J segment, the protein can translocate through an engineered nanopore that can specifically bind to the sequence in J. In some other instances, when N-terminally modified protein has the tag of the A segment exhibits a characteristic trace of ionic flow blockade during translocation of the N-terminally modified protein and can reveal extra information about the protein translocation. The characteristic trace refers to a pattern that shows as a duration of a specific blockade level, characteristic spikes or steps, or a sequence of current compared to an unmodified protein of the same sequence. The N-terminally modified protein is designed to show signal characteristic traces that are distinct and easy to detect. Further details describing the general method or nanopore detection, fingerprinting, or sequencing of proteins is described in 11,994,508, which is incorporated by reference herein.Example method for N-terminal tagging of proteins Step 1: Preprocessing of Sample
[0089] In some embodiments, Step 1 includes one or more of the following: (a) adding denaturants and heat treatment to enhance denaturation; (b) adding metal ion chelators followed by removing small molecules by protein purification; (c) adding reducing agents to cleave disulfide bonds, followed by alkylating resultant free thiols and removing the alkylating reagents by protein purification.
[0090] (a) Taking into account whether the sample already has denaturant in it, denaturation can be achieved variously. For protein sample preparation nowadays, detergents like SDS are commonly used. Supplementing higher concentration of the same detergent or adding other denaturants like chaotropes and organic solvents can enhance the denaturation. In some preferred embodiments, the detergent concentration is kept to a minimum concentration. In other preferred embodiments, the detergent used is CHAPS or CHAPSO. It was observed that higher concentration of detergent results in more residual detergent in the final product; however, the use of CHAPS or CHAPSO can be removed more easily via purification than other detergents. Additional detergent may result in complications with protein purification. If the sample contains detergent, the detergent can be removed by purification methods such as precipitation, or detergent removal resins (e.g. Thermo Scientific™ 87780) during the process before an incompatible step occurs.
[0091] (b) Quite commonly found in nature are proteins complexed with metal ions and small molecules, e.g., hemoglobin contains heme molecules and iron ions chelated in the- 27 -4242682. vlDocket No. 5200.2432001 (INV-25057) heme. These small molecules, and metal ions, may interfere with further chemistry. To avoid this, removal of small molecules and metal ions is performed. Such removal is synergized by the denaturation of proteins in substep (a), which breaks the conformation of protein-small molecule complexes and releases the small molecules into solution. Metal chelator (e.g., EDTA, EGTA, dimercaprol) is added to the sample at high concentrations (e.g., > 0.1 mM EDTA) and incubated for a period of time ( > 10 min). The temperature of the incubation may be adjusted between r.t. to 100°C to enhance dissociation of small molecules by higher temperature. The chelated metal ions and small molecules in the solution are removed by protein purification methods.
[0092] (c) Disulfide bonds are extensively found in eukaryotic proteins, and cleavage of them can further loosen the conformation of proteins, enhancing unfolding. Reducing reagents (e.g. TCEP, DTT, 2-ME, cysteine, cysteamine) for cleavage of disulfide bond is added to the solution at high concentrations (e.g. > 1 mM TCEP) and the solution is incubated for a period of time (15 min - 4 hrs). The cleaved disulfide bonds result in free thiols which are prone to oxidation and electrophilic attack, so alkylation converts them into alkyl sulfides. An alkylation reagent (e.g. IAA, AA, IAM, N-EM, 4-VP) is added to the solution at high concentration (e.g. > 5 mM AA) and the solution incubated for a period of time (15 min - 4 hrs) in dark. Optionally to minimize over-alkylation on nucleophiles other than the thiol moiety, after the incubation with alkylation reagent, excessive thiol-containing quenching reagent (e.g. DTT, 2-ME, cysteine, cysteamine) is added to quench remaining alkylation reagent; the alkylation reagent can also be removed directly by protein purification without quenching. The concentrations of reducing reagent, alkylation reagent, and the optional quenching reagent should be coordinated so that the remaining reducing reagent does not inhibit the alkylation, and the quenching reagent can thoroughly quench the alkylation reaction, e.g. (1 eq : 1.5 eq : 2 eq), or (20 mM, 34 mM, 40 mM).
[0093] The buffer in this step can be the buffer in which the sample was acquired, or any other common buffers for biological purposes. However, the product of this step should have a buffer that contains no amino moiety, to avoid interference with the 2-PCA chemistry.Step 2: 2-PCA Linker Conjugation
[0094] The linker includes: (a) a pyridine-2-carbaldehyde (2-PCA) moiety which selectively conjugates with the N-terminus of a protein molecule; and (b) an alkyne (terminal- 28 -4242682. vlDocket No. 5200.2432001 (INV-25057) alkyne or cycloalkyne) for CuAAC conjugation to the tag. The alkyne can directly attach to the pyridine ring of the 2-PCA moiety or attach via a spacer e.g. PEG.
[0095] The 2-PCA moiety is reported to react with the N-terminus in three ways: (a) The carbaldehyde of 2-PCA forms an imidazolidinone ring with the N-terminus amine and the first peptide bond on the peptide chain; (b) with the presence of certain transition metal salts as catalysts, e.g. Cu(II) salts, the carbaldehyde of 2-PCA forms a carbon-carbon bond with the first a-carbon on the peptide chain; and (c) in the presence of Cu(II) and a 1,3- dipolarophile, e.g. maleimide, the 2-PCA participates in a cycloaddition at the N-terminus together with the dipolarophile. See Hanaya et al. Angew. Chem. Int. Ed. 2025, 64, e202417134 for further details. Either reaction can be used in step (ii) depending on the condition of the reaction. In case (c), a functionalized maleimide can be optionally used as the linker, where the click chemistry handle is on the maleimide derivative instead of the 2- PCA. Non-functionalized maleimide derivatives such as NEM can also be used.
[0096] In some embodiments of the reaction (a), Step 2 includes: adding a linker, reacting the linker protein together, and, optionally, purifying the protein to remove any free linker.
[0097] The optional purification step is included so that the free linker does not compete with conjugated proteins during the following tagging reaction and does not interfere with the reaction if Step 4 uses CuAAC. If excessive amount of tag is added, and Step 4 uses SPAAC, such removal of linker is optional.
[0098] The concentration of the 2-PCA-containing linker is usually above 0.1 mM, typically above 1 mM, where a low concentration lowers the efficiency of the reaction. The temperature of incubation is selected between room temperature (r.t) to 70°C. For typical usage of linker concentration above 10 mM, depending on the specific linker selected, solubility of the linker may be exceeded. Organic solvents are added in this case to dissolve the linker in addition to serving as protein denaturant. Incubation time is at least 1 hr. A longer incubation time helps the reaction to proceed thoroughly. An incubation time of 24 hrs to 96 hrs is typical. In some embodiments, it is possible to have incubation time of 26 days, and no adverse effect were observed.
[0099] In some embodiments, Step 2 includes: adding Cu(II) salt; adding a linker, reacting the mixture of the Cu(II) salt, linker, and protein together, quenching the reaction by adding a chelator; and, optionally, purifying the reaction to remove the linker, chelator and Cu(II).- 29 -4242682. vlDocketNo. 5200.2432001 (INV-25057)
[0100] The Cu(II) salt is typically Cu(OAc)2. The concentration is typically between 10 pM to 5 mM. Preparing Cu(OAc)2 by mixing CuSC and NaOAc may result in more active species than using commercial crystalline Cu(OAc)2. Attachment of extra linkers to the proteins is observed with high concentration of Cu(II) (FIG 13). Typical concentration is 200pM. The incubation time is selected between 15 min to 24 hrs. Attachment of extra linkers is also observed for long period of incubation. The typical period is 4 hrs.
[0101] Cu(II) salt catalysis provides some advantages to the reactivity. In reaction conditions that do not use a Cu(II) catalyst, the second nitrogen atom from the N-terminus of the protein peptide backbone is used to form an imidazolidinone ring. When the second residue of a protein is a proline (“P2” proteins), this nitrogen is tertiary, thus hindering the formation of the imidazolidinone ring. The Cu(II) catalysis mechanism utilizes the first a- carbon of the peptide chain, thus not hindered by such prolines. In one of the examples, streptavidin which has proline as its second residue could not afford product with the original route, but afforded product with Cu(II) catalysis (FIGs. 6 and 7).
[0102] Quenching by chelation is preferential because it was observed that without the chelation, further step is inhibited if 5-E-2-PCA and CuAAC is used. It is assumed that the ethynyl moiety is hindered when Cu(II) forms complex with the linker and the protein chain. The chelator is typically EDTA, at a concentration of 2 mM.
[0103] It was observed that DBCO moiety under certain conditions containing Cu(II), contributes to an unexpected band on the SDS-PAGE (FIGs. 10, 11, and 12). It is assumed that the stained alkyne structure of the DBCO could contribute to some addition reaction with protein side chains like cysteine.
[0104] All the purification methods in the purification section can be used for the purifications in this step. When ultrafiltration is used, if organic solvents have been added to the reaction, extensive dilution of the sample should be conducted before loading the sample to the ultrafilter, to avoid damage of ultrafiltration membrane by the organic solvent.
[0105] The buffer for this step should contain no amino moiety. A pH of 4 ~ 8 is used for the efficiency of the 2-PCA condensation with amine. Typically, the pH is 7.5 and the buffer is phosphate buffer or PBS, or a buffer without a primary or secondary amine (e.g. HEPES).
[0106] In some embodiments, a solid-phase tethered tag-PCA complex is used in place of the linker, and the desired protein-linker-tag product is formed while tethered to the solid phase substrate. In such cases, steps 3 and 4 are skipped, and the protocol directly proceeds to step 5.- 30 -4242682. vlDocket No. 5200.2432001 (INV-25057)Step 3: Optional Blocking of Lysine Side Chains
[0107] In some embodiments, Step 3 to block lysine sidechain amine moiety includes: adding amine-specific small molecules (<lkDa), e.g. NHS esters (NHS acetate, NHS dimethylglycinate, NHS nicotinate, etc.), pyrocarbonates (Boc anhydride, etc.); reacting the amine-specific small molecules with the protein at a time scale of hours; and purifying protein to remove small molecules.
[0108] This process converts lysine sidechain primary amines into less reactive moieties to avoid non-uniform carbamylation on these amines from degradation products from urea and urea derivatives, e.g. cyanate or methyl cyanate, when urea or DMU is used as the denaturant in one of the subsequent steps. Such a process may eliminate positive charges from the lysine sidechains depending on the blocking moiety and the pH and will compromise Step 5 the extraction of tagged protein if the extraction uses an anion exchange chromatography.
[0109] Optionally, to mitigate such alteration of charge, an additional process to block glutamic acid and aspartic acid sidechain carboxyl moiety can be appended to the above process: adding EDC / NHS to activate the carboxyl moiety; adding amine-containing smallmolecules (<lkDa), e.g. ethanolamine; reacting the mixture together at a time scale of hours; and purifying protein to remove small molecules.
[0110] With the additional process of blocking carboxyl moieties, the protein purification can be omitted, where the residual amine-reactive reagents will be depleted by excessive amounts of amine-containing reagents.Step 4: Click Reaction for Tag Conjugation
[0111] In some embodiments, Step 4 to add the tag to the protein analyte includes: adding azide-functionalized tag of choice and CuAAC catalyst mixture (CuSC>4, sodium ascorbate, THPTA), followed by reacting the protein and tag for hours. If DBCO linker is used, no CuAAC catalyst is added.
[0112] Depending on the form of the tag preparation, the addition of tag may be performed by mixing liquid tag stock or solid phase beads with the protein analyte solution, or in some embodiments, if packed solid phase beads or solid phase resin bed is used, the addition should be performed by absorbing the protein analyte solution to the solid phase.- 31 -4242682. vlDocket No. 5200.2432001 (INV-25057)Step 5: Extraction of Tagged Protein
[0113] The extraction Step 5 can be combined from several options based on the previous steps and the purpose of tagging. In some embodiments a liquid tag stock may be used for Step 4, then methods of the extraction can include: SEC to separate tagged product from the untagged protein and free tags in protein samples separable by molecular weight (e.g. lysozyme); SCX to remove free tag, then SAX to separate tagged product from untagged protein; SCX to remove the free tag, then affinity column corresponding to the affinity tag attached to the tagged protein; or separation by gel electrophoresis and the band corresponding to the tagged product is cut from the gel and recovered. In some other embodiments, a solid phase tag may be used for Step 4, then the method of extraction can include: washing the solid phase several times and eluting with buffer containing cleaving reagents so tagged product and free tag are eluted, followed by SCX to remove the free tag.
[0114] The SCX process effectively separates proteins (tagged and untagged), from free tags. It uses a substantially low pH (-2) to ensure protein positively charged, while the tag part is uncharged or slightly negatively charged due to the pKa (~3.5) of monomers being higher than the pH. The tagged proteins thus are also positively charged, summing up the charge from the protein part and the tag part. Thereby, positively charged species (untagged proteins and tagged proteins) are separated from neutral or slightly negatively charged species (free tags).
[0115] The SAX process separates tagged and untagged proteins. A pH (3.5-5.5) is used below the pl of common proteins but above the pKa of the monomers of the tag. Thus, proteins without tag neutral or slightly charged, and tags are strongly negatively charged, and tagged proteins are also negatively charged. Thereby, tagged proteins are separated from untagged proteins.
[0116] The SEC process separates tagged proteins, untagged proteins and free tags by size. This is most effective when the size of the protein is similar to, or smaller than the size of the tag, to afford significant relative difference in size. For example, lysozyme has a MW of around 14kDa, so separation with a PEG-D20 tag with a MW of 8 kDa can be achieved.
[0117] Gel excision can also be used. The gel is cut according to a reference position by a test run of the same sample, smashed, and soaked into buffer to allow diffusion of the protein. However, this method is limited in throughput.- 32 -4242682. vlDocket No. 5200.2432001 (INV-25057)Step 6: Post processing
[0118] In some embodiments, step 6 of post processing is performed to reduce disulfide bonds, which is performed through reacting a reducing agent for hours prior to testing by nanopore translocation. In non-limiting example, TCEP-HC1 is added to the product (final concentration 1 mM - 100 mM), and 1 equivalent of NaOH is also added to neutralize the acidity of TCEP-HC1. The product is incubated (15 mins - 16 hours). Optionally a buffer exchange by a protein purification method defined above can be conducted, to bring the product into the sample chamber buffer of the nanopore setup to be used. The product can also be adjusted by the addition of solvent (water) and / or corresponding solutes to match the sample chamber buffer. The product is then injected into the sample chamber of a nanopore setup. In some embodiments, there is not a limit to the selection of the sample chamber buffer, which is determined by the user based on the actual nanopore experiment of interest. An example of such buffer is (1 M KC1, 2 M GdmCl, 2 mM TCEP, 10 mM Tris, pH 7.5).EXEMPLIFICATIONExample 1: Synthesis of Tagged Protein, IEX purificationStep 1 (Preprocessing of sample) and Step 2 (2-PCA linker conjugation)
[0119] Commercial ConA (MPBio 02195283-CF) was used as the model protein. The protein was dissolved in water (lOmg / mL) as the sample. The sample was denatured by addition of 8M GdmCl stock solution, and 10X PBS, resulting in Img / mL ConA, 6M GdmCl in PBS buffer. A reaction mixture was prepared by addition of 5-E-2-PCA and Cu(OAc)2 to a final concentration of 20 mM 5-E-2-PCA, 0.2 mM Cu(OAc)2, where the Cu(OAc)2 was prepared by mixing CuSC>4 and NaOAc. DMSO was added at a volume ratio of 3 : 10. The mixture was incubated at 37°C for 5 hrs. EDTA-2Na (final cone. 2 mM) was added to the mixture and incubated for 15 min. It is noted that Step 1 and Step 2 were performed in one step as reducing disulfide bonds and alkylation was not needed for ConA. This can be done with proteins that do not require reducing disulfide bonds and alkylating.
[0120] Each ImL of the mixture was then mixed with 3mL of (PBS, 6M GdmCl) and loaded to centrifugal ultrafilter (Amicon Ultra-4 lOkDa MWCO) to spin down to less than 500uL. After each spin, the sample was refilled to 4mL with (PBS, 6M GdmCl) and mixed with pipette. After 3 repeats, the sample was buffer exchanged 2 times to (10 mM HEPES, 8M DMU) using spin desalting column (Zeba™ Spin Desalting Columns, 7K MWCO, 2 mL).- 33 -4242682. vlDocketNo. 5200.2432001 (INV-25057)Step 3 (Blocking lysine side chains)
[0121] 10 mM NHS acetate was added. Sample was then incubated for 2 hrs at 37°C, followed by desalting using spin desalting column.Step 4 (Click reaction for tag conjugation)
[0122] N3-PEG-D20 tag (Poly(L-Aspartic acid)-PEG- Azide, PEG 5kDa, pAsp 3kDa, Nanosoft polymers 9150-3000-5000-100mg) was added to the sample (final cone. 1 mM). CuSC>4 and THPTA were mixed separately immediately before use, to give (10 mM:50 mM) chelated copper stock. CuSC +TElPTA and sodium ascorbate was added to the sample (final cone. 0.5 mM:2.5 mM:5 mM). The sample was incubated at 37°C for 5 hrs.Step 5 (Extraction of tagged protein)
[0123] SCX binding solution (8M DMU, 200 mM citric acid, 2 mM EDTA-2Na) was added at a volume ratio of 10: 1 to acidify the sample. The sample was loaded to SCX spin column (Pierce™ Strong Cation Exchange Spin Columns, Mini) by 3 portions and washed with binding solution 5 times. The sample was then eluted with 250pL elution buffer (10 mM HEPES, 8M Urea, 2M GdmCl). The sample as the eluent was buffer exchanged to SAX binding buffer (10 mM DMP, 8M Urea, pH 5) (FIGs. 1, 2, and 3).
[0124] The sample was loaded to an SAX column (Bio-Rad EconoFit Macro-Prep High Q ImL) on a FPLC system (AKTA pure), and washed with at least 5CV of binding buffer, and eluted with SAX elution buffer (10 mM DMP, 1.5M GdmCl, 8M Urea, pH5) (FIGs. 23, 25C, and 25G). The peaks were collected and concentrated to 100~150pL / fraction by ultrafilter (Amicon Ultra-4) (FIGs. 1, 2, 4, 5, 8, and 9).Step 6 (Post processing)
[0125] Before nanopore experiment, the sample in the SAX elution buffer was buffer exchanged with spin desalting column (Zeba 7kDa, 0.5mL) to nanopore chamber buffer (IM KC1, 2M GdmCl, 10 mM Tris, pH7.5). The nanopore protein aHL was mixed with the sample and loaded into cis chamber. Once the aHL inserted into the DPhPC membrane formed on a custom aperture made of SU-8, the current was recorded for analysis (FIGs. 24, 25 D-F, and 25H-J).- 34 -4242682. vlDocket No. 5200.2432001 (INV-25057)SDS-PAGE
[0126] For SDS-PAGE assays of the sample at different stages, when the sample contained GdmCl, a buffer exchange to (PBS, 2% CHAPS) using miniaturized spin columns (Zeba 7kDa, 75uL) was performed to avoid precipitation of SDS with GdmCl. Sample was then mixed with loading buffer and heated at 60°C, for Ihr, loaded on 12% tris-glycine PAGE gel in SDS-PAGE running buffer, and run at 50V for lOmin, then 150V for 60min. The gel was stained with SYPRO Ruby.Variants
[0127] Multiple experiments according to the above protocol were performed, some of which used variant conditions including:- 35 -4242682. vlDocket No. 5200.2432001 (INV-25057)Example 2: Synthesis of Tagged Protein , Lysozyme as model protein, SEC purification
[0128] Commercial lysozyme was used as the model protein. A reaction mixture was prepared in PBS with Img / mL lysozyme, 10 mM PCA-PEG4-DBCO, and 6M GdmCl. The mixture was incubated at 37°C for 4 days.
[0129] Each ImL of the mixture was then mixed with 3mL of (PBS, 6M GdmCl) and loaded to centrifugal ultrafilter (Amicon Ultra-4 lOkDa MWCO) to spin down to less than 500uL. After each spin, the sample was refilled to 4mL with (PBS, 6M GdmCl) and mixed with pipette. The sample was washed 4 times.
[0130] N3-PEG-D20 tag (Poly(L-Aspartic acid)-PEG- Azide, PEG 5kDa, pAsp 3kDa, Nanosoft polymers 9150-3000-5000-100mg) was added to the sample (final cone. 1 mM). Sample was incubated overnight.
[0131] The sample was then loaded into SEC chromatographic column (Cytiva Superdex 75 300 / 10 GL) and separated at a flow rate of 0.5ml / min, in (PBS, 6 M GdmCl) buffer (FIG.22).
[0132] Before nanopore experiment, the sample was diluted into 2 M GdmCl concentration, and TCEP-HC1 (20 mM final) was added with equivalent amount of NaOH (20 mM final). The sample was incubated at 37°C for 40 mins and loaded to nanopore experiment.
[0133] SDS-PAGE was done as described in Example 1. Schematics of major reactions and tags used are shown in FIGs. 25 A and 25B.Variants
[0134] Multiple experiments according to the above protocol were performed, some of which used variant conditions including:4242682. vlDocket No. 5200.2432001 (INV-25057)Example 3: Synthesis of Tagged Protein, Revised reaction route and study of translocationStep 1 (Preprocessing of sample) and Step 2 (2-PCA linker conjugation)
[0135] Commercial ConA (MPBio 02195283-CF) and beta Casein (Sigma C6905) were used as the model protein. The protein was dissolved in a buffer of (6M GdmCl, 40mM HEPES, pH7.5) to final concentration of lOmg / mL. A reaction mixture was prepared by addition of 5-E-2-PCA, NEM and Cu(OAc)2 and the cosolvent NMP to a final concentration of 5mg / mL protein, 25 mM 5-E-2-PCA, 20 mM NEM, 2 mM Cu(OAc)2, 50% Vol NMP, where the Cu(OAc)2was acquired commercially in crystalline form. The mixture was incubated at 37°C for overnight (16hrs). EDTA-2Na (final cone. 2 mM) was added to the mixture and incubated for 30 min. It is noted that Step 1 and Step 2 were performed in one step as reducing disulfide bonds and alkylation was not needed for ConA or beta Casein. This can be done with proteins that do not require reducing disulfide bonds and alkylating.
[0136] Each 500uL of the mixture was buffer exchanged into (6M GdmCl, 40mM HEPES, pH 7.5) using spin desalting columns (Zeba™ Spin Desalting Columns, 7K MWCO, 2 mL), and then buffer exchanged again into a buffer of (6M GdmCl, lOOmM MeONH2- HCl / 50mM NaOH), to accelerate the cleavage of Schiff base intermediate form by the lysines and the PCA linker. After a brief incubation at 37°C for 3hrs, the mixture was buffer exchanged again into (6M GdmCl, 40mM HEPES, pH 7.5) and stored at 4°C, and was used up within 1 month.Step 3 (Blocking lysine side chains)
[0137] This step was omitted for this example as the subsequent separation did not use ion exchange methods.- 37 -4242682. vlDocket No. 5200.2432001 (INV-25057)Step 4 (Click reaction for tag conjugation)
[0138] The product of the upstream process was buffer exchanged into a buffer of (8M DMU, 20mM HEPES, pH 7.5) using spin columns (Zeba 7kDa, 0.5mL). N3-PEG-D20 tag (Poly(L-Aspartic acid)-PEG- Azide, PEG 5kDa, pAsp 3kDa, Nanosoft polymers 9150-3000- 5000-100mg) was added to the sample (final cone. 0.5 mM). Q1SO4 and THPTA were mixed separately immediately before use, to give (10 mM:50 mM) chelated copper stock. CuSC +THPTA and sodium ascorbate was added to the sample (final cone. 0.5 mM:2.5 mM: 10 mM). The sample was incubated at 37°C for overnight (16hrs).Step 5 (Extraction of tagged protein)
[0139] For tagged protein was extracted with SDS-PAGE with electroelution. The sample was split into portions, diluted with 1 : 1 volume of water / loading buffer mixture, and loaded to SDS-PAGE gel (Novex TG WedgeWell 12%), and run at 50V for lOmin, then 150V for 60min. The gel was briefly stained with Coomassie blue and destained. The band corresponding to the tagged protein was excised from the gel, then loaded into dialysis membrane (MWCO lOkDa) with Tris-glycine-SDS running buffer. The apparatus was sunk into Tris-glycine-SDS running buffer in a horizontal electrophoretic tank. Voltage (-150V for ~15cm horizontal tanks, or as high as possible depending on Joule heating) was applied to drive electroelution of the protein out of the gel, until the blue color completely exited from the gel. The liquid content of the dialysis apparatus was then collected and concentrated using ultrafilter (Amicon Ultra 4 MWCO lOkDa). SDS-PAGE of the extracted sample showed clear and pure tagged product (FIG. 28).Step 6 (Post processing and nanopore testing)
[0140] Before nanopore experiment, the sample in the SDS running buffer was buffer exchanged with spin desalting column (Zeba 7kDa, 0.5mL) to 8M DMU, 20 mM HEPES pH 7.5 and then buffer exchanged again into 6M GdmCl, lOmM HEPES, pH7.5.The sample was then tested on a nanopore setup. A PTFE thin film was perforated by electric arc to form an aperture of - 50 pm diameter and clamped between two chambers. The chambers were filled with 3M GdmCl, 0.5M KC1, lOmM Tris, ImM EDTA, pH 8.5, with lipid and Ag / AgCl electrodes added to form a Montal -Mueller setup [Montal and Mueller, Proc. Nat. Acad. Sci. Vol. 69, No. 12, pp. 3561-3566], where the lipid formed a bilayer in the- 38 -4242682. vlDocketNo. 5200.2432001 (INV-25057) aperture between the two chambers and the electrodes were connected to a patch clamp amplifier to apply desired voltages between the two chambers while monitoring the ionic current across the bilayer. The current was digitized and recorded at a sample rate of 250kHz with a PC equipped with a data acquisition card. The nanopore (Bacillus cereus Cytotoxin K o Mycobacterium smegmatis porin A mutant M2 (MspA M2) ) was added to the chamber. After an insertion of the nanopore onto the lipid bilayer, shown as an onset of ionic current, the sample was added. Due to diffusion and driving forces from the nanopore including electrophoretic force and electroosmotic flow, analyte molecules in the sample were captured by and translocated through the nanopore, shown as events of blockade of ionic current. Such events carry information about the analyte and are the data of interest for nanopore research. A clip of such current recording data is exemplified in FIG. 29.Translocation study
[0141] As seen in FIG. 29, there were events of fast, shallow spikes as well as patterned, deep and long dwells. The former were capture without translocation, while the latter was capture and translocation. The latter contained a shallow part and a deep part, which corresponded to the PEG segment of the tag and the protein, respectively. To gain statistical information about the events, the events were extracted from the full current trace by thresholding. For events with two parts, each event was segmented into multiple steps based on the edges of steep blockade changes within. The first step was the prominent level of the first part of the event, while the other steps were ascribed to undulation of the current during the translocation of the protein, i.e. the second part (FIG. 30A). The first and second parts were measured separately to give statistics on a scatter plot (FIG. 30B). While the software of this analysis is available online at an open-source database, it should be noted that this analysis could be achieved with different algorithms given the distinct level of the tag and is not limited to a single algorithm. Separate statistics showed a clearly defined shallow step corresponding to the PEG segment in the tag, with a deep step corresponding to the protein segment Such phenomenon supported that the translocation was single-file translocation, and the tag was successfully added to the terminus of the protein (FIGs. 29, 31 A-F, 32A, and 32B) (FIG. 29).
[0142] While the tag initiated the translocation in most cases, in certain cases, the tag could also be used to show the direction of translocation when translocation from both termini is possible, which was experimentally difficult to determine for each individual event- 39 -4242682. vlDocket No. 5200.2432001 (INV-25057) prior to this disclosure (e.g. [Yu et al. Nature Biotechnology, 41, 1130-1139 (2023)]). For example, when the tagged protein was applied to MspA M2 pore, under certain conditions, protein could also initiate the translocation from the C-terminus (FIGs. 33A-D). This phenomenon could be overlooked without the characteristic level of the tag clearly indicating the direction of translocation, demonstrating the value of tagging in not only facilitating the translocation but also revealing previously undisclosed behavior of translocation.
[0143] The translocation study further indicates that the characteristic current trace facilitates analysis can be used to study capture behavior, translocation dynamics, calibration of conductivity, probing the status of a nanopore, barcoding for sample multiplexing, identification of translocation direction or any combination thereof.
[0144] Translocation testing or studies can be used for the detection, fingerprinting, or sequencing of proteins based on the blockade of ionic flow as the protein translocates through the nanopore, where the ionic flow blockade is read out by means used in conventional nanopore experiments, including but not limited to electrochemical measurement of the electric current carried by the ionic flow, or optical measurement of the fluorescence of a ionsensitive fluorophore (e.g. Fluo-3).INCORPORATION BY REFERENCES; EQUIVALENTS
[0145] The teachings of all patents, published applications and references cited herein are incorporated by reference in their entirety.
[0146] While example embodiments have been particularly shown and described, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the scope of the embodiments encompassed by the appended claims.4242682. vl
Claims
Docket No. 5200.2432001 (INV-25057)CLAIMSWhat is claimed is:
1. A method of tagging a protein, the method comprising: a) optionally denaturing a protein; b) reacting a linker and the protein to form a linker-functionalized protein,wherein Y is = or , n is an integer from 1 to 12, and m is an integer from 1 to 6; optionally denaturing the linker-functionalized protein, wherein at least step a) or step c) is performed prior to step d); and- 41 -4242682. vlDocket No. 5200.2432001 (INV-25057) d) reacting the linker-functionalized protein with a tag, wherein the tag compriseswherein a is an integer from 1 to 100, X is a canonical poly amino acid repeating in x residues, wherein the canonical poly amino acid residue comprises negatively charged amino acids or a combination of negatively charged amino acid and neutral amino acids, wherein x is an integer from 1 to 200, J is an amino acid comprising SEQ ID NO: 1, and E is a solid phase bead2. The method of claim 1, further comprising reacting the linker-functionalized protein with an amine specific protecting compound of a molecular weight less than 1 kDa after step c) and before step d).
3. The method of claim 2, further comprising reducing the linker-functionalized protein and / or reacting the linker-functionalized protein with an alkylating agent before reacting the linker functionalized protein with the amine specific protecting compound.
4. The method of claim 3, wherein step d) further comprises reacting a Cu(II) salt and quenching.
5. The method of claim 1, wherein A ishas a molecular weight greater than IkDa.
6. The method of claim 2, wherein the amine specific protecting compound comprises an N-hydroxysuccinimide (NHS) ester, a pyrocarbonate, or a carboxylic anhydride.
7. The method of claim 1, wherein steps a) and b) are performed at the same time.- 42 -4242682. vlDocket No. 5200.2432001 (INV-25057)9. The method of claim 1, wherein the tag is Ns-poly ethylene glycol (PEG)5K- poly(aspartic acid)3K, N3-poly(glutamic acid) 1 OK, or Ns-PEGlOK.
10. The method of claim 9, wherein the tag further comprises SEQ ID NO: 1.
11. The method of claim 1, wherein the denaturant has a concentration from about 0.1% to about 96%.
12. An N-terminally modified protein comprising the following formula:- 43 -4242682. vlDocketNo. 5200.2432001 (INV-25057)wherein n is an integer from 1 to 12, m is an integer from 1 to 6, R is a canonical amino acid side chain of the protein,- 44 -4242682. vlDocket No. 5200.2432001 (INV-25057)wherein A is, wherein a is an integer from 1 to 100, X is a canonical poly amino acid repeating in x residues, wherein x is an integer from 1 to 200, J is an amino acid sequence comprising SEQ ID NO: 1, and E is a solid phase bead.
13. The protein of claim 12, wherein A comprisesand has a molecular weight greater than IkDa.
14. A kit for chemically modifying proteins to make a tagged protein, the kit comprising:a denaturant for a protein;4242682. vlDocket No. 5200.2432001 (INV-25057)m is an integer from 1 to ; and iii) a tag, wherein the tag compriseswherein A is, wherein a is an integer from 1 to100, X is a canonical poly amino acid repeating in x residues, wherein x is an integer from 1 to 200, J is an amino acid sequence comprising SEQ ID NO: 1, and E is a solid phase bead.
15. The kit of claim 14, further comprising a Cu(II) salt, sodium ascorbate, and a chelator.
16. The kit of claim 14, further comprising an N-Hydroxysuccinimide (NHS) ester of a molecular weight less than 1 kDa or a pyrocarbonate of a molecular weight less than 1 kDa.- 46 -4242682. vlDocket No. 5200.2432001 (INV-25057)17. The kit of claim 14, wherein the linker comprises18. The kit of claim 14, wherein the tag comprises Ns-polyethylene glycol (PEG)5K- poly(aspartic acid)3K (PEG-D20 tag), N3-poly(glutamic acid)10K (E76 tag), or N3- PEG10K (PEG10K tag).
19. The kit of claim 14, wherein the tag further comprises SEQ ID NO: 1.
20. The kit of claim 14, wherein the denaturant comprises a chaotrope, a detergent, an organic solvent, an ionic liquid, a deep eutectic solvent, or any combination thereof.
21. The kit of claim 20 wherein the chaotrope comprises a guanidinium salt, a perchlorate salt, urea, a urea derivative, thiourea, a thiourea derivative, imidazolidinone, a bromide salt, a hexafluorophosphorate salt, a tetrafluoroborate salt, a thiocyanate salt, a tetraalkylammonium salt, a lithium salt, or any combination thereof.
22. The kit of claim 20, wherein the detergent comprises sodium dodecyl sulfate (SDS), polyethylene glycol tert-octyl phenyl ether, polyoxyethylene (10) lauryl ether, polyoxyethylene (20) cetyl ether, dodecyulguanidine acetate (dodin), cetyltrimethylammonium salts, Nonyl phenoxypolyethoxylethanol (NP-40), or any combination thereof.
23. The kit of claim 20, wherein the organic solvent comprises dimethyl sulfoxide (DMSO), dimethyl formamide (DMF), N-methylpyrrolidone (NMP), formamide, N- methylformamide, acetamide, N-methylacetamide, dimethylacetamide (DMAc), hexafluoro-2-propanol (HFIP), trifluoroethanol (TFE), acetonitrile, phenol, or any combination thereof.- 47 -4242682. vlDocket No. 5200.2432001 (INV-25057)24. A method for nanopore detection, fingerprinting, or sequencing of proteins, the method comprising: containing a fluid in a structure, the structure defining an interior cavity and a barrier dividing the interior cavity into a first chamber and a second chamber, the barrier containing a channel structure defining a nanopore therethrough fluidically coupling the first and second chambers, the nanopore having a flow path bounded by an interior surface of the channel structure; wherein the fluid includes an N-terminally modified protein of claim 12; applying a voltage gradient between the first chamber and the second chamber via the flow path, the voltage gradient inducing an ionic flow through the flow path of the nanopore; observing an electronic signature produced by interactions between ions comprising the ionic flow and the linear series of amino acid residues of the N- terminally modified protein passing through the flow path of the nanopore altering the ionic current through the flow path of the nanopore; and determining a property of the N-terminally modified protein based on the observed electronic signature.
25. The method of claim 24, wherein the N-terminally modified protein passing through the flow path of the nanopore is subject to an electroosmotic force in the flow path of the nanopore induced by the nanopore26. The method of claim 24, wherein the property of the N-terminally modified protein is: an identity of the N-terminally modified protein, detection of the N-terminally modified protein, an amino acid residue sequence of the N-terminally modified protein, direction of movement through the nanopore of the N-terminally modified protein, or any combination thereof.
27. The method of claim 24, wherein the nanopore is alpha-hemolysin (aHL), Bacillus cereus Cytotoxin K, Mycobacterium smegmatis porin A (MspA) or mutants thereof.- 48 -4242682. vl