Methods of polypeptide sequencing
Patent Information
- Application Number
- PCT/US2025/028808
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-11
- Filing Date
- 2025-05-11
- Publication Date
- 2026-02-19
AI Technical Summary
Existing protein sequencing technologies face challenges in accurately distinguishing amino acids due to background noise and signal fluctuations, particularly in nanopore sequencing, which is exacerbated by the different properties of peptide chains compared to DNA and RNA.
A method and system for sequencing polypeptides involves generating a chimeric nucleic acid chain with amino acids linked to a nucleic acid strand, using blockers to control translocation through a nano-scale recording device, and detecting signals to identify amino acids based on these translocations.
This approach enhances the signal-to-noise ratio and improves amino acid identification resolution by selectively extracting amino acid information, allowing for higher quality data acquisition.
Abstract
Description
METHODS OF POLYPEPTIDE SEQUENCINGSEQUENCE LISTING
[0001] The sequence listing that is contained in the file named “088177-8003W001”, which is 18,506 bytes and was created on May 8, 2025, is filed herewith by electronic submission and is incorporated by reference herein.CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to US provisional application 63 / 645,885, filed May 11, 2024, the disclosure of which is incorporated herein by reference.FIELD OF THE INVENTION
[0003] The present disclosure generally relates to molecular biology and analyte analysis. In particular, the present disclosure relates to methods and systems of sequencing polypeptides.BACKGROUND
[0004] Proteins, the molecular workhorses of living organisms, play vital roles in almost every biological process. Understanding their structure and function is essential for deciphering the complexities of life. Protein sequencing, as a fundamental technique in molecular biology, provides invaluable insights into the composition and arrangement of amino acids (AA) within a protein molecule.
[0005] At its core, protein sequencing involves determining the precise order of A As in a protein chain. This sequential arrangement holds the key to unraveling the protein's structure, function, and evolutionary relationships. By elucidating the AA sequence, people can uncover crucial information about protein folding, interactions with other molecules, and involvement in biological pathways. The applications are across but not limited to various sectors, including new antigens discovery, immune repertoire cataloging, and disease related post-translational modification identification.
[0006] The journey of protein sequencing began several decades ago with pioneering techniques such as Edman degradation and Sanger sequencing. These methods laid the groundwork for modem protein sequencing technologies, which have since evolved to become faster, more accurate, and more versatile. For example, Quantum-Si employs a proprietary semiconductor chip-based platform that utilizes single-molecule fluorescence detection to sequence proteins at the single-molecule level. This approach allows for highly sensitive and quantitative protein analysis. Encodia, Inc., another protein sequencing company, utilizesDNA-encoding strategies combined with next-generation sequencing (NGS) to enable massively parallel protein sequencing. However, both approaches involve the use of specific antibodies for identifying different AAs. The specificity and sensitivity of the antibodies to the 20 AAs become a large challenge, even more to post-translational modification of AAs.
[0007] Nanopore Sequencing, which is primarily used for DNA and RNA sequencing, has been explored for protein sequencing. Nanopores can potentially detect without labeling and characterize individual protein molecules or peptides as they pass through nanoscale pores, offering real-time analysis. Achieving a high signal-to-noise ratio (SNR) is critical for accurate sequence determination, but it can be challenging due to background noise and fluctuations in signal intensity. Meanwhile, the properties of peptide chain are entirely different from DNA and RNA, such as charge and rigidity. Thus, the difficulty of single AA identification in a polypeptide jumps up. Glyphic Biotechnologies Inc. used their special chemistry to dissemble peptide to single AA and conjugate them to a DNA backbone. Then the DNA-AA chimera was sequenced by commercial ONT platform. However, the SNR is not ideal enough to distinguish different AAs. Therefore, there is a need for protein sequencing technologies that can overcome such challenges.SUMMARY OF INVENTION
[0008] The present disclosure in one aspect provides a method of sequencing a polypeptide having a sequence of amino acids.
[0009] In some embodiments, the method of polypeptide sequencing comprises: (a) generating from the polypeptide a chimeric nucleic acid chain comprising (i) a nucleic acid strand comprising a first end and a second end, and (ii) the amino acids dissembled from the polypeptide, wherein the amino acids are linked to the nucleic acid strand sequentially along the nucleic acid strand from the first end to the second end according to the sequence of the amino acids in the polypeptide; (b) subjecting the chimeric nucleic acid chain to a recording device comprising (i) a first compartment, and (ii) a second compartment connected to the first compartment, wherein the nucleic acid strand non-covalently binds to at least a blocker in the first compartment, wherein the blocker prevents an amino acid of the chimeric nucleic acid chain from translocating from the first compartment to the second compartment; (c) disintegrating the blocker from the nucleic acid strand to allow the amino acid to translocate from the first compartment to the second compartment; (d) detecting via the recording device a signal that is generated from translocation of the amino acid from the first compartment tothe second compartment; and (e) identifying the amino acid based on the signal detected in step (d).
[0010] In some embodiments, the method of polypeptide sequencing comprises: (a) generating from the polypeptide a chimeric nucleic acid chain comprising (i) a nucleic acid strand comprising a first end and a second end, and (ii) the amino acids dissembled from the polypeptide, wherein the amino acids are linked to the nucleic acid strand sequentially along the nucleic acid strand from the first end to the second end according to the sequence of the amino acids in the polypeptide; (b) subjecting the chimeric nucleic acid chain to a recording device comprising (i) a first compartment, and (ii) a second compartment connected to the first compartment, wherein the nucleic acid strand comprises a blocker forming a structure that prevents an amino acid of the chimeric nucleic acid chain from translocating from the first compartment to the second compartment; (c) changing the structure of the blocker to allow the first amino acid to translocate from the first compartment to the second compartment; (d) detecting via the recording device a signal that is generated from translocation of the amino acid from the first compartment to the second compartment; and (e) identifying of the amino acid based on the signal detected in step (d).
[0011] In another aspect, the present disclosure provides a system of sequencing a polypeptide comprising a sequence of amino acids.
[0012] In some embodiments, the system for polypeptide sequencing comprises: (a) a module of generating from the polypeptide a chimeric nucleic acid chain comprising (i) a nucleic acid strand comprising a first end and a second end, and (ii) the amino acids dissembled from the polypeptide, wherein the amino acids are linked to the nucleic acid strand sequentially along the nucleic acid strand from the first end to the second end according to the sequence of the amino acids in the polypeptide; (b) a recording device comprising (i) a first compartment, and (ii) a second compartment connected to the first compartment, wherein the nucleic acid strand non-covalently binds to at least a blocker in the first compartment, wherein the blocker prevents an amino acid of the chimeric nucleic acid chain from translocating from the first compartment to the second compartment; (c) a module of disintegrating the blocker from the nucleic acid strand to allow the amino acid to translocate from the first compartment to the second compartment; (d) a module of detecting a signal that is generated from translocation of the amino acid from the first compartment to the second compartment; and (e) a module of identifying the amino acid based on the signal detected in (d).
[0013] In some embodiments, the system for polypeptide sequencing comprises: (a) a module of generating from the polypeptide a chimeric nucleic acid chain comprising (i) anucleic acid strand comprising a first end and a second end, and (ii) the amino acids dissembled from the polypeptide, wherein the amino acids are linked to the nucleic acid strand sequentially along the nucleic acid strand from the first end to the second end according to the sequence of the amino acids in the polypeptide; (b) a recording device comprising (i) a first compartment, and (ii) a second compartment connected to the first compartment, wherein the nucleic acid strand comprises a blocker forming a structure that prevents an amino acid of the chimeric nucleic acid chain from translocating from the first compartment to the second compartment;(c) a module of changing the structure of the blocker to allow the amino acid to translocate from the first compartment to the second compartment; (d) a module of detecting a signal that is generated from translocation of the amino acid from the first compartment to the second compartment; and (e) a module of identifying the amino acid based on the signal detected in(d).BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The following drawings form part of the present specification and are included to further demonstrate certain aspects of the present disclosure. The disclosure may be better understood by reference to one or more of these drawings in combination with the detailed description of specific embodiments presented herein.
[0015] FIG. 1 shows an exemplary embodiment of the chimeric nucleic acid chain described in the present disclosure.
[0016] FIG. 2 shows the schematic of an example workflow of generating a chimeric nucleic acid chain from a polypeptide using an iterative process.
[0017] FIG. 3 shows the schematic of the type I mechanism of continuous and blocker- controlled decoding of amino acids on a chimeric nucleic acid chain by nanopore.
[0018] FIG. 4 shows an exemplary embodiment of blocker design.
[0019] FIG. 5 shows the schematic of the type II mechanism of continuous and blocker- controlled decoding of amino acids on a chimeric nucleic acid chain by nanopore.
[0020] FIG. 6 shows an exemplary embodiment of blocker design.
[0021] FIG. 7 shows an exemplary embodiment of continuous and blocker-controlled decoding of amino acids on a chimeric nucleic acid chain by nanopore via iterative reading.
[0022] FIG. 8 shows an exemplary embodiment of continuous and blocker-controlled decoding of amino acids on a chimeric nucleic acid chain by nanopore via iterative reading.
[0023] FIG. 9 shows an exemplary embodiment of barcoding the chimeric nucleic acid chain.
[0024] FIG. 10A shows an exemplary amino acid base unit (BU).
[0025] FIG. 10B shows the synthesis of an amino acid base unit via click chemistry.
[0026] FIG. 10C shows a typical pulse electric single of an amino acid base unit.
[0027] FIG. 11A shows representative signals of nanopore detection from 5 BUs and their residual current.
[0028] FIG. 11B shows that the profiles of some amino acids (such as R, K, H,W, G) can be identified based on current blockage and duration times.
[0029] FIG. 11C shows that 15 amino acids can be differentiated based on their respective residual currents and duration times.
[0030] FIG. 11D shows amino acids A, G, F, Q displayed different current noise pattern of IRMS despite that these four amino acids showed similar average residual currents.
[0031] FIG. 12A shows an exemplary chimeric nucleic acid comprising three amino acids.
[0032] FIG. 12B shows a diagram illustrating sequential translocation of three amino acids on a chimeric nucleic acid and their detection.
[0033] FIG. 12C shows the current traces of two polymers (E-W-R vs. E-R-W) with the blockers.
[0034] FIG. 12D shows the ONT detection of the trimer polymer (E-W-R) without the blockers.
[0035] FIG. 13 shows the PTMs of the side chain in lysine(K) and their nanopore detection.
[0036] FilG. 14A shows the nanopore rotaxane for multiple reads of two codes.
[0037] FIG. 14B shows the 13 repeated reads of two codes by the polarity control.
[0038] FIG. 15A shows the schematic of amino acid detection using the method described herein with the ONT platform.
[0039] FIG. 15B shows the two amino acids (R&H) detection using the method described herein with the ONT platform.
[0040] FIG. 15C shows the sequential detection of three codons from two barcodes.DETAILED DESCRIPTION OF THE INVENTION
[0041] Before the present disclosure is described in greater detail, it is to be understood that this disclosure is not limited to particular embodiments described, and as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the present disclosure will be limited only by the appended claims.
[0042] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of the ordinary skills in the art to which this disclosure belongs. Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present disclosure, the preferred methods and materials are now described.
[0043] All publications and patents cited in this specification are herein incorporated by reference as if each individual publication or patent were specifically and individually indicated to be incorporated by reference and are incorporated herein by reference to disclose and describe the methods and / or materials in connection with which the publications are cited. The citation of any publication is for its disclosure prior to the filing date and should not be construed as an admission that the present disclosure is not entitled to antedate such publication by virtue of prior disclosure. Further, the dates of publication provided could be different from the actual publication dates that may need to be independently confirmed.
[0044] As will be apparent to those of skill in the art upon reading this disclosure, each of the individual embodiments described and illustrated herein has discrete components and features which may be readily separated from or combined with the features of any of the other several embodiments without departing from the scope or spirit of the present disclosure. Any recited method can be carried out in the order of events recited or in any other order that is logically possible.
[0045] Definitions
[0046] The following definitions are provided to assist the reader. Unless otherwise defined, all terms of art, notations and other scientific or medical terms or terminology used herein are intended to have meanings commonly understood by those of skill in the chemical and medical arts. In some cases, terms with commonly understood meanings are defined herein for clarity and / or for ready reference, and the inclusion of such definitions herein should not necessarily be construed to represent a substantial difference over the definition of the term as generally understood in the art.
[0047] As used herein, the term “or” is an inclusive “or” operator and is equivalent to the term “and / or”, unless the context clearly dictates otherwise. The term “based on” is not exclusive and allowed for being based on additional factors not described unless the context clearly dictates otherwise. In addition, the singular forms “a,” “an” and “the” include plural references unless the content clearly dictates otherwise. The meaning of “in . . .” includes “within . . .” and “on . . .”.
[0048] As used herein, the term “amino acid” generally refers to an organic compound that combines to form a protein or peptide. An amino acid generally comprises an amine group, a carboxylic acid group, and a side-chain specific to each amino acid, which serves as a monomeric subunit of a peptide. An amino acid may include the 20 standard, naturally occurring or canonical amino acids as well as non-standard amino acids. The standard, naturally-occurring or canonical amino acids include Alanine (A or Ala), Cysteine (C or Cys), Aspartic Acid (D or Asp), Glutamic Acid (E or Glu), Phenylalanine (F or Phe), Glycine (G or Gly), Histidine (H or His), Isoleucine (I or He), Lysine (K or Lys), Leucine (L or Leu), Methionine (M or Met), Asparagine (N or Asn), Proline (P or Pro), Glutamine (Q or Gin), Arginine (R or Arg), Serine (S or Ser), Threonine (T or Thr), Valine (V or Vai), Tryptophan (W or Trp), and Tyrosine (Y or Tyr). An amino acid may be an L-amino acid or a D-amino acid. Non-standard amino acids may be modified amino acids, amino acid analogs, amino acid mimetics, non-standard proteinogenic amino acids, or non-proteinogenic amino acids that occur naturally or are chemically synthesized. Examples of nonstandard amino acids include, but are not limited to, selenocysteine, pyrrolysine, and N- formylmethionine, (3 -amino acids, Homo-amino acids, Proline and Pyruvic acid derivatives, 3- substituted alanine derivatives, glycine derivatives, ring-substituted phenylalanine and tyrosine derivatives, linear core amino acids, and N-methyl amino acids. Moreover, nonstandard amino acids include chemical modifications appear in the post-translational modifications (PTMs), which refer to any alteration in the amino acid sequence of the protein after its synthesis. The modifications include, but not limited to, phosphorylation, glycosylation, ubiquitination, S-nitrosylation, methylation, N-acetylation, lipidation and proteolysis. The amino acid can also be a derivative of natural and non-natural amino acids resulting from a chemical process such as Edman degradation (e.g. PTH amino acid).
[0049] As used herein, the terms “antibody” and “immunoglobulin” may generally refer to proteins that can recognize and bind to a specific antigen. An antibody or immunoglobulin may refer to an antibody isotype, fragments of antibodies including, but not limited to, Fab, Fv, scFv, and Fd fragments, chimeric antibodies, humanized antibodies, single-chain antibodies, and fusion proteins including an antigen-binding portion of an antibody and a non-antibody protein. The antibodies may be detectably labeled, e.g., with a fluorophore, radioisotope, enzyme (e.g., a peroxidase) which generates a detectable product, fluorescent protein, nucleic acid barcode sequence, and the like. The antibodies may be further conjugated to other moieties, such as members of specific binding pairs, e.g., biotin (member of biotin-avidin specific binding pair), and the like. Also encompassed by the terms are Fab', Fv, F(ab')2, and other antibody fragmentsthat retain specific binding to antigen. Antibodies may exist in a variety of other forms including, for example, Fv, Fab, and (Fab)2, as well as bi-functional (i.e., bi-specific) hybrid antibodies (e.g., Lanzavecchia et al., Eur. J. Immunol. 17, 105 (1987)) and in single chains (e.g., Huston et al., Proc. Natl. Acad. Sci. U.S.A., 85, 5879-5883 (1988) and Bird et al., Science, 242, 423-426 (1988), which are incorporated herein by reference). (See, generally, Hood et al., Immunology, Benjamin, N.Y., 2nd ed. (1984), and Hunkapiller and Hood, Nature, 323, 15-16 (1986), which are herein incorporated by reference).
[0050] As used herein, the term “barcode” generally refers to an identifying feature that may be used to distinguish similar items. A barcode may comprise a nucleic acid molecule of about 2 to about 30 bases. A barcode may comprise a nucleic acid molecule of about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150 or more bases, which may provide a unique identifier tag or origin information for a molecule (e.g., protein, polypeptide, peptide), a binding agent, a set of binding agents from a binding cycle, a sample molecule, a set of samples, molecules within a compartment (e.g., droplet, bead, partition or separated location), macromolecules within a set of compartments, a fraction of macromolecules, a set of macromolecule fractions, a spatial region or set of spatial regions, a library of macromolecules, or a library of binding agents. A barcode can be an artificial sequence or a naturally occurring sequence including peptides, proteins, protein complexes, carbohydrates, and synthetic polymeric materials. In certain embodiments, each barcode within a population of barcodes is different. In other embodiments, a portion of barcodes in a population of barcodes is different, e.g., at least about 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, or 99% of the barcodes in a population of barcodes is different. A population of barcodes may be randomly generated or non-randomly generated. A population of barcodes may comprise error correcting barcodes. Barcodes can be used to computationally deconvolute sequence reads derived from an individual molecule, sample, library, etc. Barcodes may comprise multiplexed information, e.g., arising from different samples, compartments, individual molecules, etc. A barcode can also be used for deconvolution of a collection of molecules that have been distributed into small compartments for enhanced mapping. For example, rather than mapping a peptide back to the proteome, the peptide can be mapped back to its originating protein molecule or protein complex, a sample or partition from which itoriginated, etc. A barcode may comprise any useful sequence, including repeat sequences (e.g., a poly -A, poly-T, poly-C, poly-G region) or the barcode may comprise non-repeat sequences. As used herein, a “sample barcode”, also referred to as “sample tag” generally refers to a barcode molecule comprising identifying information of a sample from which a barcoded molecule derives.
[0051] As used herein, the term “hybridization” is used in reference to the pairing of complementary nucleic acids. Hybridization and the strength of hybridization (e.g., the strength of the association between the nucleic acids) is influenced by such factors as the degree of complementary between the nucleic acids, stringency of the conditions involved, and the melting temperature of the formed hybrid. Hybridization methods involve the annealing of one nucleic acid to another, complementary nucleic acid, e.g., based on Watson-Crick base pairing.
[0052] The terms “nucleic acid”, “nucleic acid molecule”, “oligonucleotide” and “polynucleotide” may be used interchangeably herein and generally refer to a polymeric form of naturally occurring or synthetic nucleotides, or analogs thereof, of any length. A nucleic acid molecule may comprise one or more deoxyribonucleotides, deoxynucleotide triphosphates, dideoxynucleotide triphosphates, ribonucleotides, hexitol nucleotides, cyclohexane nucleotides, or analogs or combinations thereof. A nucleic acid molecule may comprise, e.g., DNA, RNA, HNA, CeNA, and modified forms thereof. A nucleic acid molecule may comprise nucleotides that are linked by phosphodiester bonds. A nucleic acid molecule may have any two- or three- dimensional structure, and may perform any function, known or unknown. A nucleic acid molecule may be single stranded, double stranded, or partially double stranded. Non-limiting examples of polynucleotides include a gene, a gene fragment, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, noncoding RNA, small interfering RNA, short hairpin RNA, micro RNA, scaRNA, ribozymes, riboswitches, viral RNA, complementary DNA (cDNA), cosmid DNA, mitochondrial DNA, chromosomal or genomic DNA, viral DNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, control regions, isolated RNA of any sequence, nucleic acid probes, nucleic acid adapters, and primers. The nucleic acid molecule may be linear, circular, or any other geometry. Examples of polynucleotide analogs include but are not limited to xeno nucleic acid (XNA), bridged nucleic acid (BNA), glycol nucleic acid (GNA), hexitol nucleic acid (HNA), cyclohexane nucleic acid (CeNA), peptide nucleic acids (PNAs), yPNAs, morpholino polynucleotides, locked nucleic acids (LNAs), threose nucleic acid (TNA), 2'-O- Methyl polynucleotides, 2'-O-alkyl ribosyl substituted polynucleotides, phosphorothioate polynucleotides, and boronophosphate polynucleotides. A polynucleotide analog may possesspurine or pyrimidine analogs, including for example, 7-deaza purine analogs, 8-hal opurine analogs, 5-halopyrimidine analogs, or universal base analogs that can pair with any base, including hypoxanthine, nitroazoles, isocarbostyril analogues, azole carboxamides, and aromatic triazole analogues, or base analogs with additional functionality, such as a biotin moiety for affinity binding.
[0053] As used herein, the term “peptide” may refer to any short, single peptide chain. A peptide may be no more than about 100, 95, 90, 85, 80, 75, 70, 65, 60, 55, 50, 45, 40, 35, 30, 25, 20, 15, 10, 5, or less than about 5 amino acids in length. A peptide may have a known or unknown biological function or activity. Peptides can include natural, synthetic, modified, or degraded proteins or peptides, or a combination thereof.
[0054] As used herein, “polypeptide” and “protein” can be used interchangeably. The term “polypeptide” refers to two or more amino acids linked together by a peptide bond. The term “polypeptide” includes proteins that have a C-terminal end and an N- terminal end as generally known in the art and may be synthetic in origin or naturally occurring. As used herein “at least a portion of the polypeptide” refers to 2 or more amino acids of the polypeptide. A polypeptide may comprise one or more peptides. Optionally, a portion of the polypeptide includes at least: 1, 5, 10, 20, 30 or 50 amino acids, either consecutive or with gaps, of the complete amino acid sequence of the polypeptide, or the full amino acid sequence of the polypeptide.
[0055] As used herein, “sequencing” generally refers to determining the order of: (A) nucleotides (base sequences) in a nucleic acid sample, e.g., DNA or RNA; or determining the order of (B) amino acids in all or part of a polymer, such as a protein, peptide, or other multimeric molecule. Many techniques are available, such as Sanger sequencing or High Throughput Sequencing technologies (HTS). Sanger sequencing may involve sequencing via detection through (capillary) electrophoresis, in which up to 384 capillaries may be sequence analyzed in one run. High throughput sequencing involves the parallel sequencing of thousands or millions or more sequences at once. HTS can be defined as Next Generation sequencing (NGS), i.e. techniques based on solid phase pyrosequencing or as Next-Next Generation sequencing based on single nucleotide real time sequencing (SMRT). HTS technologies are available such as offered by Roche, Illumina and Applied Biosystems (Life Technologies). Further high throughput sequencing technologies are described by and / or available from Helicos, Pacific Biosciences, Complete Genomics, Ion Torrent Systems, Oxford Nanopore Technologies, Nabsys, ZS Genetics, GnuBio.
[0056] As used herein, “next generation sequencing” refers to high-throughput sequencing methods that allow the sequencing of millions to billions of molecules in parallel. Examples ofnext generation sequencing methods include sequencing by synthesis, sequencing by ligation, sequencing by hybridization, polony sequencing, ion semiconductor sequencing, nanopore sequencing, and pyrosequencing. By attaching primers to a solid substrate and a complementary sequence to a nucleic acid molecule, a nucleic acid molecule can be hybridized to the solid substrate via the primer and then multiple copies can be generated in a discrete area on the solid substrate by using polymerase to amplify (these groupings are sometimes referred to as polymerase colonies or polonies). Consequently, during the sequencing process, a nucleotide at a particular position can be sequenced multiple times (e.g., hundreds or thousands of times) — this depth of coverage is referred to as “deep sequencing.” Examples of high throughput nucleic acid sequencing technology include platforms provided by Illumina, BGI, Qiagen, ThermoFisher, and Roche, including formats such as parallel bead arrays, sequencing by synthesis, sequencing by ligation, capillary electrophoresis, electronic microchips, “biochips,” microarrays, parallel microchips, and single-molecule arrays, as reviewed by Service (Science 311 : 1544-1546, 2006).
[0057] As used herein, the term “stopper” means the compound used to stop the movement of the chimera from escaping the nanopore. The stopper can be selected from, but not limited to, small molecules such as streptavidin, which have the high affinity to bind the terminus, such as the biotinylated terminus, of the polymer; and proper form of nucleotides, such as single strand DNA, which can form a duplex or structure to stop the movement.
[0058] As used herein, the term “translocase” and “unfoldase” may be used interchangeably as one type of motor proteins in the present disclosure. A translocase may be used to control the speed of an analyte through a pore. A translocase may be used to increase the speed of an analyte through a pore. A translocase may be used to decrease the speed of an analyte through a pore. In some embodiments, a translocase comprises helicases, exonucleases, proteases translocases, or topoisomerases. In some embodiments, a translocase can translocate a polypeptide in the N-to-C direction, the C-to-N direction, or both. In some embodiments, a translocase may translate a polypeptide in the C-to-N or N-to-C direction. In some embodiments, a translocase binds to the N-terminus or the C-terminus of a polypeptide.
[0059] As used herein, “analyzing” the macromolecule means to quantify, characterize, distinguish, or a combination thereof, all or a portion of the components of a molecule (e.g., a macromolecule, a biological molecule such as a protein, amino acid, nucleic acid molecule, etc.). For example, analyzing a peptide, polypeptide, or protein may comprise determining all or a portion of the amino acid sequence (contiguous or non-continuous) of the peptide. Analyzing a macromolecule may include partial identification of a component of themacromolecule. For example, partial identification of amino acids in a protein sequence can identify an amino acid in the protein as belonging to a subset of possible amino acids. Analysis may be performed sequentially, e.g., beginning with analysis of the n NTAA (N-terminal amino acid), and then proceeding to the next amino acid of the peptide (i.e., n-1, n-2, n-3, and so forth). In such instances, sequencing may be performed by cleavage of the n NTAA, thereby converting the n-1 amino acid of the peptide to an N-terminal amino acid (referred to herein as the “n-1 NTAA”). Similarly, analysis of a peptide may begin from C-terminus towards the N- terminus with each round of cleavage from the C- terminus creating a new CTAA (C-terminal amino acid). Cleavage of the n CTAA converts the n-1 amino acid of the peptide to a C-terminal amino acid, referred to herein as an “n-1 CTAA”. Analyzing the peptide may also include determining the presence and frequency of post-translational modifications on the peptide, which may or may not include information regarding the sequential order of the post- translational modifications on the peptide. Analyzing the peptide may also include determining the presence and frequency of epitopes in the peptide, which may or may not include information regarding the sequential order or location of the epitopes within the peptide. Analyzing the peptide may include combining different types of analysis, for example obtaining epitope information, amino acid sequence information, post-translational modification information, or any combination thereof.
[0060] Overview
[0061] Proteins perform a vast array of biological functions within organisms, including providing structure to cells and organism, transporting molecules from one location to another, responding to stimuli, catalyzing metabolic reactions, etc. Proteins differ from one another primarily in their sequence of amino acids, which is dictated by the nucleotide sequence of their genes. As a result, the protein information in a cell or organism can to some degree be deciphered through the transcriptome, i.e., the mRNAs expressed by the cell or organism. However, protein complexity surpasses that of the transcriptome, mainly due to post- translational modifications (PTMs) and dynamic changes in response to various factors. Therefore, there is a need to analyze the protein information, e.g., proteomics, of a cell or organism directly.
[0062] Despite their information richness, proteomic data remains largely unexplored compared to genomics. While genomics has benefited from next-generation sequencing (NGS), which enables massive DNA sequence analysis, proteomics lags behind due to limited throughput in current technologies. Nanopore sequencing is emerging as a promisingtechnology for protein sequencing. However, the application of nanopore sequencing in protein is still limited, with one major hurdle of relying on the electrophoretic force to translocate the protein strands through the pores in a controllable fashion. Therefore, to overcome such challenge, the present disclosure in one aspect provides novel methods and systems to sequence a polypeptide via a nano-scale recording device, e.g., nanopore, nanogap, or field effect transistor. In some embodiments, the methods and systems disclosed herein involve generating a chimeric nucleic acid chain from the polypeptide, wherein the chimeric nucleic acid chain contains the amino acid residues dissembled from the polypeptide and encodes the amino acid sequence of the polypeptide. The chimeric nucleic acid chain is then subject to a nano-scale recording device, wherein the chimeric nucleic acid chain binds to some blockers or contains some blockers capable of preventing the amino acid residues on the chimeric nucleic acid chain from translocating in the nano-scale recording device. The methods and systems disclosed herein then dissociate the blockers from the chimeric nucleic acid chain or change the structure of the blockers to allow the amino acid residues to translocate in the nano-scale recording device, thus detecting a signal generated from the translocation of the amino acid residues and determining the identity of the amino acid residues based on such signal.
[0063] Unlike the method of Encodia and Quantum-Si, the methods and systems disclosed herein for polypeptide sequencing don’t need binder or antibodies specific for individual amino acids.
[0064] The use of the blockers can significantly improve the resolution of amino acid identification by increasing the signal to noise ratio. All the signals generated by nucleic acids on the backbone of the chimeric nucleic acid chain are screened by the blockers, which means that the blockers can selectively extract the information of amino acids. Meanwhile, the blockers can increase the duration of single amino acid in the constriction of nanopore, which increases the quality of signals and generates more information for amino acid identification, such as noise, pattern and so on.
[0065] Unlike the conventional nanopores, as the sequential decoding of amino acids of the present disclosure is controlled by blockers, not motor proteins like helicase, the recording condition of the methods and systems disclosed herein can be harsher to get higher quality data.
[0066] Chimeric Nucleic Acid Chain
[0067] In certain embodiments, the methods and systems disclosed herein involve in the first step converting a polypeptide to be sequenced into a chimeric nucleic acid chain, which comprises a nucleic acid strand and the amino acids dissembled from the polypeptide, whereinthe amino acids are linked to the nucleic acid strand sequentially along the nucleic acid strand according to the sequence of the amino acids in the polypeptide. The structure of the chimeric nucleic acid chain and the method of producing the same is described in detail below. Unless specified separately, all the compounds and components mentioned / used can be commercially purchased from diverse suppliers.
[0068] Chimeric Nucleic Acid Chain
[0069] As used herein, the term “chimeric nucleic acid chain” refers to a chain of nucleotides coupled to at least an analyte other than a nucleotide. In particular, the chimeric nucleic chain described herein refers to a chain of nucleotides wherein a plurality of amino acids is sequentially linked to the chain of nucleotides. As used herein, the term “sequentially” in the context of linking amino acids to a nucleic acid chain means that each amino acid is linked to individual nucleotide along the nucleic acid chain. FIG. 1 illustrates an exemplary embodiment of the chimeric nucleic acid chain described in the present disclosure. Referring to FIG. 1, a chimeric nucleic acid chain 100 contains a chain of nucleotides 101. In some examples, the chain of nucleotides 101 is single stranded, double stranded, or partially double stranded. Referring to FIG. 1, a plurality of amino acids 104 is sequentially coupled to the chain of nucleotides 101 via linkers 105. In some embodiments, the amino acids 104 are dissembled from a polypeptide and are sequentially coupled along the chain of nucleotides 101 according to the sequence of the amino acids in the polypeptide, i.e., the order of the amino acids 104 coupled to the chain of nucleotides 101 following the order of the amino acids in the polypeptide from which the amino acids are dissembled.
[0070] In some examples, the chain of nucleotides can be DNA or RNA, which can comprise any useful number and type of nucleotides, e.g., including canonical and noncanonical bases, and the number of nucleotides may be modulated based on the intended purpose. For example, the length of the chain of nucleotides may be modulated to alter a property of the analytes (e.g., volume, aspect ratio, charge, etc.) to which the chain of nucleotides is linked. The chain of nucleotides may additionally or alternatively comprise any useful functional sequences, including but not limited to barcode sequences or other identifying sequences, UMI sequences, enzyme recognition sites (e.g., transposition sites, restriction sites), spacer sequences, sequencing primer sequences, read sequences, or primer sequences. The chain of nucleotides may comprise canonical bases, noncanonical bases, naturally occurring bases, synthetic bases, abasic sites, or a combination thereof.
[0071] The amino acids may be covalently or non-covalently coupled to the chain of nucleotides. The coupling may be performed using any suitable chemistry and reaction conditions and may comprise the use of a linker. In an example, a chain of nucleotides may comprise a first reactive group, e.g., a first click chemistry moiety, as described elsewhere herein, and may be contacted with a linker comprising a second reactive group, e.g., a second click chemistry moiety that is able to react with the first reactive group. The linker may also comprise an additional reactive group that is able to tether to an analyte. For example, the additional reactive group may be a thiocyanate conjugate, e.g., an isothiocyanate (ITC) such as phenyl isothiocyanate (PITC) or naphthylisothiocyanate (NITC), or an aldehyde group, e.g., ortho-phthalaldehyde (OP A), 2,3 -naphthal enedicarboxyaldehyde (ND A), a guanidinylating agent, dinitrofluorobenzene (DNFB), dansyl chloride, or other amino acid-reactive group. The linker may be reacted with an analyte. Use of such a linker comprising at least two reactive groups may allow for (i) tethering of the analyte to the linker and (ii) tethering of the linker to the nucleic acid segment. In some instances, the linker may be provided pre-tethered to the nucleic acid segment prior to contacting with the analyte. In some instances, the conjugation of the analyte to the nucleic acid segment, either via a linker or without a linker, may change the chemical structure of the analyte. For example, if using a linker comprising an isothiocyanate moiety, the analyte may be derivatized to a thiocarbamyl group (e.g., under alkaline conditions), a thiazolinone group (e.g., under acid conditions), a thiohydantoin group, or other chemical moiety, thereby generating a chimeric nucleic acid chain comprising a modified analyte coupled thereto.
[0072] Iterative Process to Generate a Chimeric Nucleic Acid Chain
[0073] In some instances, the chimeric nucleic acid chain disclosed herein is generated using an iterative process as described in W02024 / 030919A1, which is hereby incorporated by reference in its entirety. In short, in the iterative process, individual amino acids of a polypeptide may be sequentially removed and re-tethered together, such that the distance between the individual analytes is increased.
[0074] In an example, an iterative process described herein may comprise providing a polypeptide comprising a plurality of amino acids, a linker (e.g., as described elsewhere herein), a nucleic acid segment, and a capture nucleic acid chain. The linker may be configured to couple to (i) a terminal amino acid (e.g., N-terminal amino acid (NTAA) or C-terminal amino acid (CTAA)) of the peptide and (ii) a nucleic acid segment. The iterative process may further comprise contacting the linker with the amino acid and the nucleic acid segment. Alternatively,or in addition, the linker may be provided pre-tethered to the nucleic acid segment and subsequently reacted with the amino acid. The linker may couple to the amino acid of the peptide to generate an amino acid-linker complex. The amino acid-linker complex may then be coupled to a capture nucleic acid chain via the nucleic acid segment. For example, the capture nucleic acid chain may comprise nucleic acid molecules, which may be coupled to the nucleic acid segment via hybridization, ligation, or both. In some instances, the method may further comprise, after coupling the amino acid-linker complex to the capture nucleic acid chain via the nucleic acid segment, cleaving the amino acid from the peptide to yield an amino acidlinker-capture nucleic acid chain (AALC) complex, and optionally repeating the process. In an example in which the process is repeated, an additional linker may be provided which is configured to couple to (i) an additional amino acid of the peptide (e.g., the n-1 NTAA or n-1 CTAA) and (ii) an additional nucleic acid segment. The method may further comprise contacting the additional linker with the additional amino acid to generate an additional amino acid-linker complex. The additional nucleic acid segment may be coupled to the linker prior to, during, or subsequent to the coupling of the linker to the additional amino acid. The additional nucleic acid segment may be configured to couple to the AALC complex (e.g., via the nucleic acid segment of the AALC complex). As such, in some examples, subsequent to generation of the additional linker- additional amino acid complex, the additional linker-additional amino acid complex may couple to the AALC complex, thereby generating a stacked AALC complex, and the additional amino acid may be cleaved from the peptide prior to, during, or subsequent to generation of the stacked AALC complex.
[0075] In some instances, the iterative process for the polypeptide may occur across a plurality of capture nucleic acid chains. For instance, use of a substrate may facilitate the iterative process for a substrate-bound peptide. For example, a substrate may be provided that comprises a plurality of capture nucleic acid chains, and, in some instances, the capture nucleic acid chains are located adjacent to the peptide or protein. A first amino acid (e.g., n NTAA or n CTAA) of the peptide or protein may be coupled to a first capture nucleic acid chain (e.g., via a first linker and a first nucleic acid segment), a second amino acid (e.g., n-1 NTAA or n-1 CTAA) may be coupled to a second capture nucleic acid chain (e.g., via a second linker and a second nucleic acid segment), and a third amino acid (e.g., n-2 NTAA or n-2 CTAA) may be coupled to a third capture nucleic acid chain (e.g., via a third linker and third nucleic acid segment). In another example, a first amino acid (e.g., n NTAA) may be coupled to a first capture nucleic acid chain, a second amino acid (e.g., n-1 NTAA) may be coupled to the AALC complex (e.g., from the n NTAA), thereby generating a stacked AALC complex, and a thirdamino acid (e.g., n-2 NTAA) may be coupled to a second capture nucleic acid chain. As will be appreciated, any number of amino acids (or modified amino acids) may be coupled to any number of capture nucleic acid chains (or resultant AALC or stacked AALC complexes).
[0076] FIG. 2 schematically illustrates an example workflow of generating a chimeric nucleic acid chain from a polypeptide using an iterative process. In such an example workflow 200, a polypeptide 201 linked to a capture nucleic acid chain 202 is provided, wherein the capture nucleic acid chain 202 is coupled to a bead 203. The capture nucleic acid chain 202 comprises a first nucleic acid molecule (e.g., DNA molecule). In process 204, a terminal amino acid (TAA) 205 of the polypeptide 201 is modified to contain a first reactive group 206. In process 207, a nucleic acid segment 208 pre-tethered to a linker 209 is provided, wherein the linker 209 contains a second reactive group 210 that is capable of reacting with the first reactive group 206. In process 207, the linker 209 is coupled to the TAA 205 via the reaction between the first reactive group 206 and the second reactive group 210, and the nucleic acid segment 208 is coupled to the capture nucleic acid chain 202, e.g., via ligation. In process 211, the TAA 205 is cleaved from the polypeptide 201 to generate an amino acid-linker-capture (AALC) complex 212, which comprises the TAA 205, the linker 209, the nucleic acid segment 208, and the capture nucleic acid chain 202. Processes 204, 207, and 211 may be iterated and repeated any number of times (“rounds”) using additional nucleic acid segment 208 pre-tethered with linkers 309 to tether to the AALC complex 212). Multiple rounds may continue until all or a subset of the amino acids in the polypeptide 201 are removed from the polypeptide and tethered together. For example, a chimeric nucleic acid chain 213 comprising a first nucleic acid chain 214 linked to n amino acids may result from n rounds of the workflow 200 (e.g., iterations of processes 204, 207, and 211).
[0077] In some embodiments, the chimeric nucleic acid chain generated via an iterative process described above and coupled to a substrate may be cleaved after the completion of the iterative process. The cleavage may be mediated by a photo-cleavable linker. Accordingly, the system described herein may also comprise a UV lamp to cleave the chimeric nucleic acid chain from the substrate.
[0078] Capture Nucleic Acid Chain
[0079] In the context of the iterative process described above, a capture nucleic acid chain refers to a nucleic acid chain capable of coupling to an analyte via a nucleic acid segment linking to the analyte. In some embodiments, the capture nucleic acid chain may be coupled to other molecules, such as an amino acid, a peptide, a linker, etc. The coupling of other moleculesto the capture moiety may comprise a covalent interaction or a noncovalent interaction. The coupling may occur by interaction of binding pairs, e.g., biotin and avidin (or streptavidin), antigen or epitope and antibody or antibody fragment, cyclodextrins and small hydrophobic molecules (e.g., alkanes, benzene, polycyclics), cucurbiturils and adamantaneammonium or trimethylamm oniomethyl ferrocene, cyclophane (e.g., calixarenes, cavitands, pillararenes, tetralactams), etc. In some instances, the capture nucleic acid chain comprises an additional nucleic acid segments or analytes from the previous iterative process.
[0080] The capture nucleic acid chain can comprise any naturally occurring, non -naturally occurring or engineered nucleotide base. For example, the nucleic acid molecule may comprise a pseudo-complementary base, a bridged nucleic acid, a xenonucleic acid, a locked nucleic acid, a peptide nucleic acid (PNA), a gamma-PNA, a morpholino, etc., as is described elsewhere herein. The capture nucleic acid chain may comprise one or more functional sequences, including, but not limited to, a priming sequence, sequencing sequence (e.g., P5 or P7 sequence), sequencing read sequence (e.g., R1 or R2 sequence), a mosaic end sequence, a transposase recognition sequence, a cleavage site (e.g., restriction site), a UMI, a blocking group, a spacer sequence, a barcode sequence, or other functional sequence. In some instances, the capture nucleic acid chain comprises a cleavable or releasable moiety, e.g., a restriction enzyme recognition site, an abasic site, a uracil which can be cleaved using USER® or uracil DNA glycosylase, a disulfide bond that can be releasable upon addition of a reducing agent, a photo-cleavable linker etc. In some instances, the capture nucleic acid chain comprises a barcode sequence that comprises any useful information, e.g., the identity of the peptide that is to be analyzed, temporal information, spatial information, etc.
[0081] In some instances, the capture nucleic acid chain is provided coupled to a substrate. In one example, the substrate comprises one or more identical capture nucleic acid chain; these identical capture nucleic acid chain may be used for generating one or more AALC complexes, e.g., for a terminal amino acid, an n-1 amino acid, etc. In some instances, commercially available substrates, e.g., beads (e.g., DNA beads or barcoded beads), flow cells, or chips, e.g., Illumina® HiSeq, iSeq, MiniSeq, NextSeq, NovaSeq, etc. may be used as the substrates described herein.
[0082] The capture nucleic acid chain may be coupled to a substrate using any useful approach. In some instances, the capture moiety comprises a substrate-tethering group or linker or additional functional group. In some examples, the capture moiety comprises a nucleic acid molecule that comprises a substrate-tethering group, e.g., biotin, a click chemistry moiety suchas an azide, that can couple to a substrate, e.g., a substrate comprising streptavidin or a complementary click chemistry moiety that can react with that of the substrate-tethering group. The capture nucleic acid chain may additionally comprise a binding sequence, to which another nucleic acid segment that is part of or coupled to the modified amino acid. In some instances, the capture nucleic acid chain comprises a single-stranded oligonucleotide or a single-stranded region in which a complementary oligonucleotide can hybridize.
[0083] Linker
[0084] One or more linkers may be used to couple the analyte to the chain of nucleotides to thereby generate the chimeric nucleic acid chain. In some instances, the linker comprises a click chemistry moiety. The click chemistry moiety may comprise any suitable bioorthogonal moieties, as described elsewhere herein, e.g., alkenes, alkynes (e.g., alkyne, cycloalkynes such as DBCO and BCN), azides, epoxides, amines, thiols, nitrones, isonitriles, isocyanides, aziridines, activated esters, and tetrazines, and combinations, variations, or derivatives thereof. The linker may be subjected to conditions sufficient to react the first click chemistry moiety to the second click chemistry moiety, e.g., provision of metal catalysts, appropriate solvents, pH, temperature, ionic concentration, or light / energy for any useful duration of time.
[0085] In some embodiments where the analytes to be identified and / or quantified are amino acids, the linker may comprise an amino acid-reactive moiety. The amino acid- reactive moiety of the linker may be any useful moiety that enables the reactive moiety to conjugate to and optionally cleave an amino acid. In some examples, the reactive moiety can react with a terminal amino acid (e.g., NTAA or CTAA). In such examples, the reactive moiety may comprise any primary amine or carboxylic group reactive group, including but not limited to isocyanates, acyl azides, NHS esters, sulfonyl chlorides, aldehydes, glyoxals, epoxides, oxiranes, carbonates, aryl halides, imidoesters, carbodiimides, anhydrides, phenyl esters, isothiocyanates (e.g., phenyl isothiocyanate, sodium isothiocyanate, ammonium isothiocyanates (e.g., tetrabutylammonium isothiocyanate, tetrabutylammonium isothiocyanate), diphenylphosphoryl isothiocyanate), acetyl chloride, cyanogen bromide, carboxypeptidases, azide, alkyne, DBCO, maleimide, succinimide, thiol-thiol disulfide bonds, tetrazine, TCO, vinyl, methylcyclopropene, acryloyl, allyl, among others. Additional examples of amino acid reactive groups are provided in U.S. Pat. Pub. No. 2020 / 0217853A1, which is incorporated by reference herein in its entirety.
[0086] The linker may comprise any additional useful moieties. For example, the linker may comprise a releasable or cleavable moiety, which may facilitate removal of the aminoacid-linker complex from the nucleic acid segment, or portion thereof, or from the substrate. Such a releasable or cleavable moiety may comprise, for example, a disulfide bond, which may be releasable by contacting a reducing agent (e.g., DTT, TCEP). In some examples, the linker may couple to the nucleic acid segment via the releasable or cleavable moiety, alternatively or in addition to the coupling via click chemistry moieties. As such, the coupling between the nucleic acid segment and the linker may be reversible. The linker may additionally comprise any number of spacing moieties, e.g., polymers (e.g., PEG, PVA, polyacrylamide), aminohexanoic acid, nucleic acids, alkyl chains, etc. Such spacing moieties may increase the distance between any other moieties of the linker, e.g., the amino acid-reactive group and the nucleic acid segment-reactive group.
[0087] The linker may comprise any number of spacing moieties, e.g., alkyl chains, polymer spacers (e.g., PEG), nucleic acid or oligo spacers, or other useful spacing moieties which may be useful in modulating the size or molecular weight of the linker. For example, the linker may comprise at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, or a greater number of spacing moieties (e.g., hydrocarbon units, PEG units, nucleotides or spacer sequences etc.). The linker may comprise at most about 100, at most about 10, at most about 9, at most about 8, at most about 7, at most about 6, at most about 5, at most about 4, at most about 3, at most about 2, or at most 1 spacing moieties. The linker may comprise any useful number of functional groups, e.g., for attachment to multiple analytes. The linker may comprise at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, or a greater number of functional groups.
[0088] Nano-Scale Recording Device
[0089] The method and system of polypeptide sequencing disclosed herein involves using a nanoscale recording device, such as nanopores, nanogaps, nanochannels, or a field effect transistor, to identify the amino acids coupled along the chimeric nucleic acid chain. Typically, the nanoscale recording device comprises two compartments connected via a nanopore, nanogap, or nanochannel. In some instances, a nanopore, nanogap, or nanochannel may be provided on a membrane in an ionic solution. A signal may be measured from the nanopore, membrane, or surrounding solution. For example, conductance, current, current blockage, or other parameter within the nanopore may be monitored as a function of time. As a molecule enters the nanopore, nanogap, or nanochannel, a change in the conductance, current, or other parameter may occur and provide information (e.g., size, charge, aspect ratio, volume) on themolecule. Each amino acid or modified amino acid, or a subset of amino acids or modified amino acids, may generate a unique signal that is distinguishable from other amino acids or modified amino acids. As such, the unique signal signatures may be assigned to the amino acids or modified amino acids in order to determine the identity of the amino acids, or modified amino acids.
[0090] In some instances, a modified amino acid may be analyzed numerous times, e.g., via translocation and measuring of a current or conductance of the modified amino acid, through the same nanopore or nanogap. For example, iterative reading of a modified amino acid may be beneficial in improving the accuracy of the reads or identification of the modified amino acid. In such cases, the modified amino acid may be translocated through one or more nanopores at least 2 times, at least 3 times, at least 4 times, at least 5 times, at least 6 times, at least 7 times, at least 8 times, at least 9 times, at least 10 times, at least 20 times, at least 50 times, at least 100 times, at least 200 times, or even greater.
[0091] The nanopore, nanogap, or nanochannel may be generated from an organic material, e.g., a pore-forming protein or a transmembrane protein. Such a protein may be naturally occurring, synthetic, or engineered. Examples of naturally occurring organic nanopores include wild-type aerolysin, alpha-hemolysin, mycobacterial porins (e.g., MspA porin), Phi29 connector channels, Fragaceatoxin C, Cytolysin A, Ferric hydroxyamate uptake component A, Curb specific gene G, outer membrane porin G, viral DNA packaging motors, etc. Alternatively or in addition to, the nanopore, nanochannel, or nanogap may comprise an engineered variant of a naturally occurring nanopore. In some embodiments, the nanopore, nanogap, or nanochannel is comprised of an inorganic material. For example, solid-state nanopores may be made from dielectric materials such as a silicon compound (e.g., silicon nitride, silicon dioxide), an aluminum compound (e.g., aluminum oxide), a titanium compound (e.g., titanium oxide), a molybdenum compound (e.g., molybdenum sulfate), hafnium, graphene, etc. The nanopore, nanochannel, or nanogap may assume any useful form factor or geometry, e.g., gaps or channels within membranes, capillaries, etc., and may be generated using any suitable process, e.g., ion beam sculpting, electron beam exposure. The nanopore, nanochannel, or nanogap may comprise an elastomeric material.
[0092] In some instances, the nanopore, nanogap, or nanochannel is coupled to a protein. The protein may be a molecular motor, which may facilitate movement of the modified amino acid or portion thereof (e.g., the polymerizable material) through the nanopore. In a nonlimiting example, the modified amino acid may comprise an amino acid coupled, optionally via a linker, to a nucleic acid molecule, and the nanopore may be coupled to a helicase or amolecule comprising helicase activity. The helicase may be used to translocate the nucleic acid molecule and the coupled amino acid through the nanopore. In some instances, the molecular motor may increase, decrease, or otherwise change the translocation velocity of the modified amino acid through the nanopore as compared to an unmodified amino acid. In other examples, the molecular motor may comprise a topoisomerase, a polymerase, a nuclease (e.g., endonucleases such as restriction endonucleases or Cas proteins, or exonuclease) an unfoldase, mitotic spindle protein (e.g., nuclear mitotic apparatus protein, kinesin, dynein), or other motor protein (e.g., myosin), a variant thereof, or other polymer-processing protein. In some instances, the protein comprises a protease or proteosome, which can enable “chop-n-drop” or cleaving of the modified amino acid or portion thereof prior to translocation in the nanopore.
[0093] Alternatively or in addition to, translocation of the amino acids through a medium or through a nanopore or nanogap may be facilitated by application of a force. For example, a molecule may be translocated by application of pressure (e.g., pressure-driven flow), an electric field (e.g., via electrophoresis, electroosmotic flow, isoelectric focusing), a magnetic field (e.g., using a ferromagnetic fluid, magnetic particles), light (e.g., optoelectronics).
[0094] A commercially available nanopore system may be used in the methods described herein. For example, a nanopore system from Oxford Nanopore Technologies (ONT), such as the MinlON, VolTRAX, GridlON, PromethlON, MinIT, Flongle, or Q-Line products may be used to identify and characterize the modified amino acids described herein.
[0095] In some embodiments, the nano-scale recording device, e.g., nanopore, may be engineered to improve the resolution of amino acid identification. The change of properties of pore protein, such as charge, size and hydrophilicity, can generate distinguishable current pattern of different amino acids. In some embodiments, pore proteins can be designed by Al to generate ideal candidates with more stable structure and specific characters for the amino acid identification.
[0096] Blockers
[0097] The methods and systems described herein involve using a blocker capable of preventing the amino acids coupled along the chimeric nucleic acid chain from translocation in the nano-scale recording device.
[0098] The blocker described herein may have various forms. In some embodiments, the blocker (type I) is designed to non-covalently bind the chimeric nucleic acid chain to form a structure, such as duplex, triplex, pseudoknot, G-quadruplex, or DNA / RNA origami. In some embodiments, the blocker can be DNA, RNA or artificial nucleic acid homologs. In someembodiments, the blocker can be a protein binding complex, or a nanoparticle. In some embodiments, binding of the blocker to the chimeric nucleic acid chain is mediated by a molecule selected from a mercury ion, a silver ion, anthracycline and a peptide. In some embodiments, the blocker (type II) is a similar structure formed by a nucleic acid sequence contained in the chimeric nucleic acid chain. Both forms of the blockers can prevent the chimeric nucleic acid chain from translocation, e.g., stalling the chimeric nucleic acid chain at the channel of nanopore protein. Meanwhile, an amino acid is stuck at the constriction site of the nanopore. The specific features of the current pattern can be read out during the long-time stalling, including the duration, current level, noise level, transition of multiple status and so on. After stalling, the blocker can be disassociated from the chimera nucleic acid chain or linearize spontaneously, controlled by programmable voltage or toehold displacement, and next blocker stalls the nanopore for the reading of next amino acid. In some embodiments, the stability of the association between the blocker and the chimera nucleic acid chain can be increased by the introduction of artificial nucleic acid homologs, such as PNA, LNA, BNA, and XNA. The interaction between the blocker and the chimeric nucleic acid chain can also be modulated by nucleic acid binding chemicals, such as mercury ion, silver ion, anthracycline, and peptide. The sequential dissociation of the blockers from the chimeric nucleic acid chain generates a stepwise current trace with all the information of amino acids from a polypeptide in order. Thus, the decoding of amino acids on a chimeric nucleic acid chain by nanopore can fulfill the sequencing of target polypeptide and the identification of post-translational modifications (PTMs).
[0099] FIG. 3 illustrates an embodiment of the blocker that non-covalently binds the chimeric nucleic acid chain. In the schematic workflow 300, a chimeric nucleic acid 301 is loaded at a nanopore device 302, wherein an amino acid 303 of the chimeric nucleic acid chain 301 is at positioned at the nanopore channel 304 of the nanopore. A plurality of blockers binds to the chimeric nucleic acid chain 301, preventing the chimeric nucleic acid chain from translocation. In particular, the blocker 305 binding to a region of the chimeric nucleic acid 301 following the amino acid 303 stalls the amino acid 303 at the nanopore channel 304. The specific features of current pattern can be read out during the long-time stalling, including the duration, current level, noise level, transition of multiple status and so on. In the process 306, the blocker 305 dissociates from the chimeric nucleic acid 301, spontaneously, controlled by programmable voltage or toehold displacement, which allows the amino acid 303 to translocate through the channel 304. The nanopore device 302 then reads a change of the current pattern caused by such translocation, which reflects the identity of the amino acid 303. The nextblocker 307 then stalls the amino acid 308 at the channel 304, preventing the nanopore device 302 from reading the amino acid 308. The sequential dissociation of the blockers from the chimeric nucleic acid chain generates a stepwise current trace with all the information of amino acids from a polypeptide in order. Combining with the method of generating the chimeric nucleic acid chain 301 from a polypeptide as described elsewhere in the present disclosure, the amino acids on the chimeric nucleic acid chain 301 can be deciphered by the nanopore device 302 to sequence the target polypeptide.
[0100] The position of the blocker fitted to the nanopore geometry can be controlled by size and shape. As shown in FIG. 4, disassociation of the blocker from the chimeric nucleic acid chain can occur inside or outside the nanopore channel, which is controlled by the blocker.
[0101] FIG. 5 illustrates an embodiment of the blocker that is a structure formed by a nucleic acid sequence contained in the chimeric nucleic acid chain. In the schematic workflow 500, a chimeric nucleic acid 501 is loaded at a nanopore device 502, wherein an amino acid 503 of the chimeric nucleic acid 501 is at positioned at the nanopore channel 504 of the nanopore. A plurality of blockers forms a structure of the chimeric nucleic acid chain 501 that prevents the chimeric nucleic acid chain 501 from translocation. In particular, the blocker 505 forms a structure of the chimeric nucleic acid 501 following the amino acid 503 stalls the amino acid 503 at the nanopore channel 504. The specific features of current pattern can be read out during the long-time stalling, including the duration, current level, noise level, transition of multiple status and so on. In the process 506, the blocker 505 linearizes, spontaneously, controlled by programmable voltage or toehold displacement, which allows the amino acid 503 to translocate through the channel 504. The nanopore device 502 then reads a change of the current pattern caused by such translocation, which reflects the identity of the amino acid 503. The next blocker 507 then stalls the amino acid 508 at the channel 504, preventing the nanopore device 502 from reading the amino acid 508. The sequential linearization of the blockers generates a stepwise current trace with all the information of the amino acids from a polypeptide in order. Combining with the method of generating the chimeric nucleic acid chain 501 from a polypeptide as described elsewhere in the present disclosure, the amino acids on the chimeric nucleic acid chain 501 can be deciphered by the nanopore device 502 to sequence the target polypeptide.
[0102] As shown in FIG. 6, the blocker that binds to a chimeric nucleic acid chain (FIG. 6(a)) or the chimeric nucleic acid chain (FIG. 6(b)) can be modified with chemicals, which interact with the amino acids, to increase the resolution to distinguish the signal of amino acidsand their modifications. Examples of such modification include without limitation chelating agents.
[0103] In some embodiments, the dissociation of the blocker and the chimeric nucleic acid chain can be controlled by the voltage of the nano-scale device, thus controlling the duration of data acquisition in each step. In some embodiments, the blockers may exist on both sides of the nanopore, and the voltage of the nanopore device may be controlled to sequence the chimeric nucleic acid chain in the nanopore multiple times, until a special voltage is applied to release the chimeric nucleic chain from the nanopore.
[0104] Iterative Sequencing
[0105] Also disclosed herein are methods and systems that perform iterative sequencing of a chimeric nucleic acid chain. In such iterative sequencing, after a chimeric nucleic acid chain translocate stepwise via a nanopore channel to be read, the chimeric nucleic acid chain is not completely released from the nanopore channel. Instead, the chimeric nucleic acid is manipulated to translocate back through the nanopore channel and is read again by the nanopore device.
[0106] FIG. 7 illustrates a schematic workflow of an iterative sequencing of a chimeric nucleic acid chain. Now referring to FIG. 7, in process (1), a chimeric nucleic acid chain 701 is loaded at a nanopore device 702, wherein the chimeric nucleic acid chain 701 has a first end 703 and a second end 704. A plurality of blockers binds to the chimeric nucleic acid chain 701, preventing the chimeric nucleic acid chain 701 from translocating through the nanopore channel. The sequential dissociation of the blockers from the chimeric nucleic acid chain generates a stepwise current trace for identifying the amino acids on the chimeric nucleic acid chain 701. In process (2), the chimeric nucleic acid chain 701 is prevented from completely translocated through the nanopore channel by a stopper 705 coupled to the chimeric nucleic acid chain 701 at the second end 704. As used herein, the stopper 705 can be an agent larger than the diameter of the nanopore channel, such as a protein, a nanoparticle, or a bead. Meanwhile, after translocating through the nanopore channel, the first end 701 is coupled to a second stopper 706 larger than the diameter of the nanopore channel. The second stopper 706 can be the same as or different from the first stopper 705. In process (3), the chimeric nuclei acid chain 701 is forced to translocate back through the nanopore channel. The blockers then again bind to the chimeric nucleic acid chain 701. The second stopper 706 prevents the chimeric nucleic acid chain 701 from completely translocating through the nanopore channel. In process (4), the chimeric nucleic acid chain 701 again translocate through the nanopore channel as inprocess (1), thus reading the amino acids of the chimeric nucleic acid chain 701 in a second time. The process (2)-(4) can iterate, thus sequencing the amino acids multiple times. The iterative sequencing process may be stopped by cleaving the stopper 705 from the first end 703. The nanopore is then ready to capture a new chimeric nucleic acid chain for sequencing.
[0107] In some embodiments, the blocker may be tethered on the nanopore protein to increase the binding change when performing iterative sequencing. FIG. 8 illustrates a schematic workflow of an iterative sequencing of a chimeric nucleic acid chain using a blocker tethered on the nanopore protein. Now referring to FIG. 8, in process 801, a chimeric nucleic acid chain 802 is loaded at a nanopore device 803, wherein a blocker 804 tethered on the nanopore protein binds to the chimeric nucleic acid chain 802 to prevent the amino acid 805from translocating through the nanopore channel. The chimeric nucleic acid chain 801 has a first end 806 and a second end 807, each coupled to a stopper 808 and 809, respectively. In process 810, the blocker 804 is dissociated from the chimeric nucleic acid chain 802, allowing the amino acid 804 to translocate through the nanopore channel. The blocker 804 then binds to the chimeric nucleic acid chain 802, now preventing the amino acid 811 from translocating through the nanopore channel. The sequential binding and dissociation of the blocker 804 to the chimeric nucleic acid chain 802 allows the sequencing of the amino acids on the chimeric nucleic acid chain 802. In process 812, the chimeric nucleic acid chain 802 is prevented from completely existing from the nanopore channel by the stopper 809. The chimeric nuclei acid chain 802 is then forced to translocate back through the nanopore channel, while the stopper 808 prevents the complete translocation. The process 801, 810 and 812 can iterate, thus sequencing the amino acids multiple times.
[0108] Multiplex Sequencing
[0109] In some embodiments, the chimeric nucleic acid chain may be modified to allow multiplex sequencing, i.e., sequencing polypeptides derived from multiple samples. In some embodiments, chimeric nucleic acid chains generated from the polypeptides from different samples contain unique chimeric barcode. In some embodiments, the chimeric barcode comprises a nucleic acid chain coupled with barcode blocker. FIG. 9 illustrates a schematic of a chimeric nucleic acid chain containing a chimeric barcode. Referring to FIG. 9, the chimeric nucleic acid chain 901 comprises two segments: the first segment 902 contains a nucleic acid chain coupled with amino acids to be sequenced as described elsewhere in this disclosure; the second segment 903 is a chimeric barcode that contains a nucleic acid chain coupled with barcode molecules. The barcode molecules generate unique current trace when the chimericbarcode segment is read by a nano-scale device, thus generating a barcode signal for the chimeric nucleic acid chain. The chimeric barcode described herein may comprise any useful information, e.g., spatial, temporal, partition-identifying, sample-identifying sequences, a UMI, or other functional sequence.
[0110] Systems, Kits and Compositions
[0111] Also provided herein are systems, kits, and compositions for sequencing polypeptides. The systems, kits, and compositions provided herein may be useful in implementing any of the described methods or may be provided in complement to the described methods.
[0112] In one aspect, a kit of the present disclosure may comprise a substrate, a capture nucleic acid chain, a linker, a nucleic acid segment, a peptide-conjugation reagent, an enzyme (e.g., a polymerizing or ligating enzyme, a cleaving enzyme, a restriction enzyme, a nicking enzyme, an exonuclease, a repair enzyme such as a uracil DNA glycosylase), or any combination thereof. The kits may comprise buffers, reagents, binding agents, catalysts, or other chemicals or biological molecules (e.g., enzymes) necessary for conducting a chemical or enzymatic reaction. The kits of the present disclosure may further comprise instructions for using the components of the kit or for implementing any of the methods and processes described herein. For example, the kit may comprise instructions for conjugating a capture nucleic acid chain onto a substrate using a peptide-conjugation reaction. The kit may comprise instructions for conjugating a capture nucleic acid chain to a substrate, or alternatively, the substrate may comprise the capture nucleic acid chain coupled thereto. Similarly, the kit may comprise instructions for performing an iterative process, as described herein, e.g., to generate a chimeric nucleic acid chain, an amino acid-linker complex, an AALC complex, etc.
[0113] In another aspect, disclosed herein are compositions that may be used to characterize a polypeptide. A composition may comprise a linker covalently attached to a nucleic acid molecule, which may, for example, be useful in the iterative process of a polypeptide for polypeptide sequencing. In an example, a composition may comprise (A) a linker comprising (i) a first moiety that can couple to an amino acid (e.g., a CTAA or NTAA of a peptide), (ii) a second moiety that can couple to a nucleic acid molecule (e.g., DNA), and optionally, (iii) a releasable or cleavable moiety, which may be the same or different moiety as (ii), and also optionally, (iv) a spacer moiety, and (B) a nucleic acid molecule. In another example, the composition may comprise a linker that is covalently coupled to a nucleic acid molecule; such a linker may comprise a moiety that can couple to and optionally cleave anamino acid (e.g., a CTAA or NTAA of a peptide), and the covalently coupled nucleic acid molecule may be configured to tether to another nucleic acid molecule (e.g., a capture moiety), which may in some instances, be provided attached to a substrate.
[0114] A system of the present disclosure may comprise a nano-scale recording device that is configured to record a signal generated by the translocation of a chimeric nucleic acid chain, as described herein, and to provide sequencing reads of the amino acids coupled along the chimeric nucleic acid chain. Alternatively, or in addition to, a system of the present disclosure may be configured to process, prepare, or sequence a chimeric nucleic acid chain. The system may be configured to provide the polypeptide, the capture nucleic acid chain, and a linker comprising a nucleic acid molecule; couple the linker to the polypeptide; and optionally, iterate one or more operations or processes. Accordingly, systems of the present disclosure may comprise any useful apparatuses or tools, including but not limited to mixers, liquid handlers, vortexes, centrifuges, heating or cooling elements, mechanical stages, microfluidic compartments or devices and fluidic controls.
[0115] A system of the present disclosure may also comprise an enzyme (e.g., a polymerizing or ligating enzyme, a cleaving enzyme, a restriction enzyme, a nicking enzyme, an exonuclease, a repair enzyme such as a uracil DNA glycosylase), a detection or labeling agent, buffers, reagents, binding agents, catalysts, or other chemicals or biological molecules (e.g., enzymes) necessary for conducting a chemical or enzymatic reaction, or any combination thereof. The system may further comprise one or more detection (e.g., imaging or mass spectrometry) systems, separation systems (e.g., HPLC), or other analytical instruments.
[0116] A system of the present disclosure may also comprise a UV lamp that cleaves the chimeric nucleic acid from a substrate via a photo-cleavable linker.EXAMPLES
[0117] Unless otherwise specified, all reagents mentioned in the following examples are commercially available.
[0118] Example 1: Single amino acids profiling
[0119] This example illustrates the profiling of single amino acids using the method disclosed herein.
[0120] To profile the single amino acids, a basic unit (BU) composed of a Phenylthiohydantoin (PTH)-amino acid coupled to a DNA backbone (SEQ ID NO: 1, FIG. 10A) is generated by linking an azido-PTH-amino acid to the DNA backbone via copper-click chemistry (FIG. 10B). As shown in FIG. 10A, the DNA backbone contains abasic (1,2-dideoxyribose) (illustrated as “X”) surrounding the modified nucleotide linked to the PTH- amino acid (illustrated as “U”, wherein R represents the amino acid chains or PTM side chains; when R is -CH3 (amino acid is alanine), U is l-((2R,4S,5R)-4-hydroxy-5-(hydroxymethyl) tetrahydrofuran-2-yl)-5-(6-(l-(4-(4-methyl-5-oxo-2-thioxoimidazolidin-l-yl)phenyl)-lH- l,2,3-triazol-4-yl)hex-l-yn-l-yl)pyrimidine-2,4(lH,3H)-dione). The location of amino acid in the BU chimera backbone was engineered to position the amino acid in the sensitive area of a MspA nanopore inserted in a bilipid layer membrane. The DNA blocker (indicated as black box in FIG. 10A and FIG. IOC) in the BU comprises a hairpin structure. This hairpin structure is large and strong enough to pause the translocation until the structure gets changed / deformed to go through the nanopore of a MspA nanopore system by the applied field. The pause provides enough time to acquire the augmented signal data of the amino acid (FIG. IOC). A person skilled in the art would understand that any type(s) of blocker / blockers can be used in the method as long as they can form a large and strong enough structure to pause the translocation until the structure gets deformed to go through the nanopore.
[0121] BU generates electrical pulse signals on the MspA nanopore system (Yan et al., Rapid and multiplex preparation of engineered Mycobacterium smegmatis porin A (MspA) nanopores for single molecule sensing and sequencing, Chem. Sci. (2021) 12: 9339-9346) when individual amino acid of BU passed through the nanopore. The electrical pulse has various features, such as the open nanopore current (Io), the current blockage (lb), the residual current (Ib / Io), the blockage current noise (iRMs-signai), and the duration of the blockage (time) (FIG. 11 A). These features can be used to distinguish different amino acids conjugated to the BU. For example, FIG. 11B showed typical nanopore pulse signals for a few amino acids (H, R, K, W, G). Each signal shown is just one of a few hundred signals collected from the nanopore detection with each BU. The features were extracted from collected signals and statistically mapped in 2D plotting to discern various amino acids. For instance, the above features can distinguish 20 natural amino acids. The exemplary differentiation of 15 natural amino acids is illustrated in FIG. 11C. Various duration (dwell time) of the BU’ blockage between 30 and 140 ms suggests different interaction between the amino acids in BU and the nanopore. While some of the amino acids, such as R, K, H, W, can be easily identified by the distinct level of their current blockages, the noise (IRMs-signai) provides additional dimension to identifying amino acids (such as A, G, F, and Q) with similar current blockages (FIG. 11D).
[0122] Example 2: Sequential detection of three amino acids
[0123] This example illustrates the sequential amino acid detection using the sequencing methods disclosed herein.
[0124] Three amino acids (E, R and W) were conjugated into three different positions (1, 2, 3) respectively in a backbone polymer (SEQ ID NO: 2, FIG. 12A) and the translocation of the polymer was monitored in the detection system with a MspA nanopore (FIG. 12B). The single-channel detection system was assembled using a house-made Teflon chamber, divided by a Teflon film containing a single 100 pm aperture in the center. A lipid bilayer was formed across the aperture using l,2-diphytanoyl-sn-glycero-3 -phosphocholine (DPhPC) at a concentration of 10 mg / mL in pentane. Both the cis and trans compartments were filled with 1.0M KC1 (pH 7.4). The nanopore M2 MspA was introduced into the cis compartment to facilitate single-pore insertion. Data collection was performed at room temperature using an eOne Light amplifier and Elements Data Reader software (Elements). A transmembrane voltage of 150 mV was applied, and data were acquired at a sampling rate of 5 kHz.
[0125] In the presence of the blockers (SEQ ID NOs: 3 and 4, and hairpin of SEQ ID NO: 2, illustrated as the squares in FIG. 12A and FIG. 12B), three different levels of current were detected in a sequential manner (E-R-W). Each step is the registration of one amino acid in the polymer: the first is E with an average current blockage of 104.53 pA, the second is R with an average current blockage of 63.98 pA, and the third W with an average current blockage of 100.28 pA (FIG. 12C, left).
[0126] In another experiment, when the sequence of the amino acids was changed to E-W- R, the average of current blockage are 104.66 pA, 100.44 pA, and 61.32 pA for E, W and R, respectively (FIG. 12C, right). The current levels of the last two were swapped while the first was not affected. This data demonstrates not only the sequential readout of amino acids but also the fact that the signal of one amino acid is not susceptible to any interreferences by the neighboring amino acids.
[0127] For the comparison, the same polymer was applied to the ONT nanopore system (Oxford Nanopore Technologies, Flow Cell (R10.4.1)) without the blocker. An enzyme in the ONT chemistry played a role in the translocation, which obscured the readout of the amino acids among the backbone DNA base signals. The commercial software assigned the same DNA sequence “ATG” to the three different amino acids E, W, and R. This data demonstrated that, without the blocker, the ONT system failed to recognize three amino acids (FIG. 12D).
[0128] Example 3: Detection of three post-transitional modifications (PTMs) of lysine
[0129] This example demonstrates the capability of the methods disclosed herein to accurately distinguish between the different types of modifications to the same amino acid.
[0130] The side chain of lysine often undergoes post-translation modification, which leads to diverse functional groups. Three lysine PTMs (acyl, methyl, dimethyl) were conjugated tothe polymer of SEQ ID NO: 1. The conjugation methods are the same as described in Example 1. The location of the modified lysine in the polymer and its blocker chemistry were engineered by the guidance of the BU design described above. Their translocation was examined separately by the nanopore system using a MspA nanopore as described in the above examples. Each PTM showed a unique electrical signal: the current blockage, the duration time and the nose pattern (IRMS) (FIG. 13).
[0131] As illustrated in FIG. 13, both natural lysine and every lysine PTM (i.e. acyl, methyl, dimethyl) are distinguished, respectively. This unique signal pattern could not be produced without the blocker nor the BU molecular design.
[0132] Example 4: Multiple repeated reads of the same polymer
[0133] This example demonstrates that our sequencing methods could improve the accuracy of sequencing by adding sequencing cycles if necessary.
[0134] The translocation of a single polymer chain through the nanopore can be reversed before it escapes. For this approach, two molecular stoppers (cis stopper and trans stopper) were located at the opposite end of a single chimera polymer (SEQ ID NO: 5, FIG. 14A). The cis stopper is a Streptavidin that binds to the biotinylated terminus of the polymer. The trans stopper is a double stranded DNA formed by the nucleotides at the 5’ end of the polymer (SEQ ID NO: 5) complemented with an external ssDNA (SEQ ID NO: 6). The chimera threaded through the nanopore resembles the rotaxane system, such as a vectorial system where the molecular translocation can be controlled by a biased potential. The chimera carrying two different codes (code 1 and code2, which are natural polynucleotides of T or A followed by blockers (SEQ ID NO: 7) respectively) was trapped after the first read of the two codes. Next, the polarity was reversed so the blocker reassociated with the chimera to make sure the next round detection of the codes. After the second read, the reversible association / dissociation of the blocker repeated, which led to 13 rereads of the two codes in the same chimera (FIG. 14B).
[0135] Example 5: A commercial platform (ONT) feasibility study
[0136] This example demonstrates that the methods described herein can be implemented in a commercially available nanopore platform, e.g., the ONT platform, for the protein sequencing and is flexible enough to accommodate any other nanopore detection platform.
[0137] To test the feasibility of applying the methods described herein to commercially available nanopore systems, a base unit chimera backbone was designed to target a nanopore on the ONT platform (FIG. 15A). R or H amino acid was then conjugated to the base unit chimera backbone (SEQ ID NO: 1) and presented for ONT nanopore analysis. The nanopore detection of R and H showed distinct signal patterns and levels of current blockage (FIG. 15B).
[0138] In addition, three DNA codes in a barcode molecule were sequentially detected and even in the mixture of two different barcode molecules. Among the three DNA codes, code “0” had a higher current blockage than code “1”, and code “1” had a higher current blockage than code “2”. When three-code DNA barcode molecules “012” and “201” were presented for a sequential detection on the ONT system, the corresponding currents were observed for the three DNA codons in the expected orders (FIG. 15C).EMBODIMENTS
[0139] Embodiment 1. A method of sequencing a polypeptide comprising a sequence of amino acids, the method comprising:(a) generating from the polypeptide a chimeric nucleic acid chain comprising(i) a nucleic acid strand comprising a first end and a second end, and(ii) the individual amino acids dissembled from the polypeptide, wherein the individual amino acids are linked to the nucleic acid strand sequentially along the nucleic acid strand from the first end to the second end according to the sequence of the amino acids in the polypeptide;(b) subjecting the chimeric nucleic acid chain to a recording device comprising(i) a first compartment, and(ii) a second compartment connected to the first compartment, wherein the nucleic acid strand non-covalently binds to at least an external blocker in the first compartment, wherein the blocker prevents an amino acid of the chimeric nucleic acid chain from translocating from the first compartment to the second compartment;(c) dissociating the blocker from the nucleic acid strand to allow the amino acid to translocate from the first compartment to the second compartment;(d) detecting via the recording device a signal that is generated from translocation of the amino acid from the first compartment to the second compartment; and(e) identifying the amino acid based on the signal detected in step (d).
[0140] Embodiment 2. The method of embodiment 1, wherein the chimeric nucleic acid chain comprises a second amino acid linked to the nucleic acid strand in a position subsequent to the amino acid, and the nucleic acid strand non-covalently binds to a second external blocker that prevents the second amino acid from translocating from the first compartment to the second compartment, wherein the method further comprises:(f) dissociating the second blocker from the nucleic acid strand to allow the second amino acid to translocate from the first compartment to the second compartment;(g) detecting via the recording device a second signal that is generated from translocation of the second amino acid from the first compartment to the second compartment; and(h) identifying the second amino acid based on the second signal detected in step (g).
[0141] Embodiment 3. The method of embodiment 1 or embodiment 2, wherein the blocker forms a structure selected from duplex, triplex, pseudoknot, G-quadruplex, or DNA / RNA origami;
[0142] optionally, wherein the blocker comprises a DNA, an RNA or a nucleic acid homolog;
[0143] optionally, wherein the blocker comprises a protein or a nanoparticle.
[0144] Embodiment 4. A method of sequencing a polypeptide comprising a sequence of amino acids, the method comprising:(a) generating from the polypeptide a chimeric nucleic acid chain comprising(i) a nucleic acid strand comprising a first end and a second end, and(ii) the individual amino acids dissembled from the polypeptide, wherein the individual amino acids are linked to the nucleic acid strand sequentially along the nucleic acid strand from the first end to the second end according to the sequence of the amino acids in the polypeptide;(b) subjecting the chimeric nucleic acid chain to a recording device comprising(i) a first compartment, and(ii) a second compartment connected to the first compartment, wherein the nucleic acid strand comprises a structure that acts as an internal blocker to prevent an amino acid of the chimeric nucleic acid chain from translocating from the first compartment to the second compartment;(c) changing the structure of the blocker to allow the amino acid to translocate from the first compartment to the second compartment;(d) detecting via the recording device a signal that is generated from translocation of the amino acid from the first compartment to the second compartment; and(e) identifying of the amino acid based on the signal detected in step (d).
[0145] Embodiment 5. The method of embodiment 4, wherein the chimeric nucleic acid chain comprises a second amino acid linked to the nucleic acid strand in a position subsequent to the amino acid, and the nucleic acid strand contains a second internal blocker having a structure preventing the second amino acid from translocating from the first compartment to the second compartment, wherein the method further comprises:(f) changing the structure of the second internal blocker to allow the second amino acid to translocate from the first compartment to the second compartment;(g) detecting via the recording device a second signal that is generated from translocation of the second amino acid from the first compartment to the second compartment; and(h) identifying the second amino acid based on the second signal detected in step (g).
[0146] Embodiment 6. The method of any of embodiments 1-5, wherein a motor protein is not used in step (c) to facilitate the amino acid to translocate from the first compartment to the second compartment;
[0147] optionally, the motor protein is a translocase.
[0148] Embodiment 7. The method of any of embodiments 1-6, wherein the blocker comprises an agent capable of interacting with the amino acid to adjust the signal generated from translocation of the amino acid from the first compartment to the second compartment.
[0149] Embodiment 8. The method of any of embodiments 1-7, wherein the structure of the blocker is changed by controlling voltage in the recording device.
[0150] Embodiment 9. The method of any of embodiments 1-8, wherein the blocker is dissociated from the nucleic acid strand by controlling voltage in the recording device;
[0151] optionally, wherein the blocker is dissociated from the nucleic acid strand spontaneously.
[0152] Embodiment 10. The method of embodiment 9, wherein the recording device is a nanopore device, and the second compartment is connected to the first compartment via a nanopore channel.
[0153] Embodiment 11. The method of embodiment any of embodiments 1-10, wherein the blocker is dissociated from the nucleic acid strand inside the nanopore channel or outside the nanopore channel.
[0154] Embodiment 12. The method of embodiment 1 or 2, wherein the blocker is tethered to the nanopore channel.
[0155] Embodiment 13. The method of embodiment 12, wherein the nanopore channel comprises an agent capable of interacting with the amino acid to adjust the signal generated from translocation of the amino acid from the first compartment to the second compartment.
[0156] Embodiment 14. The method of any of embodiments 1-5, further comprising a step (e’) following step (e) or a step (h’) following step (h): adjusting the recording device to translocating the amino acid from the second compartment to the first compartment through the nanopore channel.
[0157] Embodiment 15. The method of embodiment 14, wherein the second compartment of the recording device comprises at least a second blocker capable of preventing the amino acid from translocating from the second compartment to the first compartment;
[0158] optionally, wherein the second blocker is disintegrated from the nucleic acid strand to allow the amino acid to translocate from the second compartment to the first compartment.
[0159] Embodiment 16. The method of any of embodiments 1-5, wherein the chimeric nucleic acid chain further comprises a stopper linked at the second end of the nucleic acid strand to prevent the nucleic acid strand from completely passing from the first compartment to the second compartment;
[0160] optionally, wherein the second compartment contains a second stopper capable of binding to the first end of the nucleic acid strand to prevent the nucleic acid strand from completely passing from the second compartment to the first compartment.
[0161] Embodiment 17. The method of any of embodiments 1-5, wherein the nucleic acid strand comprises a barcode sequence.
[0162] Embodiment 18. A system of sequencing a polypeptide comprising a sequence of amino acids, the system comprising:(a) a module of generating from the polypeptide a chimeric nucleic acid chain comprising(i) a nucleic acid strand comprising a first end and a second end, and(ii) the individual amino acids dissembled from the polypeptide, wherein the individual amino acids are linked to the nucleic acid strand sequentially along the nucleic acid strand from the first end to the second end according to the sequence of the amino acids in the polypeptide;(b) a recording device comprising(i) a first compartment, and(ii) a second compartment connected to the first compartment, wherein nucleic acid strand non-covalently binds to at least a blocker in the first compartment, wherein the blocker prevents an amino acid of the chimeric nucleic acid chain from translocating from the first compartment to the second compartment;(c) a module of disintegrating the blocker from the nucleic acid strand to allow the amino acid to translocate from the first compartment to the second compartment;(d) a module of detecting a signal that is generated from translocation of the amino acid from the first compartment to the second compartment; and(e) a module of identifying the amino acid based on the signal detected in step (d).
[0163] Embodiment 19. The system of embodiment 18, wherein the chimeric nucleic acid chain comprises a second amino acid linked to the nucleic acid strand in a position subsequent to the amino acid, and the nucleic acid strand non-covalently binds to a second external blocker that prevents the second amino acid from translocating from the first compartment to the second compartment, wherein the system further comprises:(f) a module of dissociating the second blocker from the nucleic acid strand to allow the second amino acid to translocate from the first compartment to the second compartment;(g) a module of detecting via the recording device a second signal that is generated from translocation of the second amino acid from the first compartment to the second compartment; and (h) a module of identifying the second amino acid based on the second signal detected in step (g).
[0164] Embodiment 20. A system of sequencing a polypeptide comprising a sequence of amino acids, the system comprising:(a) a module of generating from the polypeptide a chimeric nucleic acid chain comprising(i) a nucleic acid strand comprising a first end and a second end, and(ii) the individual amino acids dissembled from the polypeptide, wherein the individual amino acids are linked to the nucleic acid strand sequentially along the nucleic acid strand from the first end to the second end according to the sequence of the amino acids in the polypeptide;(b) a recording device comprising(i) a first compartment, and(ii) a second compartment connected to the first compartment, wherein the nucleic acid strand comprises a blocker forming a structure that prevents an amino acid of the chimeric nucleic acid chain from translocating from the first compartment to the second compartment;(c) a module of changing the structure of the blocker to allow the amino acid to translocate from the first compartment to the second compartment;(d) a module of detecting a signal that is generated from translocation of the amino acid from the first compartment to the second compartment; and(e) a module of identifying the amino acid based on the signal detected in (d).
[0165] The method of embodiment 21, wherein the chimeric nucleic acid chain comprises a second amino acid linked to the nucleic acid strand in a position subsequent to the amino acid, and the nucleic acid strand contains a second internal blocker having a structure preventing thesecond amino acid from translocating from the first compartment to the second compartment, wherein the system further comprises:(f) a module of changing the structure of the second internal blocker to allow the second amino acid to translocate from the first compartment to the second compartment;(g) a module of detecting via the recording device a second signal that is generated from translocation of the second amino acid from the first compartment to the second compartment; and (h) a module of identifying the second amino acid based on the second signal detected in step (g).
[0166] The above examples and embodiments are provided to better illustrate the claimed invention and are not to be interpreted as limiting the scope of the invention. All specific compositions, materials, and methods described above, in whole or in part, fall within the scope of the present invention. These specific compositions, materials, and methods are not intended to limit the invention, but merely to illustrate specific embodiments falling within the scope of the invention. A person skilled in the art may develop equivalent compositions, materials, and methods without the exercise of inventive capacity and without departing from the scope of the invention. It will be understood that many variations can be made in the procedures herein described while still remaining within the bounds of the present invention. It is the intention of the inventors that such variations are included within the scope of the invention.References1. Derrington IM, Butler TZ, Collins MD, Manrao E, Pavlenok M, Niederweis M, Gundlach JH. Nanopore DNA sequencing with MspA. Proc Natl Acad Sci U S A. 2010 Sep 14; 107(37): 16060-5. doi: 10.1073 / pnas.1001831107. Epub 2010 Aug 26. PMID: 20798343; PMCID: PMC2941267.2. Y. Ding, A. M. Fleming, H. S. White, C. J. Burrows, Internal vs fishhook hairpin DNA: unzipping locations and mechanisms in the alpha-hemolysin nanopore. J Phys Chem B 118, 12873- 12882 (2014).3. S . Yan et al. , Non-binary Encoded Nucleic Acid Barcodes Directly Readable by a Nanopore . Angewandte Chemie International Edition 61, e202116482 (2022)4. A. H. Laszlo et al., Decoding long nanopore sequencing reads of natural DNA. Nat Biotech 32, 829-833 (2014).5. US20230230636A1 Nanopore unzipping-sequencing for DNA data storage6. Zhang, M., Tang, C., Wang, Z., Chen, S., Zhang, D., Li, K., Sun, K., Zhao, C., Wang, Y ., Xu, M., Dai, L., Lu, G., Shi, H., Ren, H., Chen, L., & Geng, J. (2024). Real-time detection of 20 amino acids and discrimination of pathologically relevant peptides with functionalized nanopore. Nature Methods.7. Wang, K., Zhang, S., Zhou, X., Yang, X., Li, X., Wang, Y., Fan, P., Xiao, Y., Sun, W., Zhang, P., Li, W., & Huang, S. (2023). Unambiguous discrimination of all 20 proteinogenic amino acids and their modifications by nanopore. Nature Methods.8. Brinkerhoff, H., Kang, A. S. W., Liu, J., Aksimentiev, A., & Dekker, C. (n.d.). Multiple rereads of single proteins at single-amino acid resolution using nanopores.
Claims
WHAT IS CLAIMED IS:
1. A method of sequencing a polypeptide comprising a sequence of amino acids, the method comprising:(a) generating from the polypeptide a chimeric nucleic acid chain comprising(i) a nucleic acid strand comprising a first end and a second end, and(ii) the individual amino acids dissembled from the polypeptide, wherein the individual amino acids are linked to the nucleic acid strand sequentially along the nucleic acid strand from the first end to the second end according to the sequence of the amino acids in the polypeptide;(b) subjecting the chimeric nucleic acid chain to a recording device comprising(i) a first compartment, and(ii) a second compartment connected to the first compartment, wherein the nucleic acid strand non-covalently binds to at least an external blocker in the first compartment, wherein the blocker prevents an amino acid of the chimeric nucleic acid chain from translocating from the first compartment to the second compartment;(c) dissociating the blocker from the nucleic acid strand to allow the amino acid to translocate from the first compartment to the second compartment;(d) detecting via the recording device a signal that is generated from translocation of the amino acid from the first compartment to the second compartment; and(e) identifying the amino acid based on the signal detected in step (d).
2. The method of claim 1, wherein the chimeric nucleic acid chain comprises a second amino acid linked to the nucleic acid strand in a position subsequent to the amino acid, and the nucleic acid strand non-covalently binds to a second external blocker that prevents the second amino acid from translocating from the first compartment to the second compartment, wherein the method further comprises:(f) dissociating the second blocker from the nucleic acid strand to allow the second amino acid to translocate from the first compartment to the second compartment;(g) detecting via the recording device a second signal that is generated from translocation of the second amino acid from the first compartment to the second compartment; and(h) identifying the second amino acid based on the second signal detected in step (g).
3. The method of claim 1, wherein the blocker forms a structure selected from duplex, triplex, pseudoknot, G-quadruplex, or DNA / RNA origami; optionally, wherein the blocker comprises a DNA, an RNA or a nucleic acid homolog; optionally, wherein the blocker comprises a protein or a nanoparticle.
4. A method of sequencing a polypeptide comprising a sequence of amino acids, the method comprising:(a) generating from the polypeptide a chimeric nucleic acid chain comprising(i) a nucleic acid strand comprising a first end and a second end, and(ii) the individual amino acids dissembled from the polypeptide, wherein the individual amino acids are linked to the nucleic acid strand sequentially along the nucleic acid strand from the first end to the second end according to the sequence of the amino acids in the polypeptide;(b) subjecting the chimeric nucleic acid chain to a recording device comprising(i) a first compartment, and(ii) a second compartment connected to the first compartment, wherein the nucleic acid strand comprises a structure that acts as an internal blocker to prevent an amino acid of the chimeric nucleic acid chain from translocating from the first compartment to the second compartment;(c) changing the structure of the blocker to allow the amino acid to translocate from the first compartment to the second compartment;(d) detecting via the recording device a signal that is generated from translocation of the amino acid from the first compartment to the second compartment; and(e) identifying of the amino acid based on the signal detected in step (d).
5. The method of claim 4, wherein the chimeric nucleic acid chain comprises a second amino acid linked to the nucleic acid strand in a position subsequent to the amino acid, and the nucleic acid strand contains a second internal blocker having a structure preventing the second amino acid from translocating from the first compartment to the second compartment, wherein the method further comprises:(f) changing the structure of the second internal blocker to allow the second amino acid to translocate from the first compartment to the second compartment;(g) detecting via the recording device a second signal that is generated from translocation of the second amino acid from the first compartment to the second compartment; and(h) identifying the second amino acid based on the second signal detected in step (g).
6. The method of any of claims 1-5, wherein a motor protein is not used in step (c) to facilitate the amino acid to translocate from the first compartment to the second compartment; optionally, the motor protein is a translocase.
7. The method of any of claims 1-5, wherein the blocker comprises an agent capable of interacting with the amino acid to adjust the signal generated from translocation of the amino acid from the first compartment to the second compartment.
8. The method of any of claims 1-5, wherein the structure of the blocker is changed by controlling voltage in the recording device.
9. The method of any of claims 1-5, wherein the blocker is dissociated from the nucleic acid strand by controlling voltage in the recording device; optionally, wherein the blocker is dissociated from the nucleic acid strand spontaneously.
10. The method of claim 9, wherein the recording device is a nanopore device, and the second compartment is connected to the first compartment via a nanopore channel.
11. The method of claim 9, wherein the blocker is dissociated from the nucleic acid strand inside the nanopore channel or outside the nanopore channel.
12. The method of claim 1 or 2, wherein the blocker is tethered to the nanopore channel.
13. The method of claim 12, wherein the nanopore channel comprises an agent capable of interacting with the amino acid to adjust the signal generated from translocation of the amino acid from the first compartment to the second compartment.
14. The method of any of claims 1-5, further comprising a step (e’) following step (e) or a step (h’) following step (h): adjusting the recording device to translocating the amino acid from the second compartment to the first compartment through the nanopore channel.
15. The method of claim 14, wherein the second compartment of the recording device comprises at least a second blocker capable of preventing the amino acid from translocating from the second compartment to the first compartment; optionally, wherein the second blocker is disintegrated from the nucleic acid strand to allow the amino acid to translocate from the second compartment to the first compartment.
16. The method of any of claims 1-5, wherein the chimeric nucleic acid chain further comprises a stopper linked at the second end of the nucleic acid strand to prevent the nucleic acid strand from completely passing from the first compartment to the second compartment;optionally, wherein the second compartment contains a second stopper capable of binding to the first end of the nucleic acid strand to prevent the nucleic acid strand from completely passing from the second compartment to the first compartment.
17. The method of any of claims 1-5, wherein the nucleic acid strand comprises a barcode sequence.
18. A system of sequencing a polypeptide comprising a sequence of amino acids, the system comprising:(a) a module of generating from the polypeptide a chimeric nucleic acid chain comprising(i) a nucleic acid strand comprising a first end and a second end, and(ii) the individual amino acids dissembled from the polypeptide, wherein the individual amino acids are linked to the nucleic acid strand sequentially along the nucleic acid strand from the first end to the second end according to the sequence of the amino acids in the polypeptide;(b) a recording device comprising(i) a first compartment, and(ii) a second compartment connected to the first compartment, wherein nucleic acid strand non-covalently binds to at least a blocker in the first compartment, wherein the blocker prevents an amino acid of the chimeric nucleic acid chain from translocating from the first compartment to the second compartment;(c) a module of disintegrating the blocker from the nucleic acid strand to allow the amino acid to translocate from the first compartment to the second compartment;(d) a module of detecting a signal that is generated from translocation of the amino acid from the first compartment to the second compartment; and(e) a module of identifying the amino acid based on the signal detected in step (d).
19. A system of sequencing a polypeptide comprising a sequence of amino acids, the system comprising:(a) a module of generating from the polypeptide a chimeric nucleic acid chain comprising(i) a nucleic acid strand comprising a first end and a second end, and(ii) the individual amino acids dissembled from the polypeptide, wherein the individual amino acids are linked to the nucleic acid strand sequentially along the nucleic acid strand from the first end to the second end according to the sequence of the amino acids in the polypeptide;(b) a recording device comprising(i) a first compartment, and(ii) a second compartment connected to the first compartment, wherein the nucleic acid strand comprises a blocker forming a structure that prevents an amino acid of the chimeric nucleic acid chain from translocating from the first compartment to the second compartment;(c) a module of changing the structure of the blocker to allow the amino acid to translocate from the first compartment to the second compartment;(d) a module of detecting a signal that is generated from translocation of the amino acid from the first compartment to the second compartment; and(e) a module of identifying the amino acid based on the signal detected in (d).