Methods of multiplexed polypeptide sequencing
The method addresses the challenge of high noise in nanopore protein sequencing by using a polymerizable chain-amino acid complex with a barcode to enhance sequencing accuracy and compatibility with NGS and nanopore systems.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-11
- Publication Date
- 2026-03-19
AI Technical Summary
Existing methods for protein sequencing, particularly using nanopore technology, face challenges in achieving a high signal-to-noise ratio due to background noise and signal fluctuations, and there is a need for a barcode technology compatible with both NGS and nanopore sequencing systems to accurately sequence amino acids and nucleic acid barcodes simultaneously.
A method involving coupling a barcode to a polypeptide, forming a polymerizable chain-amino acid complex (PCAA), and using a nanopore device to determine the sequence of amino acids and barcode identity, enabling simultaneous sequencing of both through a nanopore sensor.
The method achieves accurate sequencing of polypeptides by determining amino acid sequences and identifying the polypeptide origin based on the barcode, enhancing the signal-to-noise ratio and compatibility with multiple sequencing systems.
Smart Images

Figure 00000049_0000 
Figure 00000049_0001 
Figure 00000050_0000
Abstract
Description
088177-8007W001METHODS OF MULTIPLEXED POLYPEPTIDE SEQUENCINGSEQUENCE LISTING
[0001] The sequence listing that is contained in the file named “088177- 8007W001_SeqL.xml”, which is 16,902 bytes and was created on September 11, 2025, is filed herewith by electronic submission and is incorporated by reference herein.CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to US provisional application 63 / 693,705, filed September 11, 2024, the disclosure of which is incorporated herein by reference.FIELD OF THE INVENTION
[0003] The present disclosure generally relates to molecular biology and analyte analysis. In particular, the present disclosure relates to methods and systems of sequencing polypeptides.BACKGROUND
[0004] Proteins play vital roles in almost every biological process. While the structure of a protein is largely dictated by the amino acid sequence of the protein, protein sequencing provides invaluable insights into the protein's structure, function, and evolutionary relationships. By elucidating the amino acid sequence of a protein, people can uncover crucial information about the protein’s folding, interactions with other molecules, and involvement in biological pathways.
[0005] The two major direct methods of protein sequencing are mass spectrometry and Edman degradation, which laid the foundation of modem protein sequencing technologies. More recently, nanopore sequencing, which is primarily used for DNA and RNA sequencing, has been explored for protein sequencing. Nanopores can potentially label-free detect and characterize individual protein molecules or peptides as they pass through nanoscale pores, offering real-time analysis. Achieving a high signal-to-noise ratio (SNR) is critical for accurate protein sequence determination, but it can be challenging due to background noise and fluctuations in signal intensity. Meanwhile, the properties of peptide chain are entirely different from DNA and RNA, such as charge and rigidity. Thus, the difficulty of single amino acid identification in a polypeptide jumps up.
[0006] To develop an improved polypeptide sequencing, a method of dissembling polypeptide to single amino acids and conjugating the amino acids to a DNA backbone has been disclosed (see W02024030919A1). The DNA backbone may comprise a nucleic acid088177-8007W001 barcode. The DNA-amino acid chimera was then sequenced by commercial sequencing platform, such as NGS or nanopore platform, and the identity of the DNA-amino acid chimera may be determined by the nucleic acid barcode. However, if they are using nanopore-based polypeptide sequencing, there will be a challenge to sequence the amino acids on the DNA- amino acid chimera and the nucleic acid barcode on the DNA backbone via a nanopore reader system at the same time. Therefore, there is a need to develop a barcode technology that is not only compatible with NGS sequencing system, but also compatible with a nanopore protein sequencing system.SUMMARY OF INVENTION
[0007] The present disclosure in one aspect provides a method of sequencing a plurality of polypeptides.
[0008] In some embodiments, the method comprises:(a) coupling a first barcode to a first polypeptide in a first sample, wherein the first polypeptide comprises a sequence of amino acids, wherein the first barcode comprises at least a first barcode moiety, and wherein the first barcode moiety comprises(i) a first codon capable of being recognized by a nanopore device having a nanopore sensor, and(ii) a first time-control element capable of temporarily holding the first codon at the nanopore sensor;(b) generating a first polymerizable chain-amino acid (PCAA) complex comprising(i) a first stacked polymerizable chain,(ii) the amino acids dissembled from the first polypeptide, wherein the amino acids dissembled from the first polypeptide are sequentially linked to the first stacked polymerizable chain, and(iii) the first barcode linked to the first stacked polymerizable chain;(c) subj ecting the first PCAA complex to the nanopore device to determine:(i) the sequence of amino acids along the first stacked polymerizable chain, and(ii) the identity of the first barcode; and(d) determining088177-8007W001(i) the sequence of the first polypeptide based on the sequence of amino acids along the first stacked polymerizable chain, and(ii) that the first polypeptide is from the first sample based on the identity of the first barcode. In some embodiments, the first barcode comprises a first polymerizable chain; optionally, wherein the first polymerizable chain is a nucleic acid or PEG; optionally, wherein the nucleic acid is a DNA, a RNA or a modified variant thereof.
[0009] In some embodiments, the first codon comprises a polynucleotide. In some embodiments, the first codon comprises an amino acid derivative; optionally, wherein the amino acid derivative is a phenylthiohydantoin amino acid. In some embodiments, the first codon comprises a polynucleotide and an amino acid derivative. In some embodiments, the first codon comprises a compound selected from bicyclo[6.1.0]nonyne (BCN), a crown ether molecule, a sugar molecule (e.g., a sialic acid).
[0010] In some embodiments, the first barcode further comprises a second barcode moiety comprising a second codon and a second time-control element.
[0011] In some embodiments, the first time-control element comprises a polynucleotide. In some embodiments, the first time-control element forms a structure capable of temporarily holding the first codon at the nanopore sensor. In some embodiments, the first time-control element binds to a reagent to form a complex capable of temporarily holding the first codon at the nanopore sensor.
[0012] In some embodiments, nanopore device does not comprise a motor protein; optionally, the motor protein is a translocase; optionally, the translocase is a helicase, a DNA polymerase or a RNA polymerase.
[0013] In some embodiments, the method further comprises:(b’) mixing the first PCAA complex with a second PCAA complex, wherein the second PCAA complex is generated from a second polypeptide in a second sample, wherein the second PCAA complex comprises(i) a second stacked polymerizable chain and the amino acids dissembled from the second polypeptide, wherein the amino acids dissembled from the second polypeptide are sequentially linked to the second stacked polymerizable chain, and(ii) a second barcode linked to the second stacked polymerizable chain, wherein the second barcode is different from the first barcode;(c’) subjecting the second PCAA complex to the nanopore device to determine:088177-8007W001(i) the sequence of amino acids along the second stacked polymerizable chain, and(ii) the identity of the second barcode; and(d’) determining(i) the sequence of the second polypeptide based on the sequence of amino acids along the second stacked polymerizable chain, and(ii) that the second polypeptide is from the second sample based on the identity of the second barcode.
[0014] In another aspect, the present disclosure provides a method comprising:(a) coupling a first barcode to each of a first plurality of polypeptides in a first sample, wherein the first barcode comprises at least a first barcode moiety, and wherein the first barcode moiety comprises(i) a first codon capable of being recognized by a nanopore device having a nanopore sensor, and(ii) a first time-control element capable of temporarily holding the first codon at the nanopore sensor;(b) generating a first plurality of PCAA complexes, each PCAA complex comprising(i) a first stacked polymerizable chain,(ii) amino acids dissembled from a polypeptide in the first plurality of polypeptides, wherein the amino acids dissembled from the polypeptide are sequentially linked to the first stacked polymerizable chain, and(iii) the first barcode, wherein the first barcode is linked to the first stacked polymerizable chain;(c) mixing the first plurality of PCAA complexes with a second plurality of PCAA complexes, wherein the second plurality of PCAA complexes is generated from a second plurality of polypeptides in a second sample, wherein each of the second plurality of PCAA complexes comprises(i) a second stacked polymerizable chain,(ii) amino acids dissembled from a polypeptide in the second plurality of polypeptides, wherein the amino acids are sequentially linked to the second stacked polymerizable chain, and(iii) a second barcode linked to the second stacked polymerizable chain, wherein the second barcode is different from the first barcode;088177-8007W001(d) subjecting a mixture of the first plurality of PCAA complexes and the second plurality of PCAA complexes to the nanopore device to determine:(i) the sequence of amino acids along the first stacked polymerizable chain,(ii) the identity of the first barcode,(i1) the sequence of amino acids along the second stacked polymerizable chain,(ii) the identity of the second barcode; and(e) determining(i) the sequence of the polypeptide in the first plurality of polypeptides based on the sequence of amino acids along the first stacked polymerizable chain,(i1) the sequence of the polypeptide in the second plurality of polypeptides based on the sequence of amino acids along the second stacked polymerizable chain,(ii) that the polypeptide in the first plurality of polypeptides is from the first sample based on the identity of the first barcode, and(ii’) that the polypeptide in the second plurality of polypeptides is from the second sample based on the identity of the second barcode.
[0015] In another aspect, the present disclosure provides a system comprising:(a) a module for coupling a first barcode to a first polypeptide in a first sample, wherein the first polypeptide comprises a sequence of amino acids, wherein the first barcode comprises at least one barcode moiety, and wherein the first barcode moiety comprises(i) a first codon capable of being recognized by a nanopore device having a nanopore sensor, and(ii) a first time-control element capable of temporarily holding the first codon at the nanopore sensor;(b) a module for generating a first PCAA complex comprising(i) a first stacked polymerizable chain and the amino acids dissembled from the first polypeptide, wherein the amino acids dissembled from the first polypeptide are sequentially linked to the first stacked polymerizable chain, and(ii) the first barcode linked to the first stacked polymerizable chain;088177-8007W001(c) a module for subjecting the first PCAA complex to the nanopore device to determine:(i) the sequence of amino acids along the first stacked polymerizable chain, and(ii) the identity of the first barcode; and(d) a module for determining(i) the sequence of the first polypeptide based on the sequence of amino acids along the first stacked polymerizable chain, and(ii) that the first polypeptide is from the first sample based on the identity of the first barcode.
[0016] In some embodiments, the system further comprises:(b’) a module for mixing the first PCAA complex with a second PCAA complex, wherein the second PCAA complex is generated from a second polypeptide in a second sample, wherein the second PCAA complex comprises(i) a second stacked polymerizable chain and the amino acids dissembled from the second polypeptide, wherein the amino acids dissembled from the second polypeptide are sequentially linked to the second stacked polymerizable chain, and(ii) a second barcode linked to the second stacked polymerizable chain, wherein the second barcode is different from the first barcode;(c’) a module for subjecting the second PCAA complex to the nanopore device to determine:(i) the sequence of amino acids along the second stacked polymerizable chain, and(ii) the identity of the second barcode; and(d’) a module for determining(i) the sequence of the second polypeptide based on the sequence of amino acids along the second stacked polymerizable chain, and(ii) that the second polypeptide is from the second sample based on the identity of the second barcode.
[0017] In another aspect, the present disclosure provides a barcode composition comprising at least a first barcode moiety, wherein the first barcode moiety comprises(i) a first codon capable of being recognized by a nanopore device having a nanopore sensor, and088177-8007W001(ii) a first time-control element capable of temporarily holding the codon at the nanopore sensor.
[0018] In some embodiments, the codon comprises a polynucleotide. In some embodiments, the codon comprises an amino acid derivative; optionally, wherein the amino acid derivative is a phenylthiohydantoin amino acid. In some embodiments, the codon comprises a polynucleotide and an amino acid derivative. In some embodiments, the codon comprises a compound selected from bicyclo[6.1.0]nonyne (BCN), a crown ether molecule, a sugar molecule (e.g., a sialic acid).
[0019] In some embodiments, the barcode further comprises a second barcode moiety comprising a second codon and a second time-control element.
[0020] In some embodiments, the first time-control element comprises a polynucleotide. In some embodiments, the first time-control element forms a structure capable of temporarily holding the codon at the nanopore sensor. In some embodiments, the first time-control element binds to a reagent to form a complex capable of temporarily holding the codon at the nanopore sensor.
[0021] In some embodiments, nanopore device does not comprise a motor protein; optionally, the motor protein is a translocase; optionally, the translocase is a helicase, a DNA polymerase or a RNA polymerase.BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The following drawings form part of the present specification and are included to further demonstrate certain aspects of the present disclosure. The disclosure may be better understood by reference to one or more of these drawings in combination with the detailed description of specific embodiments presented herein.
[0023] FIG. 1 shows the schematic of a Type I barcode.
[0024] FIG. 2 shows the schematic of a Type II barcode.
[0025] FIG. 3 shows the schematic of the mechanism of decoding a barcode having a codon and a time-control element on a PCAA complex by a nanopore device, wherein the timecontrolling element forms a structure capable of temporarily holding the codon at the nanopore sensor.
[0026]
[0027] FIG. 4 shows the schematic of the mechanism of decoding a barcode having a codon and a time-control element on a PCAA complex by a nanopore device, wherein the time-088177-8007W001 control element binds to a reagent to form a complex capable of temporarily holding the codon at the nanopore sensor.
[0028] FIG. 5 shows an exemplary embodiment of the methods described in the present disclosure.
[0029] FIG. 6 shows the schematic of an exemplary embodiment of the PCAA complex described in the present disclosure.
[0030] FIG. 7 shows the schematic of an example workflow of generating a PCAA complex from a barcoded polypeptide using an iterative process.
[0031] FIG. 8 shows the backbones of two three-codon barcodes made of DNA (SEQ ID NO: 1 and 2, respectively). Each barcode backbone contains three codons (ATC and TCA, respectively), wherein each codon is made of 4-mer nucleotides. The cholesterol tag sequence of the barcodes complements to a cholesterol tag (SEQ ID NO: 4). Each codon of the barcodes is followed by a blocker sequence, which complements to a blocker (SEQ ID NO: 3).
[0032] FIG. 9 shows the plot of concentration ratio vs frequency ratio. The dash line shows the linear fit. The plot demonstrates reliable relative quantification of these two barcodes mixed at a large range of ratios.
[0033] FIG. 10 that PTH-AA (phenylthiohydantoin amino acid) can be used to conjugate to alkyne-DNA for Type ILA chimera barcodes or conjugate to DNA-sBCN for Type ILB chimera barcodes.
[0034] FIG. 11 shows the assembly of a DNA-AA trimer by annealing three monomer units (A, B, and C) in the presence of two complementary splint DNA oligos (X and Y).
[0035] FIG. 12 shows a schematic of the solid-phase synthesis method.
[0036] FIG. 13 shows sequences of the components for an exemplary solid-phase synthesis method.
[0037] FIG. 14 shows the purified three-codon chimera barcode in a gel electrophoresis visualized on a GelDoc EZ imager (Bio-Rad).
[0038] FIG. 15 shows the further purification of the assembled chimera barcodes by HPLC chromatogram and the fractions targeting three-codon barcodes were collected for the nanopore experiment.
[0039] FIG. 16 shows that the three different three-codon chimera Type II-A barcodes exhibited three highly distinguishable nanopore signal patterns using a single channel MspA nanopore system.
[0040] FIG. 17 shows that the three different three-codon Type II-B chimera barcodes exhibited three distinguishable nanopore signal patterns using a MspA nanopore.088177-8007W001DETAILED DESCRIPTION OF THE INVENTION
[0041] Before the present disclosure is described in greater detail, it is to be understood that this disclosure is not limited to particular embodiments described, and as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the present disclosure will be limited only by the appended claims.
[0042] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present disclosure, the preferred methods and materials are now described.
[0043] All publications and patents cited in this specification are herein incorporated by reference as if each individual publication or patent were specifically and individually indicated to be incorporated by reference and are incorporated herein by reference to disclose and describe the methods and / or materials in connection with which the publications are cited. The citation of any publication is for its disclosure prior to the filing date and should not be construed as an admission that the present disclosure is not entitled to antedate such publication by virtue of prior disclosure. Further, the dates of publication provided could be different from the actual publication dates that may need to be independently confirmed.
[0044] As will be apparent to those of skill in the art upon reading this disclosure, each of the individual embodiments described and illustrated herein has discrete components and features which may be readily separated from or combined with the features of any of the other several embodiments without departing from the scope or spirit of the present disclosure. Any recited method can be carried out in the order of events recited or in any other order that is logically possible.
[0045] Definitions
[0046] The following definitions are provided to assist the reader. Unless otherwise defined, all terms of art, notations and other scientific or medical terms or terminology used herein are intended to have the meanings commonly understood by those of skill in the chemical and medical arts. In some cases, terms with commonly understood meanings are defined herein for clarity and / or for ready reference, and the inclusion of such definitions herein should not necessarily be construed to represent a substantial difference over the definition of the term as generally understood in the art.088177-8007W001
[0047] As used herein, the term “or” is an inclusive “or” operator and is equivalent to the term “and / or”, unless the context clearly dictates otherwise. The term “based on” is not exclusive and allowed for being based on additional factors not described unless the context clearly dictates otherwise. In addition, the singular forms “a,” “an” and “the” include plural references unless the content clearly dictates otherwise. The meaning of “in . . .” includes “within . . .” and “on . . .”.
[0048] As used herein, the term “amino acid” generally refers to an organic compound that combines to form a protein or peptide. An amino acid generally comprises an amine group, a carboxylic acid group, and a side-chain specific to each amino acid, which serve as a monomeric subunit of a peptide. An amino acid may include the 20 standard, naturally occurring or canonical amino acids as well as non-standard amino acids. The standard, naturally-occurring or canonical amino acids include Alanine (A or Ala), Cysteine (C or Cys), Aspartic Acid (D or Asp), Glutamic Acid (E or Glu), Phenylalanine (F or Phe), Glycine (G or Gly), Histidine (H or His), Isoleucine (I or He), Lysine (K or Lys), Leucine (L or Leu), Methionine (M or Met), Asparagine (N or Asn), Proline (P or Pro), Glutamine (Q or Gin), Arginine (R or Arg), Serine (S or Ser), Threonine (T or Thr), Valine (V or Vai), Tryptophan (W or Trp), and Tyrosine (Y or Tyr). An amino acid may be an L-amino acid or a D-amino acid. Non-standard amino acids may be modified amino acids, amino acid analogs, amino acid mimetics, non-standard proteinogenic amino acids, or non-proteinogenic amino acids that occur naturally or are chemically synthesized. Examples of nonstandard amino acids include, but are not limited to, selenocysteine, pyrrolysine, and N- formylmethionine, (3 -amino acids, Homo-amino acids, Proline and Pyruvic acid derivatives, 3- substituted alanine derivatives, glycine derivatives, ring-substituted phenylalanine and tyrosine derivatives, linear core amino acids, and N-methyl amino acids. Moreover, nonstandard amino acids include chemical modifications appear in the post-translational modifications (PTMs), which refer to any alteration in the amino acid sequence of the protein after its synthesis. The modifications include, but not limited to, phosphorylation, glycosylation, ubiquitination, S-nitrosylation, methylation, N-acetylation, lipidation and proteolysis. The amino acid can also be a derivative of natural and non-natural amino acids resulting from a chemical process such as edman degradation (e.g. PTH amino acid).
[0049] As used herein, the terms “antibody” and “immunoglobulin” may generally refer to proteins that can recognize and bind to a specific antigen. An antibody or immunoglobulin may refer to an antibody isotype, fragments of antibodies including, but not limited to, Fab, Fv, scFv, and Fd fragments, chimeric antibodies, humanized antibodies, single-chain antibodies, and088177-8007W001 fusion proteins including an antigen-binding portion of an antibody and a non-antibody protein. The antibodies may be detectably labeled, e.g., with a fluorophore, radioisotope, enzyme (e.g., a peroxidase) which generates a detectable product, fluorescent protein, nucleic acid barcode sequence, and the like. The antibodies may be further conjugated to other moieties, such as members of specific binding pairs, e.g., biotin (member of biotin-avidin specific binding pair), and the like. Also encompassed by the terms are Fab', Fv, F(ab')2, and other antibody fragments that retain specific binding to antigen. Antibodies may exist in a variety of other forms including, for example, Fv, Fab, and (Fab)2, as well as bi-functional (i.e., bi-specific) hybrid antibodies (e.g., Lanzavecchia et al., Eur. J. Immunol. 17, 105 (1987)) and in single chains (e.g., Huston et al., Proc. Natl. Acad. Sci. U.S.A., 85, 5879-5883 (1988) and Bird et al., Science, 242, 423-426 (1988), which are incorporated herein by reference). (See, generally, Hood et al., Immunology, Benjamin, N.Y., 2nd ed. (1984), and Hunkapiller and Hood, Nature, 323, 15-16 (1986), which are herein incorporated by reference).
[0050] As used herein, the term “binder” or “binding agent” refers to a molecule, e.g., a nucleic acid molecule, a peptide, a polypeptide, a protein, carbohydrate, a synthetic molecule, or a small molecule that binds to, associates with, unites with, recognizes, or combines with another molecule. The binding agent may bind to a macromolecule or a component or feature of a macromolecule. A binding agent may form a covalent association or non-covalent association with a molecule, a macromolecule, or a component or feature of a macromolecule. A binding agent may also be a chimeric binding agent, composed of two or more types of molecules, such as a nucleic acid molecule-peptide chimeric binding agent, a carbohydratepeptide chimeric binding agent, or a lipid-peptide chimeric binding agent. A binding agent may be a naturally occurring, synthetically produced, or recombinantly expressed molecule. A binding agent may bind to a single monomer or subunit of a macromolecule (e.g., a single amino acid of a peptide) or bind to a plurality of linked subunits of a macromolecule (e.g., a di-peptide, tri-peptide, or higher order peptide of a longer peptide, polypeptide, or protein molecule). A binding agent may bind to a linear molecule or a molecule having a three- dimensional structure (also referred to as conformation). For example, an antibody binding agent may bind to linear peptide, polypeptide, or protein, or bind to a conformational peptide, polypeptide, or protein. A binding agent may bind to an N-terminal peptide, a C-terminal peptide, or an intervening peptide of a peptide, polypeptide, or protein molecule. A binding agent may bind to an N-terminal amino acid, C-terminal amino acid, or an intervening amino acid of a peptide molecule. A binding agent may preferably bind to a chemically modified or labeled amino acid over a non -modified or unlabeled amino acid. For example, a binding agent088177-8007W001 may preferably bind to an amino acid that has been modified with an acetyl moiety, guanyl moiety, dansyl moiety, PTC moiety, DNP moiety, SNP moiety, etc., over an amino acid that does not possess such a moiety. A binding agent may bind to a post-translational modification of a peptide molecule. A binding agent may exhibit selective binding to a component or feature of a macromolecule (e.g., a binding agent may selectively bind to one of the 20 possible natural amino acid residues and bind with very low affinity or not at all to the other 19 natural amino acid residues). A binding agent may exhibit less selective binding, where the binding agent is capable of binding a plurality of components or features of a macromolecule (e.g., a binding agent may bind with similar affinity to two or more different amino acid residues). A binding agent may comprise a tag, which may be coupled to the binding agent via a linker.
[0051] As used herein, the term “barcode” generally refers to an identifying feature that may be used to distinguish similar items. A barcode may comprise a nucleic acid molecule of about 2 to about 30 bases. A barcode may comprise a nucleic acid molecule of about 2, 3, 4, 5,6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31,32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56,57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81,82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150 or more bases, which may provide a unique identifier tag or origin information for a molecule (e.g., protein, polypeptide, peptide), a binding agent, a set of binding agents from a binding cycle, a sample molecule, a set of samples, molecules within a compartment (e.g., droplet, bead, partition or separated location), macromolecules within a set of compartments, a fraction of macromolecules, a set of macromolecule fractions, a spatial region or set of spatial regions, a library of macromolecules, or a library of binding agents. A barcode can be an artificial sequence or a naturally occurring sequence including peptides, proteins, protein complexes, carbohydrates, and synthetic polymeric materials. In certain embodiments, each barcode within a population of barcodes is different. In other embodiments, a portion of barcodes in a population of barcodes is different, e.g., at least about 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, or 99% of the barcodes in a population of barcodes is different. A population of barcodes may be randomly generated or non-randomly generated. A population of barcodes may comprise error-correcting barcodes. Barcodes can be used to computationally deconvolute sequence reads derived from an individual molecule, sample, library, etc. Barcodes may comprise multiplexed information, e.g., arising from different samples, compartments, individual molecules, etc. A barcode can also be used for deconvolution of a collection of088177-8007W001 molecules that have been distributed into small compartments for enhanced mapping. For example, rather than mapping a peptide back to the proteome, the peptide can be mapped back to its originating protein molecule or protein complex, a sample or partition from which it originated, etc. A barcode may comprise any useful sequence, including repeat sequences (e.g., a poly -A, poly-T, poly-C, poly-G region) or the barcode may comprise non-repeat sequences. As used herein, a “sample barcode”, also referred to as “sample tag” generally refers to a barcode molecule comprising identifying information of a sample from which a barcoded molecule derives.
[0052] As used herein, the term “hybridization” is used in reference to the pairing of complementary nucleic acids. Hybridization and the strength of hybridization (e.g., the strength of the association between the nucleic acids) is influenced by such factors as the degree of complementary between the nucleic acids, stringency of the conditions involved, and the melting temperature of the formed hybrid. Hybridization methods involve the annealing of one nucleic acid to another, complementary nucleic acid, e.g., based on Watson-Crick base pairing.
[0053] The terms “nucleic acid”, “nucleic acid molecule”, “oligonucleotide” and “polynucleotide” may be used interchangeably herein and generally refer to a polymeric form of naturally occurring or synthetic nucleotides, or analogs thereof, of any length. A nucleic acid molecule may comprise one or more deoxyribonucleotides, deoxynucleotide triphosphates, dideoxynucleotide triphosphates, ribonucleotides, hexitol nucleotides, cyclohexane nucleotides, or analogs or combinations thereof. A nucleic acid molecule may comprise, e.g., DNA, RNA, HNA, CeNA, and modified forms thereof. A nucleic acid molecule may comprise nucleotides that are linked by phosphodiester bonds. A nucleic acid molecule may have any two- or three- dimensional structure, and may perform any function, known or unknown. A nucleic acid molecule may be single stranded, double stranded, or partially double stranded. Non-limiting examples of polynucleotides include a gene, a gene fragment, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, noncoding RNA, small interfering RNA, short hairpin RNA, micro RNA, scaRNA, ribozymes, riboswitches, viral RNA, complementary DNA (cDNA), cosmid DNA, mitochondrial DNA, chromosomal or genomic DNA, viral DNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, control regions, isolated RNA of any sequence, nucleic acid probes, nucleic acid adapters, and primers. The nucleic acid molecule may be linear, circular, or any other geometry. Examples of polynucleotide analogs include but are not limited to xeno nucleic acid (XNA), bridged nucleic acid (BNA), glycol nucleic acid (GNA), hexitol nucleic acid (HNA), cyclohexane nucleic acid (CeNA), peptide nucleic acids (PNAs), yPNAs,088177-8007W001 morpholino polynucleotides, locked nucleic acids (LNAs), threose nucleic acid (TNA), 2'-O- Methyl polynucleotides, 2'-O-alkyl ribosyl substituted polynucleotides, phosphorothioate polynucleotides, and boronophosphate polynucleotides. A polynucleotide analog may possess purine or pyrimidine analogs, including for example, 7-deaza purine analogs, 8-hal opurine analogs, 5-halopyrimidine analogs, or universal base analogs that can pair with any base, including hypoxanthine, nitroazoles, isocarbostyril analogues, azole carboxamides, and aromatic triazole analogues, or base analogs with additional functionality, such as a biotin moiety for affinity binding.
[0054] As used herein, the term “peptide” may refer to any short, single peptide chain. A peptide may be no more than about 100, 95, 90, 85, 80, 75, 70, 65, 60, 55, 50, 45, 40, 35, 30, 25, 20, 15, 10, 5, or less than about 5 amino acids in length. A peptide may have a known or unknown biological function or activity. Peptides can include natural, synthetic, modified, or degraded proteins or peptides, or a combination thereof.
[0055] As used herein, “polypeptide” refers to two or more amino acids linked together by a peptide bond. The term “polypeptide” includes proteins that have a C-terminal end and an N- terminal end as generally known in the art and may be synthetic in origin or naturally occurring. As used herein “at least a portion of the polypeptide” refers to 2 or more amino acids of the polypeptide. A polypeptide may comprise one or more peptides. Optionally, a portion of the polypeptide includes atleast: 1, 5, 10, 20, 30 or 50 amino acids, either consecutive or with gaps, of the complete amino acid sequence of the polypeptide, or the full amino acid sequence of the polypeptide.
[0056] As used herein, “sequencing” generally refers to determining the order of: (A) nucleotides (base sequences) in a nucleic acid sample, e.g., DNA or RNA; or determining the order of (B) amino acids in all or part of a polymer, such as a protein, peptide, or other multimeric molecule. Many techniques are available, such as Sanger sequencing or High Throughput Sequencing technologies (HTS). Sanger sequencing may involve sequencing via detection through (capillary) electrophoresis, in which up to 384 capillaries may be sequence analyzed in one run. High throughput sequencing involves the parallel sequencing of thousands or millions or more sequences at once. HTS can be defined as Next Generation sequencing (NGS), i.e. techniques based on solid phase pyrosequencing or as Next-Next Generation sequencing based on single nucleotide real time sequencing (SMRT). HTS technologies are available such as offered by Roche, Illumina and Applied Biosystems (Life Technologies). Further high throughput sequencing technologies are described by and / or available from088177-8007W001Helicos, Pacific Biosciences, Complete Genomics, Ion Torrent Systems, Oxford Nanopore Technologies, Nabsys, ZS Genetics, GnuBio.
[0057] The term “sequentially” as used herein means that the components coupled a polymerizable chain are juxtaposed from one end of the polymerizable chain to the other end. For one example, in a polymerizable chain-amino acid (PCAA) complex, when the polymerizable chain is a nucleic acid chain, the amino acids coupled sequentially to the polymerizable chain means that the amino acids are juxtaposed along the nucleic acid chain from one end of the nucleic acid chain to the other, e.g., in a 5’ end to 3’ end, or 3’ to 5’ order.
[0058] As used herein, “next generation sequencing (NGS)” refers to high-throughput sequencing methods that allow the sequencing of millions to billions of molecules in parallel. Examples of next generation sequencing methods include sequencing by synthesis, sequencing by ligation, sequencing by hybridization, polony sequencing, ion semiconductor sequencing, nanopore sequencing, and pyrosequencing. By attaching primers to a solid substrate and a complementary sequence to a nucleic acid molecule, a nucleic acid molecule can be hybridized to the solid substrate via the primer and then multiple copies can be generated in a discrete area on the solid substrate by using polymerase to amplify (these groupings are sometimes referred to as polymerase colonies or polonies). Consequently, during the sequencing process, a nucleotide at a particular position can be sequenced multiple times (e.g., hundreds or thousands of times) — this depth of coverage is referred to as “deep sequencing.” Examples of high throughput nucleic acid sequencing technology include platforms provided by Illumina, BGI, Qiagen, ThermoFisher, and Roche, including formats such as parallel bead arrays, sequencing by synthesis, sequencing by ligation, capillary electrophoresis, electronic microchips, “biochips,” microarrays, parallel microchips, and single-molecule arrays, as reviewed by Service (Science 311 : 1544-1546, 2006).
[0059] As used herein, “analyzing” the macromolecule means to quantify, characterize, distinguish, or a combination thereof, all or a portion of the components of a molecule (e.g., a macromolecule, a biological molecule such as a protein, amino acid, nucleic acid molecule, etc.). For example, analyzing a peptide, polypeptide, or protein may comprise determining all or a portion of the amino acid sequence (contiguous or non-continuous) of the peptide. Analyzing a macromolecule may include partial identification of a component of the macromolecule. For example, partial identification of amino acids in a protein sequence can identify an amino acid in the protein as belonging to a subset of possible amino acids. Analysis may be performed sequentially, e.g., beginning with analysis of the n NTAA (N-terminal amino acid), and then proceeding to the next amino acid of the peptide (i.e., n-1, n-2, n-3, and so forth).088177-8007W001In such instances, sequencing may be performed by cleavage of the n NTAA, thereby converting the n-1 amino acid of the peptide to an N-terminal amino acid (referred to herein as the “n-1 NTAA”). Similarly, analysis of a peptide may begin from C-terminus towards the N- terminus with each round of cleavage from the C- terminus creating a new CTAA (C-terminal amino acid). Cleavage of the n CTAA converts the n-1 amino acid of the peptide to a C-terminal amino acid, referred to herein as an “n-1 CTAA”. Analyzing the peptide may also include determining a presence and frequency of post-translational modifications on the peptide, which may or may not include information regarding the sequential order of the post- translational modifications on the peptide. Analyzing the peptide may also include determining the presence and frequency of epitopes in the peptide, which may or may not include information regarding the sequential order or location of the epitopes within the peptide. Analyzing the peptide may include combining different types of analysis, for example obtaining epitope information, amino acid sequence information, post-translational modification information, or any combination thereof.
[0060] The term “sample” as used herein refers to a collected substance or material that comprises or is suspected to comprise one or more analytes of interest (e.g., biomolecules, e.g., polypeptides). A sample may be modified from its collection for purposes such as storage or stability. A sample may be naturally occurring or synthetic. A sample may be processed to separate or remove unwanted fractions or impurities from the analyte(s) of interest. A sample may be enriched or purified. For example, a sample may comprise a fraction of a separation process (e.g., chromatography, fractionation, electrophoresis, etc.). Alternatively, a sample may not be subjected to processing that separates or removes any unwanted fractions or impurities from the analyte(s) of interest. A sample may be obtained from any suitable source or location, including from organisms, cells, tissues, cell preparations, cell-free compositions, the environment (e.g., air, water, dirt, soil, agriculture, soil, dust). A sample may be obtained from an organism or part of an organism, such as from a fluid, tissue, or cell. A sample may include biological and / or non- biological sample. As used herein, the term “biological sample” refers to a sample that is derived from a predominantly biological system or organism, such as one or more viral particles, cells (e.g. individualized cells), organelles (e.g. individualized organelles), tissues, bodily fluids, bone, cartilage, and exoskeleton. A biological sample may comprise a majority of biological material on a mass basis, excluding the weight of fluid within the sample. Biological samples may comprise one or more proteins, referred to herein as protein samples. Biological samples can be acquired from various sources, e.g., from a clinical patient sample, such as blood, serum, plasma, Cerebral Spinal Fluid (CSF), saliva, mucosal secretions,088177-8007W001 urine, lymph, perspiration, vaginal fluid, semen, etc. A biological sample may be processed to purify and retain one or more biomolecules (e.g., proteins, nucleic acids, carbohydrates, lipids, glycoproteins, lipoproteins, metabolites, etc.) from the biological sample. A biological sample (e.g., a protein sample) may be derived from cultured cells, which may be treated or untreated. A biological sample (e.g., a protein sample) can also result from tissue specimens, such as biopsy samples, which may optionally be processed to liberate biomolecules (e.g., proteins) contained therein. Tissue samples may also be derived from in vivo specimens, including fresh, frozen, acute, and fixed tissues.
[0061] Barcode
[0062] In one aspect, the method and composition described herein involves a barcode used to label a polypeptide before subjecting the polypeptide or its derivative, e.g., a polymerizable chain-amino acid (PCAA) derived from the polypeptide (see below for detailed description) to a nanopore device for sequencing.
[0063] As used herein, the term “barcode” refers to a molecule providing a sample, e.g., polypeptide, nucleic acid, etc., with a unique identifier or information to trace the origin of the sample. As used herein, a barcode can comprise a polymerizable chain, such as a nucleic acid or PEG. In some embodiments, the nucleic acid comprised in the barcode can be a DNA, a RNA or a modified variant thereof.
[0064] In some embodiments, a barcode may further comprise a unique identifier tag or origin information for a molecule (e.g., protein, polypeptide, peptide), a binding agent, a set of binding agents from a binding cycle, a sample molecule, a set of samples, molecules within a compartment (e.g., droplet, bead, partition or separated location), macromolecules within a set of compartments, a fraction of macromolecules, a set of macromolecule fractions, a spatial region or set of spatial regions, a library of macromolecules, or a library of binding agents. A barcode can be an artificial sequence or a naturally occurring sequence including peptides, proteins, protein complexes, carbohydrates, and synthetic polymeric materials. In certain embodiments, each barcode within a population of barcodes is different. In other embodiments, a portion of barcodes in a population of barcodes is different, e.g., at least about 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, or 99% of the barcodes in a population of barcodes is different. A population of barcodes may be randomly generated or non-randomly generated. A population of barcodes may comprise error correcting barcodes. Barcodes can be used to computationally deconvolute sequence reads derived from an individual molecule, sample, library, etc. Barcodes may comprise multiplexed information, e.g., arising from different samples, compartments,088177-8007W001 individual molecules, etc. A barcode can also be used for deconvolution of a collection of molecules that have been distributed into small compartments for enhanced mapping. For example, rather than mapping a peptide back to the proteome, the peptide can be mapped back to its originating protein molecule or protein complex, a sample or partition from which it originated, etc. A barcode may comprise any useful sequence, including repeat sequences (e.g., a poly -A, poly-T, poly-C, poly-G region) or the barcode may comprise non-repeat sequences. As used herein, a “sample barcode”, also referred to as “sample tag” generally refers to a barcode molecule comprising identifying information of a sample from which a barcoded molecule derives.
[0065] In some embodiments, the methods in the present disclosure does not have to identify the specific sequence information (i.e. the specific nucleic acid or amino acid information) of either the barcode moiety or the barcode. In some embodiments, the methods in the present disclosure only need to identify the signal signatures of the barcode moiety or the barcode. The signal signatures have been disclosed in US Provisional Application No. 63 / 645,885 and PCT application PCT / US2025 / 028808, which are herein incorporated by reference. To be specific, the signal signature may be measured from the nanopore, membrane, or surrounding solution. For example, conductance, current, current blockage, or other parameter within the nanopore may be monitored as a function of time. As a molecule enters the nanopore, nanogap, or nanochannel, a change in the conductance, current, or other parameter may occur and provide information (e.g., size, charge, aspect ratio, volume) on the molecule. Each amino acid or modified amino acid, or a subset of amino acids or modified amino acids, may generate a unique signal that is distinguishable from other amino acids or modified amino acids. As such, the unique signal signatures may be assigned to the amino acids or modified amino acids in order to determine the identity of the amino acids or modified amino acids.
[0066] In some embodiments, there are some benefits of mixing amino acid codons with nucleic acid codons in a barcode. In some embodiments, it creates a larger combination of possible barcodes which increases the throughput capacity.
[0067] In some embodiments, the barcode described herein contains at least one barcode moiety, which is a basic unit that can be recognized and distinguished by a detector, e.g. a nanopore device. In some embodiments, a barcode moiety comprises a codon and a timecontrol element. In some embodiments, a barcode comprises a concatenation of barcode moieties linearly.
[0068] Codons088177-8007W001
[0069] As used herein, a “codon” refers to a sequence or a chemical structure that is presented to a detector to generate a unique signal. As described herein, a “codon” comprises any compound that can be recognized by a detector, e.g., a nanopore device. In some embodiments, the codon in a barcode moiety described herein comprises polynucleotides (k- mer). In some embodiments, the codon described herein comprises one or more amino acids or amino acid derivatives. In some embodiments, the amino acid derivatives comprised in a codon are phenylthiohydantoin amino acids (PTH-AA). In some embodiments, the codon described herein comprises a chimera of polynucleotide and amino acid (or amino acid derivative). In some embodiments, the codon described herein comprises a compound such as bicyclo[6.1.0]nonyne (BCN), a crown ether molecule, or a sugar molecule (e.g., a sialic acid). In some embodiments, the arrangement and combination of these nucleotides, amino acids, peptides and / or chemical compound can be recognized and distinguished by a detector, e.g., a nanopore device.
[0070] In some embodiments, a codon is designed to be recognized as an integral unit of a barcode moiety when it is associated with a time-control element (as described in detail below). In some embodiments, this conditional recognition mechanism ensures that the codon's information is accessed or activated under specific temporal and identification parameters. Coupling the codon and the time-control element in the barcode moiety can facilitate a precise and controlled recognition event.
[0071] In some embodiments, a codon comprises a nucleotide sequence ranging from 1 to 50 nucleic acids, optionally ranging from 3 to 20 nucleic acids, and the nucleotide may have one or more chemical modifications including methylation, phosphorylation, or specific amino acid conjugation. The combination of these elements forms a unique signature that is recognized holistically when the time-control element (e.g., a unique DNA / RNA sequence) enables its interaction with blocker (e.g., a complementary DNA / RNA oligo). This system ensures high specificity and temporal control over the codon's function.
[0072] In some embodiments, the codons each include 1 or more unpaired polynucleotides. In some embodiments, the codons each include 2 or more unpaired polynucleotides. In some embodiments, the codons each include 1 to 5, 1 to 4, 1 to 3, 2 to 5, 2 to 4, 2 to 3, 3 to 5, or 3 to 4 unpaired polynucleotides. In some embodiments, the codons each include 1, 2, 3, 4, 5 polynucleotides, or combinations thereof. In some embodiments, the codons each include up to 10 polynucleotides. In some embodiments, the nanopore conductance was found to be sensitive to the first 4 unpaired nucleotides of the nucleic acid chain upstream of the blocker. Therefore, particularly suitable codons each include 4 unpaired polynucleotides.088177-8007W001
[0073] In some embodiments, the number of the amino acids within a codon can be one or more amino acid. In some embodiments, the number of the amino acids within a codon can be one or two amino acids.
[0074] Barcode Moiety
[0075] As described in the present disclosure, two types of barcodes are provided for nanopore peptide sequencing, which are called Type I barcode (as shown in FIG. 1), and Type II barcode (as shown in FIG. 2), which can be Type II-A and Type II-B barcodes.
[0076] To be specific, Type I barcode comprises nucleic acid but no amino acid. Type I barcode may comprise a nucleic acid molecule of more than 1 bases, optionally about 1 to about 50 bases. In some embodiments, Type I barcode has a k-mer DNA section and the k-mer DNA section provides a codon as a single identifier. The number of k is an integer between 1 and 50, optionally between 3 and 10, such as 3, 4, 5, 6, 7, 8, 9, 10.
[0077] In some embodiments, a Type I barcode consists of a series of nucleic acid codons. In some embodiments, a Type I barcode consists of a series of natural nucleic acid codons. In some embodiments, a barcode comprising 3-mer, 4-mer or 5-mer, which belongs to Type I barcode, is regarded as one codon.
[0078] Besides, Type II barcode can be regarded as a chimera barcode where a chemical structure (such as an amino acid) is conjugated to the nucleic acid backbone; the conjugated structure generates a unique signal for barcoding. The chimera barcode can use AA-readout only as an identifier (Type A). It can also include a k-mer DNA readout as a secondary identifier and combine both signals for barcoding a single sample (Type B).
[0079] In some embodiments, Type II barcode contains amino acid codons only, or it contains chimera codons with both amino acid and nucleic acid. In some embodiments, Type II barcode contains modified codons, including chemical conjugation or base / ribose removal.
[0080] In some embodiments, a barcode can have at least 1 barcode moiety. In some embodiments, a barcode can have more than 1, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 barcode moi eties, if necessary.
[0081] Time-control elements
[0082] In the present disclosure, a time-control element is a sequence or a chemical structure in the barcode moiety that contributes to form a structure that holds the codon in the nanopore for a time duration. The structure dissembles upon applying additional force or by an applied voltage and let the polymer chain translocate through nanopore. In some embodiments, the additional force is a time-varying waveform.088177-8007W001
[0083] In some embodiments, the structure is a hairpin structure, duplex DNA or triplex DNA.
[0084] In the present disclosure, the time-control element described herein may have various forms. In some embodiments, the time-control elements described herein is a section of PCAA complex capable of holding the PCAA complex from translocation in the nano-scale recording device that has been disclosed in US Provisional Application No. 63 / 645,885 and PCT application PCT / US2025 / 028808, which are herein incorporated by reference.
[0085] In some embodiments, the time-control element described herein forms a structure capable of temporarily holding the first codon at the nanopore sensor. FIG. 3 illustrates an embodiment of the time-control element that is a structure formed by a nucleic acid sequence contained in the barcode. In the schematic workflow 600, a PCAA complex 601 coupled to a barcode is loaded at a nanopore device 602, wherein a codon 603 of the PCAA complex 601 is positioned at the nanopore channel 604 of the nanopore. A time control element 605 forms a structure that prevents the barcode from translocation. In particular, the time-control element 605 forms a structure following the codon 603 stalls the codon 603 at the nanopore channel 604. The specific features of current pattern can be read out during the long-time stalling, including the duration, current level, noise level, transition of multiple status and so on. In the process 606, the time-control element 605 linearizes, spontaneously, controlled by programmable voltage or toehold displacement, which allows the codon 603 to translocate through the channel 604. The nanopore device 602 then reads a change of the current pattern caused by such translocation, which reflects the identity of the codon 603. The next timecontrol element 607 then stalls the codon 608 at the channel 604, preventing the nanopore device 602 from translocating the codon 608. The sequential dissociation of time-control element generates a stepwise current trace with all the information of the codons in order. Combining with the method of generating the PCAA complex 601 from a polypeptide as described elsewhere in the present disclosure, the amino acids on the PCAA complex 601 can be deciphered by the nanopore device 602 to decode the barcode and determine the sequence of a target polypeptide.
[0086] In some embodiments, the time-control element described herein is designed to non-covalently bind a reagent to form a complex or structure, such as duplex, triplex, pseudoknot, G-quadruplex, or DNA / RNA origami, to temporarily hold the codon at the nanopore sensor. In some embodiments, the reagent can be DNA, RNA or artificial nucleic acid homologs. In some embodiments, the reagent can be a nucleic acid binding complex, or a nanoparticle. In some embodiments, binding of the time-control element to the reagent is088177-8007W001 mediated by a molecule selected from a mercury ion, a silver ion, anthracycline and a peptide. In some embodiments, the time-control element is similar to a blocker structure or complex contained in the PCAA complex. Both time-control element and the blocker can prevent the PCAA complex from translocation, e.g., stalling the PCAA complex at the channel of nanopore protein. Meanwhile, the codon or the amino acid (or amino acid derivative) is held fixed at the constriction site of the nanopore. The specific features of the current pattern can be read out during the long-time stalling, including the duration, current level, noise level, transition between multiple states and so on. After stalling, the time-control element or blocker can change the structure or dissociate from the binding reagent, controlled by programmable voltage or toehold displacement, to allow the translocation of the codon or the amino acid on the PCAA complex. And the next time-control element or blocker may stall the barcode or PCAA complex for the reading of next codon or amino acid. In some embodiments, the stability of the association between the time-control element and the binding reagent can be modulated by the introduction of artificial nucleic acid homologs, such as PNA, LNA, BNA, and XNA. The interaction between the time-control element and the binding reagent can also be modulated by nucleic acid binding chemicals, such as mercury ion, silver ion, anthracycline, and peptide. The sequential dissociation of the binding reagent from the barcode generates a stepwise current trace with all the information of the codons in order. Thus, the decoding of codon on a barcode by nanopore can fulfill the identification of the barcode.
[0087] In some embodiments, the length of the time-control element has a polynucleotide sequence between 1 and 200 nucleotides, preferably between 2 and 100 nucleotides, and more preferably between 5 and 30 nucleotides. In some embodiments, the time-control element each comprise about 5 to about 30, about 5 to about 25, about 5 to about 20, about 10 to about 30, about 10 to about 25, about 10 to about 20, about 15 to about 30, about 15 to about 25, or about 15 to about 20 nucleotides. In some embodiments, the time-control element each comprise 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 nucleotides or combinations thereof. In some embodiments, the sequence of the time-control element is complementary to a reagent of polynucleotide. In some embodiments, the time-control element does not haveto be 100% complementary, such as 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80% or 90% complementary, as long as the time-control element can bind to the binding reagent sequence.
[0088] FIG. 4 illustrates an embodiment of the time-control element that non-covalently binds a reagent to form a complex capable of temporarily holding the codon at the nanopore sensor. In the schematic workflow 400, a PCAA complex 401 is loaded at a nanopore device088177-8007W001402, wherein a codon 403 of a barcode coupled to the PCAA complex 401 is positioned at the nanopore channel 404 of the nanopore. A time-control element following the codon 403 binds to a reagent 405 and stalls the codon 403 at the nanopore channel 404. The specific features of current pattern can be read out during the long-time stalling, including the duration, current level, noise level, transition between multiple states and so on. In the process 406, the reagent 405 dissociates from the time-control element, spontaneously, controlled by programmable voltage or toehold displacement, which allows the codon 403 to translocate through the channel 404. The nanopore device 402 then reads a change of the current pattern caused by such translocation, which reflects the identity of the codon 403. The next time-control element binding to a reagent 407 then stalls the codon 408 at the channel 404, preventing the nanopore device 402 from translocating the codon 408. The sequential dissociation of the reagents from the time-control elements generates a stepwise current trace with all the information of codons from a barcode in order. Combining with the method of generating the PCAA complex 401 from a polypeptide as described elsewhere in the present disclosure, the barcode and the amino acids on the PCAA 401 can be inferred by the nanopore device 402 to decipher the barcode and sequence the target polypeptide.
[0089] In some embodiments, the dissociation of the time-control element and the reagent can be controlled by the voltage of the nano-scale device, thus controlling the duration of data acquisition in each step. In some embodiments, the binding reagents may exist on both sides of the nanopore, and the voltage of the nanopore device may be controlled to sequence the barcode and PCAA complex in the nanopore multiple times, until a special voltage is applied to release the PCAA complex from the nanopore.
[0090] In the present disclosure, the time-control element’s role is a molecular function that enables to control or modulate the translocation behaviors (such as physical or chemical behaviors) of analytes through a nanopore. Such a function can manifest the slow-down of the translocation rate to lead their extended residence time in the nanopore. The extended time will provide better chance to design chemistry between an analyte and the internal side of the nanopore, which can produce any readout to distinguish one from others.
[0091] In some embodiments, while ssDNA can pass through a nanopore beyond the sampling speed, dsDNA cannot translocate through some nanopores so the translocation of ssDNA can be stalled by the dsDNA that is created by the hybridization in a part of the ssDNA. Such hybridization can happen either with exogenous complementary ssDNA or the internal complementary sequence. In either case, the hybridization sequence is a time-control element,088177-8007W001 which still can get affected by the concentration of salt and chemical components of a solution along with an applied voltage.
[0092] Barcoded Sequencing Method
[0093] High-throughput sequencing technologies, enabling millions of DNA molecules to be sequenced at a time, have revolutionized the studies on genomics, epigenomics, and transcriptomics. On the other hand, although the protein information in a cell or organism can to some degree be deciphered through the transcriptome, i.e., the mRNAs expressed by the cell or organism, protein complexity surpasses that of the transcriptome, mainly due to post- translational modifications (PTMs) and dynamic changes in response to various factors. Therefore, there is a need to analyze the protein information, e.g., proteomics, of a cell or organism directly.
[0094] Despite their information richness, proteomic data remain largely unexplored compared to genomics. While genomics has benefited from next-generation sequencing (NGS), which enables massive DNA sequence analysis, proteomics lags behind due to limited throughput in current technologies. In particular, multiplexed protein sequencing technology, which allows proteins from multiple samples to be sequenced at a time, has been an unmet need. One proposed approach to perform multiplexed protein sequencing is to use nucleic acid barcodes (see, e.g., W02024030919A1). However, it has been a challenge to sequence by using nucleic acid barcodes through amino acid decoder at a time in one assay.
[0095] Therefore, to overcome such challenges, the present disclosure in one aspect provides novel methods and systems to sequence a plurality of proteins using barcodes, e.g., chimeric barcodes comprising amino acid residues (or amino acids for short). FIG. 5 illustrates an exemplary embodiment of the methods described in the present disclosure. Referring to FIG. 5, in the step 101, the chimeric barcode 102 are provided, which comprises at least one amino acid residue 103 linked to a polymerizable chain 104. In FIG. 5, three amino acid residues (hereinafter referred to as amino acids, and can be used interchangeably) 103 are shown. It is understood that in some embodiments, the chimeric barcode can comprise only one amino acid residue 103. Also provided is a biological sample 105 comprises at least a polypeptide analyte 106. The chimeric barcode 102 is mixed with the biological sample 105, coupling to an end of the polypeptide analyte 106 and generating a barcoded polypeptide complex 107.
[0096] In the next step 108, the barcoded polypeptide complex 107 is combined with additional barcoded polypeptide complexes from other samples, generating a barcoded sample pool 109. Preferably, the chimeric barcodes used in barcoded polypeptide complexes from different samples are different from each other.088177-8007W001
[0097] In step 110, the barcoded polypeptide complexes in the barcoded sample pool 109 are converted to a pool of polymerizable chain-amino acid (PCAA) complexes. Each PCAA complex comprises a stacked polymerizable chain 111 and the amino acids dissembled from a polypeptide analyte. The stacked polymerizable chain 111 includes the polymerizable chain 104 stacked with additional polymerizable molecules. The dissembled amino acids are linked to the stacked polymerizable chain 111 sequentially along the first stacked polymerizable chain 111 according to the sequence of the amino acids in the polypeptide analyte. The chimeric barcode 102 is linked at one end of the stacked polymerizable chain 111. As the polymerizable chain 104 becomes a part of the first stacked polymerizable chain 111, the amino acids 103 are linked to the end of the stacked polymerizable chain 111.
[0098] In step 112, the pool of PCAA complexes is subject to a sequencing process to determine the sequence of amino acids on each PCAA complex. As the sequence of amino acids on the PCAA complex corresponds to the sequence of amino acids on the polypeptide analyte, the sequencing process also determines the sequences of the polypeptide analytes. Meanwhile, the amino acids 103 are also sequenced during the sequencing of the PCAA complex, the identification or sequencing of the amino acids 103 on each PCAA complex can be used to identify the originating sample from which the PCAA complex is generated.
[0099] Coupling Chimeric Barcode to Polypeptide Analyte
[0100] In certain embodiments, the methods and systems disclosed herein involve in the first step coupling a barcode to an end, e.g., C-terminal end, of a polypeptide analyte. In certain embodiments, the barcode comprises a polymerizable chain coupled with at least one barcode moiety, which is not a nucleic acid molecule.
[0101] Chimeric barcode moiety and chimeric barcode
[0102] As used herein, a chimeric barcode moiety (hereinafter referred to as barcode moiety, and can be used interchangeably) refers to a compound which provides an identifying feature that may be used to distinguish the molecule, preferably a compound capable of being recognized and distinguished by proper devices, such as a nanopore device. In some embodiments, the compound comprises one amino acid. In some embodiments, the amino acid residue is selected from 20 standard, naturally occurring or canonical amino acids. In some embodiments, the amino acid is a non-natural, non-standard or modified amino acid.
[0103] As used herein, a chimeric barcode comprises at least one chimeric barcode moiety. That is to say, the chimeric barcode may comprise one chimeric barcode moiety only, or it may comprise two or more chimeric barcode moieties.088177-8007W001
[0104] In some embodiments, the amino acid is coupled to the polymerizable chain via a linker. In some embodiments, the linker may comprise an amino acid-reactive moiety. The amino acid-reactive moiety of the linker may be any useful moiety that enables the reactive moiety to conjugate to and optionally cleave an amino acid. In some examples, the reactive moiety may comprise any primary amine or carboxylic group reactive group, including but not limited to isocyanates, acyl azides, NHS esters, sulfonyl chlorides, aldehydes, glyoxals, epoxides, oxiranes, carbonates, aryl halides, imidoesters, carbodiimides, anhydrides, phenyl esters, isothiocyanates (e.g., phenyl isothiocyanate, sodium isothiocyanate, ammonium isothiocyanates (e.g., tetrabutylammonium isothiocyanate, tetrabutylammonium isothiocyanate), diphenylphosphoryl isothiocyanate), acetyl chloride, cyanogen bromide, carboxypeptidases, azide, alkyne, DBCO, maleimide, succinimide, thiol-thiol disulfide bonds, tetrazine, TCO, vinyl, methylcyclopropene, acryloyl, allyl, among others. Additional examples of amino acid reactive groups are provided in U.S. Pat. Pub. No. 2020 / 0217853A1, which is incorporated by reference herein in its entirety.
[0105] In some instances, the linker comprises a second reactive group, e.g., a click chemistry moiety that can react with a polymerizable molecule. In some embodiments, the polymerizable molecule is selected from the group consisting of nucleic acid, peptide nucleic acid (PNA), and polyethylene glycol (PEG). The second reactive group may comprise any suitable bioorthogonal moieties, e.g., alkenes, alkynes (e.g., alkyne, cycloalkynes such as DBCO andBCN), azides, epoxides, amines, thiols, nitrones, isonitriles, isocyanides, aziridines, activated esters, and tetrazines, and combinations, variations, or derivatives thereof. The linker may be subjected to conditions sufficient to react the amino acid and / or the polymerizable molecule, e.g., provision of metal catalysts, appropriate solvents, pH, temperature, ionic concentration, or light / energy for any useful duration of time.
[0106] The linker may comprise any number of spacing moieties, e.g., alkyl chains, polymer spacers (e.g., PEG), nucleic acid or oligo spacers, or other useful spacing moieties which may be useful in modulating the size or molecular weight of the linker. For example, the linker may comprise at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, or a greater number of spacing moieties (e.g., hydrocarbon units, PEG units, nucleotides or spacer sequences etc.). The linker may comprise at most about 100, at most about 10, at most about 9, at most about 8, at most about 7, at most about 6, at most about 5, at most about 4, at most about 3, at most about 2, or at most 1 spacing moieties. The linker may comprise any useful number of functional groups, e.g., for attachment to multiple088177-8007W001 analytes. The linker may comprise at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, or a greater number of functional groups.
[0107] The linker may comprise any additional useful moieties. For example, the linker may comprise a releasable or cleavable moiety, which may facilitate removal of the amino acid-linker complex from the polymerizable molecule, or portion thereof. Such a releasable or cleavable moiety may comprise, for example, a disulfide bond, which may be releasable by contacting with a reducing agent (e.g., DTT, TCEP). In some examples, the linker may couple to the polymerizable molecule via the releasable or cleavable moiety, alternatively or in addition to the coupling via click chemistry moieties. As such, the coupling between the polymerizable molecule and the linker may be reversible. The linker may additionally comprise any number of spacing moieties, e.g., polymers (e.g., PEG, PVA, polyacrylamide), aminohexanoic acid, nucleic acids, alkyl chains, etc. Such spacing moieties may increase the distance between any other moieties of the linker, e.g., the amino acid-reactive group and the polymerizable molecule-reactive group.
[0108] Polymerizable Chain
[0109] As used herein, a polymerizable chain refers to a chain of polymerizable molecules. In some examples, the polymerizable chain can be any polymerizable nucleic acid molecules, e.g., DNA or RNA, which can comprise any useful number and type of nucleotides, e.g., including canonical and noncanonical bases, and the number of nucleotides may be modulated based on the intended purpose. For example, the length of the polymerizable chain may be modulated to alter a property of the analytes (e.g., volume, aspect ratio, charge, etc.) to which the polymerizable chain is linked. The chain of polymerizable molecule may additionally or alternatively comprise any useful functional sequences, including but not limited to barcode sequences or other identifying sequences, UMI sequences, enzyme recognition sites (e.g., transposition sites, restriction sites), spacer sequences, sequencing primer sequences, read sequences, or primer sequences. The chain of polymerizable molecule may comprise canonical bases, noncanonical bases, naturally occurring bases, synthetic bases, abasic sites, or a combination thereof.
[0110] In some examples, the polymerizable molecules can also be selected from other known molecules, such as polyethylene glycol (PEG), as long as it can form a polymerizable chain.
[0111] In some embodiments, the polymerizable chain is capable of coupling to a barcode moiety as well as to a terminal, e.g., C-terminal, amino acid of a polypeptide analyte. The coupling of the polymerizable chain to the barcode moiety has been disclosed elsewhere herein.088177-8007W001The coupling of the polymerizable chain to the terminal amino acid of the polypeptide analyte may comprise a covalent interaction or a noncovalent interaction. The coupling may occur by interaction of binding pairs, e.g., biotin and avidin (or streptavidin), antigen or epitope and antibody or antibody fragment, cyclodextrins and small hydrophobic molecules (e.g., alkanes, benzene, polycyclics), cucurbiturils and adamantaneammonium or trimethylammoniomethyl ferrocene, cyclophane (e.g., calixarenes, cavitands, pillararenes, tetralactams), etc.
[0112] In some instances, the polymerizable chain may be coupled to the polypeptide analyte without being coupled to a substrate.
[0113] Alternatively, the polymerizable chain is provided coupled to a substrate. In one example, the substrate comprises one or more identical capture moieties. In some instances, commercially available substrates, e.g., beads (e.g., DNA beads or barcoded beads), flow cells, or chips, e.g., Illumina® HiSeq, iSeq, MiniSeq, NextSeq, NovaSeq, etc. may be used as the substrates described herein. The polymerizable chain may be coupled to a substrate using any useful approach. In some instances, the polymerizable chain comprises a substrate-tethering group or linker or additional functional group. In some examples, the polymerizable chain comprises a nucleic acid molecule that comprises a substrate-tethering group, e.g., biotin, a click chemistry moiety such as an azide, that can couple to a substrate, e.g., a substrate comprising streptavidin or a complementary click chemistry moiety that can react with that of the substrate-tethering group.
[0114] In some instances, the capture moiety comprises a polymerizable molecule, such as nucleic acid molecules. The nucleic acid molecules may be coupled to one another, e.g., via complementary base pairing directly or via a splint molecule and optional ligation. The capture moiety can comprise any naturally occurring, non-naturally occurring or engineered nucleotide base. For example, the nucleic acid molecule may comprise a pseudo-complementary base, a bridged nucleic acid, a xenonucleic acid, a locked nucleic acid, a peptide nucleic acid (PNA), a gamma-PNA, a morpholino, etc., as is described elsewhere herein.
[0115] The polymerizable chain may comprise one or more functional sequences, including, but not limited to, a priming sequence, sequencing sequence (e.g., P5 or P7 sequence), sequencing read sequence (e.g., R1 or R2 sequence), a mosaic end sequence, a transposase recognition sequence, a cleavage site (e.g., restriction site), a UMI, a blocking group, a spacer sequence, a barcode sequence, or other functional sequence. In some instances, the capture moiety comprises a cleavable or releasable moiety, e.g., a restriction enzyme recognition site, an abasic site, a uracil which can be cleaved using USER® or uracil DNA088177-8007W001 glycosylase, a disulfide bond that can be releasable upon addition of a reducing agent, a photo- cleavable linker etc. In some instances, the polymerizable chain comprises a barcode sequence that comprises any useful information, e.g., the identity of the peptide that is to be analyzed, temporal information, spatial information, etc.
[0116] Forming Polymerizable Chain-Amino Acid (PCAA) Complex
[0117] In certain embodiments, the methods and systems disclosed herein involve in a step of converting a polypeptide analyte into a polymerizable chain-amino acid (PCAA) complex, which comprises a polymerizable chain and the amino acids dissembled from the polypeptide analyte, wherein the dissembled amino acids are linked to the polymerizable chain sequentially along the polymerizable chain according to the sequence of the amino acids in the polypeptide analyte. The structure of the PCAA complex and the method of producing the same is described in detail below. Unless specified separately, all the compounds and components mentioned / used can be commercially purchased from or synthesized by diverse suppliers.
[0118] PCAA Complex
[0119] As used herein, the term “polymerizable chain-amino acid complex” or “PCAA complex” refers to a chain of polymerizable molecules coupled to a series of amino acids. In particular, the PCAA complex described herein refers to a chain of polymerizable molecules wherein a plurality of amino acids is sequentially linked to the chain of polymerizable molecules. FIG. 6 illustrates an exemplary embodiment of the PCAA complex described in the present disclosure. Referring to FIG. 6, a PCAA complex 200 contains a chain of polymerizable molecule 201. In some examples, the chain of a polymerizable molecule is a nucleic acid. In some examples, the chain of nucleic acid is single stranded, double stranded, or partially double stranded. Referring to FIG. 6, a plurality of amino acids 202 is coupled to the chain of polymerizable molecule 201 via linkers 203. In some embodiments, the amino acids 202 are disassembled from a polypeptide analyte and are coupled along the chain of polymerizable molecule 201 via an iterative process as described elsewhere herein. Referring to FIG. 6, the chain of polymerizable molecule 201 is also coupled with barcode moi eties 204. It is understood that there can be only one barcode moiety 204 in PCAA complex.
[0120] In some examples, the chain of polymerizable molecule can be DNA or RNA, which can comprise any useful number and type of nucleotides, e.g., including canonical and noncanonical bases, and the number of nucleotides may be modulated based on the intended purpose. For example, the length of the chain of polymerizable molecule may be modulated to alter a property of the analytes (e.g., volume, aspect ratio, charge, etc.) to which the chain of polymerizable molecule is linked. The chain of polymerizable molecule may additionally or088177-8007W001 alternatively comprise any useful functional sequences, including but not limited to barcode sequences or other identifying sequences, UMI sequences, enzyme recognition sites (e.g., transposition sites, restriction sites), spacer sequences, sequencing primer sequences, read sequences, or primer sequences. The chain of polymerizable molecule may comprise canonical bases, noncanonical bases, naturally occurring bases, synthetic bases, abasic sites, or a combination thereof.
[0121] The amino acids may be covalently or non-covalently coupled to the chain of polymerizable molecule. The coupling method may be performed using any suitable chemistries and reaction conditions and may comprise the usage of a linker. In an example, a chain of polymerizable molecule may comprise a first reactive group, e.g., a first click chemistry moiety, as described elsewhere herein, and may be contacted with a linker comprising a second reactive group, e.g., a second click chemistry moiety that is able to react with the first reactive group. The linker may also comprise an additional reactive group that can tether to an amino acid or a polypeptide analyte. For example, the additional reactive group may be a thiocyanate conjugate, e.g., an isothiocyanate (ITC) such as phenyl isothiocyanate (PITC) or naphthylisothiocyanate (NITC), or an aldehyde group, e.g., ortho-phthalaldehyde (OP A), 2,3 -naphthalenedicarboxy aldehyde (ND A), a guanidinylating agent, dinitrofluorobenzene (DNFB), dansyl chloride, or other amino acid-reactive group. The linker may be reacted with a polypeptide analyte. Use of such a linker comprising at least two reactive groups may allow for (i) tethering of the amino acid or the polypeptide analyte to the linker and (ii) tethering of the linker to the polymerizable molecule. In some instances, the linker may be provided pre-tethered to the polymerizable molecule prior to contacting with the polypeptide analyte. In some instances, the conjugation of the polypeptide analyte to the polymerizable molecule, either via a linker or without a linker, may change the chemical structure of the polypeptide analyte. For example, if using a linker comprising an isothiocyanate moiety, the analyte may be derivatized to a thiocarbamyl group (e.g., under alkaline conditions), a thiazolinone group (e.g., under acid conditions), a thiohydantoin group, or other chemical moiety, thereby generating a PCAA complex comprising a modified polypeptide analyte coupled thereto.
[0122] Iterative Process to Generate a PCAA Complex
[0123] In some instances, the PCAA complex disclosed herein is generated using an iterative process as described in W02024 / 030919A1, which is hereby incorporated by reference in its entirety. In short, in the iterative process, individual amino acids of a088177-8007W001 polypeptide may be sequentially removed and re-tethered together, such that the distance between the individual analytes is increased.
[0124] In an example, an iterative process described herein may comprise providing a barcoded polypeptide comprising a plurality of amino acids and the chimeric barcode described elsewhere herein, a linker (e.g., as described elsewhere herein), and a polymerizable molecule. The linker may be configured to couple to (i) a terminal amino acid (e.g., N-terminal amino acid (NTAA) or C-terminal amino acid (CTAA)) of the polypeptide and (ii) a polymerizable molecule. The iterative process may further comprise contacting the linker with the terminal amino acid and the polymerizable molecule. Alternatively, or in addition, the linker may be provided pre-tethered to the polymerizable molecule and subsequently reacted with the terminal amino acid. The linker may be coupled to the terminal amino acid of the barcoded polypeptide to generate an amino acid-linker complex. The amino acid-linker complex may then couple to a polymerizable chain (as part of the chimeric barcode) linked to the polypeptide analyte. For example, the polymerizable chain may comprise nucleic acid molecules, which may be coupled to a nucleic acid segment of the polymerizable molecule via hybridization, ligation, or both. In some instances, the method may further comprise, after coupling the amino acid-linker complex to the polymerizable chain, cleaving the terminal amino acid from the polypeptide analyte to yield a polymerizable chain-amino acid (PCAA) complex, and optionally repeating the process. In an example in which the process is repeated, an additional linker may be provided which is configured to couple to (i) an additional amino acid of the polypeptide (e.g., the n-1 NTAA or n-1 CTAA) and (ii) an additional polymerizable molecule. The method may further comprise contacting the additional linker with the additional amino acid to generate an additional amino acid-linker complex. The additional polymerizable molecule may be coupled to the linker prior to, during, or subsequent to the coupling of the linker to the additional amino acid. The additional polymerizable molecule may be configured to couple to the PCAA complex (e.g., via the stacked polymerizable chain of the stacked PCAA complex). As such, in some examples, subsequent to generation of the additional linker- additional amino acid complex, the additional linker-additional amino acid complex may couple to the PCAA complex, thereby generating a stacked PCAA complex, and the additional amino acid may be cleaved from the peptide prior to, during, or subsequent to generation of the stacked PCAA complex.
[0125] FIG. 7 schematically illustrates an example workflow of generating a PCAA complex from a barcoded polypeptide using an iterative process. In such an example workflow 300, a polypeptide analyte 301 linked to a chimeric barcode 302 is provided, wherein the088177-8007W001 chimeric barcode 302 comprises a polymerizable chain 304 coupled with barcode moieties 303. It is understood that the polymerizable chain may be coupled with one barcode moiety 303. In step 305, a terminal amino acid (TAA) 306 of the polypeptide 301 is modified to contain a first reactive group. A polymerizable molecule 307 pre-tethered to a linker 308 is provided, wherein the linker 308 contains a second reactive group that can react with the first reactive group. The linker 308 is then couple to the TAA 306 via the reaction between the first reactive group and the second reactive group. In step 309, the polymerizable molecule 307 is coupled to the polymerizable chain 304, e.g., via ligation. In step 310, the TAA 306 is cleaved from the polypeptide 301 to generate a polymerizable chain-amino acid (PCAA) complex 311, which comprises the TAA 306, the linker 308, the polymerizable molecule 307, and the polymerizable chain 304. Steps 305, 309, and 310 may be iterated and repeated any number of times (“rounds”) using additional polymerizable molecule 307 pre-tethered with linkers 308 to tether to the PCAA complex 311. Multiple rounds may continue until all or a subset of the amino acids in the polypeptide 301 are removed from the polypeptide and tethered together. For example, a PCAA complex 312 comprising a polymerizable chain 313 linked to n amino acids may result from n rounds of the workflow 300 (e.g., iterations of processes 305, 309, and 310).
[0126] In some embodiments, the iterative process described above may be performed on a substrate to which the capture polymerizable molecule is tethered. The PCAA complex generated via the iterative process and coupled to the substrate may be cleaved after the completion of the iterative process. The cleavage may be mediated by a photo-cleavable linker through which the capture polymerizable molecule is tethered to the substrate. Accordingly, the system described herein may also comprise a UV lamp to cleave the PCAA complex from the substrate.
[0127] Sequencing of PCAA Complexes
[0128] In some embodiments, the method and system disclosed herein involves identifying the amino acids on the PCAA complexes. In some embodiments, the identification is performed using a nanoscale recording device, e.g., a nanopore, a nanogap or nanochannel device. In some embodiments, the identification is performed using a binder-associated decoding approach.
[0129] Nanoscale Device Sequencing approach
[0130] The method and system of polypeptide sequencing using a nanoscale recording device, such as nanopores, nanogaps, nanochannels, or a field effect transistor, to identify the amino acids coupled along a polymerizable chain has been disclosed in W02024030919A1, US Provisional Application No. 63 / 645,885 and PCT application PCT / US2025 / 028808, which are herein incorporated by reference.088177-8007W001
[0131] Typically, the nanoscale recording device comprises two chambers connected via a nanopore, nanogap, or nanochannel. In some instances, a nanopore, nanogap, or nanochannel may be provided on a membrane in an ionic solution. A signal may be measured from the nanopore, membrane, or surrounding solution. For example, a conductance, current, current blockage, or other parameter within the nanopore may be monitored as a function of time. As a molecule enters the nanopore, nanogap, or nanochannel, a change in the conductance, current, or other parameter may occur and provide information (e.g., size, charge, aspect ratio, volume) on the molecule. Each amino acid or modified amino acid, or a subset of amino acids or modified amino acids, may generate a unique signal that is distinguishable from other amino acids or modified amino acids. As such, the unique signal signatures may be assigned to the amino acids or modified amino acids in order to determine the identity of the amino acids, or modified amino acids.
[0132] In some instances, a modified amino acid may be analyzed numerous times, e.g., via translocation and measuring of a current or conductance of the modified amino acid, through the same nanopore or nanogap. For example, iterative reading of a modified amino acid may be beneficial in improving the accuracy of the reads or identification of the modified amino acid. In such cases, the modified amino acid may be translocated through one or more nanopores at least 2 times, at least 3 times, at least 4 times, at least 5 times, at least 6 times, at least 7 times, at least 8 times, at least 9 times, at least 10 times, at least 20 times, at least 50 times, at least 100 times, at least 200 times, or even greater.
[0133] The nanopore, nanogap, or nanochannel may be generated from an organic material, e.g., a pore-forming protein or a transmembrane protein. Such a protein may be naturally occurring, synthetic, or engineered. Examples of naturally occurring organic nanopores include wild-type aerolysin, alpha-hemolysin, mycobacterial porins (e.g., MspA porin), Phi29 connector channels, Fragaceatoxin C, Cytolysin A, Ferric hydroxyamate uptake component A, Curb specific gene G, outer membrane porin G, viral DNA packaging motors, etc. Alternatively or in addition to, the nanopore, nanochannel, or nanogap may comprise an engineered variant of a naturally occurring nanopore. In some embodiments, the nanopore, nanogap, or nanochannel is comprised of an inorganic material. For example, solid-state nanopores may be made from dielectric materials such as a silicon compound (e.g., silicon nitride, silicon dioxide), an aluminum compound (e.g., aluminum oxide), a titanium compound (e.g., titanium oxide), a molybdenum compound (e.g., molybdenum sulfate), hafnium, graphene, etc. The nanopore, nanochannel, or nanogap may assume any useful form factor or geometry, e.g., gaps or channels within membranes, capillaries, etc., and may be generated using any suitable process,088177-8007W001 e.g., ion beam sculpting, electron beam exposure. The nanopore, nanochannel, or nanogap may comprise an elastomeric material.
[0134] In some instances, the nanopore, nanogap, or nanochannel is coupled to a protein. The protein may be a molecular motor, which may facilitate movement of the modified amino acid or portion thereof (e.g., the polymerizable material) through the nanopore. In a nonlimiting example, the modified amino acid may comprise an amino acid coupled, optionally via a linker, to a nucleic acid molecule, and the nanopore may be coupled to a helicase or a molecule comprising helicase activity. The helicase may be used to translocate the nucleic acid molecule and the coupled amino acid through the nanopore. In some instances, the molecular motor may increase, decrease, or otherwise change the translocation velocity of the modified amino acid through the nanopore as compared to an unmodified amino acid. In other examples, the molecular motor may comprise a topoisomerase, a polymerase, a nuclease (e.g., endonucleases such as restriction endonucleases or Cas proteins, or exonuclease) an unfoldase, mitotic spindle protein (e.g., nuclear mitotic apparatus protein, kinesin, dynein), or other motor protein (e.g., myosin), a variant thereof, or other polymer-processing protein. In some instances, the protein comprises a protease or proteosome, which can enable “chop-n-drop” or cleaving of the modified amino acid or portion thereof prior to translocation in the nanopore.
[0135] Alternatively or in addition to, translocation of the amino acids through a medium or through a nanopore or nanogap may be facilitated by application of a force. For example, a molecule may be translocated by application of pressure (e.g., pressure-driven flow), an electric field (e.g., via electrophoresis, electroosmotic flow, isoelectric focusing), a magnetic field (e.g., using a ferromagnetic fluid, magnetic particles), light (e.g., optoelectronics).
[0136] A commercially available nanopore system may be used in the methods described herein. For example, a nanopore system from Oxford Nanopore Technologies (ONT), such as the MinlON, VolTRAX, GridlON, PromethlON, MinIT, Flongle, or Q-Line products may be used to identify and characterize the modified amino acids described herein.
[0137] In some embodiments, the nano-scale recording device, e.g., nanopore, may be engineered to improve the resolution of amino acid identification. The change of properties of pore protein, such as charge, size and hydrophilicity, can generate distinguishable current pattern of different amino acids. In some embodiments, pore proteins can be designed by Al to generate ideal candidates with more stable structure and specific characters for the amino acid identification.
[0138] In some embodiments, the methods and systems described herein involve using a blocker capable of preventing the amino acids on the PCAA complex from translocation in the088177-8007W001 nano-scale recording device as disclosed in US Provisional Application No. 63 / 645,885 and PCT application PCT / US2025 / 028808, which are herein incorporated by reference.
[0139] In one embodiment, after generating a polymerizable chain-amino acid (PCAA) complex comprising a stacked polymerizable chain and the amino acids dissembled from the polypeptide, the method of sequencing a polypeptide comprising a sequence of amino acids comprises the following steps:(a) generating from the polypeptide a PCAA complex comprising a polymerizable chain comprising a first end and a second end, and the amino acids dissembled from the polypeptide, and the amino acids are linked to the polymerizable chain sequentially along the polymerizable chain from the first end to the second end according to the sequence of the amino acids in the polypeptide;(b) subjecting the PCAA complex to a recording device comprising a first chamber, and a second chamber connected to the first chamber, and the polymerizable chain non-covalently binds to at least a blocker in the first chamber, and the blocker prevents an amino acid of the PCAA complex from translocating from the first chamber to the second chamber;(c) dissociating the blocker from the polymerizable chain to allow the amino acid to translocate from the first chamber to the second chamber;(d) detecting via the recording device a signal that is generated from translocation of the amino acid from the first chamber to the second chamber; and(e) identifying the amino acid based on the signal detected in (d).
[0140] In another embodiment, after generating a polymerizable chain-amino acid (PCAA) complex comprising a stacked polymerizable chain and the amino acids dissembled from the polypeptide, the method of sequencing a polypeptide comprising a sequence of amino acids comprises the following steps:(a) generating from the polypeptide a PCAA complex comprising a polymerizable chain comprising a first end and a second end, and the amino acids dissembled from the polypeptide, and the amino acids are linked to the polymerizable chain sequentially along the nucleic acid strand from the first end to the second end according to the sequence of the amino acids in the polypeptide;(b) subjecting the PCAA complex to a recording device comprising a first chamber, and a second chamber connected to the first chamber, and the nucleic acid strand comprises a blocker sequence forming a structure that prevents an amino acid of the chimeric nucleic acid chain from translocating from the first chamber to the second chamber;088177-8007W001(c) changing the structure of the blocker sequence to allow the first amino acid to translocate from the first chamber to the second chamber;(d) detecting via the recording device a signal that is generated from translocation of the amino acid from the first chamber to the second chamber; and(e) identifying of the amino acid based on the signal detected in (d).
[0141] Binder-Associated Decoding approach
[0142] In some embodiments, the methods and systems described herein involve using a binder-associated decoding assay. The binder-associated decoding assay converts the PCAA complex to a dsDNA, which can be sequenced using nucleic acid sequencer. The binder- associated decoding assay has been described in US Provisional Application No. 63 / 645,884, which are herein incorporated by reference.
[0143] In short, a binder-associated decoding assay involves incubating a PCAA complex with a plurality of binder-DNA agents, wherein each binder-DNA agent comprises (i) a binder capable of specifically binding to an amino acid, and (ii) a DNA comprising a first barcode sequence, thereby generating a complex comprising the PCAA complex and the plurality of binder-DNA agents. The plurality of binder-DNA agents is then ligated to generate a nucleic acid strand comprising a plurality of barcodes, which encodes the information of the amino acids on the PCAA complex. The nucleic acid strand is then sequenced to identify the amino acids based on the barcodes information.
[0144] In one embodiment, after generating a polymerizable chain-amino acid (PCAA) complex comprising a stacked polymerizable chain and the amino acids dissembled from the polypeptide, the method of sequencing a polypeptide comprising a sequence of amino acids comprises the following steps:(a) generating from the polypeptide a PCAA complex comprising a first nucleic acid strand and the amino acids disassembled from the polypeptide, and the amino acids are sequentially linked to the first nucleic acid strand according to the sequence of the amino acids in the polypeptide;(b) incubating the PCAA with a plurality of binder-DNA agents, and each of the binder- DNA agent comprises a binder specifically binding to an amino acid and a DNA comprising a barcode sequence, thereby generating a complex comprising the chimeric nucleic acid chain and the plurality of binder-DNA agents;(c) ligating the plurality of binder-DNA agents to generate a second nucleic acid strand;(d) sequencing the second nucleic acid strand to determine the order of the barcode sequences in the second nucleic acid strand; and088177-8007W001(e) determining the sequence of the amino acids in the polypeptide based on the order of the barcode sequences in the second nucleic acid strand.
[0145] Systems, Kits and Compositions
[0146] Also provided herein are systems, kits, and compositions for sequencing polypeptides. The systems, kits, and compositions provided herein may be useful in implementing any of the described methods or may be provided in complement to the described methods.
[0147] In one aspect, a kit of the present disclosure may comprise a chimeric barcode, a linker, a polymerizable molecule, a peptide-conjugation reagent, an enzyme (e.g., a polymerizing or ligating enzyme, a cleaving enzyme, a restriction enzyme, a nicking enzyme, an exonuclease, a repair enzyme such as a uracil DNA glycosylase), a substrate, or any combination thereof. The kits may comprise buffers, reagents, binding agents, catalysts, or other chemicals or biological molecules (e.g., enzymes) necessary for conducting a chemical or enzymatic reaction. The kits of the present disclosure may further comprise instructions for using the components of the kit or for implementing any of the methods and processes described herein. For example, the kit may comprise instructions for coupling a chimeric barcode to a polypeptide analyte using a polypeptide-conjugation reaction. Similarly, the kit may comprise instructions for performing an iterative process, as described herein, e.g., to generate a PCAA complex, etc.
[0148] In another aspect, disclosed herein are compositions that may be used to characterize a polypeptide. A composition may comprise a linker covalently attached to a polymerizable molecule, which may, for example, be useful in iterative process of a polypeptide for polypeptide sequencing. In an example, a composition may comprise (A) a linker comprising (i) a first moiety that can couple to an amino acid (e.g., a CTAA or NTAA of a peptide), (ii) a second moiety that can couple to a polymerizable molecule (e.g., DNA), and optionally, (iii) a releasable or cleavable moiety, which may be the same or different moiety as (ii), and also optionally, (iv) a spacer moiety, and (B) a polymerizable molecule. In another example, the composition may comprise a linker that is covalently coupled to a polymerizable molecule; such a linker may comprise a moiety that can couple to and optionally cleave an amino acid (e.g., a CTAA or NTAA of a peptide), and the covalently coupled polymerizable molecule may be configured to tether to another polymerizable molecule (e.g., a capture moiety), which may in some instances, be provided attached to a substrate.088177-8007W001
[0149] A system of the present disclosure may comprise a nano-scale recording device that is configured to record a signal generated by the translocation of a chimeric nucleic acid chain, as described herein, and to provide sequencing reads of the amino acids coupled along the PCAA complex. Alternatively, or in addition to, a system of the present disclosure may be configured to process, prepare, or sequence a PCAA complex. The system may be configured to provide the polypeptide, the capture moiety, and a linker comprising a polymerizable molecule; couple the linker to the polypeptide; and optionally, iterate one or more operations or processes. Accordingly, systems of the present disclosure may comprise any useful apparatuses or tools, including but not limited to mixers, liquid handlers, vortexes, centrifuges, heating or cooling elements, mechanical stages, microfluidic chambers or devices and fluidic controls.
[0150] A system of the present disclosure may also comprise an enzyme (e.g., a polymerizing or ligating enzyme, a cleaving enzyme, a restriction enzyme, a nicking enzyme, an exonuclease, a repair enzyme such as a uracil DNA glycosylase), a detection or labeling agent, buffers, reagents, binding agents, catalysts, or other chemicals or biological molecules (e.g., enzymes) necessary for conducting a chemical or enzymatic reaction, or any combination thereof. The system may further comprise one or more detection (e.g., imaging or mass spectrometry) systems, separation systems (e.g., HPLC), or other analytical instruments.
[0151] A system of the present disclosure may also comprise a UV lamp that cleaves the chimeric nucleic acid from a substrate via a photo-cleavable linker.EXAMPLES
[0152] Unless otherwise specified, all reagents mentioned in the following examples are commercially available.EXAMPLE 1
[0153] This example illustrates the generation of two three-codon barcodes made of DNA (Type I codon) and their detection in an exemplary method disclosed herein.
[0154] As shown in FIG. 8, the backbones of two three-codon barcodes made of DNA were synthesized (SEQ ID NO: 1 and 2, respectively). Each barcode backbone contains three codons (ATC and TCA, respectively), wherein each codon is made of 4-mer nucleotides. The cholesterol tag sequence of the barcodes complements to a cholesterol tag (SEQ ID NO: 4, see Yusko, E., et al. Controlling protein translocation through nanopores with bio-inspired fluid walls. Nature Nanotech 6, 253-260 (2011)). Each codon of the barcodes is followed by a blocker sequence, which complements to a blocker (SEQ ID NO: 3).088177-8007W001
[0155] To prepare a barcode composition, the barcode backbone, blocker and cholesterol tag were mixed with the molar ratio of 1 : 5 : 1.
[0156] The mixture was annealed at 95 degree water bath. The annealed barcode composition with ATC-codon (Barcode 1) was then diluted 100X by water, while the barcode composition with TCA-codon (Barcode 2) was diluted 10X, 100X, 1000X, 10000X by water. The annealed barcodes were used for the detection.
[0157] To quantify the two three-codon barcodes, Barcode 1 and Barcode 2 were mixed at a concentration ratio of 1 :0.01, 1 :0.1, 1 : 1, or 1 : 10, respectively. After 1 hour recording, the events of Barcode 1 or 2 were counted. The measured event frequency ratios between Barcode 1 and Barcode 2 were 1 :0.01, 1 :0.10, 1 : 1.02, and 1 :9.35, respectively. FIG. 9 shows the plot of concentration ratio vs frequency ratio. The dash line shows the linear fit. The plot demonstrates reliable relative quantification of these two barcodes mixed at a large range of ratios. Table 1 summarized the data analysis of the barcode events of the four different barcode mixtures from the detection.
[0158] Table 1. Data analysis of the barcode eventsConcentration ratio (X axis) (Barcodel:Barcode2) 1 : 0.01 1 : 0.1 1:1 1:10Frequency (Barcode 1) 1627 345 157 119Frequency (Barcode2) 17 33 160 1113Frequency ratio (Y axis) (Barcode l:Barcode2) 1 : 0.01 1 : 0.10 1 : 1.02 1 : 9.35Capture rate (Barcodel) (1 / s) 0.4292 0.2500 0.1250 0.0606Capture rate (Barcode2) (1 / s) 0.0054 0.0242 0.1235 0.4762Rate ratio (Barcodel :Barcode2) 1 : 0.01 1 : 0.10 1:0.99 1:7.86EXAMPLE 2
[0159] This example illustrates the methods to conjugate modified amino acids to the DNA backbones for Type II chimera barcodes.
[0160] As shown in FIG. 10, PTH-AA (phenylthiohydantoin amino acid) can be used to conjugate to alkyne-DNA for Type ILA chimera barcodes or conjugate to DNA-sBCN for Type II-B chimera barcodes.
[0161] Conjugation of PTH-AA to DNA was achieved via copper-catalyzed azide-alkyne cycloaddition (CuAAC), using a reaction mixture containing 5 D pM alkyne-DNA, 15 DmM PTH-AA, lOODmM sodium phosphate buffer, 5 DmM aminoguanidine hydrochloride,088177-8007W0012.5 DmM BTTAA ligand, 0.5 DmM copper sulfate, lODmM sodium ascorbate (as a reducing agent), and 50% (v / v) DMSO. The reaction was performed at room temperature for 3 hours. DNA-amino acid conjugates were purified using the Zymo Oligo Clean & Concentrator kit.
[0162] To conjugate PTH-AA to DNA-sBCN, a 20 > pL reaction was prepared containing 10D pL of bicyclo[6.1.0]nonyne(BCN)-DNA (200D pM), I D pL of PTH-AA (phenylthiohydantoin amino acid), 2D pL of 10* PBS, and 7D pL of nuclease-free water. The reaction was incubated overnight at room temperature with gentle rotation. The product was purified using a Cytiva gel filtration spin column.EXAMPLE 3
[0163] This example illustrates a method of synthesizing a Type ILA chimera barcode in solution.
[0164] As shown in FIG. 11, DNA-AA trimer was assembled by annealing three monomer units (A, B, and C) in the presence of two complementary splint DNA oligos (X and Y). The ratio for oligo A, B, C, and splint oligo = 1 :3:5 :5. Following annealing, ligation was performed using T4 DNA ligase at 25 D°C for 1 hour. The desired trimer product was separated by electrophoresis using a TBU polyacrylamide gel. The gels were stained with SybrGold and visualized on a gel imager (Azure Biosystems). The trimer products were subsequently recovered using a small-RNA PAGE Recovery Kit.EXAMPLE 4
[0165] This example illustrates a method of synthesizing a Type II-B chimera barcode on solid phase (e.g., magnetic beads).
[0166] FIG. 12 shows a schematic of the solid-phase synthesis method, and FIG. 13 shows sequences of the components for the solid-phase synthesis method.
[0167] the Oligo d(T)25 magnetic beads (New England Biolabs (NEB), Cat # S1419S) were thoroughly mixed and 20pl of the suspended beads were transferred to a microcentrifuge tube. The tube placed on a magnetic stand and the supernatant was decanted after the beads have gathered. The beads were washed 2 times with 50ul of Binding Buffer (100 mM Tris-HCl, pH 8.0, 500 mM NaCl, 1 mM EDTA). Last, the beads were washed with lOuL lxT4 DNA Ligase buffer (NEB, Cat # M0202S).
[0168] For the 1st Unit conjugation, the bead binding DNA, splint DNA and chimer-DNA unit were mixed with the washed beads unit at a certain ratio. The suspension was incubated at room temperature for 5 minutes to allow DNA assembling and binding to the Oligo d(T)25 on the beads. 10xT4 DNA Ligase Buffer and 2pl T4 ligase were added (final concentration of T4088177-8007W001DNA Ligase Buffer is IX, NEB M0202S) and incubated at room temperature for 30 min. The ligated beads were washed with 50uL of Wash Buffer (20 mM Tris-HCl, pH 8.0, 500 mM NaCl, 1 mM EDTA) twice for the next step,
[0169] For the 2nd Unit conjugation, the 1st unit ligated beads were washed with lOpl IXrCutSmart buffer (NEB B6004S). The beads were resuspended withl8pl of IXrCutSmart buffer and 2 pl of the restriction enzyme (NEB) was to the beads. The enzyme reaction was incubated at 37°C for 30min. The reaction was washed from the beads with 50 pl of Wash Buffer twice. For the ligation with another chimera DNA unit, the beads were primed withlOpl 1XT4 DNA Ligase Buffer. Second chimera-DNA Unit, 10xT4 DNA Ligase Buffer, 2pl T4 ligase were mixed and incubated at room temperature for 30 min. The same washing step was repeated. For the 3rd to n-th Unit conjugation, repeat the steps of 2nd Unit conjugation to make 3-mer chimera barcode.
[0170] After the unit ligation, the beads were resuspended in 20ul Elution Buffer (20 mM Tris-HCl, pH 8.0, 20mM NaOH, 1 mM EDTA) and incubated 60°C for 1 min to elute the barcode from the beads. The supernatant was collected immediately on a magnetic stand.
[0171] To purify the barcode, the multi-codon chimera barcode was eluted in water from the spin column, followed by the instruction of Oligo purification kit (Zymo Research, Oligo Clean & Concentrator, D4061).
[0172] To visualize the synthesis products, Ipl of each purified synthesized barcode sample was mixed with 1.5pl H2O and 2.5ul Novex™ TBE-Urea sample buffer (2X). The samples were incubated at 70 °C for 2min before being loaded to 15% TBU PAGE gel. The gel was run at 200V for 1 hour. Afterwards, the gel was stained with SybrGold and visualized on a GelDoc EZ imager (Bio-Rad) (FIG. 14. 3mer_l : F-polyA-F-polyT-F, 3mer_2: G-polyA- G-polyT-G, 3mer_3: H-polyA-E-polyT-Q, 3mer_4: H-polyA-H-polyT-H, 3mer_5: Q-polyA- Q-polyT-Q). The assembled chimera barcodes were further purified by HPLC chromatogram and the fractions targeting three-codon barcodes were collected for the nanopore experiment (FIG. 15)EXAMPLE 5
[0173] This example illustrates the detection of three-codon Type II-A chimera barcodes.
[0174] A single-channel recording using a MspA nanopore was used to detect three different three-codon chimera barcodes (ERW, HRG, and RHG). Data collection was performed with 1 M KC1 (10 mM Tris pH 7.4) at both cis and trans side. The chimera barcode was annealed with cholesterol tag and blocker with ratio of 5: 10:3 (barcode: blocker:088177-8007W001 cholesterol tag). The annealed sample was added to the cis side at the final concentration of approximately 3 nM. Data collection was performed at room temperature using the eOne Light amplifier and Elements Data Reader software (Elements). A transmembrane voltage of 135 mV was applied, and data were acquired at a sampling rate of 5 kHz. All representative traces are shown with 0-0.5 Ib / Io; Io (open pore current) and lb (block current). Reference for MspA nanopore (Yan et al., Rapid and multiplex preparation of engineered Mycobacterium smegmatis porin A (MspA) nanopores for single molecule sensing and sequencing, Chem. Sci. (2021) 12: 9339-9346).
[0175] As shown in FIG. 16, the three different three-codon chimera barcodes exhibited three highly distinguishable nanopore signal patterns.EXAMPLE 6
[0176] This example illustrates the detection of three-codon Type ILB chimera barcodes using a MspA nanopore.
[0177] The chimera barcodes were annealed with the cholesterol tag and the blocker with ratio of 1 : 1 :40 (barcode: blocker: cholesterol tag). The annealed sample was added to the cis side at the final concentration of approximately 10 nM. Data collection was performed with 1 M KC1 (10 mM Tris, pH 7.4) at both cis and trans side at room temperature using the eOne Light amplifier and Elements Data Reader software (Elements). A transmembrane voltage of 150 mV was applied, and data were acquired at a sampling rate of 5 kHz. All representative traces are shown with 0.6 second time window and 0-0.6 Ib / Io; Io (open pore current) and lb (block current). As shown in FIG. 17, the three different three-codon chimera barcodes exhibited three distinguishable nanopore signal patterns.
[0178] While the disclosure has been particularly shown and described with reference to specific embodiments (some of which are preferred embodiments), it should be understood by those having skill in the art that various changes in form and detail may be made therein without departing from the spirit and scope of the present disclosure as disclosed herein.
Claims
088177-8007W001WHAT IS CLAIMED IS:
1. A method comprising:(a) coupling a first barcode to a first polypeptide in a first sample, wherein the first polypeptide comprises a sequence of amino acids, wherein the first barcode comprises at least a first barcode moiety, and wherein the first barcode moiety comprises(i) a first codon capable of being recognized by a nanopore device having a nanopore sensor, and(ii) a first time-control element capable of temporarily holding the first codon at the nanopore sensor;(b) generating a first polymerizable chain-amino acid (PCAA) complex comprising(i) a first stacked polymerizable chain,(ii) the amino acids dissembled from the first polypeptide, wherein the amino acids dissembled from the first polypeptide are sequentially linked to the first stacked polymerizable chain, and(iii) the first barcode linked to the first stacked polymerizable chain;(c) subjecting the first PCAA complex to the nanopore device to determine:(i) the sequence of amino acids along the first stacked polymerizable chain, and(ii) the identity of the first barcode; and(d) determining(i) the sequence of the first polypeptide based on the sequence of amino acids along the first stacked polymerizable chain, and(ii) that the first polypeptide is from the first sample based on the identity of the first barcode.
2. The method of claim 1, wherein the first barcode comprises a first polymerizable chain; optionally, wherein the first polymerizable chain is a nucleic acid or PEG; optionally, wherein the nucleic acid is a DNA, a RNA or a modified variant thereof.
3. The method of claim 1, wherein the first codon comprises: a) a polynucleotide; b) an amino acid derivative; optionally, wherein the amino acid derivative is a phenylthiohydantoin amino acid;088177-8007W001 c) a compound selected from bicyclo[6.1.0]nonyne (BCN), a crown ether molecule, a sugar molecule (e.g., a sialic acid); or d) a polynucleotide and an amino acid derivative.
4. The method of claim 1, wherein the first barcode further comprises a second barcode moiety comprising a second codon and a second time-control element.
5. The method of claim 1, wherein the first time-control element comprises a polynucleotide.
6. The method of claim 1, wherein the first time-control element forms a structure capable of temporarily holding the first codon at the nanopore sensor.
7. The method of claim 1 , wherein the first time-control element binds to a reagent to form a complex capable of temporarily holding the first codon at the nanopore sensor.
8. The method of claim 1, wherein the nanopore device does not comprise a motor protein; optionally, the motor protein is a translocase; optionally, the translocase is a helicase, a DNA polymerase or a RNA polymerase.
9. The method of claim 1, wherein the method further comprises:(b’) mixing the first PCAA complex with a second PCAA complex, wherein the second PCAA complex is generated from a second polypeptide in a second sample, wherein the second PCAA complex comprises(i) a second stacked polymerizable chain,(ii) the amino acids dissembled from the second polypeptide, wherein the amino acids dissembled from the second polypeptide are sequentially linked to the second stacked polymerizable chain, and(iii) a second barcode linked to the second stacked polymerizable chain, wherein the second barcode is different from the first barcode;(c’) subjecting the second PCAA complex to the nanopore device to determine:(i) the sequence of amino acids along the second stacked polymerizable chain, and(ii) the identity of the second barcode; and(d’) determining(i) the sequence of the second polypeptide based on the sequence of amino acids along the second stacked polymerizable chain, and088177-8007W001(ii) that the second polypeptide is from the second sample based on the identity of the second barcode.
10. A method comprising:(a) coupling a first barcode to each of a first plurality of polypeptides in a first sample, wherein the first barcode comprises at least a first barcode moiety, and wherein the first barcode moiety comprises(i) a first codon capable of being recognized by a nanopore device having a nanopore sensor, and(ii) a first time-control element capable of temporarily holding the first codon at the nanopore sensor;(b) generating a first plurality of PCAA complexes, each PCAA complex comprising(i) a first stacked polymerizable chain,(ii) amino acids dissembled from a polypeptide in the first plurality of polypeptides, wherein the amino acids dissembled from the polypeptide are sequentially linked to the first stacked polymerizable chain, and(iii) the first barcode, wherein the first barcode is linked to the first stacked polymerizable chain;(c) mixing the first plurality of PCAA complexes with a second plurality of PCAA complexes, wherein the second plurality of PCAA complexes is generated from a second plurality of polypeptides in a second sample, wherein each of the second plurality of PCAA complexes comprises(i) a second stacked polymerizable chain,(ii) amino acids dissembled from a polypeptide in the second plurality of polypeptides, wherein the amino acids are sequentially linked to the second stacked polymerizable chain, and(iii) a second barcode linked to the second stacked polymerizable chain, wherein the second barcode is different from the first barcode;(d) subjecting a mixture of the first plurality of PCAA complexes and the second plurality of PCAA complexes to the nanopore device to determine:(i) the sequence of amino acids along the first stacked polymerizable chain,(ii) the identity of the first barcode,088177-8007W001(i1) the sequence of amino acids along the second stacked polymerizable chain,(ii) the identity of the second barcode; and(e) determining(i) the sequence of the polypeptide in the first plurality of polypeptides based on the sequence of amino acids along the first stacked polymerizable chain,(i1) the sequence of the polypeptide in the second plurality of polypeptides based on the sequence of amino acids along the second stacked polymerizable chain,(ii) that the polypeptide in the first plurality of polypeptides is from the first sample based on the identity of the first barcode, and(ii’) that the polypeptide in the second plurality of polypeptides is from the second sample based on the identity of the second barcode.
11. A system comprising:(a) a module for coupling a first barcode to a first polypeptide in a first sample, wherein the first polypeptide comprises a sequence of amino acids, wherein the first barcode comprises at least one barcode moiety, and wherein the first barcode moiety comprises(i) a first codon capable of being recognized by a nanopore device having a nanopore sensor, and(ii) a first time-control element capable of temporarily holding the first codon at the nanopore sensor;(b) a module for generating a first PCAA complex comprising(i) a first stacked polymerizable chain,(ii) the amino acids dissembled from the first polypeptide, wherein the amino acids dissembled from the first polypeptide are sequentially linked to the first stacked polymerizable chain, and(iii) the first barcode linked to the first stacked polymerizable chain;(c) a module for subjecting the first PCAA complex to the nanopore device to determine:(i) the sequence of amino acids along the first stacked polymerizable chain, and(ii) the identity of the first barcode; and088177-8007W001(d) a module for determining(i) the sequence of the first polypeptide based on the sequence of amino acids along the first stacked polymerizable chain, and(ii) that the first polypeptide is from the first sample based on the identity of the first barcode.
12. The system of claim 11, further comprising:(b’) a module for mixing the first PCAA complex with a second PCAA complex, wherein the second PCAA complex is generated from a second polypeptide in a second sample, wherein the second PCAA complex comprises(i) a second stacked polymerizable chain and the amino acids dissembled from the second polypeptide, wherein the amino acids dissembled from the second polypeptide are sequentially linked to the second stacked polymerizable chain, and(ii) a second barcode linked to the second stacked polymerizable chain, wherein the second barcode is different from the first barcode;(c’) a module for subjecting the second PCAA complex to the nanopore device to determine:(i) the sequence of amino acids along the second stacked polymerizable chain, and(ii) the identity of the second barcode; and(d’) a module for determining(i) the sequence of the second polypeptide based on the sequence of amino acids along the second stacked polymerizable chain, and(ii) that the second polypeptide is from the second sample based on the identity of the second barcode.
13. A barcode composition comprising at least a first barcode moiety, wherein the first barcode moiety comprises(i) a first codon capable of being recognized by a nanopore device having a nanopore sensor, and(ii) a first time-control element capable of temporarily holding the codon at the nanopore sensor.
14. The barcode composition of claim 13, wherein the barcode comprises a polymerizable chain; optionally, wherein the polymerizable chain is a nucleic acid or PEG; optionally, wherein the nucleic acid is a DNA, a RNA or a modified variant thereof.088177-8007W00115. The barcode composition of claim 13, wherein the first codon comprises: a) a polynucleotide ; b) an amino acid derivative; optionally, wherein the amino acid derivative is a phenylthiohydantoin amino acid; c) a compound selected from bicyclo[6.1.0]nonyne (BCN), a crown ether molecule, a sugar molecule (e.g., a sialic acid); or d) a polynucleotide and an amino acid derivative.
16. The barcode composition of claim 13, wherein the first barcode further comprises a second barcode moiety comprising a second codon and a second time-control element.
17. The barcode composition of claim 13, wherein the first time-control element comprises a polynucleotide.
18. The barcode composition of claim 13, wherein the first time-control element forms a structure capable of temporarily holding the first codon at the nanopore sensor.
19. The barcode composition of claim 13, wherein the first time-control element binds to a reagent to form a complex capable of temporarily holding the first codon at the nanopore sensor.
Citation Information
Patent Citations
Translocation control elements, reporter codes, and further means for translocation control for use in nanopore sequencing
US20220411458A1
Protein sequencing via coupling of polymerizable molecules
WO2024030919A1