Protein sequencing via polymeric molecule attachment
By attaching polymerizable molecules to amino acids for nanopore sequencing, the method addresses selectivity and sensitivity issues in protein sequencing, achieving accurate and high-throughput amino acid identification.
Patent Information
- Application Number
- JP2025505933
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-02
- Filing Date
- 2023-08-01
- Publication Date
- 2025-08-20
AI Technical Summary
Current protein sequencing techniques face challenges in selectivity, sensitivity, and yield due to the folded nature of proteins and intramolecular spacing of amino acids, leading to overlapping signals in nanopore readouts.
Attaching polymerizable molecules to amino acids via an intramolecular expansion process, allowing for nanopore- or nanogap-mediated sequencing that distinguishes individual amino acids by altering transport rates and signal-to-noise ratios.
Enhances the accuracy and throughput of protein sequencing by enabling the identification of amino acids in their correct order, improving signal separation and detection.
Smart Images

Figure 2025527266000001_ABST
Abstract
Description
[Technical Field]
[0001] cross reference This application claims the benefit of U.S. Provisional Application No. 63 / 394,475, filed August 2, 2022, which is incorporated herein by reference in its entirety. [Background technology]
[0002]
[0003] Advances in biomolecular analysis and characterization have proven crucial in understanding biological and pathological mechanisms that influence disease diagnosis and modeling, the development of therapeutics and treatments, and improved health outcomes. Among these advances, nucleic acid sequencing has emerged as a key tool in the genomic and transcriptomic analysis of biological samples.
[0003] Protein signaling supports various cellular processes and plays important functions in viruses, cells, and organisms. However, current techniques for studying proteins are limited in selectivity, sensitivity, and yield, or require a priori knowledge. Therefore, novel approaches for protein characterization and analysis are needed. Summary of the Invention
[0004] The present disclosure recognizes a need for a technology for de novo protein studies with improved accuracy and in a high-throughput format. The use of nanopores in protein or peptide sequencing by conventional methods faces challenges. Transport of peptides or proteins through nanopores and subsequent readout is hindered by the properties of proteins: they are folded and not uniformly charged, and the intramolecular spacing of amino acids is small, causing many amino acids to enter the nanopore simultaneously, resulting in overlapping signals that are difficult to separate from one another. Provided herein are systems, compositions, kits, and methods for analyzing proteins that address the above needs. The disclosed methods may include attaching multiple polymerizable molecules to amino acids of a peptide via an intramolecular expansion process and analyzing the polymerizable molecule-amino acid conjugates. One or more of the processes described herein may include nanopore- or nanogap-mediated sequencing, which allows for the identification of individual amino acids of a peptide in the order in which they appear or occur in the peptide.
[0005] In one aspect, provided herein is a method for processing a modified amino acid, comprising: (a) providing a modified amino acid, wherein the modified amino acid comprises a polymerizable molecule; and (b) transporting the modified amino acid or derivative thereof through a nanopore or nanogap, wherein the polymerizable molecule causes a change in the transport rate of the modified amino acid or derivative thereof through the nanopore or nanogap compared to the transport rate of an amino acid that does not comprise the polymerizable molecule.
[0006] In another aspect, a method is provided comprising: (a) providing a modified amino acid, wherein the modified amino acid comprises a polymerizable molecule; (b) transporting the modified amino acid or derivative thereof through a nanopore or nanogap; and (c) measuring a signal from the nanopore or nanogap, wherein the polymerizable molecule causes an increase in the measured signal-to-noise ratio (SNR) of the signal in (c) compared to the SNR of an amino acid that does not comprise the polymerizable molecule.
[0007] In some embodiments, the modified amino acids comprise or are derived from proteinogenic amino acids attached to the polymerizable molecule.
[0008] In some embodiments, the modified amino acid or derivative thereof further comprises a binding agent, hi some embodiments, the binding agent comprises an antibody, an antibody fragment, a nanobody, an aptamer, a peptide, a polymer, an inorganic compound, a small molecule, or a derivative thereof.
[0009] In some embodiments, the modified amino acid comprises a non-naturally occurring chemical modification. In some embodiments, the non-naturally occurring chemical modification is a protecting group.
[0010] In some embodiments, the nanopore binds to a helicase, hi some embodiments, the helicase translocates the modified amino acid or derivative thereof through the nanopore.
[0011] In some embodiments, (b) is performed using a topoisomerase or a polymerase.
[0012] In some embodiments, (b) is carried out using an electric or magnetic field.
[0013] In some embodiments, the polymerizable molecule comprises a nucleic acid molecule, hi some embodiments, the nucleic acid molecule comprises a deoxyribonucleic acid (DNA) molecule, a xenonucleic acid (XNA) molecule, a ribonucleic acid (RNA) molecule, or modified variants thereof.
[0014] In some embodiments, the nanopore comprises a transmembrane protein.
[0015] In some embodiments, the nanogap comprises an inorganic material, hi some embodiments, the nanogap comprises silicon nitride or molybdenum sulfate.
[0016] In some embodiments, the polymerizable molecule is covalently attached to the modified amino acid.
[0017] The method further comprises, prior to (a), generating the modified amino acid from an amino acid of the peptide. In some embodiments, the amino acid is located at a terminal end of the peptide. In some embodiments, the terminal end is the N-terminus. In some embodiments, the modified amino acid comprises the polymerizable molecule attached to the terminal amino acid. In some embodiments, the polymerizable molecule is attached to the terminal amino acid via an isothiocyanate moiety.
[0018] In some embodiments, the method further comprises, prior to (a), (i) providing a linker, wherein the linker comprises an amino acid reactive moiety; and (ii) contacting the amino acid reactive moiety with a proteinogenic amino acid to obtain the modified amino acid. In some embodiments, the linker is attached to a polymerizable molecule. In some embodiments, the method further comprises attaching the linker to the polymerizable molecule. In some embodiments, the linker comprises a first reactive moiety and the polymerizable molecule comprises a second reactive moiety, and the method further comprises reacting the first reactive moiety with the second reactive moiety to generate the linker attached to the polymerizable molecule. In some embodiments, the first reactive moiety or the second reactive moiety comprises a click chemistry moiety. In some embodiments, the polymerizable molecule comprises a nucleic acid linker, wherein the nucleic acid linker comprises a modified nucleobase and a click chemistry moiety. In some embodiments, the amino acid reactive moiety comprises phenylisothiocyanate (PITC). In some embodiments, the proteinogenic amino acid is a terminal amino acid of a peptide. In some embodiments, the method further comprises tethering the modified amino acid to a capture moiety and cleaving the terminal amino acid from the peptide to obtain a remaining peptide and an amino acid-linker-capture moiety (AALC) complex. In some embodiments, the capture moiety is attached to a substrate. In some embodiments, the capture moiety comprises a first nucleic acid molecule and the polymerizable molecule comprises a second nucleic acid molecule. In some embodiments, the first nucleic acid molecule or the second nucleic acid molecule comprises a nucleic acid barcode molecule. In some embodiments, the nucleic acid barcode molecule identifies the peptide. In some embodiments, the nucleic acid barcode molecule comprises temporal information. In some embodiments, the capture moiety comprises a cleavable moiety. In some embodiments, the tethering comprises linking the first nucleic acid molecule to the second nucleic acid molecule. In some embodiments, the linking is performed using a ligase.In some embodiments, the linking comprises (i) providing a splint oligonucleotide comprising a first sequence complementary to at least a portion of the first nucleic acid molecule and a second sequence complementary to at least a portion of the second nucleic acid molecule, and (ii) hybridizing the first sequence to the portion of the first nucleic acid molecule and the second sequence to the portion of the second nucleic acid molecule. In some embodiments, the method further comprises (I) providing an additional linker attached to an additional nucleic acid molecule, (II) contacting the additional linker with the terminal amino acid of the remaining peptide to obtain an additional modified amino acid, and (III) tethering the additional modified amino acid to the AALC complex to obtain the derivative of the modified amino acid. In some embodiments, the method further comprises repeating the steps of providing another additional linker, contacting the another additional linker with the terminal amino acid of the remaining peptide to obtain an additional modified amino acid, and tethering the additional modified amino acid to the AALC complex to obtain the derivative of the modified amino acid, wherein the derivative of the modified amino acid comprises multiple modified amino acids. In some embodiments, the nucleic acid molecule or the additional nucleic acid molecule comprises a barcode sequence. In some embodiments, the nucleic acid molecule comprises a first sequence, and the additional nucleic acid molecule comprises a second sequence, and the first sequence is different from the second sequence. In some embodiments, the first sequence or the second sequence comprises temporal information or spatial information. In some embodiments, the temporal information is the cycle or repeat number index in which the linker or the additional linker is provided.
[0019] In some embodiments, the transport rate of the modified amino acid is slower than the transport rate of the amino acid without the polymerizable molecule.
[0020] In some embodiments, the method further comprises determining the transport rate by measuring a current signal of the nanopore or the nanogap.
[0021] In some embodiments, the method further comprises using the nanopore or nanogap to identify the modified amino acid or derivative thereof, hi some embodiments, the identifying step comprises measuring a current signal from the nanopore or nanogap and determining the amino acid type or post-translationally modified variant of the modified amino acid or derivative thereof.
[0022] In some embodiments, (b) is carried out at a temperature below ambient temperature.
[0023] In another aspect, (a) providing the peptide and a linker, wherein the linker is capable of binding to an amino acid of the peptide; (b) attaching the linker to the amino acid of the peptide; (c) attaching the linker to a capture moiety; (d) cleaving the amino acid from the peptide to obtain an amino acid-linker-capture moiety (AALC) conjugate; (e) providing an additional linker, said additional linker being capable of binding to another amino acid of said peptide; (f) forming an additional amino acid-linker conjugate by attaching the additional linker to the other amino acid; (g) binding the additional amino acid-linker complex to the AALC complex to form a stacked AALC complex; Disclosed herein is a method for processing a peptide, comprising: (a) attaching the peptide to a polymerizable molecule;
[0024] In some embodiments, the linker comprises a reactive moiety. In some embodiments, the polymerizable molecule comprises an additional reactive moiety capable of reacting with the reactive moiety of the linker, and prior to (a), the polymerizable molecule is attached to the linker by reacting the reactive moiety with the additional reactive moiety. In some embodiments, the polymerizable molecule comprises an additional linker comprising the additional reactive moiety, the additional linker comprising a modified nucleobase, and the additional reactive moiety comprises a click chemistry moiety.
[0025] In some embodiments, the method further comprises, prior to (c), providing the polymerizable molecule.
[0026] In some embodiments, the capture moiety is attached to a substrate, in some embodiments, the substrate is substantially planar, in some embodiments, the substrate is a bead.
[0027] In some embodiments, the method further comprises providing the capture moiety, wherein the capture moiety comprises a nucleic acid molecule. In some embodiments, the nucleic acid molecule comprises a DNA molecule, an RNA molecule, an XNA molecule, or a modified variant thereof. In some embodiments, the DNA molecule is single-stranded. In some embodiments, the polymerizable molecule comprises an additional nucleic acid molecule. In some embodiments, (c) comprises attaching the additional nucleic acid molecule to the nucleic acid molecule of the capture moiety. In some embodiments, the attachment is performed using a ligase. In some embodiments, at least a portion of the additional nucleic acid molecule is complementary to at least a portion of the nucleic acid molecule of the capture moiety. In some embodiments, the attachment is performed by hybridizing the at least a portion of the additional nucleic acid molecule to the at least a portion of the nucleic acid molecule. In some embodiments, attaching the additional nucleic acid molecule to the nucleic acid molecule of the capture moiety is performed using a splint oligonucleotide. In some embodiments, the capture moiety comprises a nucleic acid barcode molecule. In some embodiments, the nucleic acid barcode molecule identifies the peptide. In some embodiments, the capture moiety binds to the peptide.
[0028] In some embodiments, the amino acid or the other amino acid is the N-terminal amino acid or the C-terminal amino acid.
[0029] In some embodiments, (b) occurs before (c).
[0030] In some embodiments, (c) occurs before (b).
[0031] In some embodiments, the method further comprises repeating steps (a) through (g).
[0032] In some embodiments, the method further comprises identifying the AALC complex or the stacked AALC complex. In some embodiments, the identifying step comprises using a nanopore or a nanogap. In some embodiments, the identifying step comprises transporting the AALC complex or the stacked AALC complex, or a portion thereof, through the nanopore or the nanogap. In some embodiments, the transporting is performed at a temperature below ambient temperature. In some embodiments, the identifying step comprises determining the amino acid type of the AALC complex or the amino acid type of the stacked AALC complex.
[0033] In some embodiments, the polymerizable molecule comprises a nucleic acid barcode molecule that includes temporal information.
[0034] In some embodiments, the capture moiety comprises a cleavable moiety.
[0035] In yet another aspect, (a) providing a nucleic acid molecule linked to an amino acid or modified amino acid; (b) sequencing the nucleic acid molecule; (c) identifying said amino acid or said modified amino acid; wherein (b) and (c) are performed within 1 minute of each other.
[0036] In some embodiments, (b) comprises generating sequencing reads.
[0037] In some embodiments, (b) or (c) is performed using a nanopore, hi some embodiments, the nanopore is coupled to a helicase, a topoisomerase, an unfoldase, or a polymerase.
[0038] In some embodiments, (b) and (c) are performed using a nanopore. In some embodiments, the nucleic acid molecule is bound to the amino acid or the modified amino acid via a linker. In some embodiments, prior to (a), the linker comprises an amino acid-responsive moiety and is bound to the nucleic acid molecule. In some embodiments, prior to (a), the linker comprises an additional reactive moiety. In some embodiments, the modified amino acid comprises a proteinogenic amino acid, the linker, and the nucleic acid molecule. In some embodiments, the linker comprises at least two atoms. In some embodiments, the linker comprises a phenylisothiocyanate moiety. In some embodiments, the method further comprises, prior to (a), binding the linker to the amino acid or the modified amino acid to form an amino acid-linker conjugate bound to the nucleic acid molecule. In some embodiments, the method further comprises, prior to (b) or (c), binding the amino acid-linker conjugate to a capture moiety to form an amino acid-linker-capture moiety (AALC) conjugate. In some embodiments, the capture moiety comprises a nucleic acid molecule. In some embodiments, the method further comprises, prior to (a), binding the amino acid or the modified amino acid to a peptide and cleaving the AALC complex from the peptide. In some embodiments, the capture moiety comprises a nucleic acid barcode molecule that identifies the peptide. In some embodiments, the method further comprises providing an additional linker, binding the additional linker to an additional amino acid of the peptide to form an additional amino acid-linker complex, and binding the additional amino acid-linker complex to the AALC complex to form a stacked AALC complex. In some embodiments, the capture moiety comprises a cleavable moiety.
[0039] In some embodiments, the nucleic acid molecule comprises temporal information.
[0040] In some embodiments, the identifying step comprises measuring a current signal from the amino acid or the modified amino acid using a nanopore or nanogap, and determining the amino acid type or post-translationally modified variant of the modified amino acid or the modified amino acid. In some embodiments, the measuring is performed at a temperature below ambient temperature.
[0041] In some embodiments, the nucleic acid molecule comprises a modified nucleobase and a click chemistry moiety.
[0042] In another aspect, disclosed herein is a composition comprising a linker covalently attached to a nucleic acid barcode molecule, wherein said linker comprises an amino acid reactive group, and wherein said nucleic acid barcode molecule comprises temporal information.
[0043] In some embodiments, the amino acid reactive group comprises an isothiocyanate. In some embodiments, the isothiocyanate is PITC. In some embodiments, the nucleic acid barcode molecule comprises a modified nucleobase and a click chemistry moiety.
[0044] Another aspect of the present disclosure provides a non-transitory computer-readable medium containing machine-executable code that, when executed by one or more computer processors, performs any of the methods described above or elsewhere herein.
[0045] Another aspect of the present disclosure provides a system comprising one or more computer processors and a computer memory coupled thereto, the computer memory including machine-executable code that, when executed by the one or more computer processors, performs any of the methods described above or elsewhere herein.
[0046]
[0013] Further aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, wherein only illustrative embodiments of the present disclosure have been shown and described. As will be understood, the present disclosure is capable of other and different embodiments, and its various details can be modified in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.
[0047] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent that publications and patents or patent applications incorporated by reference conflict with the disclosure contained herein, the present specification is intended to supersede and / or take precedence over such conflicting material. [Brief explanation of the drawings]
[0048] The novel features of the invention are set forth with particularity in the appended claims. The features and advantages of the present invention will be better understood by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also referred to herein as "figures" and "FIGs"), in which: [Figure 1] 1 illustrates a schematic diagram of an example peptide sequencing workflow disclosed herein. [Figure 2] 1 illustrates a schematic diagram of another example workflow for peptide sequencing as disclosed herein. [Figure 3] 1 illustrates a schematic diagram of another example workflow for peptide sequencing as disclosed herein. [Figure 4] 1A and 1B illustrate schematic diagrams of examples of compositions comprising linkers and polymerizable molecules. [Figure 5] 1 illustrates a computer system that is programmed or otherwise configured to perform the methods provided herein. [Figure 6A] FIG. 1 is a schematic diagram showing examples of polymerizable molecules of three different amino acid types (Asp, Trp, and Tyr) and modified amino acids or derivatives thereof, including amino acid-linker conjugates. [Figure 6B] Example current traces obtained from the nanopore sequencing system for two control molecules and three amino acid-linker-polymerizable molecule complexes are shown. [Figure 6C] 1 shows an example of a current trace as a function of time. [Figure 6D] 1 shows an example of current traces as a function of base pair position. [Figure 6E] Examples of analytical data for classification of different molecular types are shown. [Figure 7A] FIG. 1 is a schematic diagram of another example of a modified amino acid or derivative thereof comprising a polymerizable molecule. [Figure 7B] 1 shows an example of a current trace as a function of time. [Figure 7C] 1 shows an example of current traces as a function of base pair position. [Figure 8A] FIG. 1 is a schematic diagram of another example of a modified amino acid or derivative thereof comprising a polymerizable molecule. [Figure 8B] 1 shows an example of current traces as a function of time and base pair position. [Figure 8C] 1 shows an example of current traces as a function of time comparing two polymerizable molecules containing different linkers. [Figure 8D] t-SNE plot of the dynamic time warping correlation matrix. [Figure 9] Example data on the residence times of two control molecules and three amino acid-linker-polymerizable molecule conjugates are shown. [Figure 10A] FIG. 1 is a schematic diagram of a stacked amino acid-linker-polymerizable molecule conjugate described herein. [Figure 10B]FIG. 1 shows a schematic diagram of an exemplary model of a stacked amino acid-linker-polymerizable molecule complex. [Figure 10C] 1 shows a table of exemplary amino acid types included in a model stacked amino acid-linker-polymerizable molecule complex. [Figure 10D] 1 shows an example of current traces as a function of time and base pair position. [Figure 10E] Examples of classification data for different molecule types are shown. [Figure 11A] 1 shows examples of current traces obtained from polymerizable molecules transporting through a nanopore sequencing system at different temperatures. [Figure 11B] 1 shows histograms of transport duration of molecules through a nanopore at different temperatures. [Figure 11C] 1 shows histograms of transport duration for different sequences of polymerizable molecules at different temperatures. [Figure 11D] 1 shows an example of a current trace of a portion of a polymerizable molecule. [Figure 11E] 1 shows an example of the residence time of some of the polymerizable molecules under two different temperature conditions. DETAILED DESCRIPTION OF THE INVENTION
[0049] While various embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It is understood that various alternatives to the embodiments of the invention described herein may be utilized.
[0050] definition Whenever the terms "at least," "greater than," or "greater than or equal to" precede the first number in a series of two or more numbers, the terms "at least," "greater than," or "greater than or equal to" apply to each and every number in the series. For example, 1, 2, or 3 or more is equivalent to 1 or more, 2 or more, or 3 or more.
[0051] Whenever the term "no more than," "less than," or "less than or equal to" precedes the first number in a series of two or more numbers, the term "no more than," "less than," or "less than or equal to" applies to each number in the series. For example, 3, 2, or 1 or less is equivalent to 3 or less, 2 or less, or 1 or less.
[0052] References to "one embodiment," "an embodiment," "exemplary embodiment," "some embodiments," "certain particular embodiments," "various embodiments," etc. indicate that the embodiments of the disclosed technology so described may include particular features, structures, or characteristics, but not all embodiments necessarily include the particular feature, structure, or characteristic. Furthermore, repeated use of the phrase "in one embodiment" does not necessarily refer to the same embodiment, but may.
[0053] Ranges may be expressed herein as "about," "approximately," or "substantially" from one particular value and / or to another particular value. When such a range is expressed, other exemplary embodiments include from one particular value and / or to the other particular value. Furthermore, the term "about" means within an acceptable range of error for a particular value, as determined by one of ordinary skill in the art, which depends in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, "about" can mean within an acceptable standard deviation, according to practice in the art. Alternatively, "about" can mean within a range of ±20%, preferably ±10%, more preferably ±5%, and even more preferably ±1% of a given value. Alternatively, particularly with respect to biological systems or processes, the term can mean within an order of magnitude, preferably within two-fold, of a value. When particular values are described in this application and claims, unless otherwise specified, the term "about" is implicit and, in this context, means within an acceptable range of error for the particular value.
[0054] "Comprising" or "containing" or "including" means that at least the named compounds, elements, particles, or method steps are present in a composition or product or method, but does not exclude the presence of other such compounds, materials, particles, or method steps, even if they perform the same function as the named ones.
[0055] Throughout this description, various components may be identified having specific values or parameters, but these items are provided as exemplary embodiments. Indeed, many equivalent parameters, sizes, ranges, and / or values may be implemented, and the exemplary embodiments do not limit the various aspects and concepts of the present disclosure. Terms such as "first," "second," "primary," "secondary," etc. do not denote any order, quantity, or importance, but rather are used to distinguish one element from another.
[0056] As used herein, the term "protein" generally refers to a molecule containing two or more amino acids linked by peptide bonds. Proteins may also be referred to as "polypeptides," "oligopeptides," or "peptides." Proteins can be natural or synthetic. Proteins can contain one or more non-natural amino acids, modified amino acids, or non-amino acid linkers. Proteins can contain D-amino acid enantiomers, L-amino acid enantiomers, or both. The amino acids of a protein can be naturally or synthetically modified, such as by post-translational or chemical modification. In some circumstances, different proteins can be distinguished from one another based on different genes expressed in an organism, different primary sequence lengths, or different primary sequence compositions. Nevertheless, proteins expressed from the same gene can be different proteoforms, distinguished, for example, by non-identical length, non-identical amino acid sequence, or non-identical post-translational modifications. Different proteins can be distinguished based on one or both of their gene of origin and proteoform state.
[0057] As used herein, the term "peptide" can refer to any short, single peptide chain. Peptides can be about 100, 95, 90, 85, 80, 75, 70, 65, 60, 55, 50, 45, 40, 35, 30, 25, 20, 15, 10, 5 or less amino acids in length. Peptides can have known or unknown biological functions or activities. Peptides can include natural, synthetic, modified, or degraded proteins or peptides, or combinations thereof.
[0058] As used herein, the term "single analyte" may refer to an analyte that is individually manipulated or distinguished from other analytes. A single analyte may include a biological molecule or a synthetic molecule. A single analyte may include a small molecule. A single analyte may be a single molecule (e.g., a single biological molecule such as a single protein, nucleic acid molecule, affinity reagent, lipid, carbohydrate, etc.), a single complex of two or more molecules (e.g., a multimeric protein having two or more separable subunits, a single protein bound to a nucleic acid molecule, or a single protein bound to an affinity reagent), a single particle, etc. In the context of a composition, system, or method herein, a reference to a "single analyte" herein does not necessarily exclude the application of the composition, system, or method to multiple single analytes that are individually manipulated or distinguished, unless the context or explicitly indicates otherwise.
[0059] As used herein, a "polypeptide" refers to two or more amino acids linked together by a peptide bond. The term "polypeptide," as generally known in the art, includes proteins that have a C-terminus and an N-terminus and can be of synthetic origin or naturally occurring. As used herein, "at least a portion of a polypeptide" refers to two or more amino acids of a polypeptide. A polypeptide can include one or more peptides. Optionally, a portion of a polypeptide includes at least 1, 5, 10, 20, 30, or 50 amino acids, contiguous or spaced, of the complete amino acid sequence of the polypeptide or the entire amino acid sequence of the polypeptide.
[0060] As used herein, "immobilized" generally refers to an association between a polypeptide and a substrate such that at least a portion of the polypeptide and the substrate are held in physical proximity. The term "immobilized" encompasses both indirect or direct association, and may be reversible or irreversible; e.g., the association is optionally covalent or non-covalent.
[0061] As used herein, the term "sample" generally refers to a collected substance or material that contains or is suspected of containing one or more analytes (e.g., a biomolecule, e.g., a polypeptide). A sample may be modified for purposes such as preservation or stability. Lipids may be natural or synthetic. A sample may be treated to separate or remove unwanted fractions or impurities from the analytes. A sample may be concentrated or purified. For example, a sample may include a fraction from a separation process (e.g., chromatography, fractionation, electrophoresis, etc.). Alternatively, a sample is not subjected to treatment to separate or remove unwanted fractions or impurities from the analytes. A sample may be obtained from any suitable source or location, including an organism, a cell, a tissue, a cell preparation, a cell-free composition, or the environment (e.g., air, water, mud, soil, farmland, earth, dust). A sample may be obtained from an organism or part of an organism, such as from a bodily fluid, tissue, or cell. A sample may contain biological and / or non-biological components. As used herein, the term "biological sample" or "biological source" refers to a sample derived primarily from a biological system or organism, such as one or more virus particles, cells (e.g., individualized cells), organelles (e.g., individualized organelles), tissues, body fluids, bone, cartilage, and exoskeleton. A biological sample may contain a majority of biological material on a mass basis, excluding the weight of fluids in the sample. A biological sample may contain one or more proteins and is referred to herein as a protein sample. Biological samples can be obtained from a variety of sources, such as clinical patient samples, including blood, serum, plasma, cerebrospinal fluid (CSF), saliva, mucosal secretions, urine, lymph, sweat, vaginal fluid, and semen. Biological samples may be processed to purify and preserve one or more biomolecules (e.g., proteins, nucleic acids, carbohydrates, lipids, glycoproteins, lipoproteins, metabolites, etc.) from the biological sample. A biological sample (e.g., a protein sample) may be derived from cultured cells, which may be treated or untreated. A biological sample (e.g., a protein sample) may also be obtained from a tissue specimen, such as a biopsy sample, which may optionally be treated to release biological molecules (e.g., proteins) contained therein.Tissue samples may also be derived from in vivo specimens, including fresh, frozen, acute, and fixed tissues.
[0062] As used herein, the terms "antibody" and "immunoglobulin" generally refer to proteins capable of recognizing and binding to specific antigens. Antibodies or immunoglobulins can refer to antibody isotypes, antibody fragments including, but not limited to, Fab, Fv, scFv, and Fd fragments, chimeric antibodies, humanized antibodies, single-chain antibodies, and fusion proteins comprising the antigen-binding portion of an antibody and a non-antibody protein. Antibodies may be detectably labeled with, for example, fluorophores, radioisotopes, enzymes (e.g., peroxidase) that generate a detectable product, fluorescent proteins, nucleic acid barcode sequences, and the like. Antibodies may be further conjugated to other moieties, such as members of specific binding pairs, for example, biotin (a member of the biotin-avidin specific binding pair). Fab', Fv, F(ab')2, and other antibody fragments that retain specific binding to an antigen are also encompassed by the term. Antibodies can exist in a variety of other forms, including, for example, Fv, Fab, and (Fab)2, as well as bifunctional (i.e., bispecific) hybrid antibodies (e.g., Lanzavecchia et al., Eur. J. Immunol. 17, 105 (1987)) and single chains (e.g., Huston et al., Proc. Natl. Acad. Sci. USA, 85, 5879-5883 (1988) and Bird et al., Science, 242, 423-426 (1988), both of which are incorporated herein by reference). (See generally, Hood et al., Immunology, Benjamin, NY, 2nd ed. (1984), and Hunkapiller and Hood, Nature, 323, 15-16 (1986), both of which are incorporated herein by reference.)
[0063] As used herein, "binding" generally refers to a covalent or non-covalent interaction between two molecules (referred to herein as "binding partners," e.g., a substrate and an enzyme or an antibody and an epitope). The binding between binding partners can be specific or non-specific.
[0064] As used herein, "specifically binds" or "binds specifically" generally refers to an interaction between binding partners (e.g., a binding partner and a cognate molecule) such that the binding partners bind to each other but not to another molecule that may be present in the environment (e.g., in a biological sample, in a tissue, in an in vitro assay) under a set of conditions. A specific binding interaction may involve a binding partner binding to a cognate molecule. A specific binding interaction may involve a binding partner binding to its cognate molecule at a significantly or substantially higher level or with higher affinity compared to the binding of the binding partner to a non-cognate molecule. A specific binding interaction may involve a first binding partner that has a higher selectivity for binding to a cognate molecule compared to a non-cognate molecule.
[0065] The terms "nucleic acid," "nucleic acid molecule," "oligonucleotide," and "polynucleotide" can be used interchangeably and generally refer to a polymeric form of any length of natural or synthetic nucleotides, or analogs thereof. Nucleic acid molecules can contain one or more deoxyribonucleotides, deoxyribonucleotide triphosphates, dideoxynucleotide triphosphates, ribonucleotides, hexitol nucleotides, cyclohexane nucleotides, or analogs or combinations thereof. Nucleic acid molecules can include, for example, DNA, RNA, HNA, CeNA, and modified forms thereof. Nucleic acid molecules can contain nucleotides linked by phosphodiester bonds. Nucleic acid molecules can have any two- or three-dimensional structure and can perform any function, known or unknown. Nucleic acid molecules can be single-stranded, double-stranded, or partially double-stranded. Non-limiting examples of polynucleotides include genes, gene fragments, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, non-coding RNA, small interfering RNA, short hairpin RNA, microRNA, scaRNA, ribozymes, riboswitches, viral RNA, complementary DNA (cDNA), cosmid DNA, mitochondrial DNA, chromosomal or genomic DNA, viral DNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, control regions, isolated RNA of any sequence, nucleic acid probes, nucleic acid adapters, and primers.Nucleic acid molecules can be linear, circular, or any other shape. Examples of polynucleotide analogs include, but are not limited to, xenonucleic acids (XNA), bridged nucleic acids (BNA), glycol nucleic acids (GNA), hexitol nucleic acids (HNA), cyclohexane nucleic acids (CeNA), peptide nucleic acids (PNA), gamma PNA, morpholino polynucleotides, locked nucleic acids (LNA), threose nucleic acids (TNA), 2'-O-methyl polynucleotides, 2'-O-alkylribosyl substituted polynucleotides, phosphorothioate polynucleotides, and boronophosphate polynucleotides.Polynucleotide analogs can have purine or pyrimidine analogs, including, for example, 7-deazapurine analogs, 8-halopurine analogs, 5-halopyrimidine analogs, or universal base analogs that can pair with any base, including hypoxanthine, nitroazole, isocarbostyril analogs, azolecarboxamide, and aromatic triazole analogs, or base analogs with additional functionality, such as a biotin moiety for affinity binding.
[0066] As used herein, the term "amino acid" generally refers to organic compounds that combine to form proteins or peptides. Amino acids generally contain an amine group, a carboxylic acid group, and a side chain specific to each amino acid that functions as a monomeric subunit of a peptide. Amino acids can include the 20 naturally occurring standard or canonical amino acids as well as non-standard amino acids. Naturally occurring standard or canonical amino acids include alanine (A or Ala), cysteine (C or Cys), aspartic acid (D or Asp), glutamic acid (E or Glu), phenylalanine (F or Phe), glycine (G or Gly), histidine (H or His), isoleucine (I or Ile), lysine (K or Lys), leucine (L or Leu), methionine (M or Met), asparagine (N or Asn), proline (P or Pro), glutamine (Q or Gln), arginine (R or Arg), serine (S or Ser), threonine (T or Thr), valine (V or Val), tryptophan (W or Trp), and tyrosine (Y or Tyr). Amino acids may be L- or D-amino acids. Non-standard amino acids can be naturally occurring or chemically synthesized modified amino acids, amino acid analogs, amino acid mimetics, non-standard proteinogenic amino acids, or non-proteinogenic amino acids. Examples of non-standard amino acids include, but are not limited to, selenocysteine, pyrrolysine, and N-formylmethionine, 3-amino acids, homoamino acids, proline and pyruvate derivatives, 3-substituted alanine derivatives, glycine derivatives, ring-substituted phenylalanine and tyrosine derivatives, linear core amino acids, and N-methylamino acids.
[0067] As used herein, the term "amino acid type" generally refers to one of the naturally occurring standard or canonical amino acids, e.g., one member of the group consisting of alanine (A or Ala), cysteine (C or Cys), aspartic acid (D or Asp), glutamic acid (E or Glu), phenylalanine (F or Phe), glycine (G or Gly), histidine (H or His), isoleucine (I or Ile), lysine (K or Lys), leucine (L or Leu), methionine (M or Met), asparagine (N or Asn), proline (P or Pro), glutamine (Q or Gln), arginine (R or Arg), serine (S or Ser), threonine (T or Thr), valine (V or Val), tryptophan (W or Trp), and tyrosine (Y or Tyr), derivatives thereof, and modified forms of any of the foregoing amino acids. The term "amino acid type" may be used herein to distinguish between amino acids that contain different side chain groups rather than identical amino acids (e.g., amino acids at different positions in a single peptide that have the same side chain). Amino acid types may include naturally occurring standard amino acids or modified canonical amino acids, e.g., post-translationally, epigenetically, or chemically or enzymatically modified.
[0068] As used herein, the term "post-translational modification" refers to a modification that occurs on a peptide after translation. Post-translational modifications may be covalent modifications or enzymatic modifications. Examples of post-translational modifications include acylation, acetylation, alkylation (including methylation), biotinylation, butyrylation, carbamylation, carbonylation, deamidation, deimination, diphthamide formation, disulfide bridge formation, eliminylation, flavin attachment, formylation, gamma carboxylation, glutamylation, glycylation, glycosylation, glycosylphosphatidylinositol anchoring (glypiation), heme C attachment, hydroxylation, Modifications include, but are not limited to, hypusine formation, iodination, isoprenylation, lipid addition, lipoylation, malonylation, methylation, myristoylation, oxidation, palmitoylation, PEGylation, phosphopantetheinylation, phosphorylation, prenylation, propionylation, retinylidene Schiff base formation, S-glutathionylation, S-nitrosylation, S-sulfenylation, selenation, succinylation, sulfination, ubiquitination, and C-terminal amidation. Post-translational modifications include modifications of the amino and / or carboxyl termini of peptides. Modifications of terminal amino groups include, but are not limited to, desamino, N-lower alkyl, N-di-lower alkyl, and N-acyl modifications. Modifications of the terminal carboxy group include amide, lower alkyl amide, dialkyl amide, and lower alkyl ester modifications (e.g., where lower alkyl is C1-C4 alkyl). Post-translational modifications also include, but are not limited to, modifications of amino acids between the amino and carboxy termini, such as those described above. The term post-translational modification can also include peptide modifications that include one or more detectable labels. Post-translational modifications can be natural or synthetic.
[0069] As used herein, the term "binding agent" refers to a molecule that binds, associates, integrates with, recognizes, or binds to another molecule, e.g., a nucleic acid molecule, peptide, polypeptide, protein, carbohydrate, synthetic molecule, or small molecule. A binding agent may bind to a macromolecule, or a component or feature of a macromolecule. A binding agent may form a covalent or non-covalent association with a molecule, macromolecule, or component or feature of a macromolecule. A binding agent may also be a chimeric binding agent composed of two or more types of molecules, such as a nucleic acid molecule-peptide chimeric binding agent, a carbohydrate-peptide chimeric binding agent, or a lipid-peptide chimeric binding agent. A binding agent may be a naturally occurring molecule, a synthetically produced molecule, or a recombinantly expressed molecule. A binding agent may bind to a single monomer or subunit of a macromolecule (e.g., a single amino acid of a peptide) or to multiple binding subunits of a macromolecule (e.g., a dipeptide, tripeptide, or higher-order peptide of a longer peptide, polypeptide, or protein molecule). A binding agent may bind to a linear molecule or a molecule with a three-dimensional structure (also called a conformation). For example, an antibody binder may bind to a linear peptide, polypeptide, or protein, or to a conformational peptide, polypeptide, or protein. The binder may bind to the N-terminal peptide, C-terminal peptide, or intervening peptide of a peptide, polypeptide, or protein molecule. The binder may bind to the N-terminal amino acid, C-terminal amino acid, or intervening amino acid of a peptide molecule. The binder preferably binds to chemically modified or labeled amino acids over unmodified or unlabeled amino acids. For example, the binder may preferentially bind amino acids modified with acetyl, guanyl, dansyl, PTC, DNP, SNP, etc. over amino acids lacking such moieties. The binder may bind to post-translational modifications of a peptide molecule. The binder may exhibit selective binding to a component or property of a macromolecule (e.g., a binder may selectively bind to one of the 20 possible naturally occurring amino acid residues and bind with very low affinity or not at all to the other 19 naturally occurring amino acid residues).A binder may exhibit less selective binding, where the binder is capable of binding to multiple components or properties of a macromolecule (e.g., the binder may bind to two or more different amino acid residues with similar affinity). The binder may include a tag, which may be attached to the binder via a linker.
[0070] As used herein, the term "linker" generally refers to a molecule or moiety involved in the attachment of two or more molecules. A linker can facilitate covalent or non-covalent interactions between two or more molecules. A linker may be a cross-linking agent. A linker may be monofunctional, bifunctional, trifunctional, tetrafunctional, or multifunctional. A linker may be or include a nucleotide, nucleotide analog, amino acid, peptide, polypeptide, or non-nucleotide chemical moiety, such as an organic or inorganic compound. A linker may include a polymer, such as polyethylene glycol (PEG), polyethylene, polypropylene, polyvinyl chloride, polystyrene, or other organic or inorganic polymer. A linker may include one or more reactive ends, such as an amine-reactive group, a carboxyl-reactive group, a sulfhydryl-reactive group, a hydroxyl-reactive group, or the like. Alternatively, a linker may not include a reactive end. In some examples, a linker may be used to attach different molecule types, such as peptides and nucleic acid molecules, lipids and peptides, different biomolecule types, such as carbohydrates and peptides, non-biomolecule types, or biomolecules and non-biomolecules. For example, a linker may be used to connect a binder to a tag, a tag to a polymer (e.g., a peptide, a nucleic acid molecule), a polymer to a solid support, a tag to a solid support, etc. The linker may connect two molecules via an enzymatic or chemical reaction (e.g., click chemistry). The linker may connect more than two molecules, for example, via an enzymatic or chemical reaction.
[0071] As used herein, the term "conjugated" generally refers to a covalent or ionic interaction between two entities, e.g., molecules, compounds, or combinations thereof.
[0072] As used herein, the term "tag" generally refers to a molecule or moiety conjugated to a molecule. A tag can include a detectable label, such as a fluorophore or fluorescent protein, a radioisotope, an enzyme (e.g., a chromogenic or fluorescent protein, a protein that can catalyze a chromogenic substrate), a mass tag, a hapten (e.g., biotin, digoxigenin, urushiol, fluorescein), a vibrational or FTIR tag (e.g., an alkyne group). A tag can include a biomolecule, such as a nucleic acid molecule, a protein, a lipid, a carbohydrate, or a combination thereof. A tag can include one or more nucleic acid molecules and can optionally encode information about the tag or the molecule (e.g., a binding agent such as an antibody) to which the tag is conjugated. For example, a tag can include a nucleic acid barcode molecule. A tag can include an organic or inorganic compound.
[0073] As used herein, the term "barcode" generally refers to an identifying characteristic that can be used to distinguish between similar entities. A barcode can include a nucleic acid molecule of about 2 to about 30 bases. A barcode can include about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114 9, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, The barcodes may include nucleic acid molecules of 97, 98, 99, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, or 150 or more bases, which may provide unique identifier tags or origin information for molecules (e.g., proteins, polypeptides, peptides), binders, sets of binders from a binding cycle, sample molecules, sets of samples, molecules in a compartment (e.g., a droplet, bead, partition, or separate location), macromolecules in a set of compartments, fractions of macromolecules, sets of fractions of macromolecules, spatial regions or sets of spatial regions, libraries of macromolecules, or libraries of binders. Barcodes may be artificial or naturally occurring sequences, including peptides, proteins, protein complexes, carbohydrates, and synthetic polymer materials. In certain embodiments, each barcode in a barcode population is different. In other embodiments, a portion of the barcodes in the population of barcodes are different, for example, at least about 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, or 99% of the barcodes in the population of barcodes are different. The population of barcodes may be randomly or deliberately generated. The population of barcodes may include error-correcting barcodes.Barcodes can be used to computationally deconvolute sequence reads derived from individual molecules, samples, libraries, etc. Barcodes can contain multiplexed information, e.g., originating from different samples, compartments, individual molecules, etc. Barcodes can also be used to deconvolute collections of molecules distributed into small compartments to enhance mapping. For example, rather than mapping peptides to a proteome, peptides can be mapped to protein molecules or protein complexes, the sample or partition from which the peptides originated, etc. Barcodes can include any useful sequence, including repetitive sequences (e.g., poly-A, poly-T, poly-C, poly-G regions), and barcodes can also include non-repetitive sequences.
[0074] As used herein, a "sample barcode," also referred to as a "sample tag," generally refers to a barcode molecule that contains identifying information for the sample from which the barcoded molecule is derived.
[0075] As used herein, "spatial barcode" generally refers to a barcode molecule that contains information that identifies the region of a 2D or 3D sample (e.g., a tissue section) from which a molecule originates or is derived.Spatial barcodes can be used for molecular pathology on tissue sections.Spatial barcodes can enable multiplex sequencing of multiple samples or libraries from tissue sections.
[0076] As used herein, "temporal barcode" generally refers to a barcode molecule that contains time-based information about the barcoded molecule. The types of time-based data encoded in a temporal barcode can include, among others, the lifespan of the barcode molecule, the time of sample collection, the time or duration from the start of an experiment or induction by a stimulus, information about the age of a cell or tissue, the sequence of interactions between molecules, the time or cycle or round (e.g., of a repetitive process) at which the barcode molecule is provided, and the like. Different types of barcodes (e.g., spatial, temporal, cell-specific) can be combined into one multiplexed barcode.
[0077] As used herein, the term "nucleic acid sequence" or "oligonucleotide sequence" generally refers to a contiguous string of nucleotide bases and may refer to a particular arrangement of the nucleotide bases in relationship to each other when they appear in an oligonucleotide. Similarly, the term "polypeptide sequence" or "amino acid sequence" may refer to a contiguous string of amino acids and may refer to a particular arrangement of the amino acids in relationship to each other when they appear in a polypeptide.
[0078] A "nucleotide sequence" according to the present invention may comprise any polymer or oligomer of nucleotides, including pyrimidine and purine bases, such as cytosine, thymine, and uracil, as well as adenine and guanine, and combinations thereof. The nucleotide sequence may comprise any deoxyribonucleotide, ribonucleotide, hexitol-nucleotide, cyclohexane-nucleotide, peptide nucleic acid component, and any chemical variant thereof, such as methylated, 7-deazapurine analogs, 8-halopurine analogs, hydroxymethylated, or glycosylated forms of these bases. The polymer or oligomer may be heterogeneous or homogeneous in composition, isolated from natural sources, or artificially or synthetically produced. Furthermore, the nucleotide sequence may be DNA, RNA, HNA, CeNA, or a mixture thereof, and may exist permanently or transiently in single- or double-stranded form, including homoduplexes, heteroduplexes, and hybrid states.
[0079] The terms "complementary" or "complementarity" generally refer to polynucleotides (i.e., a sequence of nucleotides) related by the base-pairing rules. For example, the sequence "5'-AGT-3'" is complementary to the sequence "5'-ACT-3'." Complementarity can be "partial," in which only a portion of the nucleic acids' bases match according to the base-pairing rules, or there can be "complete" or "total" complementarity between nucleic acids. The degree of complementarity between nucleic acid strands can significantly affect the efficiency and strength of hybridization between nucleic acid strands under defined conditions.
[0080] As used herein, the term "hybridization" generally refers to the pairing of complementary nucleic acids.Hybridization and the strength of hybridization (for example, the strength of the association between nucleic acids) are affected by factors such as the degree of complementarity between nucleic acids, the stringency of the conditions involved, and the melting temperature of the formed hybrid.Hybridization methods involve the annealing of a certain nucleic acid with another complementary nucleic acid, for example, based on Watson-Crick base pairing.
[0081] As used herein, the term "proteomics" generally refers to the quantitative and / or qualitative analysis of the proteome within a sample, such as a biological sample, e.g., a cell, tissue, or bodily fluid. Proteomics can include the analysis of the spatial distribution of proteins within a sample (e.g., a cell and / or tissue). Proteomics can include the examination of the dynamic state of the proteome, e.g., how one or more proteins change over time.
[0082] The terminal amino acid at one end of a peptide chain having a free amino group may be referred to herein as the "N-terminal amino acid" (NTAA). The terminal amino acid at the other end of the chain having a free carboxyl group may be referred to herein as the "C-terminal amino acid" (CTAA). The amino acids constituting a peptide may be numbered sequentially, and the length of the peptide is "n" amino acids. As used herein, in some instances, an NTAA may be considered the nth amino acid (also referred to herein as the "nth NTAA"). In such cases, the next amino acid is the n-1 amino acid, then n-2 amino acids, etc., continuing down the length of the peptide from the N-terminus to the C-terminus. Alternatively, a CTAA may be considered the nth amino acid (also referred to herein as the "CTAA"). In such cases, the next amino acid is the n-1 amino acid, then n-2 amino acids, etc., continuing down the length of the peptide from the C-terminus to the N-terminus. The NTAA, CTAA, or both may be modified or labeled with a chemical moiety.
[0083] As used herein, the terms "determining," "measuring," "assessing," and "assaying" are used interchangeably and include both quantitative and qualitative determinations.
[0084] As used herein, the term "unique molecular identifier" or "UMI" generally refers to a molecular barcode that contains indexing information. A UMI is a unique molecular identifier that is about 3 to about 150 bases in length (3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 1 , 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, or 150 bases). A UMI can provide a unique identifier tag for each molecule (e.g., peptide, binding agent, nucleic acid molecule) that contains or associates with the UMI. A UMI can comprise a random sequence (e.g., a random N-mer).
[0085] As used herein, a "derivative" of a nucleic acid molecule generally refers to a nucleic acid molecule derived from an original nucleic acid molecule. A derivative may have the same or substantially the same nucleotide sequence as the original nucleic acid molecule, or may contain a complementary or partially complementary sequence to the original nucleic acid molecule. A derivative may be the same type of nucleic acid (e.g., DNA or RNA) as the original nucleic acid molecule, or a different type of nucleic acid (e.g., cDNA generated from an RNA molecule). A derivative of a nucleic acid molecule may exhibit sequence identity with the original nucleic acid molecule. A derivative nucleic acid molecule may also be subjected to additional processing from the original nucleic acid molecule, such as chemical or enzymatic modification, splicing, ligation, polymerization, fragmentation, tagging (e.g., using a transposase), digestion, etc.
[0086] A derivative polypeptide or peptide may be derived from a parent polypeptide (or peptide). A derivative may contain the same amino acid sequence as the original polypeptide, or the sequence may differ. A derivative polypeptide results from or is subject to additional processing from the parent polypeptide, such as chemical or enzymatic modification. A derivative polypeptide may include one or more tags, nucleic acid molecules, barcode molecules, labels (e.g., detectable labels), fluorophores, probes, linkers, post-translational modifications, chemical protecting groups, or other chemical moieties.
[0087] As used herein, the term "compartment" generally refers to a physical region or volume that separates or isolates a subset of molecules from a sample of molecules. For example, a compartment may separate individual cells from other cells, or a subset of a sample's proteome from the remainder of the sample's proteome. A compartment may be an aqueous compartment (e.g., a microfluidic droplet), a solid compartment (e.g., a plate, a tube, a vial, a picotiter well on a gel bead, or a microtiter well), or a separate region on a surface. A compartment may contain one or more beads to which macromolecules can be immobilized.
[0088] As used herein, the terms "solid support," "solid surface," or "solid substrate" or "substrate" refer to any solid material, including porous and non-porous materials, with which molecules can be directly or indirectly associated. Molecules may be associated with the substrate through covalent or non-covalent interactions, or a combination thereof. Substrates may be two-dimensional (e.g., planar) or three-dimensional (e.g., gel matrices or beads). Non-limiting examples of solid supports include beads, microbeads, arrays, glass surfaces, silicon surfaces, plastic surfaces, filters, membranes, nylon or other polymers, silicon wafer chips, flow-through chips, flow cells, biochips containing signal transduction electronics, channels, microtiter wells, ELISA plates, spin interference disks, nitrocellulose membranes, nitrocellulose-based polymer surfaces, polymer matrices, nanoparticles, or microspheres. Materials for solid supports include, but are not limited to, acrylamide, agarose, cellulose, nitrocellulose, glass, gold, quartz, polystyrene, polyethylene vinyl acetate, polypropylene, polymethacrylate, polyethylene, polyethylene oxide, polysilicate, polycarbonate, Teflon, fluorocarbon, nylon, silicone rubber, polyanhydrides, polyglycolic acid, polylactic acid, polyorthoesters, functionalized silanes, polypropyl fumarate, collagen, glycosaminoglycans, polyamino acids, dextran, or any combination thereof. Solid supports further include thin films, membranes, bottles, dishes, shaped polymers such as fibers, woven fibers, tubes, particles, beads, microspheres, microparticles, or any combination thereof. For example, when the solid surface is a bead, the bead can include, but is not limited to, a ceramic bead, a polystyrene bead, a polymer bead, a methylstyrene bead, an agarose bead, an acrylamide bead, a solid core bead, a porous bead, a paramagnetic bead, a glass bead, or a controlled pore bead. The bead can be spherical or irregular. The size of the bead can range from nanometers, e.g., 100 nm, to millimeters, e.g., 1 mm.In certain embodiments, the beads range in size from about 0.2 microns to about 200 microns, or from about 0.5 microns to about 5 microns. In some embodiments, the beads have diameters of about 1 μm, 1.5 μm, 2 μm, 2.5 μm, 2.8 μm, 3 μm, 3.5 μm, 4 μm, 4.5 μm, 5 μm, 5.5 μm, 6 μm, 6.5 μm, 7 μm, 7.5 μm, 8 μm, 8.5 μm, 9 μm, 9.5 μm, 10 μm, 11 μm, 12 μm, 13 μm, 14 μm, 15 μm, 16 μm, 17 μm, 18 μm, 19 μm, 20 μm, 21 μm, 22 μm, 23 μm, 24 μm, 25 μm, 26 μm, 27 μm, 28 μm, 29 μm, 30 μm, 31 μm, 32 μm, 33 μm, 34 μm, 35 μm, 36 μm, 37 μm, 38 μm, 39 μm, 40 μm, 41 μm, 42 μm, 43 μm, 44 μm, 45 μm, 46 μm, 47 μm, 48 μm, 49 μm, 50 μm, 51 μm, 52 μm, 53 μm, 54 μm, 55 μm, 56 μm, 57 μm, 58 μm, 59 μm, 60 μm, 61 μm, 62 μm, 63 μm, 64 μm, 65 μm, 66 μm, m, 16μm, 17μm, 18μm, 19μm, 20μm, 21μm, 22μm, 23μm, 24μm, 25μm, 26μm, 27μm, 28μm, 29μm, 30μm, 31μm, 32μm, 33μm, 34μm, 35μm, 36μm, 37μm, 38μm, 39μm, 40μm, 41μm, 42μm, 43μm, 44 μm, 45μm, 46μm, 47μm, 48μm, 49μm, 50μm, 51μm, 52μm, 53μm, 54μm, 55μm, 56μm, 57μm, 58μm m, 59μm, 60μm, 61μm, 62μm, 63μm, 64μm, 65μm, 66μm, 67μm, 68μm, 69μm, 70μm, 71μm, 72μm, 73 μm, 74 μm, 75 μm, 76 μm, 77 μm, 78 μm, 79 μm, 80 μm, 81 μm, 82 μm, 83 μm, 84 μm, 85 μm, 86 μm, 87 μm, 88 μm, 89 μm, 90 μm, 91 μm, 92 μm, 93 μm, 94 μm, 95 μm, 96 μm, 97 μm, 98 μm, 99 μm, or 100 μm. In certain embodiments, a "bead" solid support may refer to an individual bead or a plurality of beads.
[0089] As used herein, "sequencing" generally refers to (A) determining the order of nucleotides (base sequences) in a nucleic acid sample, e.g., DNA or RNA, or (B) determining the order of amino acids in all or part of a polymer, such as a protein, peptide, or other multimeric molecule. Many techniques are available, including Sanger sequencing and high-throughput sequencing (HTS). Sanger sequencing can involve sequencing via detection by (capillary) electrophoresis, in which up to 384 capillaries can be sequenced in a single run. High-throughput sequencing involves the simultaneous parallel sequencing of tens of thousands or more sequences. HTS can be defined as next-generation sequencing (NGS), a technology based on solid-phase pyrosequencing, or as next-generation sequencing based on single-nucleotide real-time sequencing (SMRT). HTS technologies are available, such as those offered by Roche, Illumina, and Applied Biosystems. Additional high-throughput sequencing technologies are described by and / or available from Helicos, Pacific Biosciences, Complete Genomics, Ion Torrent Systems, Oxford Nanopore Technologies, Nabsys, ZS Genetics, GnuBio.
[0090] As used herein, "next-generation sequencing" generally refers to a high-throughput sequencing method that allows millions to billions of molecules to be sequenced in parallel. Examples of next-generation sequencing methods include sequencing by synthesis, sequencing by ligation, sequencing by hybridization, polony sequencing, ion semiconductor sequencing, nanopore sequencing, and pyrosequencing. By attaching a primer to a solid substrate and a complementary sequence to the nucleic acid molecule, the nucleic acid molecule can be hybridized to the solid substrate via the primer, and then amplified by using a polymerase in separate regions on the solid substrate to generate multiple copies (these groups are sometimes referred to as polymerase colonies or polony). Therefore, during the sequencing process, a nucleotide at a specific position can be sequenced multiple times (e.g., hundreds or thousands of times), and this depth of coverage is referred to as "deep sequencing." Examples of high-throughput nucleic acid sequencing technologies include platforms offered by Illumina, BGI, Qiagen, ThermoFisher, and Roche, including formats such as parallel bead arrays, sequencing-by-synthesis, sequencing-by-ligation, capillary electrophoresis, electronic microchips, "biochips," microarrays, parallel microchips, and single molecule arrays, as reviewed by Service (Science 311:1544-1546, 2006).
[0091] As used herein, "analyzing" a macromolecule refers to quantifying, characterizing, distinguishing, or a combination thereof, all or part of the components of the molecule (e.g., a biomolecule, such as a polymer, protein, amino acid, or nucleic acid molecule). For example, analyzing a peptide, polypeptide, or protein can include determining all or part of the peptide's (contiguous or non-contiguous) amino acid sequence. Analyzing a macromolecule may also include partially identifying the components of the macromolecule. For example, partially identifying amino acids in a protein sequence can identify amino acids in the protein as belonging to a subset of possible amino acids. Analysis may be performed sequentially, for example, starting with analysis of the nth NTAA and then proceeding to the next amino acid of the peptide (i.e., n-1, n-2, n-3, etc.). In such an example, sequencing may be performed by cleaving the nth NTAA, thereby converting the n-1th amino acid of the peptide to the N-terminal amino acid (referred to herein as the "n-1th NTAA"). Similarly, analysis of a peptide may begin from the C-terminus toward the N-terminus, with each round of cleavage from the C-terminus creating a new CTAA. Cleavage of the nth CTAA converts the n-1th amino acid of the peptide to the C-terminal amino acid, referred to herein as the "n-1th CTAA." Analyzing a peptide may also include determining the presence and frequency of post-translational modifications on the peptide, which may or may not include information about the sequential order of post-translational modifications on the peptide. Analyzing a peptide may also include determining the presence and frequency of epitopes within the peptide, which may or may not include information about the sequential order or location of the epitopes within the peptide. Analysis of a peptide may involve combining various types of analyses (e.g., obtaining epitope information, amino acid sequence information, post-translational modification information, or any combination thereof).
[0092] As used herein, the term "array" generally refers to a collection of molecules attached to one or more solid supports such that molecules at one address can be distinguished from molecules at other addresses. An array can include different molecules, each located at a different address on the solid support. Alternatively, an array can include separate solid supports that function as addresses, each carrying a different molecule, and the different molecules can be identified according to the location of the solid support on the surface to which it is attached or in a liquid, such as a fluid stream. Molecules in an array can be, for example, nucleic acids such as SNAP, polypeptides, proteins, peptides, oligopeptides, enzymes, ligands, or receptors such as antibodies, functional fragments of antibodies, or aptamers. Addresses in an array can optionally be optically observable, and in some configurations, adjacent addresses can be optically distinguishable when detected using the methods or apparatus described herein.
[0093] As used herein, the term "functionalized" refers to any material or substance that has been modified to contain a functional group. A functionalized material or substance can be naturally or synthetically functionalized. For example, a polypeptide can be naturally functionalized with a phosphate group, an oligosaccharide (e.g., glycosyl, glycosylphosphatidylinositol, or phosphoglycosyl), nitrosyl, methyl, acetyl, a lipid (e.g., glycosylphosphatidylinositol, myristoyl, or prenyl), ubiquitin, or other naturally occurring post-translational modifications. A functionalized material or substance can be functionalized for any given purpose, including altering its chemical nature (e.g., altering its hydrophobicity or changing its surface charge density) or its reactivity (e.g., being able to react with a moiety or reagent to form a covalent bond with the moiety or reagent).
[0094] As used herein, the terms "click reaction," "click chemistry," or "bioorthogonal reaction" refer to a single-step, thermodynamically favorable conjugation reaction utilizing a biocompatible reagent. Click reactions may not utilize toxic or bioincompatible reagents (e.g., acids, bases, heavy metals) and may not produce toxic or bioincompatible by-products. Click reactions may utilize aqueous solvents or buffers (e.g., phosphate buffer, Tris buffer, saline buffer, MOPS, etc.). Click reactions may be thermodynamically favorable, with a negative Gibbs free energy of reaction of, for example, less than about -5 kilojoules / mole (kJ / mole), less than -10 kJ / mole, less than -25 kJ / mole, less than -50 kJ / mole, less than -100 kJ / mole, less than -200 kJ / mole, less than -300 kJ / mole, less than -400 kJ / mole, or less than -500 kJ / mole. Exemplary bioorthogonal and click reactions are described in detail in International Publication No. WO 2019 / 195633 A1, which is incorporated herein by reference in its entirety. Examples of click reactions include metal-catalyzed azide-alkyne cycloaddition, strain-promoted azide-alkyne cycloaddition, strain-promoted azide-nitrone cycloaddition, strained alkene reaction, thiolene reaction, Diels-Alder reaction, inverse electron demand Diels-Alder reaction, [3 + 2] cycloaddition, [4 + 1] cycloaddition, nucleophilic substitution, dihydroxylation, thiophosphorus reaction, photoclick, nitrone dipol cycloaddition, norbornene cycloaddition, oxanobornadiene cycloaddition, tetrazine ligation, and tetrazole photoclick reaction.Exemplary functional groups or reactive handles utilized to perform the click reaction include alkenes (e.g., linear alkenes or cyclic alkenes such as trans-cyclooctene (TCO)), alkynes (e.g., linear alkynes or cycloalkynes such as cyclooctynes or derivatives thereof, such as aza-dimethoxycyclooctyne (DIMAC), symmetrical pyrrolocyclooctyne (SYPCO), pyrrolocyclooctyne (PYRROC), difluorocyclooctyne (DIFO), α,α-bis(trifluoromethyl)pyrrolocyclooctyne (TRIPCO), bicyclo[6.1.0]nonyne (BCN), dibenzocyclooctyne (DIBO), difluorinated cyclooctynes (DIFO, difluorobenzocyclooctyne (DIFBO), dibenzoazacyclo-octyne (DBCO), difluoro-aza-dibenzocyclo octyne (F2-DIBAC), biaryl-azacyclooctyne (BARAC), difluorodimethoxydibenzocyclooctynol (FMDIBO), difluoromethoxydibenzocyclooctione (keto-FMDIBO), and 3,3,6,6-tetramethylthiacycloheptyne (TMTH), TMTH-sulfoximine (TMTHSI), azide, epoxide, amine, thiol, nitrone, isonitrile, isocyanide, aziridine, activated ester, and tetrazine, triazole, and combinations, variations, or derivatives thereof. The click chemistry moieties may be subjected to conditions sufficient to react a first click chemistry moiety with a second click chemistry moiety, such as the provision of a metal catalyst, a suitable solvent, pH, temperature, ion concentration, or light / energy, for any useful period of time.
[0095] As used herein, the terms "group" and "moiety" are intended to be synonymous when used in reference to the structure of a molecule. The term refers to a component or portion of a molecule. The term does not necessarily represent the relative size of the component or portion compared to the molecule, unless otherwise indicated. The term does not necessarily represent the relative size of the component or portion compared to other components or portions of the molecule, unless otherwise indicated. A group or moiety may contain one or more atoms.
[0096] As used herein, "primer" generally refers to a nucleic acid molecule capable of priming the synthesis of a nucleic acid molecule (e.g., DNA or RNA). A primer may be single-stranded. A primer may contain one or more recognition sites for proteins (e.g., polymerizing enzymes, restriction enzymes, cleaving enzymes, etc.) to bind to the primer or to a primer hybridized to a template strand. A primer may contain DNA, RNA, or other nucleic acid analogs, as well as non-standard bases (e.g., spacer moieties, uracil, abasic sites). A primer may optionally contain any number of functional sequences, such as a sequencing primer sequence (e.g., P5 or P7 sequence), a sequencing primer binding sequence, a lead sequence (e.g., R1 or R2 sequence), a restriction site, a transposition site (e.g., mosaic end sequence), etc.
[0097] "Amplification" or "amplifying" generally refers to a polynucleotide amplification reaction, i.e., a population of polynucleotides replicated from one or more starting sequences. Amplifying can refer to various amplification reactions, including, but not limited to, polymerase chain reaction (PCR), linear polymerase reaction, nucleic acid sequence-based amplification, rolling circle amplification, and similar reactions. An amplification reaction can produce an amplicon.
[0098] As referred to herein, "adapter" generally refers to a short nucleic acid molecule (e.g., about 10-100 base pairs in length). Adapters may comprise short double-stranded DNA molecules. Adapters may be attached to the ends of DNA fragments or amplicons, for example, via polymerization or ligation. Adapters may comprise synthetic oligonucleotides, e.g., oligonucleotides having nucleotide sequences at least partially complementary to each other. Adapters may have blunt or staggered ends (also referred to herein as 3' or 5' "overhang sequences" or "sticky ends," or blunt and staggered ends). Adapters may be attached to fragments (e.g., by ligation) to provide adapter-ligated fragments, which may serve as starting points for subsequent manipulations, such as amplification or sequencing. The adapter may be functionalized, for example, conjugated with a tag, a probe, a detectable label, an affinity capture reagent (eg, biotin or streptavidin).
[0099] As used herein, the terms "transport" and "transport" generally refer to the movement of molecules through a medium (e.g., a gas, liquid, or solid). Molecular transport can occur spontaneously (e.g., by diffusion, Brownian motion, etc.). Alternatively, or in addition, molecular transport can occur by the application of force or pressure, such as frictional force, tension force, normal force, air resistance force, spring force, gravity, electric force, magnetic force, acoustic force (e.g., acoustophoresis), etc. In some examples, molecular transport can be achieved by the application of pressure-driven flow or electrophoretic force. Transport can occur through a liquid or through a solid substrate (e.g., pores or gaps).
[0100] As used herein, abbreviations for natural 1-enantiomer amino acids are conventional and are as follows: alanine (A, Ala), arginine (R, Arg), asparagine (N, Asn), aspartic acid (D, Asp), cysteine (C, Cys), glutamic acid (E, Glu), glutamine (Q, Gln), glycine (G, Gly), histidine (H, His), isoleucine (I, Ile), leucine (L, Leu), lysine (K, Lys), methionine (M, Met), phenylalanine (F, Phe), proline (P, Pro), serine (S, Ser), threonine (T, Thr), tryptophan (W, Trp), tyrosine (Y, Tyr), valine (V, Val). Unless otherwise specified, X can represent any amino acid. In some embodiments, X can be asparagine (N), glutamine (Q), histidine (H), lysine (K), or arginine (R). References to these amino acids are also in the form "[amino acid][residue / residues)]" (e.g., lysine residue, lysine residues, leucine residue, leucine residues, etc.).
[0101] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of this disclosure, suitable methods and materials are described below.
[0102] Protein sequencing by binding to polymerizable molecules Provided herein are methods, systems, compositions, and kits for characterizing proteins. The disclosed methods, systems, compositions, and kits enable the analysis of individual amino acids in a protein, thereby providing information about the composition or sequence of amino acids in the protein (also referred to herein as "protein sequencing"). In some aspects, the disclosed methods include providing an analyte, such as an amino acid, attached to a polymerizable molecule and identifying the polymerizable molecule, the analyte (e.g., amino acid), or both. In some aspects, the methods, systems, compositions, and kits provided herein involve the use of a nanoscale object (e.g., a nanopore or nanogap) to identify amino acids that can be attached to the polymerizable molecule. Advantageously, the polymerizable molecule can modify or adjust the properties of the amino acid, for example, during transport adjacent to or through the nanoscale object, so that the amino acid can be more accurately identified. Thus, the disclosed methods, systems, and compositions provide a simpler and more accurate approach to protein sequencing.
[0103] In one aspect of the present disclosure, provided herein is a method for processing a modified amino acid, the method comprising: providing a modified amino acid, wherein the modified amino acid comprises a polymerizable molecule; and transporting the modified amino acid or a derivative thereof through a nanopore or nanogap. The method may further comprise measuring a signal from the nanopore or nanogap and using the signal to determine the identity of the modified amino acid or a derivative thereof. In some examples, the polymerizable molecule facilitates or modulates the transport rate of the modified amino acid through the nanopore or nanogap, making the modified amino acid more detectable compared to an unmodified amino acid. In some examples, transport of the modified amino acid through the nanopore or nanogap may result in an improved measured signal (e.g., current blockade) or an increased signal-to-noise ratio of the measured signal compared to an unmodified amino acid or an amino acid not bound to a polymerizable molecule.
[0104] Modified amino acid: A modified amino acid or derivative thereof may be derived from or part of a protein or peptide. For example, a modified amino acid may include or be derived from an amino acid located at the end (N-terminus or C-terminus) of a peptide, or may include or be derived from an amino acid located within a peptide. A modified amino acid may include a protein-producing amino acid or a derivative thereof that has undergone any number of modifications. Non-limiting examples of modifications include chemical modifications (e.g., protecting groups), biological modifications (e.g., modifications brought about by post-translational modifications, enzymatic treatment, or digestion), physical modifications (e.g., mutations brought about by irradiation, heat, etc.), and the like. In some examples, a modified amino acid or derivative thereof comprises or is bound to a binding agent such as an antibody, antibody fragment, nanobody, aptamer, peptide, small molecule, inorganic compound, polymer, or any variant or combination thereof. In some examples, a modified amino acid comprises a non-naturally occurring chemical modification. For example, the modified amino acid may contain a protecting group, non-limiting examples of which include methyl, formyl, ethyl, acetyl, t-butyl, anisyl, benzyl, trifluoroacetyl, N-hydroxysuccinimide, t-butyloxycarbonyl, benzoyl, 4-methylbenzyl, thioanidyl, thiocresyl, benzyloxymethyl, 4-nitrophenyl, benzyloxycarbonyl, 2-nitrobenzoyl, 2-nitrophenylsulfenyl, 4-toluenesulfonyl, pentafluorophenyl, diphenylmethyl, 2-chlorobenzyloxycarbonyl, 2,4,5-trichlorophenyl, 2-bromobenzyloxycarbonyl, 9-fluorenylmethyloxycarbonyl, triphenylmethyl, or 2,2,5,7,8-pentamethyl-chroman-6-sulfonyl groups. In some examples, the modified amino acid comprises a proteinogenic amino acid or a derivative thereof linked to a polymerizable molecule. In some examples, the modified amino acid comprises a proteinogenic amino acid or a derivative thereof, a linker, and a polymerizable molecule.
[0105] The modified amino acids may include any useful modification. The modification may be natural (e.g., post-translational) or non-natural, such as by labeling or tagging with an amino acid or amine-reactive agent, or a linker containing an amino acid or amine-reactive agent. Examples of amino acid or amine-reactive agents include isothiocyanates (e.g., PITC, NITC), 1-fluoro-2,4-dinitrobenzene (DNFB), dansyl chloride, 4-sulfonyl-2-nitrofluorobenzene (SNFB), acetylating agents, acylating agents, alkylating agents, guanidinating agents, thioacetylating agents, thioacylating agents, thiobenzoylating agents, or derivatives or combinations thereof. Alternatively, or in addition, one or more modified amino acids may include an adduct (e.g., a polymer such as PEG, a polymerizable molecule such as a nucleic acid molecule, a nanoparticle or nanotube, a peptide, or a protein), a lipid, a carbohydrate, a metabolite, a fluorophore, a hapten, a quencher, a tag (e.g., a fluorescent tag, a magnetic tag, a radioactive tag), a barcode, or other moiety. In some examples, the modified amino acid may include a modification that facilitates the recruitment of an enzyme (or ribozyme or DNAzyme) that recognizes or cleaves a terminal amino acid, such as the NTAA or CTAA of the peptide. For example, the terminal amino acid of the peptide may be modified with a sugar to recruit a lectin or lectin-binding protease. In another example, one or more modified amino acids may comprise or be bound to a nucleic acid molecule having a first sequence complementary to a second sequence contained in an oligo-binding protease. Hybridization between the first and second sequences may facilitate localized recruitment of the protease to the amino acid to be cleaved. In yet another example, the peptide may be modified with phenylisothiocyanate (PITC), which may enable recruitment and cleavage of the modified amino acid by edmanase. In some examples, the amino acid modification may include an epitope tag that can facilitate binding of a binding agent to the modified amino acid. Examples of such epitope tags include fluorophores, nucleic acid molecules, peptides, haptens, polymers, chemical moieties, or other additional molecules.
[0106] The methods described herein may further include a step of generating a modified amino acid. In some examples, the modified amino acid includes a proteinogenic amino acid or a derivative thereof, a linker, and a polymerizable molecule. In one example, a peptide including a terminal amino acid may be contacted with (i) a first reactive moiety capable of reacting with the amino acid and (ii) a linker including a second reactive moiety. A polymerizable molecule including a third reactive moiety capable of reacting with the second reactive moiety may be provided before, during, or after the reaction of the first reactive moiety with the terminal amino acid. The second reactive moiety and the third reactive moiety may include click chemistry moieties capable of reacting with each other (e.g., azide and DBCO, azide and BCN, alkyne and DBCO, TCO and tetrazine, etc.). Reaction of the terminal amino acid with the linker and the linker with the polymerizable molecule may yield a modified amino acid including the terminal amino acid, the linker, and the polymerizable molecule. In some examples, cleavage of the terminal amino acid from the peptide may be performed, and the modified amino acid may include a cleavage product including the cleaved, optionally derivatized amino acid, the linker, and the polymerizable molecule.
[0107] The polymerizable molecule of the modified amino acid may alter the properties of the amino acid attached thereto, compared to the unmodified amino acid or the amino acid to which the polymerizable molecule is attached. Such altered properties (e.g., size, charge, aspect ratio, etc.) may affect the interaction or residence time of the modified amino acid with the nanopore or nanogap. In some embodiments, the polymerizable molecule alters the transport rate of the modified amino acid when it is transported through the nanopore or nanogap, compared to the transport rate of the unmodified amino acid. The transport rate of the modified amino acid may be faster or slower than the transport rate of the unmodified amino acid. In some examples, the transport rate of the modified amino acid is slower than the transport rate of the unmodified amino acid. For example, the transport rate of the modified amino acid may be reduced by at least about 0.1%, at least about 1%, at least about 2%, at least about 3%, at least about 4%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, at least about 10%, at least about 20%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 100%, at least about 200%, at least about 500%, at least about 1000%, at least about 5000%, or more. The transport rate of the modified amino acid may be reduced by various percentages, such as about 5% to 20%, or about 500% to 3000%. In some cases, the transport rate of the modified amino acid is not constant, e.g., may be within a range of varying percentages over which the transport rate varies, and may vary depending on the properties (e.g., size, charge, polarity, etc.) of the polymerizable molecule, the amino acids surrounding the modified amino acid, the linker, or any other components of the molecule.
[0108] Altered transport rates of modified amino acids compared to unmodified amino acids (e.g., amino acids without attached polymerizable molecules) can make the modified amino acids more detectable. For example, polymerizable molecules attached to amino acids can slow or alter their transport rate through a nanopore or nanogap, allowing for more accurate current measurements. Alternatively, or in addition, polymerizable molecules can alter the charge, size, or aspect ratio of the attached amino acids, thereby altering the current signature and improving the detectability of the amino acids, polymerizable molecules, or both. Alternatively, or in addition, the medium (e.g., liquid, buffer, solution) in contact with the nanopore may contain one or more agents that can alter or modulate the transport rate of modified amino acids. For example, the medium may contain ions, salts, or other molecules that can selectively or preferentially alter the interaction of modified amino acids with the nanopore compared to unmodified amino acids.
[0109] In some cases, upon transport through a nanopore or nanogap, the measured signal-to-noise ratio (SNR) of the modified amino acid is higher than the measured SNR of the unmodified amino acid (e.g., a proteinogenic amino acid, an amino acid without a polymerizable molecule attached). The increase in SNR can be at least about 0.1%, at least about 1%, at least about 2%, at least about 3%, at least about 4%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, at least about 10%, at least about 20%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 100%, at least about 200%, at least about 500%, at least about 1000%, at least about 5000% or more. The SNR of the modified amino acid can be increased by various percentages, such as, for example, about 5%-20%, about 500%-1000%, etc.
[0110] In some examples, the measured signal (e.g., current blockade) of the modified amino acid when it is transported through the nanopore or nanogap is substantially different from the measured signal of the unmodified amino acid. In some examples, the measured signal of the modified amino acid when it is transported through the nanopore or nanogap is substantially different from the measured signal of the polymerizable molecule alone. The difference in the measured signals can be measured by any useful metric, such as fold change, percentage change, absolute current measurement, signal-to-noise ratio, etc.
[0111] The polymerizable molecule may comprise any useful polymerizable moiety. The polymerizable moiety may comprise a natural or synthetic polymer (organic or inorganic) or a biopolymer. The polymer may comprise one or more monomer types. When more than one monomer type is used, the polymer may form an alternating copolymer structure, a periodic copolymer structure, a random copolymer structure, a block copolymer structure, a chain or graft copolymer, or other useful structure. The polymer may be linear or non-linear. The polymer may comprise polyethylene glycol (PEG), PEG-diacrylate, PEG-acrylate, PEG-thiol, PEG-azide, PEG-alkyne, polyacrylamide, agarose, collagen, fibrin, gelatin, chitosan, hyaluronic acid, alginate, polyvinyl alcohol, or another polymer.
[0112] In some examples, the polymerizable molecule comprises a biomolecule, such as a protein or peptide, a nucleic acid molecule, such as DNA or RNA, a carbohydrate, or a lipid chain. In some examples, the polymerizable molecule comprises a nucleic acid molecule, which can comprise any useful number and type of nucleotides, e.g., standard and non-standard bases, and the number of nucleotides can be adjusted based on the intended purpose. For example, the length of the nucleic acid molecule can be adjusted to change the properties (e.g., volume, aspect ratio, charge, etc.) of the amino acid to which the polymerizable molecule is tethered. The nucleic acid molecule may additionally or alternatively comprise any useful functional sequence, including, but not limited to, a barcode sequence or other identification sequence, a UMI sequence, an enzyme recognition site (e.g., a translocation site, a restriction site), a spacer sequence, a sequencing primer sequence, a lead sequence, or a primer sequence. The nucleic acid molecule may comprise standard bases, non-standard bases, natural bases, synthetic bases, abasic sites, or a combination thereof. In some examples, modified amino acids or derivatives thereof are produced using a cyclic process, as described elsewhere herein. Thus, the nucleic acid molecule may comprise information regarding the number of rounds or cycles of the cyclic process. In some cases, the nucleic acid molecule comprises a nucleic acid barcode molecule, which may contain useful information regarding amino acid identity, temporal information, spatial information, etc.
[0113] The polymerizable molecule (e.g., a nucleic acid molecule) may be covalently or non-covalently attached to the amino acid. Attachment may be performed using any suitable chemistries and reaction conditions, and may include the use of a linker. In one example, the nucleic acid molecule may include a first reactive group, e.g., a first click chemistry moiety, as described elsewhere herein, and may be contacted with a linker including a second reactive group, e.g., a second click chemistry moiety, capable of reacting with the first reactive group. The linker may also include an additional reactive group that is attached to an amino acid (e.g., a terminal amino acid) and can optionally cleave the amino acid from the peptide. For example, the additional reactive group can be a thiocyanate conjugate, e.g., an isothiocyanate (ITC) such as phenylisothiocyanate (PITC) or naphthylisothiocyanate (NITC), or an aldehyde group, e.g., ortho-phthalaldehyde (OPA), 2,3-naphthalenedicarboxaldehyde (NDA), a guanidinylating agent, dinitrofluorobenzene (DNFB), dansyl chloride, or other amino acid reactive group. The linker may be reacted with an amino acid of a peptide, e.g., NTAA or CTAA. The use of a linker containing at least two reactive groups can enable (i) tethering of the amino acid to the linker and (ii) tethering of the linker to the nucleic acid molecule (see, e.g., Figure 4). In some cases, the linker may be provided pre-tethered to the nucleic acid molecule before contacting with the amino acid. In some cases, conjugation of a polymerizable molecule to an amino acid, with or without a linker, can change the chemical structure of the amino acid. For example, when using a linker containing an isothiocyanate moiety, an amino acid can be derivatized to a thiocarbamyl group (e.g., under alkaline conditions), a thiazolone group (e.g., under acid conditions), a thiohydantoin group, or other chemical moiety, thereby generating a modified amino acid having a polymerizable molecule attached thereto.
[0114] Linker: One or more linkers may be used to attach a polymerizable molecule to an amino acid or its derivative to generate a modified amino acid. In some examples, the linker comprises a click chemistry moiety. The click chemistry moiety may include any suitable bioorthogonal moiety, such as an alkene, an alkyne (e.g., a cycloalkyne such as alkyne, DBCO, and BCN), an azide, an epoxide, an amine, a thiol, a nitrone, an isonitrile, an isocyanide, an aziridine, an activated ester, and a tetrazine, as described elsewhere herein, and combinations, variations, or derivatives thereof. The linker may be subjected to conditions sufficient to react a first click chemistry moiety with a second click chemistry moiety, such as the provision of a metal catalyst, a suitable solvent, pH, temperature, ion concentration, or light / energy, for any useful period of time.
[0115] The linker may comprise an amino acid reactive moiety. The amino acid reactive moiety of the linker may be any useful moiety that allows the reactive moiety to be conjugated to an amino acid and optionally cleave the amino acid. In some examples, the third reactive moiety may react with the terminal amino acid (e.g., NTAA or CTAA). In such examples, the third reactive moiety may comprise any primary amine or carboxyl group reactive group, including, but not limited to, isocyanate, acyl azide, NHS ester, sulfonyl chloride, aldehyde, glyoxal, epoxide, oxirane, carbonate, aryl halide, imido ester, carbodiimide, anhydride, phenyl ester, isothiocyanate (e.g., phenyl isothiocyanate, sodium isothiocyanate, ammonium isothiocyanate (e.g., tetrabutylammonium isothiocyanate, diphenylphosphoryl isothiocyanate)), acetyl chloride, cyanogen bromide, carboxypeptidase, azide, alkyne, DBCO, maleimide, succinimide, thiol-thiol disulfide bond, tetrazine, TCO, vinyl, methylcyclopropene, acryloyl, aryl, etc. Other examples of amino acid reactive groups are described in U.S. Patent Publication No. 2020 / 0217853, which is incorporated herein by reference in its entirety.
[0116] The linker may contain additional useful moieties. For example, the linker may contain a releasable or cleavable moiety that can facilitate removal of the amino acid-linker complex from the polymerizable molecule, or a portion thereof, or from the substrate. Such a releasable or cleavable moiety may include, for example, a disulfide bond, which may be releasable upon contact with a reducing agent (e.g., DTT, TCEP). In some cases, the linker may be attached to the polymerizable molecule via a releasable or cleavable moiety instead of, or in addition to, attachment via a click chemistry moiety. In this way, the bond between the polymerizable molecule and the linker may be reversible. The linker may further contain any number of spacing moieties, such as polymers (e.g., PEG, PVA, polyacrylamide), aminohexanoic acid, nucleic acids, alkyl chains, etc. Such spacing moieties may increase the distance between any other moiety of the linker, such as the amino acid reactive group and the polymerizable molecule reactive group.
[0117] In some examples, the polymerizable molecule includes a linker. The linker may be used, for example, to attach a reactive moiety to the polymerizable molecule, which can react with the reactive moiety of another linker. In one such example, the polymerizable molecule, e.g., a nucleic acid molecule, may include a linker containing a click chemistry moiety. The linker containing the click chemistry moiety can be attached to the polymerizable molecule using any useful technique, e.g., by incorporating a linker-conjugated nucleotide or nucleoside, and can be located at any useful position (e.g., the 5' end, the 3' end, or the center of the polymerizable molecule). For example, click-functionalized nucleotides or nucleosides, e.g., ethynyldeoxyuridine, octadiynyldeoxyuridine, can be incorporated into the backbone of a DNA or RNA molecule. In this way, the click chemistry moiety of the polymerizable molecule can then be attached to another linker containing a complementary click chemistry moiety and an amino acid reactive group (e.g., isothiocyanate, dansyl chloride, DNFB, etc.).
[0118] The linker may contain any number of spacing moieties, such as alkyl chains, polymer spacers (e.g., PEG), nucleic acid or oligospacers, or other useful spacing moieties that may be useful for adjusting the size or molecular weight of the linker. For example, the linker may contain at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, or more spacing moieties (e.g., hydrocarbon units, PEG units, nucleotides, or spacer sequences, etc.). The linker may contain up to about 100, up to about 10, up to about 9, up to about 8, up to about 7, up to about 6, up to about 5, up to about 4, up to about 3, up to about 2, or up to about 1 spacing moiety. The linker may contain any useful number of functional groups, for example, for attaching multiple molecules. The linker may include at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, or more functional groups.
[0119] Nanopore / Nanogap: Characterization of modified amino acids may be performed using nanoscale technologies such as nanopores, nanogaps, and nanochannels. In some examples, the nanopore, nanogap, or nanochannel may be located on a membrane in an ionic solution. Signals may be measured from the nanopore, membrane, or surrounding solution. For example, the conductivity, current, current blockage, or other parameter within the nanopore may be monitored as a function of time. When a molecule enters the nanopore, nanogap, or nanochannel, a change in conductivity, current, or other parameter may occur, providing information about the molecule (e.g., size, charge, aspect ratio, volume). Each amino acid or modified amino acid, or a subset of amino acids or modified amino acids, may generate a unique signal that is distinguishable from other amino acids or modified amino acids. In this way, a unique signal signature may be assigned to an amino acid or modified amino acid to determine the identity of the amino acid, polymerizable molecule, or both.
[0120] In some cases, modified amino acids may be analyzed multiple times, for example, by transporting the modified amino acids through the same or different nanopores or nanogaps and measuring the current or conductivity. For example, repeated reading of the modified amino acids can be beneficial for improving the accuracy of the readings and for identifying the modified amino acids (or polymerizable molecules). In such cases, the modified amino acids may be transported through one or more nanopores at least two times, at least three times, at least four times, at least five times, at least six times, at least seven times, at least eight times, at least nine times, at least ten times, at least 20 times, at least 50 times, at least 100 times, at least 200 times, or even more times. Alternatively, or in addition, the modified amino acids may be ratcheted back and forth within the nanopore using, for example, a processive enzyme such as phi29 polymerase, as described, for example, in Cherf et al. 2012. Nat. Biotechnol. 30(4):344-348, which is incorporated herein by reference in its entirety.
[0121] Nanopores, nanogaps, or nanochannels may be generated from organic materials, such as pore-forming proteins or transmembrane proteins. Such proteins may be natural, synthetic, or engineered. Examples of natural organic nanopores include wild-type aerolysin, α-hemolysin, mycobacterial porins (e.g., MspA porin), Phi29 connector channels, fragaceatoxin C, cytolysin A, ferric hydroxyamic acid uptake component A, curli-specific gene G, outer membrane porin G, viral DNA packaging motors, and the like. Alternatively, or in addition, the nanopore, nanochannel, or nanogap may comprise an engineered variant of a natural nanopore. In some embodiments, the nanopore, nanogap, or nanochannel is composed of inorganic materials. For example, solid-state nanopores may be fabricated from dielectric materials such as silicon compounds (e.g., silicon nitride, silicon dioxide), aluminum compounds (e.g., aluminum oxide), titanium compounds (e.g., titanium oxide), molybdenum compounds (e.g., molybdenum sulfate), hafnium, graphene, etc. Nanopores, nanochannels, or nanogaps may assume any useful form factor or shape, such as gaps or channels in membranes, capillaries, etc., and may be generated using any suitable process, such as ion beam sculpting, electron beam exposure, etc. Nanopores, nanochannels, or nanogaps may comprise elastomeric materials.
[0122] In some examples, the nanopore, nanogap, or nanochannel is coupled to a protein. The protein may be a molecular motor, which may facilitate the movement of the modified amino acid, or a portion thereof (e.g., a polymeric material), through the nanopore. In a non-limiting example, the modified amino acid may comprise an amino acid attached to a nucleic acid molecule, optionally via a linker, and the nanopore may be attached to a helicase or a molecule having helicase activity. The helicase may be used to transport the nucleic acid molecule and the attached amino acid through the nanopore. In some examples, the molecular motor may increase, decrease, or otherwise alter the transport rate of the modified amino acid through the nanopore compared to the unmodified amino acid. In other examples, the molecular motor may comprise a topoisomerase, polymerase, nuclease (e.g., an endonuclease, such as a restriction endonuclease or a Cas protein, or an exonuclease), unfoldase, mitotic spindle protein (e.g., a nuclear division apparatus protein, kinesin, dynein), or other motor protein (e.g., myosin), variants thereof, or other polymer-processing protein. In some examples, the protein may comprise a protease or proteosome, which may allow for "chop-n-drop" or cleavage of modified amino acids or portions thereof prior to transport within the nanopore.
[0123] Alternatively, or in addition, transport of a molecule described herein (e.g., modified amino acids, polymeric molecules) through a medium or through a nanopore or nanogap may be facilitated by the application of a force. For example, molecules may be transported by the application of pressure (e.g., pressure-driven flow), an electric field (e.g., via electrophoresis, electroosmotic flow, isoelectric focusing), a magnetic field (e.g., using ferrofluids, magnetic particles), or light (e.g., optoelectronics).
[0124] Commercially available nanopore systems may be used in the methods described herein, for example, nanopore systems from Oxford Nanopore Technologies (ONT), such as the MinION, VolTRAX, GridION, PromethION, MinIT, Flongle, or Q-Line products, may be used to identify and characterize the modified amino acids described herein.
[0125] Intramolecular extension: Modified amino acids or derivatives thereof may be produced using an intramolecular extension process, for example, using one or more linkers and polymerizable molecules. In intramolecular extension, individual amino acids, clusters of amino acids, or small peptides (e.g., dipeptides, tripeptides, or quadripeptides) of a protein (e.g., a peptide or protein analyte) may be sequentially removed and retethered to increase the distance between individual amino acids or clusters of amino acids. Advantageously, performing an intramolecular extension process on one or more amino acids of a peptide may avoid or overcome some problems associated with nanopore sequencing of peptides. For example, increasing the spacing between amino acids may reduce the number of amino acids entering the nanopore at a given instance, thereby reducing the amount of overlapping signal resulting from the number of amino acids in the nanopore. Furthermore, increasing the spacing between amino acids may disrupt intramolecular interactions that convolute current blocking signals, thereby enabling more accurate signal output from the nanopore.
[0126] In one example, a method for intramolecular extension may include providing a peptide comprising a plurality of amino acids, a linker (e.g., as described elsewhere herein), a polymerizable molecule (e.g., as described elsewhere herein), and a capture moiety. The linker may be configured to bind to (i) an amino acid of the peptide (e.g., NTAA or CTAA) and (ii) the polymerizable molecule. The method may further include contacting the linker with the amino acid and the polymerizable molecule. Alternatively, or in addition, the linker may be provided pre-tethered to the polymerizable molecule and then reacted with the amino acid. The linker may bind to an amino acid of the peptide to form an amino acid-linker complex. The amino acid-linker complex may then be bound to the capture moiety via the polymerizable molecule. For example, both the polymerizable molecule and the capture moiety may comprise nucleic acid molecules, which may be bound via hybridization, ligation, or both. In some examples, the method may further include attaching a polymerizable molecule to the capture moiety, cleaving an amino acid from the peptide to obtain an amino acid-linker-capture moiety (AALC) complex, and optionally repeating these steps. In examples where the steps are repeated, an additional linker configured to attach to (i) an additional amino acid of the peptide (e.g., n-1 NTAA or n-1 CTAA) and (ii) an additional polymerizable molecule may be provided. The method may further include contacting the additional linker with an additional amino acid to produce an additional amino acid-linker complex. The additional polymerizable molecule may be attached to the linker before, during, or after attaching the linker to the additional amino acid. The additional polymerizable molecule may be configured to attach to the AALC complex (e.g., via the polymerizable molecule of the AALC complex). Thus, in some instances, following the generation of the additional linker-additional amino acid conjugate, an additional linker-additional amino acid conjugate may be attached to the AALC conjugate, thereby generating a stacked AALC conjugate, and the additional amino acid may be cleaved from the peptide before, during, or after generation of the stacked AALC conjugate.Thus, a "modified amino acid" can refer to an amino acid-linker conjugate, an amino acid-linker-polymerizable molecule, an AALC conjugate, a stacked AALC conjugate (comprising two or more amino acids), or a truncated AALC conjugate, or a combination or portion thereof (e.g., only the included amino acid portion, only the amino acid-linker conjugate portion, etc.).
[0127] In some cases, intramolecular extension of a peptide or protein may occur across multiple capture moieties. For example, a substrate may be used to facilitate intramolecular extension of a peptide bound to the substrate. For example, a substrate containing multiple capture moieties may be provided, and in some cases, the capture moieties are arranged adjacent to the peptide or protein. The first amino acid (e.g., the nth NTAA or the nth CTAA) of the peptide or protein may be bound to the first capture moiety (e.g., via a first linker and a first polymerizable molecule), the second amino acid (e.g., the n-1th NTAA or the n-1th CTAA) may be bound to the second capture moiety (e.g., via a second linker and a second polymerizable molecule), and the third amino acid (e.g., the n-2th NTAA or the n-2th CTAA) may be bound to the third capture moiety. In another example, a first amino acid (e.g., the nth NTAA) can bind to a first capture moiety, a second amino acid (e.g., the n-1th NTAA) can bind to an AALC complex (e.g., from the nth NTAA), thereby generating a stacked AALC complex, and a third amino acid (e.g., the n-2th NTAA) can bind to a second capture moiety. As will be understood, any number of amino acids (or modified amino acids) can be bound to any number of capture moieties (or the resulting AALC or stacked AALC complex).
[0128] Capture moiety: The capture moiety may be bound to an amino acid, linker, or polymerizable molecule. Binding of the amino acid, linker, or polymerizable molecule to the capture moiety may involve covalent or non-covalent interactions. Binding may occur through interactions between binding pairs, such as biotin and avidin (or streptavidin), antigens or epitopes and antibodies or antibody fragments, cyclodextrins and small hydrophobic molecules (e.g., alkanes, benzenes, polycyclics), cucurbiturils and adamantanammonium or trimethylammoniomethylferrocene, cyclophanes (e.g., calixarenes, cavitands, pyrararene, tetralactams), etc. In some embodiments, binding of the amino acid, linker, or polymerizable molecule to the capture moiety occurs through binding of nucleic acid molecules (e.g., hybridization to each other or to splint molecules).
[0129] In some cases, the capture moiety includes an additional polymerizable molecule (e.g., a nucleic acid molecule or a peptide). In one such example, both the polymerizable molecule of the modified amino acid and the capture moiety may include a nucleic acid molecule. The nucleic acid molecules may be bound to each other, for example, directly via complementary base pairing or via a splint molecule and optional ligation.
[0130] The nucleic acid molecules of the capture moiety may contain natural, non-natural, or engineered nucleotide bases. For example, the nucleic acid molecules may contain pseudo-complementary bases, bridged nucleic acids, xenonucleic acids, locked nucleic acids, peptide nucleic acids (PNAs), γ-PNAs, morpholinos, etc., as described elsewhere herein. The capture moiety may contain one or more functional sequences, including, but not limited to, a priming sequence, a sequencing sequence (e.g., a P5 or P7 sequence), a sequencing lead sequence (e.g., an R1 or R2 sequence), a mosaic end sequence, a transposase recognition sequence, a cleavage site (e.g., a restriction site), a UMI, a blocking group, a spacer sequence, a barcode sequence, or other functional sequence. In some examples, the capture moiety contains a cleavable or releasable moiety, such as a restriction enzyme recognition site, an abasic site, a uracil cleavable using USER® or uracil DNA glycosylase, a disulfide bond releasable upon the addition of a reducing agent, etc. In some examples, the capture moiety includes a partial restriction site; for example, the capture moiety can include a first partial restriction site and the polymerizable molecule can include a second partial restriction site; upon binding or ligation of the polymerizable molecule to the capture moiety, the two partial restriction sites can generate a complete restriction site such that the individual molecules (capture moiety and polymerizable molecule) are not individually cleaved by restriction digestion, but the ligated or combined product is. In some examples, the capture moiety includes a barcode sequence containing any useful information, such as the identity of the peptide being analyzed, temporal information, spatial information, etc.
[0131] In some examples, the capture moiety is provided bound to a substrate. In one example, the substrate includes one or more identical capture nucleic acid molecules, and these identical capture nucleic acid molecules can function as capture moieties for generating one or more AALC complexes, for example, for the terminal amino acid, the n-1th amino acid, etc. In some examples, commercially available substrates, such as beads (e.g., DNA beads or barcode beads), flow cells, or chips, such as Illumina® HiSeq, iSeq, MiniSeq, NextSeq, NovaSeq, etc., can be used as the substrates described herein. In some examples, the capture moiety can include additional useful sequences, such as primer sequences (e.g., P5 or P7 sequences) or lead sequences (e.g., R1 or R2).
[0132] The capture moiety may be attached to the substrate using any useful technique. In some examples, the capture moiety comprises a substrate-tethering group, or a linker, or an additional functional group. In some examples, the capture moiety comprises a substrate-tethering group, e.g., biotin, a substrate, including streptavidin, a click chemistry moiety such as an azide that can bind to the substrate, or a nucleic acid molecule that includes a complementary click chemistry moiety that can react with the click chemistry moiety of the substrate-tethering group. The capture moiety may further comprise a binding sequence to which another nucleic acid molecule (e.g., a polymerizable molecule that is part of or attached to a modified amino acid) is attached. In some examples, the capture moiety comprises a single-stranded oligonucleotide or a single-stranded region to which a complementary oligonucleotide can hybridize. The complementary oligonucleotide may comprise a detectable label (e.g., a fluorophore) that allows for detection of the capture moiety.
[0133] Alternatively, the capture moiety may not be bound to a substrate. For example, the capture moiety may directly bind to the peptide to be analyzed or the peptide undergoing intramolecular expansion. In such instances, the capture moiety may further comprise a nucleic acid barcode molecule that encodes the identity of the peptide or the sample or partition from which the peptide originates.
[0134] Cleavage: In some examples, the method further includes cleaving the amino acid or modified amino acid from the peptide. Cleavage of the amino acid or modified amino acid can be achieved using any suitable mechanism, such as through the application of a stimulus. The stimulus can be, for example, a chemical stimulus, a biological stimulus, a thermal stimulus (e.g., application of heat), a light stimulus, a physical or mechanical stimulus, or other types of stimuli or combinations of stimuli. In some examples, the stimulus includes a chemical stimulus, such as a change in pH, application of an acid or base, addition of a dissolving agent, an initiator, a radical generator, a reducing agent, etc. In some examples, the chemical stimulus includes application of a Lewis acid (e.g., boron trifluoride, boron trifluoride etherate, boron trichloride, boron tribromide, boron triiodide, scandium trifluoride). In some examples, the stimulus includes a biological stimulus, such as an enzyme (e.g., edmanase, protease, endonuclease) or a ribozyme or DNAzyme that can cleave the modified amino acid or catalyze cleavage.
[0135] In some examples, the method can include using a linker containing an amino acid reactive group (e.g., PITC), coupling the amino acid reactive group of the linker to an amino acid, and cleaving the amino acid from the peptide using a stimulus (e.g., a change in pH or temperature). In one example, a linker containing a PITC moiety can be attached to NTAA under mildly alkaline conditions to generate a phenylthiocarbamoyl (PTC) derivative of NTAA, and cleavage of NTAA from the peptide can be achieved using the Edman degradation reaction (e.g., application of an acid such as trifluoroacetic acid or boron triflate, optionally with heat), to generate a thiazolinone (ATZ) derivative or a phenylthiohydantoin (PTH) derivative. As described elsewhere herein, the linker can also include a moiety or molecule (e.g., a polymeric molecule such as a nucleic acid molecule) that can be attached to a capture moiety, such that an amino acid or modified amino acid can be attached to the capture moiety, thereby generating an AALC complex.
[0136] Considering the harsh reaction conditions of standard Edman degradation, the polymerizable molecules (e.g., nucleic acid molecules, peptides, lipids) described herein may contain alterations or modifications to make them more resistant to the reaction conditions. For example, the nucleic acid molecule may contain primarily pyrimidines (e.g., thymine, cytosine, uracil), which are more resistant to acid decomposition and heat than purines (e.g., adenine and guanine). For example, the nucleic acid molecule may contain at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% thymine or cytosine. Alternatively, or in addition, the canonical nucleotides may be substituted or contain acid-resistant nucleotide analogs, such as hexitol nucleic acids.
[0137] Alternative degradation chemistries may also be utilized. Milder degradation under basic conditions for N-terminal amino acid removal may involve the use of triethylamine acetate in acetonitrile or other solvents such as water, N,N-dimethylformamide (DMF), or mixtures of solvents. Alternatively, degradation may be achieved using a thioacylation approach, i.e., weaker acid reagents such as trichloroacetic acid (pKa 0.66) or dichloroacetic acid (pKa 1.35), or alternative basic reaction conditions, such as acid-base pairs such as N,N-diisopropylethylamine (DIPEA), pyridine, or acetic acid derivatives.
[0138] C-terminal degradation methods are also provided herein. C-terminal degradation may include Edman-like degradation techniques. C-terminal degradation may involve the use of an activating reagent that reacts with the C-terminal carboxyl group of a peptide and a derivatizing agent (e.g., thiocyanate to generate peptide-thiocyanate or peptide-thiohydantoin). Non-limiting examples of activating reagents include acetyl chloride and acetic anhydride. Alternatively, or in addition, single-step C-terminal derivatization of a peptide to a peptidyl-thiohydantoin may be performed, for example, using the Schlack-Kumpf technique, in which the peptide is reacted with thiocyanic acid (e.g., in acetone) to generate the peptidyl-thiohydantoin. The peptide-thiohydantoin may be cleaved, for example, using basic conditions, to generate an amino acid thiohydantoin and the remaining peptide.
[0139] Amino acid cleavage may also be achieved using enzymatic or enzyme analog (e.g., ribozyme or DNAzyme) techniques. Examples of enzymatic cleavage include the use of edmanase (e.g., modified cruzain), aminopeptidases (e.g., Pfu aminopeptidase I, PhTET aminopeptidase, P. horikoshii aminopeptidase), metalloenzyme aminopeptidase, acylpeptide hydrolase, tRNA synthetase, endopeptidase, carboxypeptidase, etc. Enzymes, ribozymes, or DNAzymes may be modified or engineered to recognize modified amino acids, for example, amino acids conjugated with chemical moieties (e.g., PITC, NITC, dansyl chloride, SNFB, DNP, SNP, biotin, streptavidin, nucleic acid molecules, lipids, carbohydrates, acetyl groups, acyl groups, guandinylating agents, etc.).
[0140] One or more reactions can be accelerated by applying energy or radiation, such as electromagnetic radiation. For example, the degradation or cleavage of terminal amino acids of peptides can be accelerated by applying microwave energy to accelerate the reaction kinetics. For example, the hydrolysis of proteins can be accelerated by applying microwave energy, as described in, for example, Margolis et al., 1991, Journal of Automatic Chemistry, Vol. 13, No. 3, pp. 93-95, which is incorporated herein by reference.
[0141] In some cases, more than one amino acid may be cleaved from a peptide per cleavage event. Cleavage may include cleaving two, three, four, five, six, seven, eight, nine, ten, or more amino acids. For example, a polymeric analyte may include a peptide containing multiple amino acids, and single amino acids, dipeptides, tripeptides, quadripeptides, or more may be cleaved using the methods described herein. In some cases, up to about 10 amino acids, up to about 9 amino acids, up to about 8 amino acids, up to about 7 amino acids, up to about 6 amino acids, up to about 5 amino acids, up to about 4 amino acids, up to about 3 amino acids, or fewer amino acids may be cleaved in a given cleavage event. In some cases, cleavage of more than one amino acid may be mediated using an enzyme (e.g., edmanase, protease) or ribozyme or DNAzyme capable of recognizing or cleaving more than one amino acid.
[0142] Cleavage of amino acids or modified amino acids can be carried out using a biological stimulus such as an enzyme, a ribozyme, or a DNAzyme. The enzyme can be any useful cleavage enzyme, for example, a protease such as edmanase or cruzain, a cleavage protein (e.g., ClpS, ClpX), proteinase K, an exopeptidase, an aminopeptidase, a diaminopeptidase, a serine protease, a cysteine protease, a threonine protease, an aspartic acid protease, an aspartic acid protease, a glutamic acid protease, a metalloprotease, an asparagine peptide lyase, a pepsin, or a trypsin. , pancreatin, Lys-C, Glu-C, Asp-N, chymotrypsin, carboxypeptidase (e.g., carboxypeptidase A, carboxypeptidase B, carboxypeptidase Y), SUMO protease, elastase, papain, endoprotease, proteinase, TrypZean®, bromelain, collagenase, hyaluronase, sarcoprotein ... Moricin, ficin, keratinase, tryptase, fibroblast activation, enterokinase, chymotrypsinogen, chymase, clostripain, calpain, alpha-degrading protease, proline-specific endopeptidase, furin, thrombin, subtilisin, genenase, PCSK9, cathepsin, prolidase, methionine aminopeptidase, cathepsin C, 1-cyclohexen-L-yl-boronic acid pinacol ester, pyroglutamic acid aminopeptidase, lecithin The cleavage enzyme may be tyrosine, kininogen, kallikrein, DPPIV / CD26, thimetoxin, prolyl oligopeptidase, leucine aminopeptidase, dipeptidyl peptidase, or other enzymes or proteases, or a combination or modification thereof (e.g., an engineered variant or variant). In some examples, the cleavage enzyme, ribozyme, or DNAzyme is configured or engineered to cleave a terminal amino acid or amino acids.Alternatively, the cleavage enzyme or ribozyme or DNAzyme is configured or engineered to cleave outside the internal amino acids at non-terminal positions of the peptide, for example, positions n-1, n-2, n-3, n-4, n-5, n-6, n-7, n-8, n-9, n-10, etc., where n is the number of amino acids in the peptide.
[0143] In the case of enzymatic cleavage, an additional reagent may be provided to catalyze or induce cleavage. For example, metalloproteases, aminopeptidases, or exopeptidases may promote cleavage of an amino acid or multiple amino acids in the presence of a catalyst, such as a metal or metal ion (e.g., cobalt). Thus, a catalyst may be provided to promote binding of the enzyme to an amino acid and subsequent cleavage of the amino acid from the peptide. In some cases, cleavage may be mediated by an apoenzyme, which is inactive in the absence of a cofactor metal catalyst, and cleavage may be controlled by the addition of a metal or metal ion.
[0144] Other examples of cleavage stimuli include the following: optical stimuli (e.g., application of UV, X-rays, gamma rays, or other wavelengths of light), mechanical stimuli (e.g., ultrasonication, high pressure, electromagnetic energy), thermal stimuli (e.g., application of heat), or chemical stimuli. In some examples, peptides contain or can be modified to contain cleavable or unstable bonds that can be cleaved upon application of an appropriate stimulus, such as disulfide bonds (e.g., cleavable upon application of a chemical stimulus such as a reducing agent), ester bonds (e.g., cleavable by a change in pH), vicinal diol bonds (e.g., cleavable by sodium periodate), Diels-Alder bonds (e.g., cleavable by application of heat), sulfone bonds (e.g., cleavable via base), silyl ether bonds (e.g., cleavable via acid), glycosidic bonds (e.g., cleavable via amylase), peptide bonds (e.g., cleavable via proteases), or phosphodiester bonds (e.g., cleavable via nucleases (e.g., DNase)).
[0145] Similarly, in some cases, the capture moiety can be cleaved from peptide or substrate.Cleavage can be carried out at any useful or convenient stage, for example, after the formation of AALC complex or stacked AALC complex, after multiple iterations of the workflow for producing multi-stacked AALC complex.In some cases, the cleavage of the capture moiety can be carried out after the formation of multi-stacked AALC complex, and the cleaved product can be sequenced, for example, using nanopore.
[0146] FIG. 1 illustrates a schematic example of a workflow for generating modified amino acids using an intramolecular extension process, followed by nanopore sequencing. In this exemplary workflow (100), a peptide (103) and a capture moiety (105) are provided, which are optionally bound to a substrate (101). The capture moiety (105) may include a first nucleic acid molecule (e.g., a DNA molecule). In process (106), a linker (107) and a polymerizable molecule (109), such as a second nucleic acid molecule, are provided. In some examples, the linker (107) is pre-tethered to the polymerizable molecule (109), or the linker (107) and the polymerizable molecule (109) may be provided separately. In process (106), the linker (107) may be bound to an amino acid (e.g., NTAA) of the peptide (103) to generate an amino acid-linker complex. In process (110), the amino acid-linker complex may be bound to the capture moiety (105). Attachment of the amino acid-linker conjugate to the capture moiety (105) may be mediated by a polymerizable molecule (109). Optionally, the amino acid-linker conjugate and the capture moiety may be covalently attached (e.g., the polymerizable molecule (109) may be covalently attached to the capture moiety (105)) using chemical (e.g., click chemistry) or enzymatic (e.g., ligase) techniques. Alternatively, or in addition, the polymerizable molecule (109) may comprise a first sequence that is complementary to and capable of hybridizing with a second sequence (not shown) of the capture moiety (105), or the polymerizable molecule (109) may be attached to the capture moiety (105) via a splint or bridge molecule that may comprise a sequence that is complementary to the first sequence of the polymerizable molecule (109) and the second sequence (not shown) of the capture moiety (105). In process (112), an amino acid may be removed or cleaved from peptide (103) to generate an amino acid-linker-capture (AALC) complex (113) comprising a modified amino acid (e.g., an NTAA that has been modified as a result of cleavage from the peptide), a linker (107), a polymerizable molecule (109), and a capture moiety (105).Processes 106, 110, and 112 may be repeated any number of times ("rounds") using additional linkers 107 and polymerizable molecules 109 and tethering additional polymerizable molecules together (e.g., tethering additional polymerizable molecules to the polymerizable molecules of the AALC complex 113). Multiple rounds may continue until all or a subset of the amino acids in peptide 103 have been removed from the peptide and tethered together. For example, a modified amino acid derivative, such as complex 115 containing multiple modified amino acids, may result from n rounds of workflow 100 (e.g., iterations of processes 106, 110, and 112).
[0147] As described elsewhere herein, a modified amino acid may include an amino acid having one or more modifications. For example, with reference to FIG. 1 , a modified amino acid may be used to refer to an amino acid-linker complex (e.g., resulting from process 106) or a portion thereof (e.g., only an amino acid of the amino acid-linker complex), an amino acid-linker-polymerizable molecule complex, an amino acid-linker-polymerizable molecule-capture moiety complex (e.g., resulting from process 110) or a portion thereof, a truncated amino acid of an AALC complex (113) or a portion thereof, or one of the amino acids of complex (115) of an AALC complex (stacked AALC complex). It will be understood that a modified amino acid may include other additional modifications not depicted in FIG. 1 , such as post-translational modifications or chemical modifications (e.g., protecting groups) described elsewhere herein. In some examples, a modified amino acid includes a derivatized amino acid. For example, a linker may include a PITC moiety as an amino acid reactive group, and conjugation of the PITC to an amino acid (e.g., NTAA) of a peptide under mildly basic conditions generates a phenylthiocarbamoyl (PTC) derivative of the amino acid. The PTC-derivatized amino acid may be treated with an acid (e.g., TFA or Lewis acid) to generate a cleaved cyclic 2-anilino-5(4)-thiazolinone (ATZ)-derivatized amino acid, leaving a new N-terminus on the remaining peptide. The ATZ-derivatized amino acid may be converted to a phenylthiohydantoin (PTH) derivative or a PTC derivative.
[0148] A "derivative" of a modified amino acid generally refers to a molecule derived from the modified amino acid. A derivative may be the product of a reaction or interaction (e.g., chemical, enzymatic) of the modified amino acid with another molecule. For example, with reference to FIG. 1 , the modified amino acid may be an amino acid in an amino acid-linker complex (e.g., resulting from process (106)), and a derivative of the modified amino acid may refer to all or part of the following: (i) an amino acid-linker complex bound to a capture moiety (e.g., resulting from process (110)), (ii) an AALC complex (113), or (iii) a complex (115) of an AALC complex (stacked AALC complex). Optionally, the modified amino acid may be further processed to obtain a "derivative" of the modified amino acid. For example, the modified amino acid may be subjected to further chemical or enzymatic treatment, cleavage of the capture moiety, extension or amplification reaction, removal from the substrate, or physical treatment such as mechanical shearing or fragmentation to obtain a derivative of the amino acid.
[0149] Substrate: One or more molecules described herein may be bound to a substrate. The one or more molecules bound to the substrate may be the same type of molecule or different types of molecules. For example, a capture moiety (e.g., a nucleic acid molecule) and a peptide may be bound to the substrate via a covalent or non-covalent interaction. The capture moiety and peptide may be bound to the substrate using any suitable chemical, such as a click chemistry moiety (e.g., an alkyne-azide bond), a photoreactive group (e.g., benzophenone), l-ethyl-3-(3-dimethylaminopropyl)carbodiimide hydrochloride (EDC) (e.g., for binding amino oligos or peptides), N-hydroxysulfosuccinimide (NHS), sulfo-NHS, or NHS-ester (e.g., for binding sulfhydryl oligos), maleimide, thiol, biotin-streptavidin interaction, cystamine, glutaraldehyde, formaldehyde, SMCC, sulfo-SMCC, silane (e.g., aminosilane), or a combination thereof. In some examples, the substrate may be functionalized to include coupling chemistries for attaching peptides or capture moieties. In one non-limiting example, the substrate (e.g., bead or surface) includes an alkyne such as dibenzocyclooctyne (DBCO), which can be configured to react with amines (e.g., DBCO-alcohol, DBCO-Boc), carboxyls or carbonyls (e.g., DBCO, DBCO-silane), sulfhydryls, etc. Azide-functionalized nucleic acids or proteins can be reacted with DBCO to attach the nucleic acid or protein to the DBCO substrate. In other examples, a linker, such as a bifunctional linker, can be used to attach the molecule to the substrate. Such bifunctional linkers can contain the same reactive moiety at both ends or different moieties at each end (e.g., heterobifunctional linkers).
[0150] Referring again to FIG. 1 , the complex (115) of the AALC complex (stacked AALC complex) may be further processed to prepare the complex (115) for further characterization, e.g., via sequencing. For example, if the capture moiety (105) is tethered to the substrate (101), the complex (115) may be removed from the substrate. For example, the capture moiety (105) may include a releasable or cleavable moiety, e.g., a restriction site recognized by a restriction enzyme, a chemically cleavable bond, e.g., a disulfide bond, a photolabile bond cleavable upon application of a light stimulus, etc., as described elsewhere herein. Alternatively, or in addition, the capture moiety (105) may include a primer binding site so that the stack of polymerizable molecules (e.g., nucleic acid molecules) can be copied or amplified, e.g., using a primer extension reaction, for downstream analysis or sequencing. Additional steps may be performed, such as purification or concentration, clean-up, nucleic acid reactions (e.g., ligation, extension, amplification, tagmentation, restriction enzyme digestion), fragmentation, barcoding, addition of adapters (e.g., sequencing adapters, read sequences, etc.), enzymatic treatment, etc.
[0151] Nanopore sequencing: Modified amino acids or their derivatives may be subjected to nanopore sequencing to determine the identity of the amino acid (e.g., as a proteinogenic amino acid, as one of the naturally occurring amino acids), and optionally the identity of the polymerizable molecule. Nanopore sequencing may be performed using commercially available nanopore systems, such as Oxford Nanopore Technologies, Genia Technologies, NobleGen, or Quantum Biosystem. Nanopore sequencing may be performed to determine the identity of different components of modified amino acids. For example, modified amino acids, including derivatized amino acids (e.g., thiocyanate-conjugated amino acids, or thiocarbamyl, thiazolinone, or thiohydantoin derivatives) attached to polymerizable molecules (e.g., nucleic acid molecules), can be subjected to nanopore sequencing, which can output which amino acid type the modified amino acid comprises or is derived from (e.g., whether the modified amino acid comprises one of the 20 proteinogenic amino acids or a post-translationally modified amino acid), and optionally the identity of the individual monomers of the polymerizable molecule (e.g., the nucleic acid sequence of the nucleic acid molecule). If the polymerizable molecule includes a nucleic acid molecule encoding additional information (e.g., including a barcode sequence, UMI, cycle information, spatial information, etc.), sequencing of both the derivatized amino acid and the nucleic acid molecule can generate multiplexed information.
[0152] Multiplexed sequencing: Advantageously, the use of nanopore sequencing to analyze peptides using the methods provided herein can result in a streamlined process for generating multiplexed data. For example, the methods provided herein can enable the identification of amino acid residues and polymerizable molecules (e.g., identification of individual monomers, nucleic acid sequences of nucleic acid molecules). As described herein, modified amino acids can include polymerizable molecules (e.g., nucleic acid molecules) and amino acids or derivatives thereof. Polymerizable molecules can contain multiple types of information. For example, polymerizable molecules, such as nucleic acid molecules, can contain sequences that encode not only spatial information but also cyclic or other temporal information, as described above. In one such example, an array of peptides and capture moieties can be provided on a substrate. The array can include multiple individually addressable units, each (or a subset) of the individually addressable units of the array containing a peptide to be sequenced and a capture moiety. Linkers capable of binding to amino acids of the peptides can be provided throughout the array. Multiple polymerizable molecules can be provided before, during, or after providing the linkers. The plurality of polymerizable molecules may contain spatial information (e.g., a spatial barcode sequence), which uniquely identifies the individually addressable unit and thus its location in the array. Alternatively, or in addition, the polymerizable molecule or capture moiety may contain information about the peptide, such as a barcode that identifies the peptide, a partition, the sample from which the peptide originated, the population of cells, etc. The polymerizable molecule or capture moiety may further contain temporal information (e.g., a cycle barcode indicating the round or iteration in which the polymerizable molecule or capture moiety is provided). Subsequent sequencing may be used to reveal spatial or temporal information (e.g., the position of origin in the array of peptides or amino acids, the period or timing in which the polymerizable molecule or capture moiety is provided, etc.). Alternatively, or in addition, the capture moiety may contain multiplexing information (e.g., a barcode sequence, spatial information, temporal information, UMI, etc., to identify the peptide or the sample or partition from which the peptide originated).
[0153] Similarly, binding agents may be used to add multiple analytical approaches to peptide sequences. The use of binding agents bound to modified amino acids is useful for regulating the transport of modified amino acids through a nanopore. The binding agent may contain additional information that can be read out using nanopore sequencing. For example, different binding agents that recognize specific amino acid residues (or modified amino acids, or amino acid-linker complexes) may be used. Different binding agents, alone or in combination with their bound modified amino acids, may each associate with a unique current signature when transported through a nanopore and thus may be useful for identifying specific residues. In some examples, the binding agent may further contain additional coding information, for example, via an attached barcode molecule or detectable tag. For example, the binding agent may have a nucleic acid molecule attached, and the barcode may identify the binding agent or the binding partner (e.g., amino acid residue) of the binding agent. Alternatively, or in addition, the binding agent may contain spatial or temporal information, as described above. Thus, when a modified amino acid associated with a binding agent is transported through a nanopore, multiple sets of information can be obtained, for example, spatial, temporal, encoded barcode information can be obtained from the binding agent, the amino acid or modified amino acid, the polymerizable molecule, the capture moiety, or a combination thereof.
[0154] Advantageously, the methods provided herein can provide multiplexed information in a streamlined workflow that does not require multiple analytical instruments. In this way, obtaining multiplexed data can be performed substantially simultaneously. For example, using nanopore sequencing, the identity of an amino acid (or modified amino acid) and the identity of the polymerizable molecule attached thereto can be determined in up to about 5 minutes, about 1 minute, about 50 seconds, about 40 seconds, about 30 seconds, about 20 seconds, about 10 seconds, about 1 second, about 900 milliseconds, about 800 milliseconds, about 700 milliseconds, about 600 milliseconds, about 500 milliseconds, about 400 milliseconds, about 300 milliseconds, about 200 milliseconds, about 100 milliseconds, about It may be obtained in 900 microseconds, about 800 microseconds, about 700 microseconds, about 600 microseconds, about 500 microseconds, about 400 microseconds, about 300 microseconds, about 200 microseconds, about 100 microseconds, about 900 nanoseconds, about 800 nanoseconds, about 700 nanoseconds, about 600 nanoseconds, about 500 nanoseconds, about 400 nanoseconds, about 300 nanoseconds, about 200 nanoseconds, about 100 nanoseconds, or less.
[0155] Arrays: The methods and systems provided herein may include the use of arrays, for example, for massively parallel sample processing or sequencing. In one example, a substrate may include an array of individually addressable units comprising multiple peptide analytes. In some examples, a subset of the individually addressable units comprises a single peptide analyte. The individually addressable units may further comprise single or multiple capture moieties. In some examples, the subset of individually addressable units comprising a single peptide each comprises at least one capture moiety. Arrays may be patterned (e.g., the individually addressable units are arranged in a pattern) or random (e.g., the individually addressable units are randomly distributed across the substrate).
[0156] An array may contain any useful number of individually addressable units, for example, an array may have at least 5, at least 10, at least 100, at least 1,000, at least 10,000, at least 100,000, at least 1,000,000, at least 10,000,000, at least 100,000,000, at least 1,000,000,000, or more individually addressable units. Alternatively, the array may have up to about 1,000,000,000, up to about 100,000,000, up to about 10,000,000, up to about 1,000,000, up to about 100,000, up to about 10,000, up to about 1,000, up to about 100, up to about 10, or up to about 5 individually addressable units. Arrays may have various numbers of individually addressable units, for example, from about 1,000 to about 100,000, or from about 50,000 to about 1,000,000.
[0157] Binding Agents: In some embodiments, the methods, compositions, systems, and kits provided herein may include the use of a binding agent. The binding agent may be or include a protein or peptide (e.g., an antibody, antibody fragment, nanobody), a nucleic acid molecule (e.g., an aptamer), a polymer, an inorganic compound, a small molecule, or a derivative (e.g., an engineered variant), or a combination thereof. The binding agent can bind to a modified amino acid or a portion thereof. The use of a binding agent may be useful for identifying individual amino acids (or modified amino acids) of a peptide or for modulating the transport of modified amino acids through a nanopore or nanogap. The binding agent may include a recognition site that specifically recognizes an amino acid, a modified amino acid, or a derivatized (and optionally modified) amino acid. For example, the binding agent may be configured to recognize a portion of a modified amino acid, such as a particular amino acid residue, a residue-linker complex, or a derivatized amino acid (e.g., a thiocarbomyl-derivatized residue, a thiazolone-derivatized residue, a thiohydantoin-derivatized residue, etc.), or a portion of a modified amino acid. In some cases, binding agents may be derived from or engineered from naturally occurring enzymes, such as aminopeptidases or tRNA synthetases.
[0158] In some cases, a binding agent may be provided as part of the intramolecular extension process. For example, a binding agent may specifically recognize and bind a modified amino acid (e.g., an AALC complex, an amino acid bound to a polymerizable molecule) or a portion thereof. In some cases, a binding agent may recognize and bind a portion of a modified amino acid, such as an amino acid-linker conjugate or a derivative thereof (e.g., an isothiocyanate-amino acid, or a thiocarbamyl, thiazolone, or thiohydantoin derivative thereof). The binding agent may contain a detectable moiety (e.g., a fluorophore, a barcode molecule, a radioisotope, or other tag), which may facilitate detection and identification of individual amino acids. Alternatively, or in addition, the binding agent may be recognized by an additional binding agent (e.g., a secondary antibody) containing a detectable moiety. Binding of the additional binding agent to the binding agent may indicate the presence of the modified amino acid.
[0159] FIG. 2 schematically illustrates a workflow for another example of intramolecular peptide extension, including the use of a binding agent. In workflow (200), a peptide (203) and a capture moiety (205) are provided, which are optionally bound to a substrate (201). The capture moiety (205) may include a first nucleic acid molecule (e.g., DNA). In process (206), a linker (207) and a polymerizable molecule (209), such as a second nucleic acid molecule, are provided. In some examples, the linker (207) is pre-tethered to the polymerizable molecule (209). Alternatively, the linker (207) and the polymerizable molecule (209) may be provided separately. In process (206), the linker (207) may be bound to an amino acid (e.g., NTAA) of the peptide (203) to generate an amino acid-linker conjugate. In process (210), the amino acid-linker conjugate may be bound to the capture moiety (205). Attachment of the amino acid-linker conjugate to the capture moiety (205) can be mediated by a polymerizable molecule (209). Optionally, the amino acid-linker conjugate and the capture moiety can be covalently attached (e.g., covalently attaching the polymerizable molecule (209) to the capture moiety (205)) using chemical (e.g., click chemistry) or enzymatic (e.g., ligase) techniques. Alternatively, or in addition, the polymerizable molecule (209) can include a first sequence that is complementary to and capable of hybridizing with a second sequence (not shown) of the capture moiety (205), or the polymerizable molecule (209) can be attached to the capture moiety (205) via a splint or bridge molecule that can include a sequence that is complementary to the first sequence of the polymerizable molecule (209) and the second sequence (not shown) of the capture moiety (205). In process (212), an amino acid may be removed from peptide (203) to produce an amino acid-linker-capture (AALC) complex (213) comprising a modified amino acid (e.g., an NTAA modified as a result of cleavage from the peptide), a linker (207), a polymerizable molecule (209), and a capture moiety (205). In process (214), a binding agent (217) (e.g., an antibody, binding protein, etc.) may be contacted with the AALC complex (213). The binding agent may be configured to recognize all or a portion of the AALC complex (213).For example, the binding agent may recognize a residue of a modified amino acid, or the binding agent may recognize all or part of the residue and linker (207). In one example, the linker may include a PITC moiety, and the binding agent may recognize a PITC-amino acid, or a derivative thereof, such as phenylthiocarbamyl, thiazolone, or phenylthiohydantoin. In some examples, the binding agent (217) may include a detectable moiety (not shown) or may be contacted with an additional binding agent (e.g., a secondary antibody) that may optionally include a detectable moiety (not shown).
[0160] Processes 206, 210, 212, and 214 may be repeated any number of times ("rounds"), using additional linkers 207 and polymerizable molecules 209 (optionally including cycle / round information) to tether additional polymerizable molecules together (e.g., tethering additional polymerizable molecules to the polymerizable molecules of AALC complex 113). Multiple rounds may continue until all or a subset of the amino acids in peptide 203 have been removed from the peptide and tethered together. For example, a modified amino acid derivative, such as complex 215 containing multiple modified amino acids, may result from n rounds of workflow 200 (e.g., repeating processes 206, 210, 212, and optionally process 214).
[0161] Intramolecular extension can be performed in solution or using a substrate-free approach. Figure 3 illustrates another example workflow for generating modified amino acids using an intramolecular extension process in solution, followed by nanopore sequencing. In this example workflow (300), a peptide (303) and a capture moiety (305) are provided. The capture moiety (305) may include a first nucleic acid molecule (e.g., a DNA molecule) and may include identification information for the peptide (303), such as a peptide identification barcode sequence. The capture moiety (305) may further include a releasable or cleavable moiety. In process (306), a linker (307) and a polymerizable molecule (309), e.g., a second nucleic acid molecule, are provided. In some examples, the linker (307) is pre-tethered to the polymerizable molecule (309). Alternatively, the linker (307) and the polymerizable molecule (309) may be provided separately. The polymerizable molecule (309) may include temporal information, for example, identifying the cycle or round in which the polymerizable molecule is provided. In process (306), a linker (307) may be attached to an amino acid (e.g., NTAA) of the peptide (303) to generate an amino acid-linker conjugate. In process (310), the amino acid-linker conjugate may be attached to a capture moiety (305). Attachment of the amino acid-linker conjugate to the capture moiety (305) may be mediated by the polymerizable molecule (309) and, optionally, an additional polymerizable molecule (311). Optionally, the amino acid-linker conjugate and the capture moiety may be covalently attached together (e.g., via ligation). Alternatively, or in addition, the polymerizable molecule (309) may comprise a first sequence that is complementary to and capable of hybridizing with a second sequence (not shown) of the capture moiety (305), or the polymerizable molecule (309) may be attached to the capture moiety (305) via a splint or bridge molecule that may comprise a sequence that is complementary to the first sequence of the polymerizable molecule (309) and the second sequence (not shown) of the capture moiety (305).In process 312, an amino acid may be removed or cleaved from peptide 312 to produce an amino acid-linker-capture (AALC) complex 313 comprising a modified amino acid (e.g., an NTAA modified as a result of cleavage from the peptide), a linker 307, a polymerizable molecule 309, and a capture moiety 305. Processes 306, 310, and 312 may be repeated or iterated any number of times ("rounds") using additional linkers 307 and polymerizable molecules 309 to tether additional polymerizable molecules together (e.g., tethering additional polymerizable molecules to the polymerizable molecules of AALC complex 313). Multiple rounds may continue until all or a subset of the amino acids in peptide 303 have been removed from the peptide and tethered together. For example, modified amino acid derivatives, such as stacked AALC complexes (315) containing multiple modified amino acids, can result from n rounds of workflow (300) (e.g., repeating processes (106), (310), and (312)). After any useful number of rounds, the completed stacked AALC complexes can be cleaved from or at the capture portion (305), e.g., using a cleavable moiety. The cleaved products can then be sequenced using a nanopore or nanogap.
[0162] In some cases, the polymerizable molecule 309 includes temporal information about the cycle in which the molecule was generated. In such cases, the temporal information can be used to perform quality control. For example, a missing cycle number can indicate that an amino acid was missing or not present in the peptide, that an amino acid cleavage was not performed, or other abnormalities.
[0163] The order of operations and steps provided herein is listed for illustrative purposes only, and it will be understood that any of the operations and steps can be performed at any convenient or useful step or time. For example, the polymerizable molecule can be provided pre-attached to a linker (e.g., as shown in process (106) of FIG. 1, process (206) of FIG. 2, and process (306) of FIG. 3), or the polymerizable molecule can be provided before or after providing the linker. Similarly, attachment of the polymerizable molecule to the capture moiety can occur before, during, or after attachment of the linker to the amino acid. In another example, removal (e.g., cleavage) of an amino acid from the peptide can occur before, during, or after attachment of the polymerizable molecule to the capture moiety. In yet another example, if a binding agent and intramolecular extension are performed, the binding agent can be introduced before, during, or after the intramolecular process. For example, workflows 100, 200, and 300 may be performed to contact the stacked AALC complex (115, 215, and 315) with multiple binders (e.g., 10, 20, or more different binders, each with specificity for a different amino acid or modified amino acid) that have specificity for a single or oligomer of different modified amino acids (e.g., dipeptides or tripeptides). The multiple binders may include detectable labels that can be used to identify the amino acids (e.g., using single-molecule imaging). Alternatively, or in addition, the AALC complex and binders may be introduced into a nanopore or nanogap for sequencing.
[0164] Additional manipulations of the methods described herein are also contemplated. For example, with reference to Figures 1-3, the AALC complex (113), (213), or (313), or a portion of the AALC complex, may be cleaved or removed from the substrate or peptide (e.g., via cleavage, enzymatic digestion, or chemical dissociation) before repeating workflows (100), (200), or (300). In such examples, the cleaved molecule (e.g., including the modified amino acid and optionally the polymerizable molecule) may be transported through a nanopore for protein sequencing. In other examples, an additional cleavage or dissociation reaction may be performed to separate the amino acid from the polymerizable molecule. In such examples, the separated polymerizable molecule or amino acid may be transported through a nanopore. Additional useful manipulations, methods, systems, and compositions for processing or using peptides, linkers, polymerizable molecules, and the like may also be found in International Application No. PCT / US2023 / 017954, which is incorporated herein by reference in its entirety.
[0165] Any of the operations described herein can be carried out under useful physical conditions.For example, one or more of the operations described herein can be carried out at high or low temperatures, which can aid reaction rate, reaction completion, the transport speed or residence time of molecules through nanopore or nanogap, etc.In some cases, nanopore or nanogap sequencing system can be carried out on ice or at low temperatures, which can beneficially increase the measured signal-to-noise ratio of current signal, adjust the transport speed of molecules through nanopore or nanogap, or otherwise improve reading or detection.
[0166] Detection: Additional methods and devices for sequencing peptides and / or polymerizable molecules can be used instead of or in addition to nanopore sequencing or nanogap sequencing.For example, the methods provided herein can utilize one or more imaging systems to identify molecules (e.g., modified amino acids or polymerizable molecules).Examples of optical imaging systems and methods include, but are not limited to, fluorescence microscopes, confocal microscopes, total internal reflection fluorescence microscopes (TIRF), magnification microscopes, two-photon microscopes, integrated correlation microscopes, stimulated emission depletion (STED) microscopes, stochastic optical reconstruction microscopes (STORM), reversible saturable optical linear fluorescence transition (RESOLFT), spatially modulated illumination, spectroscopic precision distance microscopes (SPDM), photoactivated localization microscopes (PALM), fluorescence PALM (FPALM), structured illumination microscopes (SIM), saturation SIM (SSIM), etc.
[0167] In some examples, imaging may be used to identify modified amino acids or polymerizable molecules. For example, with reference to FIG. 2, a binding agent (e.g., an antibody) may carry a fluorophore or other optically detectable tag, and the binding agent may selectively bind to a specific residue of the modified amino acid (e.g., in process (217)). The fluorophore may identify the binding agent, and the presence of the fluorophore indicates the presence of the binding agent (and thus the amino acid residue to which the binding agent binds). Alternatively, a secondary binding agent (e.g., a secondary antibody) carrying a fluorophore or other detectable agent may bind to the binding agent (e.g., the primary antibody) to indicate the presence of the binding agent (and thus the amino acid residue to which the primary antibody binds). Similarly, with respect to polymerizable molecules (e.g., nucleic acid molecules), the sequence of the nucleic acid molecule may be determined using any suitable imaging technique, such as fluorescence in situ hybridization or sequencing-by-synthesis.
[0168] Alternatively, or in addition, other proteomics tools may be used to identify modified amino acids and polymerizable molecules. For example, modified amino acids (e.g., amino acids in contact with PITC linkers attached to DNA molecules) may be analyzed using mass spectrometry, immunochemistry (e.g., via ELISA, fluorescence imaging), or other analytical techniques. In some cases, the modified amino acids may be further processed before analysis. For example, such further processing may include stripping, enrichment, or purification of the modified amino acids from the substrate, or the addition of tags or other moieties. In some cases, the modified amino acids may be subjected to a separation technique, such as chromatography (e.g., liquid chromatography, HPLC, nanoLC), electrophoretic separation (e.g., electrophoresis, isoelectric focusing, electroosmotic flow), filtration, evaporation, distillation, or other separation technique, before analysis.
[0169] In another aspect of the present disclosure, disclosed herein is a method of analyzing an analyte, comprising: (a) providing a polymerizable molecule (e.g., a nucleic acid molecule) bound to the analyte; (b) sequencing the polymerizable molecule (e.g., the nucleic acid molecule); and (c) identifying the analyte, wherein (b) and (c) are performed substantially simultaneously.
[0170] The analyte may include a non-nucleic acid biomolecule (e.g., a protein or peptide, a lipid, a carbohydrate, or a combination thereof) or a component thereof. In some examples, the analyte is a peptide comprising at least two amino acids. In some examples, the analyte includes modified amino acids (e.g., modified amino acids generated from peptides through intramolecular extension and the use of linkers) as described elsewhere herein. In some examples, the analyte includes modified amino acids, a linker, and a polymerizable molecule (e.g., a nucleic acid molecule).
[0171] As described herein, analytical methods may include the use of nanopores or nanogaps for sequencing or identification. Nanopore or nanogap sequencing techniques can be useful for generating multiplexed data, such as (i) sequencing of polymerizable molecules (e.g., sequencing of nucleic acid molecules to generate sequencing reads) and (ii) identification of analytes (e.g., identification of amino acid residues). Such multiplexed data generation can occur substantially simultaneously, e.g., within about 5 minutes, about 1 minute, about 1 second, about 1 millisecond, about 1 microsecond, about 1 picosecond, or less. Alternatively, or in addition, sequencing and identification may not occur substantially simultaneously; rather, signal generation from the polymerizable molecules and amino acids may occur substantially simultaneously. For example, a modified amino acid or a derivative thereof (e.g., an AALC complex, a stacked AALC complex, or a portion thereof) may be transported through a nanopore. Transport of the modified amino acid or a derivative thereof through the nanopore can generate and collect one or more signals (e.g., current signals) generated from both the polymerizable molecule and the amino acid. Such one or more signals may then be deconvoluted using computational methods to determine the identity of the polymerizable molecule (e.g., nucleic acid sequence) and the amino acid type of the modified amino acid or derivative thereof. Advantageously, the use of nanopore or nanogap sequencing to generate multiplexed data may eliminate the need for multiple analytical techniques or instruments.
[0172] Systems, Kits, Compositions: Also provided herein are systems, kits, and compositions for characterizing analytes (e.g., peptides). The systems, kits, and compositions provided herein may be useful in practicing any of the described methods or may be provided in complement to the described methods.
[0173] In one embodiment, the kits of the present invention may include a substrate, a capture moiety, a linker, a polymerizable molecule, a peptide conjugation reagent, an enzyme (e.g., a polymerizing or ligating enzyme, a cleaving enzyme, a restriction enzyme, a nicking enzyme, an exonuclease, a repair enzyme such as uracil-DNA glycosylase), a detection or labeling agent, or any combination thereof. The kit may also include buffers, reagents, binders, catalysts, or other chemicals or biomolecules (e.g., enzymes) necessary to perform chemical or enzymatic reactions. The kits of the present disclosure may further include instructions for using the components of the kit or for performing any of the methods and processes described herein. For example, the kit may include instructions for conjugating a peptide onto a substrate using a peptide conjugation reaction. The kit may include instructions for attaching a capture moiety to a substrate, or the substrate may have a capture moiety attached to it. Similarly, the kit may include instructions for performing intramolecular extensions as described herein to generate modified amino acids, including, for example, polymerizable molecules (e.g., DNA) attached to the modified amino acids, amino acid-linker conjugates, AALC conjugates, etc.
[0174] In another aspect, disclosed herein are compositions that can be used to characterize analytes (e.g., peptides). The compositions can include a linker covalently attached to a polymerizable molecule (e.g., a nucleic acid molecule), which can be useful, for example, for intramolecular extension of peptides for peptide sequencing. In one example, the composition includes (A) a linker comprising: (i) a first portion capable of binding to an amino acid (e.g., CTAA or NTAA of a peptide); (ii) a nucleic acid molecule (e.g., DNA); and, optionally, (iii) a releasable or cleavable portion, which can be the same or different from (ii); and, optionally, (iv) a spacer portion; and (B) a nucleic acid molecule. In another example, the composition can include a linker covalently attached to a nucleic acid molecule. Such a linker can bind to an amino acid (e.g., CTAA or NTAA of a peptide) and optionally include a cleavable portion; the covalently attached nucleic acid molecule can be configured to tether to another nucleic acid molecule (e.g., a capture moiety), which, in some examples, can be provided attached to a substrate.
[0175] Figure 4 schematically illustrates one example of a composition described herein. Panel A of Figure 4 shows a linker (403) that can be attached to a polymerizable molecule (401) (e.g., a DNA molecule). The polymerizable molecule (401) contains an azide moiety (N = N + =N -The linker (403) may comprise a first reactive group, such as a thiocyanate (denoted as N=C=S) or PITC, which may be attached to any useful position on the polymerizable molecule (401), e.g., the 5'-end, 3'-end, or an internal position in the DNA molecule in the case of a DNA molecule. The linker (403) may comprise (i) an amino acid reactive moiety, such as a thiocyanate (denoted as N=C=S) or PITC, and (ii) a second reactive group (e.g., an alkyne) capable of reacting with the first reactive group (e.g., an azide) on the polymerizable molecule (401). Panel B of Figure 4 shows an example of a polymerizable molecule covalently bonded to a linker. The first reactive group (e.g., an azide) on the polymerizable molecule can react with the second reactive group (e.g., an alkyne) on the linker, thereby generating a covalently bonded structure comprising the polymerizable molecule and the linker, which can react with and tether to an amino acid. As described herein, the amino acid reactive group (e.g., ITC, PITC) can also cleave an amino acid from a peptide. In some cases, the first reactive group is attached to the polymerizable molecule via an additional linker. For example, the additional linker may include a nucleic acid linker that includes a modified nucleic acid base or sugar backbone, such as a nucleoside, nucleoside monophosphate, nucleoside diphosphate, or nucleoside triphosphate, which can be incorporated into a nucleic acid molecule. The additional linker may include a first reactive group, such as a click chemistry moiety. In some cases, the additional linker includes ethynyl deoxyuridine or octadiynyl deoxyuridine.
[0176] Systems, Compositions, Kits: Also provided herein are systems, compositions, and kits for processing peptides or producing or processing modified amino acids. The systems of the present disclosure may include a sequencing instrument configured to receive the modified amino acids or derivatives thereof described herein and provide sequencing reads of the modified amino acids or derivatives thereof. Alternatively, or in addition, the systems of the present disclosure may be configured to process, prepare, or sequence the modified amino acids. The systems may be configured to provide a peptide, a capture moiety, and a linker comprising a polymerizable molecule, attach the linker to an amino acid of the peptide, cleave the amino acid from the peptide, and optionally repeat one or more operations or steps. Accordingly, the systems of the present disclosure may include any useful device or tool, including, but not limited to, a mixer, a liquid handler, a vortex, a centrifuge, a heating or cooling element, a mechanical stage, a microfluidic chamber or device, and fluid control.
[0177] The disclosed systems may include a nanopore or nanochannel sequencing system configured to sequence or analyze modified amino acids. The modified amino acids may include or be linked to polymerizable molecules, as described elsewhere herein. In this manner, the systems may be configured to output multiplexed data regarding the modified amino acids (e.g., whether the modified amino acids are or are derived from one of the 20 proteinogenic amino acid residues) and the polymerizable molecules (e.g., nucleic acid sequences). The disclosed systems may include substrates, capture moieties, linkers, polymerizable molecules, peptide conjugation reagents, enzymes (e.g., polymerizing or ligating enzymes, cleaving enzymes, restriction enzymes, nicking enzymes, exonucleases, repair enzymes such as uracil-DNA glycosylase), detection or labeling agents, buffers, reagents, binders, catalysts, or other chemicals or biomolecules (e.g., enzymes) required to perform chemical or enzymatic reactions, or any combination thereof. The systems may further include one or more detection (e.g., imaging or mass spectrometry) systems, separation systems (e.g., HPLC), or other analytical equipment.
[0178] The compositions of the present disclosure may include any useful article or reagent for processing or sequencing peptides or for generating modified amino acids or derivatives thereof. The composition may include a capture moiety configured to bind to a polymerizable molecule. The composition may include a linker (e.g., for binding to an amino acid), which may include a polymerizable molecule, a cleavage agent, a ligation agent (e.g., a ligase), a capture moiety, a reagent for performing a reaction, or a combination thereof. In some examples, the composition includes a linker comprising (e.g., covalently attached to) a nucleic acid barcode molecule, wherein the linker includes an amino acid reactive group. The nucleic acid barcode molecule may include any useful barcode sequence, such as a spatial sequence, a temporal sequence, a peptide identification sequence, a partition identification sequence, a sample identification sequence, a UMI, or other functional sequence.
[0179] The kits of the present disclosure may include any useful reagents for processing, analyzing, or sequencing peptides. The kits may include reagents for providing peptides and capture moieties, reagents for binding peptides and capture moieties to substrates, reagents for cleaving amino acids, reagents for binding peptides to linkers containing either an amino acid reactive group and an additional reactive group or a polymerizable molecule, reagents for binding capture moieties to polymerizable molecules, reagents for providing binders, cleavage reagents, stabilization reagents, wash buffers, or combinations thereof. The kits may further include reagents for removing or separating amino acids from peptides, or reagents for removing or separating AALC complexes from substrates or capture moieties. In some examples, the kits may further include reagents for removing or separating AALC complexes or stacked AALC complexes from substrates, or the kits may include reagents for performing nucleic acid extension reactions. Kits may include any relevant reagents, such as buffers, detergents, chelators, cofactors, enzymes, ribozymes or DNAzymes, acids, bases, salts, metal ions, primers, nucleic acid molecules, nucleotides, proteins, polynucleotides, binding agents (e.g., antibodies, aptamers, nanobodies, antibody fragments), lipids, carbohydrates, ribozymes, riboswitches, probes, fluorophores, oxidizing agents, reducing agents, nuclease or protease inhibitors, dyes, organic molecules, inorganic molecules, emulsifiers, surfactants, stabilizers, polymers, water, small molecules, therapeutic agents, radioactive materials, preservatives, or other useful reagents. Kits of the present disclosure may also provide instructions for use of the contents of the kit.
[0180] Substrate conjugation The present disclosure provides methods for attaching molecules (e.g., macromolecular analytes, e.g., biomolecules such as nucleic acid molecules, peptides, lipids, carbohydrates, etc.) to substrates. The substrates may be functionalized to allow covalent or non-covalent attachment of molecules to the substrate. The substrates may include any useful functional moiety, e.g., reactive moieties, that can bind or conjugate with a molecule or another reactive moiety. In non-limiting examples, the reactive moiety may include click chemistry moieties, e.g., azides, alkynes, nitrones, alkenes (e.g., strained alkenes), tetrazines, methyltetrazines, triazoles, tetrazoles, phosphites, phosphines, etc. The click chemistry moiety may be reactive in a copper-catalyzed Huisgen cycloaddition reaction, or a 1,3-dipolar cycloaddition reaction between an azide and a terminal alkyne, a Diels-Alder reaction (e.g., a cycloaddition reaction between a diene and a dienophile), or a nucleophilic substitution reaction in which one of the reactive species is an epoxy or aziridine. The molecule that binds to the substrate may include a click chemistry moiety that is complementary to the substrate. For example, the substrate may contain an alkyne moiety and the molecule to be attached may contain an azide moiety that can react with the alkyne moiety of the substrate to form a covalent bond. In one such example, the substrate may contain a dibenzocyclooctyne (DBCO) moiety to which an azide-containing molecule (e.g., azide-DNA, azide-polymer, azide-peptide) can react and be attached.
[0181] Alternatively, or in addition, the reactive moiety may comprise a photoreactive moiety that can be activated when exposed to a light stimulus (e.g., light such as UV light or visible light). Examples of photoreactive moieties include aryl(phenyl)azides (e.g., phenylazide, ortho-hydroxyphenylazide, meta-hydroxyphenylazide, tetrafluorophenylazide, ortho-nitrophenylazide, meta-nitrophenylazide), diazirine, azido-methyl-coumarin, benzophenone, anthraquinone, diazo compounds, diazirine, psoralen, 3-cyanovinylcarbazole phosphoramidite (CNVK), and analogs or derivatives thereof.
[0182] The reactive moiety can include a carboxyl-reactive crosslinker group, such as a diazo compound such as diazomethane and diazoacetyl, carbonyldiimidazole, a carbodiimide (e.g., l-ethyl-3-(3-dimethylaminopropyl)carbodiimide hydrochloride (EDC)), dicyclohexylcarbodiimide (DCC)), or an amine-reactive group (e.g., N-hydroxysulfosuccinimide (NHS), sulfo-NHS, or NHS-ester). The reactive group can also include a crosslinker, which can include an NHS group, an EDC group, a maleimide (e.g., for attachment to a Michael acceptor), a thiol, a cystamine, an aldehyde, a succinimidyl group, an epoxide, or an acrylate. Examples of crosslinking agents include, for example, NHS (N-hydroxysuccinimide), sulfo-NHS (N-hydroxysulfosuccinimide), EDC (1-ethyl-3-[3-dimethylaminopropyl], carbodiimide hydrochloride, SMCC (succinimidyl 4-(N-maleimidomethyl)cyclohexane-1-carboxylate), sulfo-SMCC, DSS (disuccinimidyl suberate), DSG (disuccinimidyl glutarate), DFDNB (1,5-difluoro-2,4-dinitrobenzene), BS3 (bis(sulfosuccinimidyl)suberate), TSAT (tris-(succinimidyl)aminotriacetate), ), BS(PEG)5 (PEGylated bis(sulfosuccinimidyl)suberate), BS(PEG)9 (PEGylated bis(sulfosuccinimidyl)suberate), DSP (dithiobis(succinimidyl propionate), DTSSP (3,3'-dithiobis(succinimidyl propionate)), DST (disuccinimidyl tartrate), BSOCOES (bis(2-(succinimidooxycarbonyloxy)ethyl)sulfone), EGS (ethylene glycol bis(succinimidyl succinate)), DMA (dimethyl adipimidate), DMP (dimethyl pimelimidate), DMS (dimethyl suberimidate), DTBP (Wang and Richard's Reagent), BM(PEG)2 (1,8-bismaleimide-diethylene glycol), BM(PEG)3 (1,11-bismaleimide-triethylene glycol), BMB (1,4-bismaleimide butane),DTME (dithiobismaleimidoethane), BMH (bismaleimidohexane), BMOE (bismaleimidoethane), TMEA (tris(2-maleimidoethyl)amine), SPDP (succinimidyl 3-(2-pyridyldithio)propionate), SMCC (succinimidyl trans-4-(maleimidylmethyl)cyclohexane-1-carboxylate), SIA (succinimidyl iodoacetate), SBAP (succinimidyl 3-(bromoacetamido)propionate), STAB (succinimidyl(4-iodo)propionate) cetyl)aminobenzoate), sulfo-SIAB (sulfosuccinimidyl (4-iodoacetyl)aminobenzoate), AMAS (N-α-maleimidoaceto-oxysuccinimide ester), BMPS (N-β-maleimidopropyl-oxysuccinimide ester), GMBS (N-γ-maleimidobutyryl-oxysuccinimide ester), sulfo-GMBS (N-γ-maleimidobutyryl-oxysulfosuccinimide ester), MBS (m-maleimidobenzoyl-N-hydroxysuccinimide ester), sulfo Ho-MBS (m-maleimidobenzoyl-N-hydroxysulfosuccinimide (ester), SMCC (succinimidyl 4-(N-maleimidomethyl)cyclohexane-1-carboxylate), sulfo-SMCC (sulfosuccinimidyl 4-(N-maleimidomethyl)cyclohexane-1-carboxylate), EMCS (N-ε-maleimidocaproyl-oxysuccinimide ester), sulfo-EMCS (N-ε-maleimidocaproyl-oxysulfosuccinimide) (ester), SMPB (succinimidyl 4-(p -maleimidophenyl)butyrate), sulfo-SMPB (sulfosuccinimidyl 4-(N-maleimidophenyl)butyrate), SMPH (succinimidyl 6-((β-maleimidopropionamidohexanoate)), LC-SMCC (succinimidyl 4-(N-maleimidomethyl)cyclohexane-1-carboxy-(6-amidocaproate)), sulfo-KMUS (N-κmaleimidoundecanoyl-oxysulfosuccinimide ester), SPDP (succinimidyl 3-(2-pyridyldithio)propionate),LC-SPDP (succinimidyl 6-(3(2-pyridyldithio)propionamido)hexanoate), LC-SPDP (succinimidyl 6-(3(2-pyridyldithio)propionamido)hexanoate), sulfo-LC-SPDP (sulfosuccinimidyl 6-(3'-(2-pyridyldithio)propionamido)hexanoate), SMPT (4-succinimidyloxycarbonyl-α-methyl-α(2-pyridyldithio)toluene), PEG4-SPDP (PEGylated, long-chain SPDP crosslinker), P EG12-SPDP (PEGylated, long-chain SPDP crosslinker), SM(PEG)2 (PEGylated SMCC crosslinker), SM(PEG)4 (PEGylated SMCC crosslinker), SM(PEG)6 (PEGylated, long-chain SMCC crosslinker), SM(PEG)8 (PEGylated, long-chain SMCC crosslinker), SM(PEG)12 (PEGylated, long-chain SMCC crosslinker), SM(PEG)24 (PEGylated, long-chain SMCC crosslinker), BMPH (N-β-maleimidopropionic acid hydrazide), EMCH (N-ε -maleimidocaproic acid hydrazide), MPBH (4-(4-N-maleimidophenyl)butyric acid hydrazide), KMUH (N-κ-maleimidoundecanoic acid hydrazide), PDPH (3-(2-pyridyldithio)propionyl hydrazide), ATFB-SE (4-azido-2,3,5,6-tetrafluorobenzoic acid, succinimidyl ester), ANB-NOS (N-5-azido-2-nitrobenzoyloxysuccinimide), SDA (NHS-diazirine) (succinimidyl 4,4'-adipentanoate), LC- SDA (NHS-LC-diazirine) (succinimidyl 6-(4,4'-azipentanamido)hexanoate), SDAD (NHS-SS-diazirine) (succinimidyl 2-(4,4'-azipentanamido)ethyl)-1,3-dithiopropionate), sulfo-SDA (sulfo-NHS-diazirine) (sulfosuccinimidyl 4,4'-azipentanoate), sulfo-LC-SDA (sulfo-NHS-LC-diazirine) (sulfosuccinimidyl 6-(4,4'-azipentanamido)hexanoate),Sulfo-SDAD (sulfo-NHS-SS-diazirine) (sulfosuccinimidyl 2-(4,4'-azipentanamido)ethyl)-1,3-dithiopropionate), SPB (succinimidyl-[4-(psoralen-8-yloxy)]-butyrate), sulfo-SANPAH (sulfosuccinimidyl 6-(4'-azido-2'-nitrophenylamino)hexanoate), DCC (dichlorohexylcarbodiimide), EDC (1-ethyl-3-(3-dimethylaminopropyl)carbodiimide hydrochloride), glutaraldehyde, formaldehyde, and combinations or derivatives thereof.
[0183] Molecules can also be attached to substrates using linkers. Linkers can have any useful number of functional or reactive groups and can be monofunctional (having one functional group), bifunctional, trifunctional, tetrafunctional, or contain more functional groups. In some examples, molecules (e.g., nucleic acid molecules, peptides, or polymers) can be attached to substrates using heterobifunctional linkers. Heterobifunctional linkers can contain any useful functional group, as described herein. Non-limiting examples of heterobifunctional linkers include p-azidobenzoylhydrazide hydrazide (ABH), N-5-azido-2-nitrobenzoyloxysuccinimide (ANB-NOS), N-[4-(p-azidosalicylamido)butyl]-3'-(2'-pyridyldithio)propionamide (APDP), p-azidophenylglyoxal monohydrate (APG), bis[B-(4-azidosalicylamido)butyl]-3'-(2'-pyridyldithio)propionamide (APDP), p-azidophenylglyoxal monohydrate (APG), bis[B-(4-azidosalicylamido)butyl]-3'-(2'-pyridyldithio)propionamide (APDP), and bis[B-(4-azidosalicylamido)butyl]-3'-(2'-pyridyldithio)propionamide (APDP). Bis[2-(succinimidyl)amino]ethyl disulfide (BASED), bis[2-(succinimidooxycarbonyloxy)ethyl]sulfone (BSOCOES), BMPS, 1,4-di[3'-(2'-pyridyldithio)propionamido]butane (DPDPB), dithiobis(succinimidyl propionate) (DSP), disuccinimidyl suberate (DSS), disuccinimidyl tartrate (DST), 3,3'-Dithiobis(sulfosuccinimidyl propionate (DTSSP), EDC, ethylene glycol bis(succinimidyl succinate) (EGS), N-(E-maleimidocaproic acid hydrazide (EMCH), N-(E-maleimidocaproyloxy)-succinimide ester (EMCS), N-maleimidobutyryloxysuccinimide ester (GMBS), hydroxylamine-HCl, MAL-PEG-SCM, m-maleimidobenzoyl-N-hydrazide Methylsuccinimide ester (MBS), N-hydroxysuccinimidyl-4-azidosalicylate (NHS-ASA), PDPH, N-succinimidyl bromoacetate (SBA), SIA, sulfo-SIA, succinimidyl-4-(N-maleimidomethyl)cyclohexane-1-carboxylate (SMCC), succinimidyl 4-(p-maleimidophenyl)butyrate (SMPB), succinimidyl-6-[s-maleimidopropionamido]hexanoate Sulfo-MBS, Sulfo-SANPAH, Sulfo-SMCC, Sulfo-DST, Sulfo-EMCS, Sulfo-GMBS, N-hydroxysulfosuccinimidyl-4-azidobenzoate (Sulfo-MBS), N-succinimidyl 3-[2-pyridyldithio]-propionate (SPDP), Sulfo-LC-SPDP, N-(p-maleimidophenyl isocyanate) (PMPI), N-succinimidyl (4-iodoacetyl)aminobenzoate (SIAB), Sulfo-MBS, Sulfo-SANPAH, Sulfo-SMCC, Sulfo-DST, Sulfo-EMCS, Sulfo-GMBS, N-hydroxysulfosuccinimidyl-4-azidobenzoate (Sulfo-MBS), Examples include sulfosuccinimidyl (4-azidophenyl)-1,3-dithiopropionate (sulfo-SADP), sulfosuccinimidyl 2-(m-azido-o-nitrobenzamido)-ethyl-1,3'-dithiopropionate (sulfo-SAND), sulfosuccinimidyl-2-(p-azidosalicylamido)ethyl-1,3-dithiopropionate (sulfo-SASD), sulfo-SIAB, sulfo-SMCC, sulfo-SMPB, etc.
[0184] Additional examples of conjugation reactions that can be used to attach molecules to substrates include the Ullmann reaction, Heck reaction, Negishi reaction, Stille reaction, Suzuki reaction, Buchwald-Hartwig coupling, Stephens coupling, Glaser coupling, Kumada coupling, Larocco indole synthesis, Miyaura borylation, Sonogashira cross-coupling, and Grubbs reaction.
[0185] More than one type of molecule may be attached to the substrate. For example, a substrate may be attached to a nucleic acid molecule and a peptide. Alternatively, a substrate may be attached to only one type of molecule (e.g., only nucleic acid molecules, only peptides, only lipids, only carbohydrates, etc.). The substrate may be attached to any useful combination of molecules, linkers, reactive moieties, or functional groups, as described elsewhere herein, which may be attached at any useful density. For example, a multifunctional linker may be used to attach both nucleic acid barcode molecules and peptides to the substrate. Alternatively, the substrate may include multiple bifunctional linkers that can attach to different molecules. In another example, the substrate may include a linker and a reactive site. The linker is used to attach one type of molecule (e.g., a peptide or nucleic acid molecule), while the reactive site is used to attach another type of molecule (e.g., a nucleic acid molecule or a peptide).
[0186] The linker may include other functional moieties, such as a spacer (e.g., a polymer chain, e.g., PEG, an alkyl chain, etc.), a cleavage site (e.g., a disulfide bridge cleavable upon application of a chemical stimulus, a photocleavable or thermally cleavable moiety, etc.), an enzyme recognition site, etc.
[0187] The proximity of a substrate-bound molecule to its nearest neighbor (e.g., another molecule) can be controlled using various techniques, such as self-assembled monolayers, patterning techniques, binding moieties, etc. In some cases, it may be advantageous for two molecules to be in close proximity (e.g., two polymerizable molecules, such as a peptide and a nucleic acid molecule, or two nucleic acid molecules). For example, with respect to the sequencing techniques described herein, a capture moiety may be used to bind a monomer of a polymerizable analyte; following monomer cleavage, an additional polymerizable molecule or multiple polymerizable molecules may need to be in close proximity to the capture moiety to enable transfer of information encoded by the polymerizable molecule of the binder. The proximity of the molecules (e.g., capture moiety and polymerizable molecule) may be mediated using tethering molecules, such as nucleic acid molecule "staples" or multifunctional linkers.
[0188] The nucleic acid molecule may be attached to the substrate by direct binding. In such cases, the substrate or the nucleic acid molecule may contain functional moieties that can interact with each other. For example, the substrate and nucleic acid molecule may contain complementary click chemistry pairs, such as an alkyne and an azide. In one such example, the substrate may contain an alkyne moiety (e.g., DBCO), which can be reacted with an azide-functionalized nucleic acid molecule. The nucleic acid molecule may be covalently attached to the substrate by reacting with the alkyne moiety in a click chemistry reaction. In another example, the substrate may contain an avidin or streptavidin moiety, to which a biotinylated nucleic acid molecule can interact and non-covalently bind. Alternatively, or in addition, the substrate may contain a nucleic acid molecule to which an additional nucleic acid molecule (e.g., a nucleic acid analyte, a nucleic acid linker) has been attached using hybridization, ligation, click chemistry, or crosslinking (e.g., photocrosslinking such as CNVK).
[0189] Alternatively, or in addition, the nucleic acid molecule may be attached to the substrate using a linker, for example, as described elsewhere herein. The linker may contain at least two functional groups (e.g., a heterobifunctional linker) that can bind to both the substrate and the nucleic acid molecule. In one example, the substrate contains an amine group, and an alkyne-functionalized DNA primer (e.g., DBCO-DNA primer) may be attached using a linker such as azidoacetic acid NHS ester. In another example, an amine-functionalized substrate may be attached to an azide-functionalized DNA primer using a DBCO-NHS ester or DBCO-PEG-NHS ester linker. As described elsewhere herein, the linker may contain additional functional moieties (e.g., a cleavage site, a spacer such as a polymer chain, or an alkyl chain).
[0190] Similarly, peptides may be attached to substrates by direct bonding or by using a linker. Peptides may be attached to substrates at the peptide termini (e.g., C-terminus or N-terminus), at internal residues or amino acids of the peptide, or at multiple positions along the peptide. In an example of direct bonding, the peptide may be functionalized with a moiety that can interact with a moiety of the substrate (e.g., a click chemistry pair, avidin-biotin). For example, the substrate and peptide may comprise complementary click chemistry pairs, such as binding partners of an alkyne and an azide, or avidin and biotin. In one example of a click chemistry pair, the substrate may comprise an alkyne moiety (e.g., DBCO), which can react with an azide-functionalized peptide. In a click chemistry reaction, the peptide may be reacted with the alkyne moiety to covalently bond the substrate to the peptide. In another example, the substrate may comprise an avidin or streptavidin moiety, to which the biotinylated peptide interacts and non-covalently binds.
[0191] Alternatively, or in addition, the peptide may be attached to the substrate using a linker, e.g., as described elsewhere herein. The linker may contain at least two functional groups (e.g., a heterobifunctional linker) that can bind to both the substrate and the nucleic acid molecule. In one example, the substrate contains an amine group, and a linker such as azidoacetic acid NHS ester may be used to attach an alkyne-functionalized peptide. In another example, a DBCO-NHS ester or DBCO-PEG-NHS ester linker may be used to attach an amine-functionalized substrate to an azide-functionalized peptide. In yet another example, a substrate containing an amine group may be attached to an azide-functionalized peptide using EDC and Sulfo-NHS.
[0192] Peptides can be functionalized with functional moieties to allow them to be bonded to substrates. Functional moieties include silanes, such as aminosilanes (e.g., APTES), amino-PEG-silanes, click chemistry moieties, or other binding moieties, and can be attached to peptides at the peptide termini (N-terminus or C-terminus), internal amino acids, or multiple positions (e.g., multiple internal amino acids, one or both termini, etc.). Chemical methods for functionalizing peptides can include C-terminus-specific conjugation (e.g., via C-terminus decarboxyalkylation) using photoredox catalysis, or amide bonding to amine-functionalized surfaces, as described, for example, by Bloom et al., Nature Chemistry 10, 205-211.2018 and Zhang et al., ACS Chem.Biol. 2021, 16, 11, 2595-2603, each of which is incorporated herein by reference in its entirety. N-terminal attachment can involve amide coupling of the N-terminal amine group to a carboxyl-functionalized surface or the use of 2-pyridinecarboxaldehyde variants. Alternatively, or in addition, terminal functionalization of peptides can be achieved enzymatically or using enzyme analogs such as ribozymes or DNAzymes. In one example of enzymatic functionalization and conjugation, carboxypeptidases or amidases can be used for C-terminal functionalization (e.g., as described in Xu et al., ACS Chem Biol. 2011 Oct 21;6(10):1015-1020; Zhu et al., Chinese Chemical Letters. 2018, Vol. 29 Issue 7, Pages 1116-1118; and Zhu et al., ACS Catal. 2022, 12, 13, 8019-8026), each of which is incorporated herein by reference in its entirety, thereby enabling the addition of click chemistry moieties to peptides.The click chemistry-functionalized peptide may then be directly conjugated to a substrate via another clickable group (e.g., a BCN-azide or DBCO-azide bond), or in other instances, may be reacted with another linker or polymerizable molecule (e.g., a bait nucleic acid molecule bearing a clickable group) that can be directly or indirectly conjugated to a substrate (e.g., using a capture nucleic acid molecule to hybridize the bait nucleic acid molecule). Additional examples of enzymes that can be used for functionalization or conjugation include sortase A, subtiligase, butelase I, or trypsiligase. In some instances, ubiquitin ligase can be used to attach a ubiquitin protein bearing a linker moiety to a substrate. These linker moieties can then be used to chemically conjugate proteins to the ubiquitin-conjugated substrate. In some examples, glycosylation enzymes can be used to attach functionalized sugar groups (e.g., click chemistry-functionalized sugars, polymer-linked sugars, biotinylated sugars) to amino acid residues, enabling attachment to substrates (e.g., via click chemistry, polymer cross-linking or nucleic acid hybridization, avidin-biotin interactions, etc.). Internal amino acid residues or post-translationally modified residues can also be attached to substrates using, for example, thiol labeling, amide linkages to glutamic or aspartic acid residues using EDC / NHS chemistry or DMT-MM, esterification of glutamic or aspartic acid residues, alkylation or disulfide bridge labeling of cysteines, or amide linkages to lysine residues.
[0193] The peptide may be treated before, during, or after binding to the substrate. In some examples, the peptide is conjugated with a tag that allows the peptide to bind to the substrate, such as a His tag, a SNAP tag, a CLIP tag, a SpyCatcher, a SpyTag, or a nucleic acid tag (e.g., a bait oligo that can bind to a capture oligo of the substrate).
[0194] In some cases, it may be advantageous to block or protect primary amine or carboxyl groups, and optionally deblock or deprotect the N-terminal primary amine or C-terminal carboxyl group, to facilitate attachment of the N-terminus or C-terminus to a substrate. In one example, selective attachment of a peptide to a single point (e.g., the C-terminus) can be achieved by reacting the peptide with a linker containing an amine-reactive group (e.g., an isothiocyanate such as PITC) and a reactive group (e.g., a click chemistry group). The linker can be, for example, a PITC-conjugated click chemistry moiety such as PITC-azide or PITC-alkyne, optionally with a spacer moiety therebetween, e.g., PITC-alkyl-azide, PITC-PEG-azide, PITC-alkyl-alkyne, or PITC-PEG-azide. The linker reacts with primary amines, including the N-terminus, to "block" them (e.g., to modify lysines). Cleavage of the N-terminal amino acid can then be performed (e.g., using an Edman reagent such as an acid), and one of the remaining modified lysines can be attached to a substrate (e.g., using a click chemistry moiety linked to an amine-reactive group). Optionally, the peptide can be treated with a protease, such as LysC, which cleaves the peptide so that the remaining peptide has a C-terminal lysine and contains only a primary amine at the C-terminal lysine residue and the N-terminus; such cleavage can be performed before reacting the amine-reactive group, as shown, for example, by Xie et al., Langmuir 2022, 38, 30, 9119-9128, which is incorporated herein by reference in its entirety.
[0195] Similarly, carboxyl groups can be reacted in a manner that allows for C-terminal or internal residue attachment. In an example of C-terminal conjugation, if the carboxyl group is treated with an activating reagent (e.g., acetic anhydride), it can be labeled with a C-terminal sequencing reagent, such as an isothiocyanate, to generate a peptide-thiohydantoin (at the C-terminus) and a "blocked" carboxyl group on the aspartic acid and glutamic acid residues. The thiohydantoin can then be reacted and attached to a substrate. Alternatively, cleavage of the C-terminal amino acid via a single round of C-terminal sequencing degradation or via a protease exposes only a single reactive carboxyl group at the C-terminal amino acid. The single reactive C-terminal carboxyl group can then be used as a reactive moiety for a single attachment site.
[0196] In another approach, peptides or proteins can be attached via their N-terminus, taking advantage of the specific reactivity of N-terminal amine groups. Amine-based reactions, such as amide bonds, can be performed at low pH, where only N-terminal amine groups are active. Additionally, 2-pyridinecarboxaldehyde and its variants can be used to react with N-terminal amine groups.
[0197] In some cases, peptides may be attached to substrates using, for example, polymerization reactions, e.g., PEGylated peptides, methacrylamide-modified peptides, Michael-type addition of maleimide-terminated oligo-NIPAAM-conjugated peptides, photocrosslinking of azophenyl-conjugated peptides, and other polymerization reactions of monomer-conjugated peptides, as described, for example, in Krishna et al., Biopolymers. 2010;94(1):32-48, which is incorporated herein by reference in its entirety.
[0198] Multiple types of molecules may be attached to a substrate. The substrate may be attached with any combination of molecules, including, but not limited to, peptides, proteins (e.g., enzymes, ribozymes, DNAzymes, antibodies, nanobodies, antibody fragments), nucleic acid molecules, lipids, carbohydrates or sugars, metabolites, small molecules, polymers, metals, viral particles, biotin, avidin, streptavidin, neutravidin, etc. Multiple types of molecules may be attached to the substrate simultaneously or sequentially. For example, the substrate may be treated to conjugate nucleic acid molecules and then treated to conjugate peptides, or the substrate may be treated to conjugate peptides before nucleic acid molecules. Any number of conjugation or attachment chemistries may be used. For example, when multiple types of molecules are attached to a substrate, any number of conjugation chemistries may be used for each type of molecule.
[0199] The substrate, or a portion thereof, may be subjected to conditions sufficient to passivate the substrate or a portion thereof. Passivation of the substrate may be useful for various purposes, such as preventing nonspecific binding of binders, altering the surface density of molecules (e.g., increasing the density of nucleic acid molecules or peptides), or blocking reactive sites (e.g., blocking available click chemistry moieties following conjugation of molecules on the substrate). Passivation may be achieved using chemical methods, such as deposition of a blocking agent, such as a protein (e.g., albumin), Tween-20, polymer, metal, or metal oxide, or biochemical methods, such as metallomicrons. Substrates containing reactive moieties may also be passivated after molecular conjugation (e.g., conjugation of nucleic acid molecules, peptides, etc.) by reacting unreacted sites with an appropriate molecule. For example, substrates containing click chemistry moieties, such as DBCO beads, may be conjugated to molecules of interest (e.g., polymerizable molecules, e.g., nucleic acid molecules, peptides) at useful densities using click chemistry (e.g., azido-nucleic acid molecules, azido-peptides). Unreacted sites may be passivated by supplying and reacting complementary click chemistry molecules, such as azide-polymers (e.g., PEG-azide), which may reduce downstream nonspecific interactions.
[0200] Passivation of the substrate may be carried out at any useful time or stage. For example, passivation to block unreacted DBCO sites may be carried out before, during, or after conjugation of an analyte or other molecule of interest (e.g., peptides and nucleic acid molecules). Passivation can be controlled by the stoichiometry or density of the passivating agent relative to the molecule of interest, or by physical techniques such as photopatterning, self-assembled monolayers, etc.
[0201] Sample processing The present disclosure also provides a method for processing a sample.One or more methods for processing a sample can include preparing biological sample for analysis, and in some cases, include dividing cells for performing single-cell analysis.The method for processing a biological sample can include extracting or isolating one or more peptides or proteins from biological sample for further processing and analysis, as described elsewhere herein.
[0202] Preparation of cell suspensions for single-cell analysis: The methods described herein may include preparation of a single-cell suspension from a biological sample. A single-cell suspension may be prepared from a biological sample by dissociating cells and, optionally, culturing them in a liquid medium. In some examples, the biological sample includes a liquid sample. For example, the biological sample may include a bacterial liquid culture, a mammalian liquid culture, a blood, plasma, or serum sample. Processing of such liquid samples may include centrifugation (e.g., to isolate cells), resuspension of the cells in an appropriate medium, such as Dulbecco's phosphate-buffered saline (DPBS), and optional culturing of the isolated cells.
[0203] Biological samples may include cultured cells, e.g., cells cultured in suspension, or cells attached to a solid surface, such as a petri dish or tissue culture dish. Cultured adherent cell samples may be treated to detach cells from the surface, e.g., with a protease such as trypsin, to generate a cell suspension. Biological samples may also include tissue samples or biopsy samples. Cell suspensions may be prepared by mechanically or enzymatically treating the tissue or biopsy sample. Such treatments include sonication (mechanical treatment) or enzymatic treatments, such as the use of pronase, collagenase, hyaluronidase, metalloproteases, trypsin, or other enzymes that digest extracellular matrix components. Dissociated cells can then be stored in an appropriate buffer, such as DPBS.
[0204] Cell sorting: Biological samples or cell suspensions may be subjected to sorting to isolate cells of interest. Sorting can be performed to select or separate cells based on cellular qualities or characteristics, such as target protein expression, size, deformability, fluorescence or other optical properties, or other physical properties of the cells. Sorting can be achieved using any number of techniques, including immunosorting (e.g., fluorescence-activated cell sorting (FACS) or magnetic-activated cell sorting (MACS)), electrophoretic techniques, chromatography, microfluidic techniques (e.g., using inertial focusing, cell trapping, electrophoresis), acoustic sorting, optical sorting (e.g., optoelectronic tweezers), mechanical cell picking (e.g., using manual or robotic pipettes), or passive techniques (e.g., gravity sedimentation).
[0205] Partitioning: Cells of a biological sample or cell suspension can be divided into individual partitions such that at least a subset of the individual partitions contain single cells. Each individual partition may contain a barcode molecule (e.g., a fluorophore or set of fluorophores, a nucleic acid barcode molecule, etc.). The barcode molecule may be unique to the partition, such that each partition contains a different barcode sequence than the other partitions. Barcode molecules may be loaded into individual partitions at any useful ratio of barcode molecules to sample type (e.g., cells, proteins, nucleic acid molecules). Barcode molecules may be loaded into partitions such that approximately 0.0001, 0.001, 0.1, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10,000, or 200,000 barcodes are loaded per sample type. In some cases, the barcodes are loaded into partitions such that there are more than about 0.0001, 0.001, 0.1, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10,000, or 200,000 barcodes loaded per sample type. In some cases, the barcodes are loaded into partitions such that there are less than about 0.0001, 0.001, 0.1, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10,000, or 200,000 barcodes loaded per sample type.
[0206] A partition can assume any useful shape, such as a droplet, a microwell, a solid substrate, a gel (e.g., cells encapsulated in a gel bead), a bead, a flask, a tube, a spot, a capsule, a channel, a chamber, or other compartment or container. A partition may also be part of an array of partitions, such as, for example, a droplet in a microfluidic device, a microwell in a microwell plate, or a spot in a multi-spot array.
[0207] Lysis, Permeabilization, and Analyte Extraction: Single cells (e.g., individual compartments) may be processed to obtain one or more analytes contained therein. Methods for processing single cells may include lysing the cells to release their contents into individual compartments or partitions. Lysis may be performed using detergents (e.g., Triton-X 100, sodium dodecyl sulfate, sodium deoxycholate, CHAPS), RIPA buffer, temperature changes (e.g., increasing or decreasing temperature, freezing, freeze-thawing), enzymes, ribozymes, DNAzymes, mechanical lysis (e.g., sonication, application of mechanical force), electrical lysis, or a combination thereof. Lysis may be performed in the presence of protease inhibitors to prevent degradation or digestion of proteins from the cells. The contents may optionally be further processed, such as by purification or extraction, protein or peptide denaturation, enzymatic or chemical digestion, etc. In some cases, the contents may be subjected to enzymatic digestion to remove nucleic acid molecules, for example, using nucleases such as DNAse or RNAse. Alternatively, or in addition, cells may be fixed and / or permeabilized (e.g., using a fixative). Examples of fixatives include aldehydes (e.g., glutaraldehyde, formaldehyde, paraformaldehyde), alcohols (e.g., methanol, ethanol), acetone, acids (e.g., acetic acid, Davidson's AFA), oxidizing agents (e.g., osmium tetroxide, potassium dichromate, chromate, permanganate), Zenker's fixative, picrate, Hepes-glutamate buffer-mediated organic solvent protective effect (HOPE), or Karnovsky's fixative. Cell permeabilization can be achieved mechanically (e.g., using sonication, electroporation, shearing) or chemically (e.g., using organic solvents such as methanol or acetone or detergents such as saponin, Tween-20, Triton X-100).
[0208] Protein Treatment: Further processing of a biological sample (or single-cell suspension or partitioned cells) may enable proteomic analysis. For example, disaggregation of proteins in a sample may be performed using, for example, chemical or mechanical techniques. Chemical disaggregation methods include, but are not limited to, sodium dodecyl ester (SDS), Triton-X 100, 3-((3-cholamidopropyldimethylaminio)-1-propanesulfonic acid (CHAPS), ethylene carbonate, or formamide. Mechanical disaggregation methods may include, but are not limited to, sonication or high temperature treatment. The biological sample (or single-cell suspension or partitioned cells) may be subjected to conditions sufficient to denature one or more proteins. Denaturation can be achieved using heat, chemicals (e.g., SDS, urea, guanidine, formamide, organometallic compounds), reducing agents (e.g., dithiothreitol (DTT), β-mercaptoethanol, TCEP), or other suitable methods. This can be achieved using urea, chaotropes, enzymes (e.g., ClpX, ClpS, unfoldases), ribozymes, or DNAzymes. Similarly, peptides or proteins may be subjected to conditions to solubilize them in solution, for example, with detergents, organic solvents, spermidine, or by tagging the peptide or protein with a polyionic tag (e.g., DNA, PEG, or other polymer). Alternatively, or in addition, peptides or proteins may be concentrated or purified. In one example, the peptide or protein of interest is precipitated using trichloroacetic acid, chloroform, TRIzol, or other chemical reagents. Other biological or chemical agents, such as lysozyme, papain, cruzain, trypsin, protease inhibitors, nucleases or nuclease-containing proteins (e.g., DNAse, RNAse, DNA glycosylase, restriction endonucleases, transposase, micrococcal nuclease, Cas proteins), etc., may be included during protein processing. In some instances, minimal protein processing is performed, for example, to maintain the native state or conformation of the protein.
[0209] Peptides or proteins may be fragmented prior to analysis. Protein fragmentation, as described elsewhere herein, can be useful for reducing the size of proteins and enabling efficient processing of peptides. Fragmentation may be carried out using proteases, such as trypsin, chymotrypsin, pepsin, Lys-C, Glu-C, proteinase K, furin, thrombin, endopeptidase, papain, subtilisin, elastase, enterokinase, genenanase, endoproteases, metalloproteases, or chemical treatments, such as cyanogen bromide, hydrazine, hydroxylamine, formic acid, BNPS-skatole, iodobenzoic acid, 2-nitro-5-thiocyanobenzoic acid, etc. Alternatively, or in addition, fragmentation may be carried out using mechanical methods such as sonication, vortexing, mechanical agitation, application of temperature changes (e.g., freeze / thaw, heat), or other fragmentation techniques.
[0210] Enrichment of proteins or peptides in a biological sample may be performed, for example, to separate proteins and peptides from cellular debris or other types of analytes (e.g., nucleic acids, lipids, carbohydrates, metabolites). Such enrichment may include, for example, using affinity columns (e.g., ion exchange), size exclusion columns, affinity precipitation (e.g., immunoprecipitation), chemical precipitation (e.g., using trichloroacetic acid, chloroform, TRIzol), chromatography (e.g., HPLC), or electrophoresis. If cells are partitioned prior to enrichment, enrichment may be performed using microbeads, affinity microcolumns, affinity beads, etc. In some examples, fractionation is performed on proteins or peptides, which can be used to separate proteins by size, hydrophobicity, charge, affinity, size, mass, density, etc.
[0211] Peptides may be barcoded in bulk or in partitions. Peptides may be barcoded with any useful type of barcode molecule, such as a spectral or fluorescent barcode, a mass tag, a nucleic acid barcode molecule, or the like. The barcode molecule may enable identification of the peptide, partition, sample, cell, or cellular compartment of origin. For example, a cell sample may be partitioned such that a partition contains at most one cell, and the partition may contain a unique barcode molecule (e.g., a nucleic acid barcode molecule) that identifies that partition, i.e., cell. Subsequent labeling of peptides within a partition with a barcode molecule (e.g., by permeabilizing or lysing the cell) may be useful for identifying whether the peptides originate from or are derived from the same cell or partition. In other examples, a substrate may contain a nucleic acid molecule containing a unique barcode sequence that differs from the barcode sequences of other substrates. In this manner, the barcode sequence may be used to identify the substrate. In some examples, a barcoded substrate may be partitioned with a cell sample such that at least some of the partitions contain a single cell and a single barcoded substrate. In this way, peptides originating from a single cell and transferred to the barcode substrate may all be identifiable as originating from a single cell. Barcode molecules may also contain additional useful functional sequences, such as UMIs, primer sites, restriction sites, cleavage sites, transposition sites, sequencing sites, read sites, etc.
[0212] Attachment of a barcode molecule to a peptide can be achieved using any suitable chemical reaction. For example, C-terminal conjugation of a nucleic acid barcode molecule can be achieved by amide linkage of an amine-linked DNA barcode molecule to the peptide, or by thiol alkylation, e.g., by reacting a thiolated peptide with an alkylated (e.g., iodoacetamide) DNA barcode molecule. N-terminal conjugation can be achieved, for example, by reacting a 2-pyridinecarboxyaldehyde label of the DNA barcode with the N-terminus of the peptide. Internal residues, such as glutamic acid, can also be labeled with amine-linked DNA barcode molecules or carboxylated DNA barcodes (e.g., by reacting with the primary amine of lysine).
[0213] Individual peptides may be barcoded at multiple sites for a given peptide. For example, peptides may be partitioned into partitions containing multiple identical barcode molecules containing barcode sequences unique to the partition. Peptides may be labeled with unique partition barcode sequences at single or multiple sites, optionally each containing a unique molecular identifier (UMI), so that subsequent downstream analysis (e.g., sequencing) is attributable to the same peptide using the barcode sequence. In some cases, the terminus (e.g., N-terminus or C-terminus) or an internal amino acid of a peptide may be labeled with a barcode. In some cases, peptides may be fragmented before analysis or sequencing, and thus, attaching identical barcode molecules to the same peptide upstream may result in sequence analysis attributing a single peptide. Peptide barcoding may occur before, during, or after fragmentation. Peptides may be labeled with barcodes (e.g., nucleic acid barcode molecules) using any suitable chemical, such as those described above, or bifunctional or trifunctional linkers containing multiple linking moieties, such as click chemistry moieties, NHS-esters, EDC, etc., as described elsewhere herein. For example, attachment to the C-terminus may include an amide bond to the C-terminal carboxyl group or photooxidative tagging of the C-terminal carboxyl group (e.g., to add an electrophilic tag). Attachment to the N-terminus may include an amide bond to an N-terminal amine group, specific attachment may be achieved at low pH, or variants of 2-pyridinecarboxaldehyde may be used for specific attachment to the N-terminus. Internal attachment may include, for example, an amide bond to glutamic or aspartic acids using EDC / NHS chemistry or DMT-MM, alkylation or disulfide bridge labeling of cysteines, or an amide bond to a lysine residue.
[0214] In some examples, peptides may be labeled with different barcode molecules, and these barcode molecules may be indexed by their proximity to each other, for example, using primers that can anneal to adjacent barcode molecules. In one such approach, a protein may be labeled with multiple barcodes with different barcode sequences, and then proximity-based polymerase extension may be used to copy and associate the sequences of adjacent barcodes. For example, each barcode molecule may contain a primer binding site to which a dual primer linker sequence containing two sequences is annealed. The dual primer linker sequence may bind to the primer binding sites of two adjacent barcodes. For example, an extension reaction using a polymerase may extend and copy the barcode sequences of adjacent barcodes. Subsequently, the dual primer linker sequence (now bearing copies of two adjacent barcodes) may be extracted and sequenced. From the sequencing reads, a barcode sequence adjacency matrix may be generated (e.g., barcode sequences on a single dual primer linker may be matched as spatially adjacent). Thus, each barcode sequence may be associated with its nearest neighbor, and the peptide portions may therefore be aligned or reduced as adjacent. Such an approach is useful when the peptide is fragmented, and the barcode sequences can be used to match individual fragments of the peptide with their nearest neighbors.
[0215] In another example, bridge amplification can be used to barcode peptides at multiple positions for a given peptide. In such an approach, peptides or proteins may be labeled at multiple sites using nucleic acid primers. Nucleic acid barcode molecules may be provided that can anneal to or be ligated to nucleic acid primers (not shown). Subsequent rounds of bridge amplification can be performed to copy the nucleic acid barcode molecules to other primers located at other sites on the given peptide. In some examples, peptides may be tagged with multiple copies of nucleic acid primers, and the barcode sequences may be provided sparsely so that only one nucleic acid primer per peptide is extended by polymerase extension. Subsequent rounds of bridge amplification can result in peptides with the same barcode sequence in each nucleic acid primer. Subsequent peptide fragmentation can be performed so that peptide fragments contain, on average, one single barcode. Thus, in some cases, the output of such an amplification approach can be peptides with individual barcodes generated from fragmentation of a multiply labeled protein, in which peptides from the same protein have the same barcode.
[0216] A sample of cells may be partitioned into individual partitions or compartments (e.g., droplets, microwells) such that at least some of the partitions contain single cells. The partitions may then be treated with a lysing agent to lyse the cells and release proteins from the cells into the partitions. Proteins may then be labeled with partition-specific barcodes (e.g., using barcode beads) so that all peptides or proteins originating from a single compartment contain the same barcode. In some cases, the barcode may comprise a nucleic acid barcode molecule, and the barcode sequence can be used in downstream processing, e.g., sequencing, to identify the partition or cell from which the peptide originated. The nucleic acid barcode molecule may also comprise any additional useful sequences, e.g., a UMI, a primer sequence, etc.
[0217] Bulk processing: Biological samples may be processed in bulk. For example, a biological sample may be processed to obtain a suspension of cells, which may be directly lysed without partitioning the cells into individual compartments. Cells may be lysed in bulk using any useful technique, such as those described above, and may optionally be subjected to further processing, such as homogenization, protease inhibition, denaturation, protein processing (e.g., chemical treatment, fragmentation), or a combination thereof. Biological samples may be subjected to pretreatment before cell lysis or protein extraction. Such pretreatment may include debris removal, purification, filtration, concentration, fractionation, etc.
[0218] Spatial barcoding: A biological sample may include a tissue sample containing multiple cells. The tissue sample may be processed using a technique (e.g., using spatial barcodes) to preserve spatial information (e.g., identify peptides from individual cells). For example, a 2D or 3D tissue sample may be provided, and individual cells or locations within the tissue sample may be contacted with multiple spatial barcodes (e.g., nucleic acid barcode molecules) containing different barcode sequences. The different barcode sequences may be attributed to specific locations within the 2D or 3D tissue sample, which may correspond to the location of cells. For example, spatial barcodes can be provided using deterministic methods such as two-photon patterning or probabilistic methods such as PCR to assign unique spatial barcodes to different segments of a 2D or 3D tissue sample. Thus, a peptide labeled with a spatial barcode may be assigned to a single location or a single cell within the tissue sample.
[0219] In another example of a workflow for spatial barcoding a tissue sample, a tissue sample containing multiple cells (illustrated as a 2x2 array of cells) may be provided. The tissue sample may be subjected to lysis or fixation and permeabilization to provide access to proteins contained within multiple cells. Spatial barcodes, such as nucleic acid barcode molecules, may be provided. The spatial barcodes may include coordinate or location information. In one example, each cell may be contacted with a different spatial barcode, or portions of cells may be contacted with different spatial barcodes, which may optionally be pre-indexed (e.g., using imaging or deterministic spatial barcoding techniques). Further processing of the peptides may be performed as described elsewhere herein. When peptides are labeled with spatial barcodes, each peptide bearing a spatial barcode may be assigned its coordinate or location of origin, which may help identify the cell of origin from which the peptide originates.
[0220] In one example, the spatial barcode array may be provided on a substrate (e.g., a microscope slide, a hydrogel mesh). In some cases, the spatial barcode may be directly conjugated to the substrate or provided on a barcoded bead. In one example, a plurality of beads, each containing a different barcode sequence, may be arranged in an array on a substrate. Each bead may include a spatial barcode containing the spatial barcode sequence and, optionally, a unique molecular identifier (UMI). A tissue sample (e.g., a fixed tissue sample) may be placed adjacent to (e.g., overlaid on) the spatial barcode array. The tissue sample may then be subjected to conditions sufficient to transfer peptides or proteins to the spatial barcode array. For example, the peptides or proteins may be transported via passive transport, e.g., diffusion or Brownian motion, or active transport, e.g., electrophoresis, pressure-driven flow, etc. Peptides or proteins can be attached to the spatial barcode using, for example, a linker (e.g., an amine-reactive group, or a click chemistry group, such as an azide, alkyne, or other functional moiety, such as an aldehyde group, NHS, or carboxyl group), conjugation chemistry, or an anchoring agent to generate a barcode-tagged (barcoded) protein. Examples of anchoring agents include fixatives such as formaldehyde, paraformaldehyde, glutaraldehyde, or monomers for incorporation into hydrogels, such as acryloyl-X, acrylamide, N-(3-aminopropyl)methacrylamide, N-(3-aminoethyl)methacrylamide, or benzophenone. Anchoring agents can also include conjugated linkers, such as acryloyl-X, biotin-NHS, biotin-PEG-amine, DBCO-NHS, DBCO-amine, etc. For bead arrays, multiple beads can collect samples for further processing.
[0221] Computer Systems The present disclosure provides a computer system programmed to implement the disclosed method. Figure 5 shows a computer system (501) programmed or otherwise configured to receive sequencing data and output information regarding the identity of amino acid residues of polymerizable molecules or modified amino acids. The computer system (501) can control various aspects of generating sequencing reads of the present disclosure, such as receiving one or more sets of sequencing data, processing the sequencing data using an algorithm, and outputting one or more sequencing results. The computer system (501) can be an electronic device of a user or a computer system, and can be located remotely from the electronic device. The electronic device can be a mobile electronic device.
[0222] The computer system (501) includes a central processing unit (CPU, also referred to herein as a "processor" and a "computer processor") (505), which may be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system (501) also includes memory or storage locations (510) (e.g., random access memory, read-only memory, flash memory), electronic storage devices (515) (e.g., hard disks), communication interfaces (520) (e.g., network adapters) for communicating with one or more other systems, and peripherals (525), such as cache, other memory, data storage devices, and / or electronic display adapters. The memory (510), storage devices (515), interfaces (520), and peripherals (525) communicate with the CPU (505) via a communication bus (solid lines), such as a motherboard. The storage devices (515) may be a data storage device (or data repository) for storing data. The computer system (501) may be operatively connected to a computer network ("network") (530) with the aid of the communication interface (520). The network (530) may be the Internet and / or an extranet, an intranet and / or an extranet in communication with the Internet. In some cases, the network (530) is a telecommunications and / or data network. The network (530) may include one or more computer servers, which may enable distributed computing, such as cloud computing. In some cases, the network (530) may implement a peer-to-peer network with the aid of the computer system (501), thereby enabling devices coupled to the computer system (501) to act as clients or servers.
[0223] The CPU (505) can execute a series of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a storage location, such as memory (510). The instructions may be directed to the CPU (505), which may then program or otherwise configure the CPU (505) to implement the methods of the present disclosure. Examples of operations performed by the CPU (505) include fetch, decode, execute, and writeback.
[0224] The CPU 505 may be part of a circuit, such as an integrated circuit. One or more other components of the system 501 may also be included in the circuit. In some cases, the circuit is an application-specific integrated circuit (ASIC).
[0225] The storage device (515) can store files such as drivers, libraries, and saved programs. The storage device (515) can store user data, such as user preferences and user programs. The computer system (501) may optionally include one or more additional data storage devices external to the computer system (501), such as located on a remote server in communication with the computer system (501) via an intranet or the Internet.
[0226] The computer system (501) can communicate with one or more remote computer systems via the network (530). For example, the computer system (501) can communicate with a user's remote computer system. Examples of remote computer systems include a personal computer (e.g., a portable PC), a slate or tablet PC (e.g., an Apple® iPad®, a Samsung® Galaxy Tab), a telephone, a smartphone (e.g., an Apple® iPhone®, an Android-enabled device, a Blackberry®), or a personal digital assistant. A user can access the computer system (501) via the network (530).
[0227] The methods described herein can be performed by machine (e.g., a computer processor) executable code stored in an electronic storage location of the computer system (501), such as, for example, on memory (510) or electronic storage (515). The machine-executable or machine-readable code can be provided in the form of software. During use, the code can be executed by the processor (505). In some cases, the code can be retrieved from storage (515) and stored in memory (510) for immediate access by the processor (505). In some situations, the electronic storage (515) can be omitted, and machine-executable instructions are stored in memory (510).
[0228] The code may be pre-compiled and configured for use with a machine having a processor adapted to execute the code, or may be compiled at run time. The code may be supplied in a programming language that can be selected to render the code executable in a pre-compiled or as-compiled fashion.
[0229] Aspects of the systems and methods provided herein, such as the computer system (501), may be embodied in programming. Various aspects of this technology may be considered as "products" or "articles of manufacture," typically in the form of machine- (or processor-) executable code and / or associated data carried on or embedded in a type of machine-readable medium. The machine-executable code may be stored in electronic storage devices such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. "Storage" type media may include any or all of the tangible memory of a computer or processor, or its associated modules, such as various semiconductor memories, tape drives, disk drives, etc., which may provide non-transitory recording media at any time for programming the software. All or portions of the software are sometimes communicated via the Internet or various other telecommunications networks. Such communication may enable, for example, loading of the software from one computer or processor to another, e.g., from a management server or host computer to an application server computer platform. Thus, other types of media that may bear software elements include light waves, radio waves, and electromagnetic waves, such as those used across physical interfaces between local devices, via wired and optical landline networks, and on various air-links. Physical elements that carry such waves, such as wired or wireless links, optical links, etc., may also be considered software-bearing media. As used herein, unless limited to non-transitory, tangible "storage" media, terms such as computer or machine "readable medium" refer to media that participate in providing instructions to a processor for execution.
[0230] Thus, machine-readable media such as computer-executable code may take many forms, including, but not limited to, tangible storage media, carrier wave media, or physical transmission media. Non-volatile storage media include, for example, optical or magnetic disks, any of the storage devices in a computer, such as those that may be used to implement the databases shown in the figures. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that comprise a bus within a computer system. Carrier-wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer-readable media include, for example: floppy disks, flexible disks, hard disks, magnetic tape, other magnetic media, CD-ROMs, DVDs or DVD-ROMs, other optical media, punch cards, paper tape, other physical storage media with patterns of holes, RAM, ROM, PROMs and EPROMs, FLASH-EPROMs, other memory chips or cartridges, carrier waves carrying data or instructions, cables or links which transmit such carrier waves, or other media from which a computer can read programming code and / or data. Many of these forms of computer-readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0231] The computer system (501) may include or be in communication with an electronic display (535) that includes a user interface (UI) (540) for providing, for example, sequencing data results, identities of modified amino acids, identities of polymerizable molecules, or nucleic acid sequences of nucleic acid molecules. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0232] The disclosed methods and systems can be implemented by one or more algorithms. The algorithms can be implemented by software when executed by a central processing unit (505). The algorithms can, for example, process sequencing data (e.g., nanopore or nanogap ionic current signals) to generate one or more sequencing outputs (e.g., identification of amino acids or monomers of a polymerizable molecule, e.g., a nucleic acid sequence of a nucleic acid molecule). [Example]
[0233] Example 1 - Preparation of a linker for attaching polymerizable molecules to amino acids As described herein, a linker may contain two functional components. The first functional component may be a moiety capable of reacting with an amino acid. This moiety may be, for example, an isothiocyanate (e.g., PITC), dinitrofluorobenzene, dansyl chloride, or other amino acid reactive group. The second functional component may, in some instances, contain a reactive group capable of binding to a polymerizable molecule, which may be a nucleic acid molecule (e.g., DNA, RNA, LNA, PNA), a peptide, or a synthetic organic compound. In the case of a DNA-based polymerizable group, the DNA may contain sequences for primers, sites for enzymatic ligation and chemical conjugation, barcode sequences, or combinations thereof. In some instances, the linker may be pre-functionalized to contain a polymerizable molecule (e.g., as shown in panel B of Figure 4):
[0234] Method 1: Preparation of linker-DNA sequence conjugates DNA sequences ranging in length from 15 to 150 bp are synthesized using either internal azide modification (e.g., azido-dT modification, Integrated DNA Technology) or external azide modification at the 5' or 3' end. Linkers containing (i) a PITC moiety (amino acid reactive group) and (ii) an alkyne moiety may be conjugated to the DNA sequence via copper-catalyzed click chemistry as follows: 100 μM linker and 10 μM DNA sequence are prepared in a reaction buffer consisting of 200 μM Tris((1-benzyl-1H-1,2,3-triazol-4-yl)methyl)amine (TBTA), 500 μM TCEP, and 200 μM CuSO4, 1x PBS, pH 7.86. The reaction is incubated at room temperature for 2 hours. The linker-DNA sequence conjugate is then purified from the reaction using reverse-phase high-pressure liquid chromatography (HPLC).
[0235] Method 2: Preparation of linker-peptide conjugates Peptide sequences (ranging from 10 to 50 aa in length) may be synthesized with internal cysteine residues. Peptides may be alkylated with iodoacetamide azide as follows: Peptides are reduced by preparing them at 1 mg / ml in a solution containing 1% SDS and 100 mM ammonium bicarbonate (pH 8.0). To this solution, Tris(2-carboxyethylphosphine hydrochloride) (TCEP-HCl) is added to a final concentration of 10 mM and incubated at 55°C for 1 hour. Iodoacetamide azide is then added to a final concentration of 10 mM and incubated at room temperature for 1 hour. The alkylated peptides are then purified by two buffer exchanges using a buffer exchange column. The alkylated peptides are then converted via copper-catalyzed click chemistry to peptides containing (i) a PITC moiety (an amino acid reactive group) and (ii) an alkyne moiety. The linker is reacted with the alkylated peptide sequence. 100 μM linker and 10 μM alkylated peptide sequence are prepared in a reaction buffer consisting of 200 μM Tris((1-benzyl-1H-1,2,3-triazol-4-yl)methyl)amine (TBTA), 500 μM TCEP, and 200 μM CuSO4, 1x PBS, pH 7.86. The reaction is incubated at room temperature for 2 hours. The linker (linker-peptide sequence conjugate) is then purified from the reaction using reverse-phase high-pressure liquid chromatography (HPLC).
[0236] Example 2 - Nanopore detection of amino acids conjugated to polymeric molecules To determine whether modified amino acids or their derivatives, such as AALC complexes or stacked AALC complexes, or portions thereof, can be detected by nanopores, amino acid-linker-polymerizable molecule complexes are generated and transported through a nanopore sequencing system. The amino acid-linker-polymerizable molecule complexes are generated by attaching a polymerizable molecule to the amino acid-linker complex. The polymerizable molecule includes an alkyne-linked DNA molecule. More specifically, the polymerizable molecule includes an alkyne linker to the T236 position within a 16-nt poly-dT DNA oligo scaffold located in the middle of one strand of a 450-nt dsDNA backbone, as illustrated schematically in Figure 6A. The amino acid-linker complex is generated using 1-(2-azidoethyl)-4-isothiocyanatobenzene, a bifunctional linker containing (1) a PITC moiety (an amino acid reactive group) and (2) an azide moiety capable of reacting with the alkyne of the polymerizable molecule. As also shown in the bottom panel of Figure 6A, the bifunctional linker is reacted with three different amino acids: Asp, Trp, and Tyr via the PITC moiety to generate an amino acid-linker conjugate, which is then reacted with an alkyne-linked DNA molecule to generate an amino acid-linker-polymerizable molecule conjugate.
[0237] The amino acid-linker-polymerizable molecule conjugates were then subjected to a nanopore sequencing system (Oxford Nanopore Technologies MinION) to obtain information on whether the generated current signal could be reassigned to an amino acid type. An unmodified ("unmodified") DNA backbone was used as a baseline control, and an alkyne-linked DNA molecule ("alkyne") was used as a control for alkyne linker-specific current blockade. Figure 6B shows the current traces and the assignment of the nucleotide sequence of the DNA backbone. It is noteworthy that all the different molecules passed through the nanopore, and discernible qualitative differences or perturbations in the current signal could be visualized from the amino acid-linker-polymerizable molecule conjugates (labeled "linker-Asp," "linker-Trp," or "linker-Tyr") that were not present in the current signal of the control molecule. Figure 6C shows the current traces (measured current signals) as a function of time for five conditions: unmodified, alkyne, linker-Asp, linker-Trp, and linker-Tyr. The combined overlay traces for all readouts from the nanopores in the device are shown on the left, and three representative individual traces are shown on the right. Figure 6D shows the current traces (measured current signals) as a function of base pair (nucleotide) position for five samples: unmodified, alkyne, linker-Asp, linker-Trp, and linker-Tyr. The combined overlay traces for all readouts from the nanopores in the device are shown on the left, with the average signal shown in the center and error bars representing standard deviation. The radial plots on the right show the base or nucleotide position (along the radial axis) versus the median measured current for each of the five conditions tested. Interestingly, the presence of the amino acid-linker complex on the polymerizable molecule affected the current signal 2 nt before and 5 nt after the conjugation site, suggesting that the amino acid-linker polymerizable molecule spacing should be greater than 8 nt to minimize interference from neighboring amino acids.Next, we use a custom machine learning algorithm model to determine how different the current traces of the amino acid-linker-polymerizable molecule complexes are from control molecules. Figure 6E shows a normalized confusion matrix between the classified molecules (unmodified, alkyne, linker-Asp, linker-Trp, and linker-Tyr) on the X-axis and the actual molecules on the Y-axis. As can be visualized, we achieve a reasonable degree of accuracy in classifying the correct molecules.
[0238] Example 3 - Multiplexed readout of amino acids conjugated to barcoded polymeric molecules To determine whether nanopore detection of modified amino acids or their derivatives containing multiplexed information, e.g., barcoded polymerizable molecules (e.g., nucleic acid barcode molecules), is possible, amino acid-linker-barcoded polymerizable molecule conjugates are generated in a similar manner to that described above. The barcoded polymerizable molecules contain a 4-nt barcode sequence located 5' of an alkyne linker located at position T236 within a 450-nt ssDNA backbone, as illustrated schematically in Figure 7A. Two types of ssDNA backbones are tested. The first type contains a polydT region, as described above and illustrated schematically in Figure 6A, and the second type contains a mixed-base region, as shown in Figure 7A. Next, amino acid-linker conjugates are generated using 1-(2-azidoethyl)-4-isothiocyanatobenzene, a bifunctional linker containing (1) a PITC moiety (an amino acid reactive group) and (2) an azide moiety capable of reacting with the alkyne of the polymerizable molecule. The bifunctional linker is reacted with two different amino acids, Asp and Tyr, via the PITC moiety to generate an amino acid-linker conjugate (identical to that shown in Figure 6A), which is then reacted with an alkyne-linked DNA molecule to generate an amino acid-linker-polymerizable molecule conjugate.
[0239] The amino acid-linker-polymerizable molecule complexes were then subjected to a nanopore sequencing system (Oxford Nanopore Technologies MinION) to determine whether the generated current signals could be reassigned to amino acid types (Asp or Tyr) and whether the barcode sequences of the barcoded polymerizable molecules could be identified. Alkyne-linked DNA molecules ("alkyne") without the amino acid-linker complex were used as a control for alkyne linker-specific current blockade. Figure 7B shows current traces (measured current signals) as a function of time (ms) for amino acid-linker-polymerizable molecule complexes of barcoded polymerizable molecules containing poly-dT sequences ("poly-dT backbone") and mixed-base sequences ("mixed-base backbone"). For both types of polymerizable molecules, distinguishable differences or perturbations in the current signals could be visualized from the amino acid-linker-polymerizable molecule conjugates ("linker-Asp" and "linker-Tyr"). Figure 7C shows current traces (measured current signals) as a function of base (nucleotide) position for three conditions: linker-Tyr, linker-Asp, and alkyne (control) for both types of polymerizable molecules. Interestingly, mixed-base polymerizable molecules yielded more distinguishable current traces compared to poly-dT polymerizable molecules, suggesting that using mixed-base polymerizable molecules instead of repetitive nucleotide sequences may yield more distinguishable current traces. These results also suggest that multiplexed readout of both DNA sequence and amino acid type is possible.
[0240] Example 4 - Effect of polymerizable molecule linker length on nanopore current readout of amino acids conjugated with polymerizable molecules As described herein, a polymerizable molecule of a modified amino acid or its derivative (e.g., an AALC complex or a stacked AALC complex or portion thereof) can include a linker containing a click chemistry moiety, which can be attached to the amino acid-linker complex or another linker containing (i) a complementary click chemistry moiety and (ii) an amino acid reactive group. To determine whether the intramolecular distance between the click chemistry moiety and the polymerizable molecule (i.e., linker length) can affect nanopore detection of a modified amino acid or its derivative, such as an amino acid-linker-polymerizable complex, two different amino acid-linker-polymerizable complexes were generated. As shown in Figure 7A, the first amino acid-linker-polymerizable complex contains an octadiynyl dU linker, and as shown in Figure 8A, the second amino acid-linker-polymerizable complex contains an ethynyl dU linker. Each amino acid-linker-polymerizable conjugate contains the same DNA backbone, which contains a 4-nt barcode sequence located 5' of a modified nucleotide (octadiynyl dU or ethynyl dU) at position 199 of the 400-nt DNA backbone. The amino acid-linker conjugates are generated using 1-(2-azidoethyl)-4-isothiocyanatobenzene, a bifunctional linker containing (1) a PITC moiety (an amino acid reactive group) and (2) an azide moiety capable of reacting with an alkyne on the polymerizable molecule. The bifunctional linker is reacted with two different amino acids, Asp and Tyr, via the PITC moiety to generate the amino acid-linker conjugate (identical to that shown in Figure 6A). The amino acid-linker conjugate is then reacted with an alkyne-linked DNA molecule via copper-catalyzed click chemistry to generate the amino acid-linker-polymerizable molecule conjugate.
[0241] Figure 8B shows current traces (background-subtracted measured current signals) as a function of time (top) or base pair (nucleotide) position (bottom) for amino acid-linker-polymerizable molecule conjugates containing an ethynyl dU linker for five different molecules: a control polymerizable molecule ("alkyne") with an ethynyl dU linker but no amino acid-linker conjugate, and amino acid-linker-polymerizable molecule conjugates containing linker-Glu, linker-Phe, linker-Gly, and linker-Trp. The current traces for the conjugates containing amino acids (linker-Glu, linker-Phe, linker-Gly, and linker-Trp) show clear differences compared to the control molecule without an amino acid-linker conjugate.
[0242] Figure 8C shows current traces (measured current signals) as a function of time for amino acid-linker-polymerizable molecule complexes containing either an ethynyl dU linker (top) or an octadiynyl dU linker (bottom), for four or five different molecules: a control polymerizable molecule with an ethynyl dU linker or an octadiynyl dU linker but no amino acid-linker conjugate ("alkyne"), and amino acid-linker-polymerizable molecule complexes containing linker-Glu, linker-Phe, linker-Gly, linker-Trp (for an ethynyl dU linker), or linker-Asp, linker-Tyr, or linker-Trp (for an octadiynyl dU linker). The current traces seem to suggest that the longer the linker size (octadiynyl dU), the greater the difference between the modified molecule (amino acid-linker-polymerizable molecule complex) and the unmodified control molecule (polymerizable molecule without the amino acid-linker conjugate attached).
[0243] Figure 8D shows a t-SNE plot of the dynamic time warping (DTW) correlation matrix containing nanopore current traces for amino acid-linker-polymerizable molecule complexes containing either an ethynyl dU linker (left) or an octadiynyl dU linker (right) compared to a control polymerizable molecule with either an ethynyl dU linker or an octadiynyl dU linker but without the amino acid-linker conjugate ("alkyne"). The linker-Trp ("W") amino acid-linker-polymerizable molecule conjugate is shown. The t-SNE plot shows that the longer the linker size (octadiynyl dU), the greater the difference between the modified molecule (amino acid-linker-polymerizable molecule complex, "W") and the unmodified control molecule (polymerizable molecule without the amino acid-linker conjugate attached, "alkyne").
[0244] Example 5 - Effect of modified amino acids or their derivatives on transport rate To determine the effect of modified amino acids or their derivatives (e.g., amino acids containing polymerizable molecules, AALC complexes, stacked AALC complexes, or portions thereof) on the nanopore sequencer readout, residence times are measured. The residence time of a base is the duration of a single base's presence in the nanopore and corresponds to the transport rate of the molecule through the nanopore. Several amino acid-linker-barcoded polymerizable molecule conjugates are measured. Amino acid-linker-polymerizable molecules are synthesized as described above, using an octadiynyl dU linker on the polymerizable molecule, as illustrated schematically in Figure 6A. The amino acid-linker-polymerizable molecule conjugates are run through a nanopore sequencing system (Oxford Nanopore Technologies MinION) to obtain information on whether modifications to the polymerizable molecule, such as the addition of an amino acid-linker complex, affect the transport rate. An unmodified ("unmodified") DNA backbone is used as a baseline control, and an alkyne-linked DNA molecule ("alkyne") is used as a control for alkyne linker-specific current blockade. Figure 9 shows a heat map of the measured residence times for a control molecule and three different amino acid-linker-polymerizable molecule complexes (labeled "Linker-Asp," "Linker-Trp," or "Linker-Tyr"). The x-axis represents base pair position, which is cropped to show a 25 nt region containing and surrounding the poly-dT sequence, with the alkyne linker at position 13 nt. The y-axis represents individual current traces. The relative intensity at each time point represents the relative residence time at that time point (each corresponding to 0.25 ms). Notably, the measured residence times are different across the amino acid-linker-polymerizable molecule complexes compared to the control molecule. Interestingly, the modifications (addition of an alkyne linker or an alkyne linker conjugated to an amino acid-linker complex) appear to affect the dwell time of the DNA backbone at several positions, including base pair positions 0–5 and 10–25 surrounding the modification site, suggesting that some molecular stretching may occur and that the presence of AALC may affect a large window of nucleotides surrounding the modification site.
[0245] Example 6 - Nanopore sequencing of stacked amino acid-polymeric complexes To determine whether modified amino acids or their derivatives, such as stacked AALC complexes or portions thereof, can be detected through nanopores, we first generated amino acid-linker-polymerizable molecule conjugates containing two amino acids, as illustrated schematically in Figure 10A. The polymerizable molecule contains an alkyne linker (octadiynyl dU) at two positions within a 252-nt DNA backbone, as shown in Figure 10B. Next, we generated amino acid-linker conjugates using 1-(2-azidoethyl)-4-isothiocyanatobenzene, a bifunctional linker containing (1) a PITC moiety (an amino acid reactive group) and (2) an azide moiety capable of reacting with the alkyne of the polymerizable molecule. As also shown in the bottom panel of Figure 6A, we generated amino acid-linker conjugates by reacting the bifunctional linker with three different amino acids: Asp, Trp, and Tyr via the PITC moiety. The amino acid-linker conjugate is then reacted with an alkyne-linked DNA molecule to generate a double amino acid-linker-polymerizable molecule conjugate containing a pair of amino acids. Figure 10C shows a table of different pairs of amino acids conjugated at the first linker position or the second linker position. An unmodified ("unmodified") DNA backbone is used as a baseline control, and an alkyne-linked DNA molecule ("alkyne") is used as a control for alkyne linker-specific current blockade.
[0246] The dual amino acid-linker-polymerizable molecule conjugates are then subjected to a nanopore sequencing system (Oxford Nanopore Technologies MinION) to obtain information on whether the generated current signal can be reassigned to the amino acid type present in the dual amino acid-linker-polymerizable molecule complex. Figure 10D shows the current traces (measured current signals) as a function of time (top) or base pair position (bottom). Similar to the dual amino acid-linker-polymerizable molecule complexes containing a single amino acid, dual amino acid-linker-polymerizable molecule complexes containing two amino acids pass through the nanopore, and each amino acid at its respective position (represented by an arrow) in the amino acid-linker-polymerizable molecule conjugate (labeled "Asp / Trp," "Asp / Tyr," or "Trp / Tyr") can visualize discernible differences or perturbations not present in the current signal of the control molecule. A custom machine learning algorithm model is then used to determine the extent to which the current traces of the dual amino acid-linker-polymerizable molecule complexes containing two amino acids differ from those of the control molecule. Figure 10E shows the normalized confusion matrix between the predicted molecules (unmodified, alkyne, Asp / Trp, Asp / Tyr, Trp / Tyr) on the X-axis and the actual molecules on the Y-axis. As can be visualized, we achieve high accuracy in classifying correct molecules, ranging from 92% to 98%. Overall, these results suggest that processing amino acids to generate modified amino acids, including polymerizable molecules, may be a promising approach for performing high-accuracy peptide sequencing using nanopore or nanochannel approaches.
[0247] Example 7 - Temperature regulation to improve current signal-to-noise ratio As described herein, the transport rate of a modified amino acid or its derivative may differ from that of an unmodified amino acid; similarly, the transport rate of a modified amino acid or its derivative containing a polymerizable molecule may differ from that of a polymerizable molecule without an attached amino acid. One hypothesis is that the transport rate of a molecule through a nanopore may be reduced by lowering the ambient temperature. To test this hypothesis, unmodified polymerizable molecules (those without an amino acid-linker complex) can be subjected to a nanopore sequencing system at ambient temperature (room temperature) or on ice.
[0248] As illustrated schematically in Figure 11A, the polymerizable molecules contain a 450-nt dsDNA backbone containing a 16-nt poly-dT sequence (selected sequences of approximately 225 bp to 245 bp are shown). The polymerizable molecules are run through a nanopore sequencing system (Oxford Nanopore Technologies MinION) to obtain information on the effect of temperature on sequencing output. The flow cell temperature for measurements performed at room temperature is 34°C, while measurements performed on ice are approximately 19°C to 20°C. Figure 11A shows current traces (measured current signals) as a function of base pair position and identified nucleotides for selected sequences from the polymerizable molecules. Interestingly, the signal-to-noise ratio for the 16-nt poly-dT sequence appears qualitatively higher.
[0249] Figure 11B shows a histogram of transport duration, a measure of the time a region or entire polymerizable molecule is present within the nanopore, corresponding to the region or entire transport rate of the polymerizable molecule. The transport duration of the polymerizable molecule is measured in milliseconds (ms) at two different temperature conditions (room temperature or ice). The plot on the left shows the transport duration of the entire polymerizable molecule, while the plot on the right shows the transport duration of selected regions (positions 200–275 nt) surrounding the 16-nt poly-dT sequence. As can be visualized from the histogram, ice conditions result in a slower transport rate (longer transport durations) and a greater variance or range in the measured duration values.
[0250] Figure 11C shows histograms of transport durations (in ms) of polymerizable molecules under two different temperature conditions (room temperature or ice). The left plot shows the transport durations of polymerizable molecules 200-225 nt upstream of the 16-nt poly-dT sequence. The middle plot shows the transport durations of selected sites (225-250 nt) immediately adjacent to and surrounding the 16-nt poly-dT sequence. The right plot shows the transport durations of polymerizable molecules 250-275 nt downstream of the 16-nt poly-dT sequence. As can be visualized in all three plots, ice conditions result in a slower transport rate (longer transport durations) and a greater variability or range of measured duration values for both the poly-dT region and the adjacent regions.
[0251] Figure 11D shows the current traces (measured current signals) as a function of base position from 200 nt to 280 nt. The plot on the left shows the current traces at room temperature, and the plot on the right shows the current traces at ice conditions. Ice conditions improve the signal-to-noise ratio of the current, as evidenced by the reduced size of the error bars in the poly-dT region (approximately 225 nt to 245 nt).
[0252] Figure 11E shows the transport duration in ms as a function of base position, from the 200-280 nt position. The plot on the left shows the current trace at room temperature, and the plot on the right shows the current trace at ice. Interestingly, while ice conditions result in an increase in the dwell time per base position, they also result in increased noise in the transport duration, except for the poly-dT region (approximately 225-245 bp).
[0253] Overall, the results in Figures 11A-11E suggest that lowering the temperature can result in an improvement in the signal-to-noise ratio of the current trace read from the polymerizable molecules in a nanopore sequencing system, which may be applicable to improving the readout accuracy of modified amino acids or their derivatives.
[0254] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the present invention be limited by the specific examples provided within the specification. While the present invention has been described with reference to the foregoing specification, the description and illustration of the embodiments herein are not meant to be construed in a limiting sense. Numerous modifications, changes, and substitutions will occur to those skilled in the art without departing from the invention. Furthermore, it will be understood that all aspects of the present invention are not limited to the specific depictions, configurations, or relative proportions set forth herein, which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the present invention described herein may be utilized in practicing the invention. It is therefore contemplated that the present invention shall cover any such alternatives, modifications, variations, or equivalents. The following claims define the scope of the present invention, and methods and structures within the scope of these claims and their equivalents are intended to be covered thereby.
Claims
1. (a) providing a modified amino acid, wherein the modified amino acid comprises a polymerizable molecule; (b) transporting the modified amino acid or derivative thereof through a nanopore or nanogap; A method for processing the modified amino acid, comprising: wherein the polymerizable molecule causes a change in the transport rate of the modified amino acid or its derivative through the nanopore or nanogap compared to the transport rate of an amino acid that does not contain the polymerizable molecule.
2. (a) providing a modified amino acid, wherein the modified amino acid comprises a polymerizable molecule; (b) transporting the modified amino acid or derivative thereof through a nanopore or nanogap; (c) measuring a signal from the nanopore or the nanogap; wherein the polymerizable molecule causes an increase in the measured signal-to-noise ratio (SNR) of the signal in (c) compared to the SNR of an amino acid that does not contain the polymerizable molecule.
3. 3. The method of claim 1 or 2, wherein the modified amino acids comprise or are derived from proteinogenic amino acids attached to the polymerizable molecule.
4. The method of any one of claims 1 to 3, wherein the modified amino acid or derivative thereof further comprises a binding agent.
5. The method of claim 4 , wherein the binding agent comprises an antibody, an antibody fragment, a nanobody, an aptamer, a peptide, a polymer, an inorganic compound, a small molecule, or a derivative thereof.
6. The method of any one of claims 1 to 5, wherein the modified amino acid comprises a non-naturally occurring chemical modification.
7. The method of claim 6 , wherein the non-natural chemical modification is a protecting group.
8. The method of any one of claims 1 to 7, wherein the nanopore binds a helicase.
9. 9. The method of claim 8, wherein the helicase directs the modified amino acid or derivative thereof through the nanopore.
10. The method of any one of claims 1 to 9, wherein (b) is carried out using a topoisomerase or a polymerase.
11. The method of any one of claims 1 to 9, wherein (b) is carried out using an electric or magnetic field.
12. The method of any one of claims 1 to 11, wherein the polymerizable molecule comprises a nucleic acid molecule.
13. 13. The method of claim 12, wherein the nucleic acid molecule comprises a deoxyribonucleic acid (DNA) molecule, a xenonucleic acid (XNA) molecule, a ribonucleic acid (RNA) molecule, or a modified variant thereof.
14. The method of any one of claims 1 to 13, wherein the nanopore comprises a transmembrane protein.
15. The method of any one of claims 1 to 13, wherein the nanogap comprises an inorganic material.
16. The method of claim 15 , wherein the nanogap comprises silicon nitride or molybdenum sulfate.
17. The method of any one of claims 1 to 16, wherein the polymerizable molecule is covalently attached to the modified amino acid.
18. The method of any one of claims 1 to 17, further comprising, before (a), a step of generating the modified amino acid from an amino acid of the peptide.
19. 19. The method of claim 18, wherein the amino acid is located at a terminus of the peptide.
20. 20. The method of claim 19, wherein the terminus is the N-terminus.
21. 20. The method of claim 19, wherein the modified amino acid comprises the polymerizable molecule attached to a terminal amino acid.
22. 22. The method of claim 21, wherein the polymerizable molecule is attached to the terminal amino acid via an isothiocyanate moiety.
23. 23. The method of any one of claims 1 to 22, further comprising, prior to (a), (i) providing a linker, wherein the linker comprises an amino acid reactive moiety, and (ii) contacting the amino acid reactive moiety with a proteinogenic amino acid to obtain the modified amino acid.
24. The method of claim 23 , wherein the linker is attached to the polymerizable molecule.
25. 24. The method of claim 23, further comprising attaching the linker to the polymerizable molecule.
26. 26. The method of claim 25, wherein the linker comprises a first reactive moiety and the polymerizable molecule comprises a second reactive moiety, the method further comprising reacting the first reactive moiety with the second reactive moiety to produce the linker attached to the polymerizable molecule.
27. 27. The method of claim 26, wherein the first reactive moiety or the second reactive moiety comprises a click chemistry moiety.
28. 27. The method of claim 26, wherein the polymerizable molecule comprises a nucleic acid linker, the nucleic acid linker comprising a modified nucleobase and a click chemistry moiety.
29. 28. The method of any one of claims 23 to 27, wherein the amino acid reactive moiety comprises phenylisothiocyanate (PITC).
30. 30. The method of any one of claims 23 to 29, wherein the proteinogenic amino acid is the terminal amino acid of a peptide.
31. 31. The method of claim 30, further comprising tethering the modified amino acid to a capture moiety and cleaving the terminal amino acid from the peptide to yield a remaining peptide and an amino acid-linker-capture moiety (AALC) conjugate.
32. The method of claim 31 , wherein the capture moiety is bound to a substrate.
33. 32. The method of claim 31 , wherein the capture moiety comprises a first nucleic acid molecule and the polymerizable molecule comprises a second nucleic acid molecule.
34. 34. The method of Claim 33, wherein the first nucleic acid molecule or the second nucleic acid molecule comprises a nucleic acid barcode molecule.
35. 35. The method of claim 34, wherein the nucleic acid barcode molecule identifies the peptide.
36. 36. The method of any one of claims 33 to 35, wherein the nucleic acid barcode molecule comprises temporal information.
37. The method of any one of claims 31 to 36, wherein the capture moiety comprises a cleavable moiety.
38. 37. The method of any one of claims 33 to 36, wherein the tethering step comprises binding the first nucleic acid molecule to the second nucleic acid molecule.
39. 39. The method of claim 38, wherein the linking is performed using a ligase.
40. 39. The method of claim 38, wherein the binding comprises (i) providing a splint oligonucleotide comprising a first sequence complementary to at least a portion of the first nucleic acid molecule and a second sequence complementary to at least a portion of the second nucleic acid molecule, and (ii) hybridizing the first sequence to the portion of the first nucleic acid molecule and the second sequence to the portion of the second nucleic acid molecule.
41. 34. The method of claim 33, further comprising the steps of: (I) providing an additional linker attached to an additional nucleic acid molecule; (II) contacting the additional linker with the terminal amino acid of the remaining peptide to obtain an additional modified amino acid; and (III) tethering the additional modified amino acid to the AALC complex to obtain the derivative of the modified amino acid.
42. 42. The method of claim 41, further comprising the step of repeating (I) to (III) to obtain the derivative of the modified amino acid, wherein the derivative of the modified amino acid comprises a plurality of modified amino acids.
43. 42. The method of claim 41 , wherein the nucleic acid molecule or the additional nucleic acid molecule comprises a barcode sequence.
44. 44. The method of any one of claims 41 to 43, wherein the nucleic acid molecule comprises a first sequence and the further nucleic acid molecule comprises a second sequence, and the first sequence is different from the second sequence.
45. 45. The method of claim 44, wherein the first sequence or the second sequence comprises temporal or spatial information.
46. 46. The method of claim 45, wherein the temporal information is an index of the cycle or repeat number in which the linker or the additional linker is provided.
47. The method of claim 1 , wherein the transport rate of the modified amino acid is slower than the transport rate of the amino acid that does not contain the polymerizable molecule.
48. The method of claim 1 , further comprising determining the transport rate by measuring a current signal of the nanopore or the nanogap.
49. 49. The method of any one of claims 1 to 48, further comprising using the nanopore or nanogap to identify the modified amino acid or derivative thereof.
50. 50. The method of claim 49, wherein the identifying step comprises measuring a current signal from the nanopore or nanogap and determining the amino acid type or post-translationally modified variant of the modified amino acid or derivative thereof.
51. 3. The method of claim 1 or 2, wherein (b) is carried out at a temperature below ambient temperature.
52. 1. A method for processing a peptide, comprising: (a) providing the peptide and a linker, wherein the linker is capable of binding to an amino acid of the peptide; (b) attaching the linker to the amino acid of the peptide; (c) attaching the linker to a capture moiety; (d) cleaving the amino acid from the peptide to obtain an amino acid-linker-capture moiety (AALC) conjugate; (e) providing an additional linker, said additional linker being capable of binding to another amino acid of said peptide; (f) forming an additional amino acid-linker conjugate by attaching the additional linker to the other amino acid; (g) attaching the additional amino acid-linker conjugate to the AALC conjugate to form a stacked AALC conjugate; wherein the bonding step in (c) is carried out using a polymerizable molecule.
53. 53. The method of claim 52, wherein the linker comprises a reactive moiety.
54. 54. The method of claim 53, wherein the polymerizable molecule comprises an additional reactive moiety capable of reacting with the reactive moiety of the linker, and prior to (a), the polymerizable molecule is attached to the linker by reacting the reactive moiety with the additional reactive moiety.
55. 55. The method of claim 54, wherein the polymerizable molecule comprises an additional linker comprising the additional reactive moiety, the additional linker comprising a modified nucleobase, and the additional reactive moiety comprises a click chemistry moiety.
56. 54. The method of claim 53, further comprising, prior to (c), providing the polymerizable molecule.
57. 57. The method of any one of claims 52 to 56, wherein the capture moiety is attached to a substrate.
58. 58. The method of claim 57, wherein the substrate is substantially planar.
59. 58. The method of claim 57, wherein the substrate is a bead.
60. 60. The method of any one of claims 52 to 59, further comprising providing said capture moiety, wherein said capture moiety comprises a nucleic acid molecule.
61. 61. The method of claim 60, wherein the nucleic acid molecule comprises a DNA molecule, an RNA molecule, an XNA molecule, or a modified variant thereof.
62. 62. The method of claim 61 , wherein the DNA molecule is single-stranded.
63. 62. The method of claim 61 , wherein the polymerizable molecule comprises an additional nucleic acid molecule.
64. 64. The method of claim 63, wherein (c) comprises binding the additional nucleic acid molecule to the nucleic acid molecule of the capture moiety.
65. 65. The method of claim 64, wherein the linking is performed using a ligase.
66. 65. The method of claim 64, wherein at least a portion of the additional nucleic acid molecule is complementary to at least a portion of the nucleic acid molecule of the capture moiety.
67. 67. The method of claim 66, wherein said binding is effected by hybridizing said at least a portion of said additional nucleic acid molecule to said at least a portion of said nucleic acid molecule.
68. 65. The method of claim 64, wherein binding the additional nucleic acid molecule to the nucleic acid molecule of the capture moiety is performed using a splint oligonucleotide.
69. 69. The method of any one of claims 60-68, wherein the capture moiety comprises a nucleic acid barcode molecule.
70. 70. The method of claim 69, wherein said nucleic acid barcode molecule identifies said peptide.
71. The method of any one of claims 60 to 70, wherein the capture moiety is bound to the peptide.
72. 72. The method of any one of claims 52 to 71, wherein the amino acid or the further amino acid is the N-terminal amino acid or the C-terminal amino acid.
73. 73. The method of any one of claims 52 to 72, wherein (b) is performed before (c).
74. 73. The method of any one of claims 52 to 72, wherein (c) is performed before (b).
75. 75. The method of any one of claims 52 to 74, further comprising repeating (a) to (g).
76. 76. The method of any one of claims 52 to 75, further comprising identifying the AALC complex or the stacked AALC complex.
77. 77. The method of claim 76, wherein the identifying step comprises using a nanopore or a nanogap.
78. 78. The method of claim 77, wherein the identifying step comprises transporting the AALC complex or the stacked AALC complex, or a portion thereof, through the nanopore or the nanogap.
79. 79. The method of claim 78, wherein said transporting occurs at a temperature below ambient temperature.
80. 77. The method of claim 76, wherein the step of identifying the AALC complex or the stacked AALC complex comprises determining the amino acid type of the AALC complex or the amino acid type of the stacked AALC complex.
81. 81. The method of any one of claims 52 to 80, wherein the polymerizable molecule comprises a nucleic acid barcode molecule containing temporal information.
82. 82. The method of any one of claims 52 to 81, wherein the capture moiety comprises a cleavable moiety.
83. (a) providing a nucleic acid molecule linked to an amino acid or modified amino acid; (b) sequencing the nucleic acid molecule; (c) identifying the amino acid or the modified amino acid; wherein (b) and (c) are performed within 1 minute of each other.
84. 84. The method of claim 83, wherein (b) comprises generating sequencing reads.
85. 85. The method of any one of claims 83-84, wherein (b) or (c) is performed using a nanopore.
86. 86. The method of claim 85, wherein the nanopore binds a helicase, a topoisomerase, an unfoldase, or a polymerase.
87. 87. The method of any one of claims 83 to 86, wherein (b) and (c) are performed using a nanopore.
88. 88. The method of any one of claims 83 to 87, wherein the nucleic acid molecule is linked to the amino acid or modified amino acid via a linker.
89. 89. The method of claim 88, wherein prior to (a), the linker comprises an amino acid reactive moiety and binds to the nucleic acid molecule.
90. 90. The method of claim 89, wherein prior to (a), the linker comprises an additional reactive moiety.
91. 89. The method of claim 88, wherein the modified amino acids comprise proteinogenic amino acids, the linker, and the nucleic acid molecule.
92. 89. The method of claim 88, wherein the linker comprises at least two atoms.
93. 89. The method of claim 88, wherein the linker comprises a phenylisothiocyanate moiety.
94. 89. The method of claim 88, further comprising, prior to (a), attaching the linker to the amino acid or modified amino acid to produce an amino acid-linker conjugate attached to the nucleic acid molecule.
95. 95. The method of claim 94, further comprising, prior to (b) or (c), forming an amino acid-linker-capture moiety (AALC) conjugate by attaching the amino acid-linker conjugate to a capture moiety.
96. 96. The method of claim 95, wherein the capture moiety comprises a nucleic acid molecule.
97. 96. The method of claim 95, wherein prior to (a), the amino acid or the modified amino acid is attached to a peptide, and the method further comprises cleaving the AALC complex from the peptide.
98. 98. The method of Claim 97, wherein said capture moiety comprises a nucleic acid barcode molecule that identifies said peptide.
99. 98. The method of claim 97, further comprising the steps of providing an additional linker, attaching the additional linker to an additional amino acid of the peptide to form an additional amino acid-linker conjugate, and attaching the additional amino acid-linker conjugate to the AALC conjugate to form a stacked AALC conjugate.
100. 96. The method of claim 95, wherein the capture moiety comprises a cleavable moiety.
101. 100. The method of any one of claims 91 to 99, wherein the nucleic acid molecule comprises temporal information.
102. 84. The method of Claim 83, wherein the identifying step comprises measuring a current signal from the amino acid or the modified amino acid using a nanopore or nanogap, and determining the amino acid type or post-translationally modified variant thereof of the amino acid or the modified amino acid.
103. 103. The method of claim 102, wherein said measuring is performed at a temperature below ambient temperature.
104. 84. The method of Claim 83, wherein the nucleic acid molecule comprises a modified nucleobase and a click chemistry moiety.
105. a linker covalently attached to a nucleic acid barcode molecule, said linker comprising an amino acid reactive group, said nucleic acid barcode molecule comprising temporal information; A composition comprising:
106. 106. The composition of claim 105, wherein the amino acid reactive group comprises an isothiocyanate.
107. 107. The composition of claim 106, wherein the isothiocyanate is PITC.
108. 106. The composition of Claim 105, wherein said nucleic acid barcode molecule comprises a modified nucleobase and a click chemistry moiety.