Methods, kits, and devices for preparing samples for multiplexed polypeptide sequencing.

Polypeptide barcoding and real-time sequencing reactions address the limitations of conventional methods, enabling efficient and accurate multiplexed proteomics analysis for improved diagnostic and therapeutic strategies.

JP2026071233APending Publication Date: 2026-04-28QUANTUM SI INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
QUANTUM SI INC
Filing Date
2026-01-08
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing approaches to multiplexed proteomics analysis are limited, and conventional polypeptide sequencing methods are cumbersome and prone to errors due to the complexity of amino acids and post-translational variations.

Method used

A method involving polypeptide barcoding is employed to prepare multiplexed samples for sequencing, where polypeptides are tagged with barcode components, combined with supplemental samples, and analyzed using real-time degradation reactions to determine amino acid sequences and origins.

Benefits of technology

This approach enhances the efficiency and accuracy of proteomics analysis by enabling simultaneous sequencing of multiple samples, reducing costs and errors, and providing insights into cellular processes for improved diagnostics and therapeutics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026071233000001_ABST
    Figure 2026071233000001_ABST
Patent Text Reader

Abstract

This invention provides a method for preparing samples for polypeptide sequencing, utilizing polypeptide barcodes to facilitate multiplexed proteomics analysis. It also provides useful compositions, kits, and devices for this purpose. [Solution] A method for preparing multiplexed samples for polypeptide sequencing using barcodes. A method for multiplexed samples for polypeptide sequencing in which polypeptide populations are physically separated. A device for preparing samples, comprising a kit containing a population of barcodes, and a sample preparation module containing an immobilized capture probe configured to interact with a cartridge containing barcodes and a reservoir.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] Proteomics has emerged as an important and necessary complement to genomics and transcriptomics in the study of biological systems. However, approaches to multiplexed proteomics analysis have been limited until now. [Overview of the Initiative]

[0002] Provided herein are methods for preparing samples for polypeptide sequencing, which utilize polypeptide barcodes to facilitate multiplexed proteomics analysis. Also provided herein are compositions, kits, and devices useful for this purpose.

[0003] In some embodiments, the disclosure relates to a method for preparing a multiplexed sample. In some embodiments, the method includes (i) contacting a population of polypeptides with a barcode component to produce a sample containing one or more barcoded polypeptides; and (ii) combining the sample from (i) with one or more supplemental samples to produce a multiplexed sample for parallel polypeptide sequencing.

[0004] In some embodiments, (i) provides (a) a group of polypeptides; and (b) contacts the group of polypeptides in (a) with a barcode component comprising a plurality of barcode molecules, wherein contacting the plurality of polypeptides with the barcode component generates a sample comprising one or more barcoded polypeptides.

[0005] In some embodiments, one or more of the supplemental samples in (ii) are (a) provided a population of polypeptides; (b) produced by contacting the population of polypeptides in (a) with a barcode component containing multiple barcode molecules, thereby producing a sample containing one or more barcoded polypeptides.

[0006] In some embodiments, the group of polypeptides in (a) consists of a single polypeptide. In some embodiments, the group of polypeptides in (a) includes polypeptide fragments derived from a single polypeptide. In some embodiments, the group of polypeptides in (a) includes multiple polypeptides.

[0007] In some embodiments, (a) includes lysing a cell population to produce a lysed sample containing multiple polypeptides expressed in the cell population. In some embodiments, the cell population consists of a single cell, multiple homogeneous cells, or multiple heterogeneous cells. In some embodiments, the cell population is isolated from a subject. In some embodiments, the subject is a human, mouse, rat, or non-human primate.

[0008] In some embodiments, (a) further comprises generating a sample containing a modified polypeptide by contacting the dissolved sample with a modifier. In some embodiments, (a) further comprises isolating a fraction of the polypeptide from the lysated sample to produce a high-concentration sample containing a subset of the polypeptide expressed in the cell population. In some embodiments, isolating a fraction of the polypeptide from the lysated sample comprises i. contacting the lysated sample with a plurality of high-concentration molecules, wherein at least a subset of the high-concentration molecules in the plurality of high-concentration molecules binds to a subset of the polypeptide in the lysated sample to produce a bound subset of the polypeptide and an unbound subset of the polypeptide; and ii. isolating a bound subpopulation of the polypeptide or an unbound subpopulation of the polypeptide.

[0009] In some embodiments, each of the multiple concentration molecules is an antibody, aptamer, or enzyme, or the concentration molecules in a subset of the multiple concentration molecules include an antibody, aptamer, or enzyme.

[0010] In some embodiments, each of the multiple concentration molecules is immobilized on the substrate, or the concentration molecules in a subset of the multiple concentration molecules are immobilized on the substrate. In some embodiments, contact between the multiple polypeptides and the multiple concentration molecules occurs when a dissolved sample containing the multiple polypeptides comes into contact with the substrate. In some embodiments, the substrate is selected from the group consisting of surfaces, beads, particles, and gels, and optionally, the surface is a solid surface, the beads are magnetic beads, or the particles are magnetic particles.

[0011] In some embodiments, each of the multiple high-concentration molecules is bound to two or more polypeptides containing different amino acid sequences, or the high-concentration molecules in a subset of the multiple high-concentration molecules are bound to two or more polypeptides containing different amino acid sequences.

[0012] In some embodiments, each of the multiple concentration molecules in a plurality of concentration molecules binds to a post-translational modification of an amino acid, or a subset of the multiple concentration molecules binds to a post-translational modification of an amino acid. In some embodiments, the post-translational modification is selected from the group consisting of acetylation, ADP-ribosylation, caspase cleavage, citrullination, formylation, hydroxylation, methylation, myristoylation, N-linked glycosylation, nedylation, nitration, O-linked glycosylation, oxidation, palmitoylation, phosphorylation, prenylation, S-nitrosylation, sulfation, smoylation, and ubiquitination.

[0013] In some embodiments, the method further comprises generating a sample containing a modified polypeptide by contacting the polypeptide of a highly concentrated sample with a modifier. In some embodiments, the modifier comprises a denaturing agent, and at least one polypeptide is modified by denaturation. In some embodiments, the modifier blocks free carboxylic acid groups, and at least one polypeptide is modified by blocking the free carboxylic acid groups of the polypeptide. In some embodiments, the modifier blocks free thiol groups, and at least one polypeptide is modified by blocking the free thiol groups of the polypeptide. In some embodiments, the modifier comprises a cleaving agent, and at least one polypeptide is modified by cleavage.

[0014] In some embodiments, the barcode component of (i) comprises a barcode molecule containing a polynucleic acid moiety. In some embodiments, the polynucleic acid moiety is 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, or 60 nucleotides long. In some embodiments, (ii) further comprises depositing the multiplexed sample on or within a solid substrate, the solid substrate comprising immobilized detection molecules corresponding to one or more polynucleic acid moieties of a barcode molecule comprising a polynucleic acid moiety, and optionally, the detection molecules comprising polynucleic acids complementary to one or more polynucleic acid moieties of the barcode molecule comprising a polynucleic acid moiety. In some embodiments, the solid substrate is a chip array.

[0015] In some embodiments, the barcode component of (i) comprises a barcode molecule containing a polypeptide moiety. In some embodiments, the polypeptide moiety is 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids long. In some embodiments, the polypeptide moiety is the amino acid sequence of an antibody. In some embodiments, (ii) further comprises depositing a multiplexed sample on or within a solid substrate, the solid substrate containing immobilized antigens corresponding to one or more polypeptide moieties of the barcode molecule containing the amino acid sequence of an antibody. In some embodiments, the solid substrate is a chip array.

[0016] In some embodiments, the barcode component of (i) comprises a barcode molecule including a small molecule portion such as a fluorescent molecular portion. In some embodiments, the fluorescent molecular portion includes aromatic or heteroaromatic compounds, such as pyrene, anthracene, naphthalene, acridine, stilbene, indole, benzindol, oxazole, carbazole, thiazole, benzothiazole, phenantholidine, phenoxazine, porphyrin, quinoline, ethidium, benzamide, cyanine, carbocyanine, salicylate, anthranilate, coumarin, fluorescein, rhodamine, and the like. In some embodiments, the fluorescent molecular portion includes a dye selected from the group consisting of xanthene dyes, naphthalene dyes, coumarin dyes, acridine dyes, cyanine dyes, benzoxazole dyes, stilbene dyes, pyrene dyes, phthalocyanine dyes, phycobiliprotein dyes, squalane dyes, and BODIPY dyes.

[0017] In some embodiments, the sample produced in (i) comprises polypeptides each having a barcode molecule covalently bonded to up to 10 amino acids at its N-terminus or C-terminus.

[0018] In other embodiments, the method comprises (i) providing two or more populations of polypeptides; and (ii) depositing two or more populations of polypeptides of (i) on or within a solid substrate, wherein each population of polypeptides remains physically separated from the other populations of polypeptides in (i), thereby preparing a multiplexed sample for parallel polypeptide sequencing. In some embodiments, the solid substrate is a chip array. In some embodiments, each population of polypeptides is deposited at a different injection port of the solid substrate.

[0019] In some embodiments, at least one of the polypeptide group in (a) consists of a single polypeptide. In some embodiments, at least one of the polypeptide group in (a) includes a polypeptide fragment derived from a single polypeptide. In some embodiments, at least one of the polypeptide group in (a) includes multiple polypeptides.

[0020] In some embodiments, (i) comprises lysing a cell population to produce a lysed sample containing a plurality of polypeptides expressed in the cell population. In some embodiments, the cell population consists of a single cell, a plurality of homogeneous cells, or a plurality of heterogeneous cells. In some embodiments, the cell population is isolated from a subject. In some embodiments, the subject is a human, mouse, rat, or non-human primate. In some embodiments, (i) further comprises contacting each of the lysed samples produced in (c)(b) with a modifier to produce a sample containing a modified polypeptide.

[0021] In some embodiments, (a) further comprises isolating a fraction of polypeptides from a lysed sample to generate a highly concentrated sample containing a subset of polypeptides expressed in a cell population.

[0022] In some embodiments, (c) includes contacting each of the soluble samples produced in i.(b) with a plurality of concentration molecules, wherein at least a subset of the concentration molecules in the plurality of concentration molecules binds to a subset of polypeptides in each soluble sample, thereby generating a bound subset of polypeptides and an unbound subset of polypeptides; and ii. isolating the bound subgroup of polypeptides or the unbound subgroup of polypeptides.

[0023] In some embodiments, each of the multiple high-concentration molecules is immobilized on the substrate, or a subset of the multiple high-concentration molecules is immobilized on the substrate.

[0024] In some embodiments, each of the multiple concentration molecules is immobilized on the substrate, or a subset of the multiple concentration molecules is immobilized on the substrate. In some embodiments, contact between the multiple polypeptides and the multiple concentration molecules occurs when a dissolved sample containing the multiple polypeptides comes into contact with the substrate. In some embodiments, the substrate is selected from the group consisting of surfaces, beads, particles, and gels, and optionally, the surface is a solid surface, the beads are magnetic beads, or the particles are magnetic particles.

[0025] In some embodiments, each of the multiple high-concentration molecules is bound to two or more polypeptides containing different amino acid sequences, or the high-concentration molecules in a subset of the multiple high-concentration molecules are bound to two or more polypeptides containing different amino acid sequences.

[0026] In some embodiments, each of the multiple concentration molecules in a plurality of concentration molecules binds to a post-translational modification of an amino acid, or a subset of the multiple concentration molecules binds to a post-translational modification of an amino acid. In some embodiments, the post-translational modification is selected from the group consisting of acetylation, ADP-ribosylation, caspase cleavage, citrullination, formylation, hydroxylation, methylation, myristoylation, N-linked glycosylation, nedylation, nitration, O-linked glycosylation, oxidation, palmitoylation, phosphorylation, prenylation, S-nitrosylation, sulfation, smoylation, and ubiquitination.

[0027] In some embodiments, (i) further comprises generating a sample containing the modified polypeptide by contacting each polypeptide of the highly concentrated sample produced in (d)(c) with a modifier. In some embodiments, the modifier comprises a denaturing agent, and at least one polypeptide is modified by denaturation. In some embodiments, the modifier blocks free carboxylic acid groups, and at least one polypeptide is modified by blocking the free carboxylic acid groups of the polypeptide. In some embodiments, the modifier blocks free thiol groups, and at least one polypeptide is modified by blocking the free thiol groups of the polypeptide. In some embodiments, the modifier comprises a cleaving agent, and at least one polypeptide is modified by cleavage.

[0028] In some embodiments, the Disclosure relates to a method for determining at least a partial amino acid sequence and origin of polypeptides in a multiplexed sample. In some embodiments, the method comprises (i) preparing a multiplexed sample according to the method herein; (ii) determining the origin of polypeptides in the multiplexed sample by detecting barcode identity of barcoded polypeptides in the multiplexed sample; and (iii) determining at least a partial amino acid sequence of polypeptides in the multiplexed sample by sequencing polypeptides in the multiplexed sample in parallel, where (iii) is performed before, after, or simultaneously with (ii).

[0029] In some embodiments, the barcode identity of a barcoded polypeptide is detected in (ii) by DNA sequencing, polypeptide sequencing, hybridization, luminescence, binding kinetics, and / or physical position on or within a solid substrate.

[0030] In some embodiments, (iii) includes (a) contacting a single polypeptide molecule of a multiplexed sample with one or more terminal amino acid recognition molecules; and (b) sequencing the single polypeptide molecule by detecting a series of signal pulses indicating the association of one or more terminal amino acid recognition molecules with consecutive amino acids exposed at the terminal end of the single polypeptide while the single polypeptide is being degraded.

[0031] In some embodiments, (iii) includes (a) contacting a single polypeptide molecule of a multiplexed sample with a composition comprising one or more terminal amino acid recognition molecules and a cleavage reagent; and (b) detecting a series of signal pulses in the presence of the cleavage reagent that indicate association between one or more terminal amino acid recognition molecules and the ends of the single polypeptide molecule, wherein the series of signal pulses indicates a series of amino acids exposed at the ends over time as a result of terminal amino acid cleavage by the cleavage reagent.

[0032] In some embodiments, (iii) comprises (a) identifying a first amino acid at the terminus of a single polypeptide molecule of the multiplexed sample; (b) removing the first amino acid to expose a second amino acid at the terminus of the single polypeptide molecule; and (c) identifying a second amino acid at the terminus of the single polypeptide molecule, wherein (a) to (c) are carried out in a single reaction mixture.

[0033] In some embodiments, (iii) includes (a) contacting a single polypeptide molecule of a multiplexed sample with one or more amino acid recognition molecules bound to the single polypeptide molecule; (b) detecting a series of signal pulses indicating association between the one or more amino acid recognition molecules and the single polypeptide molecule under polypeptide degradation conditions; and (c) identifying a first type of amino acid in the single polypeptide molecule based on a first characteristic pattern of the series of signal pulses.

[0034] In some embodiments, (iii) includes (a) acquiring data during the polypeptide degradation process; (b) analyzing the data to determine the portion of the data corresponding to amino acids sequentially exposed at the ends of the polypeptide during the degradation process; and (c) outputting an amino acid sequence representing the polypeptide.

[0035] In some embodiments, (iii) includes (a) contacting the polypeptide of a multiplexed sample with one or more label affinity reagents that selectively bind to one or more terminal amino acids at the end of the polypeptide; and (b) identifying the terminal amino acids at the end of the polypeptide by detecting the interaction between the polypeptide and the one or more label affinity reagents.

[0036] In some embodiments, (iii) comprises (a) contacting the polypeptide of a multiplexed sample with one or more label affinity reagents that selectively bind to one or more terminal amino acids at the end of the polypeptide; (b) identifying the terminal amino acids at the end of the polypeptide by detecting the interaction between the polypeptide and the one or more label affinity reagents; (c) removing the terminal amino acids; and (d) determining the amino acid sequence of the polypeptide by repeating (a) to (c) one or more times at the end of the polypeptide. In some embodiments, the method further comprises removing one of the one or more label affinity reagents that do not selectively bind to terminal amino acids after (a) and before (b); and / or removing one of the one or more label affinity reagents that selectively bind to terminal amino acids after (b) and before (c). In some embodiments, (c) includes modifying a terminal amino acid by contacting the terminal amino acid with an isothiocyanate, and: contacting the modified terminal amino acid with a protease that selectively binds to and removes the modified terminal amino acid; or exposing the modified terminal amino acid to conditions sufficiently acidic or basic to remove the modified terminal amino acid.

[0037] In some embodiments, identifying a terminal amino acid includes identifying the terminal amino acid as one of one or more terminal amino acids to which one or more labeling affinity reagents bind; or identifying the terminal amino acid as a type other than one or more terminal amino acids to which one or more labeling affinity reagents bind. In some embodiments, one or more labeling affinity reagents include one or more labeled aptamers, one or more labeled peptidases, one or more labeled antibodies, one or more labeled degradation pathway proteins, one or more aminotransferases, one or more tRNA synthetases, or a combination thereof. In some embodiments, one or more labeled peptidases are modified to inactivate their cleavage activity, or one or more labeled peptidases retain cleavage activity for removal in (c).

[0038] In some embodiments, this disclosure relates to a kit for carrying out the methods described herein. In some embodiments, the kit includes a barcode component comprising multiple barcode molecules. In some embodiments, the barcode component further includes a reaction component comprising one or more reagents for covalently bonding the barcode molecules to a polypeptide. In some embodiments, the barcode component comprises one or more barcode molecules comprising a polynucleic acid moiety, a polypeptide moiety, and / or a fluorescent molecule moiety.

[0039] In some embodiments, the polynucleic acid portion is 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, or 60 nucleotides long. In some embodiments, the polynucleic acid portion includes an aptamer.

[0040] In some embodiments, the polypeptide portion is 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids long. In some embodiments, the polypeptide portion is an antibody or aptamer.

[0041] In some embodiments, the fluorescent molecular portion includes aromatic or heteroaromatic compounds, such as pyrene, anthracene, naphthalene, acridine, stilbene, indole, benzindol, oxazole, carbazole, thiazole, benzothiazole, phenantholidine, phenoxazine, porphyrin, quinoline, ethidium, benzamide, cyanine, carbocyanine, salicylate, anthranilate, coumarin, fluorescein, rhodamine, and the like. In some embodiments, the fluorescent molecular portion includes dyes selected from the group consisting of xanthene dyes, naphthalene dyes, coumarin dyes, acridine dyes, cyanine dyes, benzoxazole dyes, stilbene dyes, pyrene dyes, phthalocyanine dyes, phycobiliprotein dyes, squalane dyes, and BODIPY dyes.

[0042] In some embodiments, the kit further comprises a solid support. In some embodiments, the solid support comprises an immobilized detection molecule (or a plurality of immobilized detection molecules). In some embodiments, the detection molecule comprises a polynucleic acid moiety corresponding to the barcode molecule of the barcode component. In some embodiments, the detection molecule comprises a polypeptide moiety corresponding to the barcode molecule of the barcode component.

[0043] In some embodiments, the kit includes a solid support that allows for the physical separation of a population of polypeptides of different origins. In some embodiments, this disclosure relates to a device for performing the method described herein.

[0044] In some embodiments, the device includes at least one hardware processor; and at least one non-temporary computer-readable storage medium that stores processor-executable instructions, which, when executed by the at least one hardware processor, cause the at least one hardware processor to perform the method described herein.

[0045] In some embodiments, the device includes at least one non-temporary computer-readable storage medium that stores processor-executable instructions, which, when executed by at least one hardware processor, cause the at least one hardware processor to perform the method described herein.

[0046] In some embodiments, the device comprises (i) a sample preparation module configured to interface with one or more cartridges, each cartridge comprising (a) one or more reservoirs or reaction vessels configured to accept complex samples; (b) one or more sequence sample preparation reagents comprising a plurality of barcode molecules; and (c) a matrix comprising one or more immobilized capture probes; and (ii) a sequencing module comprising an array of pixels, each pixel configured to receive a sequencing sample from the sample preparation module, and comprising (a) a sample well; and (b) at least one photodetector.

[0047] In some embodiments, the sample preparation reagent further comprises a plurality of concentration molecules. In some embodiments, at least one subset of the concentration molecules in the plurality of concentration molecules is covalently bound to an immobilized capture probe. In some embodiments, at least one subset of the concentration molecules is covalently bound to beads or particles that can be bound by the immobilized capture probe. In some embodiments, each of the concentration molecules in the plurality of concentration molecules comprises an antibody, aptamer, or enzyme. In some embodiments, the concentration molecules in a subset of the plurality of concentration molecules comprise an antibody, aptamer, or enzyme.

[0048] In some embodiments, the sample preparation reagent includes a modifier. In some embodiments, the modifier mediates polypeptide fragmentation, polypeptide modification, addition of post-translational modifications, and / or blocking of one or more functional groups.

[0049] In some embodiments, the sequencing module further includes a reservoir or reaction vessel configured to deliver sequencing reagents to the sample wells of each pixel. In some embodiments, the sequencing reagent includes a label affinity reagent. In some embodiments, the label affinity reagent includes one or more labeled aptamers, one or more labeled peptidases, one or more labeled antibodies, one or more labeled degradation pathway proteins, one or more aminotransferases, one or more tRNA synthetases, or a combination thereof. [Brief explanation of the drawing]

[0050] Those skilled in the art will understand that the drawings provided herein are for illustrative purposes only. It should be understood that in some cases, various aspects of the invention may be exaggerated or enlarged to facilitate understanding of the invention. Generally, similar reference letters in the drawings refer to similar features, functional similarities, and / or structural similarities across the various drawings. The drawings are not necessarily to scale and are rather focused on illustrating the principles of the teachings. The drawings are not intended to limit the scope of these teachings in any way.

[0051] The features and advantages of the present invention will become more apparent from the detailed description below, when interpreted in conjunction with the drawings. When describing embodiments with reference to drawings, directional references (such as "up," "down," "top," "bottom," "left," "right," "horizontal," and "vertical") may be used. Such references are intended solely as an aid to readers viewing the drawings in their normal orientation. These directional references are not intended to describe a preferred or sole orientation of the embodied invention. The invention may be embodied in other orientations as well.

[0052] As will be apparent from the detailed description, the examples shown in the figures and further described throughout this application for illustrative purposes illustrate non-limiting embodiments and, in some cases, may simplify certain processes or omit features or steps for the purpose of clearer explanation. [Figure 1] The diagram provides illustrative images of two samples before (left) and after (right) barcode labeling. The barcode molecule in the first sample can be distinguished from the barcode molecule in the second sample. [Figure 2] This provides an exemplary embodiment of a workflow following protein barcoding. Barcoded samples are pooled into multiplexed samples (1). The sequences and barcode identities (i.e., sample origins) of polypeptides in the multiplexed samples are then determined / identified (simultaneously or sequentially) (2). Finally, the sequences are grouped based on their barcode identities (i.e., sample origins) (3). [Figure 3-1] Figures 3A–3E provide exemplary barcode molecules and methods for detecting exemplary barcodes. Figure 3A. A barcode molecule may include a polynucleic acid moiety ("DNA barcode"), which is identified via hybridization using a detection molecule containing the polynucleic acid moiety (which may also include a luminescent molecule). Figure 3B. A barcode molecule may include a polynucleic acid moiety identified by DNA sequencing. Figure 3C. A barcode molecule may include a polypeptide moiety (e.g., a short polypeptide tag) identified by polypeptide sequencing. Figure 3D. A barcode molecule may include a chemically modified (e.g., tyrosine phosphorylated) sample identified by chemical modification by polypeptide sequencing. Figure 3E. A barcode molecule may include a polypeptide moiety (e.g., an antibody; here "antibody A" or "antibody B") identified by localization on a chip (e.g., via binding with a detection molecule; here "antigen A" or "antigen B"). [Figure 3-2] Same as above. [Figure 4]This provides an exemplary embodiment of barcoding by physical separation. The chip can be physically separated and optionally contain other barcode molecules as needed. [Figure 5] A diagram is provided illustrating an exemplary workflow for preparing multiplexed samples for polypeptide sequencing. [Figure 6] A diagram is provided illustrating an exemplary workflow for preparing multiplexed samples for polypeptide sequencing. [Figure 7] A diagram is provided illustrating an exemplary workflow for preparing high-concentration samples. [Figure 8] A diagram is provided illustrating an exemplary workflow for preparing high-concentration samples. [Figure 9] A diagram is provided illustrating an exemplary workflow for preparing high-concentration samples. [Figure 10] The present invention provides a diagram illustrating an exemplary apparatus for preparing highly concentrated and / or multiplexed samples. [Modes for carrying out the invention]

[0053] Detailed explanation As described herein, the inventors have recognized that differential binding interactions can provide an additional or alternative approach to conventional labeling strategies in polypeptide sequencing. Conventional polypeptide sequencing may involve labeling each type of amino acid with a specifically identifiable label. This process is cumbersome and prone to errors, given that there are at least 20 naturally occurring amino acids, in addition to numerous post-translational variations. In some embodiments, the disclosure relates to the discovery of techniques involving the use of amino acid recognition molecules that bind to different types of amino acids in different ways to generate a detectable, characteristic signature indicating the amino acid sequence of a polypeptide.

[0054] In some embodiments, the disclosure relates to the discovery that polypeptide sequencing reactions can be monitored in real time using only a single reaction mixture (e.g., without requiring repeating reagents circulating through a reaction vessel). Conventional polypeptide sequencing reactions may involve exposing the polypeptide to different reagent mixtures to cycle between amino acid detection and amino acid cleavage steps. Thus, in some embodiments, the disclosure relates to advances in next-generation sequencing that enable the analysis of polypeptides by amino acid detection through ongoing degradation reactions in real time.

[0055] Proteomics analysis of individual organisms can provide insights into cellular processes and response patterns, leading to improved diagnostic and therapeutic strategies. The ability to sequence multiple samples simultaneously (i.e., multiplexed) increases the efficiency of proteomics analysis of individual samples and reduces associated costs. Accordingly, in some embodiments, this disclosure relates to a method for preparing multiplexed samples for polypeptide sequencing, which leverages polypeptide barcoding to facilitate multiplexed proteomics analysis.

[0056] In some embodiments, the present disclosure relates to a method for preparing multiplexed samples for polypeptide sequencing. In some embodiments, the method includes (i) providing multiple samples (e.g., from different subjects / patients); (ii) tagging the polypeptides in each sample with different barcodes; and (iii) combining the tagged polypeptides to generate a single multiplexed sample for polypeptide sequencing.

[0057] In some embodiments, the Disclosure relates to a method for determining at least a partial amino acid sequence and origin of polypeptides in a multiplexed sample, the method comprising (i) preparing a multiplexed sample containing barcoded polypeptides; (ii) detecting barcode identity of barcoded polypeptides in the multiplexed sample; (iii) sequencing the polypeptides in the multiplexed sample in parallel, where (iii) is performed before, after, or simultaneously with (ii). The barcode detected in (ii) may be used to extract sample-specific sequence information from the multiplexed data.

[0058] Useful compositions, kits, and devices for this purpose are also provided herein. I. Methods for preparing complex samples In some embodiments, this disclosure relates to methods for preparing complex samples (e.g., complex polypeptide samples). As used herein, the term “complex sample” means a sample comprising multiple molecules (e.g., polypeptides, polynucleic acids, metabolites, etc.) of which at least two are chemically distinctive. In some embodiments, the complex sample comprises multiple polypeptides, the multiple comprising at least two polypeptides having different amino acid sequences.

[0059] Typically, complex samples originate from (or are generated by) a population of cells. In some embodiments, the population of cells consists of a single cell. In other embodiments, the population of cells includes two or more cells.

[0060] For example, in some embodiments, the cell population is at least 5, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1×10 3 cells, at least 1×10 4 cells, at least 1×10 5 cells, at least 1×10 6 cells, at least 1×10 7 cells, at least 1×10 8 cells, at least 1×10 9 cells, or at least 1×10 10 cells.

[0061] In some embodiments, the population is 1 - 5, 1 - 10, 1 - 20, 1 - 30, 1 - 50, 1 - 60, 1 - 70, 1 - 80, 1 - 90, 1 - 100, 1 - 150, 1 - 200, 1 - 250, 1 - 300, 1 - 350, 1 - 400, 1 - 450, 1 - 500, 1 - 600, 1 - 700, 1 - 800, 1 - 900, 1 - 1×10 3 cells, 1 - 1×10 4 cells, 1 - 1×10 5 cells, 1 - 1×10 6 cells, 1 - 1×10 7 cells, 1 - 1×10 8 cells, 1 - 1×10 9 cells, 1 - 1×10 10 cells, 100 - 150, 100 - 200, 100 - 250, 100 - 300, 100 - 350, 100 - 400, 100 - 450, 100 - 500, 100 - 600, 100 - 700, 100 - 800, 100 - 900, 100 - 1×10 3 cells, 100 - 1×10 4 cells, 100 - 1×10 5 cells, 100 - 1×106 pieces, 100~1×10 7 pieces, 100~1×10 8 pieces, 100~1×10 9 pieces, 100~1×10 10 pieces, 1×10 3 ~1 × 10 4 pieces, 1×10 3 ~1 × 10 5 pieces, 1×10 3 ~1 × 10 6 pieces, 1×10 3 ~1 × 10 7 pieces, 1×10 3 ~1 × 10 8 pieces, 1×10 3 ~1 × 10 9 pieces, 1×10 3 ~1 × 10 10 pieces, 1×10 4 ~1 × 10 5 pieces, 1×10 4 ~1 × 10 6 pieces, 1×10 4 ~1 × 10 7 pieces, 1×10 4 ~1 × 10 8 pieces, 1×10 4 ~1 × 10 9 pieces, 1×10 4 ~1 × 10 10 pieces, 1×10 5 ~1 × 10 6 pieces, 1×10 5 ~1 × 10 7 pieces, 1×10 5 ~1 × 10 8 pieces, 1×10 5 ~1 × 10 9 pieces, or 1 x 10 5 ~1 × 10 10 Contains individual cells.

[0062] A population of cells may include prokaryotic cells and / or eukaryotic cells. A population of cells may include multiple allocellular groups. Alternatively, a population of cells may include multiple heterocellular groups. A population of cells can be isolated from a subject (e.g., a multicellular organism or a symbiotic organism). In some embodiments, the subject is a mouse, rat, rabbit, guinea pig, hamster, pig, sheep, dog, primate, cat, or human.

[0063] Methods for isolating cell populations are known to those skilled in the art. For example, methods for preparing complex samples may include biopsy, dissection (e.g., microdissection such as laser capture), limited dilution, micromanipulation, immunomagnetic cell separation, fluorescence-activated cell sorting, density gradient centrifugation, immunodensity cell separation, microfluidic cell sorting, sedimentation, adhesion, or combinations thereof.

[0064] In some embodiments, a method for preparing a complex sample involves lysing a population of cells to produce a lysed sample containing multiple molecules (e.g., polypeptides, polynucleic acids, metabolites, etc.). Methods for lysing a population of cells are known to those skilled in the art. In some embodiments, a sample containing cells is lysed using either a known physical or chemical methodology to release the target molecules from the cells. In some embodiments, the sample may be lysed using electrolysis, enzymatic methods, surfactant-based methods, and / or mechanical homogenization. In some embodiments, if the sample contains neither cells nor tissue (e.g., a sample containing purified polypeptides), the lysis step may be omitted.

[0065] Alternatively, or in addition to the above, a method for preparing a complex sample may involve the isolation of one or more intracellular compartments, such as endosomes, synaptosomes, cytoplasm, nucleoplasm, chromatin, mitochondria, peroxisomes, lysosomes, melanosomes, exosomes, Golgi apparatus, endoplasmic reticulum, centrosomes, pseudopods, or combinations thereof.

[0066] Molecules originating from the same cell population are described herein as having the same “origin.” II. Method for preparing multiplexed samples In some embodiments, this disclosure relates to methods for preparing multiplexed samples. As used herein, the term “multiplexed sample” means a sample comprising at least two subsamples having different origins (e.g., two or more samples, each prepared from a different cell population or multiple molecules).

[0067] In some embodiments, the multiplexed sample includes at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, or at least 1000 subsamples, each having a different origin.

[0068] In some embodiments, the multiplexed samples are 2-3, 2-4, 2-5, 2-6, 2-7, 2-8, 2-9, 2-10, 2-11, 2-12, 2-13, 2-14, 2-15, 2-16, 2-17, 2-18, 2-19, 2-20, 2-25, 2-30, 2-35, 2-40, 2-45, 2-50, 2-60, 2-70, 2-80, 2-90, 2-100, 2-200, 2-300, 2-400, 2-500, 2-600, each having a different origin. ~700, 2~800, 2~900, 2~1000, 5~10, 5~15, 5~20, 5~25, 5~30, 5~35, 5~40, 5~45, 5~50, 5~60, 5~70, 5~80, 5~90, 5~100, 5~200, 5~300, 5~400, 5~500, 5~600, 5~700, 5~800, 5~900, 10~15, 10~20, 10~25, 10~30, 10~35, 10~40, 10~45, 10~50, 10~60, 10~70, 10~8 0, 10-90, 10-100, 10-200, 10-300, 10-400, 10-500, 10-600, 10-700, 10-800, 10-900, 10-1000, 20-30, 20-40, 20-50, 20-60, 20-70, 20-80, 20-90, 20-100, 20-200, 20-300, 20-400, 20-500, 20-600, 20-700, 20-800, 20-900, 20-1000, 50-60, 50-70, 50- Includes subsamples of 80, 50-90, 50-100, 50-200, 50-300, 50-400, 50-500, 50-600, 50-700, 50-800, 50-900, 50-1000, 100-200, 100-300, 100-400, 100-500, 100-600, 100-700, 100-800, 100-900, 100-1000, 500-600, 500-700, 1500-800, 500-900, or 500-1000.

[0069] In some embodiments, the multiplexed sample includes 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 subsamples, each having a different origin.

[0070] Each subsample of a multiplexed sample may contain multiple molecules. In some embodiments, one or more subsamples of a multiplexed sample contain molecules (e.g., polypeptides) of a complex sample prepared from a cell population (which may be single cells) (see "Method for preparing a complex sample") or molecules (e.g., polypeptides) of a high-concentration sample (see "Method for preparing a high-concentration sample"). In some embodiments, the multiple molecules of a subsample originate from a single molecule (e.g., via fragmentation of a single polypeptide).

[0071] Each subsample of a multiplexed sample may contain a single molecule (e.g., a single polypeptide). In some embodiments, one or more subsamples of a multiplexed sample contain a single molecule (e.g., a single polypeptide).

[0072] Typically, at least a subset of molecules in each subsample of a multiplexed sample can be distinguished from molecules in other subsamples of the multiplexed sample. For example, in some embodiments, at least a subset of polypeptides in each subsample of a multiplexed sample can be distinguished from polypeptides in other subsamples of the multiplexed sample. In this way, the origin of at least a subset of molecules in the multiplexed sample can be identified.

[0073] Therefore, in some embodiments, at least one of the subsamples of the multiplexed sample contains a barcode molecule, and each barcode molecule contains a barcode specific to the subsample (i.e., a unique barcode). If no barcode is found in the molecules of other subsamples of the multiplexed sample, that barcode is considered specific to the subsample.

[0074] In some embodiments, two or more subsamples of the multiplexed sample contain barcoded molecules. In some embodiments, each subsample of the multiplexed sample contains a barcoded molecule. In some embodiments, all but one subsample in the multiplexed sample contain a barcoded molecule.

[0075] Within a multiplexed sample, each barcoded molecule in each subsample containing a barcoded molecule (i.e., each “labeled subsample”) contains a unique barcode. In some embodiments, each barcoded molecule in a labeled subsample contains the same barcode. In some embodiments, the barcoded molecules in a labeled subsample contain a unique combination of barcodes. For example, in some embodiments, a labeled subsample contains a unique combination of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 barcoded molecules.

[0076] In some embodiments, the labeled subsample comprises a barcoded polypeptide and: a barcoded DNA molecule, a barcoded RNA molecule, a barcoded cDNA molecule, a barcoded metabolite, or a combination thereof, wherein the barcoded polypeptide comprises a first barcode (or a first combination of barcodes), the barcoded DNA molecule comprises a second barcode (or a second combination of barcodes), the barcoded RNA molecule in the subsample comprises a third barcode (or a third combination of barcodes), the barcoded cDNA molecule comprises a fourth barcode (or a fourth combination of barcodes), and the barcoded metabolite comprises a fifth barcode (or a fifth combination of barcodes), or a combination thereof.

[0077] In some embodiments, a method for preparing a multiplexed sample includes (i) contacting a population of cells with a barcode component to generate a sample containing a barcoded molecule (e.g., a barcoded polypeptide) (i.e., a first labeled subsample); and (ii) combining the sample from (i) with one or more supplemental samples (i.e., one or more additional subsamples) to generate a multiplexed sample for parallel molecular sequencing (e.g., polypeptide sequencing).

[0078] In some embodiments, a method for preparing a multiplexed sample includes (i) contacting multiple molecules with a barcode component to generate a sample containing barcoded molecules (e.g., barcoded polypeptides) (i.e., a first labeled subsample); and (ii) combining the sample from (i) with one or more supplemental samples (i.e., one or more additional subsamples) to generate a multiplexed sample for parallel molecular sequencing (e.g., polypeptide sequencing).

[0079] In some embodiments described in the preceding two paragraphs, step (ii) further includes depositing the multiplexed sample on or within a solid substrate. In some embodiments, the solid substrate comprises a plurality of immobilized (e.g., covalently bonded) detection molecules, one or more of which interact with the barcodes of the barcode molecules in the multiplexed sample. In some embodiments, the solid substrate is a chip array.

[0080] In some embodiments, a method for preparing a multiplexed sample comprises (i) providing at least two populations of molecules (e.g., polypeptides); (ii) depositing the at least two populations of molecules from (i) on or within a solid substrate, such that each population of molecules remains physically separated from the other population of molecules from (i); thereby preparing a multiplexed sample for parallel polypeptide sequencing.

[0081] A. Method for attaching a polypeptide barcode In some embodiments, this disclosure relates to methods for barcoding molecules in a sample (e.g., polypeptides, DNA, RNA, cDNA, metabolites, etc.). In some embodiments, the sample includes living cells. In some embodiments, the sample is a complex sample prepared from a population of cells (which may be a single cell) (see “Method for preparing a complex sample”). In some embodiments, the sample is a highly concentrated sample (see “Method for preparing a highly concentrated sample”). In some embodiments, the sample includes a single molecule (e.g., a polypeptide) or a fragment derived from a single molecule (e.g., a polypeptide fragment).

[0082] Of particular interest here is the present disclosure relating to a method for barcoding polypeptides. Polypeptides can be barcoded by chemical modification and / or physical separation.

[0083] (i) Chemical modification Polypeptides (or multiple polypeptides) can be barcoded by chemical modification. Chemical modification of a polypeptide alters its chemical composition and can occur during polypeptide synthesis (in vivo or in vitro) or after polypeptide synthesis (i.e., post-translational). Polypeptides can be modified at any position in their amino acid sequence. Methods for generating polypeptide conjugates (to reach barcoded polypeptides) have been previously described and are known to those skilled in the art. For example, Corey et al,Science,1987;238:1401-1403;Kukolka et al,Org.Biomol.Chem.,2004;2:2203-2206;Debets et al,Chem.Commun.,2010;46:97-99;Takeda et al. al.Bioorg.Med.Chem.Lett.,2004;14:2407-2410;Yang et al,Bioconjug.Chem.,2015;26:1381-1395;Rosen et al.,Nat.Chem.,2014;6:804-809;Cong et al. al.,Bioconjug.Chem.,2012;23:248-263;Mattson,G.,et al.Molecular Biology See Reports, 1993;17:167-183.

[0084] In some embodiments, a polypeptide (or a group of polypeptides) is barcoded by a method comprising contacting a population of cells with a barcode component to produce a sample containing the barcoded polypeptide. In such cases, the polypeptide (or a group of polypeptides) may be modified during or after synthesis (i.e., post-translation).

[0085] In some embodiments, a polypeptide (or plurality of polypeptides) is barcoded by a method comprising contacting the polypeptide (or plurality of polypeptides) with a barcode component to produce a sample containing the barcoded polypeptide. In such cases, the polypeptide (or plurality of polypeptides) will be modified after synthesis (i.e., after translation).

[0086] The barcode component may include a modifier. The modifier may include an endoprotease having a distinct cleavage pattern. Examples of endoproteases, known to those skilled in the art, include, but are not limited to, trypsin, chymotrypsin, elastase, thermolysin, pepsin, glutamyl endopeptidase, neprilysin, Lys-C, Arg-C, Asp-N, Lys-N, Glu-C, WaLP, and MaLP. See, for example, Giansanti et al., Nat. Protoc., 28 April 2016; 11(5):993-1006. The polypeptide modifier may include an enzyme capable of modifying the polypeptide in post-translational modification. Examples of post-translational modifications are known to those skilled in the art and are not limited to, but include acetylation, adenylylation, ADP-ribosylation, alkylation (e.g., methylation), amidation, arginylation, biotinylation, butyrylation, carbamylation, carbonylation, carboxylation, citrullination, deamidation, eliminylation, formylation, glycosylation (e.g., N-linked glycosylation, O-linked glycosylation), glipyatyon, glycation, hydroxylation, iodation, ISG formation, isoprenylation, lipoylation, malonylation, myristoylation, nephrylation, nitration, oxidation, and palmitoylation. These include pegylation, phosphorylation, phosphopantetheinylation, polyglycylation, polyglutamylation, prenylation, propionation, pupyrulation, S-glutathioneation, S-nitrosylation, S-sulfenylation, S-sulfinylation, S-sulfonylation, succinylation, sulfation, SUMOylation, and ubiquitination. The enzymes involved in modifying polypeptides by these methods are also known to those skilled in the art.

[0087] Alternatively, or in addition to the above, the barcode component may comprise multiple barcode molecules. In some embodiments, the barcode component consists of multiple barcode molecules. In some embodiments, the barcode component may further comprise one or more reagents (e.g., enzymes, compounds, small molecules, buffers, etc.) to facilitate the covalent bonding of the barcode molecules to the polypeptide. The barcode molecules may be covalently bonded to the polypeptide at any position. In some embodiments, the barcode molecules are covalently bonded to the polypeptide at an amino acid position within 10, 9, 8, 7, 6, 5, 4, 3, or 2 amino acids from its end (N-terminus or C-terminus). In some embodiments, the barcode molecules are covalently bonded to the polypeptide at its N-terminus. In some embodiments, the barcode is covalently bonded to the polypeptide at its C-terminus.

[0088] In some embodiments, each barcode molecule in the barcode component is chemically identical. In some embodiments, the barcode component comprises two or more chemically distinct barcode molecules. For example, the barcode component may comprise 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 chemically distinct barcode molecules.

[0089] The barcode molecule of the barcode component may be a non-natural amino acid (i.e., a non-standard amino acid). Examples of non-natural amino acids are known to those skilled in the art and include, but are not limited to, homoallylglycine (Hag), homopropargylglycine (Hpg), azidohomoalanine (Aha), azidonorleucine (Anl), azidophenylalanine (Azf), acetylphenylalanine (Acf), and propargyloxyphenylalanine (Pxf). In some embodiments in which the barcode component comprises a non-natural amino acid barcode molecule, the barcode component further comprises one or more non-natural tRNAs (or nucleic acids encoding an expressible form of a non-natural tRNA). Examples of non-natural tRNAs are known to those skilled in the art.

[0090] Alternatively, or in addition thereto, the barcode molecule of the barcode component may include a polynucleic acid moiety, a polypeptide moiety, a small molecule moiety, a linker (e.g., a PEG-like linker), a dendrimer, a scaffold, or a combination thereof. In some embodiments, the barcode molecule of the barcode component includes a polynucleic acid moiety, a polypeptide moiety, a small molecule moiety, a linker (e.g., a PEG-like linker), a dendrimer, a scaffold, or a combination thereof.

[0091] In some embodiments, the barcode molecule includes a polynucleic acid moiety. In some embodiments, the barcode molecule includes two or more polynucleic acid moieties. In embodiments in which the barcode molecule includes multiple polynucleic acid moieties, each polynucleic acid moiety may be identical, a subset of polynucleic acid moieties may be identical, or each polynucleic acid moiety may be chemically different.

[0092] In some embodiments, the polynucleic acid portion is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, or 60 nucleotides long.

[0093] In some embodiments, the polynucleotide portion has a length of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, or at least 500 nucleotides.

[0094] When you introduce some of them, the polynucleotide portion will be 5-10, 5-15, 5-20, 5-25, 5-30, 5-40, 5-50, 5-60, 5-70, 5-80, 5-90, 5-100, 5-150, 5-200, 5-250, 5-300, 5-350, 5-400, 5-450, 5-500, 10-15, 10-20, 10-25, 10-30, 10-40, 10-50, 10-60, 10-70, 10-80, 10-90, 10-100, 10-150, 10-200, 10-250, 10-300, 10-350, 10-400, 10-450, 1 The subnet lengths are 0-500, 20-30, 20-40, 20-50, 20-60, 20-70, 20-80, 20-90, 20-100, 20-150, 20-200, 20-250, 20-300, 20-350, 20-400, 20-450, 20-500, 50-75, 50-100, 50-150, 50-200, 50-250, 50-500, 50-350, 50-400, 50-450, 50-500, 100-200, 100-250, 100-500, 100-350, 100-400, 100-450, or 100-500.

[0095] In some embodiments, the polynucleic acid portion is an aptamer. In some embodiments, the barcode molecule includes a polypeptide moiety. In some embodiments, the barcode molecule includes two or more polypeptide moieties. In embodiments in which the barcode molecule includes multiple polypeptide moieties, each polypeptide moiety may be identical. A subset of polypeptide moieties may be identical. Alternatively, each polypeptide moiety may be chemically different.

[0096] In some embodiments, the polypeptide moiety is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acid lengths. In some embodiments, the polypeptide moiety is at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, or at least 500 amino acid lengths. When you start with a few options, the polypeptide portion will be 5-10, 5-15, 5-20, 5-25, 5-30, 5-40, 5-50, 5-60, 5-70, 5-80, 5-90, 5-100, 5-150, 5-200, 5-250, 5-300, 5-350, 5-400, 5-450, 5-500, 10-15, 10-20, 10-250, 10-60, 10-70, 10-80, 10-90, 10-100, 10-150, 10-200, 10-250, 10-300, 10-350, 10-400, 10-450 The amino acid lengths are 10-500, 20-30, 20-40, 20-50, 20-60, 20-70, 20-80, 20-90, 20-100, 20-150, 20-200, 20-250, 20-300, 20-350, 20-400, 20-450, 20-500, 50-75, 50-100, 50-150, 50-200, 50-250, 50-500, 50-350, 50-400, 50-450, 50-500, 100-200, 100-250, 100-500, 100-350, 100-400, 100-450, or 100-500.

[0097] In some embodiments, the polypeptide portion is an aptamer. In some embodiments, the polypeptide portion is an antibody. In some embodiments, the polypeptide portion is an antigen.

[0098] In some embodiments, the barcode molecule includes small molecule portions. In some embodiments, the barcode molecule includes two or more small molecule portions. In embodiments in which the barcode molecule includes multiple small molecule portions, each small molecule portion may be identical, a subset of small molecule portions may be identical, or each small molecule portion may be chemically different.

[0099] In some embodiments, the small molecule portion contains biotin. In some embodiments, the small molecule portion comprises a drug or a luminescent molecule (or fluorescent molecule). Examples of drugs and luminescent molecules suitable for the methods described herein are known to those skilled in the art. As used herein, a luminescent molecule is a molecule capable of absorbing one or more photons and then emitting one or more photons after one or more time intervals.

[0100] In some embodiments, the luminescent molecule may comprise first and second chromophores. In some embodiments, the excited state of the first chromophore can be relaxed via energy transfer to the second chromophore. In some embodiments, the energy transfer is Forster resonance energy transfer (FRET). Such FRET pairs may be useful in providing luminescent labels having properties that facilitate the distinction of a label from among multiple luminescent labels in a mixture. In yet another embodiment, the FRET pair comprises a first chromophore of the first luminescent label and a second chromophore of the second luminescent label. In certain embodiments, the FRET pair may absorb excitation energy in a first spectral range and emit light in a second spectral range.

[0101] In some embodiments, the luminescent molecule refers to a fluorophore or dye. Typically, the luminescent molecule includes aromatic or heteroaromatic compounds and may be pyrene, anthracene, naphthalene, naphthylamine, acridine, stilbene, indole, benzindol, oxazole, carbazole, thiazole, benzothiazole, benzoxazole, phenanthidine, phenoxazine, porphyrin, quinoline, ethidium, benzamide, cyanine, carbocyanine, salicylate, anthranilate, coumarin, fluorescein, rhodamine, xanthene, or other similar compounds.

[0102] In some embodiments, the luminescent molecule comprises a dye selected from one or more of the following: 5 / 6-carboxyrhodamine 6G, 5-carboxyrhodamine 6G, 6-carboxyrhodamine 6G, 6-TAMRA, Abberior® STAR440SXP, Abberior® STAR470SXP, Abberior® STAR488, Abberior® STAR512, Abberior® STAR520SXP, Abberior® STAR580, Abberior® STAR600, Abberior® STAR635, Abberior® STAR635P, Abberior® STAR RED, AlexaFluor(TM) 350, AlexaFluor(TM) 405, AlexaFluor(TM) 430, AlexaFluor(TM) 480, AlexaFluor(TM) 488, AlexaFluor(TM) 514, AlexaFluor(TM) 532, Alexa Fluor(TM) 546, AlexaFluor(TM) 555, AlexaFluor(TM) 568, AlexaFluor(TM) 594, AlexaFluor(TM) 610-X, AlexaFluor(TM) 633, AlexaFluor(TM) 647, AlexaFluor(TM) )660, AlexaFluor(TM) 680, AlexaFluor(TM) 700, AlexaFluor(TM) 750, AlexaFluor(TM) 790, AMCA, ATTO390, ATTO425, ATTO465, ATTO488, ATTO495, ATTO514, ATTO5 20, ATTO532, ATTO542, ATTO550, ATTO565, ATTO590, ATTO610, ATTO620, ATTO633, ATTO647, ATTO647N, ATTO655, ATTO665, ATTO680, ATTO700, ATTO725, ATTO740, ATTO Oxa12, ATTORho101, ATTORho11, ATTORho12, ATTORho13, ATTORho14, ATTORho3B, ATTORho6G, ATTOThio12, BD Horizon(TM) V450, BODIPY(TM) 493 / 501, BODIPY(TM) 530 / 550,BODIPY (trademark) 558 / 568, BODIPY (trademark) 564 / 570, BODIPY (trademark) 576 / 589, BODIPY (trademark) 581 / 591, BODIPY (trademark) 630 / 650, BODIPY (trademark) 650 / 665, BODIPY (trademark) FL, BODIPY (trademark) FL-X, BODIPY (trademark) R6G, BODIPY (trademark) TMR, BODIPY (trademark) TR, CAL Fluor (trademark) Gold540, CAL Fluor (trademark) Green510, CAL Fluor (trademark) Orange560, CAL Fluor (trademark) Red590, CAL Fluor (trademark) Red610, CAL Fluor (trademark) Red615, CAL Fluor (trademark) Red635, Cascade (trademark) Blue, CF (trademark) 350, CF (trademark) 405M, CF (trademark) 405S, CF (trademark) 488A, CF (trademark) 514, CF (trademark) 532, CF (trademark) 543, CF (trademark) 546, CF (trademark) 555, CF (trademark) 568, CF (trademark) 594, CF (trademark) 620R, CF (trademark) 633, CF (trademark) 633-V1, CF (trademark) 640R, CF (trademark) 640R-V1, CF (trademark) 640R-V2, CF (trademark) 660C, CF (trademark) 660R, CF (trademark) 680, CF (trademark) 680R, CF (trademark) 680R-V1, CF (trademark) 750, CF (trademark) 770, CF (trademark) 790, Chromeo (trademark) 642, Chromis 425N, Chromis 500N, C Chromis 515N, Chromis 530N, Chromis 550A, Chromis 550C, Chromis 550Z, Chromis 560N, Chromis 570N, Chromis 577N, Chromis 600N, Chromis 630N, Chromis 645A, Chromis 645C, Chromis 645Z, Chromis 678A, Chromis 678C, Chromis 678Z, Chromis 770A, Chromis 770C, Chromis 800A, Chromis 800C, Chromis 830A, Chromis 830C, Cy(trademark) 3, Cy(trademark) 3.5, Cy(trademark) 3B, Cy(trademark) 5, Cy(trademark) 5.5, Cy(trademark) 7, DyLight(trademark) 350, DyLight(trademark) 405.DyLight (trademark) 415-Col, DyLight (trademark) 425Q, DyLight (trademark) 485-LS, DyLight (trademark) 488, DyLight (trademark) 504Q, DyLight (trademark) 510-LS, DyLight (trademark) 515-LS, DyLight (trademark) 521-LS, DyLight (trademark) 530-R2, DyLight (trademark) 543Q, DyLight (trademark) 550, DyLight (trademark) 554-R0, DyLight (trademark) 554-R1, DyLight (trademark) 590-R2, DyLight ( Trademarks) 594, DyLight 610-B1, DyLight 615-B2, DyLight 633, DyLight 633-B1, DyLight 633-B2, DyLight 650, DyLight 655-B1, DyLight 655-B2, DyLight 655-B3, DyLight 655-B4, DyLight 662Q, DyLight 675-B1, DyLight 675-B2, DyLight 675-B3 DyLight (trademark) 675-B4, DyLight (trademark) 679-C5, DyLight (trademark) 680, DyLight (trademark) 683Q, DyLight (trademark) 690-B1, DyLight (trademark) 690-B2, DyLight (trademark) 696Q, DyLight (trademark) 700-B1, DyLight (trademark) 700-B1, DyLight (trademark) 730-B1, DyLight (trademark) 730-B2, DyLight (trademark) 730-B3, DyLight (trademark) 730-B4, DyLight (trademark) 747, DyLight t(trademark)747-B1, DyLight(trademark)747-B2, DyLight(trademark)747-B3, DyLight(trademark)747-B4, DyLight(trademark)755, DyLight(trademark)766Q, DyLight(trademark)775-B2, DyLight(trademark)775-B3, DyLight(trademark)775-B4, DyLight(trademark)780-B1, DyLight(trademark)780-B2, DyLight(trademark)780-B3, DyLight(trademark)800, DyLight(trademark)830-B2, Dyomics-350,Dyomics-350XL、Dyomics-360XL、Dyomics-370XL、Dyomics-375XL、Dyomics-380XL、Dyomics-390XL、Dyomics-405、Dyomics-415、Dyomics-430、Dyomics-431、Dyomics-478、Dyomics-480XL、Dyomics-481XL、Dyomics-485XL、Dyomics-490、Dyomics-495、Dyomics-505、Dyomics-510XL、Dyomics-511XL、Dyomics-520XL、Dyomics-521XL、Dyomics-530、Dyomics-547、Dyomics-547Pl、Dyomics-548、Dyomics-549、Dyomics-549P1、Dyomics-550、Dyomics-554、Dyomics-555、Dyomics-556、Dyomics-560、Dyomics-590、Dyomics-591、Dyomics-594、Dyomics-601XL、Dyomics-605、Dyomics-610、Dyomics-615、Dyomics-630、Dyomics-631、Dyomics-632、Dyomics-633、Dyomics-634、Dyomics-635、Dyomics-636、Dyomics-647、Dyomics-647P1、Dyomics-648、Dyomics-648P1、Dyomics-649、Dyomics-649P1、Dyomics-650、Dyomics-651、Dyomics-652、Dyomics-654、Dyomics-675、Dyomics-676、Dyomics-677、Dyomics-678、Dyomics-679P1、Dyomics-680、Dyomics-681、Dyomics-682、Dyomics-700、Dyomics-701、Dyomics-703、Dyomics-704、Dyomics-730、Dyomics-731、Dyomics-732、Dyomics-734、Dyomics-749、Dyomics-749P1、Dyomics-750、Dyomics-751、Dyomics-752、Dyomics-754、Dyomics-776、Dyomics-777, Dyomics-778, Dyomics-780, Dyomics-781, Dyomics-782, Dyomics-800, Dyomics-831, eFluor (trademark) 450, EoShin, FITC, Full Oreshine, HiLyte (trademark) Fluor405, HiLyte (trademark) Fluor488, HiLyte (trademark) Fluor532, HiLyte (trademark) Fluor555, HiLyte (trademark) Fluor594, HiLyte (trademark) Fluo r647, HiLyte Fluor680, HiLyte Fluor750, IRDye 680LT, IRDye 750, IRDye 800CW, JOE, LightCycler 640R, LightCycler Red610, LightCycler Red640, LightCycler Red670, LightCycler Red705, Risamin Roadmin B, Naftful Oresine, Oregon Green 488, Oregon Green 514, Pacific Blue, Pacific Green, Pacific Orange (trademark), PET, PF350, PF405, PF415, PF488, PF505, PF532, PF546, PF555P, PF568, PF594, PF610, PF633P, PF647P, Quasar (trademark) 570, Quasar (trademark) 670, Quasar (trademark) 705, ローダミン123, ローダミン6G, ローダミングリーン, ローダミングリーン-X, ローダミンレッド, ROX, Seta (trademark) 375, Seta (trademark) 470, Seta (trademark) 555, Se Seta (trademark) 632, Seta (trademark) 633, Seta (trademark) 650, Seta (trademark) 660, Seta (trademark) 670, Seta (trademark) 680, Seta (trademark) 700, Seta (trademark) 750, Seta (trademark) 780, Seta (trademark) APC-780, Seta (trademark) PerCP-680, Seta (trademark) R-PE-670, Seta (trademark) 646, SeTau380, SeTau425, SeTau647, SeTau405, Square635, Square650, Square660,Square672, Square680, Sulfolodamine 101, TAMRA, TET, Texas Red (trademark), TMR, TRITC, Yakima Yellow (trademark), Zenon (trademark), Zy3, Zy5, Zy5.5, and Zy7.

[0103] (ii) physical separation; Polypeptides (or multiple polypeptides) can be barcoded by physical separation. In some embodiments, polypeptides (or multiple polypeptides) are deposited on or within a solid substrate such that the polypeptide (or multiple polypeptides) remains physically separated from additional polypeptides (or additional multiple polypeptides).

[0104] In some embodiments, the solid substrate is a chip array. In some embodiments, the chip array includes multiple compartments (e.g., wells) and / or injection ports. For example, in some embodiments, the chip array includes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 compartments. In the planet embodiment, the chip array is arranged as follows: 1-2, 1-3, 1-4, 1-5, 1-6, 1-7, 1-8, 1-9, 1-10, 1-11, 1-12, 1-13, 1-14, 1-15, 1-16, 1-17, 1-18, 1-19, 1-20, 2-3, 2-4, 2-5, 2-6, 2-7, 2-8, 2-9, 2-10, 2-11, 2-12, 2-13, 2-14, 2-15, 2-16, 2-17, 2-18, 2- The chip array includes 19, 2-20, 3-4, 3-5, 3-6, 3-7, 3-8, 3-9, 3-10, 3-11, 3-12, 3-13, 3-14, 3-15, 3-16, 3-17, 3-18, 3-19, 3-20, 5-6, 5-7, 5-8, 5-9, 5-10, 5-11, 5-12, 5-13, 5-14, 5-15, 5-16, 5-17, 5-18, 5-19, 5-20, 10-15, or 15-20 compartments. In some embodiments, the chip array includes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 injection ports. In the planet embodiment, the chip array is arranged as follows: 1-2, 1-3, 1-4, 1-5, 1-6, 1-7, 1-8, 1-9, 1-10, 1-11, 1-12, 1-13, 1-14, 1-15, 1-16, 1-17, 1-18, 1-19, 1-20, 2-3, 2-4, 2-5, 2-6, 2-7, 2-8, 2-9, 2-10, 2-11, 2-12, 2-13, 2-14, 2-15, 2-16, 2-17, 2-18, 2 Includes 19, 2-20, 3-4, 3-5, 3-6, 3-7, 3-8, 3-9, 3-10, 3-11, 3-12, 3-13, 3-14, 3-15, 3-16, 3-17, 3-18, 3-19, 3-20, 5-6, 5-7, 5-8, 5-9, 5-10, 5-11, 5-12, 5-13, 5-14, 5-15, 5-16, 5-17, 5-18, 5-19, 5-20, 10-15, or 15-20 injection ports.

[0105] In some embodiments, the chip array comprises a plurality of physically separated spots (or regions) containing immobilized detection molecules, as described herein. For example, in some embodiments, the chip array comprises at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, and fewer The system includes at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 150, at least 200, at least 250, at least 300, at least 400, at least 450, at least 500, at least 550, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 5000, or at least 10,000 physically separated spots.In the globe embodiment, the chip array can be 2-10, 2-20, 2-30, 2-40, 2-50, 2-60, 2-70, 2-80, 2-90, 2-100, 10-20, 10-30, 10-40, 10-50, 10-60, 10-70, 10-80, 10-90, 10-100, 50-100, 50-150, 50-200, 50-250, 50-300, 50-350, 50-400, or 50-450. The array includes 50-500, 50-550, 50-600, 50-650, 50-700, 50-750, 50-800, 50-850, 50-900, 50-950, 50-1000, 500-1000, 500-2000, 500-3000, 500-4000, 500-5000, 500-6000, 500-7000, 500-8000, 500-9000, or 500-10,000 physically separated spots. In some embodiments, the immobilized detection molecules are covalently bonded to the chip array.

[0106] B. Method for determining the origin of barcoded molecules in multiplexed samples In some embodiments, this disclosure relates to a method for determining the origin of barcoded molecules (e.g., polypeptides, DNA, RNA, cDNA, metabolites) in a multiplexed sample. The origin of a barcoded molecule (or the origin of multiple barcoded molecules) is determined through identification of the molecular barcode. Barcode identity can be detected by sequencing (e.g., polypeptide and / or polynucleic acid sequencing), luminescence, hybridization, binding kinetics, physical location on or within a solid substrate, or a combination thereof.

[0107] In some embodiments, the barcoded polypeptides (or multiple barcoded polypeptides) of a multiplexed sample may be sequenced to determine the amino acid sequence of the polypeptide (for example, they may be sequenced in parallel). In such embodiments, the origin of the barcoded polypeptides may be determined before, after, or simultaneously with the sequencing of the polypeptides of the multiplexed sample. In some embodiments, the origin of the barcoded polypeptides is determined before the sequencing of the polypeptides. In some embodiments, the origin of the barcoded polypeptides is determined after the sequencing of the polypeptides. In some embodiments, the origin of the barcoded polypeptides is determined simultaneously with the sequencing of the polypeptides. In some embodiments, the amino acid sequences of the barcoded polypeptides of the multiplexed sample are grouped according to their origin (determined by their barcode identity).

[0108] (i) Methodology for determining polynucleic acid sequences In some embodiments, a method for determining the origin of a barcode molecule (or the origin of multiple barcode molecules) includes detecting the barcode identity of a molecule (or the barcode identity of a barcoded molecule) by sequencing the barcode of that molecule. Thus, in some embodiments, this disclosure relates to methods for sequencing polypeptides and / or polynucleic acids (e.g., deoxyribonucleic acid or ribonucleic acid). Methods for sequencing polypeptides are described below (see "Polypeptide Sequence Methodology"). Polynucleic acid sequence methodologies are also described herein.

[0109] In some embodiments, a polynucleotide sequencing method includes the steps of (i) exposing a complex in a target volume, comprising a target polynucleotide or a plurality of polynucleotides present in a sample, at least one primer, and a polymerase, to one or more labeled nucleotides; (ii) directing one or more excitation energies, or a series of pulses of one or more excitation energies, near the target volume; (iii) detecting a plurality of photons emitted from one or more labeled nucleotides during the sequential incorporation into the polynucleotide containing one of the at least one primers; and (iv) identifying the sequence of the incorporated nucleotides by determining one or more characteristics of the emitted photons.

[0110] In some embodiments, the primer is a sequencing primer. In some embodiments, the sequencing primer can be annealed to a polynucleic acid (e.g., target polynucleic acid) which may or may not be immobilized on a solid support. The solid support may include, for example, a sample well (e.g., a nanoaperture, reaction chamber) on a chip or cartridge used for polynucleic acid sequencing. In some embodiments, the sequencing primer may be immobilized on a solid support, and hybridization of the polynucleic acid (e.g., target nucleic acid) further immobilizes the nucleic acid molecule on the solid support. In some embodiments, a polymerase (e.g., RNA polymerase) is immobilized on a solid support, and the soluble sequencing primer and polynucleic acid are brought into contact with the polymerase. In some embodiments, a complex comprising the polymerase, polynucleic acid (e.g., target nucleic acid) and primer is formed in solution, and the complex is immobilized on a solid support (e.g., via immobilization of the polymerase, primer, and / or target polynucleic acid). In some embodiments, none of the components are immobilized on a solid support. For example, in some embodiments, the complex comprising polymerase, target polynucleic acid, and sequencing primer is formed in situ and the complex is not immobilized on a solid support.

[0111] In some embodiments, according to aspects of the present disclosure, multiple single-molecule sequencing reactions are carried out in parallel (e.g., on a single chip or cartridge). For example, in some embodiments, multiple single-molecule sequencing reactions are carried out in separate sample wells (e.g., nanoapertures, reaction chambers) on a single chip or cartridge.

[0112] Additional polynucleotide sequencing methodologies are known to those skilled in the art. (ii) Molecules for detection In some embodiments, a method for determining the origin of a barcoded molecule (or the origin of multiple barcoded molecules) includes indirectly detecting the barcode identity of the molecule (or the barcode identity of the barcoded molecules) using a detection molecule. For example, in some embodiments, barcode identity is detected by a method including (i) bringing a barcoded molecule (or multiple barcoded molecules) into contact with multiple detection molecules and causing one or more of the multiple detection molecules to interact with the barcode of the barcoded molecule (or with one or more barcodes of the multiple barcoded molecules); and (ii) detecting the interaction between the barcoded molecule and the detection molecule. The interaction between the barcoded molecule and the detection molecule can be identified by luminescence, hybridization, binding kinetics, or physical location.

[0113] In some embodiments, each of the detection molecules in the plurality of detection molecules is chemically identical. In some embodiments, the plurality of detection molecules comprises two or more chemically different detection molecules.

[0114] For example, in some embodiments, the multiple detection molecules include 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 chemically distinct detection molecules.

[0115] In some embodiments, the detection molecules include at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, or at least 1000 chemically different detection molecules.

[0116] In some embodiments, the number of detection molecules is 2-3, 2-4, 2-5, 2-6, 2-7, 2-8, 2-9, 2-10, 2-11, 2-12, 2-13, 2-14, 2-15, 2-16, 2-17, 2-18, 2-19, 2-20, 2-25, 2-30, 2-35, 2-40, 2-45, 2-50, 2-60, 2-70, 2-80, 2-90, 2-100, 2-200, 2-300, 2-400, 2-500, 2-600, 2-700, 2-8 00 pieces, 2~900 pieces, 2~1000 pieces, 5~10 pieces, 5~15 pieces, 5~20 pieces, 5~25 pieces, 5~30 pieces, 5~35 pieces, 5~40 pieces pieces, 5~45 pieces, 5~50 pieces, 5~60 pieces, 5~70 pieces, 5~80 pieces, 5~90 ​​pieces, 5~100 pieces, 5~200 pieces, 5~300 pieces, 5~400 pieces, 5~500 pieces, 5~600 pieces, 5~700 pieces, 5~800 pieces, 5~900 pieces, 10~15 pieces, 10~20 pieces, 10~ 25 pieces, 10~30 pieces, 10~35 pieces, 10~40 pieces, 10~45 pieces, 10~50 pieces, 10~60 pieces, 10~70 pieces, 10~80 pieces, 10~90 pieces, 10~100 pieces, 10~200 pieces, 10~300 pieces, 10~400 pieces, 10~500 pieces, 10~600 pieces, 10~7 00 pieces, 10~800 pieces, 10~900 pieces, 10~1000 pieces, 20~30 pieces, 20~40 pieces, 20~50 pieces, 20~60 pieces, 20 ~70 pieces, 20~80 pieces, 20~90 pieces, 20~100 pieces, 20~200 pieces, 20~300 pieces, 20~400 pieces, 20~500 pieces, 20~600 pieces, 20~700 pieces, 20~800 pieces, 20~900 pieces, 20~1000 pieces, 50~60 pieces, 50~70 pieces, 50~80 pieces Contains 1, 50-90, 50-100, 50-200, 50-300, 50-400, 50-500, 50-600, 50-700, 50-800, 50-900, 50-1000, 100-200, 100-300, 100-400, 100-500, 100-600, 100-700, 100-800, 100-900, 100-1000, 500-600, 500-700, 1500-800, 500-900, or 500-1000 chemically different detection molecules.

[0117] The detection molecule may include a polynucleic acid portion, a polypeptide portion, a small molecule portion, or a combination thereof. In some embodiments, the detection molecule includes a polynucleic acid moiety. In some embodiments, the detection molecule includes two or more polynucleic acid moieties. In embodiments in which the detection molecule includes multiple polynucleic acid moieties, each polynucleic acid moiety may be identical, a subset of polynucleic acid moieties may be identical, or each polynucleic acid moiety may be chemically different.

[0118] In some embodiments, the polynucleic acid portion is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, or 60 nucleotides long.

[0119] In some embodiments, the polynucleotide portion has at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 40, at least 50, at least 60, at least 70 nucleotides, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, or at least 500 nucleotide lengths.

[0120] When you introduce some of them, the polynucleotide portion will be 5-10, 5-15, 5-20, 5-25, 5-30, 5-40, 5-50, 5-60, 5-70, 5-80, 5-90, 5-100, 5-150, 5-200, 5-250, 5-300, 5-350, 5-400, 5-450, 5-500, 10-15, 10-20, 10-25, 10-30, 10-40, 10-50, 10-60, 10-70, 10-80, 10-90, 10-100, 10-150, 10-200, 10-250, 10-300, 10-350, 10-400, 10-450, 1 The subnet lengths are 0-500, 20-30, 20-40, 20-50, 20-60, 20-70, 20-80, 20-90, 20-100, 20-150, 20-200, 20-250, 20-300, 20-350, 20-400, 20-450, 20-500, 50-75, 50-100, 50-150, 50-200, 50-250, 50-500, 50-350, 50-400, 50-450, 50-500, 100-200, 100-250, 100-500, 100-350, 100-400, 100-450, or 100-500.

[0121] In some embodiments, the polynucleic acid portion is an aptamer. In some embodiments, the detection molecule includes a polypeptide moiety. In some embodiments, the detection molecule includes two or more polypeptide moieties. In embodiments in which the detection molecule includes multiple polypeptide moieties, each polypeptide moiety may be identical. A subset of polypeptide moieties may be identical. Alternatively, each polypeptide moiety may be chemically different.

[0122] In some embodiments, the polypeptide portion has a length of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids.

[0123] In some embodiments, the polypeptide portion has a length of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, or at least 500 amino acids.

[0124] When you start with a few options, the polypeptide portion will be 5-10, 5-15, 5-20, 5-25, 5-30, 5-40, 5-50, 5-60, 5-70, 5-80, 5-90, 5-100, 5-150, 5-200, 5-250, 5-300, 5-350, 5-400, 5-450, 5-500, 10-15, 10-20, 10-250, 10-60, 10-70, 10-80, 10-90, 10-100, 10-150, 10-200, 10-250, 10-300, 10-350, 10-400, 10-450 The amino acid lengths are 10-500, 20-30, 20-40, 20-50, 20-60, 20-70, 20-80, 20-90, 20-100, 20-150, 20-200, 20-250, 20-300, 20-350, 20-400, 20-450, 20-500, 50-75, 50-100, 50-150, 50-200, 50-250, 50-500, 50-350, 50-400, 50-450, 50-500, 100-200, 100-250, 100-500, 100-350, 100-400, 100-450, or 100-500.

[0125] In some embodiments, the polypeptide portion is an aptamer. In some embodiments, the polypeptide portion is an antibody. In some embodiments, the polypeptide portion is an antigen. In some embodiments, the polypeptide portion is avidin, streptavidin, or other avidin-like polypeptide, such as traptavidin, tamavidin, bradavidin, xenavidine, and their homologs and variants.

[0126] In some embodiments, the detection molecule includes a small molecule portion, such as a drug portion or a luminescent molecule portion (of a fluorescent molecule portion). In some embodiments, the detection molecule includes two or more small molecule portions. In embodiments in which the detection molecule includes multiple small molecule portions, each small molecule portion may be identical, a subset of small molecule portions may be identical, or each small molecule portion may be chemically different.

[0127] Examples of drugs and luminescent molecules suitable for the methods described herein are known to those skilled in the art. As used herein, a luminescent molecule is a molecule that can absorb one or more photons and then emit one or more photons after one or more time intervals.

[0128] In some embodiments, the luminescent molecule may comprise first and second chromophores. In some embodiments, the excited state of the first chromophore can be relaxed via energy transfer to the second chromophore. In some embodiments, the energy transfer is Forster resonance energy transfer (FRET). Such FRET pairs may be useful in providing luminescent labels having properties that facilitate the distinction of a label from among multiple luminescent labels in a mixture. In yet another embodiment, the FRET pair comprises a first chromophore of the first luminescent label and a second chromophore of the second luminescent label. In certain embodiments, the FRET pair may absorb excitation energy in a first spectral range and emit light in a second spectral range.

[0129] In some embodiments, the luminescent molecule refers to a fluorophore or dye. Typically, the luminescent molecule includes aromatic or heteroaromatic compounds and may be pyrene, anthracene, naphthalene, naphthylamine, acridine, stilbene, indole, benzindol, oxazole, carbazole, thiazole, benzothiazole, benzoxazole, phenanthidine, phenoxazine, porphyrin, quinoline, ethidium, benzamide, cyanine, carbocyanine, salicylate, anthranilate, coumarin, fluorescein, rhodamine, xanthene, or other similar compounds.

[0130] In some embodiments, the luminescent molecule comprises a dye selected from one or more of the following: 5 / 6-carboxyrhodamine 6G, 5-carboxyrhodamine 6G, 6-carboxyrhodamine 6G, 6-TAMRA, Abberior® STAR440SXP, Abberior® STAR470SXP, Abberior® STAR488, Abberior® STAR512, Abberior® STAR520SXP, Abberior® STAR580, Abberior® STAR600, Abberior® STAR635, Abberior® STAR635P, Abberior® STAR RED, AlexaFluor(TM) 350, AlexaFluor(TM) 405, AlexaFluor(TM) 430, AlexaFluor(TM) 480, AlexaFluor(TM) 488, AlexaFluor(TM) 514, AlexaFluor(TM) 532, Alexa Fluor(TM) 546, AlexaFluor(TM) 555, AlexaFluor(TM) 568, AlexaFluor(TM) 594, AlexaFluor(TM) 610-X, AlexaFluor(TM) 633, AlexaFluor(TM) 647, AlexaFluor(TM) )660, AlexaFluor(TM) 680, AlexaFluor(TM) 700, AlexaFluor(TM) 750, AlexaFluor(TM) 790, AMCA, ATTO390, ATTO425, ATTO465, ATTO488, ATTO495, ATTO514, ATTO5 20, ATTO532, ATTO542, ATTO550, ATTO565, ATTO590, ATTO610, ATTO620, ATTO633, ATTO647, ATTO647N, ATTO655, ATTO665, ATTO680, ATTO700, ATTO725, ATTO740, ATTO Oxa12, ATTORho101, ATTORho11, ATTORho12, ATTORho13, ATTORho14, ATTORho3B, ATTORho6G, ATTOThio12, BD Horizon(TM) V450, BODIPY(TM) 493 / 501, BODIPY(TM) 530 / 550,BODIPY (trademark) 558 / 568, BODIPY (trademark) 564 / 570, BODIPY (trademark) 576 / 589, BODIPY (trademark) 581 / 591, BODIPY (trademark) 630 / 650, BODIPY (trademark) 650 / 665, BODIPY (trademark) FL, BODIPY (trademark) FL-X, BODIPY (trademark) R6G, BODIPY (trademark) TMR, BODIPY (trademark) TR, CAL Fluor (trademark) Gold540, CAL Fluor (trademark) Green510, CAL Fluor (trademark) Orange560, CAL Fluor (trademark) Red590, CAL Fluor (trademark) Red610, CAL Fluor (trademark) Red615, CAL Fluor (trademark) Red635, Cascade (trademark) Blue, CF (trademark) 350, CF (trademark) 405M, CF (trademark) 405S, CF (trademark) 488A, CF (trademark) 514, CF (trademark) 532, CF (trademark) 543, CF (trademark) 546, CF (trademark) 555, CF (trademark) 568, CF (trademark) 594, CF (trademark) 620R, CF (trademark) 633, CF (trademark) 633-V1, CF (trademark) 640R, CF (trademark) 640R-V1, CF (trademark) 640R-V2, CF (trademark) 660C, CF (trademark) 660R, CF (trademark) 680, CF (trademark) 680R, CF (trademark) 680R-V1, CF (trademark) 750, CF (trademark) 770, CF (trademark) 790, Chromeo (trademark) 642, Chromis 425N, Chromis 500N, C Chromis 515N, Chromis 530N, Chromis 550A, Chromis 550C, Chromis 550Z, Chromis 560N, Chromis 570N, Chromis 577N, Chromis 600N, Chromis 630N, Chromis 645A, Chromis 645C, Chromis 645Z, Chromis 678A, Chromis 678C, Chromis 678Z, Chromis 770A, Chromis 770C, Chromis 800A, Chromis 800C, Chromis 830A, Chromis 830C, Cy(trademark) 3, Cy(trademark) 3.5, Cy(trademark) 3B, Cy(trademark) 5, Cy(trademark) 5.5, Cy(trademark) 7, DyLight(trademark) 350, DyLight(trademark) 405.DyLight (trademark) 415-Col, DyLight (trademark) 425Q, DyLight (trademark) 485-LS, DyLight (trademark) 488, DyLight (trademark) 504Q, DyLight (trademark) 510-LS, DyLight (trademark) 515-LS, DyLight (trademark) 521-LS, DyLight (trademark) 530-R2, DyLight (trademark) 543Q, DyLight (trademark) 550, DyLight (trademark) 554-R0, DyLight (trademark) 554-R1, DyLight (trademark) 590-R2, DyLight ( Trademarks) 594, DyLight 610-B1, DyLight 615-B2, DyLight 633, DyLight 633-B1, DyLight 633-B2, DyLight 650, DyLight 655-B1, DyLight 655-B2, DyLight 655-B3, DyLight 655-B4, DyLight 662Q, DyLight 675-B1, DyLight 675-B2, DyLight 675-B3 DyLight (trademark) 675-B4, DyLight (trademark) 679-C5, DyLight (trademark) 680, DyLight (trademark) 683Q, DyLight (trademark) 690-B1, DyLight (trademark) 690-B2, DyLight (trademark) 696Q, DyLight (trademark) 700-B1, DyLight (trademark) 700-B1, DyLight (trademark) 730-B1, DyLight (trademark) 730-B2, DyLight (trademark) 730-B3, DyLight (trademark) 730-B4, DyLight (trademark) 747, DyLight t(trademark)747-B1, DyLight(trademark)747-B2, DyLight(trademark)747-B3, DyLight(trademark)747-B4, DyLight(trademark)755, DyLight(trademark)766Q, DyLight(trademark)775-B2, DyLight(trademark)775-B3, DyLight(trademark)775-B4, DyLight(trademark)780-B1, DyLight(trademark)780-B2, DyLight(trademark)780-B3, DyLight(trademark)800, DyLight(trademark)830-B2, Dyomics-350,Dyomics-350XL、Dyomics-360XL、Dyomics-370XL、Dyomics-375XL、Dyomics-380XL、Dyomics-390XL、Dyomics-405、Dyomics-415、Dyomics-430、Dyomics-431、Dyomics-478、Dyomics-480XL、Dyomics-481XL、Dyomics-485XL、Dyomics-490、Dyomics-495、Dyomics-505、Dyomics-510XL、Dyomics-511XL、Dyomics-520XL、Dyomics-521XL、Dyomics-530、Dyomics-547、Dyomics-547Pl、Dyomics-548、Dyomics-549、Dyomics-549P1、Dyomics-550、Dyomics-554、Dyomics-555、Dyomics-556、Dyomics-560、Dyomics-590、Dyomics-591、Dyomics-594、Dyomics-601XL、Dyomics-605、Dyomics-610、Dyomics-615、Dyomics-630、Dyomics-631、Dyomics-632、Dyomics-633、Dyomics-634、Dyomics-635、Dyomics-636、Dyomics-647、Dyomics-647P1、Dyomics-648、Dyomics-648P1、Dyomics-649、Dyomics-649P1、Dyomics-650、Dyomics-651、Dyomics-652、Dyomics-654、Dyomics-675、Dyomics-676、Dyomics-677、Dyomics-678、Dyomics-679P1、Dyomics-680、Dyomics-681、Dyomics-682、Dyomics-700、Dyomics-701、Dyomics-703、Dyomics-704、Dyomics-730、Dyomics-731、Dyomics-732、Dyomics-734、Dyomics-749、Dyomics-749P1、Dyomics-750、Dyomics-751、Dyomics-752、Dyomics-754、Dyomics-776、Dyomics-777, Dyomics-778, Dyomics-780, Dyomics-781, Dyomics-782, Dyomics-800, Dyomics-831, eFluor (trademark) 450, EoShin, FITC, Full Oreshine, HiLyte (trademark) Fluor405, HiLyte (trademark) Fluor488, HiLyte (trademark) Fluor532, HiLyte (trademark) Fluor555, HiLyte (trademark) Fluor594, HiLyte (trademark) Fluo r647, HiLyte Fluor680, HiLyte Fluor750, IRDye 680LT, IRDye 750, IRDye 800CW, JOE, LightCycler 640R, LightCycler Red610, LightCycler Red640, LightCycler Red670, LightCycler Red705, Risamin Roadmin B, Naftful Oresine, Oregon Green 488, Oregon Green 514, Pacific Blue, Pacific Green, Pacific Orange (trademark), PET, PF350, PF405, PF415, PF488, PF505, PF532, PF546, PF555P, PF568, PF594, PF610, PF633P, PF647P, Quasar (trademark) 570, Quasar (trademark) 670, Quasar (trademark) 705, ローダミン123, ローダミン6G, ローダミングリーン, ローダミングリーン-X, ローダミンレッド, ROX, Seta (trademark) 375, Seta (trademark) 470, Seta (trademark) 555, Se Seta (trademark) 632, Seta (trademark) 633, Seta (trademark) 650, Seta (trademark) 660, Seta (trademark) 670, Seta (trademark) 680, Seta (trademark) 700, Seta (trademark) 750, Seta (trademark) 780, Seta (trademark) APC-780, Seta (trademark) PerCP-680, Seta (trademark) R-PE-670, Seta (trademark) 646, SeTau380, SeTau425, SeTau647, SeTau405, Square635, Square650, Square660,Square672, Square680, Sulfolodamine 101, TAMRA, TET, Texas Red (trademark), TMR, TRITC, Yakima Yellow (trademark), Zenon (trademark), Zy3, Zy5, Zy5.5, and Zy7.

[0131] In some embodiments, the detection molecule is immobilized (e.g., covalently bonded) on a substrate. The substrate may be a surface (e.g., a solid surface), beads (e.g., magnetic beads), particles (e.g., magnetic particles), or a gel.

[0132] (iii) Luminescence In some embodiments, a method for determining the origin of a barcoded molecule (or the origin of multiple barcoded molecules) includes detecting the barcode identity of the molecule (or multiple barcoded molecules) by luminescence. The detection of barcode identity may be direct or indirect (for example, by detecting the luminescence of a detection molecule).

[0133] In some embodiments, barcode identities are identified based on luminescence lifetime, luminescence intensity, brightness, absorption spectrum, emission spectrum, emission quantum yield, or a combination of two or more thereof. In some embodiments, multiple barcode identities can be distinguished from one another based on different luminescence lifetimes, luminescence intensity, brightness, absorption spectra, emission spectra, emission quantum yield, or a combination of two or more thereof.

[0134] In some embodiments, luminescence is detected by exposing a luminescent molecule to a series of distinct light pulses and evaluating the timing or other characteristics of each photon emitted from the molecule. In some embodiments, the luminescence lifetime of a molecule is determined from a series of photons emitted from the molecule, and this luminescence lifetime can be used to identify the molecule. In some embodiments, the luminescence intensity of a molecule is determined from a series of photons emitted from the molecule, and this luminescence intensity can be used to identify the molecule. In some embodiments, the luminescence lifetime and luminescence intensity of a molecule are determined from a series of photons emitted from the molecule, and the molecule can be identified using this luminescence lifetime and luminescence intensity.

[0135] In certain embodiments, a light-emitting molecule absorbs one photon and emits one photon after a certain time interval. In some embodiments, the emission lifetime of a molecule can be determined or estimated by measuring that time interval. In some embodiments, the emission lifetime of a molecule can be determined or estimated by measuring multiple time intervals for multiple pulse and emission events. In some embodiments, the emission lifetime of a molecule can be distinguished among the emission lifetimes of multiple types of molecules by measuring the above time intervals. In some embodiments, the emission lifetime of a molecule can be distinguished among the emission lifetimes of multiple types of molecules by measuring multiple time intervals for multiple pulse and emission events. In certain embodiments, a molecule is identified or distinguished among multiple types of labels by determining or estimating the emission lifetime of the label. In certain embodiments, a molecule is identified or distinguished among multiple types of molecules by distinguishing the emission lifetime of a molecule among multiple emission lifetimes of multiple types of molecules.

[0136] The luminescence lifetime of a luminescent molecule can be determined using any suitable method (e.g., by measuring the lifetime using a suitable technique, or by determining the time-dependent characteristics of the emission). In some embodiments, determining the luminescence lifetime of a molecule includes determining the lifetime relative to another label. In some embodiments, determining the luminescence lifetime of a molecule includes determining the lifetime relative to a reference. In some embodiments, determining the luminescence lifetime of a molecule includes measuring the lifetime (e.g., fluorescence lifetime). In some embodiments, determining the luminescence lifetime of a molecule includes determining one or more temporal characteristics that indicate lifetime. In some embodiments, the luminescence lifetime of a molecule can be determined based on the distribution of multiple emission events (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, or more emission events) occurring over one or more time gate windows relative to an excitation pulse. For example, the luminescence lifetime of a molecule can be distinguished from multiple molecules with different luminescence lifetimes based on the distribution of photon arrival times measured with respect to the excitation pulse.

[0137] It should be understood that the luminescence lifetime of a luminescent molecule indicates the timing of photons emitted after the label reaches an excited state, and that labels can be distinguished by information indicating the timing of the photons. Some embodiments may include distinguishing molecules from multiple molecules based on the luminescence lifetime of a label by measuring the time associated with the photons emitted by the molecule. The time distribution may provide an index of the luminescence lifetime that can be determined from the above distribution. In some embodiments, molecules can be distinguished from multiple molecules based on the time distribution, for example, by comparing the time distribution with a reference distribution corresponding to a known molecule. In some embodiments, the value of the luminescence lifetime is determined from the time distribution.

[0138] As used herein, in some embodiments, luminescence intensity refers to the number of photons emitted per unit time by a luminescent molecule excited by the delivery of pulsed excitation energy. In some embodiments, luminescence intensity refers to the number of detected emitted photons per unit time emitted by a molecule excited by the delivery of pulsed excitation energy and detected by a particular sensor or set of sensors.

[0139] Where used herein, in some embodiments, brightness refers to a parameter that reports the average luminescence intensity per luminescent molecule. Thus, in some embodiments, “luminescence intensity” may be used to refer, in most cases, to the brightness of a composition containing one or more molecules. In some embodiments, the brightness of a molecule is equal to the product of its quantum yield and extinction coefficient.

[0140] As used herein, in some embodiments, the emission quantum yield refers to the proportion of excitation events within a given wavelength or spectral range that lead to emission events, and is typically less than 1. In some embodiments, the emission quantum yields of the emission labels described herein are 0 to about 0.001, about 0.001 to about 0.01, about 0.01 to about 0.1, about 0.1 to about 0.5, about 0.5 to 0.9, or about 0.9 to 1. In some embodiments, molecules are identified by determining or estimating the emission quantum yield.

[0141] As used herein, in some embodiments, the excitation energy is a pulse of light from a light source. In some embodiments, the excitation energy is in the visible spectrum. In some embodiments, the excitation energy is in the ultraviolet spectrum. In some embodiments, the excitation energy is in the infrared spectrum. In some embodiments, the excitation energy is at or near the absorption maximum of a light emission label that emits multiple photons to be detected. In certain embodiments, the excitation energy is in the range of about 500 nm to about 700 nm (e.g., about 500 nm to about 600 nm, about 600 nm to about 700 nm, about 500 nm to about 550 nm, about 550 nm to about 600 nm, about 600 nm to about 650 nm, or about 650 nm to about 700 nm). In certain embodiments, the excitation energy may be monochromatic or limited to a spectral range. In some embodiments, the spectral range has the range of about 0.1 nm to about 1 nm, about 1 nm to about 2 nm, or about 2 nm to about 5 nm. In some embodiments, the spectral range is in the range of about 5 nm to about 10 nm, about 10 nm to about 50 nm, or about 50 nm to about 100 nm.

[0142] (iv) physical separation; In some embodiments, a method for determining the origin of a barcoded molecule (or the origin of multiple barcoded molecules) includes detecting the barcode identity of the molecule (or multiple barcoded molecules) by physical separation. Detecting barcode identity by physical separation may include determining the location of the barcoded molecules on a substrate (e.g., a microarray chip).

[0143] For example, a substrate may contain multiple detection molecules (as described herein) organized at separate locations on the substrate. In such a case, barcoded molecules containing barcodes that hybridize, bind to, or are bound to the detection molecules on the substrate can be positioned at the locations of the detection molecules. Thus, in some embodiments, a method for determining the origin of a barcoded molecule (or the origin of multiple barcoded molecules) includes contacting a polypeptide (or multiple polypeptides) with a substrate containing multiple detection molecules.

[0144] As described above, in some embodiments, polypeptides (or plurality of polypeptides) are barcoded by depositing polypeptides (or plurality of polypeptides) on or within a solid substrate such that the polypeptides (or plurality of polypeptides) remain physically separated from additional polypeptides (or additional plurality of polypeptides). In such embodiments, a method for determining the origin of a barcoded molecule (or the origin of a plurality of barcoded molecules) includes detecting the location of the barcoded molecule (or plurality of barcoded molecules) on a solid substrate.

[0145] C. Exemplary Embodiments In some embodiments, the barcode molecule includes a polynucleic acid portion that is identified by DNA sequencing (Figure 3B).

[0146] In some embodiments, the barcode molecule includes a polynucleic acid moiety, which is identified via hybridization using a detection molecule containing the polynucleic acid moiety (Figure 3A). In some embodiments, the detection molecule further includes a luminescent molecule moiety. In some embodiments, the detection molecule is immobilized (e.g., covalently bonded) on a substrate.

[0147] In some embodiments, the barcode molecule includes a polynucleic acid moiety that is identified via hybridization using a detection molecule containing a polypeptide moiety (e.g., a DNA-binding protein, an aptamer, etc.). In some embodiments, the detection molecule further includes a luminescent molecule moiety. In some embodiments, the detection molecule is covalently bonded to a substrate.

[0148] In some embodiments, the barcode molecule includes a polypeptide portion (e.g., a short polypeptide tag) that is identified by polypeptide sequencing (Figure 3C). In some embodiments, the barcode molecule comprises a polypeptide moiety (e.g., a DNA-binding protein, or a portion thereof), which is identified using a detection molecule comprising a polynucleic acid moiety (e.g., a polynucleic acid sequence, or a portion thereof, bound by the DNA-binding protein). In some embodiments, the detection molecule further comprises a luminescent molecule moiety. In some embodiments, the detection molecule is covalently bonded to a substrate.

[0149] In some embodiments, the barcode molecule comprises a polypeptide moiety, which is identified using a detection molecule comprising a polynucleic acid moiety (e.g., an aptamer). In some embodiments, the detection molecule further comprises a luminescent molecule moiety. In some embodiments, the detection molecule is covalently bonded to a substrate.

[0150] In some embodiments, the barcode molecule includes amino acid modifications made to the polypeptide after it has been translated (Figure 3D). In some embodiments, the barcode molecule comprises a polypeptide moiety (e.g., an antibody, antigen, aptamer, etc.), which is identified using a detection molecule comprising the polypeptide moiety (e.g., an antigen, antibody, or substrate, etc.). In some embodiments, the detection molecule further comprises a luminescent molecule moiety. In some embodiments, the detection molecule is covalently bonded to a substrate (Figure 3E).

[0151] In some embodiments, the barcode component includes an endoprotease having a distinct cleavage profile that can be detected by polypeptide sequencing. III. Method for preparing high-concentration samples In some embodiments, the sample is concentrated before, simultaneously with, or after barcoding (e.g., polypeptide barcoding). Therefore, in some embodiments, the disclosure relates to methods for concentrating a sample with respect to one or more molecules of interest (e.g., one or more polypeptides of interest). In particular, in some embodiments, the disclosure relates to methods for polypeptide concentration. As used herein, the term “polypeptide concentration” refers to the process of increasing the abundance of one or more polypeptides of interest compared to the abundance of one or more reference polypeptides (e.g., unintended polypeptides in a complex sample). As used herein, the term “polypeptide of interest” refers to the polypeptide to be concentrated. The polypeptide of interest may include a specific amino acid sequence. Alternatively, the polypeptide of interest may include a specific polypeptide modification (e.g., post-translational modification). These methods facilitate the proteomics analysis of complex samples composed of many different polypeptides, of which only some may be of interest.

[0152] In some embodiments, a method for concentrating polypeptides includes generating a concentrated sample containing a subset of polypeptides by selecting a subset of polypeptides from a plurality of polypeptides using a plurality of concentration molecules. In some embodiments, this method includes generating a concentrated sample containing a subset of polypeptides from a plurality of polypeptides by contacting a plurality of polypeptides with a plurality of concentration molecules.

[0153] In some embodiments, a method for concentrating polypeptides includes (a) contacting a plurality of polypeptides with a plurality of concentration molecules, wherein at least one subset of the concentration molecules in the plurality of concentration molecules binds to a subset of polypeptides in the plurality of polypeptides, thereby generating a bound subset of polypeptides and an unbound subset of polypeptides; and (b) isolating the bound subset of polypeptides to generate a concentration sample containing a subset of polypeptides in the plurality of polypeptides.

[0154] In some embodiments, a method for concentrating polypeptides includes (a) contacting a plurality of polypeptides with a plurality of concentration molecules, wherein at least one subset of the concentration molecules in the plurality of concentration molecules binds to a subset of polypeptides in the plurality of polypeptides, thereby generating a bound subset of polypeptides and an unbound subset of polypeptides; and (b) isolating the unbound subset of polypeptides to generate a concentration sample containing a subset of polypeptides in the plurality of polypeptides.

[0155] In the embodiments described in the preceding paragraphs, it is understood that the binding of the concentration-enhancing molecule to the polypeptide is equivalent to the binding of the polypeptide to the concentration-enhancing molecule. Therefore, step (a) in the above embodiments can be equivalent to the following: (a) Contacting a plurality of polypeptides with a plurality of concentration-enhancing molecules, wherein at least one subset of the concentration-enhancing molecules in the plurality of concentration-enhancing molecules is bound by a subset of polypeptides in the plurality of polypeptides, thereby generating a bound subset of polypeptides and an unbound subset of polypeptides.

[0156] It is also understood that steps (a) and (b) of the above embodiments may be repeated once or more times using additional concentration molecules to produce a more highly concentrated sample. For example, in some embodiments, this method includes (a) contacting a plurality of polypeptides with a first plurality of concentration molecules, wherein at least one subset of the concentration molecules in the first plurality of concentration molecules binds to a subset of polypeptides in the plurality of polypeptides, thereby generating a first bound subset of polypeptides and a first unbound subset of polypeptides; (b) isolating the first bound subset of polypeptides or the first unbound subset of polypeptides from (a); and (c) repeating steps (a) and (b) with one or more additional concentration molecules to produce a highly concentrated sample containing a subset of polypeptides in the plurality of polypeptides. In some embodiments, steps (a) and (b) are repeated using a second, third, fourth, fifth, sixth, seventh, eighth, ninth, tenth, or any number of additional concentration molecules.

[0157] For example, in some embodiments, the method includes (a) contacting a plurality of polypeptides with a first plurality of concentration molecules, wherein at least one subset of the concentration molecules in the first plurality of concentration molecules binds to a subset of polypeptides in the plurality of polypeptides to produce a first bound subset of polypeptides and a first unbound subset of polypeptides; (b) isolating the first bound subset of polypeptides or the first unbound subset of polypeptides from (a); (c) contacting the isolated polypeptides from (b) with a second plurality of concentration molecules, wherein at least one subset of the concentration molecules in the second plurality of concentration molecules binds to a subset of polypeptides isolated in (b) to produce a second bound subset of polypeptides and a second unbound subset of polypeptides; and (d) isolating the second bound subset of polypeptides or the second unbound subset of polypeptides from (c) to produce a concentration sample containing a subset of polypeptides from the plurality of polypeptides.

[0158] Alternatively, or in addition to the above, methods for increasing concentration may include chromatography (e.g., size exclusion, ion exchange), isoelectric focusing, membrane filtration, molecular sieve filtration, concentration, precipitation (e.g., cryoprecipitation), drying, dialysis, or a combination thereof.

[0159] In some embodiments, this method involves contacting a complex sample with a kit or device described herein. See "Sample Preparation Kits" and "Sample Preparation and Sample Sequence Determination Devices."

[0160] In some embodiments, the polypeptides in the high-concentration sample are identical (i.e., they contain the same amino acid sequence). In some embodiments, the high-concentration sample contains at least two unique polypeptides (i.e., those having different amino acid sequences). For example, in some embodiments, the high-concentration sample contains at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, or at least 100 unique polypeptides. In some embodiments, the highly concentrated sample is 1-2, 1-5, 1-10, 1-15, 1-20, 1-30, 1-40, 1-50, 1-60, 1-70, 1-80, 1-90, 1-100, 2-5, 2-10, 2-15, 2-20, 2-30, 2-40, 2-50, 2-60, 2-70, 2-80, 2-90, 2-100, 5-10, 5-15, 5-20, 5-30, 5-40, 5-50, 5-60, 5-70, 5-80, 5-90, 5-100, 10-15, 10-20, 10-30, 10-40, 10-50, 10-60, 10-70, 1 Contains specific polypeptides in the ranges of 0-80, 10-90, 10-100, 15-20, 20-30, 20-40, 20-50, 20-60, 20-70, 20-80, 20-90, 20-100, 20-30, 20-40, 20-50, 20-60, 20-70, 20-80, 20-90, 20-100, 30-40, 30-50, 30-60, 30-70, 30-80, 30-90, 30-100, 40-50, 40-60, 40-70, 40-80, 40-90, 40-100, 50-60, 50-70, 50-80, 50-90, or 50-100.

[0161] In some embodiments, the high-concentration sample contains polypeptides sharing at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% sequence identity. In some embodiments, the high-concentration sample contains polypeptides sharing one or more polypeptide modifications (e.g., post-translational modifications). Examples of post-translational modifications are known to those skilled in the art and are not limited to, but include acetylation, adenylylation, ADP-ribosylation, alkylation (e.g., methylation), amidation, arginylation, biotinylation, butyrylation, carbamylation, carbonylation, carboxylation, citrullination, deamidation, eliminylation, formylation, glycosylation (e.g., N-linked glycosylation, O-linked glycosylation), glipyatyon, glycation, hydroxylation, iodation, ISG formation, isoprenylation, lipoylation, malonylation, myristoylation, nephrylation, nitration, oxidation, and palmitoylation. This includes pegylation, phosphorylation, phosphopantetheinylation, polyglycylation, polyglutamylation, prenylation, propionation, pupyrulation, S-glutathioneation, S-nitrosylation, S-sulfenylation, S-sulfinylation, S-sulfonylation, succinylation, sulfation, SUMOylation, and ubiquitination.

[0162] A. Molecules for high concentration As used herein, the term “concentration molecule” refers to a molecule that exhibits preferential binding to (or by) one or more target polypeptides. A concentration molecule may bind to (or be bound to) a target polypeptide via direct interaction with the amino acid sequence of the target polypeptide. Alternatively, a concentration molecule may bind to (or be bound to) a target polypeptide via interaction with modifications of the target polypeptide (e.g., post-translational modifications). Binding of a concentration molecule to (or by) a target polypeptide may be mediated by electrostatic interactions, hydrophobic interactions, complementary shapes, or a combination thereof.

[0163] In some embodiments, the target polypeptide is the polypeptide of interest. In other embodiments, the target polypeptide is not the polypeptide of interest. Exemplary high-concentration molecules that preferentially bind to one or more target polypeptides (or target polypeptide variants) include immunoglobulins, anticarin, lipocalin, DARPin, aptamers, enzymes, lectins, and peptide interaction domains.

[0164] As used herein, the term “immunoglobulin” refers to a polypeptide characterized by having an immunoglobulin fold, functioning as an antibody, and binding to one or more substrates (e.g., a target polypeptide). Thus, the term “immunoglobulin” encompasses conventional immunoglobulins (i.e., IgA, IgD, IgE, IgG, and IgM), single-chain variable fragments (scFv), antigen-binding fragments (Fab), affibodies, and nanobodies, as well as single-domain antibodies (sdAb) such as VHH and VNAR.

[0165] As used herein, the term "aptamer" refers to a polynucleic acid (e.g., DNA or RNA) or polypeptide that preferentially binds to one or more target molecules (e.g., target polypeptides). While some examples are found in nature, aptamers are typically constructed through repeated in vitro selection.

[0166] As used herein, the term “enzyme” refers to a polymeric biological catalyst that, upon binding to one or more substrates (e.g., a target polypeptide), accelerates a chemical reaction. Typically, enzymes release their substrates after the completion of the chemical reaction. Therefore, in some embodiments in which the concentration molecule contains an enzyme, the enzyme is catalytically inactivated to increase the likelihood that the enzyme remains bound to the substrate. Catalytic inactivation can be achieved by mutagenesis and / or depletion of one or more enzyme cofactors (i.e., non-protein compounds or metal ions necessary for the enzyme's activity as a catalyst).

[0167] As used herein, the term “peptide interaction domain” refers to a polypeptide (or part of a polypeptide) that interacts with one or more polypeptides (e.g., a target polypeptide). For example, a peptide interaction domain may be a scaffold protein, a polypeptide of a multiprotein complex, or a part thereof.

[0168] In some embodiments, the molecule for increasing concentration includes immunoglobulins, aptamers, enzymes, and / or peptide interaction domains. Exemplary high-concentration molecules that are preferentially bound by one or more target polypeptides include oligonucleotides (e.g., double-stranded DNA, single-stranded DNA, double-stranded RNA, single-stranded RNA, etc.), oligosaccharides (or polysaccharides), lipids, glycoproteins, receptor ligands, receptor agonists, receptor antagonists, enzyme substrates, and enzyme cofactors.

[0169] In some embodiments, the concentration molecule includes oligonucleotides (e.g., double-stranded DNA, single-stranded DNA, double-stranded RNA, single-stranded RNA, etc.), oligosaccharides, lipids, receptor ligands, receptor agonists, receptor antagonists, enzyme substrates, and / or enzyme cofactors.

[0170] Preferred binding is used herein to highlight the following characteristics of the concentration molecule: (i) the concentration molecule does not need to exhibit high specificity (i.e., it is sufficient to bind to (or be bound to) a single target polypeptide at a sufficient level); (ii) the concentration molecule may exhibit some degree of off-target binding (i.e., it may bind to (or be bound to) off-target molecules at a detectable level); and (iii) the concentration molecule does not need to bind to the target polypeptide with 100% efficiency (i.e., even if an excess of the concentration molecule is present, not all target polypeptides in a complex sample necessarily need to be bound).

[0171] In some embodiments, the concentration-enhancing molecule preferentially binds to (or is preferentially bound to) a single target polypeptide. However, in other embodiments, the concentration-enhancing molecule preferentially binds to (or is preferentially bound to) two or more target polypeptides.

[0172] In some embodiments, the high-concentration molecule exhibits (or is preferentially bound to) at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, or at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, or at least 10,000 target polypeptides.

[0173] In some embodiments, the concentration-enhancing molecule exhibits (or is preferentially bound to) target polypeptides 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15.

[0174] In some embodiments, the molecules for high concentration are 1-2, 1-5, 1-10, 1-15, 1-20, 1-30, 1-40, 1-50, 1-60, 1-70, 1-80, 1-90, 1-100, 2-5, 2-10, 2-15, 2-20, 2-30, 2-40, 2-50, 2-60, 2-70, 2-80, 2-90, 2-100, 5-10, 5-15, 5-20, 5-30, 5-40, 5-50, 5-60, 5-70, 5-80, 5-90, 5-100, 10-15, 10-20, 10-30, 10-40, 10-50, 10-60, 10-70, 10-80, 10-90, 10-100, 15-20, 20-30, 20-40, 20-50, 20-60, 20-70, 20-80, 20-90, 20-100, 20-30, 20-40, 20-50, 20-60, 20 ~70, 20~80, 20~90, 20~100, 30~40, 30~50, 30~60, 30~70, 30~80, 30~90, 30~100, 40~50, 40~60, 40~70, 40~80, 40~90, 40~100, 50~60, 50~70, 50~80, 50~90, or 50~100, 100~200, 100~300, 100~400, 100~500, 100~6 It exhibits (or is preferentially bound to) target polypeptides of 00, 100-700, 100-800, 100-900, 100-1000, 100-5000, 100-10,000, 500-600, 500-700, 500-800, 500-900, 500-1000, 500-5000, 500-10,000, 1000-5000, or 1000-10,000.

[0175] In some embodiments, the concentration-enhancing molecule exhibits (or is preferentially bound to) multiple related target polypeptides (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, or more related polypeptides) that share at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% sequence homology.

[0176] In some embodiments, the molecule for high concentration is subjected to acetylation, adenylylation, ADP-ribosylation, alkylation (e.g., methylation), amidation, arginylation, biotinylation, butyrylation, carbamylation, carbonylation, carboxylation, citrullination, deamidation, eliminylation, formylation, glycosylation (e.g., N-linked glycosylation, O-linked glycosylation), glipyatyon, saccharification, hydroxylation, iodization, ISGation, isoprenylation, lipoylation, malonylation, myristoylation, nephrylation, nitration, oxidation, and palmitoylation. It exhibits (or is preferentially bound to) post-translational modifications such as pegylation, phosphorylation, phosphopantetheinylation, polyglycylation, polyglutamylation, prenylation, propionation, pupyrulation, S-glutathioneation, S-nitrosylation, S-sulfenylation, S-sulfinylation, S-sulfonylation, succinylation, sulfation, SUMOylation, and ubiquitination.

[0177] The high-concentration molecules can be immobilized (e.g., covalently bonded) onto a substrate (e.g., a capture probe as described in "Devices for Sample Preparation and Sample Sequence Determination"). The substrate may be a surface (e.g., a solid surface), beads (e.g., magnetic beads), particles (e.g., magnetic particles), or a gel.

[0178] (i) Multiple molecules for increasing concentration Typically, the concentration-enhancing methods described herein utilize multiple concentration-enhancing molecules. These multiple concentration-enhancing molecules may be chemically identical (i.e., multiple molecules may constitute a single concentration-enhancing molecule "type"). Alternatively, the multiple concentration-enhancing molecules may comprise combinations of different concentration-enhancing molecules (i.e., having two or more concentration-enhancing molecule "types").

[0179] In some embodiments, the multiple high-concentration molecules include a single high-concentration molecule type. In other embodiments, the multiple high-concentration molecules include a combination of 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, or 15 or more high-concentration molecule types. In some embodiments, the multiple high-concentration molecules include at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, or at least 100, at least 200, at least 300, at least 400, at least 500 high-concentration molecule types.

[0180] In some embodiments, the multiple high-concentration molecules include combinations of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 high-concentration molecule types.

[0181] In some embodiments, the multiple molecules for increasing concentration are 1-2, 1-5, 1-10, 1-15, 1-20, 1-30, 1-40, 1-50, 1-60, 1-70, 1-80, 1-90, 1-100, 2-5, 2-10, 2-15, 2-20, 2-30, 2-40, 2-50, 2-60, 2-70, 2-80, 2-90, 2-100, 5-10, 5-15, 5-20, 5-30, 5-40, 5-50, 5-60, 5-70, 5-80, 5-90, 5-100, 10-15, 10-20, 10-30, 10-40, 10-50, 10-60, 10-70, 10-80, 10-90, 10-100, 15 ~20, 20~30, 20~40, 20~50, 20~60, 20~70, 20~80, 20~90, 20~100, 20~30, 20~40, 20~50, 20~60, 20~70, 20~80, 20~90, 20~100, 30~40, 30~50, 30~60, 30~70, 30~80, 30 Includes combinations of molecular types for high concentration, such as ~90, 30~100, 40~50, 40~60, 40~70, 40~80, 40~90, 40~100, 50~60, 50~70, 50~80, 50~90, or 50~100, 100~200, 100~300, 100~400, or 100~500.

[0182] In some embodiments, each of the multiple concentration molecules preferentially binds to (or is preferentially bound to) a single target polypeptide. In other embodiments, one or more of the multiple concentration molecules (e.g., a subset) exhibit preferential binding to (or is preferentially bound to) two or more target polypeptides. In yet another embodiment, each of the multiple concentration molecules exhibits (or is preferentially bound to) two or more target polypeptides.

[0183] In some embodiments, one or more (e.g., a subset) of the multiple concentration molecules bind to post-translational polypeptide modifications. In other embodiments, each of the multiple concentration molecules exhibits preferential binding to two or more post-translational polypeptide modifications.

[0184] In some embodiments, each of the multiple concentration molecules is immobilized (e.g., covalently bonded) to a substrate (e.g., a capture probe as described in "Devices for Sample Preparation and Sample Sequence Determination"), a surface (e.g., a solid surface), beads (e.g., magnetic beads), particles (e.g., magnetic particles, or a gel). In some embodiments, one or more of the multiple concentration molecules (e.g., a subset) are immobilized (e.g., covalently bonded) to the substrate. Thus, in some embodiments, contact between the multiple polypeptides and the multiple concentration molecules occurs when a sample containing the multiple polypeptides comes into contact with the substrate.

[0185] For example, in some embodiments, the concentration-enhancing molecule is covalently bonded (e.g., crosslinked) within the gel, and the sample is pulled through the gel. In some embodiments, the concentration-enhancing molecule is covalently bonded to a bead (e.g., a magnetic bead) and then pulled down.

[0186] (ii) Multiple molecules for increasing concentration As described above, in some embodiments, the method comprises (a) contacting a plurality of polypeptides with a first plurality of concentration molecules, wherein at least a subset of the concentration molecules in the first plurality of concentration molecules binds to a subset of polypeptides in the plurality of polypeptides, thereby generating a first bound subset of polypeptides and a first unbound subset of polypeptides; (b) isolating the first bound subset of polypeptides or the first unbound subset of polypeptides from (a); and (c) repeating steps (a) and (b) with one or more additional plurality of concentration molecules to generate a concentration sample containing a subset of polypeptides in the plurality of polypeptides. In some embodiments, steps (a) and (b) are repeated with a second, third, fourth, fifth, sixth, seventh, eighth, ninth, tenth, or any number of additional plurality of concentration molecules.

[0187] In some embodiments, each of the multiple concentration molecules used in the polypeptide concentration method is unique (i.e., each comprises multiple different concentration molecules). In other embodiments, two or more of the multiple molecules are identical. In some embodiments, at least one of the multiple concentration molecules targets post-translational polypeptide modification, and at least one of the multiple concentration molecules does not target post-translational modification.

[0188] For example, the first concentration step (using a first set of multiple concentration molecules) can concentrate a specific post-translational polypeptide modification, and the second concentration step (using a second set of multiple concentration molecules) can concentrate a specific polypeptide (and its variants). Alternatively, the first concentration step (using a first set of multiple concentration molecules) can concentrate a specific polypeptide (and its variants), and the second concentration step (using a second set of multiple concentration molecules) can concentrate a specific post-translational modification.

[0189] B. Modification of polypeptides One or more polypeptides of a complex sample can be modified in vitro before, simultaneously with, and / or after, the polypeptide concentration described above. For example, in some embodiments, the complex sample is contacted with a modifier before, simultaneously with, and / or after, the polypeptide concentration is carried out. In particular, the modifier may mediate polypeptide fragmentation, polypeptide modification, addition of post-translational modifications, and / or blocking of one or more functional groups.

[0190] In some embodiments, one or more polypeptides of a complex sample are modified by fragmentation. In some embodiments, fragmentation includes enzymatic digestion. In some embodiments, digestion is carried out by contacting the polypeptide with an endopeptidase (e.g., trypsin) under digestive conditions. In some embodiments, fragmentation includes chemical digestion. Examples of reagents suitable for chemical and enzymatic digestion are known in the Art and, but are not limited to, trypsin, chemotrypsin, Lys-C, Arg-C, Asp-N, Lys-N, BNPS-skatole, CNBr, caspase, formic acid, glutamyl endopeptidase, hydroxylamine, iodosobenzoic acid, neutrophil elastase, pepsin, proline endopeptidase, proteinase K, staphylococcal peptidase I, thermolysin, and thrombin.

[0191] In some embodiments, one or more polypeptides of a complex sample are modified by denaturation (e.g., by thermal and / or chemical means). In some embodiments, one or more polypeptides of a complex sample are post-translationally modified in vitro, for example, by acetylation, adenylylation, ADP-ribosylation, alkylation (e.g., methylation), amidation, arginylation, biotinylation, butyrylation, carbamylation, carbonylation, carboxylation, citrullination, deamidation, eliminylation, formylation, glycosylation (e.g., N-linked glycosylation, O-linked glycosylation), glipyatyon, glycation, hydroxylation, iodation, ISGation, isoprenylation, lipoylation, malonylation, myristoylation, nephrylation, nitration, oxidation, and palmitoylation. It can be modified by pegylation, phosphorylation, phosphopantetheinylation, polyglycylation, polyglutamylation, prenylation, propionation, pupyrulation, S-glutathioneation, S-nitrosylation, S-sulfenylation, S-sulfinylation, S-sulfonylation, succinylation, sulfation, SUMOylation, or ubiquitination.

[0192] In some embodiments, one or more polypeptides of a complex sample are modified by blocking one or more functional groups (e.g., free carboxylic acid groups and / or thiol groups).

[0193] In some embodiments, blocking free carboxylic acid groups refers to the chemical modification of these groups, which alters the chemical reactivity compared to the unmodified carboxylate salt. Suitable carboxylate blocking methods are known in the art and require the modification of the side-chain carboxylate groups to be chemically distinct from the carboxy-terminal carboxylate groups of the polypeptide to be functionalized. In some embodiments, blocking free carboxylic acid groups includes esterification or amidation of the free carboxylic acid groups of the polypeptide. In some embodiments, blocking free carboxylic acid groups includes, for example, methyl esterification of the free carboxylic acid groups of the polypeptide by reacting the polypeptide with methanolic HCl. Additional examples of reagents and techniques useful for blocking free carboxylate groups include, but are not limited to, 4-sulfo-2,3,5,6-tetrafluorophenol (STP) and / or carbodiimides, e.g., N-(3-dimethylaminopropyl)-N'-ethylcarbodiimide hydrochloride (EDAC), uronium reagents, diazomethane, alcohols and acids for Fischer esterification, the use of N-hydroxylsuccinimide (NHS) to form NHS esters (potentially as an intermediate for subsequent ester or amine formation), or any other methods for modifying or blocking carboxylic acids via reaction with carbonyldiimidazole (CDI) or formation of mixed anhydrides, or potentially via ester or amide formation.

[0194] In some embodiments, blocking free thiol groups refers to chemical modification of these groups, which alters their chemical reactivity compared to unmodified thiols. In some embodiments, blocking free thiol groups includes reduction and alkylation of free thiol groups in polypeptides. In some embodiments, reduction and alkylation are carried out by contacting the polypeptide with dithiothreitol (DTT) and one or both of iodoacetamide and iodoacetic acid. Examples of additional and alternative cysteine ​​reducing reagents that can be used are well known and not limited to, but include 2-mercaptoethanol, tris(2-carboxyethyl)phosphine hydrochloride (TCEP), tributylphosphine, dithiobutylamine (DTBA), or any reagent capable of reducing thiol groups. Examples of additional and alternative cysteine ​​blocking (e.g., cysteine ​​alkylation) reagents that can be used are well known and not limited to, but include acrylamide, 4-vinylpyridine, N-ethylmalemide (NEM), N-ε-maleimidocaproic acid (EMCA), or any reagent that modifies cysteine ​​to prevent the formation of disulfide bonds.

[0195] In some embodiments, the N-terminal or C-terminal amino acids of the polypeptide are modified. In some embodiments, the carboxyl terminus of a polypeptide is modified in a manner comprising: (i) blocking a free carboxylate group of the polypeptide; (ii) denaturing the polypeptide (e.g., by thermal and / or chemical means); (iii) blocking a free thiol group of the polypeptide; (iv) digesting the polypeptide to produce at least one polypeptide fragment containing a free C-terminal carboxylate group; and (v) conjugating the functional moiety to the free C-terminal carboxylate group (e.g., chemically). In some embodiments, this method further comprises dialysis of the sample containing the polypeptide after (i) and before (ii).

[0196] In some embodiments, the carboxyl terminus of a polypeptide is modified in a manner comprising: (i) denaturing the polypeptide (e.g., by thermal and / or chemical means); (ii) blocking a free thiol group of the polypeptide; (iii) digesting the polypeptide to produce at least one polypeptide fragment containing a free C-terminal carboxylate group; (iv) blocking a free C-terminal carboxylate group to produce at least one polypeptide fragment containing a blocked C-terminal carboxylate group; and (v) conjugating a functional moiety (e.g., enzymatically) to the blocked C-terminal carboxylate group. In some embodiments, the method further comprises dialyzing the sample containing the polypeptide after (iv) and before (v).

[0197] In some embodiments, a complex sample is contacted with a modifier before concentration to mediate polypeptide fragmentation, polypeptide modification, addition of post-translational modifications, and / or blocking of one or more functional groups. Alternatively, or in addition to this, in some embodiments, a complex sample containing a modifier mediates polypeptide fragmentation, polypeptide modification, addition of post-translational modifications, and / or blocking of one or more functional groups simultaneously with concentration. Alternatively, or in addition to this, in some embodiments, a complex sample with a modifier (or a sample derived therefrom containing one or more target polypeptides) mediates polypeptide fragmentation, polypeptide modification, addition of post-translational modifications, and / or blocking of one or more functional groups after concentration.

[0198] IV. Polypeptide sequencing methodology In some embodiments, molecules (e.g., polypeptides) of multiplexed samples are sequenced. Therefore, in some embodiments, this disclosure relates to methods for polypeptide sequencing and identification. Various methods for sequencing polypeptide molecules are known to those skilled in the art, including mass spectrometry (e.g., peptide mass fingerprinting and tandem mass spectrometry) and Edman degradation. Additional, previously undescribed methods for sequencing polypeptides are described herein.

[0199] As used herein, “sequencing,” “sequencing,” “determining sequences,” and similar terms relating to polypeptides include the determination of partial and complete amino acid sequence information of a polypeptide. That is, these terms include sequence comparison, fingerprinting, and similar levels of information relating to a target molecule, as well as the explicit identification and ordering of each amino acid of the target molecule within a region of interest. These terms include identifying a single amino acid (or the probability of a single amino acid) of a polypeptide. In some embodiments, two or more amino acids (or the probability of two or more amino acids) of a polypeptide are identified. Therefore, in some embodiments, the terms “amino acid sequence” and “polypeptide sequence” as used herein may refer to the polypeptide material itself and are not limited to specific sequence information (e.g., a sequence of letters representing the order of amino acids from one end to the other) that biochemically characterizes a particular polypeptide.

[0200] In some embodiments, the probability of an amino acid at a specific position within a polypeptide is determined and shown in a probability array. For example, in the case of a polypeptide consisting of two amino acids, the terms “sequencing,” “sequencing,” and “sequencing” are used to determine the probability of the amino acids at position 1 and / or 2, such as [[0.80,0.12,0.05,0.01,0.01,0.01,0.00,0.00,0.00,0.00,0.00,0.00,0.00,0.00,0. This may include determining the probabilities as follows: [00,0.00,0.00,0.00,0.00,0.00,0.00,0.00],[0.00,0.10,0.90,0.00,0.00,0.00,0.00,0.00,0.00,0.00,0.00,0.00,0.00,0.00,0.00,0.00,0.00,0.00,0.00,0.00]], where the probabilities in the array correspond to A, R, N, D, C, Q, E, G, H, I, L, K, M, F, P, S, T, W, Y, V, respectively. Those skilled in the art will understand that this example (and exemplary probability arrays) can be extended to accommodate the analysis of additional amino acid identities (e.g., modified amino acids), such as those described herein.

[0201] In some embodiments, sequencing of a polypeptide molecule involves identifying at least two amino acids (or amino acid probabilities) in the polypeptide molecule (e.g., at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven, at least twelve, at least thirteen, at least fourteen, at least fifteen, at least sixteen, at least seventeen, at least eighteen, at least nineteen, at least twenty, at least twenty-five, at least thirty, at least thirty-five, at least forty, at least forty-five, at least fifty, at least sixty, at least seventy, at least eighty, at least ninety, at least one hundred, or more). In some embodiments, the at least two amino acids are adjacent amino acids. In some embodiments, the at least two amino acids are non-adjacent amino acids.

[0202] In some embodiments, sequencing of a polypeptide molecule involves identifying less than 100% of all amino acids in the polypeptide molecule (e.g., less than 99%, less than 95%, less than 90%, less than 85%, less than 80%, less than 75%, less than 70%, less than 65%, less than 60%, less than 55%, less than 50%, less than 45%, less than 40%, less than 35%, less than 30%, less than 25%, less than 20%, less than 15%, less than 10%, less than 5%, less than 1%, or less than that). For example, in some embodiments, sequencing of a polypeptide molecule involves identifying less than 100% of one type of amino acid in the polypeptide molecule (e.g., identifying a portion of all amino acids of one type in the polypeptide molecule). In some embodiments, sequencing of a polypeptide molecule involves identifying less than 100% of each type of amino acid in the polypeptide molecule.

[0203] In some embodiments, sequencing of a polypeptide molecule involves identifying at least one, at least five, at least ten, at least fifteen, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, or more types of amino acids in the polypeptide.

[0204] In some embodiments, the present application provides compositions and methods for sequencing a polypeptide by identifying a set of amino acids present at the terminus of the polypeptide over time (e.g., by repeated detection and cleavage of terminal amino acids). In yet other embodiments, the present application provides compositions and methods for sequencing a polypeptide by identifying the labeled amino acid content of the polypeptide and comparing it with a reference sequence database.

[0205] In some embodiments, the present application provides compositions and methods for sequencing a polypeptide by sequencing multiple fragments of the polypeptide. In some embodiments, polypeptide sequencing involves identifying and / or determining the sequence of a polypeptide by combining sequence information of multiple polypeptide fragments. In some embodiments, combining sequence information may be performed by computer hardware and software. See "Devices for Sample Preparation and Sample Sequencing". Methods described herein may enable sequencing of a set of relevant polypeptides, such as an entire biological proteome. In some embodiments, multiple single-molecule sequencing reactions are performed in parallel (e.g., on a single chip) according to embodiments of the present application. For example, in some embodiments, multiple single-molecule sequencing reactions are each performed in separate sample wells on a single chip or array.

[0206] In some embodiments, the methods provided herein may be used for sequencing and identifying individual polypeptides in a sample containing a complex or highly concentrated mixture of polypeptides. In some embodiments, the present application provides a method for uniquely identifying individual polypeptides in a complex or highly concentrated mixture of polypeptides. In some embodiments, individual polypeptides are detected in a mixed sample by determining the partial amino acid sequence of that polypeptide. In some embodiments, the partial amino acid sequence of a polypeptide is within a contiguous stretch of about 5 to 50 amino acids.

[0207] While we do not wish to be bound by any particular theory, it is believed that most human proteins can be identified by referring to proteomics databases, even with incomplete sequence information. For example, simple modeling of the human proteome has shown that detecting only four amino acids within a 6-40 amino acid stretch can uniquely identify approximately 98% of proteins (see, e.g., Swaminathan, et al. PLoS Comput Biol. 2015, 11(2):e1004080; and Yao, et al. Phys. Biol. 2015, 12(5):055003). Therefore, complex or highly concentrated mixtures of polypeptides can be broken down (e.g., chemically or enzymatically) into short polypeptide fragments of approximately 6-40 amino acids, and sequencing of this polypeptide library would reveal the identity and abundance of each polypeptide present in the original complex or highly concentrated mixture. A composition and method for selectively labeling and identifying polypeptides by determining partial sequence information is described in detail in U.S. Patent Application No. 15 / 510,962, filed on September 15, 2015, entitled "SINGLE MOLECULE PEPTIDE SEQUENCING," which is incorporated in its entirety by reference.

[0208] The embodiments enable sequencing of a single polypeptide molecule with high accuracy, such as at least about 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.9%, 99.99%, 99.999%, or 99.9999%. In some embodiments, the target molecule used in single-molecule sequencing is a polypeptide immobilized on the surface of a solid support, such as the bottom or sidewall surface of a sample well. The sample well may contain other reagents necessary for the sequencing reaction according to the present invention, such as one or more suitable buffers, cofactors, label affinity reagents, and enzymes (e.g., luminescently labeled or unlabeled catalytically active or inactive exopeptidase enzymes).

[0209] Sequencing according to this application may, in some embodiments, involve immobilizing polypeptides on the surface of a substrate (e.g., a solid support, e.g., a chip, e.g., an integrated device as described herein). In some embodiments, polypeptides may be immobilized on the surface of sample wells on the substrate (e.g., the bottom surface of the sample wells). In some embodiments, the N-terminal amino acids of the polypeptide are immobilized (e.g., attached to the surface). In some embodiments, the C-terminal amino acids of the polypeptide are immobilized (e.g., attached to the surface). In some embodiments, one or more non-terminal amino acids are immobilized (e.g., attached to the surface). The immobilized amino acids can be attached using any suitable covalent or non-covalent bonds, as described herein, for example. In some embodiments, multiple polypeptides are attached to multiple sample wells, for example, in an array of sample wells on a substrate (e.g., one polypeptide is attached to the surface of each sample well, e.g., the bottom surface).

[0210] Sequence determination according to this application may, in some aspects, be carried out using a system that enables single-molecule analysis. This system may include a sequencing device and instruments configured to interface with the sequencing device. See "Devices for Sample Preparation and Sample Sequence Determination".

[0211] A. Labeling affinity reagent and method of use In some embodiments, the methods provided herein involve contacting a polypeptide with a labeling affinity reagent (also referred herein as an amino acid recognition molecule, which may or may not contain a label) that selectively binds to one type of terminal amino acid. In some embodiments, as used herein, the terminal amino acid may refer to an amino-terminal amino acid of the polypeptide or a carboxy-terminal amino acid of the polypeptide. In some embodiments, the labeling affinity reagent binds selectively to one type of terminal amino acid rather than other types of terminal amino acids. In some embodiments, the labeling affinity reagent binds selectively to the same type of terminal amino acid rather than one type of internal amino acid. In yet another embodiment, the labeling affinity reagent selectively binds to one type of amino acid at any position in the polypeptide, for example, to the same type of amino acid as the terminal and internal amino acids.

[0212] As used herein, in some embodiments, an amino acid type refers to one of the 20 naturally occurring amino acids or a subset of that type. In some embodiments, an amino acid type refers to a modified variant of one of the 20 naturally occurring amino acids, or a subset of its unmodified and / or modified variants. Examples of modified amino acid variants include, but are not limited to, post-translational modification variants (e.g., acetylation, ADP-ribosylation, caspase cleavage, citrullination, formylation, hydroxylation, methylation, myristoylation, N-linked glycosylation, nedylation, nitration, O-linked glycosylation, oxidation, palmitoylation, phosphorylation, prenylation, S-nitrosylation, sulfation, smoylation, and ubiquitination), chemical modification variants, non-natural amino acids, and protein-constitutive amino acids such as selenocysteine ​​and pyrrolicin. In some embodiments, a subset of amino acid types includes more than 1 and less than 20 amino acids having one or more similar biochemical properties. For example, in some embodiments, the amino acid type refers to one type selected from amino acids having charged side chains (e.g., positively and / or negatively charged side chains), amino acids having polar side chains (e.g., polar uncharged side chains), amino acids having nonpolar side chains (e.g., nonpolar aliphatic and / or aromatic side chains), and amino acids having hydrophobic side chains.

[0213] In some embodiments, the methods provided herein involve contacting a polypeptide with one or more labeling affinity reagents that selectively bind to one or more terminal amino acids. In an exemplary and non-limiting example, if four labeling affinity reagents are used in the method of the present application, any one reagent selectively binds to one terminal amino acid different from the other three types of amino acids to which any of the other three selectively bind (for example, the first reagent binds to the first type, the second reagent binds to the second type, the third reagent binds to the third type, and the fourth reagent binds to the fourth type of terminal amino acid). For the purposes of this discussion, one or more labeling affinity reagents in the context of the methods described herein may alternatively be referred to as a set of labeling affinity reagents.

[0214] In some embodiments, the label affinity reagent set includes at least one to a maximum of six label affinity reagents. For example, in some embodiments, the label affinity reagent set includes one, two, three, four, five, or six label affinity reagents. In some embodiments, the label affinity reagent set includes 10 or fewer label affinity reagents. In some embodiments, the label affinity reagent set includes 8 or fewer label affinity reagents. In some embodiments, the label affinity reagent set includes 6 or fewer label affinity reagents. In some embodiments, the label affinity reagent set includes 4 or fewer label affinity reagents. In some embodiments, the label affinity reagent set includes 3 or fewer label affinity reagents. In some embodiments, the label affinity reagent set includes 2 or fewer label affinity reagents. In some embodiments, the label affinity reagent set includes 4 label affinity reagents. In some embodiments, the label affinity reagent set includes at least two to 20 label affinity reagents (e.g., at least two to 10, at least two to 8, at least four to 20, at least four to 10). In some embodiments, the label affinity reagent set includes more than 20 affinity reagents (e.g., 20-25, 20-30). However, it should be understood that any number of affinity reagents can be used according to the method of this application to accommodate the desired application.

[0215] According to the present invention, in some embodiments, one or more amino acids are identified by detecting the luminescence of a labeling affinity reagent (e.g., an amino acid recognition molecule including a luminescent label). In some embodiments, the labeling affinity reagent comprises an affinity reagent that selectively binds to one type of amino acid and a luminescent label having luminescence associated with the affinity reagent. Thus, the luminescence (e.g., luminescence lifetime, luminescence intensity, and other luminescence properties described elsewhere in this specification) may be associated with the selective binding of the affinity reagent to identify amino acids in a polypeptide. In some embodiments, multiple types of labeling affinity reagents may be used in a manner compatible with the present invention, each type comprising a luminescent label having luminescence uniquely identifiable from among multiple types. Suitable luminescent labels may include luminescent molecules such as fluorophore dyes, which are described elsewhere in this specification.

[0216] In some embodiments, one or more amino acids are identified by detecting one or more electrical properties of a labeling affinity reagent. In some embodiments, the labeling affinity reagent comprises an affinity reagent that selectively binds to one type of amino acid and a conductive label associated with the affinity reagent. Thus, one or more electrical properties (e.g., charge, current oscillation color, and other electrical properties) may be related to the selective binding of the affinity reagent to identify the amino acids of the polypeptide. In some embodiments, multiple types of labeling affinity reagents can be used in a manner consistent with the present application, each type comprising a conductive label that produces a change in electrical signal (e.g., a change in conductance such as the amplitude of conductivity and a change in conductivity transition of a characteristic pattern) that can be uniquely identified from among multiple types. In some embodiments, each of the multiple types of labeling affinity reagents comprises a conductive label having a different number of charged groups (e.g., a different number of negative and / or positively charged groups). Thus, in some embodiments, the conductive label is a charge label. Examples of charge labels include dendrimers, nanoparticles, nucleic acids, and other polymers having multiple charged groups. In some embodiments, the conductive label is uniquely identifiable by its net charge (e.g., net positive charge or net negative charge), its charge density, and / or the number of its charged groups.

[0217] In some embodiments, affinity reagents (e.g., amino acid recognition molecules) can be constructed by those skilled in the art using conventionally known techniques. In some embodiments, desirable properties may include the ability to selectively and with high affinity bind to one type of amino acid only when it is located at the terminal (e.g., N-terminus or C-terminus) of a polypeptide. In yet other embodiments, desirable properties may include the ability to selectively and with high affinity bind to one type of amino acid when it is located at the terminal (e.g., N-terminus or C-terminus) of a polypeptide, and when it is located at an internal position of the polypeptide.

[0218] As used herein, in some embodiments, the terms “selective” and “specific” (and their variations, e.g., selectively, specifically, selectivity, specificity) refer to interactions that preferentially bind. For example, in some embodiments, a label affinity reagent that selectively binds to a certain type of amino acid preferentially binds to that type of amino acid over other types. A selective binding interaction distinguishes one type of amino acid (e.g., one terminal amino acid) from another type of amino acid (e.g., other terminal amino acids) typically by about 10 to 100 times or more (e.g., about 1,000 times or 10,000 times or more). Therefore, it should be understood that a selective binding interaction can refer to any binding interaction that uniquely identifies one type of amino acid from another type of amino acid. For example, in some embodiments, the present application provides a method for polypeptide sequencing by obtaining data showing the association of one or more amino acid recognition molecules with a polypeptide molecule. In some embodiments, the data includes a series of signal pulses corresponding to a series of reversible amino acid recognition molecule binding interactions between the polypeptide molecule and amino acids, and the data can be used to determine the identity of the amino acids. Therefore, in some embodiments, “selective” or “specific” binding interactions refer to detected binding interactions that distinguish one type of amino acid from another type of amino acid. In some embodiments, the labeled affinity reagent (e.g., amino acid recognition molecule) is about 10 -6 Less than M (for example, about 10 -7 Less than M, approximately 10 -8 Less than M, approximately 10 -9 Less than M, approximately 10 -10 Less than M, approximately 10 -11 Less than M, approximately 10 -12 From less than M to 10 -16 Dissociation constant (K) up to the lowest value of M D ) selectively binds to one type of amino acid and does not significantly bind to other types of amino acids. In some embodiments, the label affinity reagent is Kn less than about 100 nM, less than about 50 nM, less than about 25 nM, about 10 nM, or less than about 1 nM. DIt selectively binds to one type of amino acid (e.g., one terminal amino acid). In some embodiments, the label affinity reagent has a K content of about 50 nM to about 50 μM (e.g., about 50 nM to about 500 nM, about 50 nM to about 5 μM, about 500 nM to about 50 μM, about 5 μM to about 50 μM, or about 10 μM to about 50 μM). D It selectively binds to one type of amino acid. In some embodiments, the amino acid recognition molecule selectively binds to one type of amino acid with a KD of approximately 50 nM.

[0219] In some embodiments, the label affinity reagent (e.g., amino acid recognition molecule) is approximately 10 -6 Less than M (for example, about 10 -7 Less than M, approximately 10 -8 Less than M, approximately 10 -9 Less than M, approximately 10 -10 Less than M, approximately 10 -11 Less than M, approximately 10 -12 From less than M to 10 -16 It binds to two or more amino acids with a KD of less than M. In some embodiments, the amino acid recognition molecule has a KD of less than about 100 nM, less than about 50 nM, less than about 25 nM, about 10 nM, or less than about 1 nM. D It binds to two or more amino acids. In some embodiments, the amino acid recognition molecule binds to two or more amino acids at a KD of about 50 nM to about 50 μM (e.g., about 50 nM to about 500 nM, about 50 nM to about 5 μM, about 500 nM to about 50 μM, about 5 μM to about 50 μM, or about 10 μM to about 50 μM). In some embodiments, the amino acid recognition molecule binds to two or more amino acids at a KD of about 50 nM.

[0220] In some embodiments, the label affinity reagent (e.g., an amino acid recognition molecule) is used for at least 0.1 seconds. -1 It binds to at least one amino acid at a dissociation rate (koff). In some embodiments, the dissociation rate is approximately 0.1 s. -1 ~about 1,000s -1 (For example, about 0.5s) -1 ~about 500s -1 Approximately 0.1 seconds -1~about 100 s -1 、about 1 s -1 ~about 100 s -1 、or about 0.5 s -1 ~about 50 s -1 ) is. In some embodiments, the dissociation rate is about 0.5 s -1 ~about 20 s -1 is. In some embodiments, the dissociation rate is about 2 s -1 ~about 20 s -1 is. In some embodiments, the dissociation rate is about 0.5 s -1 ~about 2 s -1 [[ID=xx]]is.

[0221] In some embodiments, the value of KD or koff can be a known literature value, or the value can be determined empirically. For example, the value of KD or koff can be measured by a single-molecule assay or an ensemble assay. In some embodiments, the value of koff can be determined empirically based on the signal pulse information obtained from a single-molecule assay as described elsewhere in this specification. For example, the value of koff can be approximated as the reciprocal of the average pulse width. In some embodiments, the amino acid recognition molecule binds to two or more amino acids with different KD or koff for each of the two or more. In some embodiments, the first KD or koff of the first type of amino acid is at least 10% (e.g., at least 25%, at least 50%, at least 100%, or more) different from the second KD or koff of the second type of amino acid. In some embodiments, the first and second values of KD or koff differ by about 10 - 25%, 25 - 50%, 50 - 75%, 75 - 100%, or more than 100%, e.g., by a difference of about 2-fold, 3-fold, 4-fold, 5-fold, or more.

[0222] In some embodiments, the label affinity reagent comprises a luminescent label (e.g., a label) and an affinity reagent that selectively binds to one or more terminal amino acids of the polypeptide. In some embodiments, the affinity reagent is selective for one amino acid or a subset of amino acid types (e.g., fewer than 20 common types of amino acids) at the terminal position or both terminal and internal positions.

[0223] As described herein, affinity reagents (also known as “recognition molecules”) can be any biomolecule capable of selectively or specifically binding to a particular molecule rather than another (for example, to a particular type of amino acid rather than another type of amino acid, as in the case of “amino acid recognition molecules” as referred herein). Affinity reagents (e.g., recognition molecules) include, for example, proteins and nucleic acids, which may be synthetic or recombinant. In some embodiments, affinity reagents or recognition molecules may be antibodies or antigen-binding moieties of antibodies, or enzymatic biomolecules such as peptidases, aminotransferases, ribozymes, aptazymes, or tRNA synthetases, including aminoacyl-tRNA synthetases, and related molecules described in U.S. Patent No. 5,059,059 (U.S. Patent Application No. 15 / 255,433, filed September 2, 2016, entitled “MOLECULES AND METHODS FOR ITERATIVE POLYPEPTIDE ANALYSIS AND PROCESSING”).

[0224] In some embodiments, the affinity reagent or recognition molecule of the present invention is a degradation pathway protein. Examples of degradation pathway proteins suitable for use as recognition molecules include, but are not limited to, N-endorule pathway proteins such as Arg / N-endorule pathway proteins, Ac / N-endorule pathway proteins, and Pro / N-endorule pathway proteins. In some embodiments, the recognition molecule is an N-endorule pathway protein selected from Gid4 protein, Ubr1UBR box protein, and ClpS protein (e.g., ClpS2).

[0225] Peptidases, also known as proteases or proteinases, are enzymes that catalyze the hydrolysis of peptide bonds. Peptidases digest polypeptides into shorter fragments and are generally classified into endopeptidases and exopeptidases, which cleave polypeptide chains internally and at the ends, respectively. In some embodiments, the labeling affinity reagent includes a peptidase modified to inactivate exopeptidase or endopeptidase activity. In this way, the labeling affinity reagent selectively binds to the polypeptide without cleaving amino acids. In yet other embodiments, a peptidase that has not been modified to inactivate exopeptidase or endopeptidase activity can be used. For example, in some embodiments, the labeling affinity reagent includes a labeled exopeptidase.

[0226] According to certain embodiments of the present application, polypeptide sequencing methods may include repeated detection and cleavage at the termini of polypeptides. In some embodiments, labeled exopeptidases may be used as a single reagent to perform both the amino acid detection and cleavage steps. As generally described, in some embodiments, labeled exopeptidases have aminopeptidase or carboxypeptidase activity such that they selectively bind to and cleave N-terminal or C-terminal amino acids from polypeptides. It should be understood that in certain embodiments, labeled exopeptidases may be catalytically inactivated by those skilled in the art to retain their selective binding properties for use as non-cleavage label affinity reagents, as described herein.

[0227] Exopeptidases generally require polypeptide substrates containing at least one free amino group at their amino terminus or a free carboxyl group at their carboxyl terminus. In some embodiments, the exopeptidases of the present application hydrolyze bonds at or near the terminus of a polypeptide. In some embodiments, the exopeptidases hydrolyze bonds of three residues or less from the polypeptide terminus. For example, in some embodiments, a single hydrolysis reaction catalyzed by the exopeptidase cleaves a single amino acid, dipeptide, or tripeptide from the polypeptide terminus.

[0228] In some embodiments, the exopeptidase according to the present application is an aminopeptidase or carboxypeptidase that cleaves a single amino acid from the amino-terminus or carboxy-terminus, respectively. In some embodiments, the exopeptidase according to the present application is a dipeptidylpeptidase or peptidyldipeptidase, which cleaves a dipeptide from the amino-terminus or carboxy-terminus, respectively. In yet another embodiment, the exopeptidase according to the present application is a tripeptidylpeptidase, which cleaves a tripeptide from the amino-terminus. The classification of peptidases and the activity of each class or subclass are well known and documented in the literature (see, for example, Gurupriya, VS & Roy, S Proteases and Protease Inhibitors in Male Reproduction. Proteases in Physiology and Pathology 195-216 (2017); and Brix, K. & Stoecker, W. Proteases: Structure and Function. Chapter 1).

[0229] The exopeptidases according to this application can be selected or manipulated based on the orientation of the sequencing reaction. For example, in embodiments of sequencing a polypeptide from the amino terminus to the carboxy terminus, the exopeptidase includes aminopeptidase activity. Conversely, in embodiments of sequencing a polypeptide from the carboxy terminus to the amino terminus, the exopeptidase includes carboxypeptidase activity. Examples of carboxypeptidases that recognize specific carboxy-terminal amino acids, which can be used as labeled exopeptidases or inactivated and used as non-cleavage labeling affinity reagents described herein, are described in the literature (e.g., Garcia-Guerrero, MC, et al. (2018) PNAS 115(17)).

[0230] Peptidases suitable for use as cleavage reagents and / or affinity reagents (e.g., recognition molecules) include aminopeptidases that selectively bind to one or more amino acids. In some embodiments, the aminopeptidase recognition molecule is modified to inactivate aminopeptidase activity. In some embodiments, the aminopeptidase cleavage reagent is nonspecific to cleave almost or all types of amino acids from the terminals of a polypeptide. In some embodiments, the aminopeptidase cleavage reagent is more efficient at cleaving one or more amino acids from the terminals of a polypeptide compared to other types of amino acids at the terminals of the polypeptide. For example, the aminopeptidases of this application specifically cleave alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, selenocysteine, serine, threonine, tryptophan, tyrosine, and / or valine. In some embodiments, the aminopeptidase is a proline aminopeptidase. In some embodiments, the aminopeptidase is a proline iminopeptidase. In some embodiments, the aminopeptidase is a glutamic acid / aspartic acid specific aminopeptidase. In some embodiments, the aminopeptidase is a methionine specific aminopeptidase. In some embodiments, the aminopeptidase is the aminopeptidase shown in Table 1. In some embodiments, the aminopeptidase cleavage reagent cleaves the peptide substrates shown in Table 1.

[0231] In some embodiments, the aminopeptidase is a nonspecific aminopeptidase. In some embodiments, the nonspecific aminopeptidase is a zinc metalloproteinase. In some embodiments, the nonspecific aminopeptidase is an aminopeptidase shown in Table 2. In some embodiments, the nonspecific aminopeptidase cleaves the peptide substrates shown in Table 2.

[0232] Thus, in some embodiments, the present application provides an aminopeptidase (e.g., an aminopeptidase recognition molecule, an aminopeptidase cleavage reagent) having an amino acid sequence selected from Table 1 or Table 2 (or having an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, 80-90%, 90-95%, 95-99%, or more amino acid sequence identity to the amino acid sequence selected from Table 1 or Table 2). In some embodiments, the aminopeptidase has 25-50%, 50-60%, 60-70%, 70-80%, 80-90%, 90-95%, or 95-99%, or more amino acid sequence identity to the aminopeptidase described in Table 1 or Table 2. In some embodiments, the aminopeptidase is a modified aminopeptidase and contains one or more amino acid mutations relative to the sequence shown in Table 1 or Table 2.

[0233]

Table 1-1

[0234]

Table 1-2

[0235]

Table 2-1

[0236]

Table 2-2

[0237]

Table 2-3

[0238]

Table 2-4

[0239] For the purpose of comparing two or more amino acid sequences, the percentage of "sequence identity" (also referred to as "amino acid identity" in this application) between a first amino acid sequence and a second amino acid sequence is calculated by dividing [the number of amino acid residues in the first amino acid sequence that are identical to the amino acid residues at the corresponding positions in the second amino acid sequence] by [the total number of amino acid residues in the first amino acid sequence] and then multiplying by

[0100] . Here, each deletion, insertion, substitution, or addition of an amino acid residue in the second amino acid sequence compared to the first amino acid sequence is considered a difference at a single amino acid residue (position). Alternatively, the degree of sequence identity between two amino acid sequences can be calculated using known computer algorithms (e.g., by the local homology algorithm of Smith and Waterman (1970) Adv. Appl. Math. 2:482c, by the homology alignment algorithm of Needleman and Wunsch, J. Mol. Biol. (1970) 48:443, by the similarity search method of Pearson and Lipman. Proc. Natl. Acad. Sci. USA (1998) 85:2444, or by computerized implementations of algorithms available as Blast, Clustal Omega, or other sequence alignment algorithms), for example, using standard settings. Typically, for the purpose of determining the percentage of "sequence identity" between two amino acid sequences according to the calculation methods outlined above, the amino acid sequence with the most amino acid residues is considered the "first" amino acid sequence, and the other amino acid sequence is considered the "second" amino acid sequence.

[0240] In addition to or instead of the above, two or more sequences may be evaluated for identity between them. In the context of two or more nucleic acids or amino acid sequences, the terms “identical” or “identity” percentage refer to two or more sequences or subsequences that are the same. Two sequences are “substantially identical” if, when measured using one of the sequence comparison algorithms described above, or by manual alignment and visual inspection, they have a certain percentage of the same amino acid residues or nucleotides across a particular region or across the entire sequence (e.g., at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9% identical). Optionally, identity exists over regions of at least approximately 25, 50, 75, or 100 amino acid lengths, or over regions of 100–150, 150–200, 100–200, or 200 or more amino acid lengths.

[0241] In addition to or instead of this, two or more sequences may be evaluated for alignment between sequences. In the context of two or more nucleic acids or amino acid sequences, the terms “alignment” or “alignment” percentage refer to two or more sequences or subsequences that are identical. Two sequences are “substantially aligned” if, when compared and aligned to the maximum degree of similarity across a comparison window or designated region by measurement using one of the sequence comparison algorithms described above, or by manual alignment and visual inspection, they have a certain percentage of the same amino acid residues or nucleotides across a particular region or across the entire sequence (e.g., at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9% identical). Optionally, alignments may be present over regions of at least approximately 25, 50, 75, or 100 amino acid lengths, or over regions of 100–150, 150–200, 100–200, or 200 or more amino acid lengths.

[0242] In addition to polypeptide molecules, nucleic acid molecules possess various advantageous properties for use as affinity reagents (e.g., amino acid recognition molecules) according to this invention. Nucleic acid aptamers are nucleic acid molecules engineered to bind to a desired target with high affinity and selectivity. Therefore, nucleic acid aptamers can be engineered to selectively bind to a desired type of amino acid using selection and / or concentration techniques known in the art. Accordingly, in some embodiments, affinity reagents include nucleic acid aptamers (e.g., DNA aptamers, RNA aptamers). In some embodiments, labeled affinity reagents are labeled aptamers that selectively bind to a single terminal amino acid. For example, in some embodiments, labeled aptamers selectively bind to a single amino acid (e.g., a single type of amino acid or a subset of amino acid types) at the terminus of a polypeptide, as described herein. It should be understood that, although not shown, labeled aptamers can be engineered to selectively bind to a single amino acid at any position in a polypeptide (e.g., terminal position or terminal and internal positions) according to the methods of this application.

[0243] In some embodiments, the label affinity reagent includes a label having binding-induced luminescence. For example, in some embodiments, the labeled aptamer has donor and acceptor labeling and function. In yet other embodiments, the labeled aptamer includes a quenched moiety that functions similarly to a molecular beacon, and the luminescence of the labeled aptamer is quenched internally as a free molecule and recovered as a selectively bound molecule (see, e.g., Hamaguchi et al. (2001) Analytical Biochemistry 294, 126-131). While we do not wish to be constrained by theory, these and other types of mechanisms of binding-induced luminescence are thought to be able to favorably reduce or eliminate background luminescence, thereby improving the overall sensitivity and precision of the methods described herein.

[0244] In addition to methods for identifying the terminal amino acids of a polypeptide, this application provides a method for sequencing a polypeptide using a label affinity reagent. In some embodiments, the sequencing method may include subjecting the polypeptide end to an iterative cycle of terminal amino acid detection and terminal amino acid cleavage. For example, in some embodiments, this application provides a method for determining the amino acid sequence of a polypeptide, comprising contacting the polypeptide with one or more label affinity reagents described herein and subjecting the polypeptide to Edman degradation.

[0245] Conventional Edman degradation involves repeating a cycle of modifying and cleaving the terminal amino acids of a polypeptide, where each amino acid cleaved sequentially is identified to determine the amino acid sequence of the polypeptide. As an example of conventional Edman degradation, the N-terminal amino acid of a polypeptide is modified using phenyl isothiocyanate (PITC) to form a PITC-derivative N-terminal amino acid. The PITC-derivative N-terminal amino acid is then cleaved using acidic, basic, and / or high temperatures. It has also been shown that the step of cleaving the PITC-derivative N-terminal amino acid can be achieved enzymatically using a modified cysteine ​​protease derived from the protozoan Trypanosoma cruzi, under relatively mild cleavage conditions at a neutral or near-neutral pH. A non-limiting example of a useful enzyme is described in U.S. Patent No. 5,059,059 (U.S. Patent Application No. 15 / 255,433, filed September 2, 2016, entitled "MOLECULES AND METHODS FOR ITERATIVE POLYPEPTIDE ANALYSIS AND PROCESSING").

[0246] In some embodiments, sequencing by Edman degradation involves providing a polypeptide immobilized on a solid support surface via a linker (e.g., immobilized on the bottom or sidewall surface of a sample well). In some embodiments, as described herein, the polypeptide is immobilized at one end (e.g., an amino-terminal amino acid or a carboxy-terminal amino acid), and as a result, the other end is free for detection and cleavage of the terminal amino acid. Thus, in some embodiments, the reagents used in the Edman degradation method described herein preferentially interact with the unimmobilized (e.g., free) terminal amino acids of the polypeptide. In this way, the polypeptide remains immobilized over repeated detection and cleavage cycles. For this purpose, in some embodiments, the linker may be designed according to a desired set of conditions used for detection and cleavage, for example, to limit the desorption of the polypeptide from the surface under chemical cleavage conditions. Suitable linker compositions and techniques for immobilizing polypeptides on a surface are described in detail elsewhere herein.

[0247] According to the present invention, in some embodiments, a sequencing method by Edman degradation comprises step (i) contacting a polypeptide with one or more label affinity reagents that selectively bind to one or more terminal amino acids. In some embodiments, the label affinity reagents interact with the polypeptide by selectively binding to terminal amino acids. In some embodiments, step (i) further comprises removing one or more label affinity reagents that do not selectively bind to terminal amino acids (e.g., free terminal amino acids) of the polypeptide.

[0248] In some embodiments, the method further includes identifying the terminal amino acid of the polypeptide by detecting a labeled affinity reagent. In some embodiments, the detection includes detecting luminescence from the labeled affinity reagent. As described herein, in some embodiments, the luminescence is uniquely associated with the labeled affinity reagent such that the luminescence is associated with the type of amino acid to which the labeled affinity reagent selectively binds. Thus, in some embodiments, the type of amino acid is identified by determining one or more luminescence characteristics of the labeled affinity reagent.

[0249] In some embodiments, the Edman degradation sequencing method includes step (ii) of removing the terminal amino acid of the polypeptide. In some embodiments, step (ii) includes removing a labeled affinity reagent (e.g., any one of one or more labeled affinity reagents that selectively binds to the terminal amino acid) from the polypeptide. In some embodiments, step (ii) includes modifying the terminal amino acid (e.g., free terminal amino acid) of the polypeptide by contacting the terminal amino acid with isothiocyanate (e.g., PITC) to form an isothiocyanate-modified terminal amino acid. In some embodiments, the isothiocyanate-modified terminal amino acid is more sensitive to removal by a cleavage reagent (e.g., a chemical or enzymatic cleavage reagent) than the unmodified terminal amino acid.

[0250] In some embodiments, step (ii) includes removing terminal amino acids by contacting the polypeptide with a protease that specifically binds to and cleaves isothiocyanate-modified terminal amino acids. In some embodiments, the protease includes a modified cysteine ​​protease. In some embodiments, the protease includes a modified cysteine ​​protease such as a cysteine ​​protease derived from Trypanosoma cruzi (see, for example, Borgo et al. (2015) Protein Science 24:571-579). In yet other embodiments, step (ii) includes removing terminal amino acids by exposing the polypeptide to chemical conditions (e.g., acidic, basic) sufficient to cleave the isothiocyanate-modified terminal amino acids.

[0251] In some embodiments, the Edman degradation sequencing method includes step (iii) washing the polypeptide following terminal amino acid cleavage. In some embodiments, the washing includes removing proteases. In some embodiments, the washing includes restoring the polypeptide to neutral pH conditions (e.g., after chemical cleavage under acidic or basic conditions). In some embodiments, the Edman degradation sequencing method includes repeating steps (i) to (iii) for multiple cycles.

[0252] In some embodiments, a sample containing a complex or highly concentrated mixture of polypeptides (e.g., a polypeptide mixture) can be broken down into short polypeptide fragments of about 6–40 amino acids using a common enzyme. In some embodiments, sequencing of this polypeptide library by the method of the present invention will reveal the identity and abundance of each polypeptide present in the original complex or highly concentrated mixture. As described herein and in the literature, most polypeptides in the 6–40 amino acid size range can be uniquely identified by determining the number and position of just four amino acids in the polypeptide chain.

[0253] Therefore, in some embodiments, the Edman degradation sequencing method can be performed using a set of labeled aptamers comprising four DNA aptamer types, each type recognizing a different N-terminal amino acid. Each aptamer type can be labeled with a different luminescent label, and as a result, different aptamer types can be distinguished based on one or more luminescent properties. For illustrative purposes, an exemplary set of labeled aptamers includes: a cysteine-specific aptamer labeled with a first luminescent label ("Dye 1"); a lysine-specific aptamer labeled with a second luminescent label ("Dye 2"); a tryptophan-specific aptamer labeled with a third luminescent label ("Dye 3"); and a glutamate-specific aptamer labeled with a fourth luminescent label ("Dye 4").

[0254] In some embodiments, prior to step (i), a single polypeptide molecule from a polypeptide library is immobilized on the surface of a solid support, for example, the bottom or sidewall surface of a sample well in an array of sample wells. In some embodiments, as described elsewhere herein, a portion enabling surface immobilization (e.g., biotin) or a portion improving solubility (e.g., oligonucleotide) may be chemically or enzymatically attached to the C-terminus of the polypeptide. To determine the sequence of each polypeptide, in some embodiments, the immobilized polypeptide is subjected to repeated cycles of N-terminal amino acid detection and N-terminal amino acid cleavage. In some embodiments, the process includes reagent addition and washing steps, which are performed by injection into a flow cell on the detection surface using an automated fluid system. In some embodiments, steps (i) to (iv) represent one cycle of detection and cleavage using a labeled aptamer.

[0255] In some embodiments, the Edman degradation sequencing method comprises step (i) of running a mixture of four orthogonally labeled DNA aptamers and incubating them so that the aptamers bind to any immobilized polypeptide (e.g., a polypeptide immobilized in a sample well of an array) that contains one of the four correct amino acids at its N-terminus. In some embodiments, the method further comprises washing the immobilized polypeptide to remove unbound aptamers. In some embodiments, the method further comprises imaging the immobilized polypeptide ("imaging step (i)"). In some embodiments, the acquired image contains sufficient information to determine the location of the aptamer-binding polypeptide (e.g., its location in the array of the sample well) and which of the four aptamers is bound at each location. In some embodiments, the method further comprises washing the immobilized polypeptide with a suitable buffer to remove aptamers from the immobilized polypeptide.

[0256] In some embodiments, the sequencing method includes step (ii) passing a solution containing a reactive molecule (e.g., PITC, as shown) that specifically modifies the N-terminal amine group. In some embodiments, the isothiocyanate molecule, such as PITC, modifies the N-terminal amino acid to be a substrate for cleavage by a modified protease, such as the cysteine ​​protease Kruzain derived from Trypanosoma cruzi.

[0257] In some embodiments, the sequencing method includes (iii) washing the immobilized polypeptide before flushing it with a suitable modified protease that recognizes and cleaves the modified N-terminal amino acid from the immobilized polypeptide.

[0258] In some embodiments, the method includes step (iv) of washing the immobilized polypeptide after enzymatic cleavage. In some embodiments, steps (i) to (iv) describe one cycle of Edman degradation. Thus, step (i') shown is the start of the next reaction cycle, which proceeds as steps (i') to (iv') performed as described above for steps (i) to (iv). In some embodiments, steps (i) to (iv) are repeated for about 20 to 40 cycles.

[0259] In some embodiments, labeled isothiocyanates (e.g., dye-labeled PITCs) may be used to monitor the loading of the sample. For example, in some embodiments, before subjecting the polypeptide sample to a sequencing method, the polypeptide sample is pre-bound at its ends with a luminescent label by terminal modification using a dye-labeled PITC. Thus, the loading of the polypeptide sample into the array of sample wells can be monitored by detecting luminescence from the label before step (i) above. In some embodiments, the luminescence is used to determine single occupancy of sample wells in the array (e.g., the percentage of sample wells containing a single polypeptide molecule). This can advantageously increase the amount of information that can be reliably obtained about a given sample. Once the desired sample loading state is determined by luminescence, chemical or enzymatic cleavage can be performed as described before proceeding to step (i).

[0260] In some embodiments, labeled isothiocyanates (e.g., dye-labeled PITCs) may be used to monitor the progress of the reaction of polypeptide samples in an array. For example, in some embodiments, step (ii) includes running a solution containing dye-labeled PITCs that specifically modify and label the N-terminal amine groups of polypeptides in the sample. In some embodiments, luminescence from the label may be detected during or after step (ii) to evaluate the N-terminal PITC modification of polypeptides in the sample. Thus, in some embodiments, luminescence is used to determine whether or when to proceed from step (ii) to step (iii). In some embodiments, luminescence from the label may be detected during or after step (iii) to evaluate the N-terminal amino acid cleavage of polypeptides in the sample, for example, to determine whether or when to proceed from step (iii) to step (iv).

[0261] Sequence sequencing methods may utilize separate reagents to detect and cleave the terminal amino acids of polypeptides. However, in some embodiments, the present application provides a sequence sequencing method in which a single reagent, comprising a peptidase (such as a labeled exopeptidase that selectively binds to and cleaves different types of terminal amino acids), may be used to detect and cleave the terminal amino acids of polypeptides.

[0262] Labeled exopeptidases may include lysine-specific exopeptidases containing a first luminescent label, glycine-specific exopeptidases containing a second luminescent label, aspartic acid-specific exopeptidases containing a third luminescent label, and leucine-specific exopeptidases containing a fourth luminescent label. According to certain embodiments described herein, each labeled exopeptidase selectively binds to and cleaves an amino acid only if that amino acid is at the amino-terminus or carboxy-terminus of the polypeptide. Thus, as sequencing by this approach proceeds from one end of the peptide to the other, the labeled exopeptidases are manipulated or selected so that all reagents in the set have either aminopeptidase or carboxypeptidase activity.

[0263] In some embodiments, the present invention provides a method for real-time polypeptide sequencing by evaluating the binding interaction between terminal amino acids and labeled amino acid recognition molecules (e.g., labeled affinity reagents) and labeled cleavage reagents (e.g., labeled nonspecific exopeptidases). Although we do not wish to be constrained by theory, the label affinity reagent has a binding rate, i.e., the "on" rate of binding (k on ), and the dissociation rate, i.e., the "off" rate of the binding (k off ) is defined by binding affinity (K D It selectively couples according to the rate constant k. off and k on These are important determinants of pulse duration (e.g., the time corresponding to a detectable coupling event, respectively) and pulse interval (e.g., the time between detectable coupling events). In some embodiments, these speeds can be designed to achieve pulse duration and pulse rate (e.g., signal pulse frequency) that give the best sequence determination accuracy.

[0264] The sequencing reaction mixture may further comprise a labeled nonspecific exopeptidase having a different luminescence label than that of the label affinity reagent. In some embodiments, the labeled nonspecific exopeptidase is present in the mixture at a lower concentration than that of the label affinity reagent. In some embodiments, the labeled nonspecific exopeptidase exhibits broad specificity such that it cleaves almost or all types of terminal amino acids.

[0265] In some embodiments, terminal amino acid cleavage by labeled nonspecific exopeptidase generates signal pulses, and these events occur at a lower frequency than the binding pulses of the labeled affinity reagent. In this way, amino acids of the polypeptide can be counted and / or identified in a real-time sequencing process. In some embodiments, multiple labeled affinity reagents may be used, each having a diagnostic pulse pattern (e.g., a characteristic pattern) that can be used to identify the corresponding terminal amino acids. For example, in some embodiments, different characteristic patterns correspond to the association of two or more labeled affinity reagents with different types of terminal amino acids. It should be understood that a single affinity reagent that associates with multiple types of amino acids may be used in accordance with the application, as described herein. Thus, in some embodiments, different characteristic patterns correspond to the association of one labeled affinity reagent with different types of terminal amino acids.

[0266] As detailed above, real-time sequencing processes typically involve cycles of terminal amino acid recognition and terminal amino acid cleavage, and the relative occurrence of recognition and cleavage can be controlled by the difference in concentrations between the labeled affinity reagent and the labeled nonspecific exopeptidase. In some embodiments, this concentration difference can be optimized so that the number of signal pulses detected during the recognition of individual amino acids provides a desired confidence interval for identification. For example, if the signal data provided in the initial sequencing reaction contains too few signal pulses between cleavage events to determine a characteristic pattern within the desired confidence interval, the sequencing reaction can be repeated by reducing the concentration of the nonspecific exopeptidase relative to the affinity reagent. The inventors are aware of further techniques for controlling real-time sequencing reactions, which can be used in combination with or instead of the described concentration difference approach.

[0267] In some embodiments, the sequencing reaction includes temperature-dependent cycles of terminal amino acid recognition and terminal amino acid cleavage. Each cycle of the sequencing reaction can be performed within two temperature ranges. The first temperature range ("T1") is optimized for the activity of the affinity reagent rather than the exopeptidase activity (e.g., to promote terminal amino acid recognition), and the second temperature range ("T2") is optimized for the activity of the exopeptidase activity rather than the affinity reagent activity (e.g., to promote terminal amino acid cleavage). The sequencing reaction can proceed by alternating the temperature of the reaction mixture between the first temperature range T1 (to initiate amino acid recognition) and the second temperature range T2 (to initiate amino acid cleavage). Thus, the progress of the temperature-dependent sequencing process is temperature-controllable, and alternating between the temperature ranges (e.g., between T1 and T2) can be done manually or through an automated process. In some embodiments, the affinity reagent activity within the first temperature range T1 compared to the second temperature range T2 (e.g., binding affinity to amino acids (K)) is used. DThe exopeptidase activity (e.g., rate of conversion of substrate to cleavage product) in the second temperature range T2 compared to the first temperature range T1 is at least 2 times, 10 times, at least 25 times, at least 50 times, at least 100 times, at least 1,000 times, or more.

[0268] In some embodiments, the first temperature range T1 is lower than the second temperature range T2. In some embodiments, the first temperature range T1 is about 15°C to about 40°C (e.g., about 25°C to about 35°C, about 15°C to about 30°C, about 20°C to about 30°C). In some embodiments, the second temperature range T2 is about 40°C to about 100°C (e.g., about 50°C to about 90°C, about 60°C to about 90°C, about 70°C to about 90°C). In some embodiments, the first temperature range T1 is about 20°C to about 40°C (e.g., about 30°C), and the second temperature range T2 is about 60°C to about 100°C (e.g., about 80°C).

[0269] In some embodiments, the first temperature range T1 is higher than the second temperature range T2. In some embodiments, the first temperature range T1 is about 40°C to about 100°C (e.g., about 50°C to about 90°C, about 60°C to about 90°C, about 70°C to about 90°C). In some embodiments, the second temperature range T2 is about 15°C to about 40°C (e.g., about 25°C to about 35°C, about 15°C to about 30°C, about 20°C to about 30°C). In some embodiments, the first temperature range T1 is about 60°C to about 100°C (e.g., about 80°C), and the second temperature range T2 is about 20°C to about 40°C (e.g., about 30°C).

[0270] In some embodiments, the present application provides a luminescence-dependent sequencing process using a luminescence-activating reagent. In some embodiments, the luminescence-dependent sequencing process includes a cycle of luminescence-dependent amino acid recognition and cleavage. Each cycle of the sequencing reaction can be performed by exposing the sequencing reaction mixture to two different luminescence conditions: a first luminescence condition optimized for affinity reagent activity rather than exoceptidase activity (e.g., to promote amino acid recognition), and a second luminescence condition optimized for exoceptidase activity rather than affinity reagent activity (e.g., to promote amino acid cleavage). The sequencing reaction proceeds by alternating exposure of the reaction mixture to the first luminescence condition (to initiate amino acid recognition) and exposure of the reaction mixture to the second luminescence condition (to initiate amino acid cleavage). In some embodiments, the two different luminescence conditions include a first wavelength and a second wavelength, not limiting them to examples.

[0271] In some embodiments, the present application provides a method for real-time polypeptide sequencing by evaluating the binding interactions between one or more labeled affinity reagents and terminal and internal amino acids, and the binding interactions between labeled nonspecific exopeptidases and terminal amino acids. In some embodiments, a labeled affinity reagent is used that selectively binds to and dissociates from a single amino acid at both terminal and internal positions. Selective binding generates a series of pulses in the signal output. However, in this approach, the series of pulses occurs at a rate determined by the number of different amino acid species in the entire polypeptide. Thus, in some embodiments, the rate of pulses corresponding to binding events would be a diagnostic indicator of the number of congener amino acids currently present in the polypeptide.

[0272] The labeled nonspecific peptidase may be present at a relatively lower concentration than the label affinity reagent, for example, to provide an optimal time frame between cleavage events. Furthermore, in certain embodiments, a uniquely identifiable luminescence label on the labeled nonspecific peptidase will indicate when a cleavage event occurred. As the polypeptide undergoes repeated cleavage, the pulse rate corresponding to binding by the label affinity reagent will decrease stepwise whenever a terminal amino acid is cleaved by the labeled nonspecific peptidase. Thus, in some embodiments, amino acids may be identified in this approach based on the pulse pattern and / or the pulse rate occurring within the detected pattern between cleavage events, thereby allowing the polypeptide to be sequenced.

[0273] B. Sequence determination by degradation of labeled polypeptide In some embodiments, the present application provides a method for sequencing a polypeptide by identifying unique combinations of amino acids corresponding to known polypeptide sequences. In some embodiments, the method includes detecting selectively labeled amino acids of a labeled polypeptide. In some embodiments, the labeled polypeptide comprises selectively modified amino acids such that different amino acid types contain different luminescence labels. As used herein, unless otherwise specified, labeled polypeptide refers to a polypeptide comprising one or more selectively labeled amino acid side chains. Details relating to methods of selective labeling and the preparation and analysis of labeled polypeptides are known in the art (see, for example, Swaminathan, et al. PLoS Comput Biol. 2015, 11(2):e1004080).

[0274] As described herein, in some embodiments, the present application provides a method for sequencing a polypeptide by acquiring data during a polypeptide degradation process and analyzing the data to determine portions of the data corresponding to amino acids sequentially exposed at the termini of the polypeptide during the degradation process. In some embodiments, the portions of the data include a series of signal pulses indicating the association of one or more amino acid recognition molecules with consecutive amino acids exposed at the termini of the polypeptide (e.g., during degradation). In some embodiments, the series of signal pulses corresponds to a series of reversible single-molecule bonding interactions at the termini of the polypeptide during the degradation process.

[0275] In some embodiments, polypeptide sequencing techniques described herein generate data indicating how a polypeptide interacts with a binding means (e.g., one or more amino acid recognition molecules) while the polypeptide is being degraded by a cleavage means (e.g., one or more cleavage reagents). As described above, the data may include a set of characteristic patterns corresponding to the terminal association events of the polypeptide between terminal cleavage events. In some embodiments, the sequencing method described herein involves contacting a single polypeptide molecule with a binding means and a cleavage means, wherein the binding means and the cleavage means are configured to achieve at least 10 association events before a cleavage event. In some embodiments, the means are configured to achieve at least 10 association events between two cleavage events.

[0276] As described herein, in some embodiments, multiple single-molecule sequencing reactions are carried out in parallel in an array of sample wells. In some embodiments, the array includes about 10,000 to about 1,000,000 sample wells. In some embodiments, the volume of the sample wells is about 10 -21 Liters ~ approximately 10 -15The volume can be in liters. Because the volume of the sample wells is small and only about one polypeptide can be present in the sample well at any given time, only detection of single-molecule events may be possible. Statistically, some sample wells may not contain single-molecule sequencing reactions, and some wells may contain multiple single polypeptide molecules. However, a considerable number of sample wells may each contain single-molecule reactions (e.g., at least 30% in some embodiments), and as a result, single-molecule analysis can be performed in parallel on a large number of sample wells. In some embodiments, the binding and cleaving means are configured to achieve at least 10 association events before the cleavage event in at least 10% (e.g., 10-50%, greater than 50%, 25-75%, at least 80%, or more) of the sample wells in which single-molecule reactions are occurring. In some embodiments, the binding and cleaving means are configured to achieve at least 10 association events before the cleavage event for at least 50% (e.g., greater than 50%, 50-75%, at least 80%, or more) of the amino acids of the polypeptide in the single-molecule reaction.

[0277] In some embodiments, the labeled polypeptide is immobilized and exposed to an excitation source. Aggregate luminescence from the labeled polypeptide can be detected, and in some embodiments, exposure to luminescence over time may result in a loss of the detected signal due to degradation of the luminescence label (e.g., degradation by photobleaching). In some embodiments, the labeled polypeptide contains a unique combination of selectively labeled amino acids that initially produce the detected signal. Degradation of the luminescence label over time results in a corresponding decrease in the detected signal of the photobleached labeled polypeptide. In some embodiments, the signal can be deconvoluted by analysis of one or more luminescence properties (e.g., signal deconvolution by luminescence lifetime analysis). In some embodiments, the unique combination of selectively labeled amino acids of the labeled polypeptide is pre-calculated by computer (e.g., based on known polypeptide sequences of the proteome) and empirically validated. In some embodiments, the detected amino acid label combination is compared with a database of known sequences of the organism's proteome to identify a specific polypeptide in the database corresponding to the labeled polypeptide.

[0278] In some embodiments, an optimal sample concentration is determined to perform sequencing reactions that maximize sampling in large-scale parallel analysis. In some embodiments, the concentration is selected so that a desired fraction (e.g., 30%) of the array's sample wells is occupied at any given time. While we do not wish to be constrained by theory, it is assumed that the same wells remain available for further analysis while polypeptides are bleached over a period of time. Due to diffusion, approximately 30% of the array's sample wells can be used for analysis every 3 minutes. As an example, with 1 million sample well tips, 6,000,000 polypeptides can be sampled per hour, or 24,000,000 polypeptides in 4 hours.

[0279] In some embodiments, the present application provides a method for sequencing a polypeptide by detecting the luminescence of a labeled polypeptide subjected to repeated cycles of terminal amino acid modification and cleavage. In some embodiments, the method generally proceeds as described herein, with respect to other sequencing methods by Edman degradation.

[0280] In some embodiments, the method includes (i) modifying the terminal amino acids of a labeled polypeptide. As described elsewhere herein, in some embodiments, the modification includes contacting the terminal amino acids with an isothiocyanate (e.g., PITC) to form isothiocyanate-modified terminal amino acids. In some embodiments, the isothiocyanate modification converts the terminal amino acids into a form more susceptible to removal by a cleavage reagent (e.g., a chemical or enzymatic cleavage reagent as described herein). Thus, in some embodiments, the method includes (ii) removing the modified terminal amino acids using chemical or enzymatic means as described elsewhere herein with respect to Edman degradation.

[0281] In some embodiments, the method involves repeating steps (i) to (ii) over multiple cycles, during which the luminescence of the labeled polypeptide is detected, and cleavage events corresponding to the removal of labeled amino acids from the terminus can be detected as a decrease in the detected signal. In some embodiments, the absence of a change in signal following step (ii) identifies an unknown type of amino acid. Thus, in some embodiments, partial sequence information can be determined by evaluating the signal detected following step (ii) during each consecutive round, by assigning an amino acid type based on identity determined based on the detected signal change, or by identifying the amino acid type as unknown based on the absence of a change in the detected signal.

[0282] In some embodiments, a method for sequencing a polypeptide according to the present application involves sequencing by sequential enzymatic cleavage of a labeled polypeptide. In some embodiments, the labeled polypeptide is degraded using a modified processive exopeptidase that sequentially cleaves terminal amino acids from one end to the other. Exopeptidases are described in detail elsewhere in this specification. In some embodiments, the labeled polypeptide is degraded by an immobilized processive exopeptidase. In some embodiments, the immobilized labeled polypeptide is degraded by a processive exopeptidase.

[0283] In some embodiments, the processing rate of the processive exopeptidase is known, and the timing between detected signal reductions can be used to calculate the number of unlabeled amino acids between each detection event. For example, if a 40-amino acid polypeptide is cleaved so that an amino acid is removed every second, a labeled polypeptide with three signals will initially show all three, then two, then one, and finally no signal. In this way, the order of the labeled amino acids can be determined. Thus, these methods can be used to determine partial sequence information, for example, for proteomics analysis based on polypeptide fragment sequencing.

[0284] In some embodiments, single-molecule polypeptide sequencing can be achieved using an ATP-based Forster resonance energy transfer (FRET) scheme (e.g., with one or more labeled cofactors). In some embodiments, cofactor-based FRET sequencing can be carried out using an immobilized ATP-dependent protease, donor-labeled ATP, and acceptor-labeled amino acids of the polypeptide substrate. In some embodiments, the amino acids can be labeled with acceptors, and one or more cofactors can be labeled with donors.

[0285] For example, in some embodiments, the extracted polypeptides are denatured, and cysteine ​​and lysine are labeled with fluorescent dyes. In some embodiments, an engineered version of a protein translocase (e.g., bacterial ClpX) is used to bind to individual substrate polypeptides, unfold them, and move them through its nanochannels. In some embodiments, the translocase is labeled with a donor dye, and FRET occurs between the donor on the translocase and two or more distinct acceptor dyes on the substrate as the substrate passes through the nanochannels. The order of the labeled amino acids can then be determined from the FRET signal. In some embodiments, one or more of the following non-limiting labeled ATP analogs shown in Table 3 may be used.

[0286] [Table 3-1]

[0287] [Table 3-2]

[0288] [Table 3-3]

[0289] C. Preparation of samples for sequencing Polypeptide samples (e.g., highly concentrated polypeptide samples) can be modified before sequencing.

[0290] In some embodiments, the N-terminal or C-terminal amino acids of the polypeptide are modified. In some embodiments, the ends of the polypeptide are modified with a portion that allows for immobilization to a surface (e.g., the surface of a sample well on a tip used for polypeptide analysis). In some embodiments, such a method includes modifying the ends of a labeled polypeptide to be analyzed according to the present application. In yet another embodiment, such a method includes modifying the ends of a protein or enzyme that degrades or moves the polypeptide substrate according to the present application.

[0291] In some embodiments, the carboxyl terminus of a polypeptide is modified in a manner comprising: (i) blocking a free carboxylate group of the polypeptide; (ii) denaturing the polypeptide (e.g., by thermal and / or chemical means); (iii) blocking a free thiol group of the polypeptide; (iv) digesting the polypeptide to produce at least one polypeptide fragment containing a free C-terminal carboxylate group; and (v) conjugating the functional moiety to the free C-terminal carboxylate group (e.g., chemically). In some embodiments, this method further comprises dialysis of the sample containing the polypeptide after (i) and before (ii).

[0292] In some embodiments, the carboxyl terminus of a polypeptide is modified in a manner comprising: (i) denaturing the polypeptide (e.g., by thermal and / or chemical means); (ii) blocking a free thiol group of the polypeptide; (iii) digesting the polypeptide to produce at least one polypeptide fragment containing a free C-terminal carboxylate group; (iv) blocking a free C-terminal carboxylate group to produce at least one polypeptide fragment containing a blocked C-terminal carboxylate group; and (v) conjugating a functional moiety (e.g., enzymatically) to the blocked C-terminal carboxylate group. In some embodiments, the method further comprises dialyzing the sample containing the polypeptide after (iv) and before (v).

[0293] In some embodiments, blocking free carboxylic acid groups refers to the chemical modification of these groups, which alters the chemical reactivity compared to the unmodified carboxylate salt. Suitable carboxylate blocking methods are known in the art and require the modification of the side-chain carboxylate groups to be chemically distinct from the carboxy-terminal carboxylate groups of the polypeptide to be functionalized. In some embodiments, blocking free carboxylic acid groups includes esterification or amidation of the free carboxylic acid groups of the polypeptide. In some embodiments, blocking free carboxylic acid groups includes, for example, methyl esterification of the free carboxylic acid groups of the polypeptide by reacting the polypeptide with methanolic HCl. Additional examples of reagents and techniques useful for blocking free carboxylate groups include, but are not limited to, 4-sulfo-2,3,5,6-tetrafluorophenol (STP) and / or carbodiimides, e.g., N-(3-dimethylaminopropyl)-N'-ethylcarbodiimide hydrochloride (EDAC), uronium reagents, diazomethane, alcohols and acids for Fischer esterification, the use of N-hydroxylsuccinimide (NHS) to form NHS esters (potentially as an intermediate for subsequent ester or amine formation), or any other methods for modifying or blocking carboxylic acids via reaction with carbonyldiimidazole (CDI) or formation of mixed anhydrides, or potentially via ester or amide formation.

[0294] In some embodiments, blocking free thiol groups refers to chemical modification of these groups, which alters their chemical reactivity compared to unmodified thiols. In some embodiments, blocking free thiol groups includes reduction and alkylation of free thiol groups in polypeptides. In some embodiments, reduction and alkylation are carried out by contacting the polypeptide with dithiothreitol (DTT) and one or both of iodoacetamide and iodoacetic acid. Examples of additional and alternative cysteine ​​reducing reagents that can be used are well known and not limited to, but include 2-mercaptoethanol, tris(2-carboxyethyl)phosphine hydrochloride (TCEP), tributylphosphine, dithiobutylamine (DTBA), or any reagent capable of reducing thiol groups. Examples of additional and alternative cysteine ​​blocking (e.g., cysteine ​​alkylation) reagents that can be used are well known and not limited to, but include acrylamide, 4-vinylpyridine, N-ethylmalemide (NEM), N-ε-maleimidocaproic acid (EMCA), or any reagent that modifies cysteine ​​to prevent the formation of disulfide bonds.

[0295] In some embodiments, digestion includes enzymatic digestion. In some embodiments, digestion is carried out by contacting the polypeptide with an endopeptidase (e.g., trypsin) under digestive conditions. In some embodiments, digestion includes chemical digestion. Examples of reagents suitable for chemical and enzymatic digestion are known in the art and are not limited to, but include trypsin, chemotrypsin, Lys-C, Arg-C, Asp-N, Lys-N, BNPS-skatole, CNBr, caspase, formic acid, glutamyl endopeptidase, hydroxylamine, iodosobenzoic acid, neutrophil elastase, pepsin, proline endopeptidase, proteinase K, staphylococcal peptidase I, thermolysin, and thrombin.

[0296] In some embodiments, the functional moiety comprises a biotin molecule. In some embodiments, the functional moiety comprises a reactive chemical moiety such as an alkynyl. In some embodiments, the bonding of the functional moiety comprises biotinylation of the carboxy-terminal carboxymethyl ester group by carboxypeptidase Y, as is known in the art.

[0297] In some embodiments, the solubilizing moiety is added to the polypeptide. Therefore, in some embodiments, the methods and compositions provided herein are useful for modifying the ends of polypeptides with moieties that increase their solubility. In some embodiments, the solubilizing moiety is useful for small polypeptides that are relatively insoluble, resulting from fragmentation (e.g., enzymatic fragmentation using trypsin). For example, in some embodiments, short polypeptides in a polypeptide pool can be solubilized by attaching a polymer (e.g., a short oligosaccharide, sugar, or other charged polymer) to the polypeptide.

[0298] D. Luminous Sign As used herein, luminescent labeling refers to a molecule capable of absorbing one or more photons and then emitting one or more photons after one or more time intervals. In some embodiments, this term is used interchangeably with “label” or “luminescent molecule” depending on the context. Luminescent labeling according to certain embodiments described herein may refer to luminescent labeling of label affinity reagents, luminescent labeling of labeled peptidases (e.g., labeled exopeptidases, labeled nonspecific exopeptidases), luminescent labeling of labeled peptides, luminescent labeling of labeled cofactors, or other labeling compositions described herein. In some embodiments, luminescent labeling according to this application refers to labeled amino acids in a labeled polypeptide comprising one or more labeled amino acids.

[0299] In some embodiments, the luminescent label may include first and second chromophores. In some embodiments, the excited state of the first chromophore can be relaxed via energy transfer to the second chromophore. In some embodiments, the energy transfer is Forster resonance energy transfer (FRET). Such a FRET pair may be useful in providing a luminescent label having properties that facilitate the distinction of the label from among several luminescent labels in a mixture. In yet another embodiment, the FRET pair includes a first chromophore of the first luminescent label and a second chromophore of the second luminescent label. In certain embodiments, the FRET pair may absorb excitation energy in a first spectral range and emit emission in a second spectral range.

[0300] In some embodiments, the luminescent label refers to a fluorophore or dye. Typically, the luminescent label comprises an aromatic or heteroaromatic compound and may be pyrene, anthracene, naphthalene, naphthylamine, acridine, stilbene, indole, benzindol, oxazole, carbazole, thiazole, benzothiazole, benzoxazole, phenanthidine, phenoxazine, porphyrin, quinoline, ethidium, benzamide, cyanine, carbocyanine, salicylate, anthranilate, coumarin, fluorescein, rhodamine, xanthene, or other similar compounds.

[0301] In some embodiments, the luminescent label comprises a dye selected from one or more of the following: 5 / 6-carboxyrhodamine 6G, 5-carboxyrhodamine 6G, 6-carboxyrhodamine 6G, 6-TAMRA, Abberior® STAR440SXP, Abberior® STAR470SXP, Abberior® STAR488, Abberior® STAR512, Abberior® STAR520SXP, Abberior® STAR580, Abberior® STAR600, Abberior® STAR635, Abberior® STAR635P, Abberior® STAR RED, AlexaFluor(TM) 350, AlexaFluor(TM) 405, AlexaFluor(TM) 430, AlexaFluor(TM) 480, AlexaFluor(TM) 488, AlexaFluor(TM) 514, AlexaFluor(TM) 532, Alexa Fluor(TM) 546, AlexaFluor(TM) 555, AlexaFluor(TM) 568, AlexaFluor(TM) 594, AlexaFluor(TM) 610-X, AlexaFluor(TM) 633, AlexaFluor(TM) 647, AlexaFluor(TM) )660, AlexaFluor(TM) 680, AlexaFluor(TM) 700, AlexaFluor(TM) 750, AlexaFluor(TM) 790, AMCA, ATTO390, ATTO425, ATTO465, ATTO488, ATTO495, ATTO514, ATTO5 20, ATTO532, ATTO542, ATTO550, ATTO565, ATTO590, ATTO610, ATTO620, ATTO633, ATTO647, ATTO647N, ATTO655, ATTO665, ATTO680, ATTO700, ATTO725, ATTO740, ATTO Oxa12, ATTORho101, ATTORho11, ATTORho12, ATTORho13, ATTORho14, ATTORho3B, ATTORho6G, ATTOThio12, BD Horizon(TM) V450, BODIPY(TM) 493 / 501, BODIPY(TM) 530 / 550,BODIPY (trademark) 558 / 568, BODIPY (trademark) 564 / 570, BODIPY (trademark) 576 / 589, BODIPY (trademark) 581 / 591, BODIPY (trademark) 630 / 650, BODIPY (trademark) 650 / 665, BODIPY (trademark) FL, BODIPY (trademark) FL-X, BODIPY (trademark) R6G, BODIPY (trademark) TMR, BODIPY (trademark) TR, CAL Fluor (trademark) Gold540, CAL Fluor (trademark) Green510, CAL Fluor (trademark) Orange560, CAL Fluor (trademark) Red590, CAL Fluor (trademark) Red610, CAL Fluor (trademark) Red615, CAL Fluor (trademark) Red635, Cascade (trademark) Blue, CF (trademark) 350, CF (trademark) 405M, CF (trademark) 405S, CF (trademark) 488A, CF (trademark) 514, CF (trademark) 532, CF (trademark) 543, CF (trademark) 546, CF (trademark) 555, CF (trademark) 568, CF (trademark) 594, CF (trademark) 620R, CF (trademark) 633, CF (trademark) 633-V1, CF (trademark) 640R, CF (trademark) 640R-V1, CF (trademark) 640R-V2, CF (trademark) 660C, CF (trademark) 660R, CF (trademark) 680, CF (trademark) 680R, CF (trademark) 680R-V1, CF (trademark) 750, CF (trademark) 770, CF (trademark) 790, Chromeo (trademark) 642, Chromis 425N, Chromis 500N, C Chromis 515N, Chromis 530N, Chromis 550A, Chromis 550C, Chromis 550Z, Chromis 560N, Chromis 570N, Chromis 577N, Chromis 600N, Chromis 630N, Chromis 645A, Chromis 645C, Chromis 645Z, Chromis 678A, Chromis 678C, Chromis 678Z, Chromis 770A, Chromis 770C, Chromis 800A, Chromis 800C, Chromis 830A, Chromis 830C, Cy(trademark) 3, Cy(trademark) 3.5, Cy(trademark) 3B, Cy(trademark) 5, Cy(trademark) 5.5, Cy(trademark) 7, DyLight(trademark) 350, DyLight(trademark) 405.DyLight (trademark) 415-Col, DyLight (trademark) 425Q, DyLight (trademark) 485-LS, DyLight (trademark) 488, DyLight (trademark) 504Q, DyLight (trademark) 510-LS, DyLight (trademark) 515-LS, DyLight (trademark) 521-LS, DyLight (trademark) 530-R2, DyLight (trademark) 543Q, DyLight (trademark) 550, DyLight (trademark) 554-R0, DyLight (trademark) 554-R1, DyLight (trademark) 590-R2, DyLight ( Trademarks) 594, DyLight 610-B1, DyLight 615-B2, DyLight 633, DyLight 633-B1, DyLight 633-B2, DyLight 650, DyLight 655-B1, DyLight 655-B2, DyLight 655-B3, DyLight 655-B4, DyLight 662Q, DyLight 675-B1, DyLight 675-B2, DyLight 675-B3 DyLight (trademark) 675-B4, DyLight (trademark) 679-C5, DyLight (trademark) 680, DyLight (trademark) 683Q, DyLight (trademark) 690-B1, DyLight (trademark) 690-B2, DyLight (trademark) 696Q, DyLight (trademark) 700-B1, DyLight (trademark) 700-B1, DyLight (trademark) 730-B1, DyLight (trademark) 730-B2, DyLight (trademark) 730-B3, DyLight (trademark) 730-B4, DyLight (trademark) 747, DyLight t(trademark)747-B1, DyLight(trademark)747-B2, DyLight(trademark)747-B3, DyLight(trademark)747-B4, DyLight(trademark)755, DyLight(trademark)766Q, DyLight(trademark)775-B2, DyLight(trademark)775-B3, DyLight(trademark)775-B4, DyLight(trademark)780-B1, DyLight(trademark)780-B2, DyLight(trademark)780-B3, DyLight(trademark)800, DyLight(trademark)830-B2, Dyomics-350,Dyomics-350XL、Dyomics-360XL、Dyomics-370XL、Dyomics-375XL、Dyomics-380XL、Dyomics-390XL、Dyomics-405、Dyomics-415、Dyomics-430、Dyomics-431、Dyomics-478、Dyomics-480XL、Dyomics-481XL、Dyomics-485XL、Dyomics-490、Dyomics-495、Dyomics-505、Dyomics-510XL、Dyomics-511XL、Dyomics-520XL、Dyomics-521XL、Dyomics-530、Dyomics-547、Dyomics-547Pl、Dyomics-548、Dyomics-549、Dyomics-549P1、Dyomics-550、Dyomics-554、Dyomics-555、Dyomics-556、Dyomics-560、Dyomics-590、Dyomics-591、Dyomics-594、Dyomics-601XL、Dyomics-605、Dyomics-610、Dyomics-615、Dyomics-630、Dyomics-631、Dyomics-632、Dyomics-633、Dyomics-634、Dyomics-635、Dyomics-636、Dyomics-647、Dyomics-647P1、Dyomics-648、Dyomics-648P1、Dyomics-649、Dyomics-649P1、Dyomics-650、Dyomics-651、Dyomics-652、Dyomics-654、Dyomics-675、Dyomics-676、Dyomics-677、Dyomics-678、Dyomics-679P1、Dyomics-680、Dyomics-681、Dyomics-682、Dyomics-700、Dyomics-701、Dyomics-703、Dyomics-704、Dyomics-730、Dyomics-731、Dyomics-732、Dyomics-734、Dyomics-749、Dyomics-749P1、Dyomics-750、Dyomics-751、Dyomics-752、Dyomics-754、Dyomics-776、Dyomics-777, Dyomics-778, Dyomics-780, Dyomics-781, Dyomics-782, Dyomics-800, Dyomics-831, eFluor (trademark) 450, EoShin, FITC, Full Oreshine, HiLyte (trademark) Fluor405, HiLyte (trademark) Fluor488, HiLyte (trademark) Fluor532, HiLyte (trademark) Fluor555, HiLyte (trademark) Fluor594, HiLyte (trademark) Fluo r647, HiLyte Fluor680, HiLyte Fluor750, IRDye 680LT, IRDye 750, IRDye 800CW, JOE, LightCycler 640R, LightCycler Red610, LightCycler Red640, LightCycler Red670, LightCycler Red705, Risamin Roadmin B, Naftful Oresine, Oregon Green 488, Oregon Green 514, Pacific Blue, Pacific Green, Pacific Orange (trademark), PET, PF350, PF405, PF415, PF488, PF505, PF532, PF546, PF555P, PF568, PF594, PF610, PF633P, PF647P, Quasar (trademark) 570, Quasar (trademark) 670, Quasar (trademark) 705, ローダミン123, ローダミン6G, ローダミングリーン, ローダミングリーン-X, ローダミンレッド, ROX, Seta (trademark) 375, Seta (trademark) 470, Seta (trademark) 555, Se Seta (trademark) 632, Seta (trademark) 633, Seta (trademark) 650, Seta (trademark) 660, Seta (trademark) 670, Seta (trademark) 680, Seta (trademark) 700, Seta (trademark) 750, Seta (trademark) 780, Seta (trademark) APC-780, Seta (trademark) PerCP-680, Seta (trademark) R-PE-670, Seta (trademark) 646, SeTau380, SeTau425, SeTau647, SeTau405, Square635, Square650, Square660,Square672, Square680, Sulfolodamine 101, TAMRA, TET, Texas Red (trademark), TMR, TRITC, Yakima Yellow (trademark), Zenon (trademark), Zy3, Zy5, Zy5.5, and Zy7.

[0302] E. Luminescence In some embodiments, the present application relates to polypeptide sequencing and / or identification based on one or more luminescence properties of a luminescent label.

[0303] In some embodiments, luminescent labels are identified based on luminescence lifetime, luminescence intensity, brightness, absorption spectrum, emission spectrum, emission quantum yield, or a combination of two or more thereof. In some embodiments, multiple types of luminescent labels can be distinguished from one another based on different luminescence lifetimes, luminescence intensity, brightness, absorption spectrum, emission spectrum, emission quantum yield, or a combination of two or more thereof. Identification may mean assigning the exact identity and / or quantity of one type of amino acid (e.g., a single type or a subset of types) associated with the luminescent label, and may also mean assigning the amino acid position in the polypeptide in comparison to other types of amino acids.

[0304] In some embodiments, luminescence is detected by exposing a luminescent label to a series of separate light pulses and evaluating the timing or other characteristics of each photon emitted from the label. In some embodiments, information from a plurality of photons emitted sequentially from the label is aggregated and evaluated to identify the label, thereby identifying the associated type of amino acid. In some embodiments, the luminescence lifetime of the label is determined from a plurality of photons emitted sequentially from the label, and this luminescence lifetime can be used to identify the label. In some embodiments, the luminescence intensity of the label is determined from a plurality of photons emitted sequentially from the label, and this luminescence intensity can be used to identify the label. In some embodiments, the luminescence lifetime and luminescence intensity of the label are determined from a plurality of photons emitted sequentially from the label, and the label can be identified using this luminescence lifetime and luminescence intensity.

[0305] In some embodiments of the present application, a single polypeptide molecule is exposed to a plurality of distinct light pulses, and a series of emitted photons are detected and analyzed. In some embodiments, the series of emitted photons provide information about a single polypeptide molecule that is present but does not change in the reaction sample over time. However, in some embodiments, the series of emitted photons provide information about a series of different molecules that are present in the reaction sample at different times (e.g., as the reaction or process progresses). As an example, but not limited to, such information may be used to sequence and / or identify polypeptides that undergo chemical or enzymatic degradation in accordance with the present application.

[0306] In certain embodiments, a light-emitting label absorbs one photon and emits one photon after a certain time interval. In some embodiments, the light-emitting lifetime of the label can be determined or estimated by measuring that time interval. In some embodiments, the light-emitting lifetime of the label can be determined or estimated by measuring multiple time intervals for multiple pulse and emission events. In some embodiments, the light-emitting lifetime of the label can be distinguished among multiple types of labels by measuring the above time intervals. In some embodiments, the light-emitting lifetime of the label can be distinguished among multiple types of labels by measuring multiple time intervals for multiple pulse and emission events. In certain embodiments, a label is identified or distinguished among multiple types of labels by determining or estimating the light-emitting lifetime of the label. In certain embodiments, a label is identified or distinguished among multiple types of labels by distinguishing the light-emitting lifetime of the label among multiple types of labels.

[0307] The luminescence lifetime of a luminescent label can be determined using any suitable method (e.g., by measuring the lifetime using a suitable technique, or by determining the time-dependent characteristics of the emission). In some embodiments, determining the luminescence lifetime of a label includes determining the lifetime relative to another label. In some embodiments, determining the luminescence lifetime of a label includes determining the lifetime relative to a reference. In some embodiments, determining the luminescence lifetime of a label includes measuring the lifetime (e.g., fluorescence lifetime). In some embodiments, determining the luminescence lifetime of a label includes determining one or more temporal characteristics that indicate lifetime. In some embodiments, the luminescence lifetime of a label can be determined based on the distribution of multiple emission events (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, or more emission events) occurring over one or more time-gate windows relative to an excitation pulse. For example, the luminescence lifetime of a label can be distinguished from multiple labels having different luminescence lifetimes based on the distribution of photon arrival times measured with respect to the excitation pulse.

[0308] It should be understood that the luminescence lifetime of a luminescent label indicates the timing of photons emitted after the label reaches an excited state, and that labels can be distinguished by information indicating the timing of these photons. Some embodiments may include distinguishing a label from multiple labels based on its luminescence lifetime by measuring the time associated with the photons emitted by the label. The time distribution may provide an index of the luminescence lifetime that can be determined from the above distribution. In some embodiments, a label can be distinguished from multiple labels based on its time distribution, for example, by comparing the time distribution with a reference distribution corresponding to a known label. In some embodiments, the value of the luminescence lifetime is determined from the time distribution.

[0309] As used herein, in some embodiments, luminescence intensity refers to the number of photons emitted per unit time by a luminescent label excited by the delivery of pulsed excitation energy. In some embodiments, luminescence intensity refers to the number of detected emitted photons per unit time emitted by a label excited by the delivery of pulsed excitation energy and detected by a particular sensor or set of sensors.

[0310] Where used herein, in some embodiments, brightness refers to a parameter that reports the average luminescence intensity per luminescent label. Thus, in some embodiments, “luminescence intensity” may be used to refer, in most cases, to the brightness of a composition containing one or more labels. In some embodiments, the brightness of a label is equal to the product of its quantum yield and extinction coefficient.

[0311] Where used herein, in some embodiments, the emission quantum yield refers to the proportion of excitation events within a given wavelength or spectral range that lead to emission events, and is typically less than 1. In some embodiments, the emission quantum yields of the emission labels described herein are 0 to about 0.001, about 0.001 to about 0.01, about 0.01 to about 0.1, about 0.1 to about 0.5, about 0.5 to 0.9, or about 0.9 to 1. In some embodiments, the labels are identified by determining or estimating the emission quantum yield.

[0312] As used herein, in some embodiments, the excitation energy is a pulse of light from a light source. In some embodiments, the excitation energy is in the visible spectrum. In some embodiments, the excitation energy is in the ultraviolet spectrum. In some embodiments, the excitation energy is in the infrared spectrum. In some embodiments, the excitation energy is at or near the absorption maximum of a light emission label that emits multiple photons to be detected. In certain embodiments, the excitation energy is in the range of about 500 nm to about 700 nm (e.g., about 500 nm to about 600 nm, about 600 nm to about 700 nm, about 500 nm to about 550 nm, about 550 nm to about 600 nm, about 600 nm to about 650 nm, or about 650 nm to about 700 nm). In certain embodiments, the excitation energy may be monochromatic or limited to a spectral range. In some embodiments, the spectral range has the range of about 0.1 nm to about 1 nm, about 1 nm to about 2 nm, or about 2 nm to about 5 nm. In some embodiments, the spectral range is in the range of about 5 nm to about 10 nm, about 10 nm to about 50 nm, or about 50 nm to about 100 nm.

[0313] V. Kits for sample preparation In some embodiments, the disclosure relates to a kit for preparing polypeptide samples (e.g., multiplexed samples) for sequencing. The kit may be sufficient to prepare one or more polypeptide samples (e.g., multiplexed samples) for sequencing. In some embodiments, the kit is sufficient to prepare a single polypeptide sample. In other embodiments, the kit is sufficient to prepare at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 20, at least 25, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, or at least 100 polypeptide samples.

[0314] In some embodiments, the kit includes a barcode component containing multiple barcode molecules, as described herein. See "Method for Preparing Multiplexed Samples." In some embodiments, the kit includes one or more detection molecules, as described herein. See "Method for Preparing Multiplexed Samples." In some embodiments, the kit includes a solid support that allows for the physical separation of a population of polypeptides of different origins, as described herein. See "Method for Preparing Multiplexed Samples." In some embodiments, the kit includes a concentration component containing multiple concentration molecules, as described herein. See "Method for Concentrating Polypeptides." In some embodiments, the kit includes a modifier, as described herein. See "Method for Concentrating Polypeptides." In some embodiments, the kit includes an affinity reagent, as described herein. See "Polypeptide Sequence Decoding Methodology." In some embodiments, the kit includes a labeled peptidase, as described herein. See "Polypeptide Sequence Decoding Methodology."

[0315] The kit may be specific to one or more organisms (e.g., one or more unicellular and / or multicellular organisms). In some embodiments, the kit includes components (e.g., barcode molecules, detection molecules, concentration molecules, or combinations thereof) that modify, bind, or are bound to polypeptides of one or more organisms. For example, in some embodiments, the kit includes components that modify, bind, or are bound to one or more known polypeptides in the human proteome.

[0316] In some embodiments, the kit is specific to one or more diseases or conditions. For example, the kit may be a tumor kit, a cardiac disease kit, a genetic disease kit, or a combination thereof. Tumor kits include ABL1, ABL2, ACSL3, ACVR2A, ADAMTS20, ADGRA2, ADGRB3, ADGRL3, AFF1, AFF3, AKAP9, AKT1, AKT2, AKT3, ALK, AMER1, APC, AR, ARID1A, ARID2, ARNT, ASXL1, ATF1, ATM, ATR, ATRX, AURKA, AURKB, AURKC, AXL, BAP1, BCL10, BCL11A, BCL11B, BCL2, BCL2L1, BCL2L2, BCL3, BCL6, BCL7A, BCL9, BCR, BIRC2, BIRC3, BIRC5, BLM, BLNK, BMPR1A, BRAF, BRCA1, BRCA2, BRD3, BRIP1, BTK, BUB1B, CACNA1D, CARD11, CASC5, CASP8, CBFA2T3, CBFB, CBL, CCND1, CCND2, CCNE1, CD79A, CD79B, CDC73, CDH1, CDH11, CDH2, CDH20, CDH5, CDK12, CDK4, CDK6, CDK8, CDKN2A, CDKN2B , CDKN2C, CEBPA, CHEK1, CHEK2, CIC, CKS1B, CMPK1, COL1A1, CRBN, CREB1, CREBBP, CRKL, CRLF2, CRTC1, CSF1R, CSMD3, CTNNA1, CTNNB1, CYL D, CYP2C19, CYP2D6, DAXX, DCC, DDB2, DDIT3, DDR2, DEK, DICER1, DNMT3A, DPYD, DST, EGFR, EML4, EP300, EP400, EPHA3, EPHA7, EPHB1, EPHB4 , EPHB6, ERBB2, ERBB3, ERBB4, ERCC1, ERCC2, ERCC3, ERCC4, ERCC5, ERG, ESR1, ETS1, ETV1, ETV4, EXT1, EXT2, EZH2, FANCA, FANCC, FANCD2, F ANCF, FANCG, FAS, FBXW7, FCGR2B, FGFR1, FGFR2, FGFR3, FGFR4, FH, FLCN, FLI1, FLT1, FLT3, FLT4, FN1, FOXA1, FOXL2, FOXO1, FOXO3, FOXP1,FOXP4、FZR1、G6PD、GATA1、GATA2、GATA3、GDNF、GNA11、GNAQ、GNAS、GPC3、GRM8、GUCY1A2、HCAR1、HEY1、HIF1A、HIST1H3B、HLF、HMGA1、HNF1A、HOOK3、HOXA13、HOXD11、HRAS、HSP90AA1、HSP90AB1、ICK、IDH1、IDH2、IGF1R、IGF2、IGF2R、IKBKB、IKBKE、IKZF1、IL2、IL21R、IL6ST、IL7R、ING4、IRF4、IRS2、ITGA10、ITGA9、ITGB2、ITGB3、JAK1、JAK2、JAK3、JUN、KAT6A、KAT6B、KDM5C、KDM6A、KDR、KEAP1、KIAA1549、KIT、KLF6、KMT2A、KMT2C、KMT2D、KRAS、LAMP1、LCK、LIFR、LPP、LRP1B、LTF、LTK、MAF、MAFB、MAGEA1、MAGI1、MALT1、MAML2、MAP2K1、MAP2K2、MAP2K4、MAP3K7、MAPK1、MAPK8、MARK1、MARK4、MBD1、MCL1、MDM2、MDM4、MEN1、MET、MITF、MLH1、MLLT10、MLLT4、MLLT6、MMP2、MN1、MPL、MRE11A、MSH2、MSH6、MTCP1、MTOR、MTR、MTRR、MUC1、MUTYH、MYB、MYC、MYCL、MYCN、MYD88、MYH11、MYH9、NBN、NCOA1、NCOA2、NCOA4、NF1、NF2、NFE2L2、NFKB1、NFKB2、NIN、NKX2-1、NLRP1、NOTCH1、NOTCH2、NOTCH4、NPM1、NR4A3、NRAS、NSD1、NTRK1、NTRK3、NUMA1、NUP214、NUP98、NUTM2A、NUTM2B、OMD、P2RY8、PAK3、PALB2、PARP1、PAX3、PAX5、PAX7、PAX8、PBRM1、PBX1、PDE4DIP、PDGFB、PDGFRA、PDGFRB、PER1、PGAP3、PHOX2B、PIK3C2B、PIK3CA、PIK3CB、PIK3CD、PIK3CG、PIK3R1、PIK3R2、PIM1、PKHD1、PLAG1、PLCG1、PLEKHG5、PML、PMS1、PMS2、POT1、POU5F1、PPARG, PPP2R1A, PRDM1, PRKAR1A, PRKDC, PSIP1, PTCH1, PTEN, PTGS2, PTPN11, PTPRD, PTPRT, RAD50, RAF 1, RALGDS, RAP1GDS1, RARA, RB1, RECQL4, REL, RET, RHOH, RNASEL, RNF2, RNF213, ROS1, RPS6KA2, RRM1, R UNX1, RUNX1T1, SAMD9, SBDS, SDHA, SDHB, SDHC, SDHD, SET, SETBP1, SETD2, SF3B1, SGK1, SH2D1A, SH3GL1 , SMAD2, SMAD4, SMARCA4, SMARCB1, SMO, SMUG1, SOCS1, SOX11, SOX2, SRC, SSX1, SSX2, SSX4, STAT5B, STK1 1, STK36, SUFU, SYK, SYNE1, TAF1, TAF1L, TAL1, TBL1XR1, TBX22, TCF12, TCF3, TCF7L1, TCF7L2, TCL1A, T ERT, TET1, TET2, TFE3, TGFBR2, TGM7, THBS1, TIMP3, TLR4, TLX1, TMPRSS2, TNFAIP3, TNFRSF14, TNK2, TOP 1. May contain high-concentration molecules that bind to (or are bound to) TP53, TPR, TRIM24, TRIM33, TRIP11, TRRAP, TSC1, TSC2, TSHR, TTL, UBR5, UGT1A1, USP9X, VHL, WAS, WHSC1, WRN, WT1, XPA, XPC, XPO1, XRCC2, ZNF384, ZNF521, or any combination thereof.

[0317] The kit for heart disease includes ABCC9, ABCG5, ABCG8, ACTA1, ACTA2, ACTC1, ACTN2, AKAP9, ALMS1, ANK2, ANKRD1, APOA4, APOA5, APOB, APOC2, APOE, BAG3, BRAF, CACNA1C, CACNA2D1, CACNB2, CALM1, CALR3, CASQ2, CAV3, CBL, CBS, CETP, COL3A1, COL5A1, COL5A2, COX15, CREB3L3, CRELD1, CRYAB, CSRP3, CTF1, DES, DMD, DNAJC19, DOLK, DPP6, DSC2, DSG2, DSP, DTNA, EFEMP2, ELN, EMD, EYA4, FBN1, FBN2, FHL1, FHL2, FKRP, FKTN, FXN, GAA, GATAD1, GCKR, GJA5, GLA, GPD1L, GPIHBP1, HADHA, HCN4, HFE, HRAS, HSPB8, ILK, JAG1, JPH2, JUP, KCNA5, KCND3, KCNE1, KCNE2, KCNE3, KCNH2, KCNJ2, KCNJ5, KCNJ8, KCNQ1, KLF10, KRAS, LAMA2, LAMA4, LAMP2, LDB3, LDLR, LDLRAP1, LMF1, LMNA, LPL, LTBP2, MAP2K1, MAP2K2, MIB1, MURC, MYBPC3, MYH11, MYH6, MYH7, MYL2, MYL3, MYLK, MYLK2, MYO6, MYOZ2, MYPN, NEXN, NKX2-5, NODAL, NOTCH1, NPPA, NRAS, PCSK9, PDLIM3, PKP2, PLN, PRDM16, PRKAG2, PRKAR1A, PTPN11, RAF1, RANGRF, RBM20, RYR1, RYR2, SALL4, SCN1B, SCN2B, SCN3B, SCN4B, SCN5A, SCO2, SDHA, SEPN1, SGCB, SGCD, SGCG, SHOC2, SLC25A4, SLC2A10, SMAD3, SMAD4, SNTA1, SOS1, SREBF2, TAZ, TBX20, TBX3, TBX5, TCAP, TGFB2, TGFB3, TGFBR1, TGFBR2, TMEM43, TMPO, TNNC1, TNNI3, TNNT2, TPM1, TRDN, TRIM63, TRPM4, TTN, TTR, TXNRD2, VCL, ZBTB17, ZHX3,It may also contain high-concentration molecules that bind to (or are bound to) ZIC3 or any combination thereof.

[0318] The kit for genetic diseases includes ABCA4, ABCC9, ABCD1, ACADVL, ACTA2, ACTC1, ACTN2, ADA, AIPL1, AIRE, AKAP9, ALPL, AMT, ANK2, APC, APP, APTX, ARL6, ARSA, ASL, ASPA, ATL1, ATM, ATP2A2, ATP7A, ATP7B, ATXN1, ATXN2, ATXN7, BAG3, BCKDHA, BCKDHB, BEST1, BMPR1A, BTD, BTK, CA4, CACNA1C, CACNB2, CALR3, CAPN3, CASQ2, CAV3, CCDC39, CCDC40, CDH23, CEP290, CERKL, CFTR, CHAT, CHD7, CHEK2, CHM, CHRNA1, CHRNB1, CHRND, CHRNE, CLCN1, CNGB1, COL11A1, COL11A2, COL1A1, COL1A2, COL2A1, COL3A1, COL4A1, COL4A5, COL5A1, COL5A2, COL7A1, COL9A1, CRB1, CRX, CTDP1, CTNS, CYP27A1, DBT, DCX, DES, DHCR7, DKC1, DLD, DMD, DNAH11, DNAH5, DNAH9, DNAI1, DNAI2, DNM2, DOK7, DSC2, DSG2, DSP, DYSF, ELN, EMD, ENG, EXT1, EYA1, EYS, F8, F9, FANCA, FANCC, FANCF, FANCG, FBN1, FBXO7, FGFR1, FGFR3, FMO3, FOXL2, FRG1, FRMD7, FSCN2, FXN, GAA, GALT, GATA4, GBA, GBE1, GCSH, GDF5, GJB2, GJB3, GJB6, GLA, GLDC, GNE, GNPTAB, GPC3, GPD1L, GPR143, GUCY2D, HBA2, HBB, HCN4, HEXA, HFE, HIBCH, HMBS, HR, IDS, IDUA, IKBKAP, IL2RG, IMPDH1, ITGB4, JAG1, JUP, KCNE1, KCNE2, KCNE3, KCNH2, KCNJ2, KCNQ1, KCNQ4, KIAA0196, KLHL7, KRAS, KRT14, KRT5, L1CAM, LAMB3, LAMP2, LDB3, LMNA, LRAT, LRRK2, MAPT, MC1R, MECP2, MED12, MEN1, MERTK, MFN2, MLH1, MMAA, MMAB,MMACHC, MPZ, MSH2, MTM1, MUT, MYBPC3, MYH11, MYH6, MYH7, MYL2, MYL3, MYLK, MYO7A, MYOZ2, NF1, NF2, NIPBL, NKX2-5, NME8, NPC1, NP C2, NR2E3, NRAS, NSD1, OCA2, OCRL, OTC, PABPN1, PAFAH1B1, PAH, PAX3, PAX6, PCDH15, PEX1, PEX10, PEX13, PEX14, PEX19, PEX26, PEX 3, PEX5, PINK1, PKD1, PKD2, PKHD1, PKP2, PLEC, PLN, PLOD1, PMM2, PMP22, POLG, PPT1, PRCD, PRKAG2, PROM1, PRPF31, PRPF8, PRPH2, P SEN1, PSEN2, PTCH1, PTPN11, RAF1, RAG1, RAG2, RAI1, RAPSN, RB1, RDH12, RET, RHO, ROR2, RP9, RPE65, RPGR, RPGRIP1, RPL11, RPL35A, RPS10, RPS19, RPS24, RPS26, RPS6KA3, RPS7, RS1, RSPH4A, RSPH9, RYR1, RYR2, SALL4, SCN1B, SCN3B, SCN4B, SCN5A, SCN9A, SEMA4A, S ERPINA1, SERPING1, SGCD, SH3BP2, SIX1, SIX5, SLC25A13, SLC25A4, SLC26A4, SMAD3, SMAD4, SNCA, SNRNP200, SNTA1, SOD1, SOS1, SOX 9. May contain high-concentration molecules that bind to (or are bound to) SPATA7, SPG7, STARD3, TAF1, TAZ, TBX5, TCOF1, TGFBR1, TGFBR2, TMEM43, TNNC1, TNNI3, TNNT1, TNNT2, TNXB, TOPORS, TP53, TPM1, TSC1, TSC2, TTPA, TTR, TULP1, TWIST1, TYR, USH1C, USH2A, VCL, VHL, WAS, WRN, WT1, or any combination thereof.

[0319] In some embodiments, at least one component of the kit is provided in a dried or freeze-dried form. In other embodiments, at least one component of the kit is provided in a solubilized form.

[0320] The kits provided herein are in appropriate packaging. Appropriate packaging includes, but is not limited to, vials, bottles, jars, and flexible packaging. Packaging for use in combination with specific devices is also intended. See "Devices for Sample Preparation and Sample Sequence Determination." Kits may have a sterile access port (for example, the container may be a vial with a stopper that can be pierced by an intravenous solution bag or a subcutaneous needle). Containers may have a sterile access port.

[0321] The kit may optionally provide additional components such as buffers and interpretation information. In some embodiments, the kit further includes at least one buffer. Buffers suitable for the methods described herein have already been described. In some embodiments, the kit may further include instructions for use in any of the methods described herein.

[0322] In some embodiments, the present disclosure provides a product comprising the contents of the above-described kit. VI. Devices for sample preparation and sample sequencing In some embodiments, the disclosure relates to a device for sample preparation and / or sample sequencing. In some embodiments, the device includes a sample preparation module. In some embodiments, the device includes a sample sequencing module. In some embodiments, the device includes both a sample preparation module and a sample sequencing module.

[0323] A. Devices for sample preparation Devices, including apparatus, cartridges (e.g., those containing channels (e.g., microfluidic channels)), and / or pumps (e.g., peristaltic pumps), are provided for use in the process of preparing samples for analysis. Devices can be used in accordance with this disclosure to enable the concentration, enrichment, manipulation, and / or detection of target molecules from biological samples. In some embodiments, devices and associated methods are provided for automated processing of samples to produce materials for next-generation sequencing and / or other downstream analytical techniques. Devices and associated methods may be used to carry out chemical and / or biological reactions, including reactions for nucleic acid and / or polypeptide processing according to sample preparation or sample analysis processes described elsewhere in this specification.

[0324] In some embodiments, the sample preparation device is positioned to deliver or transfer a target molecule or sample, comprising multiple molecules (e.g., target nucleic acids or target polypeptides), to a sequencing module or device. In some embodiments, the sample preparation device is directly connected to the sequencing device (e.g., physically connected) or indirectly connected to the sequencing device.

[0325] In some embodiments, the device includes a sequence preparation module configured to accept one or more cartridges. In some embodiments, the cartridge includes one or more reservoirs or reaction vessels configured to receive fluid and / or to contain one or more reagents used in the sample preparation process. In some embodiments, the cartridge includes one or more channels (e.g., microfluidic channels) configured to contain and / or transport fluid (e.g., fluid containing one or more reagents) used in the sample preparation process. Reagents include buffers, enzyme reagents, polymer matrices, barcode components (e.g., barcode molecules), detection molecules, concentration molecules, capture reagents, size-specific selection reagents, sequence-specific selection reagents, and / or purification reagents. Additional reagents for use in the sample preparation process are described elsewhere in this specification.

[0326] In some embodiments, the cartridge contains one or more stored reagents (e.g., liquids or lyophilized reagents suitable for reconstitution into liquid form). The stored reagents in the cartridge contain reagents suitable for performing a desired process and / or reagents suitable for processing a desired sample type. In some embodiments, the cartridge is a single-use cartridge (e.g., a disposable cartridge) or a multi-use cartridge (e.g., a reusable cartridge). In some embodiments, the cartridge is configured to receive user-provided samples. User-provided samples may be added to the cartridge before or after the cartridge is received by the device, for example, manually by the user or in an automated process.

[0327] In some embodiments, the device can facilitate the preparation of multiplexed samples in the process according to this disclosure. See "Method for Preparing Multiplexed Samples".

[0328] In some embodiments, the device can facilitate the concentration of a target molecule in the process according to this disclosure. See "Method for Concentrating Polypeptides." Thus, the device enables the concentration of a target polypeptide in a highly multiplexed manner by utilizing molecules.

[0329] In some embodiments, the sample is concentrated for the target molecule using electrophoretic methods. In some embodiments, the sample is concentrated for the target molecule using affinity SCODA. In some embodiments, the sample is concentrated for the target molecule using field inversion gel electrophoresis (FIGE). In some embodiments, the sample is concentrated for the target molecule using pulsed-field gel electrophoresis (PFGE).

[0330] In some embodiments, the device includes a sample preparation module comprising a matrix (e.g., a porous medium, an electrophoretic polymer gel) used during concentration, which contains immobilized capture probes that bind (directly or indirectly) to target molecules present in the sample. In some embodiments, the matrix used during concentration comprises 1, 2, 3, 4, 5, or more unique immobilized capture probes, each of which binds to a unique target molecule and / or to the same target molecule with different binding affinities.

[0331] In some embodiments, the immobilized capture probe is a polypeptide capture probe that binds to a target polypeptide or polypeptide fragment. For example, in some embodiments, the immobilized capture probe is a high-concentration molecule as described herein.

[0332] In some embodiments, the polypeptide capture probe targets the target polypeptide (or polypeptide fragment) with 10 -9 ~10 -8 M, 10 -8 ~10-7 M, 10 -7 ~10 -6 M, 10 -6 ~10 -5 M, 10 -5 ~10 -4 M, 10 -4 ~10 -3 M, or 10 -3 ~10 -2 It binds with a binding affinity of M. In some embodiments, the binding affinity is in the range of picomoles to nanomoles (e.g., about 10). -12 ~about 10 -9 M) In some embodiments, the binding affinity is in the range of nanomoles to micromoles (e.g., about 10 -9 ~about 10 -6 M) is the binding affinity. In some embodiments, the binding affinity is in the range of micromoles to millimoles (e.g., about 10). -6 ~about 10 -3 M) is the binding affinity. In some embodiments, the binding affinity is in the range of picomoles to micromoles (e.g., about 10). -12 ~about 10 -6 M) is the binding affinity. In some embodiments, the binding affinity is in the range of nanomoles to millimoles (e.g., about 10). -9 ~about 10 -3 M) is the answer.

[0333] In some embodiments, the immobilized capture probe is an oligonucleotide capture probe that hybridizes to the target nucleic acid. In some embodiments, the oligonucleotide capture probe is at least 50%, 60%, 70%, 80%, 90%, 95%, or 100% complementary to the target nucleic acid. In some embodiments, a single oligonucleotide capture probe can be used to concentrate multiple related target nucleic acids (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, or more related target nucleic acids) that share at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% sequence identity. Concentration of multiple related target nucleic acids may enable the creation of a metagenomic library. In some embodiments, the oligonucleotide capture probe may enable differential concentration of related target nucleic acids. In some embodiments, the oligonucleotide capture probe may enable concentration of a target nucleic acid compared to nucleic acids of the same sequence with different modification states (e.g., methylation state, acetylation state).

[0334] In some embodiments, for the purpose of concentrating nucleic acid target molecules with a length of 0.5 to 2 kilobases, oligonucleotide capture probes can be covalently immobilized to an acrylamide matrix using the 5' acridite moiety. In some embodiments, for the purpose of concentrating larger nucleic acid target molecules (e.g., those with a length exceeding 2 kilobases), oligonucleotide capture probes can be immobilized to an agarose matrix. In some embodiments, oligonucleotide capture probes can be immobilized to an agarose matrix using thiol epoxide chemistry (e.g., by thiol-modified oligonucleotides covalently bonded to crosslinked agarose beads). Oligonucleotide capture probes bound to agarose beads can be combined and solidified within a standard agarose matrix (e.g., with the same agarose percentage).

[0335] In some embodiments, multiple capture probes (e.g., a group of multiple capture probe types that bind to deterministic target molecules of infectious agents such as adenovirus, staphylococcus, pneumonia, or tuberculosis) can be immobilized on a concentration matrix. Applying a sample to the concentration matrix along with multiple deterministic capture probes may lead to a diagnosis of a disease or condition (e.g., the presence of an infectious agent).

[0336] In some embodiments, the device may facilitate the release of target molecules from the concentration matrix after the removal of non-target molecules in the process according to this disclosure. In some embodiments, target molecules may be released from the concentration matrix by increasing the temperature of the concentration matrix. Adjusting the matrix temperature further affects the migration rate. This is because higher temperatures increase the stringency of the capture probe and the binding affinity between the target molecule and the capture probe. In some embodiments, when concentrating a relevant target molecule, the matrix temperature may be increased stepwise to release and separate the target molecule in increasing homology steps. This may enable sequencing of the target polypeptide or target nucleic acid that gradually deviates in relation to the initial reference target molecule, and may enable the discovery of novel proteins (e.g., enzymes) or functions (e.g., enzyme function or gene function). In some embodiments, when using multiple capture probes (e.g., multiple deterministic capture probes), the matrix temperature may be increased stepwise or gradiently, which may enable temperature-dependent release of different target molecules and result in the generation of a series of barcoded release bands indicating the presence or absence of control and target molecules.

[0337] The devices according to this disclosure generally include mechanical and electronic and / or optical components that can be used to operate a cartridge as described herein. In some embodiments, the device components operate to achieve and maintain a specific temperature in or in a specific area of ​​the cartridge. In some embodiments, the device components operate to apply a specific voltage to an electrode in the cartridge for a specific length of time. In some embodiments, the device components operate to move liquid to, from, or between a reservoir and / or reaction vessel in the cartridge. In some embodiments, the device components operate to move liquid through channels in the cartridge, for example, to, from, or between a reservoir and / or reaction vessel in the cartridge. In some embodiments, the device components move liquid via a peristaltic pump mechanism (e.g., a device) that interacts with the reagent-specific reservoir or reaction vessel of the cartridge's elastomer. In some embodiments, the device components move the liquid via a peristaltic pump mechanism (e.g., a device) configured to interact with an elastomer component (e.g., an elastomer-containing surface layer) associated with the channel of the cartridge to pump the fluid through the channel. The device components may include computer resources to drive a user interface, for example, in which sample information can be entered, a specific process can be selected, and the results of the execution can be reported.

[0338] The following non-limiting examples are intended to illustrate aspects of the devices, methods, and compositions described herein. Use of the sample preparation device according to this disclosure may involve performing one or more of the steps described below. The user can open the lid of the device and insert a cartridge that supports the desired process. The user can then add a sample, which can be combined with a specific dissolution solution, to the sample port of the cartridge. The user can then close the lid of the device, enter sample-specific information via the device's touchscreen interface, select any process-specific parameters (e.g., a range of size selection for the desired size, a desired degree of homology for target molecule capture, etc.), and begin executing the sample preparation process.

[0339] After execution, the user can receive relevant execution data (e.g., confirmation of successful completion of the execution, execution-specific metrics, etc.) and process-specific information (e.g., the amount of sample generated, the presence or absence of a specific target sequence, etc.). The data generated by the execution can be used for subsequent bioinformatics analysis, which may be local or cloud-based. Depending on the process, the completed sample can be extracted from the cartridge and used later (e.g., genome sequencing, qPCR quantification, cloning, etc.). The device can then be opened and the cartridge removed.

[0340] Figure 10 provides a diagram illustrating an exemplary apparatus for preparing a sample (e.g., a highly concentrated or multiplexed sample). See, for example, U.S. Patent No. 8,608,929, which is incorporated herein by reference in its entirety.

[0341] B. Devices for sequencing Also provided collectively are devices for use in the process of sequencing a polypeptide-containing sample (e.g., a multiplexed sample), including a cartridge (e.g., one containing a channel (e.g., a microfluidic channel)) and / or a pump (e.g., a peristaltic pump). The sequencing of nucleic acids or polypeptides according to this disclosure may, in some embodiments, be carried out using a system that enables parallel single-molecule analysis and / or single-molecule sequencing. The system may include a sequencing device and instruments configured to interface with the sequencing device.

[0342] A sequencing device may include a sequencing module comprising an array of pixels, where each pixel comprises a sample well and at least one photodetector. The sample wells of the sequencing device may be formed on or through the surface of the sequencing device and may be configured to receive a sample placed on the surface of the sequencing device. In some embodiments, the sample wells are components of a cartridge (e.g., a disposable or single-use cartridge) that can be inserted into the device. In summary, a sample well can be considered an array of sample wells. Multiple sample wells may have appropriate sizes and shapes so that at least some of the sample wells receive a single target molecule, or receive a sample containing multiple molecules (e.g., a target nucleic acid or target polypeptide). In some embodiments, the number of molecules in the sample wells may be distributed among the sample wells of the sequencing device such that some sample wells contain one molecule (e.g., a target nucleic acid or target polypeptide) and other sample wells contain zero, two, or more molecules.

[0343] In some embodiments, the sequencing device is positioned to receive a sample containing multiple molecules (e.g., one or more polypeptides of interest) from a sample preparation device. In some embodiments, the sequencing device is connected directly (e.g., physically) or indirectly to the sample preparation device.

[0344] A sequencing device may include an array of pixels, where each pixel includes one sample well and at least one photodetector. The sample wells of the sequencing device may be formed on or through the surface of the sequencing device and may be configured to receive samples placed on the surface of the sequencing device. In summary, a sample well can be considered an array of sample wells. Multiple sample wells may have appropriate sizes and shapes such that at least some of the sample wells receive samples containing a single sample (e.g., a single molecule such as a polypeptide). In some embodiments, the number of samples in the sample wells may be distributed among the sample wells of the sequencing device such that some sample wells contain one sample and other sample wells contain zero, two, or more samples.

[0345] Excitation light is supplied to the sequencing device from one or more light sources, which may be located outside or inside the sequencing device. Optical elements of the sequencing device receive the excitation light from the light sources and direct the light towards an array of sample wells in the sequencing device, illuminating an illumination area within the sample wells. In some embodiments, the sample wells may have a configuration that allows the sample to be held close to the surface of the sample well, which can facilitate the delivery of excitation light to the sample and the detection of emitted light from the sample. A sample placed within the illumination area can emit light in response to being irradiated with excitation light. For example, a sample can be labeled with a fluorescent marker that emits light in response to achieving an excited state by irradiation with excitation light. The emitted light emitted by the sample can then be detected by one or more photodetectors in pixels corresponding to the sample wells containing the sample being analyzed. According to some embodiments, multiple samples can be analyzed in parallel when performed across an entire array of sample wells, which may be between approximately 10,000 and 1,000,000 pixels.

[0346] A sequencing device may include an optical system for receiving excitation light and directing the excitation light between sample well arrays. The optical system may include one or more lattice couplers configured to couple the excitation light to the sequencing device and direct the excitation light to other optical elements. The optical system may include optical elements that direct the excitation light from the lattice couplers to the sample well arrays. Such optical elements may include optical splitters, optical combiners, and waveguides. In some embodiments, one or more optical splitters can couple the excitation light from the lattice couplers and deliver the excitation light to at least one waveguide. According to some embodiments, the optical splitters may have a configuration that allows for the delivery of excitation light to be substantially uniform across all waveguides, so that each waveguide receives substantially the same amount of excitation light. Such embodiments can improve the performance of the sequencing device by improving the uniformity of the excitation light received by the sample wells of the sequencing device. For example, suitable components that may be included in a sequencing device for coupling excitation light to a sample well and / or directing emitted light to a photodetector are described in U.S. Patent Application 14 / 821,688, filed on 7 August 2015, entitled “INTEGRATED DEVICE FOR PROBING, DETECTING AND ANALYZING MOLECULES,” and U.S. Patent Application 14 / 543,865, filed on 17 November 2014, entitled “INTEGRATED DEVICE WITH EXTERNAL LIGHT SOURCE FOR PROBING, DETECTING, AND ANALYZING MOLECULES,” both of which are incorporated in their entirety by reference. Examples of suitable grid couplers and waveguides that can be implemented in array determination devices are described in U.S. Patent Application No. 15 / 844,403, filed December 15, 2017, entitled "OPTICAL COUPLER AND WAVEGUIDE SYSTEM," which is incorporated in its entirety by reference.

[0347] An additional photonic structure can be placed between the sample well and the photodetector and configured to reduce or prevent excitation light from reaching the photodetector. Otherwise, signal noise may be introduced in the detection of emitted light. In some embodiments, a metal layer that can function as a circuit in the sequencing device can also function as a spatial filter. Examples of suitable photonic structures may include spectral filters, polarizing filters, and spatial filters, which are described in U.S. Patent Application No. 16 / 042,968, filed July 23, 2018, entitled "OPTICAL REJECTION PHOTONIC STRUCTURES," and incorporated in whole by reference.

[0348] An excitation source can be positioned and aligned with respect to the alignment device using components located outside the alignment device. Such components may include optical elements, including lenses, mirrors, prisms, windows, apertures, attenuators, and / or optical fibers. Additional mechanical components may be included in the instrument to enable control of one or more alignment components. Such mechanical components may include actuators, stepping motors, and / or knobs. An example of a suitable excitation source and alignment mechanism is described in U.S. Patent Application 15 / 161,088, filed May 20, 2016, entitled “PULSED LASER AND SYSTEM,” which is incorporated herein by reference. Another example of a beam steering module is described in U.S. Patent Application 15 / 842,720, filed December 14, 2017, entitled “COMPACT BEAM SHAPING AND STEERING ASSEMBLY,” which is incorporated herein by reference. An example of adding a suitable excitation source is described in U.S. Patent Application 14 / 821,688, filed on August 7, 2015, entitled "INTEGRATED DEVICE FOR PROBING, DETECTING AND ANALYZING MOLECULES," which is incorporated in its entirety by reference.

[0349] Photodetectors positioned with individual pixels of a sequencing device can be configured and positioned to detect light emitted from a sample well corresponding to a pixel. An example of a suitable photodetector is described in U.S. Patent Application 14 / 821,656, filed August 7, 2015, entitled "INTEGRATED DEVICE FOR PROBING, DETECTING AND ANALYZING MOLECULES," which is incorporated in its entirety by reference. In some embodiments, the sample wells and their respective photodetectors can be aligned along a common axis. In this way, the photodetectors can sufficiently overlap with the sample within the pixel.

[0350] The characteristics of the detected emitted light can provide an index for identifying a marker associated with the emitted light. Such characteristics may include any suitable type of characteristics, including the arrival time of photons detected by the photodetector, the amount of photons accumulated over time by the photodetector, and / or the distribution of photons across two or more photodetectors. In some embodiments, the photodetector may have a configuration that allows detection of one or more timing characteristics (e.g., luminescence lifetime) associated with the emitted light of the sample. The photodetector can detect the distribution of photon arrival times after the excitation light pulse has propagated through the sequencing device, and the distribution of arrival times can provide an index of the timing characteristics of the emitted light of the sample (e.g., a substitute for luminescence lifetime). In some embodiments, one or more photodetectors provide an index of the probability (e.g., luminescence intensity) of emitted light emitted by a marker. In some embodiments, multiple photodetectors may be sized and arranged to capture the spatial distribution of emitted light. The output signals from one or more photodetectors can then be used to distinguish one marker from multiple markers, and multiple markers can be used to identify one sample within a sample. In some embodiments, the sample may be excited by multiple excitation energies, and the emitted light and / or timing characteristics of the emitted light emitted by the sample in response to the multiple excitation energies may distinguish one marker from multiple markers.

[0351] During operation, parallel analysis of samples in a sample well is performed by exciting some or all of the samples in the well using excitation light and detecting a signal from the sample emission with a photodetector. The light emitted from the sample is detected by the corresponding photodetector and can be converted into at least one electrical signal. The electrical signal may be transmitted along a wire in the circuit of the sequencing device, which may be connected to an instrument interfaced with the sequencing device. The electrical signal can then be processed and / or analyzed. Processing or analysis of the electrical signal may be performed by a suitable computing device located inside or outside the instrument. The instrument may include a user interface for controlling the operation of the instrument and / or the sequencing device. The user interface may be configured to allow the user to input information into the instrument, such as commands and / or settings used to control the functions of the instrument. In some embodiments, the user interface may include buttons, switches, dials, and a microphone for voice commands. The user interface may allow the user to receive feedback on the performance of the instrument and / or the sequencing device, such as information obtained by readout signals from the photodetector on the sequencing device for proper alignment and / or on the sequencing device. In some embodiments, the user interface may provide feedback using a speaker to provide audible feedback. In some embodiments, the user interface may include indicator lights and / or display screens to provide the user with visual feedback.

[0352] In some embodiments, the device may include a computer interface configured to connect to a computing device. The computer interface may be a USB interface, a FireWire® interface, or other suitable computer interface. The computing device may be any general-purpose computer, such as a laptop or desktop computer. In some embodiments, the computing device may be a server (e.g., a cloud-based server) accessible via a wireless network through a suitable computer interface. The computer interface can facilitate the communication of information between the device and the computing device. Input information for controlling and / or configuring the device may be provided to the computing device and transmitted to the device via the computer interface. Output information generated by the device may be received by the computing device via the computer interface. Output information may include feedback on the performance of the device, the performance of the sequencing device, and / or data generated from the readout signals of the photodetector.

[0353] In some embodiments, the instrument may include a processing unit configured to analyze data received from one or more photodetectors of the sequencing device and / or transmit control signals to an excitation source. In some embodiments, the processing unit may include a general-purpose processor, a specially adapted processor (e.g., a central processing unit (CPU) such as one or more microprocessors or microcontroller cores, a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a custom integrated circuit, a digital signal processor (DSP), or a combination thereof). In some embodiments, processing of data from one or more photodetectors may be performed by both the instrument's processing device and an external computing device. In other embodiments, the external computing device may be omitted, and processing of data from one or more photodetectors may be performed solely by the sequencing device's processing device.

[0354] According to some embodiments, instruments configured to analyze a sample based on luminescence characteristics can detect differences in luminescence lifetime and / or intensity between different luminescent molecules, and / or differences in lifetime and / or intensity between the same luminescent molecule in different environments. The inventors have recognized and understood that differences in luminescence lifetime can be used to identify the presence or absence of different luminescent molecules, and / or to identify different environments or conditions to which luminescent molecules are exposed. In some cases, identifying luminescent molecules based on lifetime (rather than emission wavelength) can simplify the configuration of the system. For example, when identifying luminescent molecules based on lifetime, the number of wavelength-discriminating optical systems (such as wavelength filters, dedicated detectors for each wavelength, dedicated pulsed light sources for different wavelengths, and / or diffraction optical systems) can be reduced or eliminated. In some cases, a single-pulse light source operating at a single characteristic wavelength can be used to excite different luminescent molecules that emit within the same wavelength region of the optical spectrum but have measurably different lifetimes. An analytical system that uses a single-pulse light source, rather than multiple light sources operating at different wavelengths, to excite and identify different luminescent molecules emitting in the same wavelength region is easier to operate and maintain, more compact, and can be manufactured at a lower cost.

[0355] While analytical systems based on emission lifetime analysis may have certain advantages, the amount of information obtained and / or the detection accuracy of the analytical system can be improved by adding further detection techniques. For example, some embodiments of the system may be further configured to identify one or more characteristics of a sample based on emission wavelength and / or emission intensity. In some implementations, emission intensity can be used additionally or alternatively to distinguish different emission labels. For example, some emission labels may emit at significantly different intensities or have large differences in excitation probability, even if their decay rates are similar (e.g., a difference of at least about 35%). By comparing the binned signal with the measured excitation light, it may be possible to distinguish different emission labels based on intensity levels.

[0356] According to several embodiments, different emission lifetimes can be distinguished in a photodetector configured to time-binn emission events following excitation of an emission marker. Time binning may occur during a single charge accumulation cycle of the photodetector. A charge accumulation cycle is the interval between readout events, during which photogenerated carriers are accumulated in the bins of the time-binning photodetector. An example of a time-binning photodetector is described in U.S. Patent Application 14 / 821,656, filed August 7, 2015, entitled “INTEGRATED DEVICE FOR PROBING, DETECTING AND ANALYZING MOLECULES,” which is incorporated in its entirety by reference. In some embodiments, the time-binning photodetector can generate charge carriers in a photon absorption / carrier generation region and transfer the charge carriers directly to a charge carrier storage bin in a charge carrier storage region. In such embodiments, the time-binning photodetector may not include a carrier transfer / capture region. Such time-binning photodetectors are sometimes referred to as “direct binning pixels.” An example of a time-binning photodetector including a direct binning pixel is described in U.S. Patent Application 15 / 852,571, filed December 22, 2017, entitled "INTEGRATED PHOTODETECTOR WITH DIRECT BINNING PIXEL," which is incorporated herein by reference.

[0357] In some embodiments, different numbers of fluorophores of the same type may be linked to different reagents in the sample, and as a result, each reagent may be distinguished based on its emission intensity. For example, two fluorophores may be linked to a first label affinity reagent, and four or more fluorophores may be linked to a second label affinity reagent. Because of the different numbers of fluorophores, different excitation and fluorophore emission probabilities may exist associated with different affinity reagents. For example, if there are more emission events for the second label affinity reagent during the signal accumulation interval, the apparent intensity of the bottle will be significantly higher than that of the first label affinity reagent.

[0358] The inventors have recognized and understood that distinguishing nucleotides or other biological or chemical samples based on fluorophore decay rate and / or fluorophore intensity allows for simplification of photoexcitation and detection systems. For example, photoexcitation can be performed using a single-wavelength source (e.g., a light source that produces one characteristic wavelength, rather than multiple light sources or light sources operating at multiple different characteristic wavelengths). Furthermore, wavelength discrimination optics and filters can be eliminated from the detection system. Also, a single photodetector can be used for each sample well to detect emission from different fluorophores. The phrase “characteristic wavelength” or “wavelength” is used to refer to the central or dominant wavelength within a limited bandwidth of radiation (e.g., the central or peak wavelength within a 20 nm bandwidth output by a pulsed light source). In some cases, “characteristic wavelength” or “wavelength” may be used to refer to the peak wavelength within the entire bandwidth of radiation output from the source.

[0359] Equivalents and range In a claim, articles such as “one” and “the said” may mean one or more unless otherwise indicated or the context makes it clear. A claim or statement containing “or” between one or more members of a group is deemed satisfied unless otherwise indicated or the context makes it clear that one, two or more, or all, members of the group are present in, used in or otherwise related to a given product or process. The present invention includes embodiments in which exactly one member of the group is present in, used in or otherwise related to a given product or process. The present invention includes embodiments in which two or more, or all, members of the group are present in, used in or otherwise related to a given product or process.

[0360] Furthermore, the present invention encompasses all variations, combinations, and permutations in which one or more limitations, elements, clauses, and descriptive terms from one or more of the listed claims are introduced into another claim. For example, a claim referencing another claim can be modified to include one or more limitations found in other claims referencing the same basic claim. Where components are presented as a list, for example in Markush group format, each subgroup of those components is also disclosed, and any component(s) can be removed from a group. In general, where the present invention or an aspect of the present invention is referred to as including certain components and / or features, it should be understood that a particular embodiment of the present invention or an aspect of the present invention consists of, or essentially consists of, such components and / or features. For brevity, these embodiments are not specifically described verbatim in this specification.

[0361] As used herein and in the claims, the phrase “and / or” should be understood to mean “either or both” of the elements thus combined, that is, elements that exist as a combination in some cases and as separate elements in others. Multiple elements mentioned with “and / or” should be interpreted in the same way; that is, “one or more” of the elements thus combined. Other elements other than those specifically identified by the “and / or” clause may exist at their discretion, whether or not they are related to the specifically identified elements. Therefore, as a non-restrictive example, a reference to “A and / or B” when used in combination with open-ended wording such as “including” may refer to A only in one embodiment (including, at their discretion, elements other than B); B only in another embodiment (including, at their discretion, elements other than A); and in yet another embodiment, both A and B (including, at their discretion, other elements); and so on.

[0362] Where used herein and in the claims, “or” should be understood to have the same meaning as “and / or” as defined above. For example, when separating items in a list, “or” or “and / or” shall be interpreted in an inclusive sense; that is, including at least one of the list or number of elements, but also including multiple, and optionally including additional unlisted items. Conversely, only terms that are clearly indicated, such as “one of” or “exactly one,” or, where used in the claims, “consisting of,” refer to including exactly one element of the list or number of elements. In general, where used herein, the term “or” shall be interpreted to indicate an exclusive choice (i.e., “one or the other, but not both”) only when preceded by an exclusive term such as “either,” “one of,” “one of” or “exactly one.” Where used in the claims, “essentially consisting of” shall have the usual meaning as used in the field of patent law.

[0363] As used herein and in the claims, the phrase “at least one” relating to a list of one or more elements shall be understood to mean at least one element selected from any one or more elements contained in the list of elements, but not necessarily including one from each of the elements specifically listed in the list of elements, and not excluding any combination of elements contained in the list of elements. This definition also means that elements other than those specifically identified in the list of elements referred to by the phrase “at least one” may optionally exist, whether or not they are related to the specifically identified elements. Therefore, as a non-limiting example, “at least one of A and B” (or equivalently, “at least one of A or B” or equivalently, “at least one of A and / or B”) may, in one embodiment, mean that there is at least one optionally two or more A's and no B (and optionally an element other than B); in another embodiment, mean that there is at least one optionally two or more B's and no A (and optionally an element other than A); and in yet another embodiment, mean that there is at least one optionally two or more A's and at least one optionally two or more B's (and optionally an other element), and so on.

[0364] Furthermore, unless otherwise explicitly stated, in any method claimed in this application that involves multiple steps or actions, the order of the steps or actions in that method is not necessarily limited to the order in which they are described.

[0365] In the claims and the above specification, all transitional phrases such as “include,” “contain,” “support,” “have,” “incorporate,” “involve,” “hold,” and “compose of” should be understood to be open-ended, meaning they include but are not limited to them. As described in Section 2111.03 of the U.S. Patent and Trademark Office Manual of Examination Procedure, only the transitional phrases “consist of” and “essentially consist of” are closed or semi-closed transitional phrases, respectively. Embodiments described herein using open-ended transitional phrases (e.g., “include”) should also be understood to be intended to consist of and essentially consist of the features described by the open-ended transitional phrases in alternative embodiments. For example, where this application describes “a composition comprising A and B,” this application also intends “a composition comprising A and B” and “a composition essentially consisting of A and B” as alternative embodiments.

[0366] If a range is specified, boundary values ​​are included. Furthermore, unless otherwise indicated, or unless it is obvious from the context and the understanding of those skilled in the art, values ​​expressed as a range can be considered in different embodiments of the invention as any specific value or subrange within the described range, and may be up to one-tenth of the lower limit unit of the range, unless the context clearly indicates otherwise.

[0367] This application references various issued patents, published patent applications, academic papers, and other publications, all of which are incorporated herein by reference. In the event of any conflict between any incorporated reference and this specification, this specification shall prevail. Furthermore, any particular embodiment of the Invention that constitutes prior art may be expressly excluded from any one or more claims. Such embodiments are considered to be known to those skilled in the art and may be excluded even if the exclusion is not expressly stated herein. Any particular embodiment of the Invention may be excluded from any claim for any reason, whether or not it relates to the existence of prior art.

[0368] Those skilled in the art will be able to recognize or confirm, by routine experimentation alone, many equivalents to the specific embodiments described herein. The scope of the embodiments described herein is not intended to be limited to the above description, but rather to the claims provided herein. Those skilled in the art will understand that various changes and modifications to this description can be made without departing from the spirit or scope of the invention, as defined in the following claims.

[0369] Any description of a list of chemical groups in any definition of a variable element in this specification includes the definition of that variable element as any single group or a combination of the listed groups. Any description of embodiments of a variable element in this specification includes that embodiment as any single embodiment or in combination with any other embodiment or part thereof.

Claims

[Claim 1] The invention described in the specification.