Polypeptide cleaving reagents and uses thereof
The use of aminopeptidases from Pyrococcus horikoshii with specific tag sequences improves polypeptide cleavage efficiency, addressing the challenges of proteomics in sequencing by enhancing cleavage depth and reducing rapid sequential cleavage.
Patent Information
- Application Number
- JP2025522628
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-21
- Filing Date
- 2023-10-20
- Publication Date
- 2025-12-02
AI Technical Summary
Next-generation DNA sequencing technologies have struggled to capture the complex and dynamic state of proteomics due to scale and inability to amplify sources, limiting our understanding of heredity and gene regulation.
A composition comprising aminopeptidases from Pyrococcus horikoshii with different amino acid sequences and tag sequences, such as polyhistidine and biotinylation tags, is used to improve polypeptide cleavage efficiency in sequencing reactions.
Enhances polypeptide cleavage activity, allowing for more detailed structural information to be obtained from sequencing reactions by reducing rapid sequential cleavage and increasing cleavage depth.
Smart Images

Figure 2025538856000001_ABST
Abstract
Description
[Technical Field]
[0001] Polypeptide cleavage reagents and uses thereof. The contents of the electronic sequence listing (R070870165WO-SEQ-JIB.xml; size: 64,843 bytes; created on October 20, 2023) are incorporated herein by reference in their entirety. [Background technology]
[0002] Proteins are the major structural and functional components of cells, driving crucial biological and cellular processes. Next-generation DNA sequencing technologies have revolutionized our understanding of heredity and gene regulation, but the complex and dynamic state of a cell is not fully captured by its genome and transcriptome. Applying a similar approach to proteomics has been difficult due to the scale, dynamic range, and inability to amplify sources. Summary of the Invention
[0003] In some embodiments, the present disclosure provides a composition comprising a first cleavage reagent comprising a first aminopeptidase from Pyrococcus horikoshii and a first tag sequence, and a second cleavage reagent comprising a second aminopeptidase from Pyrococcus horikoshii, wherein the first and second aminopeptidases have different amino acid sequences.
[0004] In some embodiments, the amino acid sequences of the first and second aminopeptidases share less than 80% sequence identity (e.g., less than 70%, less than 60%, less than 50%, 10-80%, 20-60%, 30-50%, or 40-50% sequence identity). In some embodiments, the first aminopeptidase has an amino acid sequence that is at least 80% identical to SEQ ID NO:3. In some embodiments, the second aminopeptidase has an amino acid sequence that is at least 80% identical to SEQ ID NO:1.
[0005] In some embodiments, the second cleavage reagent comprises a second tag sequence. In some embodiments, the composition further comprises a third cleavage reagent comprising an aminopeptidase from Yersinia pestis.
[0006] In some embodiments, the present disclosure provides a composition comprising a first cleavage reagent comprising a first aminopeptidase having an amino acid sequence at least 80% identical to SEQ ID NO:3, a second cleavage reagent comprising a second aminopeptidase having an amino acid sequence at least 80% identical to SEQ ID NO:1, and a third cleavage reagent comprising a third aminopeptidase having an amino acid sequence at least 80% identical to SEQ ID NO:5 or 7.
[0007] In some embodiments, the first cleavage reagent comprises a first tag sequence, the second cleavage reagent comprises a second tag sequence, and the third cleavage reagent comprises a third tag sequence. In some embodiments, the first tag sequence is attached to the end of the first aminopeptidase, the second tag sequence is attached to the end of the second aminopeptidase, and the third tag sequence is attached to the end of the third aminopeptidase. In some embodiments, each tag sequence is attached to the C-terminus of its respective aminopeptidase. In some embodiments, each tag sequence independently comprises at least two amino acids (e.g., 2 to 200, 2 to 100, 4 to 80, 5 to 50, 5 to 30, 5 to 20, 10 to 100, 20 to 80, or 30 to 70 amino acids).
[0008] In some embodiments, at least one of the first, second, and third tag sequences comprises a polyhistidine tag. In some embodiments, at least one of the first, second, and third tag sequences comprises a biotinylation tag. In some embodiments, the biotinylation tag comprises at least one biotin ligase recognition sequence. In some embodiments, the biotinylation tag comprises two biotin ligase recognition sequences oriented in tandem.
[0009] In some embodiments, the first, second, and third cleavage reagents are present in the composition at first, second, and third concentrations, respectively, where the first concentration is at least two-fold higher than the second concentration. In some embodiments, the first concentration is at least five-fold higher than the third concentration. In some embodiments, the second concentration is at least two-fold higher than the third concentration. In some embodiments, the molar ratio of the first cleavage reagent to the second cleavage reagent in the composition is about 2:1 to about 20:1. In some embodiments, the molar ratio of the first cleavage reagent to the third cleavage reagent in the composition is about 5:1 to about 200:1.
[0010] In some embodiments, the first cleavage reagent is a polypeptide having an amino acid sequence that is at least 80% identical to SEQ ID NO: 4. In some embodiments, the second cleavage reagent is a polypeptide having an amino acid sequence that is at least 80% identical to SEQ ID NO: 2. In some embodiments, the third cleavage reagent is a polypeptide having an amino acid sequence that is at least 80% identical to SEQ ID NO: 8.
[0011] In some aspects, the present disclosure provides a reaction mixture for polypeptide analysis. In some embodiments, the reaction mixture comprises a composition described herein and one or more amino acid binding proteins that do not have peptide cleavage activity. In some embodiments, the reaction mixture comprises a first aminopeptidase having an amino acid sequence that is at least 92% (e.g., 92-99%, 94-100%, 96-100%, 98-100%) identical to SEQ ID NO: 3, a cleavage reagent that includes a first tag sequence, and one or more amino acid binding proteins that do not have peptide cleavage activity.
[0012] In some embodiments, the present disclosure provides a method for analyzing polypeptides. In some embodiments, the method for analyzing polypeptides comprises: contacting a polypeptide with the reaction mixture described herein; monitoring a signal for a signal pulse corresponding to the interaction between one or more amino acid binding proteins and the polypeptide; and determining at least one chemical characteristic of the polypeptide based on a characteristic pattern in the signal.
[0013] In some aspects, the present disclosure provides a system comprising at least one hardware processor and at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by the at least one hardware processor, cause the at least one hardware processor to perform a method described herein. In some aspects, the present disclosure provides at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by the at least one hardware processor, cause the at least one hardware processor to perform a method described herein.
[0014] The details of certain embodiments of the present disclosure are set forth in the detailed description. Other features, objects, and advantages of the present disclosure will be apparent from the examples, drawings, and claims.
[0015] The accompanying drawings, which constitute a part of this specification, illustrate several embodiments of the present disclosure and, together with the accompanying description, serve to explain the principles of the disclosure. [Brief explanation of the drawings]
[0016] [Figure 1A] An exemplary overview of real-time dynamic protein sequencing is shown. Protein samples are digested into peptide fragments, immobilized in a nanoscale reaction chamber, and incubated with a mixture of freely diffusing N-terminal amino acid (NAA) recognition factors and aminopeptidases that perform the sequencing process. Labeled recognition factors bind to peptides on and off when one of their cognate NAAs is exposed at the N-terminus, thereby generating a characteristic pulse pattern. The NAA is cleaved by the aminopeptidase, exposing the next amino acid for recognition. The temporal order and binding kinetics of NAA recognition enable peptide identification and are sensitive to features that modulate binding kinetics, such as post-translational modifications (PTMs). [Figure 1B] 1 shows an exemplary schematic diagram of a pixel of an integrated device. [Figure 2A] Figures 2A-2D show exemplary results demonstrating improved TET aminopeptidase performance by extending the C-terminus of the peptide with a DDD motif. The bar graphs in Figures 2A-2D show cleavage depth (Figure 2A), cleavage activity (% of reads that reached the last visible RS) (Figure 2B), time required to cleave the DQQ motif and R residues (Figure 2C), and % of 4+ RS (Figure 2D). [Figure 2B] Same as above. [Figure 2C] Same as above. [Figure 2D] Same as above. [Figure 3A]Figures 3A-3B show exemplary results of AP30 expression and purification. Figure 3A shows an exemplary chromatogram showing AP30 separation and elution peaks. Figure 3B shows Talon affinity column-purified fractions separated on an SDS-PAGE gel to demonstrate AP30 enrichment (left image), as well as a native gel (right image) showing hTETII and AP30 protein profiles before and after preparation. [Figure 3B] Same as above. [Figure 4] Three bar graphs are shown comparing the intrinsic cleavage rates and substrate specificities by hTETII and AP30 for 18 individual amino acids categorized according to activity level (high activity: top chart; medium activity: middle chart; very low activity: bottom chart). [Figure 5A] Figures 5A-5C show exemplary results of a protein sequencing assay using the AP30 cutter (5 μM), the PS610 (50 nM) recognition agent, and the QP47 peptide (FAAAYPDDD) (SEQ ID NO: 46). Figure 5A shows a representative trace showing the initial and final RS recognition by PS610 as cleavage progresses. Figure 5B shows a plot showing the average cleavage time for the FA (left plot) and YP (right plot) RSs. Figure 5C shows a plot showing a population of rapid sequential cleavages and a population of ordered cleavages. [Figure 5B] Same as above. [Figure 5C] Same as above. [Figure 6A] Figures 6A-6C show exemplary results of a protein sequencing assay using 1 μM hTETII cutter, PS610 (50 nM) recognition factor, and QP47 peptide (FAAAYPDDD) (SEQ ID NO: 46). Figure 6A shows a representative trace showing the first and last RS by PS610 as cleavage progresses. Figure 6B shows a plot showing the average cleavage times for FA and YP RS. Figure 6C shows a plot showing a population of rapid sequential cleavages and a population of ordered cleavages. [Figure 6B] Same as above. [Figure 6C] Same as above. [Figure 7A]Figures 7A-7B show exemplary results of a protein sequencing assay using AP30 (1 μM) / PfuTET (40 μM) with PS610 (50 nM), PS557 (250 nM), and PS621 (250 nM) recognition factors and the QP433 peptide (RLIFAYPDDD) (SEQ ID NO: 47). Figure 7A shows a representative trace showing five RSs identified by the recognition factors as cleavage progresses. Figure 7B shows a bin ratio vs. pulse duration plot with separated clusters for recognized RSs (top plot), as well as plots showing the cleavage depth and percentage of each RS recognized in reads with gaps (middle plot) and reads without gaps (no missed recognizable residues) (bottom plot). Figure 7A also shows SEQ ID NO: 54. [Figure 7B] Same as above. [Figure 8A] Figures 8A-8B show exemplary results of protein sequencing assays using AP30 (3 μM) or hTETII (2 μM) / PfuTET (40 μM) aminopeptidase combinations with PS610 (50 nM) and PS557 (250 nM) recognition factors and QP425 (LASSIAEANRFADIADYP) (SEQ ID NO: 48) on the same chip. Figure 8A shows representative traces for reactions with AP30 and hTETII, showing RSs identified by the recognition factors as cleavage progresses. Figure 8B shows plots showing the cleavage depth and percentage of each RS recognized in ungapped reads (no recognizable residues missed). [Figure 8B] Same as above. [Figure 9A] 9A-9C show exemplary results demonstrating that AP30 improved cleavage performance by increasing cleavage activity and decreasing rapid sequential cleavage. The bar graphs in Figures 9A-9C show cleavage depth (Figure 9A), % of 4+ RS reads (Figure 9B), and rapid sequential cleavage of IA and FA RS (Figure 9C). [Figure 9B] Same as above. [Figure 9C] Same as above. [Figure 10A]Figures 10A-10B show exemplary results of AP37 expression and purification. Figure 10A shows an exemplary chromatogram showing AP37 enrichment and elution peaks. Figure 10B shows talon affinity column purified fractions separated on an SDS PAGE gel to demonstrate AP37 separation. [Figure 10B] Same as above. [Figure 11] Three bar graphs are shown comparing the intrinsic cleavage rates and substrate specificities by PfuTET and AP37 for 18 individual amino acids categorized according to activity level (high activity: top chart; medium activity: middle chart; very low activity: bottom chart). [Figure 12A] Figures 12A-12B show exemplary results of protein sequencing assays using AP37 (40 μM) or PfuTET (40 μM) aminopeptidase with PS610 (50 nM), PS961 (250 nM), and PS961 (250 nM) recognition factors, and QP514 (DQQRLIFAYPDDD) (SEQ ID NO: 49), on separate chips. Figure 12A shows representative traces for reactions with AP37 or PfuTET, showing RSs identified by the recognition factors as cleavage progresses. Figure 12B shows plots showing the cleavage depth and percentage of each RS recognized in ungapped reads (no recognizable residues missed). [Figure 12B] Same as above. [Figure 13A] Figures 13A-13C show exemplary results of protein sequencing assays performed on QP433 (RLIFAYPDDD) (SEQ ID NO: 47) with two different AP37 concentrations (20 μM / 40 μM), 4 μM hTETII, and PS610 (50 nM), PS557 (250 nM), and PS691 (100 nM) recognition factors. Figure 13A shows the improvement in cleavage depth. Figure 13B shows the improvement in % of significant reads. Figure 13C shows that AP37 reduced rapid sequential cleavage (RSC) of LI, IF, and FA RS. [Figure 13B] Same as above. [Figure 13C] Same as above. [Figure 14A]Figures 14A-14B show exemplary results showing the effect of different AP30 / AP37 ratios with the PS610 (50 nM), PS691 (100 nM), and PS961 (250 nM) recognition factors on the time to cleave the DQQ motif and R residue (Figure 14A) and the rapid sequential cleavage of RS (Figure 14B). [Figure 14B] Same as above. [Figure 15A] Figures 15A-15B show bar graphs depicting the distribution of performance metrics (cleavage depth, % of 4+ RS reads, and % of 5 RS reads) using a recognition factor mixture of PS610 (50 nM), PS961 (250 nM), and PS691 (100 nM), and AP30 / AP37 at 4 μM / 40 μM concentrations for QP514 (DQQRLIFAYPDDD) (SEQ ID NO: 49) (Figure 15A) and 10 μM / 60 μM concentrations for QP549 (DQQIASSRLAASFAAQQYPDDD) (SEQ ID NO: 50) (Figure 15B). [Figure 15B] Same as above. [Figure 16] Figure 1 shows the read resolution and abundance of ungapped reads with a single deletion allowed at a read length of 4 from one of the chip runs with the recognition factor and aminopeptidase mixtures of 250 / 100 nM PS961 / PS691 and 10 / 60 μM AP30 / hTETIII for QP549. [Figure 17A] Figures 17A-17B show exemplary results of yPIP expression and purification. Figure 17A shows an exemplary chromatogram showing yPIP separation and elution peaks. Figure 17B shows Talon affinity column purified fractions separated on an SDS PAGE gel to demonstrate yPIP enrichment. [Figure 17B] Same as above. [Figure 18A-1] Representative traces are shown showing the effect of adding yPIP (2 μM) to the AP30 / AP37 (4 / 40 μM) aminopeptidase combination in a protein sequencing assay. SEQ ID NO: 55 is shown throughout the figure. [Figure 18A-2] Same as above. [Figure 18B-1] Same as above. [Figure 18B-2]Same as above. [Figure 19A] 1 shows exemplary results of a protein sequencing assay using the combinations (a) AP30 (4 μM) / AP37 (40 μM), and (b) AP30 (4 μM) / AP37 (40 μM) / yPIP (2 μM) for the QP352 (RLAYPAFAAYPDDDF) (SEQ ID NO: 51) peptide, together with the recognition factors PS610 (50 nM) / PS961 (125 nM) / PS1122 (250 nM). [Figure 19B] Same as above. [Figure 20A] Figure 1 shows the read ratio resolution for QP352 (RLAYPAFAAYPDDDF) (SEQ ID NO: 51) protein sequencing assays using (a) AP30 (4 μM) / AP37 (40 μM), and (b) AP30 (4 μM) / AP37 (40 μM) / yPIP (2 μM) combinations with PS610 (50 nM) / PS961 (125 nM) / PS1122 (250 nM) recognition factors. [Figure 20B] Same as above. [Figure 21A] 1 shows exemplary results of a protein sequencing assay using the combinations (a) AP30 (4 μM) / AP37 (40 μM) / yPIP (0.5 μM), and (b) AP30 (4 μM) / AP37 (40 μM) / yPIP (1 μM) for the QP352 (RLAYPAFAAYPDDDF) (SEQ ID NO: 51) peptide, together with the recognition factors PS610 (50 nM) / PS961 (125 nM). [Figure 21B] Same as above. [Figure 22A] Figure 1 shows the read fraction resolution for QP352 (RLAYPAFAAYPDDDF) (SEQ ID NO: 51) protein sequencing assays using (a) AP30 (4 μM) / AP37 (40 μM) / yPIP (0.5 μM) and (b) AP30 (4 μM) / AP37 (40 μM) / yPIP (1 μM) combinations with the PS610 (50 nM) / PS961 (125 nM) recognition factor combination and the PS610 (50 nM) / PS961 (125 nM) recognition factor. [Figure 22B] Same as above. [Figure 23A]1 shows exemplary results of a protein sequencing assay using the combinations (a) AP30 (4 μM) / AP37 (40 μM) / yPIP-BC_B1 (2 μM) and (b) AP30 (4 μM) / AP37 (40 μM) / yPIP-470 (2 μM) for the QP352 (RLAYPAFAAYPDDDF) (SEQ ID NO: 51) peptide, together with the recognition factors PS610 (50 nM) / PS961 (125 nM) / PS1122 (250 nM). [Figure 23B] Same as above. [Figure 24A] Figure 1 shows the read ratio resolution for QP352 (RLAYPAFAAYPDDDF) (SEQ ID NO: 51) protein sequencing assay using combinations of (a) AP30 (4 μM) / AP37 (40 μM) / yPIP-BC_B1 (2 μM) and (b) AP30 (4 μM) / AP37 (40 μM) / yPIP-470 (2 μM) with PS610 (50 nM) / PS961 (125 nM) / PS1122 (250 nM) recognition factors. [Figure 24B] Same as above. [Figure 25A] (a) yPIP (SEQ ID NO: 6) and AP70 (SEQ ID NO: 8) protein amino acid sequence alignment showing the C-terminal 6xH tag and GGS-6xHis tag, respectively, and (b) expression constructs for yPIP and AP70, with the molecular weights of the proteins. [Figure 25B] Same as above. [Figure 26A] Figures 26A-26B show exemplary results of AP70 expression and purification. Figure 26A shows an exemplary chromatogram showing AP70 separation and elution peaks. Figure 26B shows talon affinity column purified fractions separated on an SDS PAGE gel to demonstrate AP70 enrichment. [Figure 26B] Same as above. [Figure 27-1] 1 shows exemplary results of an HPLC assay for QP734 (FPARAFAYPDDD) (SEQ ID NO: 52) peptide cleavage by AP70 at various concentrations after different conditioning treatments and an unconditioned control. [Figure 27-2] Same as above. [Figure 27-3] Same as above. [Figure 27-4] Same as above. [Figure 28-1] 1 shows exemplary results of an HPLC assay for different peptide cleavage by AP70 after different conditioning treatments and an unconditioned control. [Figure 28-2] Same as above. [Figure 28-3] Same as above. [Figure 28-4] Same as above. [Figure 29] 1 shows exemplary results comparing the cleavage activity and substrate specificity for AP70, AP30+AP37, and AP30+AP37+AP70 combinations against proline-containing PXX, XPP, and XPX peptides. [Figure 30] 1 shows exemplary results of a protein sequencing assay with AP70 (1 μM, 4 μM), a no-AP control, and AP30 (4 μM) for the QP734 (FPARAFAYPDDD) (SEQ ID NO: 52) peptide, together with PS610 (50 nM) / PS961 (125 nM) / PS1122 (250 nM) recognizers. [Figure 31A-1] Representative traces from a protein sequencing assay showing the effect of adding AP70 (1 μM) to the AP30 / AP37 (4 / 40 μM) aminopeptidase combination are shown. [Figure 31A-2] Same as above. [Figure 31B-1] Same as above. [Figure 31B-2] Same as above. [Figure 32A-1] 32A-32F show exemplary results of AP70 concentration titration to optimize AP combination performance of AP30+AP37+AP70 on a chip for QP352 (RLAYPAFAAYPDDDF) (SEQ ID NO: 51) sequencing. [Figure 32A-2] Same as above. [Figure 32B-1] Same as above. [Figure 32B-2] Same as above. [Figure 32C-1] Same as above. [Figure 32C-2] Same as above. [Figure 32D-1] Same as above. [Figure 32D-2] Same as above. [Figure 32E-1] Same as above. [Figure 32E-2] Same as above. [Figure 32F-1] Same as above. [Figure 32F-2] Same as above. [Figure 33A] 1 shows exemplary results of a protein sequencing assay using the combinations (a) AP30 (4 μM) / AP37 (40 μM) / AP70 (1 μM) and (b) AP30 (4 μM) / AP37 (40 μM) for the QP352 (RLAYPAFAAYPDDDF) (SEQ ID NO: 51) peptide, together with the recognition factors PS610 (50 nM) / PS961 (125 nM) / PS1122 (250 nM). [Figure 33B] Same as above. [Figure 34A] Figure 1 shows the read ratio resolution for QP352 (RLAYPAFAAYPDDDF) (SEQ ID NO: 51) protein sequencing assay using combinations of (a) AP30 (4 μM) / AP37 (40 μM) / AP70 (1 μM) and (b) AP30 (4 μM) / AP37 (40 μM) with PS610 (50 nM) / PS961 (125 nM) / PS1122 (250 nM) recognition factors. [Figure 34B] Same as above. [Figure 35A] 1 shows exemplary results of a protein sequencing assay using the combinations (a) AP30 (4 μM) / AP37 (40 μM) / AP70 (1 μM) and (b) AP30 (4 μM) / AP37 (40 μM) for the QP354 (LAAYPARLAYPDDDF) (SEQ ID NO: 53) peptide, together with the recognition factors PS610 (50 nM) / PS961 (125 nM) / PS1122 (250 nM). [Figure 35B] Same as above. [Figure 36A]Figure 1 shows the read ratio resolution for QP354 (LAAYPARLAYPDDDF) (SEQ ID NO: 53) protein sequencing assay using combinations of (a) AP30 (4 μM) / AP37 (40 μM) / AP70 (1 μM) and (b) AP30 (4 μM) / AP37 (40 μM) with PS610 (50 nM) / PS961 (125 nM) / PS1122 (250 nM) recognition factors. [Figure 36B] Same as above. [Figure 37A] 1 shows exemplary results of a protein sequencing assay using the combinations (a) AP30 (4 μM) / AP37 (40 μM) / AP70 BC-B1 batch (1 μM) and (b) AP30 (4 μM) / AP37 (40 μM) / yPIP BC-B1 batch (1 μM) for QP352 (RLAYPAFAAYPDDDF) (SEQ ID NO: 51) peptide, together with PS610 (50 nM) / PS961 (125 nM) / PS1122 (250 nM) recognition factors. [Figure 37B] Same as above. [Figure 38A] Figure 1 shows the read rate resolution for QP352 (RLAYPAFAAYPDDDF) (SEQ ID NO: 51) protein sequencing assay using (a) AP30 (4 μM) / AP37 (40 μM) / AP70 BC-B1 batch (1 μM) and (b) AP30 (4 μM) / AP37 (40 μM) / yPIP BC-B1 batch (1 μM) combinations with PS610 (50 nM) / PS961 (125 nM) / PS1122 (250 nM) recognition factors. [Figure 38B] Same as above. DETAILED DESCRIPTION OF THE INVENTION
[0017] Aspects of the present disclosure relate to compositions and methods for polypeptide analysis based on single-molecule binding interactions between a polypeptide and one or more reagents described herein. In some embodiments, the present disclosure provides cleavage reagents, such as aminopeptidases, with improved performance in peptide sequencing reactions. In some embodiments, the cleavage reagents of the present disclosure exhibit improved cleavage activity toward amino acids of a polypeptide, allowing more information to be obtained from a peptide sequencing reaction.
[0018] For example, Figure 1A shows an example of a dynamic peptide sequencing reaction in which individual on-off binding events result in a signal pulse of signal output. As shown on the left, a protein sample may be fragmented into peptides, which are immobilized in a reaction chamber and exposed to a mixture of amino acid recognition factors and cleavage reagents. As shown on the right, the amino acid recognition factors reversibly bind to the peptides, generating a series of changes in the signal output (e.g., signal pulses) as amino acids are progressively cleaved from the peptide termini. The temporal order of recognition and the kinetics of binding and / or cleavage can be used to determine structural information about the peptides.
[0019] Compositions and methods for performing dynamic peptide sequencing and analyzing the data obtained therefrom are fully described in PCT International Publication No. WO2020 / 102741A1, filed November 15, 2019, and PCT International Publication No. WO2021 / 236983A2, filed May 20, 2021, each of which is incorporated by reference in its entirety.
[0020] In some embodiments, the present disclosure provides cleavage reagents (e.g., aminopeptidases) with improved cleavage activity that may be advantageous in the context of peptide sequencing reactions. As described herein and shown in FIG. 1A, peptide sequencing reactions can be performed by exposing a polypeptide to a mixture of an amino acid recognition factor and an aminopeptidase and determining structural information about the polypeptide based on detectable recognition events that precede the cleavage events. In some embodiments, the aminopeptidases of the present disclosure enable more information to be obtained from the sequencing reaction. For example, in some embodiments, the aminopeptidases exhibit cleavage activity with improved speed for detecting recognition events between cleavage events (e.g., as indicated by a decrease in rapid sequential cleavage). In some embodiments, sequential cleavage of amino acids by the aminopeptidase proceeds further toward the peptide base (e.g., as indicated by an increased cleavage depth). In some embodiments, the aminopeptidase cleaves specific amino acids more efficiently compared to homologous enzymes.
[0021] Aminopeptidase In some embodiments, the present disclosure provides a cleavage reagent comprising an aminopeptidase having an amino acid sequence selected from Table 1. It is understood that the exemplary sequences in Table 1 and other examples described herein are meant to be non-limiting, and that an aminopeptidase in accordance with the present disclosure can include any homolog, variant, or fragment thereof that minimally includes the domain or subdomain responsible for amino acid cleavage.
[0022] In some embodiments, the cleavage reagent comprises an aminopeptidase having an amino acid sequence that is at least 80% identical to an amino acid sequence selected from Table 1. In some embodiments, the aminopeptidase has at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 92%, at least 94%, at least 95%, at least 96%, at least 98% or more amino acid sequence identity to an amino acid sequence selected from Table 1. In some embodiments, the aminopeptidase has 25-50%, 50-60%, 60-70%, 70-80%, 80-90%, 90-95%, 92-99%, 94-99%, 95-99%, 40-100%, 50-100%, 60-100%, 70-100%, 80-100%, 90-100%, 92-100%, 94-100%, 95-100%, 96-100%, or 100% amino acid sequence identity to an amino acid sequence selected from Table 1.
[0023] In some embodiments, the cleavage reagent comprises a synthetic or recombinant aminopeptidase. In some embodiments, the cleavage reagent comprises a monomeric aminopeptidase. In some embodiments, the cleavage reagent comprises a multimeric aminopeptidase (e.g., a multimeric complex of monomeric subunits that may be the same or different).
[0024] In some embodiments, the cleavage reagent comprises an aminopeptidase obtained or derived from a particular source (e.g., organism). As described herein, in some embodiments, an aminopeptidase identified as being from a particular organism does not impose a requirement that the aminopeptidase have an amino acid sequence that is 100% identical to a naturally occurring aminopeptidase from that organism, although in some embodiments, this requirement may be implied.
[0025] For example, in some embodiments, the cleavage reagent comprises an aminopeptidase from Pyrococcus horikoshii (e.g., Pyrococcus horikoshii TET aminopeptidase II, Pyrococcus horikoshii TET aminopeptidase III). In some embodiments, the aminopeptidase from Pyrococcus horikoshii is at least 80%, at least 85%, at least 90%, at least 95%, 80-100%, 85-95%, 90-99%, 95-99%, or 100% identical to a naturally occurring aminopeptidase from Pyrococcus horikoshii (e.g., Pyrococcus horikoshii TET aminopeptidase II, Pyrococcus horikoshii TET aminopeptidase III).
[0026] In some embodiments, the cleavage reagent comprises an aminopeptidase from Yersinia pestis (e.g., Yersinia pestis Xaa-prolyl aminopeptidase). In some embodiments, the aminopeptidase from Yersinia pestis is at least 80%, at least 85%, at least 90%, at least 95%, 80-100%, 85-95%, 90-99%, 95-99%, or 100% identical to a naturally occurring aminopeptidase from Yersinia pestis (e.g., Yersinia pestis Xaa-prolyl aminopeptidase).
[0027] In some embodiments, the cleavage reagent comprises an aminopeptidase from Pyrococcus furiosus (e.g., Pyrococcus furiosus aminopeptidase I). In some embodiments, the aminopeptidase from Pyrococcus furiosus is at least 80%, at least 85%, at least 90%, at least 95%, 80-100%, 85-95%, 90-99%, 95-99%, or 100% identical to a naturally occurring aminopeptidase from Pyrococcus furiosus (e.g., Pyrococcus furiosus aminopeptidase I).
[0028] Tag sequence In some embodiments, the cleavage reagent comprises an aminopeptidase and a tag sequence described herein. As used herein, in some embodiments, a tag sequence refers to a segment of amino acids bound to the aminopeptidase. In some embodiments, the tag sequence is attached to the terminus (e.g., terminal end) of the aminopeptidase. In some embodiments, the tag sequence is attached to the C-terminus (e.g., the C-terminal amino acid of the aminopeptidase) of the aminopeptidase. In some embodiments, the tag sequence is attached to the N-terminus (e.g., the N-terminal amino acid of the aminopeptidase) of the aminopeptidase. In some embodiments, the tag sequence is attached to an internal position of the aminopeptidase (e.g., an amino acid between the N-terminal amino acid and the C-terminal amino acid of the aminopeptidase).
[0029] In some embodiments, the tag sequence comprises at least two amino acids. For example, in some embodiments, the tag sequence comprises at least 2, at least 3, at least 4, at least 5, at least 6, at least 8, at least 10, at least 15, at least 25, at least 30, at least 40, at least 50, at least 60, at least 80, at least 100, or more amino acids. In some embodiments, the tag sequence comprises about 2 to about 200 amino acids (e.g., 2 to 150 amino acids, 2 to 100 amino acids, 50 to 200 amino acids, 50 to 150 amino acids, 50 to 100 amino acids, 4 to 80 amino acids, 5 to 50 amino acids, 5 to 30 amino acids, 5 to 20 amino acids, 10 to 100 amino acids, 20 to 80 amino acids, or 30 to 70 amino acids).
[0030] In some embodiments, the tag sequence comprises one or more functional moieties. In some embodiments, the functional moiety provides improved cleavage activity of the aminopeptidase to which it is attached. For example, in some embodiments, the cleavage reagent comprises a tag sequence that improves one or more properties associated with aminopeptidase cleavage activity (e.g., improved one or more of cleavage depth, rapid sequential cleavage, and cleavage efficiency). In some embodiments, the functional moiety of the tag sequence may provide one or more functions unrelated to aminopeptidase cleavage activity. For example, in some embodiments, the tag sequence comprises one or more of an affinity tag (e.g., a polyhistidine tag), a modification tag (e.g., a biotinylation tag), a solubility tag (e.g., a small ubiquitin-like modifier (SUMO) tag), and a linker.
[0031] In some embodiments, the tag sequence comprises a polyhistidine tag. In some embodiments, the polyhistidine tag comprises a segment of two or more histidine amino acids. In some embodiments, the polyhistidine tag comprises a segment of about 2 to about 15 (e.g., 4, 6, 8, 10, 12, 14, 15) histidine amino acids. In some embodiments, the polyhistidine tag comprises a hexahistidine tag (e.g., a 6xHis tag). In some embodiments, the polyhistidine tag comprises a decahistidine tag (e.g., a 10xHis tag). In some embodiments, the tag sequence comprises two or more (e.g., 2, 3, 4) polyhistidine tags.
[0032] In some embodiments, the tag sequence comprises a biotinylation tag. In some embodiments, the biotinylation tag comprises at least one biotin ligase recognition sequence. In some embodiments, the biotinylation tag comprises two biotin ligase recognition sequences oriented in tandem. In some embodiments, the biotin ligase recognition sequence refers to an amino acid sequence that can be recognized by a biotin ligase, which catalyzes the covalent bond between the sequence and a biotin molecule. In some embodiments, the biotin ligase recognition sequence comprises the amino acid sequence of SEQ ID NO: 36. In some embodiments, the tag sequence comprises two or more (e.g., two, three, four) biotin ligase recognition sequences.
[0033] In some embodiments, the tag sequence comprises at least one polyhistidine tag and at least one biotin ligase recognition sequence. In some embodiments, the tag sequence comprises at least one polyhistidine tag and at least two biotin ligase recognition sequences. In some embodiments, the tag sequence comprises at least two polyhistidine tags and at least one biotin ligase recognition sequence.
[0034] In some embodiments, the tag sequence comprises a solubility tag, hi some embodiments, the tag sequence comprises a tag peptide or tag protein. Examples of tag peptides include calmodulin-binding peptide (CBP) tag, FLAG epitope, human influenza hemagglutinin (HA) tag, Myc epitope, streptavidin-binding peptide, Strep tag, Strep-II tag, intrinsically disordered tag, Fasciola hepatica 8 kDa antigen (Fh8), maltose-binding protein (MBP), N-utilizing substance (NusA), thioredoxin (Trx), small ubiquitin-like modifier (SUMO), glutathione-S-transferase (GST), solubility enhancer peptide sequence (SET), IgG domain B1 of protein G (GB1), IgG repeat domain ZZ of protein A (ZZ), mutant dehalogenase (HaloTag), and solubility-enhancing ubiquitous tag. Examples of tags include, but are not limited to, SNUT (Synaptic Targeting Tag), 17-kilodalton protein (Skp), bacteriophage V5 epitope, phage T7 protein kinase (T7PK), E. coli secreted protein A (EspA), monomeric bacteriophage T7 0.3 protein (Orc protein; Mocr), E. coli trypsin inhibitor (Ecotin), calcium-binding protein (CaBP), stress-responsive arsenate reductase (ArsC), N-terminal fragment of translation initiation factor IF2 (IF2-domain I), stress-responsive proteins (e.g., RpoA, SlyD, Tsf, RpoS, PotD, Crr), and E. coli acidic proteins (e.g., msyB, yjgD, rpoD). Additional examples of tags are known in the art and can be used in accordance with the present disclosure.See, for example, Costa, S., et al. "Fusion tags for protein solubility, purification and immunogenicity in Escherichia coli: the novel Fh8 system." Front Microbiol. 2014 Feb. 19;5:63; and Kimple, ME, et al. "Overview of Affinity Tags for Protein Purification." Curr Protoc Protein Sci. 2013;73:Unit9-9 (the relevant contents of which are incorporated herein by reference).
[0035] In some embodiments, the tag sequence comprises a linker. For example, in some embodiments, the tag sequence comprises a polyhistidine tag and a linker between the polyhistidine tag and amino acids of the aminopeptidase. In some embodiments, the tag sequence comprises a biotin ligase recognition sequence and a linker between the biotin ligase recognition sequence and amino acids of the aminopeptidase. In some embodiments, the tag sequence comprises a biotin ligase recognition sequence, a polyhistidine tag, and a linker between the biotin ligase recognition sequence and the polyhistidine tag. In some embodiments, the tag sequence comprises two biotin ligase recognition sequences and a linker between the two biotin ligase recognition sequences.
[0036] In some embodiments, the linker of the tag sequence comprises one or more amino acids. In some embodiments, the linker comprises at least 2, at least 3, at least 4, at least 5, at least 6, at least 8, at least 10, at least 12, at least 15, at least 30, at least 25, at least 30, or more amino acids. In some embodiments, the linker comprises about 1 to about 50 amino acids (e.g., 1 to 50 amino acids, 1 to 30 amino acids, 2 to 25 amino acids, 3 to 30 amino acids, 6 to 25 amino acids, 2 to 15 amino acids, 1 to 12 amino acids).
[0037] In some embodiments, the linker of the tag sequence comprises at least one glycine amino acid. In some embodiments, the linker comprises at least one glycine-serine (GS) motif. In some embodiments, the linker has the following formula: (G m S) n wherein G is glycine, S is serine, m is an integer between 1 and 5 (inclusive), and n is an integer between 1 and 6 (inclusive). In some embodiments, m is an integer between 1 and 3 (inclusive). In some embodiments, m is 2 or 3. In some embodiments, m is 2. In some embodiments, m is 3. In some embodiments, n is an integer between 1 and 4 (inclusive). In some embodiments, n is an integer between 1 and 3 (inclusive). In some embodiments, n is 1 or 3. In some embodiments, n is 1. In some embodiments, n is 3.
[0038] In some embodiments, the cleavage reagent comprises a tag sequence having an amino acid sequence selected from Table 2. In some embodiments, the cleavage reagent comprises a tag sequence having an amino acid sequence that is at least 40% identical to an amino acid sequence selected from Table 2. In some embodiments, the tag sequence has at least 45%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 92%, at least 94%, at least 95%, at least 96%, at least 98% or more amino acid sequence identity to an amino acid sequence selected from Table 2. In some embodiments, the tag sequence has 25 to 50%, 50 to 60%, 60 to 70%, 70 to 80%, 80 to 90%, 90 to 95%, 92 to 99%, 94 to 99%, 95 to 99%, 40 to 100%, 50 to 100%, 60 to 100%, 70 to 100%, 80 to 100%, 90 to 100%, 92 to 100%, 94 to 100%, 95 to 100%, 96 to 100%, or 100% amino acid sequence identity to an amino acid sequence selected from Table 2.
[0039] Thus, in some embodiments, the cleavage reagent comprises an aminopeptidase and a tag sequence, wherein the aminopeptidase has an amino acid sequence selected from Table 1 and the tag sequence has an amino acid sequence selected from Table 2. In some embodiments, the cleavage reagent comprises an aminopeptidase and a tag sequence, wherein the aminopeptidase has an amino acid sequence that is at least 80% identical to an amino acid sequence selected from Table 1 and the tag sequence has an amino acid sequence that is at least 40% identical to an amino acid sequence selected from Table 2.
[0040] In some embodiments, the present disclosure provides a single polypeptide comprising an aminopeptidase described herein linked to a tag sequence described herein. For example, in some embodiments, a cleavage reagent comprises a fusion polypeptide of an aminopeptidase fused to a tag sequence. In some embodiments, the fusion polypeptide comprises an aminopeptidase and a tag sequence fused to the terminus of the aminopeptidase. In some embodiments, the fusion polypeptide comprises an aminopeptidase and a tag sequence fused to the C-terminus of the aminopeptidase. In some embodiments, the fusion polypeptide comprises an aminopeptidase and a tag sequence fused to the N-terminus of the aminopeptidase. In some embodiments, the fusion polypeptide comprises an aminopeptidase and a tag sequence fused to the N-terminus of the aminopeptidase. In some embodiments, the present disclosure provides a nucleic acid encoding a cleavage reagent described herein. In some embodiments, the nucleic acid is an expression construct encoding a fusion polypeptide of an aminopeptidase fused to a tag sequence.
[0041] Compositions and reaction mixtures In some embodiments, the present disclosure provides a composition comprising at least one cleavage reagent as described herein. In some embodiments, the composition comprises two or more cleavage reagents, wherein at least one cleavage reagent comprises an aminopeptidase and a tag sequence as described herein. In some embodiments, the composition comprises two or more cleavage reagents as described herein.
[0042] In some aspects, the present disclosure provides a composition comprising a first cleavage reagent comprising a first aminopeptidase from Pyrococcus horikoshii and a tag sequence described herein (e.g., a first tag sequence), and a second cleavage reagent comprising a second aminopeptidase from Pyrococcus horikoshii, wherein the first and second aminopeptidases have different amino acid sequences. In some embodiments, the amino acid sequences of the first and second aminopeptidases share less than 80% sequence identity (e.g., less than 70%, less than 60%, less than 50%, 10-80%, 20-60%, 30-50%, or 40-50% sequence identity). In some embodiments, the second cleavage reagent comprises a tag sequence described herein (e.g., a second tag sequence). In some embodiments, the composition comprises a third cleavage reagent. In some embodiments, the third cleavage reagent comprises an aminopeptidase from Yersinia pestis.
[0043] In some embodiments, the first and second cleavage reagents are present in the composition at first and second concentrations, respectively, where the first concentration is at least two-fold higher than the second concentration. In some embodiments, the molar ratio of the first cleavage reagent to the second cleavage reagent in the composition is about 2:1 to about 20:1 (e.g., about 2:1 to about 15:1, about 2:1 to about 10:1, about 4:1 to about 15:1, or about 5:1 to about 10:1). In some embodiments, the molar ratio of the first cleavage reagent to the second cleavage reagent in the composition is about 2:1, about 4:1, about 6:1, about 8:1, or about 10:1.
[0044] In some embodiments, the first cleavage reagent is present at a concentration of about 10 μM to about 100 μM (e.g., 20-80 μM, 20-60 μM, 10-30 μM, 30-50 μM, 50-70 μM, 70-90 μM). In some embodiments, the first cleavage reagent is present at a concentration of about 20 μM, about 40 μM, about 60 μM, or about 80 μM. In some embodiments, the first cleavage reagent is present in an amount sufficient to cleave the N-terminal amino acid from the polypeptide in an average cleavage time of about 2 to about 60 minutes (e.g., 5-50 minutes, 5-30 minutes, 5-20 minutes, 5-15 minutes, 5-10 minutes, 2-30 minutes, 10-30 minutes, 30-60 minutes), wherein the N-terminal amino acid comprises a charged side chain (e.g., arginine, lysine, glutamine, aspartic acid, glutamic acid).
[0045] In some embodiments, the second cleavage reagent is present at a concentration of about 0.1 μM to about 25 μM (e.g., 0.1 to 10 μM, 0.5 to 20 μM, 1 to 10 μM, 2 to 20 μM, 2 to 15 μM, 1 to 8 μM). In some embodiments, the second cleavage reagent is present at a concentration of about 2 μM, about 4 μM, about 6 μM, or about 8 μM. In some embodiments, the second cleavage reagent is present in an amount sufficient to cleave the N-terminal amino acid from the polypeptide in an average cleavage time of about 2 to about 60 minutes (e.g., 5 to 50 minutes, 5 to 30 minutes, 5 to 20 minutes, 5 to 15 minutes, 5 to 10 minutes, 2 to 30 minutes, 10 to 30 minutes, 30 to 60 minutes), wherein the N-terminal amino acid comprises a hydrophobic side chain (e.g., alanine, valine, isoleucine, leucine, methionine, phenylalanine, tyrosine, tryptophan).
[0046] In some embodiments, the first aminopeptidase comprises one or more substitutions compared to Pyrococcus horikoshii TET aminopeptidase III. In some embodiments, the first aminopeptidase has an amino acid sequence that is at least 80% (e.g., at least 85%, at least 90%, at least 95%, 80-100%, 85-95%, 90-99%, 95-99%, or 100%) identical to SEQ ID NO:3. In some embodiments, the first cleavage reagent is a polypeptide having an amino acid sequence that is at least 80% (e.g., at least 85%, at least 90%, at least 95%, 80-100%, 85-95%, 90-99%, 95-99%, or 100%) identical to SEQ ID NO:4.
[0047] In some embodiments, the second aminopeptidase comprises one or more substitutions compared to Pyrococcus horikoshii TET aminopeptidase II. In some embodiments, the second aminopeptidase has an amino acid sequence that is at least 80% (e.g., at least 85%, at least 90%, at least 95%, 80-100%, 85-95%, 90-99%, 95-99%, or 100%) identical to SEQ ID NO: 1. In some embodiments, the second cleavage reagent is a polypeptide having an amino acid sequence that is at least 80% (e.g., at least 85%, at least 90%, at least 95%, 80-100%, 85-95%, 90-99%, 95-99%, or 100%) identical to SEQ ID NO: 2.
[0048] In some embodiments, the present disclosure provides a composition comprising a first cleavage reagent comprising a first aminopeptidase having an amino acid sequence at least 80% identical to SEQ ID NO:3, a second cleavage reagent comprising a second aminopeptidase having an amino acid sequence at least 80% identical to SEQ ID NO:1, and a third cleavage reagent comprising a third aminopeptidase having an amino acid sequence at least 80% identical to SEQ ID NO:5 or 7.
[0049] In some embodiments, at least one of the first, second, and third cleavage reagents comprises a tag sequence described herein. For example, in some embodiments, the first cleavage reagent comprises a first tag sequence, the second cleavage reagent comprises a second tag sequence, and the third cleavage reagent comprises a third tag sequence. In some embodiments, the first tag sequence is attached to the end of the first aminopeptidase, the second tag sequence is attached to the end of the second aminopeptidase, and the third tag sequence is attached to the end of the third aminopeptidase. In some embodiments, each tag sequence is attached to the C-terminus of its respective aminopeptidase.
[0050] In some embodiments, the first, second, and third cleavage reagents are present in the composition at first, second, and third concentrations, respectively, where the first concentration is at least two-fold higher than the second concentration. In some embodiments, the first concentration is at least five-fold higher than the third concentration. In some embodiments, the second concentration is at least two-fold higher than the third concentration.
[0051] In some embodiments, the molar ratio of the first cleavage reagent to the second cleavage reagent in the composition is about 2:1 to about 20:1 (e.g., about 2:1 to about 15:1, about 2:1 to about 10:1, about 4:1 to about 15:1, or about 5:1 to about 10:1). In some embodiments, the molar ratio of the first cleavage reagent to the third cleavage reagent in the composition is about 5:1 to about 200:1 (e.g., about 5:1 to about 150:1, about 5:1 to about 100:1, about 10:1 to about 80:1, or about 10:1 to about 50:1).
[0052] In some embodiments, the first cleavage reagent is present at a concentration of about 10 μM to about 100 μM (e.g., 20-80 μM, 20-60 μM, 10-30 μM, 30-50 μM, 50-70 μM, 70-90 μM). In some embodiments, the first cleavage reagent is present in an amount sufficient to cleave the N-terminal amino acid from the polypeptide in an average cleavage time of about 2 to about 60 minutes (e.g., 5-50 minutes, 5-30 minutes, 5-20 minutes, 5-15 minutes, 5-10 minutes, 2-30 minutes, 10-30 minutes, 30-60 minutes), where the N-terminal amino acid comprises a charged side chain (e.g., arginine, lysine, glutamine, aspartic acid, glutamic acid).
[0053] In some embodiments, the second cleavage reagent is present at a concentration of about 0.1 μM to about 25 μM (e.g., 0.1 to 10 μM, 0.5 to 20 μM, 1 to 10 μM, 2 to 20 μM, 2 to 15 μM, 1 to 8 μM). In some embodiments, the second cleavage reagent is present in an amount sufficient to cleave the N-terminal amino acid from the polypeptide in an average cleavage time of about 2 to about 60 minutes (e.g., 5 to 50 minutes, 5 to 30 minutes, 5 to 20 minutes, 5 to 15 minutes, 5 to 10 minutes, 2 to 30 minutes, 10 to 30 minutes, 30 to 60 minutes), wherein the N-terminal amino acid comprises a hydrophobic side chain (e.g., alanine, valine, isoleucine, leucine, methionine, phenylalanine, tyrosine, tryptophan).
[0054] In some embodiments, the third cleavage reagent is present at a concentration of about 0.01 μM to about 25 μM (e.g., 0.1 to 10 μM, 0.5 to 20 μM, 1 to 10 μM, 2 to 20 μM, 2 to 15 μM, 1 to 8 μM). In some embodiments, the third cleavage reagent is present in an amount sufficient to cleave the N-terminal amino acid from a polypeptide in an average cleavage time of about 2 to about 60 minutes (e.g., 5 to 50 minutes, 5 to 30 minutes, 5 to 20 minutes, 5 to 15 minutes, 5 to 10 minutes, 2 to 30 minutes, 10 to 30 minutes, 30 to 60 minutes), wherein the polypeptide comprises an XP dipeptide motif, where X is the N-terminal amino acid and P is a proline amino acid.
[0055] In some embodiments, the amino acid sequence of the first aminopeptidase is at least 85%, at least 90%, at least 95%, 80-100%, 85-95%, 90-99%, 95-99%, or 100% identical to SEQ ID NO: 3. In some embodiments, the first cleavage reagent is a polypeptide having an amino acid sequence that is at least 80% (e.g., at least 85%, at least 90%, at least 95%, 80-100%, 85-95%, 90-99%, 95-99%, or 100%) identical to SEQ ID NO: 4.
[0056] In some embodiments, the amino acid sequence of the second aminopeptidase is at least 85%, at least 90%, at least 95%, 80-100%, 85-95%, 90-99%, 95-99%, or 100% identical to SEQ ID NO: 1. In some embodiments, the second cleavage reagent is a polypeptide having an amino acid sequence that is at least 80% (e.g., at least 85%, at least 90%, at least 95%, 80-100%, 85-95%, 90-99%, 95-99%, or 100%) identical to SEQ ID NO: 2.
[0057] In some embodiments, the amino acid sequence of the third aminopeptidase is at least 85%, at least 90%, at least 95%, 80-100%, 85-95%, 90-99%, 95-99%, or 100% identical to SEQ ID NO: 5 or 7. In some embodiments, the third cleavage reagent is a polypeptide having an amino acid sequence that is at least 80% (e.g., at least 85%, at least 90%, at least 95%, 80-100%, 85-95%, 90-99%, 95-99%, or 100%) identical to SEQ ID NO: 8.
[0058] In some aspects, the present disclosure provides a reaction mixture for polypeptide analysis, the reaction mixture comprising a composition described herein and one or more amino acid recognition factors (e.g., one or more amino acid binding proteins that do not have peptide-cleaving activity). In some embodiments, the amino acid recognition factor comprises an amino acid binding protein such as a ClpS protein (e.g., a Planctomycetia bacterium ClpS protein), a UBR protein (e.g., a Kluyveromyces marxianus UBR protein), an Ntaq1 protein (e.g., a Scleropages formosus Ntaq1 protein), or a variant or homolog thereof. In some embodiments, the amino acid recognition factor comprises a label (e.g., a detectable label, such as a luminescent label). Examples of amino acid recognition factors (e.g., recognition molecules) are described in detail in PCT International Publication No. WO2020 / 102741A1, filed November 15, 2019, PCT International Publication No. WO2021 / 236983A2, filed May 20, 2021, and co-pending U.S. Patent Application No. 63 / 395,328, filed August 4, 2022, the relevant contents of each of which are incorporated by reference in their entirety.
[0059] In some aspects, the disclosure provides a reaction mixture for polypeptide analysis comprising a first aminopeptidase having an amino acid sequence at least 92% identical to SEQ ID NO: 3 and comprising a first tag sequence, and one or more amino acid binding proteins that do not have peptide cleavage activity. In some embodiments, the amino acid sequence of the first aminopeptidase is at least 94%, at least 96%, at least 98%, or 100% identical to SEQ ID NO: 3. In some embodiments, the reaction mixture further comprises one or more cleavage reagents described herein.
[0060] As described herein, the compositions and reaction mixtures of the present disclosure can be used to determine at least one chemical characteristic of a polypeptide based on a characteristic pattern. In some embodiments, polypeptide sequencing reaction conditions can be configured to achieve a time interval that allows for sufficient related events to provide a desired level of confidence that the characteristic pattern has been identified. This can be achieved by configuring the reaction conditions based on various characteristics, including, for example, reagent concentration, molar ratio of one reagent to another (e.g., ratio of amino acid recognition factor to cleavage reagent, ratio of one recognition factor to another recognition factor, ratio of one cleavage reagent to another cleavage reagent), number of different reagent types (e.g., number of different types of recognition factors and / or cleavage reagents, number of recognition factor types relative to number of cleavage reagent types), cleavage activity (e.g., aminopeptidase activity), binding characteristics (e.g., kinetic and / or thermodynamic binding parameters for recognition molecule binding), reagent modifications (e.g., polyols and other recognition factor modifications that can alter interaction kinetics), reaction mixture components (e.g., one or more components such as pH, buffers, salts, divalent cations, surfactants, and other reaction mixture components described herein), temperature of the reaction, and various other parameters, as well as combinations thereof, that will be apparent to one of skill in the art. Reaction conditions can be configured based on one or more aspects described herein, including, for example, signal pulse information (e.g., pulse duration, inter-pulse duration, magnitude change), labeling strategy (e.g., number and / or type of fluorophores, linkers with or without shielding elements), surface modification (e.g., modification of sample well surface, including polypeptide immobilization), sample preparation (e.g., polypeptide fragment size, polypeptide modification for immobilization), and other aspects described herein.
[0061] In some embodiments, polypeptide sequencing reactions according to the present disclosure are performed under conditions that allow amino acid recognition and cleavage to occur simultaneously in a single reaction mixture. For example, in some embodiments, polypeptide sequencing reactions are performed in a reaction mixture having a pH that allows association and cleavage events to occur. Thus, in some embodiments, the reaction mixture has a pH of about 6.5 to about 9.0. In some embodiments, the reaction mixture has a pH of about 7.0 to about 8.5 (e.g., about 7.0 to about 8.0, about 7.5 to about 8.5, about 7.5 to about 8.0, or about 8.0 to about 8.5).
[0062] In some embodiments, the polypeptide sequencing reaction is carried out in a reaction mixture comprising one or more buffers. In some embodiments, the reaction mixture comprises a buffer at a concentration of at least 10 mM (e.g., at least 20 mM and up to 250 mM, at least 50 mM, 10-250 mM, 10-100 mM, 20-100 mM, 50-100 mM, or 100-200 mM). In some embodiments, the reaction mixture comprises a buffer at a concentration of about 10 mM to about 50 mM (e.g., about 10 mM to about 25 mM, about 25 mM to about 50 mM, or about 20 mM to about 40 mM). Examples of buffering agents include, but are not limited to, HEPES (4-(2-hydroxyethyl)-1-piperazineethanesulfonic acid), Tris (tris(hydroxymethyl)aminomethane), and MOPS (3-(N-morpholino)propanesulfonic acid).
[0063] In some embodiments, the polypeptide sequencing reaction is carried out in a reaction mixture comprising a salt concentration of at least 10 mM. In some embodiments, the reaction mixture comprises a salt concentration of at least 10 mM (e.g., at least 20 mM, at least 50 mM, at least 100 mM, or more). In some embodiments, the reaction mixture comprises a salt concentration of about 10 mM to about 250 mM (e.g., about 20 mM to about 200 mM, about 50 mM to about 150 mM, about 10 mM to about 50 mM, or about 10 mM to about 100 mM). Examples of salts include, but are not limited to, sodium salts, potassium salts, and acetate salts, such as sodium chloride (NaCl), sodium acetate (NaOAc), and potassium acetate (KOAc).
[0064] Further examples of components for use in the reaction mixture include divalent cations (e.g., Mg 2+ , Co 2+ ) and surfactants (e.g., polysorbate 20). In some embodiments, the reaction mixture contains a divalent cation at a concentration of about 0.1 mM to about 50 mM (e.g., about 10 mM to about 50 mM, about 0.1 mM to about 10 mM, or about 1 mM to about 20 mM). In some embodiments, the reaction mixture contains a surfactant at a concentration of at least 0.01% (e.g., about 0.01% to about 0.10%). In some embodiments, the reaction mixture contains one or more components useful in single molecule analysis, such as an oxygen scavenging system (e.g., a PCA / PCD system or a pyranose oxidase / catalase / glucose system) and / or one or more triplet state quenchers (e.g., Trolox®, COT, and NBA).
[0065] In some embodiments, the polypeptide sequencing reaction is performed at a temperature at which association and cleavage events can occur. In some embodiments, the polypeptide sequencing reaction is performed at a temperature of at least 10°C. In some embodiments, the polypeptide sequencing reaction is performed at a temperature of about 10°C to about 50°C (e.g., 15-45°C, 20-40°C, 25°C or near 25°C, 30°C or near 30°C, 35°C or near 35°C, 37°C or near 37°C). In some embodiments, the polypeptide sequencing reaction is performed at or near room temperature.
[0066] As detailed above, a real-time sequencing process such as that illustrated by FIG. 1A can generally involve cycles of amino acid recognition and terminal amino acid cleavage. In some embodiments, the relative occurrence of recognition and cleavage can be controlled by the concentration difference between one or more amino acid recognition factors and at least one cleavage reagent. In some embodiments, the concentration difference can be optimized so that the number of signal pulses detected during recognition of individual amino acids provides a desired confidence interval for discrimination. For example, if an initial sequencing reaction provides signal data with too few signal pulses between cleavage events to allow for the determination of a characteristic pattern with a desired confidence interval, the sequencing reaction can be repeated using a reduced concentration of non-specific exopeptidase relative to the recognition molecule.
[0067] In some embodiments, polypeptide analysis according to the present disclosure can be performed by contacting a polypeptide with a reaction mixture comprising one or more amino acid recognition factors and one or more cleavage reagents (e.g., peptidases). In some embodiments, the reaction mixture comprises the amino acid recognition factors at a concentration of about 10 nM to about 10 μM. In some embodiments, the reaction mixture comprises the cleavage reagent at a concentration of about 500 nM to about 500 μM.
[0068] In some embodiments, the reaction mixture contains an amino acid recognition factor at a concentration of about 100 nM to about 10 μM, about 250 nM to about 10 μM, about 100 nM to about 1 μM, about 250 nM to about 1 μM, about 250 nM to about 750 nM, or about 500 nM to about 1 μM. In some embodiments, the reaction mixture contains an amino acid recognition factor at a concentration of about 100 nM, about 250 nM, about 500 nM, about 750 nM, or about 1 μM. In some embodiments, the reaction mixture contains a cleavage reagent at a concentration of about 500 nM to about 250 μM, about 500 nM to about 100 μM, about 1 μM to about 100 μM, about 500 nM to about 50 μM, about 1 μM to about 100 μM, about 10 μM to about 200 μM, or about 10 μM to about 100 μM. In some embodiments, the reaction mixture comprises the cleavage reagent at a concentration of about 1 μM, about 5 μM, about 10 μM, about 30 μM, about 50 μM, about 70 μM, or about 100 μM.
[0069] In some embodiments, the reaction mixture comprises an amino acid recognition factor at a concentration of about 10 nM to about 10 μM and a cleavage reagent at a concentration of about 500 nM to about 500 μM. In some embodiments, the reaction mixture comprises an amino acid recognition factor at a concentration of about 100 nM to about 1 μM and a cleavage reagent at a concentration of about 1 μM to about 100 μM. In some embodiments, the reaction mixture comprises an amino acid recognition factor at a concentration of about 250 nM to about 1 μM and a cleavage reagent at a concentration of about 10 μM to about 100 μM. In some embodiments, the reaction mixture comprises an amino acid recognition factor at a concentration of about 500 nM and a cleavage reagent at a concentration of about 25 μM to about 75 μM. In some embodiments, the concentrations of the amino acid recognition factor and / or the cleavage reagent in the reaction mixture are as described elsewhere herein.
[0070] In some embodiments, the reaction mixture comprises the amino acid recognition factor and the cleavage reagent in a molar ratio of about 500:1, about 400:1, about 300:1, about 200:1, about 100:1, about 75:1, about 50:1, about 25:1, about 10:1, about 5:1, about 2:1, or about 1:1. In some embodiments, the reaction mixture comprises the amino acid recognition factor and the cleavage reagent in a molar ratio of about 10:1 to about 200:1. In some embodiments, the reaction mixture comprises the amino acid recognition factor and the cleavage reagent in a molar ratio of about 50:1 to about 150:1. In some embodiments, the molar ratio of amino acid recognition factor to cleavage reagent in the reaction mixture is about 1:1,000 to about 1:1 or about 1:1 to about 100:1 (e.g., 1:1,000, about 1:500, about 1:200, about 1:100, about 1:10, about 1:5, about 1:2, about 1:1, about 5:1, about 10:1, about 50:1, about 100:1). In some embodiments, the molar ratio of amino acid recognition factor to cleavage reagent in the reaction mixture is about 1:100 to about 1:1 or about 1:1 to about 10:1. In some embodiments, the molar ratio of amino acid recognition factor to cleavage reagent in the reaction mixture is as described elsewhere herein.
[0071] In some embodiments, the reaction mixture comprises one or more amino acid recognition factors and one or more cleavage reagents described herein. In some embodiments, the reaction mixture comprises at least three amino acid recognition factors and at least one cleavage reagent. In some embodiments, the reaction mixture comprises two or more cleavage reagents. In some embodiments, the reaction mixture comprises at least one and up to 10 cleavage reagents (e.g., 1-3 cleavage reagents, 2-10 cleavage reagents, 1-5 cleavage reagents, 3-10 cleavage reagents). In some embodiments, the reaction mixture comprises at least 3 and up to 30 amino acid recognition factors (e.g., 3-25, 3-20, 3-10, 3-5, 5-30, 5-20, 5-10, or 10-20 amino acid recognition factors).
[0072] In some embodiments, a reaction mixture contains more than one amino acid recognition factor and / or more than one cleavage reagent. In some embodiments, a reaction mixture described as containing more than one amino acid recognition factor or cleavage reagent refers to a mixture having more than one type of amino acid recognition factor or cleavage reagent. For example, in some embodiments, a reaction mixture contains two or more cleavage reagents, and the two or more cleavage reagents refer to two or more types of aminopeptidases. In some embodiments, one type of aminopeptidase has an amino acid sequence that is different from the amino acids or subset of amino acids cleaved by another type of cleavage reagent in the reaction mixture. In some embodiments, one type of cleavage reagent cleaves an amino acid or subset of amino acids that is different from the amino acids or subset of amino acids cleaved by another type of cleavage reagent in the reaction mixture.
[0073] Polypeptide analysis In some aspects, the present disclosure provides a method of polypeptide analysis (e.g., polypeptide sequencing). In some embodiments, the method of polypeptide analysis includes contacting a polypeptide with a reaction mixture described herein, monitoring a signal for a signal pulse corresponding to an interaction between one or more amino acid binding proteins and the polypeptide, and determining at least one chemical characteristic of the polypeptide based on a characteristic pattern in the signal.
[0074] A non-limiting example of polypeptide structural analysis by detecting single-molecule binding interactions during the polypeptide degradation process is shown in FIG. 1A. Exemplary signal traces are shown showing different association (e.g., binding) events at times corresponding to signal changes. As shown, association events between an amino acid recognition factor and the terminus of a polypeptide result in a change in signal magnitude that persists for a period of time. Different association events are shown for different amino acids exposed at the terminus of a polypeptide. As described herein, an amino acid that is "exposed" at the terminus of a polypeptide is an amino acid that is still bound to the polypeptide and that becomes the terminal amino acid upon removal of the previous terminal amino acid (e.g., alone or together with one or more additional amino acids) during degradation.
[0075] As generally indicated, association events between amino acid recognition factors and different types of amino acids at the termini of a polypeptide result in characteristic changes in signals, referred to herein as signature patterns, which can be used to determine the chemical characteristics of the polypeptide. In some embodiments, a signature pattern corresponding to one type of terminal amino acid can be used to determine structural information about the terminal amino acid and one or more amino acids adjacent to the terminal amino acid. Thus, in some embodiments, a signature pattern corresponding to one type of terminal amino acid can be used to determine structural information about at least two (e.g., at least three, at least four, at least five, two, three, four, or two to five) amino acids of a polypeptide.
[0076] In some embodiments, a transition from one characteristic pattern to another indicates an amino acid cleavage. As used herein, in some embodiments, amino acid cleavage refers to the removal of at least one amino acid from the end of a polypeptide (e.g., the removal of at least one terminal amino acid from a polypeptide). In some embodiments, amino acid cleavage is determined by inference based on the duration between characteristic patterns. In some embodiments, amino acid cleavage is determined by detecting a change in signal resulting from the association of a labeled cleavage reagent with an amino acid at the end of the polypeptide. As amino acids are sequentially cleaved from the end of the polypeptide during degradation, a series of changes in size, or a series of signal pulses, is detected.
[0077] In some embodiments, the signal data may be analyzed to extract signal pulse information by applying a threshold level to one or more parameters of the signal data. For example, in some embodiments, a threshold magnitude level may be applied to the signal data of a signal trace. In some embodiments, the threshold magnitude level is the minimum difference between the signal detected at a given time point and a baseline determined for a given data set. In some embodiments, a signal pulse that exhibits a change in magnitude exceeding the threshold magnitude level and lasts for a certain duration is assigned to each portion of the data. In some embodiments, a threshold duration may be applied to portions of the data that meet the threshold magnitude level to determine whether a signal pulse is assigned to the portion of the data that meets the threshold magnitude level. For example, experimental artifacts may produce a change in magnitude that exceeds the threshold magnitude level, but this may not last for a sufficient duration to assign a signal pulse with a desired degree of confidence (e.g., a transient related event that may be non-discriminatory for amino acid type, e.g., a non-specific detection event such as diffusion into the observation region or reagent adhesion within the observation region). Thus, in some embodiments, signal pulses are extracted from the signal data based on a threshold magnitude level and a threshold duration.
[0078] In some embodiments, the signal pulse magnitude peak is determined by averaging the detected magnitude over a duration that persists above a threshold magnitude level. It should be understood that in some embodiments, a "signal pulse" as used herein can refer to a change in signal data (e.g., raw signal data) that persists for a duration above the baseline, or signal pulse information extracted therefrom (e.g., processed signal data).
[0079] In some embodiments, the signal pulse information can be analyzed to identify different types of amino acids in a polypeptide based on different characteristic patterns in a series of signal pulses. For example, as shown in Figure 1A, the signal pulse information indicates different types of amino acids (e.g., arginine, leucine, isoleucine, phenylalanine) at the end of the polypeptide. For example, the signal pulse detected at the earliest time point provides information indicating (at least) arginine at the end of the polypeptide based on a first characteristic pattern, and the signal pulse detected at the latest time point provides information indicating at least phenylalanine at the end of the polypeptide based on a second characteristic pattern.
[0080] In some embodiments, each signal pulse of the characteristic pattern comprises a pulse duration corresponding to an association event between the amino acid recognition factor and the amino acid ligand. In some embodiments, the pulse duration is characteristic of the dissociation rate of binding. In some embodiments, each signal pulse of the characteristic pattern is separated from another signal pulse of the characteristic pattern by an inter-pulse duration. In some embodiments, the inter-pulse duration is characteristic of the association rate of binding. In some embodiments, a change in signal magnitude can be determined for the same signal pulse based on the difference between the baseline and peak of the signal pulse. In some embodiments, the characteristic pattern is determined based on the pulse duration. In some embodiments, the characteristic pattern is determined based on the pulse duration and the inter-pulse duration. In some embodiments, the characteristic pattern is determined based on any one or more of the pulse duration, the inter-pulse duration, and the change in magnitude.
[0081] 1A, in some embodiments, polypeptide analysis is performed by detecting a series of signal pulses that indicate the association of one or more amino acid recognition factors with consecutive amino acids exposed at the termini of the polypeptide in an ongoing degradation reaction. The series of signal pulses can be analyzed to determine a characteristic pattern in the series of signal pulses, and the time course of the characteristic pattern can be used to determine the chemical signature across the amino acid sequence of the polypeptide.
[0082] As described herein, signal pulse information can be used to identify amino acids based on a characteristic pattern in a series of signal pulses. In some embodiments, the characteristic pattern includes a plurality of signal pulses, each signal pulse including a pulse duration. In some embodiments, the plurality of signal pulses can be characterized by a summary statistic (e.g., mean, median, time decay constant) of the distribution of pulse durations in the characteristic pattern. In some embodiments, the average pulse duration of the characteristic pattern is about 1 millisecond to about 10 seconds (e.g., about 1 ms to about 1 s, about 1 ms to about 100 ms, about 1 ms to about 10 ms, about 10 ms to about 10 s, about 100 ms to about 10 s, about 1 s to about 10 s, about 10 ms to about 100 ms, or about 100 ms to about 500 ms). In some embodiments, the average pulse duration is about 50 milliseconds to about 2 seconds, about 50 milliseconds to about 500 milliseconds, or about 500 milliseconds to about 2 seconds.
[0083] In some embodiments, different characteristic patterns corresponding to different types of amino acids in a single polypeptide can be distinguished from one another based on statistically significant differences in summary statistics. For example, in some embodiments, one characteristic pattern can be distinguished from another characteristic pattern based on a difference in mean pulse duration of at least 10 milliseconds (e.g., about 10 ms to about 10 s, about 10 ms to about 1 s, about 10 ms to about 100 ms, about 100 ms to about 10 s, about 1 s to about 10 s, or about 100 ms to about 1 s). In some embodiments, the difference in mean pulse duration is at least 50 ms, at least 100 ms, at least 250 ms, at least 500 ms, or more. In some embodiments, the difference in mean pulse duration is about 50 ms to about 1 s, about 50 ms to about 500 ms, about 50 ms to about 250 ms, about 100 ms to about 500 ms, about 250 ms to about 500 ms, or about 500 ms to about 1 s. In some embodiments, the mean pulse duration of one characteristic pattern differs from the mean pulse duration of another characteristic pattern by about 10-25%, 25-50%, 50-75%, 75-100%, or more than 100%, e.g., about 2-fold, 3-fold, 4-fold, 5-fold, or more. It should be understood that in some embodiments, smaller differences in mean pulse duration between different characteristic patterns may require more pulse durations within each characteristic pattern to be distinguished from one another with statistical reliability.
[0084] In some embodiments, a characteristic pattern generally refers to a plurality of association events between amino acids of a polypeptide and a means for binding amino acids (e.g., amino acid recognition molecules). In some embodiments, a characteristic pattern comprises at least 10 association events (e.g., at least 25, at least 50, at least 75, at least 100, at least 250, at least 500, at least 1,000, or more association events). In some embodiments, a characteristic pattern comprises about 10 to about 1,000 association events (e.g., about 10 to about 500 association events, about 10 to about 250 association events, about 10 to about 100 association events, or about 50 to about 500 association events). In some embodiments, a plurality of association events is detected as a plurality of signal pulses.
[0085] In some embodiments, a characteristic pattern refers to a plurality of signal pulses that can be characterized by summary statistics described herein. In some embodiments, a characteristic pattern includes at least 10 signal pulses (e.g., at least 25, at least 50, at least 75, at least 100, at least 250, at least 500, at least 1,000, or more signal pulses). In some embodiments, a characteristic pattern includes about 10 to about 1,000 signal pulses (e.g., about 10 to about 500 signal pulses, about 10 to about 250 signal pulses, about 10 to about 100 signal pulses, or about 50 to about 500 signal pulses).
[0086] In some embodiments, a characteristic pattern refers to multiple association events between an amino acid recognition molecule and amino acids of a polypeptide that occur over a time interval prior to removal of an amino acid (e.g., a cleavage event). In some embodiments, a characteristic pattern refers to multiple association events that occur over a time interval between two cleavage events (e.g., before removal of an amino acid and after removal of a terminally previously exposed amino acid). In some embodiments, the time interval of a characteristic pattern is about 1 minute to about 30 minutes (e.g., about 1 minute to about 20 minutes, about 1 minute to about 10 minutes, about 5 minutes to about 20 minutes, about 5 minutes to about 15 minutes, or about 5 minutes to about 10 minutes).
[0087] In some embodiments, the series of signal pulses comprises a series of changes in the magnitude of the light signal over time. In some embodiments, the series of changes in the light signal comprises a series of changes in the luminescence generated during the associated event. In some embodiments, the luminescence is generated by a detectable label associated with one or more reagents of the sequencing reaction. For example, in some embodiments, each of the one or more amino acid recognition factors comprises a luminescent label. In some embodiments, the cleavage reagent comprises a luminescent label. Examples of luminescent labels and their uses according to the present disclosure are provided herein.
[0088] In some embodiments, the series of signal pulses comprises a series of changes in the magnitude of the electrical signal over time. In some embodiments, the series of changes in the electrical signal comprises a series of changes in conductance generated during the relevant events. In some embodiments, the conductivity is provided by a detectable label associated with one or more reagents of the sequencing reaction. For example, in some embodiments, each of the one or more amino acid recognition factors comprises a conductive label. Examples of conductive labels and their uses according to the present disclosure are provided elsewhere herein. Methods for identifying single molecules using conductive labels have been described (see, e.g., U.S. Patent Application Publication No. 2017 / 0037462).
[0089] In some embodiments, the series of conductance changes comprises a series of changes in conductance through a nanopore. For example, methods for assessing receptor-ligand interactions using nanopores have been described (see, e.g., Thakur, A.K. and Movileanu, L., (2019) Nature Biotechnology 37(1)). The inventors of the present application have recognized and appreciated that such nanopores can be used to monitor polypeptide sequencing reactions according to the present disclosure. Accordingly, in some embodiments, the present disclosure provides a method of polypeptide analysis comprising contacting a single polypeptide molecule with one or more amino acid recognition factors described herein, wherein the single polypeptide molecule is immobilized in a nanopore. In some embodiments, the method further comprises detecting a series of changes in conductance through the nanopore while the single polypeptide is being degraded, the changes indicating association of the one or more amino acid recognition factors with consecutive amino acids exposed at a terminus of the single polypeptide.
[0090] As described herein, in some embodiments, the amino acid recognition elements of the present disclosure can be used to determine at least one chemical characteristic of a polypeptide. In some embodiments, determining at least one chemical characteristic includes determining the type of amino acid present at the terminal end of the polypeptide and / or the type of amino acid present at one or more positions adjacent to the terminal amino acid. In some embodiments, determining the type of amino acid includes determining the actual amino acid identity, for example, by determining which of the 20 naturally occurring amino acids is present. In some embodiments, the type of amino acid is selected from alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, selenocysteine, serine, threonine, tryptophan, tyrosine, and valine.
[0091] In some embodiments, determining at least one chemical characteristic of the polypeptide includes determining a subset of potential amino acids that may be present in the polypeptide. In some embodiments, this can be achieved by determining that an amino acid is not one or more specific amino acids (and therefore may be any other amino acid). In some embodiments, this can be achieved by determining which of a specific subset of amino acids may be present in the polypeptide (e.g., based on size, charge, hydrophobicity, post-translational modification, binding properties) (e.g., using a recognition element that binds to a specific subset of two or more amino acids).
[0092] In some embodiments, determining at least one chemical characteristic of the polypeptide includes determining that the amino acid comprises a post-translational modification. Non-limiting examples of post-translational modifications include acetylation (e.g., acetylated lysine), ADP-ribosylation, caspase cleavage, citrullination, formylation, N-linked glycosylation (e.g., glycosylated asparagine), O-linked glycosylation (e.g., glycosylated serine, glycosylated threonine), hydroxylation, methylation (e.g., methylated lysine, methylated arginine), myristoylation (e.g., myristoylated glycine), neddylation, nitration ( For example, these include nitrated tyrosine, chlorination (e.g., chlorinated tyrosine), oxidation / reduction (e.g., oxidized cysteine, oxidized methionine), palmitoylation (e.g., palmitoylated cysteine), phosphorylation, prenylation (e.g., prenylated cysteine), S-nitrosylation (e.g., S-nitrosylated cysteine, S-nitrosylated methionine), sulfation, sumoylation (e.g., sumoylated lysine), and ubiquitination (e.g., ubiquitinated lysine).
[0093] In some embodiments, determining at least one chemical characteristic of the polypeptide includes determining that the amino acid comprises an arginine post-translational modification. For example, as described herein, the amino acid recognition element of the present disclosure can distinguish between different arginine modifications, including symmetric dimethylarginine (SDMA), asymmetric dimethylarginine (ADMA), and citrullinated arginine.
[0094] In some embodiments, determining at least one chemical characteristic of the polypeptide comprises determining that the amino acid comprises a phosphorylated side chain. For example, in some embodiments, determining at least one chemical characteristic of the polypeptide comprises determining that the amino acid comprises a phosphorylated threonine (e.g., phospho-threonine). In some embodiments, determining at least one chemical characteristic of the polypeptide comprises determining that the amino acid comprises a phosphorylated tyrosine (e.g., phosphotyrosine). In some embodiments, determining at least one chemical characteristic of the polypeptide comprises determining that the amino acid comprises a phosphorylated serine (e.g., phosphoserine).
[0095] In some embodiments, determining at least one chemical characteristic of the polypeptide includes determining that the amino acids include chemically modified variants, unnatural amino acids, or proteinogenic amino acids such as selenocysteine and pyrrolysine. Examples of unnatural amino acids include, but are not limited to, 2-naphthyl-alanine, statins, homoalanine, α-amino acids, β2-amino acids, β3-amino acids, γ-amino acids, 3-pyridyl-alanine, 4-fluorophenyl-alanine, cyclohexyl-alanine, N-alkyl amino acids, peptoid amino acids, homo-cysteine, penicillamine, 3-nitro-tyrosine, homo-phenyl-alanine, t-leucine, hydroxy-proline, 3-Abz, 5-F-tryptophan, and azabicyclo-[2.2.1]heptane.
[0096] In some embodiments, determining at least one chemical characteristic of the polypeptide includes determining that the amino acid comprises an oxidative modification. For example, as described herein, the amino acid recognition elements of the present disclosure can distinguish between oxidized methionine and its unmodified variant. In some embodiments, the oxidative modification comprises an oxidatively damaged side chain of the amino acid. In some embodiments, the oxidatively damaged side chain is selected from the group consisting of cysteine-derived products (e.g., disulfides, sulfinic acids, sulfonic acids, sulfenic acids, S-nitrosocysteine), tyrosine-derived products (e.g., di-tyrosine, 3,4-dihydroxyphenylalanine, 3-chlorotyrosine, 3-nitrotyrosine), histidine-derived products (e.g., 2-oxohistidine, 4-hydroxy-2-oxohistidine, di-histidine, asparagine, aspartic acid, urea), methionine-derived products, and the like. Oxidatively damaged amino acids include oxidatively damaged amino acids (e.g., sulfoxides, sulfones), tryptophan-derived products (e.g., di-tryptophan, N-formylkynurenine, kynurenine, 2-oxo-tryptophan, oxindolylalanine, 6-nitrotryptophan, hydroxytryptophan), phenylalanine-derived products (e.g., meta-tyrosine, ortho-tyrosine), or common side chain products (e.g., alcohols, hydroperoxides, aldehyde / ketone carbonyls). Examples of oxidatively damaged amino acids are known in the art; see, e.g., Hawkins, CL, Davies, MJ, Detection, identification, and quantification of oxidative protein modifications., J Biol Chem., 2019 Dec. 20, 294(51), 19683-19708.
[0097] In some embodiments, determining at least one chemical characteristic of the polypeptide includes determining that the amino acid comprises a side chain characterized by one or more biochemical properties. For example, the amino acid may comprise a nonpolar aliphatic side chain, a positively charged side chain, a negatively charged side chain, a nonpolar aromatic side chain, or a polar, uncharged side chain. Non-limiting examples of amino acids comprising nonpolar aliphatic side chains include alanine, glycine, valine, leucine, methionine, and isoleucine. Non-limiting examples of amino acids comprising positively charged side chains include lysine, arginine, and histidine. Non-limiting examples of amino acids comprising negatively charged side chains include aspartic acid and glutamic acid. Non-limiting examples of amino acids comprising nonpolar aromatic side chains include phenylalanine, tyrosine, and tryptophan. Non-limiting examples of amino acids comprising polar, uncharged side chains include serine, threonine, cysteine, proline, asparagine, and glutamine.
[0098] In some embodiments, a protein or polypeptide may be digested into multiple smaller polypeptides and chemical characteristics may be determined for one or more of these smaller polypeptides. In some embodiments, a first end (e.g., the N- or C-terminus) of the polypeptide is immobilized and the other end (e.g., the C- or N-terminus) is analyzed as described herein.
[0099] As used herein, sequencing a polypeptide refers to determining sequence information of a polypeptide. In some embodiments, this may include determining the identity of each consecutive amino acid for part (or all) of a polypeptide. However, in some embodiments, this may include assessing the identity of a subset of amino acids within a polypeptide (e.g., determining the relative positions of one or more amino acid types without determining the identity of each amino acid in the polypeptide). However, in some embodiments, amino acid content information can be obtained from a polypeptide without directly determining the relative positions of different types of amino acids in the polypeptide. Amino acid content alone can be used to infer the identity of existing polypeptides (e.g., by comparing amino acid content to a database of polypeptide information and determining which polypeptides have the same amino acid content).
[0100] In some embodiments, sequence information for multiple polypeptide products obtained from a longer polypeptide or protein (e.g., via enzymatic and / or chemical cleavage) can be analyzed to reconstruct or infer the sequence of the longer polypeptide or protein.
[0101] In some embodiments, the polypeptide analysis methods described herein generate data showing how a polypeptide interacts with a binding means while the polypeptide is being degraded by the cleavage means. As described above, the data can include a series of characteristic patterns corresponding to association events at the ends of the polypeptide between the cleavage events at the ends. In some embodiments, the methods of polypeptide analysis described herein include contacting a single polypeptide molecule with a binding means and a cleavage means, wherein the binding means and the cleavage means are configured to achieve at least 10 association events before the cleavage event. In some embodiments, the means are configured to achieve at least 10 association events between two cleavage events.
[0102] In some embodiments, multiple single molecule sequencing reactions are performed in parallel in an array of sample wells. In some embodiments, the array comprises about 10,000 to about 1,000,000 sample wells. The volume of a sample well, in some implementations, is about 10 -21 liters ~ approx. 10 -15 liters. Because the sample wells have small volumes, only about one polypeptide may be present in the sample well at any given time, allowing for detection of single molecule events. Statistically, some sample wells may not contain a single molecule sequencing reaction, and some may contain two or more single polypeptide molecules. However, a significant number of sample wells may each contain a single molecule reaction (e.g., in some embodiments, at least 30%), and thus single molecule analysis may be performed in parallel on a large number of sample wells. In some embodiments, the binding and cleavage means are configured to achieve at least 10 association events before a cleavage event in at least 10% (e.g., 10-50%, more than 50%, 25-75%, at least 80%, or more) of the sample wells in which a single molecule reaction is occurring. In some embodiments, the binding and cleavage means are configured to achieve at least 10 association events before a cleavage event in at least 50% (e.g., more than 50%, 50-75%, at least 80%, or more) of the amino acids of the polypeptide in the single molecule reaction.
[0103] Devices and Systems In some embodiments, methods according to the present disclosure can be implemented using a system that enables single-molecule analysis. The system can include an integrated device and an instrument configured to interface with the integrated device. The integrated device can include an array of pixels, each pixel including a sample well and at least one photodetector. The sample wells of the integrated device can be formed on or through a surface of the integrated device and configured to receive a sample disposed on the surface of the integrated device. Collectively, the sample wells can be considered an array of sample wells. The multiple sample wells can have a size and shape suitable for at least a portion of the sample wells to receive a single sample (e.g., a single molecule such as a polypeptide). In some embodiments, the number of samples in the sample wells can be distributed among the sample wells of the integrated device such that some sample wells contain one sample while other sample wells contain zero, two, or more samples.
[0104] Excitation light is provided to the integrated device from one or more light sources external to the integrated device. Optical components of the integrated device can receive the excitation light from the light source and direct the light toward the array of sample wells in the integrated device, illuminating an illumination region within the sample well. In some embodiments, the sample wells can have a configuration that allows a sample to be held in close proximity to the surface of the sample well, which can facilitate delivery of excitation light to the sample and detection of emission light from the sample. A sample positioned within the illumination region can emit emission light in response to being illuminated by the excitation light. For example, the sample can be labeled with a fluorescent label that emits light in response to achieving an excited state through illumination with the excitation light. The emission light emitted by the sample can then be detected by one or more photodetectors in pixels corresponding to the sample well with the sample being analyzed. According to some embodiments, multiple samples can be analyzed in parallel when performed across the entire array of sample wells, which can range in number from approximately 10,000 pixels to 1,000,000 pixels.
[0105] The integrated device may include an optical system for receiving excitation light and directing the excitation light among the sample well array. The optical system may include one or more grating couplers configured to couple the excitation light to other optical components of the integrated device and direct the excitation light to other optical components. For example, the optical system may include optical components that direct the excitation light from the grating coupler toward the sample well array. Such optical components may include an optical splitter, an optical combiner, and a waveguide. In some embodiments, one or more optical splitters can couple the excitation light from the grating coupler and deliver the excitation light to at least one waveguide. According to some embodiments, the optical splitter can have a configuration that enables substantially uniform delivery of excitation light across all waveguides, such that each of the waveguides receives substantially the same amount of excitation light. Such embodiments may improve the performance of the integrated device by improving the uniformity of the excitation light received by the sample wells of the integrated device. For example, examples of suitable components for coupling excitation light into a sample well and / or directing emitted light to a photodetector and for inclusion within an integrated device are described in U.S. patent application Ser. No. 14 / 821,688, filed Aug. 7, 2015, entitled "INTEGRATED DEVICE FOR PROBING, DETECTING AND ANALYZING MOLECULES," and U.S. patent application Ser. No. 14 / 543,865, filed Nov. 17, 2014, entitled "INTEGRATED DEVICE WITH EXTERNAL LIGHT SOURCE FOR PROBING, DETECTING, AND ANALYZING MOLECULES," both of which are incorporated by reference in their entireties. Examples of suitable grating couplers and waveguides that can be implemented in an integrated device are described in U.S. patent application Ser. No. 15 / 844,403, filed Dec. 15, 2017, entitled "OPTICAL COUPLER AND WAVEGUIDE SYSTEM," which is incorporated by reference in its entirety.
[0106] Additional photonic structures may be disposed between the sample well and the photodetector and configured to reduce or prevent excitation light from reaching the photodetector, which might otherwise contribute to signal noise when detecting the emission light. In some embodiments, the metal layer, which may act as a circuit for the integrated device, may also act as a spatial filter. Examples of suitable photonic structures may include spectral filters, polarization filters, and spatial filters, and are described in U.S. Patent Application No. 16 / 042,968, filed July 23, 2018, entitled "OPTICAL REJECTION PHOTONIC STRUCTURES," and U.S. Provisional Patent Application No. 63 / 124,655, filed December 11, 2020, entitled "INTEGRATED CIRCUIT WITH IMPROVED CHARGE TRANSFER EFFICIENCY AND ASSOCIATED TECHNIQUES," both of which are incorporated by reference in their entireties.
[0107] Components located remotely from the integrated device can be used to position and align the excitation source relative to the integrated device. Such components can include optical components, including lenses, mirrors, prisms, windows, apertures, attenuators, and / or optical fibers. Additional mechanical components can be included in the instrument to enable control of one or more alignment components. Such mechanical components can include actuators, stepper motors, and / or knobs. Examples of suitable excitation sources and alignment mechanisms are described in U.S. patent application Ser. No. 15 / 161,088, filed May 20, 2016, entitled "PULSED LASER AND SYSTEM," which is incorporated by reference in its entirety. Another example of a beam steering module is described in U.S. patent application Ser. No. 15 / 842,720, filed December 14, 2017, entitled "COMPACT BEAM SHAPING AND STEERING ASSEMBLY," which is incorporated by reference herein. Further examples of suitable excitation sources are described in U.S. Patent Application No. 14 / 821,688, filed August 7, 2015, entitled "INTEGRATED DEVICE FOR PROBING, DETECTING AND ANALYZING MOLECULES," which is incorporated by reference in its entirety.
[0108] The photodetector(s) associated with each pixel of the integrated device may be configured and arranged to detect light emission from the pixel's corresponding sample well. Examples of suitable photodetectors are described in U.S. Patent Application No. 14 / 821,656, filed August 7, 2015, entitled "INTEGRATED DEVICE FOR TEMPORAL BINNING OF RECEIVED PHOTONS," which is incorporated by reference in its entirety. In some embodiments, the sample wells and their respective photodetectors may be aligned along a common axis. In this manner, the photodetectors may overlap the sample wells within the pixel.
[0109] Characteristics of the detected emitted light can provide an indication for identifying a label associated with the emitted light. Such characteristics can include any suitable type of characteristic, including the arrival time of a photon detected by a photodetector, the amount of photons accumulated over time by a photodetector, and / or the distribution of photons across two or more photodetectors. In some embodiments, such characteristics can be any one or a combination of two or more of the following: luminescence lifetime, luminescence intensity, luminance, absorption spectrum, emission spectrum, luminescence quantum yield, wavelength (e.g., peak wavelength), and signal characteristics (e.g., pulse duration, inter-pulse duration, change in signal intensity).
[0110] In some embodiments, the photodetector may have a configuration that allows for detection of one or more timing characteristics (e.g., luminescence lifetime) associated with the sample's emission. The photodetector may detect a distribution of photon arrival times after a pulse of excitation light propagates through the integrated device, and the distribution of arrival times may provide an indication of the timing characteristics of the sample's emission light (e.g., a proxy for the luminescence lifetime). In some embodiments, one or more photodetectors provide an indication of the probability (e.g., luminescence intensity) of emission light emitted by a label. In some embodiments, multiple photodetectors may be sized and positioned to capture the spatial distribution of emission light. The output signal from the one or more photodetectors may then be used to distinguish one label from multiple labels, and multiple labels may be used to identify a sample within a sample. In some embodiments, the sample may be excited by multiple excitation energies, and the emission light and / or timing characteristics of the emission light emitted by the sample in response to the multiple excitation energies can distinguish one label from multiple labels.
[0111] In operation, parallel analysis of samples in the sample wells is performed by exciting some or all of the samples in the wells using excitation light and detecting signals from the sample emissions using photodetectors. The emitted light from the samples can be detected by corresponding photodetectors and converted into at least one electrical signal. The electrical signal can be transmitted along conductive lines in the circuitry of the integrated device, which can be connected to an instrument interfaced with the integrated device. The electrical signal can then be processed and / or analyzed. The processing or analysis of the electrical signal can be performed on a suitable computing device located either on or off the instrument.
[0112] The device may include a user interface for controlling the operation of the device and / or the integrated device. The user interface may be configured to allow a user to input information into the device, such as commands and / or settings used to control the device's functions. In some embodiments, the user interface may include buttons, switches, dials, and a microphone for voice commands. The user interface may allow a user to receive feedback regarding the device and / or the performance of the integrated device, such as information obtained by proper alignment and / or readout signals from a photodetector on the integrated device. In some embodiments, the user interface may provide feedback using a speaker to provide audible feedback. In some embodiments, the user interface may include indicator lights and / or a display screen to provide visual feedback to the user.
[0113] In some embodiments, the instrument may include a computer interface configured to connect to a computing device. The computer interface may be a USB interface, a FireWire interface, or any other suitable computer interface. The computing device may be any general-purpose computer, such as a laptop or desktop computer. In some embodiments, the computing device may be a server (e.g., a cloud-based server) accessible over a wireless network via a suitable computer interface. The computer interface may facilitate communication of information between the instrument and the computing device. Input information for controlling and / or configuring the instrument may be provided to the computing device and transmitted to the instrument via the computer interface. Output information generated by the instrument may be received by the computing device via the computer interface. The output information may include feedback regarding instrument performance, integrated device performance, and / or data generated from the photodetector readout signal.
[0114] In some embodiments, the instrument may include a processing device configured to analyze data received from one or more photodetectors of the integrated device and / or transmit control signals to the excitation source. In some embodiments, the processing device may comprise a general-purpose processor, a specially adapted processor (e.g., one or more central processing units (CPUs), such as microprocessors or microcontroller cores, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), custom integrated circuits, digital signal processors (DSPs), or combinations thereof). In some embodiments, processing of data from the one or more photodetectors may be performed by both the instrument's processing device and an external computing device. In other embodiments, the external computing device may be omitted, and processing of data from the one or more photodetectors may be performed solely by the integrated device's processing device.
[0115] According to some embodiments, an instrument configured to analyze a sample based on its luminescence emission characteristics may detect differences in luminescence lifetimes and / or intensities between different luminescent molecules and / or differences in lifetimes and / or intensities of the same luminescent molecule in different environments. The inventors have recognized and appreciated that differences in luminescence emission lifetimes can be used to distinguish between the presence or absence of different luminescent molecules and / or to distinguish between different environments or conditions to which the luminescent molecules are exposed. In some cases, aspects of the system can be simplified by distinguishing luminescent molecules based on lifetime (e.g., rather than emission wavelength). As an example, wavelength-discriminating optics (e.g., wavelength filters, dedicated detectors for each wavelength, dedicated pulsed light sources at different wavelengths, and / or diffractive optics) may be reduced in number or eliminated when distinguishing luminescent molecules based on lifetime. In some cases, a single pulsed light source operating at a single characteristic wavelength can be used to excite different luminescent molecules that emit within the same wavelength region of the optical spectrum but have measurably different lifetimes. Analysis systems that use a single pulsed light source, rather than multiple light sources operating at different wavelengths, to excite and distinguish different luminescent molecules that emit in the same wavelength range can be less complex to operate and maintain, more compact, and can be manufactured at lower cost.
[0116] While analytical systems based on luminescence lifetime analysis may have certain advantages, the amount of information obtained by the analytical system and / or the detection accuracy may be increased by enabling additional detection techniques. For example, some embodiments of the system may be further configured to identify one or more characteristics of a sample based on emission wavelength and / or emission intensity. In some implementations, emission intensity may additionally or alternatively be used to distinguish between different luminescent labels. For example, some luminescent labels may emit at significantly different intensities or have significant differences in excitation probability (e.g., at least about 35% difference), even if their decay rates are similar. By referencing binned signals to the measured excitation light, it may be possible to distinguish between different luminescent labels based on intensity levels.
[0117] According to some embodiments, different luminescence lifetimes can be distinguished by a photodetector configured to time-bin luminescence emission events following excitation of the luminescent labels. Time binning can occur during a single charge accumulation cycle for the photodetector. A charge accumulation cycle is the interval between readout events during which photogenerated carriers accumulate in bins of the time-binning photodetector. Examples of time-binning photodetectors are described in U.S. Patent Application No. 14 / 821,656, filed August 7, 2015, entitled "INTEGRATED DEVICE FOR TEMPORAL BINNING OF RECEIVED PHOTONS," which is incorporated herein by reference. In some embodiments, the time-binning photodetector can generate charge carriers in a photon absorption / carrier generation region and directly transfer the charge carriers to charge carrier storage bins in a charge carrier storage region. In such embodiments, the time-binning photodetector may not include a carrier transfer / capture region. Such time-binning photodetectors are sometimes referred to as "direct binning pixels." An example of a time-binning photodetector including a direct binning pixel is described in U.S. Patent Application No. 15 / 852,571, filed December 22, 2017, entitled "INTEGRATED PHOTODETECTOR WITH DIRECT BINNING PIXEL," which is incorporated herein by reference.
[0118] In some embodiments, different numbers of fluorophores of the same type can be attached to different reagents in a sample, so that each reagent can be identified based on its emission intensity. For example, two fluorophores can be attached to a first labeled recognition molecule, and four or more fluorophores can be attached to a second labeled recognition molecule. Due to the different numbers of fluorophores, there can be different excitation and emission probabilities associated with different recognition molecules. For example, there can be more emission events for the second labeled recognition molecule during a signal accumulation interval, so that the apparent intensity of the bin is significantly higher than that of the first labeled recognition molecule.
[0119] The inventors of the present application have recognized and appreciated that distinguishing biological or chemical samples based on fluorophore decay rates and / or fluorophore intensities can allow for simplification of optical excitation and detection systems. For example, optical excitation can be performed using a single wavelength source (e.g., a light source that produces one characteristic wavelength rather than multiple light sources, or a light source that operates at multiple different characteristic wavelengths). Furthermore, wavelength-discriminating optics and filters may not be required in the detection system. Also, a single photodetector may be used for each sample well to detect emissions from different fluorophores. The phrase "characteristic wavelength" or "wavelength" is used to refer to the central or dominant wavelength within a limited emission bandwidth (e.g., the central or peak wavelength within a 20 nm bandwidth output by a pulsed light source). In some cases, "characteristic wavelength" or "wavelength" can be used to refer to the peak wavelength within the full bandwidth of the emission output by a source.
[0120] According to one aspect of the present disclosure, an exemplary integrated device may be configured to perform single molecule analysis in combination with the above-described instruments. It should be understood that the exemplary integrated device described herein is intended to be exemplary, and that other integrated device configurations may be configured to perform any or all of the techniques described herein.
[0121] 1B shows a cross-sectional view of pixel 1-112 of integrated device 1-102. Pixel 1-112 includes a photodetection region, which may be a pinned photodiode (PPD), and a charge storage region, which may be a storage diode (SDO). In some embodiments, the photodetection region and the charge storage region may be formed within the semiconductor material of the pixel by doping regions of the semiconductor material. For example, the photodetection region and the charge storage region may be formed using the same conductivity type (e.g., n-type doping or p-type doping).
[0122] During operation of the pixel 1-112, excitation light can illuminate the sample well 1-108, causing incident photons, including fluorescent emission from the sample, to flow along the optical axis to the photodetection region PPD. As shown in FIG. 1B, the pixel 1-112 can include a waveguide 1-220 configured to optically couple (e.g., by evanescent coupling) excitation light from a grating coupler of an integrated device (not shown) into the sample well 1-108. In response, the sample in the sample well 1-108 can emit fluorescent light toward the photodetection region PPD. In some embodiments, the pixel 1-112 can also include one or more photonic structures 1-230, which can include one or more optical rejection structures, such as a spectral filter, a polarizing filter, and / or a spatial filter. For example, the photonic structures 1-230 can be configured to reduce the amount of excitation light reaching the photodetection region PPD and / or increase the amount of fluorescent emission reaching the photodetection region PPD. Also, as shown in pixel 1-112, pixel 1-112 may include one or more metal layers 1-240, which may be configured as filters and / or carry control signals from control circuitry configured to control the transfer gates, as further described herein.
[0123] In some embodiments, pixel 1-112 may include one or more transfer gates configured to control operation of pixel 1-112 by applying an electrical bias to one or more semiconductor regions of pixel 1-112 in response to one or more control signals. For example, when transfer gate ST0 induces a first electrical bias in the semiconductor region between photodetection region PPD and storage region SD0, a transfer path (e.g., a charge transfer channel) may be formed in the semiconductor region. Charge carriers (e.g., photoelectrons) generated in photodetection region PPD by incident photons flow along the transfer path to storage region SD0. In some embodiments, the first electrical bias may be applied during a collection period during which charge carriers from the sample are selectively directed to storage region SD0. Alternatively, when transfer gate ST0 imparts a second electrical bias to the semiconductor region between photodetection region PPD and storage region SD0, charge carriers from photodetection region PPD may be blocked from reaching storage region SD0 along the transfer path. In some embodiments, the drain gate REJ can provide a channel to the drain D to draw noise charge carriers generated in the photodetection region PPD by the excitation light away from the photodetection region PPD and the storage region SD0, such as during a rejection period before fluorescent emission photons from the sample reach the photodetection region PPD. In some embodiments, during a readout period, the transfer gate ST0 can provide a second electrical bias, and the transfer gate TX0 can provide an electrical bias to flow the charge carriers stored in the storage region SD0 to the readout region, which may be a floating diffusion (FD) region, for processing.
[0124] It should be understood that, according to various embodiments, the transfer gates described herein may comprise semiconductor materials and / or metals, and may include the gate of a field effect transistor (FET), the base of a bipolar junction transistor (BJT), etc.
[0125] In some embodiments, operation of the pixels 1-112 may include one or more collection sequences, each of which includes one or more rejection (e.g., drain) periods and one or more collection periods. In one example, a collection sequence performed in accordance with one or more pulses of an excitation light source can begin with a rejection period, such as to discard charge carriers generated within the pixels 1-112 (e.g., within the photodetection region PD) in response to excitation photons from the light source. For example, excitation photons may arrive at the pixels 1-112 prior to the arrival of fluorescent emission photons from the sample well. A transfer gate for the charge storage region can be biased to have low conductivity in a charge transfer channel connecting the charge storage region to the photodetection region, preventing the transfer and accumulation of charge carriers in the charge storage region. A drain gate for the drain region can be biased to have high conductivity in a drain channel between the photodetection region and the drain region, facilitating the draining of charge carriers from the photodetection region to the drain region. The transfer gate for any charge storage region coupled to the photodetector region may be biased to have low conductivity between the photodetector region and the charge storage region, thereby preventing charge carriers from being transferred or stored in the charge storage region during the rejection period.
[0126] The rejection period may be followed by a collection period, during which charge carriers generated in response to incident photons are transferred to one or more charge accumulation regions. During the collection period, incident photons may include fluorescent emission photons, resulting in the accumulation of fluorescent emission charge carriers in the charge accumulation region(s). For example, a transfer gate for one of the charge accumulation regions may be biased to have high conductivity between the photodetection region and the charge accumulation region, facilitating the accumulation of charge carriers in the charge accumulation region. Any drain gate coupled to the photodetection region may be biased to have low conductivity between the photodetection region and the drain region to prevent charge carriers from being discarded during the collection period.
[0127] Some embodiments may include multiple rejection and / or collection periods in a collection sequence, such as a first rejection and collection period followed by a second rejection and collection period, with each pair of rejection and collection periods occurring in response to a pulse of excitation light. In one example, charge carriers generated in the photodetection region during each collection period of a collection sequence (e.g., in response to multiple pulses of excitation light) may be collected in a single charge storage region. In some embodiments, the charge carriers collected in the charge storage region may be read out for processing before the next collection sequence. Alternatively, or additionally, in some embodiments, the charge carriers collected in the first charge storage region during the first collection sequence may be transferred to a second charge storage region sequentially coupled to the first charge storage region and read out simultaneously with the next collection sequence. In some embodiments, the processing circuitry configured to read out the charge carriers from one or more pixels may be configured to determine one or more of luminescence intensity information, luminescence lifetime information, luminescence spectrum information, and / or any other mode of luminescence information associated with performing the techniques described herein.
[0128] In some embodiments, the first collection sequence may include transferring charge carriers generated in the light detection response to the excitation pulse to the charge accumulation region at a first time following each excitation pulse, and the second collection sequence may include transferring charge carriers generated in the light detection response to the excitation pulse to the charge accumulation region at a second time following each excitation pulse. For example, the number of charge carriers collected after the first and second times can indicate luminance lifetime information of the received light.
[0129] As described further herein, the pixels of the integrated circuit can be controlled to perform one or more collection sequences using one or more control signals from the integrated circuit's control circuitry, for example, by providing control signals to the drains and / or transfer gates of the pixels of the integrated circuit. In some embodiments, charge carriers can be read out from the FD region of each pixel for processing during a readout pixel associated with each pixel and / or row or column of pixels. In some embodiments, the FD region of the pixel can be read out using a correlated double sampling (CDS) technique.
[0130] Sequence information As described herein, in some embodiments, a cleavage reagent of the present disclosure comprises an aminopeptidase having an amino acid sequence that shares a certain percentage of sequence identity with an amino acid sequence selected from Table 1. In some embodiments, a cleavage reagent comprises an aminopeptidase described herein and a tag sequence having an amino acid sequence that shares a certain percentage of sequence identity with an amino acid sequence selected from Table 2. For purposes of comparing two or more amino acid sequences, the percentage of "sequence identity" (also referred to herein as "amino acid identity") between a first amino acid sequence and a second amino acid sequence can be calculated by dividing the number of amino acid residues in the first amino acid sequence that are identical to amino acid residues at corresponding positions in the second amino acid sequence by the total number of amino acid residues in the first amino acid sequence and multiplying by 0100, where each deletion, insertion, substitution, or addition of an amino acid residue in the second amino acid sequence compared to the first amino acid sequence is considered to be a difference at a single amino acid residue (position).
[0131] Alternatively, the degree of sequence identity between two amino acid sequences can be calculated using known computer algorithms (e.g., by the local homology algorithm of Smith and Waterman (1970), Adv. Appl. Math. 2:482c, by the homology alignment algorithm of Needleman and Wunsch, J. Mol. Biol., (1970), 48:443, by the similarity search method of Pearson and Lipman, Proc. Natl. Acad. Sci. USA, (1998), 85:2444, or by computerized implementations of algorithms available as Blast, Clustal Omega, or other sequence alignment algorithms), for example, using standard settings. Typically, for purposes of determining the percentage of "sequence identity" between two amino acid sequences according to the calculation method outlined above, the amino acid sequence with the largest number of amino acid residues is designated the "first" amino acid sequence, and the other amino acid sequence is designated the "second" amino acid sequence.
[0132] Additionally or alternatively, two or more sequences may be assessed for identity between the sequences. The term "identical" or percent "identity" in the context of two or more amino acid sequences refers to two or more sequences or subsequences that are the same. Two sequences are "substantially identical" if they have a specified percentage of amino acid residues that are the same over a specified region or entire sequence (e.g., at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9% identical) when compared and aligned for maximum correspondence over a comparison window or designated region, as measured using one of the sequence comparison algorithms described above or by manual alignment and visual inspection. Optionally, identity exists over a region that is at least about 25, 50, 75, or 100 amino acids in length, or over a region that is 100-150, 150-200, 100-200, or 200 or more amino acids in length.
[0133] Additionally or alternatively, two or more sequences may be evaluated for alignment between the sequences. The term "alignment" or "percent alignment" in the context of two or more amino acid sequences refers to two or more sequences or subsequences that are the same. Two sequences are "substantially aligned" if they have a specified percentage of amino acid residues that are the same across a particular region or entire sequence (e.g., at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9% identical) when compared and aligned for maximum correspondence over a comparison window or designated region, as determined using one of the sequence comparison algorithms described above or by manual alignment and visual inspection. Optionally, the alignment exists over a region that is at least about 25, 50, 75, or 100 amino acids in length, or over a region that is 100-150, 150-200, 100-200, or 200 or more amino acids in length.
[0134] [Table 1-1]
[0135] [Table 1-2]
[0136] [Table 1-3]
[0137] [Table 1-4]
[0138] [Table 1-5]
[0139] [Table 1-6]
[0140] [Table 1-7]
[0141] [Table 1-8]
[0142] [Table 2] [Example]
[0143] Example 1. Evaluation of Pyrococcus horikoshii TET II aminopeptidase Cleavage performance was evaluated using peptide substrates extended at the C-terminus by the addition of a tripeptide DDD motif. Figures 2A-2D show the improvement in TET aminopeptidase performance by extending the C-terminus of the peptide with a DDD motif. The bar graphs in Figures 2A-2D show the beneficial effect of C-terminal extension with a DDD peptide motif on the kinetics of the hTETII / PfuTET combination (1 μM / 40 μM) using the QP514 (DQQRLIFAYPDDD) (SEQ ID NO: 49) peptide and the control QP434 (DQQRLIFAG, no DDD motif) peptide. Plots show cleavage depth (Figure 2A), cleavage activity (% of reads reaching the last visible RS) (Figure 2B), time required to cleave the DQQ motif and R residues (Figure 2C), and % of 4+ RS (Figure 2D).
[0144] The cleavage performance of AP30, a cleavage reagent containing Pyrococcus horikoshii TETII aminopeptidase (hTETII) and a C-terminal tag, was evaluated. AP30 was expressed in Escherichia coli (E. coli), and cell-free extracts were prepared after cell lysis and separated on a talon affinity column in AKTA. Figure 3A shows an exemplary chromatogram showing AP30 separation and elution peaks.
[0145] Figure 3B (left image) shows Talon affinity column-purified fractions separated on an SDS PAGE gel to demonstrate AP30 enrichment. The elution peak fraction with the majority of monomer (45.83 kDa) is indicated by an arrow. Other bands above the monomer represent different complexes of AP30 under these conditions. Figure 3B (right image) shows a native gel showing hTETII and AP30 protein profiles before and after conditioning (complexation at 65°C for 30 minutes in the presence of cobalt acetate, resulting in a mostly dodecamer complex). AP30 exhibits a more homogenous, higher-order complex formation.
[0146] Real-time cleavage kinetics assays were performed using the amino acid-AMC (7-amino-4-methylcoumarin) substrate. Cleavage kinetics were followed in real time at 30 °C with an excitation wavelength of 357 nm and an emission wavelength of 450 nm, measuring the increase in fluorescence at 441 nm. Figure 4 shows three bar graphs comparing the intrinsic cleavage rates and substrate specificity of hTETII and AP30 for 18 individual amino acids, categorized according to activity level (high activity: top chart; medium activity: middle chart; very low activity: bottom chart). Intrinsic cleavage rates were derived for each residue at the optimal concentration of aminopeptidase determined by aminopeptidase titration. Rates were calculated by exponential fitting of the intensity change over time. These results indicated that AP30 had slower intrinsic cleavage rates than hTETII for individual amino acids.
[0147] Protein sequencing assays were performed using AP30 cutter (5 μM), PS610 (50 nM) recognition factor, and QP47 peptide (FAAAYPDDD) (SEQ ID NO: 46). Figure 5A shows representative traces demonstrating the initial and final RS recognition by PS610 as cleavage progresses. Figure 5B shows plots depicting the average cleavage time for FA (left plot) and YP (right plot) RSs. Figure 5C shows plots depicting the populations of rapid sequential cleavage and regular cleavage. The RSC population with AP30 was only about 25%, which was lower than that of hTETII. These results indicated that AP30 exhibited a smaller population of rapid sequential cleavage (RSC) than hTETII.
[0148] Protein sequencing assays were performed using 1 μM hTETII cutter, PS610 (50 nM) recognition factor, and QP47 peptide (FAAAYPDDD) (SEQ ID NO: 46). Figure 6A shows a representative trace showing the first and last RS by PS610 as cleavage progresses. Figure 6B shows a plot showing the average cleavage time for FA and YP RS. Figure 6C shows a plot showing the populations of rapid sequential cleavages and regular cleavages. The RSC population is approximately 40%. These results indicated that hTETII has a larger population of rapid repeat cleavages (RSCs) than AP30.
[0149] Protein sequencing assays were performed using AP30 (1 μM) / PfuTET (40 μM), an alternative cutter combination, with PS610 (50 nM), PS557 (250 nM), and PS621 (250 nM) recognition factors and the QP433 peptide (RLIFAYPDDD) (SEQ ID NO: 47). Figure 7A shows a representative trace showing five RSs identified by the recognition factors as cleavage progresses. Figure 7B shows a bin ratio vs. pulse duration plot with separated clusters for recognized RSs (top plot), as well as plots showing the cleavage depth and percentage of each RS recognized in reads with gaps (middle plot) and reads without gaps (no missed recognizable residues) (bottom plot).
[0150] On the same chip, protein sequencing assays were performed using AP30 (3 μM) or hTETII (2 μM) / PfuTET (40 μM) aminopeptidase combinations with PS610 (50 nM) and PS557 (250 nM) recognition factors and QP425 (LASSIAEANRFADIADYP) (SEQ ID NO: 48). Figure 8A shows representative traces for reactions with AP30 and hTETII, showing the RSs identified by the recognition factors as cleavage progressed. Figure 8B shows plots showing the cleavage depth and percentage of each RS recognized in ungapped reads (no recognizable residues missed).
[0151] Figures 9A-9C show that AP30 improved cleavage performance by increasing cleavage activity and reducing rapid sequential cleavage. The bar graphs in Figures 9A-9C show the improvement in cleavage depth (Figure 9A) and % of 4+ RS reads (Figure 9B) at different AP30 concentrations using 40 μM pfuTET, PS610 (50 nM), and PS557 (250 nM) recognition factors for QP425 (LASSIAEANRFADIADYP) (SEQ ID NO: 48). Changes were calculated relative to hTETII tested at the same concentration on the same chip. AP30 reduced the rapid sequential cleavage of IA and FA RS (Figure 9C). Cleavage activity was also improved in chip readouts, as the number of useful reads (4+ RS) increased with the AP30 / PfuTET combination.
[0152] Tables 3 and 4 show exemplary results of these studies, showing that AP30 improved cleavage performance (cut depth and % useful reads) by increasing cleavage activity and reducing rapid sequential cleavage.
[0153] [Table 3]
[0154] [Table 4]
[0155] Example 2. Evaluation of Pyrococcus horikoshii TET III aminopeptidase In this example, the cleavage performance of Pyrococcus horikoshii TETIII aminopeptidase (hTETIII) and AP37, a cleavage reagent containing a C-terminal tag, was evaluated. hTETIII is a homolog of PfuTET, which was used in these studies for comparison with AP37. AP37 was expressed in Escherichia coli (E. coli), and cell-free extracts were separated on a talon affinity column in AKTA. Figure 10A shows an exemplary chromatogram demonstrating AP37 enrichment and elution peaks.
[0156] Figure 10B shows Talon affinity column-purified fractions separated on an SDS PAGE gel to demonstrate AP37 separation. The elution peak fraction containing the monomer (40.33 kDa) is indicated by an arrow. Other bands above the monomer represent different complexes of AP37 under these conditions. Results from HPLC assay data demonstrated improved cleavage activity of AP37 after cobalt acetate and heat conditioning. Table 5 shows activity data at aminopeptidase concentrations of 1 μM and 10 μM (before and after conditioning) for peptides with different N-terminal amino acids.
[0157] [Table 5]
[0158] Real-time cleavage kinetics assays were performed using an amino acid-AMC (7-amino-4-methylcoumarin) substrate. Cleavage kinetics were followed in real time at 30 °C with an excitation wavelength of 357 nm and an emission wavelength of 450 nm, measuring the increase in fluorescence at 441 nm. Figure 11 shows three bar graphs comparing the intrinsic cleavage rates and substrate specificity of PfuTET and AP37 for 18 individual amino acids, categorized according to activity level (high activity: top chart; medium activity: middle chart; very low activity: bottom chart). Intrinsic cleavage rates were derived for each residue at the optimal concentration of aminopeptidase determined by aminopeptidase titration. Rates were calculated by exponential fitting of the intensity change over time. These results indicated that AP37 has faster intrinsic cleavage rates than PfuTET for individual amino acids.
[0159] Protein sequencing assays were performed on separate chips using AP37 (40 μM) or PfuTET (40 μM) aminopeptidase with PS610 (50 nM), PS961 (250 nM), and PS961 (250 nM) recognition factors, as well as QP514 (DQQRLIFAYPDDD) (SEQ ID NO: 49). Figure 12A shows representative traces for reactions with AP37 or PfuTET, showing the RSs identified by the recognition factors as cleavage progressed. Figure 12B shows plots showing the cleavage depth and percentage of each RS recognized in ungapped reads (no missed recognizable residues). These results indicated that AP37, when used alone, exhibits better efficacy at cleaving arginines than PfuTET.
[0160] Protein sequencing assays were performed on QP433 (RLIFAYPDDD) (SEQ ID NO: 47) with two different AP37 concentrations (20 μM / 40 μM), 4 μM hTETII, and the PS610 (50 nM), PS557 (250 nM), and PS691 (100 nM) recognition factors. Figures 13A-13C show plots illustrating the results of these studies, with changes calculated relative to PfuTET tested at the same concentration on the same chip. hTETII was maintained at 4 μM in all conditions. Figure 13A shows the improvement in cleavage depth. Figure 13B shows the improvement in the % of significant reads. Figure 13C shows that AP37 reduced rapid sequential cleavage (RSC) of LI, IF, and FA RS.
[0161] Example 3. Evaluation of the combination of AP30 and AP37 aminopeptidases In this example, a protein sequencing assay was performed to evaluate the combination of AP30 and AP37 in sequencing reactions.
[0162] Figures 14A-14B show bar graphs illustrating the effect of different AP30 / AP37 ratios with the PS610 (50 nM), PS691 (100 nM), and PS961 (250 nM) recognition factors on the time to cleave the DQQ motif and R residue (Figure 14A) and the rapid sequential cleavage of RS (Figure 14B). Increasing the aminopeptidase concentration during sequencing of QP434 (DQQRLIFAG) showed an increase in the cleavage rate for both the dark amino acid (DQQ motif) and the visible amino acid R, which resulted in an increase in the rapid sequential cleavage of IF.
[0163] Figures 15A-15B show bar graphs depicting the distribution of performance metrics (cleavage depth, % of 4+ RS reads, and % of 5 RS reads) using a recognition factor mixture of PS610 (50 nM), PS961 (250 nM), and PS691 (100 nM), and AP30 / AP37 at 4 μM / 40 μM concentrations for QP514 (DQQRLIFAYPDDD) (SEQ ID NO: 49) (Figure 15A, 7 technical replicates) and 10 μM / 60 μM concentrations for QP549 (DQQIASSRLAASFAAQQYPDDD1) (SEQ ID NO: 50) (Figure 15B, 6 technical replicates). Standard recognition factor and aminopeptidase conditions were used as controls on the same chip [PS610 (50 nM), PS557 (250 nM), PS621 (250 nM), and hTETII / PfuTET at 1 μM / 40 μM for QP514 and 3 μM / 60 μM for QP549].
[0164] Figure 16 shows the read resolution and abundance of ungapped reads with a single deletion allowed at a read length of 4 from one of the chip runs with the 250 / 100 nM PS961 / PS691 and 10 / 60 μM AP30 / hTETIII recognition factor and aminopeptidase mixtures for QP549. The updated combination of aminopeptidase and recognition factor for QP549 achieved an unprecedented cleavage depth value of greater than 9 for the top 10% of openings across multiple runs. This run reached 10 / 10, with a cleavage depth of 10.64 for the top % of openings with a single, rapid, consecutive cleavage allowed at a read length of 4.
[0165] Example 4. Evaluation of Yersinia pestis Xaa-prolyl aminopeptidase (yPIP) Xaa-prolyl aminopeptidase (yPIP) from Yersinia pestis, also known as proline aminopeptidase PII, is a monomeric AP with specific activity against an N-terminal XP motif. In this example, the cleavage performance of yPIP and a cleavage reagent containing a C-terminal 6xHis tag was evaluated. yPIP (BC-B1 batch) was expressed in Escherichia coli (E. coli), and cell-free extracts were prepared after cell lysis and separated on a Talon affinity column in AKTA. Figure 17A shows an exemplary chromatogram showing yPIP separation and elution peaks. Figure 17B shows Talon affinity column-purified fractions separated on an SDS-PAGE gel to demonstrate yPIP enrichment. The 51 kDa elution peak fraction is indicated by an arrow.
[0166] Figure 18 shows representative traces demonstrating the effect of adding yPIP (2 μM) to the AP30 / AP37 (4 / 40 μM) aminopeptidase combination in a protein sequencing assay. Protein sequencing assays were performed on the QP352 (RLAYPAFAAYPDDDF) (SEQ ID NO: 51) peptide using (a) AP30 (1 μM) / AP37 (40 μM) and (b) AP30 (1 μM) / AP37 (40 μM) / yPIP (2 μM) combinations, together with PS610 (50 nM), PS961 (125 nM), and PS1122 (250 nM) recognition factors. These results indicated that yPIP in the AP30 / AP37 aminopeptidase combination aids in more efficient cleavage across the YP motif.
[0167] Figure 19 shows the results of a protein sequencing assay using the combinations (a) AP30 (4 μM) / AP37 (40 μM) and (b) AP30 (4 μM) / AP37 (40 μM) / yPIP (2 μM) for the QP352 (RLAYPAFAAYPDDDF) (SEQ ID NO: 51) peptide, together with the recognition factors PS610 (50 nM), PS961 (125 nM), and PS1122 (250 nM). A bin ratio versus pulse duration plot with separated clusters for recognized RSs is shown in the upper panel. Bar plots for the cleavage depth and percentage of each RS recognized in ungapped (middle) and gapped (lower) reads are shown. These results indicated that yPIP in the AP30 / AP37 aminopeptidase combination cleaves beyond the YP motif more efficiently to reach the FA.
[0168] Figure 20 shows the read rate resolution for QP352 (RLAYPAFAAYPDDDF) (SEQ ID NO: 51) protein sequencing assays using (a) AP30 (4 μM) / AP37 (40 μM) and (b) AP30 (4 μM) / AP37 (40 μM) / yPIP (2 μM) combinations with PS610 (50 nM) / PS961 (125 nM) / PS1122 (250 nM) recognition factors. These results further demonstrated that yPIP in the AP30 / AP37 aminopeptidase combination more efficiently cleaves beyond the YP motif to reach FAs.
[0169] Figure 21 shows the results of a protein sequencing assay using the combinations (a) AP30 (4 μM) / AP37 (40 μM) / yPIP (0.5 μM) and (b) AP30 (4 μM) / AP37 (40 μM) / yPIP (1 μM) for the QP352 (RLAYPAFAAYPDDDF) (SEQ ID NO: 51) peptide, together with the PS610 (50 nM) / PS961 (125 nM) recognition factors. A bin ratio versus pulse duration plot with separated clusters for recognized RSs is shown in the upper panel. The % of actual FAs should be higher for both yPIP concentrations because FAs are miscalled as YPs. However, the density of FA miscalls is much higher for 1 μM yPIP compared to 0.5 μM yPIP. Additionally, YP clusters were denser for 0.5 μM yPIP compared to 1 μM yPIP. Taken together, these observations indicated that the overall performance of 1 μM yPIP was preferable to 0.5 μM yPIP under these reaction conditions. Bar plots of the cleavage depth and percentage of each RS recognized in ungapped (middle panel) and gapped (lower panel) reads are shown in Figure 21. A higher initial YP and lower FA indicated poor cleavage (higher YP relative to FA; less YP cleavage), a lower initial YP and higher FA indicated YP cleavage and YP deletion (lower YP relative to FA; higher RSC), and similar YP and FA were considered ideal (good YP cleavage and fewer RSC).
[0170] Figure 22 shows the read rate resolution for QP352 (RLAYPAFAAYPDDDF) (SEQ ID NO: 51) protein sequencing assays using (a) AP30 (4 μM) / AP37 (40 μM) / yPIP (0.5 μM) and (b) AP30 (4 μM) / AP37 (40 μM) / yPIP (1 μM) combinations with the PS610 (50 nM) / PS961 (125 nM) recognition element combination and the PS610 (50 nM) / PS961 (125 nM) recognition element combination. These results indicate that 1 μM yPIP performed better than 0.5 μM in the AP30 / AP37 combination under these reaction conditions.
[0171] Figure 23 shows the results of a protein sequencing assay using the combinations (a) AP30 (4 μM) / AP37 (40 μM) / yPIP-BC_B1 (2 μM) and (b) AP30 (4 μM) / AP37 (40 μM) / yPIP-470 (2 μM) for the QP352 (RLAYPAFAAYPDDDF) (SEQ ID NO: 51) peptide, together with the recognition factors PS610 (50 nM) / PS961 (125 nM) / PS1122 (250 nM). A bin ratio versus pulse duration plot with separated clusters for recognized RSs is shown in the upper panel. Bar plots for the cleavage depth and percentage of each RS recognized in ungapped (middle) and gapped (lower) reads are shown. yPIP470 was stored for more than 18 months at -20°C and showed similar performance compared to freshly purified yPIP-BC_B1 in terms of % FA reads. However, yPIP470 showed a slight increase in the RSC of the initial YP (i.e., lower YP followed by higher FA).
[0172] Figure 24 shows the read rate resolution for QP352 (RLAYPAFAAYPDDDF) (SEQ ID NO: 51) protein sequencing assays using combinations of (a) AP30 (4 μM) / AP37 (40 μM) / yPIP-BC_B1 (2 μM) and (b) AP30 (4 μM) / AP37 (40 μM) / yPIP-470 (2 μM) with PS610 (50 nM) / PS961 (125 nM) / PS1122 (250 nM) recognition factors.
[0173] Example 5. Evaluation of AP70 Yersinia pestis Xaa-prolyl aminopeptidase In this example, the cleavage performance of yPIP (truncated by two amino acids at the C-terminus) and AP70, a cleavage reagent containing a C-terminal GGS-6xHis tag, was evaluated. Figure 25 shows (a) the yPIP and AP70 protein amino acid sequence alignments showing the C-terminal 6xH tag and GGS-6xHis tag, respectively, and (b) the expression constructs for yPIP and AP70, along with the molecular weights of the proteins.
[0174] AP70 (BC-B1 batch) was expressed in E. coli, and cell-free extracts were prepared after cell lysis and separated on a talon affinity column in AKTA. Figure 26A shows an exemplary chromatogram showing AP70 separation and elution peaks. Figure 26B shows talon affinity column-purified fractions separated on an SDS-PAGE gel to demonstrate AP70 enrichment. The 51.56 kDa elution peak fraction is indicated by an arrow.
[0175] Figure 27 shows the results of an HPLC assay for QP734 (FPARAFAYPDDD) (SEQ ID NO: 52) peptide cleavage by AP70 at various concentrations after different conditioning treatments and an unconditioned control. The reactions were separated on a hydrophobic column by reversed-phase HPLC. AP70 pretreatment conditions: no treatment; 50°C, 30 min; 50°C + Co 2+ 5mM, 30min; 50℃+Mg 2+ 5mM, 30 min. AP70 pretreatment conditions: no treatment; 50°C, 30 min; 50°C + Co 2+5mM, 30min; 50℃+Mg 2+ 5mM, 30 minutes.
[0176] Figure 28 shows the results of HPLC assays for different peptide cleavage by AP70 after different conditioning treatments and an unconditioned control. AP70 pretreatment conditions: no treatment; 50°C, 30 min; 50°C + Co 2+ 5mM, 30min; 50℃+Mg 2+ 5 mM, 30 min. HPLC reaction conditions: 200 μM peptide; 30 min at 30° C.; quench using formic acid.
[0177] A real-time cleavage kinetics assay was performed using a short peptide-AMC (7-amino-4-methylcoumarin) substrate. The cleavage kinetics was followed in real time at 30°C with an excitation wavelength of 357 nm and an emission wavelength of 450 nm, measuring the increase in fluorescence at 441 nm. Figure 29 shows three graphs comparing the cleavage activity and substrate specificity of the AP70, AP30 + AP37, and AP30 + AP37 + AP70 combinations for the proline-containing peptides PXX, XPP, and XPX. Cleavage activity was induced for each peptide at the optimal concentration of aminopeptidase, as determined by aminopeptidase titration.
[0178] Figure 30 shows the results of a protein sequencing assay for the QP734 (FPARAFAYPDDD) (SEQ ID NO: 52) peptide with AP70 (1 μM, 4 μM), a no-AP control, and AP30 (4 μM), along with the PS610 (50 nM), PS961 (125 nM), and PS1122 (250 nM) recognizers. A plot of aperture versus time shows FP, RA, and FA levels throughout the chip run (left panel). A plot of bin ratio versus pulse duration is shown, with separated clusters for recognized RS (right panel; arrows in the third panel from the top indicate data that match the PS1122 bin ratio but are miscalled as FP).
[0179] Figure 31 shows representative traces from protein sequencing assays demonstrating the effect of adding AP70 (1 μM) to the AP30 / AP37 (4 / 40 μM) aminopeptidase combination. Protein sequencing assays were performed on the QP352 (RLAYPAFAAYPDDDF) (SEQ ID NO: 51) peptide using the following combinations: (a) AP30 (1 μM), AP37 (40 μM), and AP70 (1 μM) and (b) AP30 (1 μM) and AP37 (40 μM), together with PS610 (50 nM), PS961 (125 nM), and PS1122 (250 nM) recognition factors. Table 6 shows the results of titrating AP70 concentration to optimize the AP combination performance of AP30 + AP37 + AP70 on a chip for QP352 (RLAYPAFAAYPDDDF) (SEQ ID NO: 51) sequencing.
[0180] [Table 6]
[0181] The results of additional runs performed to generate the data in Table 6 are shown in Figures 32A-32F. The runs in Figures 32A and 32B showed similar performance at both AP70 concentrations. The runs in Figures 32E and 32F showed (a) similar performance at both concentrations of AP70 in terms of reads reaching FA, and (b) an increase in the RSC of both the initial and final YP for 6 μM AP70, reflected in a lower % of YP reads compared to 2 μM AP70.
[0182] Figure 33 shows the results of a protein sequencing assay using the combinations (a) AP30 (4 μM), AP37 (40 μM), and AP70 (1 μM) and (b) AP30 (4 μM) and AP37 (40 μM) for the QP352 (RLAYPAFAAYPDDDF) (SEQ ID NO: 51) peptide, together with the recognition factors PS610 (50 nM), PS961 (125 nM), and PS1122 (250 nM). A bin ratio versus pulse duration plot with separated clusters for recognized RSs is shown in the upper panel. The presence of the "FA" cluster with 1 μM AP70 in combination with 4 / 40 μM AP30 / AP37 confirms that AP70 can cleave beyond the "YP" motif, thus increasing cleavage depth. Also shown in Figure 33 are bar plots showing the cleavage depth and percentage of each RS recognized in ungapped (middle panel) and gapped (lower panel) reads.
[0183] Figure 34 shows the read rate resolution for QP352 (RLAYPAFAAYPDDDF) (SEQ ID NO: 51) protein sequencing assays using combinations of (a) AP30 (4 μM) / AP37 (40 μM) / AP70 (1 μM) and (b) AP30 (4 μM) / AP37 (40 μM) with PS610 (50 nM) / PS961 (125 nM) / PS1122 (250 nM) recognition factors.
[0184] Figure 35 shows the results of a protein sequencing assay using the combinations (a) AP30 (4 μM), AP37 (40 μM), and AP70 (1 μM) and (b) AP30 (4 μM) and AP37 (40 μM) on the QP354 (LAAYPARLAYPDDDF) (SEQ ID NO: 53) peptide, together with the PS610 (50 nM), PS961 (125 nM), and PS1122 (250 nM) recognition factors. A bin ratio versus pulse duration plot with separated clusters for recognized RSs is shown in the upper panel. Also shown are bar plots showing the cleavage depth and percentage of each RS recognized in ungapped (middle) and gapped (lower) reads. These results demonstrated that AP70 exhibited YP cleavage on the surrogate peptide QP352.
[0185] Figure 36 shows the read ratio resolution for QP354 (LAAYPARLAYPDDDF) (SEQ ID NO: 53) protein sequencing assays using (a) AP30 (4 μM) / AP37 (40 μM) / AP70 (1 μM) and (b) AP30 (4 μM) / AP37 (40 μM) combinations with PS610 (50 nM) / PS961 (125 nM) / PS1122 (250 nM) recognition factors. These results demonstrated that AP70 showed consistent performance against surrogate peptides of the same length as QP352.
[0186] Figure 37 shows the results of a protein sequencing assay using the combinations of (a) AP30 (4 μM), AP37 (40 μM), and AP70 BC-B1 batch (1 μM) and (b) AP30 (4 μM), AP37 (40 μM), and yPIP BC-B1 batch (1 μM) for the QP352 (RLAYPAFAAYPDDDF) (SEQ ID NO: 51) peptide, together with the recognition factors PS610 (50 nM), PS961 (125 nM), and PS1122 (250 nM). A bin ratio versus pulse duration plot with separated clusters for recognized RSs is shown in the upper panel. Also shown are bar plots showing the cleavage depth and percentage of each RS recognized in ungapped (middle) and gapped (lower) reads. AP70 showed fewer RSCs for the first YP and reached the final YP for both ungapped and gapped reads compared to yPIP. Table 7 shows the performance comparison of yPIP and AP70 with the AP combination of AP37+AP70 for sequencing QP352 (RLAYPAFAAYPDDDF) (SEQ ID NO: 51).
[0187] [Table 7]
[0188] Figure 38 shows the read rate resolution for a QP352 (RLAYPAFAAYPDDDF) (SEQ ID NO: 51) protein sequencing assay using the combinations of (a) AP30 (4 μM) / AP37 (40 μM) / AP70 BC-B1 batch (1 μM) and (b) AP30 (4 μM) / AP37 (40 μM) / yPIP BC-B1 batch (1 μM) with PS610 (50 nM) / PS961 (125 nM) / PS1122 (250 nM) recognition factors. The results demonstrate that AP70 exhibits less RSC for the first YP and reaches the final YP for both ungapped and gapped reads compared to yPIP.
[0189] Equivalents and Scope In the claims, articles such as "a," "an," and "the" may mean one or more unless indicated to the contrary or clear from context. A claim or description including "or" between one or more members of a group is considered to be satisfied if one, more than one, or all of the group members are present in, utilized in, or otherwise relevant to a given product or process, unless indicated to the contrary or clear from context. The invention includes embodiments in which exactly one member of a group is present in, utilized in, or otherwise relevant to a given product or process. The invention includes embodiments in which two or more, or all of the members of a group are present in, utilized in, or otherwise relevant to a given product or process.
[0190] Furthermore, the present invention encompasses all variations, combinations, and permutations in which one or more limitations, elements, clauses, and descriptive terms from one or more of the enumerated claims are introduced into another claim. For example, any claim that depends on another claim can be modified to include one or more limitations found in other claims that depend on the same base claim. Where elements are presented as lists, for example, in Markush group format, each subgroup of the same elements is also disclosed, and any element can be removed from the group. Generally, when the invention or aspects of the invention are referred to as including certain elements and / or features, it is understood that particular embodiments of the invention or aspects of the invention include or consist essentially of such elements and / or features. For simplicity, these embodiments are not specifically described herein.
[0191] The term "and / or," as used in the specification and claims, should be understood to mean "either or both" of the elements so connected, i.e., elements present conjunctively in some cases and non-conjunctively in other cases. Elements listed with "and / or" should be construed in the same manner, i.e., as "one or more" of the elements so connected. Other elements, whether related or unrelated to the elements specifically identified by the "and / or" clause, may optionally be present. Thus, as a non-limiting example, a statement such as "A and / or B," when used in conjunction with open-ended language such as "comprising," may refer in one embodiment to A only (optionally including elements other than B); in another embodiment to B only (optionally including elements other than A); in yet another embodiment to both A and B (optionally including other elements), etc.
[0192] As used in this specification and in the claims, "or" should be understood to have the same meaning as "and / or" as defined above. For example, when separating items in a list, "or" or "and / or" shall be construed as inclusive, i.e., the inclusion of at least one, but more than one, and optionally additional unlisted items, of a number or list of elements. Terms clearly indicating the contrary, such as "only one of," "exactly one of," or, when used in the claims, "consisting of," refer to the inclusion of exactly one element of a number or list of elements. In general, the term "or" as used herein shall be construed to indicate exclusive alternatives (i.e., "one or the other but not both") only when preceded by terms of exclusivity, such as "either," "one of," "only one of," or "exactly one of." "Consisting essentially of," when used in the claims, shall have its ordinary meaning as used in the field of patent law.
[0193] As used herein and in the claims, the phrase "at least one," when referring to a list of one or more elements, should be understood to mean at least one element selected from any one or more elements in the list of elements, but not necessarily including at least one of every element specifically listed in the list of elements, nor excluding any combination of elements in the list of elements. This definition also allows for the optional presence of elements other than those specifically identified in the list of elements to which the phrase "at least one" refers, whether related to the specifically identified elements or not. Thus, as a non-limiting example, "at least one of A and B" (or, equivalently, "at least one of A or B," or, equivalently, "at least one of A and / or B") means that in one embodiment, A is at least one, optionally including more than one, and B is absent (optionally including elements other than B); in another embodiment, B is at least one, optionally including more than one, and A is absent (optionally including elements other than A); in yet another embodiment, A is at least one, optionally including more than one, and B is at least one, optionally including more than one (optionally including other elements), etc.
[0194] It should also be understood that, unless expressly indicated to the contrary, in any method claimed herein including multiple steps or acts, the order of the method steps or acts is not necessarily limited to the order in which the steps or acts are recited.
[0195] In the claims and the foregoing specification, all transitional phrases, such as "comprising," "including," "carrying," "having," "containing," "involving," "holding," "composed of," and the like, are understood to be open-ended, i.e., meaning including, but not limited to. Only the transitional phrases "consisting of" and "consisting essentially of" shall be closed or semi-closed transitional phrases, respectively, as defined in Section 2111.03 of the United States Patent Office Manual of Patent Examining Procedures. It should be recognized that embodiments described herein using open-ended transitional phrases (e.g., "comprising") also contemplate, in alternative embodiments, the features "consisting of" and "consisting essentially of" the features described by the open-ended transitional phrases. For example, if the application describes "a composition comprising A and B," the application also contemplates the alternative embodiments "a composition consisting only of A and B" and "a composition consisting essentially of A and B."
[0196] Where ranges are specified, endpoints are included. Furthermore, unless otherwise specified or apparent from the context and the understanding of one of ordinary skill in the art, values expressed as ranges are understood to contemplate any specific value or subrange within the stated range in different embodiments of the invention, to the tenth of the unit of the lower limit of that range, unless the context clearly dictates otherwise.
[0197] This application refers to various issued patents, published patent applications, journal articles, and other publications, all of which are incorporated herein by reference. In the event of a conflict between any of the incorporated references and this specification, the specification shall control. Furthermore, any particular embodiment of the present invention that falls within the prior art may be expressly excluded from any one or more of the claims. Because such embodiments are deemed known to those of skill in the art, they may be excluded even if the exclusion is not expressly set forth herein. Any particular embodiment of the present invention may be excluded from any claim for any reason, whether or not related to the existence of prior art.
[0198] Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments described herein. The scope of the embodiments described herein is not intended to be limited to the above description, but rather is set forth in the appended claims. Those skilled in the art will appreciate that various changes and modifications to this description can be made without departing from the spirit or scope of the invention, as defined in the following claims.
[0199] The recitation of a list of chemical groups in any definition of a variable herein includes a definition of that variable as any single group or combination of the listed groups. The description of an embodiment of a variable herein includes that embodiment as any single embodiment or in combination with any other embodiment or portion thereof. The description of an embodiment herein includes that embodiment as any single embodiment or in combination with any other embodiment or portion thereof.
Claims
1. a first cleavage reagent comprising a first aminopeptidase derived from Pyrococcus horikoshii and a first tag sequence; a second cleavage reagent comprising a second aminopeptidase derived from Pyrococcus horikoshii; The composition, wherein the first and second aminopeptidases have different amino acid sequences from each other.
2. 2. The composition of claim 1, wherein the amino acid sequences of the first and second aminopeptidases share less than 80% sequence identity.
3. 3. The composition of claim 2, wherein the amino acid sequences of the first and second aminopeptidase share less than 70%, less than 60%, less than 50%, 10-80%, 20-60%, 30-50%, or 40-50% sequence identity.
4. 4. The composition of claim 1, wherein the first and second cleavage reagents are present in the composition at first and second concentrations, respectively, and the first concentration is at least two-fold higher than the second concentration.
5. 5. The composition of claim 1, wherein the molar ratio of the first cleavage reagent to the second cleavage reagent in the composition is from about 2:1 to about 20:
1.
6. 6. The composition of claim 5, wherein the molar ratio of the first cleavage reagent to the second cleavage reagent in the composition is from about 2:1 to about 15:1, from about 2:1 to about 10:1, from about 4:1 to about 15:1, or from about 5:1 to about 10:
1.
7. 7. The composition of claim 1, wherein the first cleavage reagent is present at a concentration of about 10 μM to about 100 μM, about 20 μM to about 60 μM, or about 30 μM to about 50 μM.
8. the first cleavage reagent is present in an amount sufficient to cleave the N-terminal amino acid from the polypeptide in an average cleavage time of about 2 to about 60 minutes, about 5 to about 50 minutes, about 5 to 30 minutes, about 5 to 20 minutes, about 5 to 15 minutes, or about 5 to 10 minutes; The composition of any one of claims 1 to 7, wherein the N-terminal amino acid comprises a charged side chain.
9. 9. The composition of any one of claims 1 to 8, wherein the second cleavage reagent is present at a concentration of about 0.1 μM to about 25 μM, about 0.5 μM to about 20 μM, or about 1 μM to about 10 μM.
10. the second cleavage reagent is present in an amount sufficient to cleave the N-terminal amino acid from the polypeptide in an average cleavage time of about 2 to about 60 minutes, about 5 to about 50 minutes, about 5 to 30 minutes, about 5 to 20 minutes, about 5 to 15 minutes, or about 5 to 10 minutes; The composition of any one of claims 1 to 9, wherein the N-terminal amino acid comprises a hydrophobic side chain.
11. The composition of any one of claims 1 to 10, wherein the first tag sequence is attached to the C-terminus of the first aminopeptidase.
12. The composition of any one of claims 1 to 11, wherein the second cleavage reagent comprises a second tag sequence.
13. The composition of claim 12 , wherein the second tag sequence is attached to the C-terminus of the second aminopeptidase.
14. The composition of any one of claims 1 to 13, wherein the first and second tag sequences each independently comprise at least two amino acids.
15. 15. The composition of claim 14, wherein the first and second tag sequences each independently comprise from about 2 to about 200 amino acids.
16. The composition of any one of claims 1 to 15, wherein at least one of the first and second tag sequences comprises a polyhistidine tag.
17. The composition of any one of claims 1 to 16, wherein at least one of the first and second tag sequences comprises a biotinylation tag.
18. 18. The composition of claim 17, wherein the biotinylation tag comprises at least one biotin ligase recognition sequence.
19. 19. The composition of claim 18, wherein the biotinylation tag comprises two biotin ligase recognition sequences oriented in tandem.
20. 20. The composition of any one of claims 1 to 19, wherein the first aminopeptidase comprises one or more substitutions relative to Pyrococcus horikoshii TET aminopeptidase III.
21. 21. The composition of any one of claims 1 to 20, wherein the second aminopeptidase comprises one or more substitutions relative to Pyrococcus horikoshii TET aminopeptidase II.
22. 22. The composition of any one of claims 1 to 21, wherein the first aminopeptidase has an amino acid sequence that is at least 80% identical to SEQ ID NO:
3.
23. 23. The composition of claim 22, wherein the amino acid sequence of the first aminopeptidase is at least 85%, at least 90%, at least 95%, 80-100%, 85-95%, 90-99%, 95-99%, or 100% identical to SEQ ID NO:
3.
24. The composition of any one of claims 1 to 23, wherein the first cleavage reagent is a polypeptide having an amino acid sequence that is at least 80% identical to SEQ ID NO:
4.
25. 25. The composition of claim 24, wherein the amino acid sequence of the first cleavage reagent is at least 85%, at least 90%, at least 95%, 80-100%, 85-95%, 90-99%, 95-99%, or 100% identical to SEQ ID NO:
4.
26. 26. The composition of any one of claims 1 to 25, wherein the second aminopeptidase has an amino acid sequence that is at least 80% identical to SEQ ID NO:
1.
27. 27. The composition of claim 26, wherein the amino acid sequence of the second aminopeptidase is at least 85%, at least 90%, at least 95%, 80-100%, 85-95%, 90-99%, 95-99%, or 100% identical to SEQ ID NO:
1.
28. The composition of any one of claims 1 to 27, wherein the second cleavage reagent is a polypeptide having an amino acid sequence that is at least 80% identical to SEQ ID NO:
2.
29. 29. The composition of claim 28, wherein the amino acid sequence of the second cleavage reagent is at least 85%, at least 90%, at least 95%, 80-100%, 85-95%, 90-99%, 95-99%, or 100% identical to SEQ ID NO:
2.
30. 30. The composition of any one of claims 1 to 29, further comprising a third cleavage reagent comprising aminopeptidase from Yersinia pestis.
31. a first cleavage reagent comprising a first aminopeptidase having an amino acid sequence at least 80% identical to SEQ ID NO:3; a second cleavage reagent comprising a second aminopeptidase having an amino acid sequence that is at least 80% identical to SEQ ID NO:1; and a third cleavage reagent comprising a third aminopeptidase having an amino acid sequence that is at least 80% identical to SEQ ID NO:5 or 7.
32. the first cleavage reagent comprises a first tag sequence; the second cleavage reagent comprises a second tag sequence; 32. The composition of claim 31, wherein the third cleavage reagent comprises a third tag sequence.
33. the first tag sequence is attached to the end of the first aminopeptidase; the second tag sequence is attached to the end of the second aminopeptidase; 33. The composition of claim 32, wherein the third tag sequence is attached to the terminus of the third aminopeptidase.
34. 34. The composition of claim 32 or 33, wherein each tag sequence is attached to the C-terminus of its respective aminopeptidase.
35. 35. The composition of any one of claims 32 to 34, wherein the first, second, and third tag sequences each independently comprise at least two amino acids.
36. 36. The composition of claim 35, wherein the first, second, and third tag sequences each independently comprise from about 2 to about 200 amino acids.
37. 37. The composition of any one of claims 32 to 36, wherein at least one of the first, second, and third tag sequences comprises a polyhistidine tag.
38. 38. The composition of any one of claims 32 to 37, wherein at least one of the first, second, and third tag sequences comprises a biotinylation tag.
39. 39. The composition of claim 38, wherein the biotinylation tag comprises at least one biotin ligase recognition sequence.
40. 40. The composition of claim 39, wherein the biotinylation tag comprises two biotin ligase recognition sequences oriented in tandem.
41. 41. The composition of any one of claims 31 to 40, wherein the first, second, and third cleavage reagents are present in the composition at first, second, and third concentrations, respectively, and the first concentration is at least two-fold greater than the second concentration.
42. 42. The composition of claim 41, wherein the first concentration is at least 5 times greater than the third concentration.
43. 43. The composition of claim 41 or 42, wherein the second concentration is at least two times higher than the third concentration.
44. 44. The composition of any one of claims 31 to 43, wherein the molar ratio of the first cleavage reagent to the second cleavage reagent in the composition is from about 2:1 to about 20:
1.
45. 45. The composition of claim 44, wherein the molar ratio of the first cleavage reagent to the second cleavage reagent in the composition is from about 2:1 to about 15:1, from about 2:1 to about 10:1, from about 4:1 to about 15:1, or from about 5:1 to about 10:
1.
46. 46. The composition of any one of claims 31 to 45, wherein the molar ratio of the first cleavage reagent to the third cleavage reagent in the composition is from about 5:1 to about 200:
1.
47. 47. The composition of claim 46, wherein the molar ratio of the first cleavage reagent to the third cleavage reagent in the composition is from about 5:1 to about 150:1, from about 5:1 to about 100:1, from about 10:1 to about 80:1, or from about 10:1 to about 50:
1.
48. 48. The composition of any one of claims 31-47, wherein the first cleavage reagent is present at a concentration of about 10 μM to about 100 μM, about 20 μM to about 60 μM, or about 30 μM to about 50 μM.
49. the first cleavage reagent is present in an amount sufficient to cleave the N-terminal amino acid from the polypeptide in an average cleavage time of about 2 to about 60 minutes, about 5 to about 50 minutes, about 5 to 30 minutes, about 5 to 20 minutes, about 5 to 15 minutes, or about 5 to 10 minutes; 49. The composition of any one of claims 31 to 48, wherein the N-terminal amino acid comprises a charged side chain.
50. 50. The composition of any one of claims 31-49, wherein the second cleavage reagent is present at a concentration of about 0.1 μM to about 25 μM, about 0.5 μM to about 20 μM, or about 1 μM to about 10 μM.
51. the second cleavage reagent is present in an amount sufficient to cleave the N-terminal amino acid from the polypeptide in an average cleavage time of about 2 to about 60 minutes, about 5 to about 50 minutes, about 5 to 30 minutes, about 5 to 20 minutes, about 5 to 15 minutes, or about 5 to 10 minutes; 51. The composition of any one of claims 31 to 50, wherein the N-terminal amino acid comprises a hydrophobic side chain.
52. 52. The composition of any one of claims 31-51, wherein the third cleavage reagent is present at a concentration of about 0.01 μM to about 25 μM, about 0.1 μM to about 20 μM, or about 0.5 μM to about 10 μM.
53. the third cleavage reagent is present in an amount sufficient to cleave the N-terminal amino acid from the polypeptide in an average cleavage time of about 2 to about 60 minutes, about 5 to about 50 minutes, about 5 to 30 minutes, about 5 to 20 minutes, about 5 to 15 minutes, or about 5 to 10 minutes; 53. The composition of any one of claims 31 to 52, wherein the polypeptide comprises an XP dipeptide motif, where X is the N-terminal amino acid and P is a proline amino acid.
54. 54. The composition of any one of claims 31-53, wherein the amino acid sequence of the first aminopeptidase is at least 85%, at least 90%, at least 95%, 80-100%, 85-95%, 90-99%, 95-99%, or 100% identical to SEQ ID NO:
3.
55. 55. The composition of any one of claims 31 to 54, wherein the first cleavage reagent is a polypeptide having an amino acid sequence that is at least 80% identical to SEQ ID NO:
4.
56. 56. The composition of claim 55, wherein the amino acid sequence of the first cleavage reagent is at least 85%, at least 90%, at least 95%, 80-100%, 85-95%, 90-99%, 95-99%, or 100% identical to SEQ ID NO:
4.
57. 57. The composition of any one of claims 31-56, wherein the amino acid sequence of the second aminopeptidase is at least 85%, at least 90%, at least 95%, 80-100%, 85-95%, 90-99%, 95-99%, or 100% identical to SEQ ID NO:
1.
58. 58. The composition of any one of claims 31 to 57, wherein the second cleavage reagent is a polypeptide having an amino acid sequence that is at least 80% identical to SEQ ID NO:
2.
59. 59. The composition of claim 58, wherein the amino acid sequence of the second cleavage reagent is at least 85%, at least 90%, at least 95%, 80-100%, 85-95%, 90-99%, 95-99%, or 100% identical to SEQ ID NO:
2.
60. 60. The composition of any one of claims 31-59, wherein the amino acid sequence of the third aminopeptidase is at least 85%, at least 90%, at least 95%, 80-100%, 85-95%, 90-99%, 95-99%, or 100% identical to SEQ ID NO:5 or 7.
61. 61. The composition of any one of claims 31 to 60, wherein the third cleavage reagent is a polypeptide having an amino acid sequence that is at least 80% identical to SEQ ID NO:
8.
62. 62. The composition of claim 61, wherein the amino acid sequence of the third cleavage reagent is at least 85%, at least 90%, at least 95%, 80-100%, 85-95%, 90-99%, 95-99%, or 100% identical to SEQ ID NO:
8.
63. A composition according to any one of claims 1 to 62; and one or more amino acid binding proteins that do not have peptide cleavage activity.
64. a cleavage reagent comprising a first aminopeptidase having an amino acid sequence at least 92% identical to SEQ ID NO:3 and comprising a first tag sequence; and one or more amino acid binding proteins that do not have peptide cleavage activity.
65. 65. The reaction mixture of claim 64, wherein the amino acid sequence of the first aminopeptidase is at least 94%, at least 96%, at least 98%, or 100% identical to SEQ ID NO:
3.
66. contacting a polypeptide with a reaction mixture according to any one of claims 63 to 65; monitoring a signal for a signal pulse corresponding to an interaction between one or more amino acid binding proteins and the polypeptide; and determining at least one chemical characteristic of said polypeptide based on a characteristic pattern in said signal.
67. detecting a series of signal pulses while the polypeptide is being degraded; 67. The method of claim 66, wherein a characteristic pattern in the series of signal pulses is indicative of the at least one chemical characteristic of the polypeptide.
68. 68. The method of claim 66 or 67, wherein determining the at least one chemical characteristic comprises identifying at least one amino acid in the polypeptide as a naturally occurring amino acid, a non-natural amino acid, or a modified variant thereof.
69. 69. The method of any one of claims 66-68, wherein determining the at least one chemical characteristic comprises identifying at least one amino acid in the polypeptide as having a side chain that is negatively charged, positively charged, uncharged, polar, nonpolar, hydrophobic, aromatic, or a combination thereof.
70. 70. The method of any one of claims 66-69, wherein determining the at least one chemical characteristic comprises identifying at least one amino acid in the polypeptide as one type selected from alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, selenocysteine, serine, threonine, tryptophan, tyrosine, and valine.
71. 71. The method of any one of claims 66-70, wherein determining the at least one chemical characteristic comprises identifying at least one amino acid in the polypeptide as having a post-translational modification.
72. 72. The method of claim 71, wherein the post-translational modification is selected from acetylation, ADP-ribosylation, caspase cleavage, citrullination, formylation, N-linked glycosylation, O-linked glycosylation, hydroxylation, methylation, myristoylation, neddylation, nitration, oxidation, palmitoylation, phosphorylation, prenylation, S-nitrosylation, sulfation, sumoylation, and ubiquitination.
73. at least one hardware processor; and at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by said at least one hardware processor, cause said at least one hardware processor to perform the method of any one of claims 66 to 72.
74. At least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by at least one hardware processor, cause the at least one hardware processor to perform the method of any one of claims 66 to 72.