Compositions and methods for polypeptide analysis

Recombinant amino acid binding proteins with modified binding pockets address scale and dynamic range limitations in proteomics, enhancing polypeptide analysis by increasing interactions and structural information.

JP2025525911APending Publication Date: 2025-08-07QUANTUM SI INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025505985
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-08-04
Filing Date
2023-08-03
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

Existing proteomics technologies face challenges in scale, dynamic range, and the inability to amplify sources for protein analysis, limiting the understanding of complex human diseases.

Method used

Development of recombinant or synthetic amino acid binding proteins with modified binding pockets and structures, such as formulas (I), (II), and (III), that enhance interactions with amino acid ligands, allowing for improved polypeptide analysis through amino acid recognition factors.

Benefits of technology

The modified binding proteins increase the number of interactions and types of amino acid ligands detected, improving the reliability and amount of structural information obtained from polypeptide analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025525911000001_ABST
    Figure 2025525911000001_ABST
Patent Text Reader

Abstract

Provided are amino acid recognition factors with improved binding properties that allow for obtaining more structural information from a polypeptide based on the on-off binding kinetics between the recognition factor and the polypeptide. The amino acid recognition factor can include an amino acid binding protein having an altered binding pocket with one or more modifications relative to a homologous protein. The modified binding pocket increases the number of interactions formed between the binding pocket and amino acid ligands compared to the unmodified binding pocket of the homologous protein, increases the number of types of amino acid ligands that can detectably bind compared to the unmodified binding pocket of the homologous protein, and improves the binding kinetics (e.g., K) for one or more types of amino acid ligands. D , k off , k on ) can be improved, which advantageously increases the amount or reliability of structural information that can be obtained from polypeptide analysis.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to compositions and methods for polypeptide analysis. [Background technology]

[0002] Proteins are fundamental building blocks of life, driving important biological and cellular processes. Protein function is driven by its structure, including its sequence. In adjacent fields such as genomics, advances in sequencing technology have proven invaluable in improving our understanding of the progression of complex human diseases. Applying similar approaches to proteomics has been challenging due to limitations in scale, dynamic range, and the inability to amplify sources. Summary of the Invention [Means for solving the problem]

[0003] In some embodiments, recombinant or synthetic amino acid binding proteins are provided having an amino acid sequence that is at least 80% identical to SEQ ID NO:1, wherein the amino acid sequence contains an amino acid substitution at one or more positions corresponding to E22, R31, L39, N41, D42, D43, D44, H45, T46, Y47, V50, Q55, P62, E63, L68, A69, V72, D73, Q75, Y100, and M111 of SEQ ID NO:1.

[0004] In some embodiments, a recombinant or synthetic amino acid binding protein is provided that comprises the structure of formula (I): β1-α1-α2-β2-α3-β3(I) or a structural equivalent thereof, wherein: each of β1, β2, and β3 is a β strand; each of α1, α2, and α3 is an α helix; each instance of "-" is a loop; and at least a portion of each of α1, α2, the loop between β1 and α1, and the loop between α3 and β3 form a binding pocket for an amino acid ligand, wherein this binding pocket comprises one or more of the following: (i) a length of approximately 170 Å 3 volume, (ii) -3.0RTec-1 (iii) negatively charged side chains on at least 35% of the amino acids that form the binding pocket; (iv) a plurality of hydrogen bond acceptors configured to form one or more hydrogen bonds in the presence of an amino acid ligand; and (v) a plurality of van der Waals contact positions configured to form van der Waals interactions in the presence of an amino acid ligand.

[0005] In some embodiments, recombinant or synthetic amino acid binding proteins are provided having an amino acid sequence that is at least 80% identical to SEQ ID NO:2, wherein the amino acid sequence contains an amino acid substitution at one or more positions corresponding to G19, K26, S29, F30, D31, D32, T33, C34, V35, T47, G48, T53, T54, T57, E58, F59, N61, 163, D65, D68, E70, A71, H74, and T75 of SEQ ID NO:2.

[0006] In some embodiments, a recombinant or synthetic amino acid binding protein is provided that comprises the structure of formula (II): β1-α1-β2-α2-α3(II) or a structural equivalent thereof, wherein: each of β1 and β2 is a β strand; each of α1, α2, and α3 is an α helix; each instance of "-" is a loop; and at least a portion of each of α2, the loop between β1 and α1, and the loop between β2 and α2 form a binding pocket for an amino acid ligand, wherein this binding pocket comprises: (i) a length of approximately 200 Å 3 volume, (ii) -3.0RTec -1 (iii) a plurality of hydrogen bond acceptors configured to form one or more hydrogen bonds in the presence of the amino acid ligand; and (iv) a plurality of van der Waals contact locations configured to form van der Waals interactions in the presence of the amino acid ligand.

[0007] In some embodiments, recombinant or synthetic amino acid binding proteins are provided having an amino acid sequence that is at least 80% identical to SEQ ID NO:3, wherein the amino acid sequence contains an amino acid substitution at one or more positions corresponding to S22, C23, Y24, C25, E26, S39, W75, D76, Y77, H78, C85, N120, H145, and M146 of SEQ ID NO:3.

[0008] In some embodiments, a recombinant or synthetic amino acid binding protein is provided that comprises the structure of formula (III): α1-α2-α3β1-β2-β3-β4-β5-α4-β6α5-α6(III), or a structural equivalent thereof, wherein each of α1, α2, α3, α4, α5, and α6 is an α-helix, each of β1, β2, β3, β4, β5, and β6 is a β-strand, each instance of "-" is a loop, and at least a portion of each of α2, β3, β4, α5, the loop between α1 and α2, and the loop between β3 and β4 form a binding pocket for an amino acid ligand, wherein the binding pocket comprises one or more of the following: (i) a length of approximately 160 Å 3 volume, (ii) -2.0RTec -1 (iii) a plurality of hydrogen bond acceptors or donors configured to form one or more hydrogen bonds in the presence of an amino acid ligand; (iv) a plurality of van der Waals contact locations configured to form van der Waals interactions in the presence of an amino acid ligand; and (v) at least one negatively charged amino acid and at least one positively charged amino acid.

[0009] In some embodiments, an amino acid recognition factor is provided that comprises a polypeptide having at least a first amino acid binding protein and a second amino acid binding protein linked end-to-end, the first and second amino acid binding proteins being separated by a linker comprising at least two amino acids, and at least one of the first and second amino acid binding proteins being an amino acid binding protein according to any of the aspects of the technology described herein.

[0010] In some embodiments, an amino acid recognition factor is provided that comprises a polypeptide having an amino acid binding protein and a labeled protein linked end-to-end, wherein the amino acid binding protein and the labeled protein are separated by a linker comprising at least two amino acids, and the amino acid binding protein is an amino acid binding protein according to any of the aspects of the technology described herein.

[0011] In some embodiments, a composition is provided comprising two or more amino acid recognition factors, wherein at least one amino acid recognition factor is an amino acid binding protein according to any of the aspects of the technology described herein.

[0012] According to some embodiments, there is provided a method for determining at least one chemical characteristic of a polypeptide, the method comprising contacting the polypeptide with a composition according to any of the aspects of the technology described herein, monitoring a signal for a signal pulse corresponding to an interaction between one or more amino acid recognition factors and the polypeptide, and determining at least one chemical characteristic of the polypeptide based on a characteristic pattern in the signal.

[0013] In some embodiments, a system is provided that includes at least one hardware processor and at least one non-transitory computer-readable storage medium that stores processor-executable instructions that, when executed by the at least one hardware processor, cause the at least one hardware processor to perform a method according to any of the aspects of the technology described herein.

[0014] In some embodiments, at least one non-transitory computer-readable storage medium is provided that stores processor-executable instructions that, when executed by at least one hardware processor, cause the at least one hardware processor to perform a method according to any of the aspects of the technology described herein.

[0015] The details of certain embodiments of the present disclosure are set forth in the detailed description. Other features, objects, and advantages of the invention will be apparent from the examples, drawings, and claims. [Brief explanation of the drawings]

[0016] [Figure 1] An exemplary overview of real-time dynamic protein sequencing is shown. Protein samples are digested into peptide fragments, immobilized in a nanoscale reaction chamber, and incubated with a mixture of freely diffusing N-terminal amino acid (NAA) recognition factors and aminopeptidases that perform the sequencing process. Labeled recognition factors bind to peptides on and off when one of their cognate NAAs is exposed at the N-terminus, thereby generating a characteristic pulse pattern. The NAA is cleaved by the aminopeptidase, exposing the next amino acid for recognition. The temporal order and binding kinetics of NAA recognition allow for peptide identification and are sensitive to features that modulate binding kinetics, such as post-translational modifications (PTMs). SEQ ID NO: 1090 (RLIFA) is shown. [Figure 2A] Figures 2A-2G show examples of NAA recognition and dynamic sequencing. Figures 2A-2C show example traces demonstrating single-molecule N-terminal recognition by PS610 (Figure 2A), PS961 (Figure 2B), and PS691 (Figure 2C); scatter plots of the number of pulses per recognition segment (RS) versus RS mean pulse duration (PD) are shown for each peptide in Figures 2A-2C, with median PD indicated. Figure 2D shows example traces from dynamic sequencing of the synthetic peptide FAAWAAYAAAADDD (SEQ ID NO: 1034). Median PD is indicated above each RS. Figures 2E-2G show dynamic sequencing of the synthetic peptide LAQFASIAAYASDDD (SEQ ID NO: 1035) using PS610 and PS961. Figure 2E shows example traces. Figure 2F shows scatter plots of RS mean PD versus bin ratio, demonstrating discrimination of recognition factors by bin ratio and NAA by pulse duration. Figure 2G shows a scatter plot of the number of pulses per RS versus the RS mean PD, grouped by the amino acid label assigned to the RS. [Figure 2B] Same as above. [Figure 2C] Same as above. [Figure 2D] Same as above. [Figure 2E] Same as above. [Figure 2F] Same as above. [Figure 2G] Same as above. [Figure 3A] Figures 3A-3G show examples of dynamic sequencing of various peptides with high-precision kinetic output. Figures 3A-3E show the dynamic sequencing of peptide DQQRLIFAG (SEQ ID NO: 1036). Figure 3A shows an example trace. Figure 3B shows a scatter plot of RS average PD vs. bin ratio. Figure 3C shows additional example traces of dynamic sequencing of DQQRLIFAG (SEQ ID NO: 1036). Figure 3D shows the distribution of durations of each RS and non-recognition segment (NRS) obtained during sequencing, with the average duration shown. Figure 3E shows a kinetic signature plot summarizing the characteristic sequencing behavior of the DQQRLIFAG (SEQ ID NO: 1036) peptide. Figures 3F-3G show the kinetic sequencing of synthetic peptides DQQIASSRLAASFAAQQYPDDD (SEQ ID NO: 1037) (top), RLAFSALGAADDD (SEQ ID NO: 1038) (middle), and EFIAWLV (SEQ ID NO: 1039) (bottom). Figure 3F shows example traces for each peptide. Figure 3G shows the corresponding kinetic signature plots. [Figure 3B] Same as above. [Figure 3C] Same as above. [Figure 3D] Same as above. [Figure 3E] Same as above. [Figure 3F] Same as above. [Figure 3G] Same as above. [Figure 4A]Figures 4A-4E show examples of single amino acid change and PTM detection. Figures 4A-4B show dynamic sequencing of synthetic peptides differing by a single amino acid: RLAFAYPDDD (SEQ ID NO: 1040) (top), RLIFAYPDDD (SEQ ID NO: 1041) (middle), and RLVFAYPDDD (SEQ ID NO: 1042) (bottom). Figure 4A shows an example trace. Figure 4B shows a scatter plot of RS mean PD vs. bin ratio. Figures 4C-4D show detection of oxidized methionine using peptide RLMFAYPDDD (SEQ ID NO: 1043). Figure 4C shows the distribution of mean PD for leucine; labels indicate populations where leucine is followed by methionine (LM) or methionine sulfoxide (LMo). Figure 4D shows examples of traces where methionine is recognized by PS961 and leucine shows a long PD (top), or where methionine is not recognized due to oxidation (RLMoFAYPDDD (SEQ ID NO: 1044), where "Mo" is methionine sulfoxide) and leucine shows a short PD (bottom). Figure 4E shows scatter plots of RS mean PD vs. bin ratio for runs where oxidation was uncontrolled (top) or where methionine was fully oxidized (bottom). [Figure 4B] Same as above. [Figure 4C] Same as above. [Figure 4D] Same as above. [Figure 4E] Same as above. [Figure 5A]Figures 5A-5C show examples of peptide discrimination in a mixture and mapping of peptides to the human proteome. Figure 5A shows example traces from sequencing a mixture of peptides DQQRLIFAG (SEQ ID NO: 1036) and RLAFSALGAADDD (SEQ ID NO: 1038) on the same chip; the chip window indicates the position of the reaction chamber generating the sequencing readout for each peptide. Figure 5B shows example traces from dynamic sequencing of two peptides, DQQRLIFAGK (SEQ ID NO: 1045) (top) and EFIAWLVK (SEQ ID NO: 1046) (bottom), isolated from recombinant human proteins ubiquitin and GLP-1, respectively. Figure 5C shows a diagram illustrating the identification of the protein ubiquitin as a match to the kinetic signature derived from the DQQRLIFAGK (SEQ ID NO: 1045) peptide in an in silico digest of the human proteome based on kinetic information. IVNFSRLIFHHLK (SEQ ID NO: 1095), DIRLIFSNAK (SEQ ID NO: 1096), GQSRLIFTYGLTNSGK (SEQ ID NO: 1097), DQQRLLIFAGK (SEQ ID NO: 1045), and DEHCLRLIFLK (SEQ ID NO: 1098) are also shown. [Figure 5B] Same as above. [Figure 5C] Same as above. [Figure 6A]Figures 6A-6F show examples of chip operation. Figure 6A shows an exploded view of a custom semiconductor chip and a small benchtop instrument designed to support protein sequencing assays. Figure 6B shows the chip achieving electronic rejection by discarding photoelectrons from a pulsed laser, then shifting to collect fluorescence photoelectrons from the bound NAA recognition factor; the timing of the rejection and collection windows cycles between two modes (Bin 1 and Bin 0, example waveforms are shown) in alternating frames to provide bin ratio estimates of the dye's fluorescence lifetime. Figure 6C shows the chip achieving a >10,000-fold attenuation of the incident laser light within 1 ns of the start of rejection mode. Figure 6D shows example pulses for dyes with short and long fluorescence lifetimes, illustrating the difference in signal collection in Bin 0 and Bin 1. Figure 6E shows the distribution of average RS bin ratios collected for three dyes with different fluorescence lifetimes. Figure 6F shows that dye channel identification accuracy improves with the number of pulses captured per RS. [Figure 6B] Same as above. [Figure 6C] Same as above. [Figure 6D] Same as above. [Figure 6E] Same as above. [Figure 6F] Same as above. [Figure 7A]Figures 7A-7H show examples of recognition factor properties. Figures 7A-7E show recognition factor kinetic characterization using polarization assays (Example 1, Methods). Figures 7A-7B show the affinity (KD) (Figure 7A) and off-rate (koff) (Figure 7B) of PS610 for peptides with N-terminal phenylalanine, tyrosine, and tryptophan. Figure 7B shows FAKLK(FITC)DEESILKQ (SEQ ID NO: 1099), YAKLK(FITC)DEESILKQ (SEQ ID NO: 1100), and WAKLK(FITC)DEESILKQ (SEQ ID NO: 1101). Figure 7C shows the affinity of PS961 for peptides with N-terminal leucine, isoleucine, and valine. Figures 7D-7E show the affinity of PS691 for peptides with an N-terminal arginine (Figure 7D) and single-point polarization data measured for peptides with N-terminal arginine, lysine, and histidine (Figure 7E). Figure 7F shows the binding energies calculated using a computer model (Example 1, Methods) for peptides with the initial sequences LAX and LXA (where X = all 20 amino acids); boxplots show the fraction of the total binding energy contributed by amino acids at positions 1 (P1), 2 (P2), and 3 (P3), which tend to decrease exponentially from P1 to P3 (R2 > 0.97). Figure 7G shows the RS average PDs determined in single-molecule assays for LXA and LAX peptides using PS961 and for FXA and FAX peptides using PS610. Figure 7H shows that the nonpolar solvation energy terms from the computational binding model with PS961 correlate highly with the actual RS average PD values observed in single-molecule assays with peptides containing an N-terminal leucine and various amino acids at the P2 position: LVFA (SEQ ID NO: 1102), LIFA (SEQ ID NO: 1103), LVAR (SEQ ID NO: 1104), LAFA (SEQ ID NO: 1105), LQAR (SEQ ID NO: 1106), LDAA (SEQ ID NO: 1107), LCAR (SEQ ID NO: 1108), LGAA (SEQ ID NO: 1109), LMFA (SEQ ID NO: 1110), LSAR (SEQ ID NO: 1111), and LEFA (SEQ ID NO: 1112). [Figure 7B] Same as above. [Figure 7C] Same as above. [Figure 7D] Same as above. [Figure 7E] Same as above. [Figure 7F] Same as above. [Figure 7G] Same as above. [Figure 7H] Same as above. [Figure 8A] Figures 8A-8E show examples of binding and cleavage rates. Figures 8A-8B show RS mean interpulse durations (IPDs) for PS961 binding to LIF (Figure 8A) and IFA (Figure 8B) in dynamic sequencing assays at concentrations of 125 nM (orange) or 250 nM (blue); median IPD values are shown. Figure 8C shows single-exponential decay curves fitted to RS duration distributions for arginine, leucine, isoleucine, and phenylalanine obtained from dynamic sequencing of the synthetic peptide DQQRLIFAG (SEQ ID NO: 1036). Figures 8D-8E show that increasing aminopeptidase concentration in dynamic sequencing runs of the synthetic peptide DQQRLIFAG (SEQ ID NO: 1036) resulted in decreased NRS (Figure 8D) and RS (Figure 8E) durations; median RS duration values are shown. [Figure 8B] Same as above. [Figure 8C] Same as above. [Figure 8D] Same as above. [Figure 8E] Same as above. [Figure 9A]Examples of kinetic signatures derived from single amino acid changes and PTMs are shown. Figure 9A shows kinetic signature plots for three peptides: RLAFAYPDDD (SEQ ID NO: 1040) (top), RLIFAYPDDD (SEQ ID NO: 1041) (middle), and RLVFAYPDDD (SEQ ID NO: 1042) (bottom). Figures 9B-9C show incomplete RS information observed in dynamic sequencing of the RLIFAYPDDD (SEQ ID NO: 1041) peptide. Figure 9B shows the percentage of reads of each type and example traces of observed deletion of one or more RSs in traces starting with arginine recognition and ending with tyrosine recognition. RLIFY (SEQ ID NO: 1115) is shown. Figure 9C shows the percentage of reads of each type and example traces of observed truncation of one or more RSs in traces starting with arginine. Figure 9D shows the affinity of PS961 for peptides with an N-terminal methionine as measured by polarization assay (Example 1, Methods). Figure 9E shows binding energy predictions for peptides with N-terminal methionine (MFAY (SEQ ID NO: 1113)) and methionine sulfoxide (Mo) (MoFAY (SEQ ID NO: 1114)) derived from computer modeling using PS961 (Example 1, Methods). Figure 9F shows kinetic signature plots for DQQRLIFAG (SEQ ID NO: 1036) and RLAFSALGAADDD (SEQ ID NO: 1038) peptides mixed and run on the same chip. Figure 9G shows kinetic signature plots for DQQRLIFAGK (SEQ ID NO: 1045) and EFIAWLVK (SEQ ID NO: 1046) peptides obtained from digestion of recombinant human ubiquitin and GLP-1. [Figure 9B] Same as above. [Figure 9C] Same as above. [Figure 9D] Same as above. [Figure 9E] Same as above. [Figure 9F] Same as above. [Figure 9G] Same as above. [Figure 10A]Examples of peptide identification using modeled proteome-wide kinetic signatures are shown. Figures 10A-10C show heat maps of predicted pulse durations for PS961-binding tripeptide targets with leucine (Figure 10A), isoleucine (Figure 10B), or valine (Figure 10C) at the N-terminal position. Figures 10D-10F show heat maps of predicted pulse durations for PS610-binding tripeptide targets with phenylalanine (Figure 10D), tyrosine (Figure 10E), or tryptophan (Figure 10F) at the N-terminal position. Figure 10G shows heat maps of predicted pulse durations for PS1122-binding tripeptide targets with arginine at the N-terminal position. Figure 10H shows plots demonstrating high correlation between predicted and actual pulse durations from on-chip experiments for PS961 (left plot) and PS610 (right plot). Figures 10I-10K show results from an analysis of the human proteome. In Figure 10J, the sequences for IL6_HUMAN (SEQ ID NO: 1116), DGISALRK (SEQ ID NO: 1117), SNMCESSK (SEQ ID NO: 1118), EALAENNLNLPK (SEQ ID NO: 1119), DGCFQSGFNEETCLVK (SEQ ID NO: 1120), IITGLLEFEVYLEYLQNRFESSEEQARAVQMSTK (SEQ ID NO: 1121), VLIQFLQK (SEQ ID NO: 1122), DPTTNASLLTK (SEQ ID NO: 1123), and DMTTHLILRSFK (SEQ ID NO: 1124) are shown. In Figure 10K, ACLILRSEEELK (SEQ ID NO: 1125), DMTTHLILRSFK (SEQ ID NO: 1124), SDSRNTLILRK (SEQ ID NO: 1126), DSSHQISALVLRAQASEILLEELQQGLSQAK (SEQ ID NO: 1127), ARTVGIEELILRIqESK (SEQ ID NO: 1128), STLVLRCHRRRK (SEQ ID NO: 1129), DSPQEPLVLRLK (SEQ ID NO: 1130), and DLVLRATK (SEQ ID NO: 1131) are shown. Figures 10L-10M show results from an analysis of the E. coli proteome.In Figure 10M, the sequences for SSUA_ECOLI (SEQ ID NO: 1132), LALAGLLSVSTFAVAAESSPEALRIGYQK (SEQ ID NO: 1133), GSSSHNLLRALRQAGLK (SEQ ID NO: 1134), DPYYSAALLQGGVRVLK (SEQ ID NO: 1135), DLNQTGSFYLAARPYAEK (SEQ ID NO: 1136), DLFYENRLVPK (SEQ ID NO: 1137), and DIRQRIWQPLEGK (SEQ ID NO: 1138) are shown. [Figure 10B] Same as above. [Figure 10C] Same as above. [Figure 10D] Same as above. [Figure 10E] Same as above. [Figure 10F] Same as above. [Figure 10G] Same as above. [Figure 10H] Same as above. [Figure 10I] Same as above. [Figure 10J] Same as above. [Figure 10K] Same as above. [Figure 10L] Same as above. [Figure 10M] Same as above. [Figure 11A] Figures 11A-11D show exemplary results from the selection and analysis of N-terminal alanine and valine binding mutants. Figure 11A shows results from three rounds of FACS selection. Figure 11B shows results from one selection round of an error-prone PCR library mixture. Figure 11C shows an exemplary diagram of predicted binding of an alanine peptide to variants of PS557. Figure 11D shows exemplary results from a fluorescence polarization study comparing the kinetics of N-terminal alanine peptide binding for selected PS557 variants. [Figure 11B] Same as above. [Figure 11C] Same as above. [Figure 11D] Same as above. [Figure 12A]Figures 12A-12B show exemplary results from the rational design of recognition factor mutants of PS557. Figure 12A shows a heatmap illustrating the enrichment of mutations in the PS557 protein. Figure 12B shows binding assay traces of candidates selected from the Octet platform. [Figure 12B] Same as above. [Figure 13A] Figures 13A-13D show exemplary results from the development of arginine recognition factors. Figure 13A shows the polarization responses for PS621 variants binding to RA, KA, and HA peptides (top chart), as well as the polarization-determined binding affinities (Kd) for selected PS621 variants for RA and HA at 20 °C (bottom table). Figure 13B shows the Kd determination titration curves for RA binding by PS621, PS691, and PS1122. Figure 13C shows on-chip recognition of RA dipeptide by PS1122 using QP304-RAIFAG in a recognition-on-chip assay. Figure 13D shows an example of multiplexed kinetic chip analysis of PS1122 to demonstrate its enhanced range of arginine tripeptide coverage. Three peptides containing different RXA motifs (RLQFQALMAADDD (SEQ ID NO: 1139), LAQRQAFDAADDD (SEQ ID NO: 1140), FAQLQARFAADDD (SEQ ID NO: 1141)) were evaluated simultaneously in one sequencing run. [Figure 13B] Same as above. [Figure 13C] Same as above. [Figure 13D] Same as above. [Figure 14A]Figures 14A-14C show exemplary results from computer modeling of PS961, demonstrating that N41D enhances electrostatic interactions with the N-terminal amino group of a polypeptide. Figure 14A shows that the amino acid triad D73, D43, and N41 (left panel) or D41 (right panel) binds to the N-terminal amino group of an AAA-tripeptide (sticks, top center) through hydrogen bonds (dashed lines). Figure 14B shows that the PS961 binding pocket is more negatively charged than that of PS557, thus increasing the likelihood of protein-peptide interactions. Figure 14C shows that the average fa_elec energy term of the peptide N-terminus across 50 representative structures from molecular dynamics simulations is lower (more favorable) for PS961 than for PS557. [Figure 14B] Same as above. [Figure 14C] Same as above. [Figure 15] Figure 1 shows exemplary results from computer modeling of PS961, demonstrating that V72M increases hydrophobic interactions with the peptide, further compacting the protein core. Because the methionine side chain is longer, it is positioned closer to adjacent residues (van der Waals radii shown as spheres) and can interact with the side chain of the peptide N-terminus than its valine precursor (sticks and spheres, top center). [Figure 16] Figure 1 shows exemplary results from computer modeling of PS961, demonstrating that L68M further optimizes protein core packing. Similar to V72M, the longer methionine side chain extends further into the non-optimal protein cavity than leucine (van der Waals radii shown as spheres), providing stabilizing hydrophobic interactions. [Figure 17] 1 shows exemplary results from computer modeling of PS961, indicating that the Y100R mutation neutralizes and increases the charge of a small negative surface pocket distant from the binding pocket that may compete with the intended recognition pocket. [Figure 18A]Figures 18A-18B show exemplary results from computational modeling of PS961, demonstrating that Y100R enables a unique loop structure that is more susceptible to antepenultimate interactions. Figure 18A shows that R100 stabilizes an extensive hydrogen-bonding network between the penultimate backbone (AP; bar, left), R106, and other members of the loop. Figure 18B shows that the average occupancy of the R100:R106 and R106:AP hydrogen bonds shown in Figure 18A is higher in the PS961 simulation compared to PS557 (50 ns trajectories, n = 3). [Figure 18B] Same as above. [Figure 19A] Figures 19A-19C show the secondary structure, sequence, and binding pocket characteristics of PS961. Figure 19A shows the classification of the protein into secondary structure groups. Figure 19B shows a Poisson-Boltzmann electrostatic potential surface map of the binding pocket with the residues that form the binding pocket labeled and the corresponding pocket characteristics listed. Figure 19C shows the sequence of residues 32-116 of the native parent protein (PS557 (SEQ ID NO: 1)) and the engineered variant (PS961 (SEQ ID NO: 314)), highlighting the mutations, pocket positions, and secondary structure assignments per position. [Figure 19B] Same as above. [Figure 19C] Same as above. [Figure 19D]Figures 19D-19S show exemplary results from crystal structure analysis of PS961 in complex with an N-terminal methionine peptide (Figures 19D-19K) or an N-terminal alanine peptide (Figures 19L-19S). Figure 19D shows a crystal of the PS961:MAKL complex. Figure 19E shows the crystal structure of recognition factor PS961 (surface) complexed with the target peptide MAKL (SEQ ID NO: 1047) (stick). Figures 19F-19K show how PS961 binds to the target peptide MAKL (SEQ ID NO: 1047). Figure 19L shows the crystal structure of recognition factor PS961 (cartoon) complexed with the target peptide AAKL (SEQ ID NO: 1048) (stick). Figure 19M shows a superposition of the main chains of PS961 bound to MAKL (SEQ ID NO: 1047) and AAKL (SEQ ID NO: 1048). Figure 19N shows the substitution of the Asp42 side chain in the recognition factor when comparing binding to AAKL (SEQ ID NO: 1048) versus MAKL (SEQ ID NO: 1047). Figure 19O shows a comparison of the AAKL (SEQ ID NO: 1048) and MAKL (SEQ ID NO: 1047) peptides when bound to PS961. Figure 19P shows a comparison of the interactions of the AAKL (SEQ ID NO: 1048) (left panel) and MAKL (SEQ ID NO: 1047) (right panel) peptides with residues in PS961. Figure 19Q shows the reorientation of Asp12 in PS961 for binding to either the MAKL (SEQ ID NO: 1047) (with a water molecule as a mediator) or AAKL (SEQ ID NO: 1048) peptide. Figure 19R shows the reorientation of Asp42 in PS961 for binding to either the MAKL (SEQ ID NO: 1047) or AAKL (SEQ ID NO: 1048) peptide. Figure 19S shows a superposition of the AAKL (SEQ ID NO: 1048) and MAKL (SEQ ID NO: 1047) peptides, showing the 180° flip of the LYS side chain at the third position (left) and the different orientation of the Lys3 side chain in the MAKL (SEQ ID NO: 1047) and AAKL (SEQ ID NO: 1048) peptides when bound to PS961 (right). [Figure 19E] Same as above. [Figure 19F] Same as above. [Figure 19G] Same as above. [Figure 19H] Same as above. [Figure 19I] Same as above. [Figure 19J] Same as above. [Figure 19K] Same as above. [Figure 19L] Same as above. [Figure 19M] Same as above. [Figure 19N] Same as above. [Figure 19O] Same as above. [Figure 19P] Same as above. [Figure 19Q] Same as above. [Figure 19R] Same as above. [Figure 19S] Same as above. [Figure 20A] Figure 20A shows exemplary results from computer modeling of PS1122, showing the PS621 crystal structure electrostatic surface with the modeled mutations. Top right: The E70T mutation forms a novel hydrogen bond with the N-terminal arginine side chain. Right-center: The I63E mutation forms a novel hydrogen bond with the amino terminus. Bottom right: The T47L mutation was found to be a common mutation across binding preferences and likely improves protein stability by stabilizing a helical turn near the metal binding site. [Figure 20B] Figures 20B-20F show exemplary results from crystal structure analysis of PS1122 complexed with the N-terminal arginine peptide (RAKL (SEQ ID NO: 1049)). Figure 20B shows a crystal of the PS1122:RAKL complex. Figure 20C shows the crystal structure of the recognition factor PS1122 complexed with the target peptide RAKL (SEQ ID NO: 1049) (sticks). Figure 20D shows that a portion of PS1122 differs from the similar portion observed in PS621 (PS1122 precursor). Figure 20E shows that the NH3 group of the first amino acid of Arg-1 of the bound peptide interacts with amino acids Glu-63, Asp-65, and a water molecule held in place by Glu-63 (interactions are shown as dashed lines). Figure 20F shows that both the side chains of Glu-63 and Thr-70 of PS1122 interact with Arg-1 of the bound peptide. [Figure 20C] Same as above. [Figure 20D] Same as above. [Figure 20E] Same as above. [Figure 20F] Same as above. [Figure 21] Figures 21A-21C show the secondary structure, sequence, and binding pocket characteristics of PS1122. Figure 21A shows the classification of the protein into secondary structure groups. Figure 21B shows a Poisson-Boltzmann electrostatic potential surface map of the binding pocket with the residues that form the binding pocket labeled and the corresponding pocket characteristics listed. Figure 21C shows residues 1-82 of the sequence of the native parent protein (PS621 (SEQ ID NO: 2)) and the engineered variant (PS1122 (SEQ ID NO: 468)), highlighting the mutations, pocket positions, and secondary structure assignments per position. [Figure 22] An alphafold model of PS1259 is shown with a network of hydrogen bonds enabled by two mutations: C25S (shown as a stick below the N-terminal glutamine) and H78Q (shown as a stick above the N-terminal glutamine). [Figure 23A] Figures 23A-23C show the secondary structure, sequence, and binding pocket characteristics of PS1259. Figure 23A shows the classification of the protein into secondary structure groups. Figure 23B shows a Poisson-Boltzmann electrostatic potential surface map of the binding pocket with the residues that form the binding pocket labeled and the corresponding pocket characteristics listed. Figure 23C shows the sequences of the native parent protein (Ntaq1(sf) (SEQ ID NO: 3)) and the engineered mutant PS1259 (SEQ ID NO: 605), highlighting the mutations, pocket positions, and secondary structure assignments per position. [Figure 23B] Same as above. [Figure 23C] Same as above. [Figure 24A]Figures 24A-24D show exemplary results demonstrating the direct identification of arginine PTMs. Figure 24A shows different arginine PTMs (YRELRLLK (SEQ ID NO: 1077), YRADMAELRLLK (SEQ ID NO: 1078), and YRSDMAELRLLK (SEQ ID NO: 1079)), including symmetric dimethylarginine (SDMA), asymmetric dimethylarginine (ADMA), and citrullinated arginine. Figure 24B shows an exemplary workflow for collecting samples, preparing a library of digested peptides, loading onto the chip, and performing on-chip sequencing and data analysis. Figure 24C shows sequencing data demonstrating that kinetic signatures distinguish between peptides containing arginine, ADMA, and SDMA. Figure 24C-A shows example protein sequencing traces for three synthetic p38 MAPKα-derived peptides containing arginine, ADMA, or SDMA at position 2. The full-length peptide sequence is shown for each example trace. Figures 24C-B show the distribution of recognition segment (RS) mean pulse durations (PD) for RSs corresponding to the first four residue sequences of each peptide: YREL (SEQ ID NO: 1050) (left), YRADMAEL (SEQ ID NO: 1051) (center), and YRSDMAEL (SEQ ID NO: 1052) (right). The median is shown for each distribution. Figures 24C-C show the inter-pulse durations (IPD) for arginine versus ADMA detection by PS621. Figure 24D shows sequencing data demonstrating that the kinetic signature distinguishes between peptides containing arginine and citrulline. Figure 24D-A shows example protein sequencing traces for two synthetic peptides containing arginine or citrulline at position 2: the peptide sequence LRLAFAYPDDDK (SEQ ID NO: 1053) (QP707) and the citrullinated peptide sequence LRCitLAFAYPDDDK (SEQ ID NO: 1054) (QP789). The full-length peptide sequence is shown for each example trace. Figures 24D-B show the distribution of RS mean PD for RS corresponding to the first 5 residue sequence of each peptide: LRLAF (SEQ ID NO: 1055) (left) and LCitLAF (SEQ ID NO: 1056) (right). The median is shown for each distribution. [Figure 24B] Same as above. [Figure 24C-A]Same as above. [Figure 24C-B] Same as above. [Figure 24C-C] Same as above. [Figure 24D-A] Same as above. [Figure 24D-B] Same as above. [Figure 25A] Figures 25A-25C show exemplary results demonstrating the identification of threonine PTMs. Figures 25A-25B show results from sequencing reactions using recognition factors PS691, PS610, and PS961 for peptides RLTFIAYPDDD (SEQ ID NO: 1057) (Figure 25A); and RLpTFIAYPDDD (SEQ ID NO: 1058), where pT is phosphothreonine (Figure 25B). Figure 25C shows the recognition segment (RS) duration for leucine recognition in the sequencing reactions of Figures 25A (left panel) and 25B (right panel). [Figure 25B] Same as above. [Figure 25C] Same as above. [Figure 26A] Figures 26A-B show example results showing the identification of tyrosine PTMs in sequencing reactions using recognition factors PS691, PS610, and PS961 for RLYFIAYPDDD (SEQ ID NO: 1059) (Figure 26A); and RLpYFIAYPDDD (SEQ ID NO: 1060) (where pY is phosphotyrosine) (Figure 26B). [Figure 26B] Same as above. [Figure 27A] Figures 27A-B show example results showing the identification of lysine PTMs in sequencing reactions using recognition factors PS691, PS610, PS961, and PS1165 for the peptides RLYFKAYPDDD (SEQ ID NO: 1061) (Figure 27A); and RLK{acetyl}FIAYPDDD (SEQ ID NO: 1062), where K{acetyl} is acetylated lysine (Figure 27B). [Figure 27B] Same as above. [Figure 28A]Figures 28A-G show an exemplary application of the present technology to the identification of β-amyloid variants. Figure 28A shows an example of a β-amyloid variant. Figure 28B shows an example workflow for β-amyloid variant detection. Figures 28C-G show example pulse patterns of β-amyloid wild-type LVFFAE (SEQ ID NO: 1063) versus variants (LVFFAK (SEQ ID NO: 1064), LVFFGK (SEQ ID NO: 1065), LVFFAG (SEQ ID NO: 1066), LVPFAE (SEQ ID NO: 1067)). [Figure 28B] Same as above. [Figure 28C] Same as above. [Figure 28D] Same as above. [Figure 28E] Same as above. [Figure 28F] Same as above. [Figure 28G] Same as above. [Figure 29] 1 shows an exemplary schematic diagram of a pixel of an integrated device. [Figure 30A] Figures 30A-30F show exemplary results from computer modeling (Figures 30A-30E) and structures (Figure 30F) of model peptides (RLAF (SEQ ID NO: 1142) (Figure 30B), ADMA-LAF (SEQ ID NO: 1143) (Figure 30C), SDMA-LAF (SEQ ID NO: 1144) (Figure 30D), and Cit-LAF (SEQ ID NO: 1145) (Figure 30E)) evaluated with PS621 and PS1122. [Figure 30B] Same as above. [Figure 30C] Same as above. [Figure 30D] Same as above. [Figure 30E] Same as above. [Figure 30F-1] Same as above. [Figure 30F-2] Same as above. [Figure 30F-3] Same as above. [Figure 31] Amino acid frequencies across the human proteome are shown. [Figure 32A]High-throughput expression, purification, and conjugation with streptavidin of hNTAQ protein homologue variants is shown. [Figure 32B] Shown are example traces of sequencing reactions performed with six recognition factors, including the glutamate recognition factor PS1875. [Figure 32C] Expression and purification of bis-biotinylated PS2132 at 2 L scale is shown. [Figure 32D] 1 shows exemplary results from size-exclusion chromatography (SEC) of PS2132 labeled with a streptavidin-binding long-lifetime BODIPY dye. [Figure 32E] 1 shows exemplary results from a quality control SDS PAGE gel analysis of PS2132 labeled with a long-lived BODIPY dye before and after SEC column purification. [Figure 33A] 33A-33B show exemplary results from on-chip recognition of E by long-lived BODIPY dye-labeled PS1875 (FIG. 33A) and PS2132 (FIG. 33B). [Figure 33B] Same as above. [Figure 34A] 34A-34B show exemplary results from on-chip recognition of E by long-lived BODIPY dye-labeled PS1875 (FIG. 34A) and PS2121 (FIG. 34B). [Figure 34B] Same as above. [Figure 35A] 35A-35B show exemplary results from on-chip recognition of E by long-lived BODIPY dye-labeled PS1875 (FIG. 35A) and PS2123 (FIG. 35B). [Figure 35B] Same as above. [Figure 36] Shown are pulse durations (top) and inter-pulse durations (bottom) of RS corresponding to glutamic acid (E) recognition in aligned reads of a sequencing run performed using QP1165 (EIAFLKQRVWK (SEQ ID NO: 1084)) peptide with a mixture of six recognition factors containing either PS1875, PS2132, PS2121, or PS2123 as the recognition factor for E. [Figure 37A] Figures 37A-37C show exemplary results from sequencing runs of a CDNF library performed using Reagent A, which contains a mixture of five recognition factors (Figure 37A) or Reagent A combined with the E recognition factor PS2132 (Figure 37B). Figure 37C shows exemplary traces for the sequence of CDNF (SEQ ID NO: 1146) and three peptides: EFLNRFYK (SEQ ID NO: 1068), ELISFCLDTK (SEQ ID NO: 1069), and ENRLCYYLGATK (SEQ ID NO: 1070). [Figure 37B] Same as above. [Figure 37C] Same as above. [Figure 38A] Figures 38A-38C show exemplary results from sequencing runs performed on a GFAP peptide library with (Figure 38A) or without (Figure 38B) the recognition factor for glutamic acid (E) PS2132 in combination with five recognition factors. Figure 38C shows exemplary traces for the sequence of GFAP (SEQ ID NO: 1147) and two peptides: DEMARHLQEYQDLLNVK (SEQ ID NO: 1071) and LALDIEIATYRK (SEQ ID NO: 1072). [Figure 38B] Same as above. [Figure 38C] Same as above. [Figure 39A] Figures 39A-39B show exemplary results from computer modeling of PS2132. Figure 39A shows the binding pocket of PS2132 complexed with an N-terminal glutamic acid peptide. Figure 39B shows surface charge modeling for PS1259 and PS2132. [Figure 39B] Same as above. DETAILED DESCRIPTION OF THE INVENTION

[0017] The accompanying drawings, which constitute a part of this specification, illustrate several embodiments of the present disclosure and, together with the accompanying description, serve to explain the principles of the disclosure. Aspects of the present disclosure relate to compositions and methods for determining the chemical characteristics of a polypeptide based on single-molecule binding interactions between the polypeptide and one or more reagents described herein. In some embodiments, the present disclosure provides an approach for polypeptide structural analysis based on kinetic information obtained from single-molecule binding interactions between the polypeptide and one or more amino acid recognition factors described herein.

[0018] Figure 1 shows an example of a dynamic peptide sequencing reaction in which individual on-off binding events result in a signal pulse in the signal output. As shown on the left, a protein sample can be fragmented into peptides, which are immobilized in the reaction chambers of an array, and the immobilized peptides are exposed to one or more amino acid recognition factors and one or more cleavage reagents (e.g., aminopeptidases). As shown on the right, the amino acid recognition factors reversibly bind to the termini of the peptides, and a pulse in the signal output is generated while the recognition factors bind to the peptides.

[0019] Because on-off binding of recognition factors generally occurs at a faster rate than amino acid cleavage, the binding event preceding each cleavage event produces a series of changes in the signal (e.g., a signal pulse) that can be used to determine structural information about amino acids at or near the termini of the peptide. Compositions and methods for performing dynamic polypeptide sequencing and analyzing the data obtained therefrom are more fully described in International Publication Nos. WO 2020102741(A1), filed November 15, 2019, and WO 2021236983(A2), filed May 20, 2021, each of which is incorporated by reference in its entirety.

[0020] In some aspects, the present disclosure provides amino acid recognition factors with improved binding properties that allow for more structural information to be obtained from a polypeptide based on the on-off binding kinetics between the recognition factor and the polypeptide. In some embodiments, the amino acid recognition factor may comprise an amino acid binding protein having an altered binding pocket with one or more modifications compared to a homologous protein. In some embodiments, the altered binding pocket increases the number of interactions (e.g., hydrogen bond interactions, van der Waals interactions) formed between the binding pocket and the amino acid ligand compared to the unaltered binding pocket of the homologous protein. In some embodiments, the altered binding pocket increases the number of types of amino acid ligands that can be detectably bound compared to the unaltered binding pocket of the homologous protein. In some embodiments, the altered binding pocket increases the binding kinetics (e.g., K) for one or more types of amino acid ligands. D , k off , k on ), which advantageously increases the amount of, or confidence in, structural information that can be obtained from polypeptide analysis as described herein.

[0021] 1. Amino acid recognition factor In some embodiments, the present disclosure provides an amino acid recognition factor comprising an amino acid binding protein having an amino acid sequence selected from Table 1. Table 1 herein provides a list of exemplary sequences of amino acid binding proteins. It is understood that these sequences and other examples described herein are non-limiting and that an amino acid recognition factor in accordance with the present disclosure can include any homolog, variant, or fragment thereof that minimally contains the domain or subdomain involved in amino acid recognition.

[0022] In some embodiments, the disclosure provides amino acid binding proteins having an amino acid sequence that is at least 80% identical to an amino acid sequence selected from Table 1. In some embodiments, the amino acid binding protein has at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or more amino acid sequence identity to an amino acid sequence selected from Table 1. In some embodiments, the amino acid binding protein has 25-50%, 50-60%, 60-70%, 70-80%, 80-90%, 90-95%, 95-99%, 40-100%, 50-100%, 60-100%, 70-100%, 80-100%, 90-100%, or 95-100% amino acid sequence identity to an amino acid sequence selected from Table 1.

[0023] For purposes of comparing two or more amino acid sequences, the percentage of "sequence identity" (also referred to herein as "amino acid identity") between a first amino acid sequence and a second amino acid sequence can be calculated by dividing the number of amino acid residues in the first amino acid sequence that are identical to amino acid residues at corresponding positions in the second amino acid sequence by the total number of amino acid residues in the first amino acid sequence and multiplying by 0100, where each deletion, insertion, substitution, or addition of an amino acid residue in the second amino acid sequence compared to the first amino acid sequence is considered to be a difference at a single amino acid residue (position). Alternatively, the degree of sequence identity between two amino acid sequences can be calculated using known computer algorithms (e.g., by the local homology algorithm of Smith and Waterman (1970), Adv. Appl. Math. 2:482c, by the homology alignment algorithm of Needleman and Wunsch, J. Mol. Biol., (1970), 48:443, by the similarity search method of Pearson and Lipman, Proc. Natl. Acad. Sci. USA, (1998), 85:2444, or by computerized implementations of algorithms available as Blast, Clustal Omega, or other sequence alignment algorithms), for example, using standard settings. Typically, for purposes of determining the percentage of "sequence identity" between two amino acid sequences according to the calculation method outlined above, the amino acid sequence with the largest number of amino acid residues is designated the "first" amino acid sequence, and the other amino acid sequence is designated the "second" amino acid sequence.

[0024] Additionally or alternatively, two or more sequences may be assessed for identity between them. The term "identical" or percent "identity" in the context of two or more nucleotide or amino acid sequences refers to two or more sequences or subsequences that are the same. Two sequences are "substantially identical" if they have a specified percentage of amino acid residues or nucleotides that are the same over a specified region or entire sequence (e.g., at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9% identical) when compared and aligned for maximum correspondence over a comparison window or designated region, as measured using one of the sequence comparison algorithms described above or by manual alignment and visual inspection. Optionally, identity exists over a region that is at least about 25, 50, 75, or 100 amino acids in length, or over a region that is 100-150, 150-200, 100-200, or 200 or more amino acids in length.

[0025] Additionally or alternatively, two or more sequences may be assessed for alignment between the sequences. The term "alignment" or percent "alignment" in the context of two or more nucleotide or amino acid sequences refers to two or more sequences or subsequences that are the same. Two sequences are "substantially aligned" if they have a specified percentage of amino acid residues or nucleotides that are the same over a specified region or entire sequence (e.g., at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9% identical) when compared and aligned for maximum correspondence over a comparison window or designated region, as determined using one of the sequence comparison algorithms described above or by manual alignment and visual inspection. Optionally, the alignment is over a region that is at least about 25, 50, 75, or 100 amino acids in length, or over a region that is 100-150, 150-200, 100-200, or 200 or more amino acids in length.

[0026] In some embodiments, the amino acid recognition factors of the present disclosure comprise modified amino acid binding proteins that comprise one or more amino acid deletions, additions, or mutations relative to the sequences shown in Table 1. In some embodiments, the modified amino acid binding proteins comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more amino acids (which may or may not be consecutive amino acids) that have been deleted, added, or mutated relative to the sequences shown in Table 1.

[0027] A.ClpS-homologous recognition factor In some embodiments, the amino acid recognition factors of the present disclosure bind to amino acid ligands (e.g., polypeptides) comprising an N-terminal amino acid selected from leucine, isoleucine, valine, methionine, alanine, or modified variants thereof (e.g., post-translationally modified variants thereof, oxidized variants thereof). In some embodiments, the amino acid recognition factor comprises an amino acid binding protein derived from a ClpS protein, such as the Planctomycetia bacterium ClpS protein. For example, in some embodiments, the amino acid binding protein is an engineered variant comprising one or more modifications relative to SEQ ID NO: 1 described herein.

[0028] In some embodiments, the amino acid binding protein has a dissociation constant (K) of less than 2,000 nM, less than 1,500 nM, less than 1,000 nM, less than 750 nM, less than 500 nM, less than 250 nM, less than 150 nM, less than 100 nM, less than 80 nM, 10 to 2,000 nM, 25 to 1,000 nM, 50 to 500 nM, 10 to 150 nM, 25 to 75 nM, or 50 to 60 nM. DIn some embodiments, the amino acid binding protein binds to the N-terminal leucine at a K of less than 2,000 nM, less than 1,500 nM, less than 1,000 nM, less than 750 nM, less than 500 nM, less than 250 nM, less than 150 nM, less than 100 nM, less than 80 nM, 10 to 2,000 nM, 25 to 1,000 nM, 50 to 500 nM, 10 to 150 nM, 30 to 80 nM, or 60 to 75 nM. D In some embodiments, the amino acid binding protein binds to the N-terminal isoleucine at a K of less than 2,000 nM, less than 1,500 nM, less than 1,000 nM, less than 750 nM, less than 500 nM, less than 300 nM, less than 250 nM, less than 200 nM, between 10 and 2,000 nM, between 25 and 1,000 nM, between 50 and 500 nM, between 50 and 300 nM, or between 100 and 200 nM. D It binds to the N-terminal valine.

[0029] In some embodiments, the amino acid binding protein binds to one or more types of N-terminal amino acids (e.g., leucine, isoleucine, valine, methionine, and / or alanine), and each type of binding interaction is at least 0.1 s -1 Dissociation rate (k off In some embodiments, the dissociation rate is about 0.1 s -1 ~about 1,000s -1 (For example, about 0.5 seconds -1 ~about 500s -1 , about 0.1 seconds -1 ~approx. 100s -1 , about 1 s -1 ~approx. 100s -1 , or about 0.5 seconds -1 ~about 50s -1 In some embodiments, the dissociation rate is about 0.5 s -1 ~approx. 20 seconds -1 In some embodiments, the dissociation rate is about 2 s -1 ~approx. 20 seconds -1 In some embodiments, the dissociation rate is about 0.5 s -1 ~about 2s -1 is.

[0030] In some embodiments, the present disclosure provides PS635-645, PS731-732, PS759-766, PS769, PS795-870, PS896-912, PS918-1043, PS1048-1100, PS1124-1137, PS1141-1161, PS1175-1199, PS1203-1217, PS1222-1245, PS1277-1305, PS1321-1350, and PS1425-1448 (SEQ ID NOs: 22-27, 87-88, 115-122, 125, 151-226, 2 The present invention provides a recombinant or synthetic amino acid binding protein having an amino acid sequence that is at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or 100% identical to a sequence selected from any one of the following: In some embodiments, the amino acid sequences are selected from the group consisting of PS635-645, PS731-732, PS759-766, PS769, PS795-870, PS896-912, PS918-1043, PS1048-1100, PS1124-1137, PS1141-1161, PS1175-1199, PS1203-1217, PS1222-1245, PS1277-1305, PS1321-1350, and PS1425-1448 (SEQ ID NOs: 22-27, 87-88, 115-122, 125, 151 to 226, 249 to 265, 271 to 390, 395 to 446, 470 to 483, 487 to 507, 521 to 545, 549 to 563, 568 to 591, 622 to 650, 664 to 693, and 768 to 791).In some embodiments, the amino acid sequence is at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or 100% identical to PS961 (SEQ ID NO: 314). In some embodiments, the amino acid sequence is about 40% to about 100% (e.g., 80-98%, 80-95%, 80-90%, 85-95%, 90-98%, 50-100%, 60-100%, 70-100%, 80-100%, 90-100%, or 95-100%) identical to PS961 (SEQ ID NO: 314).

[0031] In some embodiments, the disclosure provides recombinant or synthetic amino acid binding proteins having an amino acid sequence that is at least 80% identical (e.g., at least 85%, at least 90%, at least 95%, at least 98%, 80-98%, 80-95%, 80-90%, 85-95%, or 90-98% identical) to SEQ ID NO:1, which amino acid sequence contains amino acid substitutions at one or more positions corresponding to E22, R31, L39, N41, D42, D43, D44, H45, T46, Y47, V50, Q55, P62, E63, L68, A69, V72, D73, Q75, Y100, and M111 of SEQ ID NO:1.

[0032] In some embodiments, the amino acid sequence comprises an amino acid substitution at a position corresponding to N41 and one or more positions corresponding to E22, R31, L39, D42, H45, V50, Q55, P62, E63, L68, V72, Q75, Y100, and M111. In some embodiments, the amino acid sequence comprises an amino acid substitution at a position corresponding to N41 and one or more positions corresponding to Q55, E63, L68, V72, and Y100. In some embodiments, the amino acid sequence comprises an amino acid substitution selected from E22V, R31H, L39M, N41D, D42L, D42P, H45C, H45F, V50A, V50F, V50Y, Q55H, Q55R, P62R, E63A, E63G, E63K, E63S, L68M, V72M, Q75L, Y100R, M111A, and M111S. In some embodiments, the amino acid substitution is selected from N41D, Q55R, E63S, L68M, V72M, and Y100R.

[0033] In some aspects, the present disclosure provides a recombinant or synthetic amino acid binding protein comprising the structure of formula (I) or a structural equivalent thereof:

[0034] [ka]

[0035] wherein: β1, β2, and β3 each is a β strand; α1, α2, and α3 each is an α helix; each instance of "-" is a loop; and at least a portion of each of α1, α2, the loop between β1 and α1, and the loop between α3 and β3 form a binding pocket for an amino acid ligand.

[0036] In some embodiments, the binding pocket comprises one or more of the following: (i) approximately 170 Å 3 volume, (ii) -3.0RTec -1(iii) negatively charged side chains on at least 35% of the amino acids forming the binding pocket; (iv) a plurality of hydrogen bond acceptors configured to form one or more hydrogen bonds (e.g., at least two, at least three, at least four, or at least five hydrogen bonds) in the presence of an amino acid ligand; and (v) a plurality of van der Waals contact positions configured to form van der Waals interactions in the presence of an amino acid ligand. In some embodiments, the binding pocket comprises one, two, three, or four of (i), (ii), (iii), (iv), and (v). In some embodiments, the binding pocket comprises (i), (ii), (iii), (iv), and (v).

[0037] In some embodiments, the binding pocket is approximately 170 Å 3 In some embodiments, the volume of the binding pocket is at least 150 Å 3 , at least 160 Å 3 , at least 170 Å 3 , at least 180 Å 3 , or at least 190 Å 3 In some embodiments, the volume of the binding pocket is 190 Å 3 Below, 180Å 3 Below, 170Å 3 Below, 160Å 3 or less, or 150Å 3 In some embodiments, the volume of the binding pocket is 150 Å or less. 3 ~170Å 3 , 150 Å 3 ~180Å 3 , 150 Å 3 ~190Å 3 , 160Å 3 ~180Å 3 , 160Å 3 ~190Å 3 , 170Å 3 ~190Å 3 , or 180 Å 3 ~190Å 3The volume of the binding pocket is in the range of 0.01 to 0.01. Methods for determining the volume of a binding pocket are known in the art and will be apparent to those skilled in the art in light of this disclosure. For example, in some embodiments, the volume of the binding pocket is determined using software configured to measure the geometric and topological properties of a protein. In some embodiments, the software can scan the protein surface using a specified probe radius to measure the volume of any cavity that directly overlaps the binding site or indirectly overlaps the binding site via adjacent cavities in van der Waals contact with each other. A non-limiting example of a suitable probe radius is the solvent probe radius (e.g., approximately 1.4 Å). A non-limiting example of suitable software is the Computed Atlas of Surface Topography of proteins (CASTp). See, for example, W. Tian et al., CASTp3.0: Computed atlas of surface topography of proteins., Nucleic Acids Res., 46, W363-W367 (2018), the relevant contents of which are incorporated herein by reference.

[0038] In some embodiments, the binding pocket has a binding affinity of -3.0 RTe c -1 In some embodiments, the electrostatic potential of the binding pocket is at least -4RTe c -1 , at least -3RTe c -1 , or at least -2 Te c -1 In some embodiments, the electrostatic potential of the binding pocket is -2RTe c -1 Below, -3RTe c -1 or less, or -4Te c -1 In some embodiments, the electrostatic potential of the binding pocket is -2RTe c-1 ~-3RTe c -1 , -2RTe c -1 ~-4RTe c -1 , or -3RTe c -1 ~-4RTe c -1 The electrostatic potential of the binding pocket is in the range of 0.01 to 0.01. Methods for determining the electrostatic potential of the binding pocket are known in the art and will be apparent to those of skill in the art in light of the present disclosure. For example, in some embodiments, the adaptive Poisson-Boltzmann solver (APBS) tool in PyMOL (PyMOL Molecular Graphics System, Version 2.0, Schrodinger, LLC) may be used (e.g., with default parameters). This tool uses pdb2pqr with the AMBER force field to calculate the electrostatic surface potential of the binding pocket and can assign protonation states. See, for example, TJ Dolinsky et al., PDB2PQR: An automated pipeline for the setup of Poisson-Boltzmann electrostatics calculations, Nucleic Acids Res., 32:W665-7 (2004); MGLerner et al., APBS plugin for PyMOL-Version 2.4 (University of Michigan, Ann Arbor, MI, 2006); JW Ponder et al., Force fields for protein simulations, Adv. Protein Chem., 66:27-85 (2003). In some cases, this solvent-accessible surface area (SASA) may be considered accessible by peptide ligands.

[0039] In some embodiments, the binding pocket comprises a plurality of hydrogen bond acceptors configured to form one or more hydrogen bonds in the presence of an amino acid ligand. In some embodiments, the binding pocket forms at least two (e.g., at least three, at least four, at least five, 2-10, 4-10, 5-15, 5-10) hydrogen bonds with the amino acid ligand. Methods for determining hydrogen bonding interactions between a binding pocket and a ligand are known in the art and will be apparent to those of skill in the art in light of the present disclosure. For example, in some embodiments, the hydrogen bonding interactions are determined by computer modeling as described in the Examples herein (e.g., using atomic coordinates for protein-ligand structural data obtained from X-ray crystallography or by computer modeling to predict protein-ligand three-dimensional structures).

[0040] In some embodiments, the multiple hydrogen bond acceptors comprise one or more atoms of the side chains of amino acid residues within the binding pocket. For example, in some embodiments, the binding pocket comprises at least one negatively charged amino acid side chain (e.g., aspartic acid, glutamic acid) that forms a bifurcated hydrogen bond with the amino acid ligand (e.g., the N-terminal amino acid of the polypeptide). In some embodiments, at least four hydrogen bonds are formed between the amino acid ligand (e.g., the N-terminal amino acid of the polypeptide) and the amino acid side chain within the binding pocket. In some embodiments, the multiple hydrogen bond acceptors comprise one or more atoms of the polypeptide backbone (e.g., backbone carbonyls) within the binding pocket.

[0041] In some embodiments, the binding pocket forms one or more hydrogen bonds with the side chain of the amino acid ligand. For example, in some embodiments, the binding pocket forms one or more hydrogen bonds with the side chain of the terminal amino acid of the amino acid ligand (e.g., a polypeptide). In some embodiments, the binding pocket forms one or more hydrogen bonds with the polypeptide backbone of the amino acid ligand (e.g., a polypeptide). In some embodiments, the binding pocket forms one or more hydrogen bonds with the terminal amino acid in the polypeptide and one or more amino acids adjacent to the terminal amino acid (e.g., the amino acids at positions 1 and 2, 3, 4, and / or 5 relative to the polypeptide terminus).

[0042] In some embodiments, the binding pocket comprises a plurality of van der Waals contact positions configured to form van der Waals interactions in the presence of an amino acid ligand. Methods for determining van der Waals interactions between a binding pocket and a ligand are known in the art and will be apparent to those of skill in the art in light of the present disclosure. For example, in some embodiments, the van der Waals interactions are determined by computer modeling as described in the Examples herein (e.g., using atomic coordinates for protein-ligand structural data derived from X-ray crystallography, or by computer modeling to predict protein-ligand three-dimensional structures).

[0043] In some embodiments, the van der Waals contact position comprises a plurality of atoms (e.g., 2-30, 5-25, 10-20, 2-10, 5-10) configured to form hydrophobic interactions with the amino acid ligand. In some embodiments, one or more atoms of the plurality of atoms are non-polar atoms. In some embodiments, the van der Waals contact position is formed by a methionine side chain within the binding pocket.

[0044] In some embodiments, the amino acid ligand is a polypeptide comprising at least three amino acids. In some embodiments, the amino acid ligand comprises the N-terminal amino acid of a polypeptide. In some embodiments, the N-terminal amino acid is selected from leucine, isoleucine, valine, methionine, and alanine. In some embodiments, the amino acid binding protein is at least 50 amino acids in length, at least 75 amino acids in length, at least 100 amino acids in length, 50-250 amino acids in length, 50-150 amino acids in length, or 100-200 amino acids in length.

[0045] In some embodiments, the loop between β1 and α1 comprises three or more negatively charged amino acids. In some embodiments, the loop between β1 and α1 comprises four or more negatively charged amino acids. In some embodiments, at least two negatively charged amino acids in the loop between β1 and α1 form hydrogen bonds with an amino acid ligand. In some embodiments, at least one negatively charged amino acid in the loop between β1 and α1 forms a bifurcated hydrogen bond with an amino acid ligand. In some embodiments, the negatively charged amino acid is selected from aspartic acid and glutamic acid.

[0046] In some embodiments, β1-α1 comprises an amino acid sequence that is at least 80% identical to the sequence of amino acids 35-58 of SEQ ID NO: 1. In some embodiments, the binding pocket is formed by amino acids at one or more positions corresponding to amino acids 41-47 and 50 of SEQ ID NO: 1. In some embodiments, the amino acid sequence comprises amino acid substitutions at one or more positions corresponding to L39, N41, D42, H45, V50, and Q55 of SEQ ID NO: 1. In some embodiments, at least one amino acid substitution is at a position corresponding to N41 of SEQ ID NO: 1. In some embodiments, the amino acid substitution is selected from N41D and Q55R.

[0047] In some embodiments, α2 comprises an amino acid sequence that is at least 80% identical to the sequence of amino acids 62-73 of SEQ ID NO: 1. In some embodiments, the binding pocket is formed by amino acids at one or more positions corresponding to amino acids 69, 72, and 73 of SEQ ID NO: 1. In some embodiments, the amino acid sequence comprises amino acid substitutions at one or more positions corresponding to P62, E63, L68, and V72 of SEQ ID NO: 1. In some embodiments, at least one amino acid substitution is at a position corresponding to V72 of SEQ ID NO: 1. In some embodiments, the amino acid substitution is selected from E63S, L68M, and V72M.

[0048] In some embodiments, the loop between α3 and β3 comprises an amino acid sequence that is at least 80% identical to the sequence of amino acids 99-112 of SEQ ID NO: 1. In some embodiments, the binding pocket is formed by amino acids at positions corresponding to amino acid 111 of SEQ ID NO: 1. In some embodiments, the amino acid sequence comprises amino acid substitutions at one or more positions corresponding to Y100 and M111 of SEQ ID NO: 1. In some embodiments, the amino acid substitution is Y100R.

[0049] In some embodiments, a structural equivalent is a structure in which at least 80% of the secondary structural α carbon atoms are aligned with the structure of Formula (I) with a root-mean-square difference of 5.0 Å or less. In some embodiments, a structural equivalent is a structure in which at least 80% of the secondary structural α carbon atoms are aligned with the structure of Formula (I) with a root-mean-square difference of 4.0 Å or less, 3.0 Å or less, 2.0 Å or less, or 1.0 Å or less.

[0050] Methods for identifying structures equivalent to the structure of formula (I) are known in the art and will be apparent to one of skill in the art in light of this disclosure. For example, in some embodiments, structural equivalents to formula (I) are structures having a root mean square difference of 5.0 Å or less, where at least 80% of the secondary structural α-carbon atoms are identical to those of PS635-645, PS731-732, PS759-766, PS769, PS795-870, PS896-912, PS918-1043, PS1048-1100, PS1124-1137, PS1141-1161, PS1175-1199, PS1203-1217, PS1218-1220, PS1224-1225, PS1226-1227, PS1228-1229, PS1230-1231, PS1232-1233, PS1234-1235, PS1236-1237, PS1238-1239, PS1240-1241, PS1242-1243, PS1244-1245, PS1246-1247, PS1248-1249, PS1250-1251, PS1252-1253, PS1254-1255, PS1256-1257, PS1258-1259, PS1260-1261, PS1262-1263, PS1264-1265, PS1266-1267, PS1268-1269, PS1270-1271, PS1272-1274, PS1276-1277, PS1278-127 and PS1425-1448 (SEQ ID NOs: 22-27, 87-88, 115-122, 125, 151-226, 249-265, 271-390, 395-446, 470-483, 487-507, 521-545, 549-563, 568-591, 622-650, 664-693, and 768-791). Comparison of protein structures was performed using PS635-645, PS731-732, PS759-766, PS769, PS795-870, PS896-912, PS918-1043, PS1048-1100, PS1124-1137, PS1141-1161, PS1175-1199, PS1203-1217, PS1222-1245, PS1277-1305, PS1321-1350, and PS1425-1448 (SEQ ID NOs: 22-27, 87-88, 115-122, 136-138, 144-146, 146-148, 148-149, 150-151, 152-153, 154-155, 156-157, 158-159, 160-161, 162-163, 164-165, 166-167, 168-169, 169-200, 169-201, 169-202, 169-203, 169-204, 169-205, 169-206, 170-171, 171-207, 172-173, 173-174, 174-175, 175-176, 176-177, 177-178, 178-179, 180-182, 181-183, 182-184, 183-185, 184-186, 25, 151-226, 249-265, 271-390, 395-446, 470-483, 487-507, 521-545, 549-563, 568-591, 622-650, 664-693, and 768-791), and comparing the three-dimensional structure of the protein to a candidate protein structure to determine whether the candidate protein is a structural equivalent.The three-dimensional protein structure can be determined, for example, using atomic coordinates derived from X-ray crystallography or by analyzing the structure of PS635-645, PS731-732, PS759-766, PS769, PS795-870, PS896-912, PS918-1043, PS1048-1100, PS1124-1137, PS1141-1161, PS1175-1199, PS1203-1217, PS1222-1245, PS1277-1305, PS1321-1350, and PS The three-dimensional structure of a protein having an amino acid sequence selected from any one of SEQ ID NOs: 1425-1448 (SEQ ID NOs: 22-27, 87-88, 115-122, 125, 151-226, 249-265, 271-390, 395-446, 470-483, 487-507, 521-545, 549-563, 568-591, 622-650, 664-693, and 768-791) can be determined by computer modeling to predict the three-dimensional structure of a protein having an amino acid sequence selected from any one of SEQ ID NOs: 22-27, 87-88, 115-122, 125, 151-226, 249-265, 271-390, 395-446, 470-483, 487-507, 521-545, 549-563, 568-591, 622-650, 664-693, and 768-791). Methods for comparing protein structures and determining root mean square difference values are known in the art (see, for example, Kufareva I, Abagyan R., Methods of protein structure comparison., Methods Mol Biol., 2012;857:231-57).

[0051] B.UBR-homologous recognition factor In some embodiments, the amino acid recognition factors of the present disclosure bind to amino acid ligands (e.g., polypeptides) comprising an N-terminal amino acid selected from arginine or modified variants thereof (e.g., post-translationally modified variants thereof, oxidized variants thereof). In some embodiments, the amino acid recognition factor comprises an amino acid binding protein derived from a UBR protein, such as the Kluyveromyces marxianus UBR protein. For example, in some embodiments, the amino acid binding protein is a modified variant comprising one or more modifications relative to SEQ ID NO: 2 described herein.

[0052] In some embodiments, the amino acid binding protein has a dissociation constant (K) of less than 2,000 nM, less than 1,500 nM, less than 1,000 nM, less than 750 nM, less than 500 nM, less than 400 nM, less than 200 nM, less than 100 nM, between 10 and 2,000 nM, between 25 and 1,000 nM, between 50 and 500 nM, between 10 and 400 nM, between 25 and 75 nM, or between 50 and 80 nM. D ) binds to the N-terminal arginine.

[0053] In some embodiments, the amino acid binding protein binds to more than one type of N-terminal amino acid (e.g., arginine or an engineered variant thereof), and each type of binding interaction is at least 0.1 s -1 Dissociation rate (k off In some embodiments, the dissociation rate is about 0.1 s -1 ~about 1,000s -1 (For example, about 0.5 seconds -1 ~about 500s -1 , about 0.1 seconds -1 ~approx. 100s -1 , about 1 s -1 ~approx. 100s -1 , or about 0.5 seconds -1 ~about 50s -1 In some embodiments, the dissociation rate is about 0.5 s -1 ~approx. 20 seconds -1 In some embodiments, the dissociation rate is about 2 s -1 ~approx. 20 seconds -1 In some embodiments, the dissociation rate is about 0.5 s -1 ~about 2s -1 is.

[0054] In some aspects, the disclosure provides recombinant or synthetic amino acid binding proteins having an amino acid sequence that is at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or 100% identical to a sequence selected from any one of PS1101-1122, PS1218-1221, and PS1351-1398 (SEQ ID NOs: 447-468, 564-567, and 694-741). In some embodiments, the amino acid sequence is about 40% to about 100% (e.g., 80 to 98%, 80 to 95%, 80 to 90%, 85 to 95%, 90 to 98%, 50 to 100%, 60 to 100%, 70 to 100%, 80 to 100%, 90 to 100%, or 95 to 100%) identical to a sequence selected from any one of PS1101 to 1122, PS1218 to 1221, and PS1351 to 1398 (SEQ ID NOs: 447 to 468, 564 to 567, and 694 to 741). In some embodiments, the amino acid sequence is at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or 100% identical to PS1122 (SEQ ID NO: 468) or PS1381 (SEQ ID NO: 724). In some embodiments, the amino acid sequence is about 40% to about 100% (e.g., 80-98%, 80-95%, 80-90%, 85-95%, 90-98%, 50-100%, 60-100%, 70-100%, 80-100%, 90-100%, or 95-100%) identical to PS1122 (SEQ ID NO: 468) or PS1381 (SEQ ID NO: 724).

[0055] In some embodiments, the disclosure provides recombinant or synthetic amino acid binding proteins having an amino acid sequence that is at least 80% identical (e.g., at least 85%, at least 90%, at least 95%, at least 98%, 80-98%, 80-95%, 80-90%, 85-95%, or 90-98% identical) to SEQ ID NO:2, which amino acid sequence contains amino acid substitutions at one or more positions corresponding to G19, K26, S29, F30, D31, D32, T33, C34, V35, T47, G48, T53, T54, T57, E58, F59, N61, I63, D65, D68, E70, A71, H74, and T75 of SEQ ID NO:2.

[0056] In some embodiments, the amino acid sequence comprises amino acid substitutions at positions corresponding to I63 and E70, and one or more positions corresponding to G19, K26, S29, D32, T47, G48, T53, T54, T57, E58, F59, N61, H74, and T75. In some embodiments, the amino acid sequence comprises amino acid substitutions at one or more positions corresponding to K26, D32, T47, I63, and E70. In some embodiments, the amino acid sequence comprises amino acid substitutions selected from G19R, K26R, S29Q, D32R, D32Y, T47K, T47L, T47R, G48R, G48Y, T53V, T54K, T57K, T57R, E58K, F59R, N61K, I63E, E70S, E70T, H74K, and T75E. In some embodiments, the amino acid substitution is selected from T47L, I63E, and E70T. In some embodiments, the amino acid substitution is selected from K26R and D32R.

[0057] In some embodiments, the present disclosure provides a recombinant or synthetic amino acid binding protein comprising the structure of formula (II) or a structural equivalent thereof:

[0058] [ka]

[0059] wherein β1 and β2 are each a β-strand; α1, α2, and α3 are each an α-helix; each instance of "-" is a loop; and at least a portion of each of α2, the loop between β1 and α1, and the loop between β2 and α2 form a binding pocket for an amino acid ligand.

[0060] In some embodiments, the binding pocket comprises one or more of the following: (i) approximately 200 Å 3 volume, (ii) -3.0RTec -1 (iii) a plurality of hydrogen bond acceptors configured to form one or more hydrogen bonds in the presence of an amino acid ligand; and (iv) a plurality of van der Waals contact locations configured to form van der Waals interactions in the presence of an amino acid ligand. In some embodiments, the binding pocket comprises one, two, or three of (i), (ii), (iii), and (iv). In some embodiments, the binding pocket comprises (i), (ii), (iii), and (iv).

[0061] In some embodiments, the binding pocket is approximately 200 Å 3 In some embodiments, the volume of the binding pocket is at least 180 Å 3 , at least 190 Å 3 , at least 200 Å 3 , at least 210 Å 3 , or at least 220 Å 3 In some embodiments, the volume of the binding pocket is 220 Å 3 Below, 210Å 3 Below, 200Å 3 Less than 190Å3v or 180Å 3 In some embodiments, the volume of the binding pocket is 180 Å or less. 3 ~200Å 3 , 180Å 3 ~210Å 3 , 180Å 3 ~220Å 3 , 190Å 3 ~210Å3 , 190Å 3 ~220Å 3 , 200Å3v~220Å 3 , or 210 Å 3 ~220Å 3 The volume of the binding pocket is in the range of 0.01 to 0.01. Methods for determining the volume of a binding pocket are known in the art and will be apparent to those skilled in the art in light of this disclosure. For example, in some embodiments, the volume of the binding pocket is determined using software configured to measure the geometric and topological properties of a protein. In some embodiments, the software can scan the protein surface using a specified probe radius to measure the volume of any cavity that directly overlaps the binding site or indirectly overlaps the binding site via adjacent cavities in van der Waals contact with each other. A non-limiting example of a suitable probe radius is the solvent probe radius (e.g., approximately 1.4 Å). A non-limiting example of suitable software is the Computed Atlas of Surface Topography of proteins (CASTp). See, for example, W. Tian et al., CASTp3.0: Computed atlas of surface topography of proteins., Nucleic Acids Res., 46, W363-W367 (2018), the relevant contents of which are incorporated herein by reference.

[0062] In some embodiments, the binding pocket has a binding affinity of -3.0 RTe c -1 In some embodiments, the electrostatic potential of the binding pocket is at least -4RTe c -1 , at least -3RTe c -1 , or at least -2 Te c -1 In some embodiments, the electrostatic potential of the binding pocket is -2RTe c -1 Below, -3RTec -1 or less, or -4Te c -1 In some embodiments, the electrostatic potential of the binding pocket is -2RTe c -1 ~-3RTe c -1 , -2RTe c -1 ~-4RTe c -1 , or -3RTe c -1 ~-4RTe c -1 The electrostatic potential of the binding pocket is in the range of 0.01 to 0.01. Methods for determining the electrostatic potential of the binding pocket are known in the art and will be apparent to those of skill in the art in light of the present disclosure. For example, in some embodiments, the adaptive Poisson-Boltzmann solver (APBS) tool in PyMOL (PyMOL Molecular Graphics System, Version 2.0, Schrodinger, LLC) may be used (e.g., with default parameters). This tool uses pdb2pqr with the AMBER force field to calculate the electrostatic surface potential of the binding pocket and can assign protonation states. See, for example, TJ Dolinsky et al., PDB2PQR: An automated pipeline for the setup of Poisson-Boltzmann ectrostatics calculations, Nucleic Acids Res., 32:W665-7 (2004); MGLerner et al., APBS plugin for PyMOL-Version 2.4 (University of Michigan, Ann Arbor, MI, 2006); JW Ponder et al., Force fields for protein simulations, Adv. Protein Chem., 66:27-85 (2003). In some cases, this solvent-accessible surface area (SASA) may be considered accessible by a peptide ligand.

[0063] In some embodiments, the binding pocket comprises a plurality of hydrogen bond acceptors configured to form one or more hydrogen bonds in the presence of an amino acid ligand. In some embodiments, the binding pocket forms at least two (e.g., at least three, at least four, at least five, 2-10, 4-10, 5-15, 5-10) hydrogen bonds with the amino acid ligand. Methods for determining hydrogen bonding interactions between a binding pocket and a ligand are known in the art and will be apparent to those of skill in the art in light of the present disclosure. For example, in some embodiments, the hydrogen bonding interactions are determined by computer modeling as described in the Examples herein (e.g., using atomic coordinates for protein-ligand structural data obtained from X-ray crystallography or by computer modeling to predict protein-ligand three-dimensional structures).

[0064] In some embodiments, the multiple hydrogen bond acceptors comprise one or more atoms of the side chains of amino acid residues within the binding pocket. For example, in some embodiments, the binding pocket comprises at least three negatively charged amino acid side chains, each of which (e.g., aspartic acid, glutamic acid) forms a hydrogen bond with the amino acid ligand. In some embodiments, at least one of the negatively charged amino acid side chains forms a hydrogen bond with the amino terminus of the amino acid ligand. In some embodiments, the binding pocket comprises at least one polar, uncharged amino acid side chain that forms a hydrogen bond with the amino acid ligand. In some embodiments, the multiple hydrogen bond acceptors comprise one or more atoms of the polypeptide backbone (e.g., backbone carbonyls) within the binding pocket.

[0065] In some embodiments, the binding pocket forms one or more hydrogen bonds with the side chain of the amino acid ligand. For example, in some embodiments, the binding pocket forms one or more hydrogen bonds with the side chain of the terminal amino acid of the amino acid ligand (e.g., a polypeptide). In some embodiments, the binding pocket forms one or more hydrogen bonds with the polypeptide backbone of the amino acid ligand (e.g., a polypeptide). In some embodiments, the binding pocket forms one or more hydrogen bonds with the terminal amino acid in the polypeptide and one or more amino acids adjacent to the terminal amino acid (e.g., the amino acids at positions 1 and 2, 3, 4, and / or 5 relative to the polypeptide terminus).

[0066] In some embodiments, the binding pocket comprises a plurality of van der Waals contact positions configured to form van der Waals interactions in the presence of an amino acid ligand. Methods for determining van der Waals interactions between a binding pocket and a ligand are known in the art and will be apparent to those of skill in the art in light of the present disclosure. For example, in some embodiments, the van der Waals interactions are determined by computer modeling as described in the Examples herein (e.g., using atomic coordinates for protein-ligand structural data derived from X-ray crystallography, or by computer modeling to predict protein-ligand three-dimensional structures).

[0067] In some embodiments, the van der Waals contact sites include a plurality of atoms (e.g., 2-30, 5-25, 10-20, 2-10, 5-10) configured to form hydrophobic interactions with the amino acid ligand. In some embodiments, one or more atoms of the plurality of atoms are non-polar atoms.

[0068] In some embodiments, the amino acid ligand is a polypeptide comprising at least three amino acids. In some embodiments, the amino acid ligand comprises the N-terminal amino acid of a polypeptide. In some embodiments, the N-terminal amino acid is arginine. In some embodiments, the amino acid binding protein is at least 50 amino acids in length, at least 75 amino acids in length, at least 100 amino acids in length, 50-250 amino acids in length, 50-150 amino acids in length, or 100-200 amino acids in length.

[0069] In some embodiments, the loop between β2 and α2 comprises three or more negatively charged amino acids. In some embodiments, the loop between β2 and α2 comprises four or more negatively charged amino acids. In some embodiments, at least three negatively charged amino acids in the loop between β2 and α2 form hydrogen bonds with an amino acid ligand. In some embodiments, at least one negatively charged amino acid in the loop between β2 and α2 forms a hydrogen bond with the amino terminus of an amino acid ligand. In some embodiments, the negatively charged amino acid is selected from aspartic acid and glutamic acid.

[0070] In some embodiments, at least one amino acid of α2 forms a hydrogen bond with an amino acid ligand. In some embodiments, α2 comprises one or more polar, uncharged amino acids. In some embodiments, at least one polar, uncharged amino acid of α2 forms a hydrogen bond with a side chain of an amino acid ligand.

[0071] In some embodiments, the loop between β1 and α1 comprises an amino acid sequence that is at least 80% identical to the sequence of amino acids 27-42 of SEQ ID NO: 2. In some embodiments, the binding pocket is formed by amino acids at one or more positions corresponding to amino acids 31, 32, and 34-36 of SEQ ID NO: 2.

[0072] In some embodiments, the loop between α1 and β2 comprises an amino acid sequence that is at least 50% identical to the sequence of amino acids 47-50 of SEQ ID NO: 2. In some embodiments, the amino acid sequence comprises an amino acid substitution at a position corresponding to T47 of SEQ ID NO: 2. In some embodiments, the amino acid substitution is T47L.

[0073] In some embodiments, β2-α2 comprises an amino acid sequence that is at least 80% identical to the sequence of amino acids 51-71 of SEQ ID NO: 2. In some embodiments, the binding pocket is formed by amino acids at one or more positions corresponding to amino acids 63, 65, 68, 70, and 71 of SEQ ID NO: 2. In some embodiments, the amino acid sequence comprises amino acid substitutions at one or more positions corresponding to I63 and E70 of SEQ ID NO: 2. In some embodiments, the amino acid substitutions are selected from I63E and E70T.

[0074] In some embodiments, a structural equivalent is a structure in which at least 80% of the secondary structural α carbon atoms align with the structure of Formula (II) with a root mean square difference of 5.0 Å or less. In some embodiments, a structural equivalent is a structure in which at least 80% of the secondary structural α carbon atoms align with the structure of Formula (II) with a root mean square difference of 4.0 Å or less, 3.0 Å or less, 2.0 Å or less, or 1.0 Å or less.

[0075] Methods for identifying structural equivalents to the structure of Formula (II) are known in the art and will be apparent to those of skill in the art in light of this disclosure. For example, in some embodiments, a structural equivalent to Formula (II) is a structure having a root-mean-square difference of 5.0 Å or less, wherein at least 80% of the secondary structural α carbon atoms align with a three-dimensional protein structure having an amino acid sequence selected from any one of PS1101-1122, PS1218-1221, and PS1351-1398 (SEQ ID NOs:447-468, 564-567, and 694-741) (e.g., PS1122 (SEQ ID NO:468), PS1381 (SEQ ID NO:724)). Comparison of protein structures can be performed by determining the three-dimensional structure of a protein having an amino acid sequence selected from any one of PS1101-1122, PS1218-1221, and PS1351-1398 (SEQ ID NOs: 447-468, 564-567, and 694-741) (e.g., PS1122 (SEQ ID NO: 468), PS1381 (SEQ ID NO: 724)), and comparing the three-dimensional structure of the protein with a candidate protein structure to determine whether the candidate protein is a structural equivalent. The three-dimensional protein structure can be determined, for example, using atomic coordinates derived from X-ray crystallography or via computer modeling to predict the three-dimensional structure of a protein having an amino acid sequence selected from any one of PS1101-1122, PS1218-1221, and PS1351-1398 (SEQ ID NOs: 447-468, 564-567, and 694-741) (e.g., PS1122 (SEQ ID NO: 468), PS1381 (SEQ ID NO: 724)). Methods for comparing protein structures and determining root mean square difference values are known in the art (see, e.g., Kufareva I, Abagyan R., Methods of protein structure comparison., Methods Mol Biol., 2012;857:231-57).

[0076] C.Ntaq1-homologous recognition factor In some embodiments, the amino acid recognition factors of the present disclosure bind to amino acid ligands (e.g., polypeptides) comprising an N-terminal amino acid selected from glutamine, asparagine, glutamic acid, aspartic acid, cysteine-S-acetamide, or modified variants thereof (e.g., post-translationally modified variants thereof, oxidized variants thereof). In some embodiments, the amino acid recognition factor comprises an amino acid binding protein derived from an Ntaq1 protein, such as the Asian arowana fish Scleropages formosus Ntaq1 protein. For example, in some embodiments, the amino acid binding protein is a modified variant comprising one or more modifications relative to SEQ ID NO: 3 described herein.

[0077] In some embodiments, the amino acid binding protein has a dissociation constant (K) of less than 2,000 nM, less than 1,500 nM, less than 1,000 nM, less than 750 nM, less than 500 nM, less than 250 nM, less than 150 nM, less than 100 nM, less than 50 nM, between 10 and 2,000 nM, between 25 and 1,000 nM, between 50 and 500 nM, between 10 and 100 nM, between 25 and 250 nM, or between 50 and 150 nM. D ) binds to the N-terminal glutamine.

[0078] In some embodiments, the amino acid binding protein binds to one or more types of N-terminal amino acids (e.g., glutamine, asparagine, glutamic acid, aspartic acid, cysteine-S-acetamide, or engineered variants thereof), and each type of binding interaction is at least 0.1 s -1 Dissociation rate (k off In some embodiments, the dissociation rate is about 0.1 s -1 ~about 1,000s -1 (For example, about 0.5 seconds -1 ~about 500s -1 , about 0.1 seconds -1 ~approx. 100s -1 , about 1 s -1 ~approx. 100s -1 , or about 0.5 seconds -1 ~about 50s -1In some embodiments, the dissociation rate is about 0.5 s -1 ~approx. 20 seconds -1 In some embodiments, the dissociation rate is about 2 s -1 ~approx. 20 seconds -1 In some embodiments, the dissociation rate is about 0.5 s -1 ~about 2s -1 is.

[0079] In some aspects, the disclosure provides recombinant or synthetic amino acid binding proteins having an amino acid sequence that is at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or 100% identical to a sequence selected from any one of PS1258-1260, PS1315-1318, PS1457-1478, PS1480-1499, PS1633-1656, PS1737-1758, PS1821-1898, PS2014-2057, and PS2116-2137 (SEQ ID NOs: 604-606, 660-663, 792-833, and 836-1025). In some embodiments, the amino acid sequences are selected from the group consisting of PS1258-1260, PS1315-1318, PS1457-1478, PS1480-1499, PS1633-1656, PS1737-1758, PS1821-1898, PS2014-2057, and PS2116-2137 (SEQ ID NOs: 604-606, 660 In some embodiments, the amino acid sequence is at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or 100% identical to PS1259 (SEQ ID NO: 605). In some embodiments, the amino acid sequence is about 40% to about 100% (e.g., 80-98%, 80-95%, 80-90%, 85-95%, 90-98%, 50-100%, 60-100%, 70-100%, 80-100%, 90-100%, or 95-100%) identical to PS1259 (SEQ ID NO: 605).In some embodiments, the amino acid sequence is at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or 100% identical to PS2132 (SEQ ID NO: 1020). In some embodiments, the amino acid sequence is about 40% to about 100% (e.g., 80-98%, 80-95%, 80-90%, 85-95%, 90-98%, 50-100%, 60-100%, 70-100%, 80-100%, 90-100%, or 95-100%) identical to PS2132 (SEQ ID NO: 1020).

[0080] In some embodiments, the disclosure provides recombinant or synthetic amino acid binding proteins having an amino acid sequence that is at least 80% identical (e.g., at least 85%, at least 90%, at least 95%, at least 98%, 80-98%, 80-95%, 80-90%, 85-95%, or 90-98% identical) to SEQ ID NO:3, which amino acid sequence contains amino acid substitutions at one or more positions corresponding to S22, C23, Y24, C25, E26, S39, W75, D76, Y77, H78, C85, N120, H145, and M146 of SEQ ID NO:3.

[0081] In some embodiments, the amino acid sequence comprises amino acid substitutions at positions corresponding to C25 and H78. In some embodiments, the amino acid sequence comprises amino acid substitutions at positions corresponding to S22, C25, H78, C85, and N120. In some embodiments, the amino acid substitution is selected from S22E, C25S, H78Q, H78K, C85T, N120R, and M146E. In some embodiments, the amino acid substitution is selected from C25S, H78Q, and M146E. In some embodiments, the amino acid substitution is selected from C25S and H78Q. In some embodiments, the amino acid sequence comprises an amino acid substitution at a position corresponding to H78, where the amino acid substitution is H78Q. In some embodiments, the amino acid sequence comprises an amino acid substitution at a position corresponding to H78, where the amino acid substitution is H78K.

[0082] In some embodiments, the present disclosure provides a recombinant or synthetic amino acid binding protein comprising the structure of formula (III) or a structural equivalent thereof:

[0083] [ka]

[0084] wherein each of α1, α2, α3, α4, α5, and α6 is an α-helix; each of β1, β2, β3, β4, β5, and β6 is a β-strand; each instance of "-" is a loop; and at least a portion of each of α2, β3, β4, α5, the loop between α1 and α2, and the loop between β3 and β4 form a binding pocket for an amino acid ligand.

[0085] In some embodiments, the binding pocket comprises one or more of the following: (i) approximately 160 Å 3 volume, (ii) -2.0RTec -1 (iii) a plurality of hydrogen bond acceptors or donors configured to form one or more hydrogen bonds in the presence of an amino acid ligand, (iv) a plurality of van der Waals contact locations configured to form van der Waals interactions in the presence of an amino acid ligand, and (v) at least one negatively charged amino acid and at least one positively charged amino acid. In some embodiments, the binding pocket comprises one, two, three, or four of (i), (ii), (iii), (iv), and (v). In some embodiments, the binding pocket comprises (i), (ii), (iii), (iv), and (v).

[0086] In some embodiments, the binding pocket is approximately 160 Å 3 In some embodiments, the volume of the binding pocket is at least 140 Å 3 , at least 150 Å 3 , at least 160 Å 3 , at least 170 Å 3, or at least 180 Å 3 In some embodiments, the volume of the binding pocket is 180 Å 3 Below, 170Å 3 Below, 160Å 3 Below, 150Å 3 or less, or 140 Å 3 In some embodiments, the volume of the binding pocket is 140 Å or less. 3 ~160Å 3 , 140Å 3 ~170Å 3 , 140Å 3 ~180Å 3 , 150 Å 3 ~170Å 3 , 150 Å 3 ~180Å 3 , 160Å 3 ~180Å 3 , or 170 Å 3 ~180Å 3 The volume of the binding pocket is in the range of 0.01 to 0.01. Methods for determining the volume of a binding pocket are known in the art and will be apparent to those skilled in the art in light of this disclosure. For example, in some embodiments, the volume of the binding pocket is determined using software configured to measure the geometric and topological properties of a protein. In some embodiments, the software can scan the protein surface using a specified probe radius to measure the volume of any cavity that directly overlaps the binding site or indirectly overlaps the binding site via adjacent cavities in van der Waals contact with each other. A non-limiting example of a suitable probe radius is the solvent probe radius (e.g., approximately 1.4 Å). A non-limiting example of suitable software is the Computed Atlas of Surface Topography of proteins (CASTp). See, for example, W. Tian et al., CASTp3.0: Computed atlas of surface topography of proteins., Nucleic Acids Res., 46, W363-W367 (2018), the relevant contents of which are incorporated herein by reference.

[0087] In some embodiments, the binding pocket is -2.0RTec -1 In some embodiments, the electrostatic potential of the binding pocket is at least -3RTec -1 , at least -2RTec -1 , or at least -1RTec -1 In some embodiments, the electrostatic potential of the binding pocket is -1 Below, -2RTec -1 Below, or -3RTec -1 In some embodiments, the electrostatic potential of the binding pocket is -1RTe c -1 ~-2RTe c -1 , -1Te c -1 ~-3Te c -1 , or -2RTe c -1 ~-3RTe c -1The electrostatic potential of the binding pocket is in the range of 0.01 to 0.01. Methods for determining the electrostatic potential of the binding pocket are known in the art and will be apparent to those of skill in the art in light of the present disclosure. For example, in some embodiments, the adaptive Poisson-Boltzmann solver (APBS) tool in PyMOL (PyMOL Molecular Graphics System, Version 2.0, Schrodinger, LLC) may be used (e.g., with default parameters). This tool uses pdb2pqr with the AMBER force field to calculate the electrostatic surface potential of the binding pocket and can assign protonation states. See, for example, TJ Dolinsky et al., PDB2PQR: An automated pipeline for the setup of Poisson-Boltzmann electrostatics calculations, Nucleic Acids Res., 32:W665-7 (2004); MGLerner et al., APBS plugin for PyMOL-Version 2.4 (University of Michigan, Ann Arbor, MI, 2006); JW Ponder et al., Force fields for protein simulations, Adv. Protein Chem. 66:27-85 (2003). In some cases, this solvent-accessible surface area (SASA) may be considered accessible by peptide ligands.

[0088] In some embodiments, the binding pocket comprises a plurality of hydrogen bond acceptors or donors configured to form one or more hydrogen bonds in the presence of an amino acid ligand. In some embodiments, the binding pocket forms at least two (e.g., at least three, at least four, at least five, 2-10, 4-10, 5-15, 5-10) hydrogen bonds with the amino acid ligand. Methods for determining hydrogen bonding interactions between a binding pocket and a ligand are known in the art and will be apparent to those of skill in the art in light of the present disclosure. For example, in some embodiments, the hydrogen bonding interactions are determined by computer modeling (e.g., using atomic coordinates for protein-ligand structural data obtained from X-ray crystallography or by computer modeling to predict protein-ligand three-dimensional structure) as described in the Examples herein.

[0089] In some embodiments, the plurality of hydrogen bond acceptors or donors comprises one or more atoms of the side chains of amino acid residues in the binding pocket. For example, in some embodiments, the binding pocket comprises at least three negatively charged amino acid side chains, each of which (e.g., aspartic acid, glutamic acid) forms a hydrogen bond with the amino acid ligand. In some embodiments, at least one of the negatively charged amino acid side chains forms a hydrogen bond with the amino terminus of the amino acid ligand. In some embodiments, the binding pocket comprises at least one polar, uncharged amino acid side chain (e.g., serine, glutamine) that forms a hydrogen bond with the amino acid ligand. In some embodiments, the plurality of hydrogen bond acceptors or donors comprises one or more atoms of the polypeptide backbone (e.g., backbone carbonyl) within the binding pocket. In some embodiments, the binding pocket comprises at least one negatively charged amino acid side chain and at least one positively charged amino acid side chain, each of which forms a hydrogen bond with the amino acid ligand. In some embodiments, at least one negatively charged amino acid side chain (e.g., aspartic acid, glutamic acid) forms a hydrogen bond with a main chain atom (e.g., nitrogen) of the amino acid ligand. In some embodiments, at least one positively charged amino acid side chain (e.g., lysine) forms a hydrogen bond with a side chain atom of the amino acid ligand.

[0090] In some embodiments, the binding pocket forms one or more hydrogen bonds with the side chain of the amino acid ligand. For example, in some embodiments, the binding pocket forms one or more hydrogen bonds with the side chain of the terminal amino acid of the amino acid ligand (e.g., a polypeptide). In some embodiments, the binding pocket forms one or more hydrogen bonds with the polypeptide backbone of the amino acid ligand (e.g., a polypeptide). In some embodiments, the binding pocket forms one or more hydrogen bonds with the terminal amino acid in the polypeptide and one or more amino acids adjacent to the terminal amino acid (e.g., the amino acids at positions 1 and 2, 3, 4, and / or 5 relative to the polypeptide terminus).

[0091] In some embodiments, the binding pocket comprises a plurality of van der Waals contact positions configured to form van der Waals interactions in the presence of an amino acid ligand. Methods for determining van der Waals interactions between a binding pocket and a ligand are known in the art and will be apparent to those of skill in the art in light of the present disclosure. For example, in some embodiments, the van der Waals interactions are determined by computer modeling as described in the Examples herein (e.g., using atomic coordinates for protein-ligand structural data derived from X-ray crystallography, or by computer modeling to predict protein-ligand three-dimensional structures).

[0092] In some embodiments, the van der Waals contact sites include a plurality of atoms (e.g., 2-30, 5-25, 10-20, 2-10, 5-10) configured to form hydrophobic interactions with the amino acid ligand. In some embodiments, one or more atoms of the plurality of atoms are non-polar atoms.

[0093] In some embodiments, the amino acid ligand is a polypeptide comprising at least three amino acids. In some embodiments, the amino acid ligand comprises the N-terminal amino acid of a polypeptide. In some embodiments, the N-terminal amino acid is selected from glutamine, asparagine, glutamic acid, aspartic acid, and cysteine-S-acetamide. In some embodiments, the amino acid ligand is a polypeptide comprising an N-terminal glutamine or asparagine. In some embodiments, the binding pocket comprises (i), (ii), (iii), and (iv), and the amino acid ligand is a polypeptide comprising an N-terminal glutamine or asparagine. In some embodiments, the amino acid ligand is a polypeptide comprising an N-terminal glutamic acid. In some embodiments, the binding pocket comprises (i), (ii), (iii), (iv), and (v), and the amino acid ligand is a polypeptide comprising an N-terminal glutamic acid. In some embodiments, the amino acid binding protein is at least 50 amino acids in length, at least 75 amino acids in length, at least 100 amino acids in length, 50-250 amino acids in length, 50-150 amino acids in length, or 100-200 amino acids in length.

[0094] In some embodiments, α2 and β4 each comprise at least one polar, uncharged amino acid that forms a hydrogen bond with an amino acid ligand. In some embodiments, the at least one polar, uncharged amino acid of α2 is serine. In some embodiments, the at least one polar, uncharged amino acid of β4 is glutamine.

[0095] In some embodiments, α2 and the loop between α1 and α2 comprise an amino acid sequence that is at least 80% identical to the sequence of amino acids 18-40 of SEQ ID NO: 3. In some embodiments, the binding pocket is formed by amino acids at one or more positions corresponding to amino acids 23-26 of SEQ ID NO: 3. In some embodiments, the amino acid sequence comprises an amino acid substitution at a position corresponding to C25 of SEQ ID NO: 3. In some embodiments, the amino acid substitution is C25S.

[0096] In some embodiments, β3-β4 comprises an amino acid sequence that is at least 80% identical to the sequence of amino acids 73-85 of SEQ ID NO: 3. In some embodiments, the binding pocket is formed by amino acids at one or more positions corresponding to amino acids 75-78 of SEQ ID NO: 3. In some embodiments, the amino acid sequence comprises an amino acid substitution at a position corresponding to H78 of SEQ ID NO: 3. In some embodiments, the amino acid substitution is H78Q.

[0097] In some embodiments, α6 comprises an amino acid sequence that is at least 66% identical to the sequence of amino acids 144-146 of SEQ ID NO: 3. In some embodiments, the binding pocket is formed by amino acids at one or more positions corresponding to amino acids 145-146 of SEQ ID NO: 3.

[0098] In some embodiments, the amino acid binding protein comprises the structure of formula (III-A) or a structural equivalent thereof:

[0099] [ka]

[0100] wherein each of α7 and α8 is an α-helix; and β7 is a β-strand. In some embodiments, the amino acid binding protein comprises the structure of formula (III-B) or a structural equivalent thereof:

[0101] [ka]

[0102] wherein each of α1, α2, and α3 is an α-helix; each of β1, β2, β3, and β4 is a β-strand; each instance of "-" is a loop; and at least a portion of each of α2, β3, β4, the loop between α1 and α2, and the loop between β3 and β4 form a binding pocket for an amino acid ligand.

[0103] In some embodiments, the binding pocket comprises (i) at least one negatively charged amino acid configured to form a hydrogen bond with the amino acid ligand, and (ii) at least one positively charged amino acid configured to form a hydrogen bond with the amino acid ligand. In some embodiments, the amino acid ligand is a polypeptide comprising at least three amino acids. In some embodiments, the amino acid ligand comprises the N-terminal amino acid of the polypeptide. In some embodiments, the N-terminal amino acid is glutamic acid.

[0104] In some embodiments, at least one negatively charged amino acid forms a hydrogen bond with a main chain atom of the amino acid ligand. In some embodiments, at least one positively charged amino acid forms a hydrogen bond with a side chain atom of the amino acid ligand. In some embodiments, at least one negatively charged amino acid forms a hydrogen bond with a main chain atom of the amino acid ligand, and at least one positively charged amino acid forms a hydrogen bond with a side chain atom of the amino acid ligand. In some embodiments, at least one side chain atom of the negatively charged amino acid forms a hydrogen bond with a main chain atom of the amino acid ligand. In some embodiments, at least one side chain atom of the positively charged amino acid forms a hydrogen bond with a side chain atom of the amino acid ligand. In some embodiments, at least one side chain atom of the negatively charged amino acid forms a hydrogen bond with a main chain atom of the amino acid ligand, and at least one side chain atom of the positively charged amino acid forms a hydrogen bond with a side chain atom of the amino acid ligand.

[0105] In some embodiments, α2 comprises at least one negatively charged amino acid. In some embodiments, β4 comprises at least one positively charged amino acid. In some embodiments, α2 comprises at least one negatively charged amino acid and β4 comprises at least one positively charged amino acid. In some embodiments, the at least one negatively charged amino acid comprises glutamic acid. In some embodiments, the at least one positively charged amino acid comprises lysine. In some embodiments, the at least one negatively charged amino acid comprises glutamic acid and the at least one positively charged amino acid comprises lysine. In some embodiments, the at least one negatively charged amino acid corresponds to E26 of SEQ ID NO:3. In some embodiments, the at least one positively charged amino acid is a lysine substitution at a position corresponding to H78 of SEQ ID NO:3. In some embodiments, the at least one negatively charged amino acid corresponds to E26 of SEQ ID NO:3 and the at least one positively charged amino acid is a lysine substitution at a position corresponding to H78 of SEQ ID NO:3.

[0106] In some embodiments, α1-α2 comprise an amino acid sequence at least 80% identical to amino acids 15-39 of SEQ ID NO:3. In some embodiments, the amino acid sequence comprises amino acid substitutions at positions corresponding to S22 and C25 of SEQ ID NO:3. In some embodiments, the amino acid substitutions are S22E and C25S. In some embodiments, β3-β4 comprise an amino acid sequence at least 80% identical to amino acids 73-85 of SEQ ID NO:3. In some embodiments, the amino acid sequence comprises amino acid substitutions at positions corresponding to H78 and C85 of SEQ ID NO:3. In some embodiments, the amino acid substitutions are H78K and C85T.

[0107] In some embodiments, a structural equivalent is a structure in which at least 80% of the secondary structural α carbon atoms align with the structure of Formula (III), (III-A), or (III-B) with a root mean square difference of 5.0 Å or less. In some embodiments, a structural equivalent is a structure in which at least 80% of the secondary structural α carbon atoms align with the structure of Formula (III), (III-A), or (III-B) with a root mean square difference of 4.0 Å or less, 3.0 Å or less, 2.0 Å or less, or 1.0 Å or less.

[0108] Methods for identifying structural equivalents to the structures of formula (III), (III-A), or (III-B) are known in the art and will be apparent to those of skill in the art in light of this disclosure. For example, in some embodiments, structurally equivalent to Formula (III), (III-A), or (III-B) is a structure having a root mean square difference of 5.0 Å or less, wherein at least 80% of the secondary structural α carbon atoms align with a three-dimensional protein structure having an amino acid sequence selected from any one of PS1258-1260, PS1315-1318, PS1457-1478, PS1480-1499, PS1633-1656, PS1737-1758, PS1821-1898, PS2014-2057, and PS2116-2137 (SEQ ID NOs: 604-606, 660-663, 792-833, and 836-1025). Protein structure comparison can be performed by determining the three-dimensional structure for a protein having an amino acid sequence selected from any one of PS1258-1260, PS1315-1318, PS1457-1478, PS1480-1499, PS1633-1656, PS1737-1758, PS1821-1898, PS2014-2057, and PS2116-2137 (SEQ ID NOs: 604-606, 660-663, 792-833, and 836-1025), and comparing the three-dimensional structure for the protein with a candidate protein structure to determine whether the candidate protein is a structural equivalent. The three-dimensional protein structure can be determined, for example, using atomic coordinates derived from X-ray crystallography or via computer modeling to predict the three-dimensional structure of a protein having an amino acid sequence selected from any one of PS1258-1260, PS1315-1318, PS1457-1478, PS1480-1499, PS1633-1656, PS1737-1758, PS1821-1898, PS2014-2057, and PS2116-2137 (SEQ ID NOs: 604-606, 660-663, 792-833, and 836-1025).Methods for comparing protein structures and determining root mean square difference values are known in the art (see, e.g., Kufareva I, Abagyan R., Methods of protein structure comparison., Methods Mol Biol., 2012;857:231-57).

[0109] D. BIR domain homologous recognition factor In some embodiments, the amino acid recognition factors of the present disclosure bind to amino acid ligands (e.g., polypeptides) containing an N-terminal alanine. In some embodiments, the amino acid recognition factor comprises an amino acid binding protein derived from a baculovirus IAP repeat-containing (BIR) protein, e.g., a Homo sapiens BIR3 domain protein. For example, in some embodiments, the amino acid binding protein is an engineered variant containing one or more modifications to SEQ ID NO: 511 described herein.

[0110] In some embodiments, the amino acid binding protein has a dissociation constant (K) of less than 2,000 nM, less than 1,500 nM, less than 1,000 nM, less than 750 nM, less than 500 nM, less than 250 nM, less than 150 nM, less than 100 nM, less than 80 nM, 10 to 2,000 nM, 25 to 1,000 nM, 50 to 500 nM, 10 to 150 nM, 25 to 75 nM, or 50 to 60 nM. D ) binds to the N-terminal alanine.

[0111] In some embodiments, the amino acid binding protein binds to an N-terminal alanine and the binding interaction is at least 0.1 s -1 Dissociation rate (k off In some embodiments, the dissociation rate is about 0.1 s -1 ~about 1,000s -1 (For example, about 0.5 seconds -1 ~about 500s -1 , about 0.1 seconds -1 ~approx. 100s -1 , about 1 s -1 ~approx. 100s-1 , or about 0.5 seconds -1 ~about 50s -1 In some embodiments, the dissociation rate is about 0.5 s -1 ~approx. 20 seconds -1 In some embodiments, the dissociation rate is about 2 s -1 ~approx. 20 seconds -1 In some embodiments, the dissociation rate is about 0.5 s -1 ~about 2s -1 is.

[0112] In some aspects, the disclosure provides recombinant or synthetic amino acid binding proteins having an amino acid sequence that is at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or 100% identical to a sequence selected from any one of PS1165-1166 (SEQ ID NOs: 511-512), PS1267 (SEQ ID NO: 613), and PS1399-1424 (SEQ ID NOs: 742-767). In some embodiments, the amino acid sequence is about 40% to about 100% (e.g., 80 to 98%, 80 to 95%, 80 to 90%, 85 to 95%, 90 to 98%, 50 to 100%, 60 to 100%, 70 to 100%, 80 to 100%, 90 to 100%, or 95 to 100%) identical to a sequence selected from any one of PS1165 to 1166 (SEQ ID NO: 511 to 512), PS1267 (SEQ ID NO: 613), and PS1399 to 1424 (SEQ ID NO: 742 to 767). In some embodiments, the amino acid sequence is at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or 100% identical to PS1165 (SEQ ID NO: 511). In some embodiments, the amino acid sequence is about 40% to about 100% (e.g., 80-98%, 80-95%, 80-90%, 85-95%, 90-98%, 50-100%, 60-100%, 70-100%, 80-100%, 90-100%, or 95-100%) identical to PS1165 (SEQ ID NO: 511).

[0113] In some aspects, the disclosure provides recombinant or synthetic amino acid binding proteins having an amino acid sequence at least 80% identical to PS1165 (SEQ ID NO: 511) and including one or more labels described herein. In some embodiments, the amino acid sequence is at least 85%, at least 90%, at least 95%, at least 98%, or 100% identical to PS1165 (SEQ ID NO: 511). In some embodiments, the amino acid sequence is about 80% to about 100% (e.g., 80-98%, 80-95%, 80-90%, 85-95%, 90-98%, 90-100%, or 95-100%) identical to PS1165 (SEQ ID NO: 511).

[0114] E. Tandem Recognition Factors In some embodiments, an amino acid recognition factor comprises a single polypeptide having tandem copies of two or more amino acid binding proteins, at least one of which is an amino acid binding protein of the present disclosure. As used herein, in some embodiments, tandem arrangement or orientation of elements in a molecule refers to joining each element to the next element end-to-end in a linear manner, such that the elements are fused in series. For example, in some embodiments, a polypeptide having tandem copies of two amino acid binding proteins refers to a fusion polypeptide in which the C-terminus of one protein is fused to the N-terminus of the other protein. Similarly, a polypeptide having tandem copies of two or more amino acid binding proteins refers to a fusion polypeptide in which the C-terminus of a first protein is fused to the N-terminus of a second protein, the C-terminus of the second protein is fused to the N-terminus of a third protein, and so on. Such fusion polypeptides can comprise multiple copies of the same amino acid binding protein or multiple copies of different amino acid binding proteins. In some embodiments, the fusion polypeptides of the present application have at least two and up to 10 amino acid binding proteins (e.g., at least two binding agents and up to 8, 6, 5, 4, or 3 binding agents). In some embodiments, the fusion polypeptides of the present application have five or fewer amino acid binding proteins (e.g., two, three, four, or five amino acid binding proteins).

[0115] In some embodiments, a fusion polypeptide is provided by expression of a single coding sequence containing segments encoding monomeric amino acid-binding protein subunits separated by a segment encoding a flexible linker, where expression of the single coding sequence produces a single full-length polypeptide with two or more independent binding sites. In some embodiments, one or more of the monomeric subunits is a ClpS homologous protein, a UBR homologous protein, or an Ntaq1 homologous protein. In some embodiments, the monomeric subunits can be identical or non-identical. If not identical, the monomeric subunits can be different variants of the same parent homologous protein, or they can be derived from different parent homologous proteins. In some embodiments, the fusion polypeptide comprises two or more ClpS homologous monomers, two or more UBR homologous monomers, or two or more Ntaq1 homologous monomers.

[0116] In some embodiments, at least one amino acid binding protein of the fusion polypeptide has an amino acid sequence selected from Table 1 (or has an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, 80-90%, 90-95%, 95-99%, or more amino acid sequence identity to an amino acid sequence selected from Table 1). In some embodiments, each amino acid binding protein of the fusion polypeptide has an amino acid sequence that is at least 80% identical (e.g., 80-90%, 90-95%, 95-99%, or more) to an amino acid sequence selected from Table 1 (or has an amino acid sequence that has at least 50%, at least 60%, at least 70%, at least 80%, 80-90%, 90-95%, 95-99%, or more amino acid sequence identity to an amino acid sequence selected from Table 1). In some embodiments, the amino acid binding proteins of the fusion polypeptide are modified to include one or more amino acid deletions, additions, or mutations compared to the sequences set forth in Table 1. In some embodiments, the amino acid binding sequences of the fusion polypeptides comprise deletions, additions or mutations of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acids (which may or may not be consecutive amino acids) compared to the sequences set forth in Table 1.

[0117] In some embodiments, the amino acid binding proteins of the fusion polypeptide recognize the same set of one or more amino acids. In some embodiments, the amino acid binding proteins of the fusion polypeptide recognize different sets of one or more amino acids. In some embodiments, the amino acid binding proteins of the fusion polypeptide recognize overlapping sets of amino acids. In some embodiments, when the amino acid binding proteins of the fusion polypeptide recognize the same amino acids, they may recognize amino acids with the same characteristic pulse pattern or amino acids with different characteristic pulse patterns.

[0118] In some embodiments, the amino acid binding proteins of a fusion polypeptide are joined end-to-end, either by a covalent bond or by a linker that covalently links the C-terminus of one protein to the N-terminus of another protein. In the context of fusion polypeptides herein, a linker refers to one or more amino acids within the fusion polypeptide that link two amino acid binding proteins and do not form part of the polypeptide sequence corresponding to either of the two proteins. In some embodiments, the linker comprises at least two amino acids (e.g., at least 2, 3, 4, 5, 6, 8, 10, 15, 25, 50, 100, or more amino acids). In some embodiments, the linker comprises up to 5, up to 10, up to 15, up to 25, up to 50, or up to 100 amino acids. In some embodiments, the linker comprises from about 2 to about 200 amino acids (e.g., from about 2 to about 100, from about 5 to about 50, from about 2 to about 20, from about 5 to about 20, or from about 2 to about 30 amino acids).

[0119] Thus, in some embodiments, the present disclosure provides an amino acid recognition factor comprising a polypeptide in which a first amino acid binding protein and a second amino acid binding protein are linked end-to-end, wherein the first and second amino acid binding proteins are separated by a linker comprising at least two amino acids.

[0120] In some embodiments, each of the first and second amino acid binding proteins independently has an amino acid sequence at least 80% identical (e.g., at least 85%, at least 90%, at least 95%, at least 98%, 80-100%, 85-100%, 90-100%, 95-100%, or 100% identical) to PS961 (SEQ ID NO: 314). In some embodiments, the amino acid recognition factor comprises a polypeptide having an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, 80-100%, 85-100%, 90-100%, 95-100%, or 100% identical to a sequence selected from any one of PS1038, PS1222, and PS1223 (SEQ ID NOs: 389, 568, and 569).

[0121] In some embodiments, each of the first and second amino acid binding proteins independently has an amino acid sequence at least 80% identical (e.g., at least 85%, at least 90%, at least 95%, at least 98%, 80-100%, 85-100%, 90-100%, 95-100%, or 100% identical) to PS1122 (SEQ ID NO: 468). In some embodiments, the amino acid recognition factor comprises a polypeptide having an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, 80-100%, 85-100%, 90-100%, 95-100%, or 100% identical to a sequence selected from any one of PS1219 to PS1221 (SEQ ID NOs: 565-567).

[0122] In some embodiments, each of the first and second amino acid binding proteins independently has an amino acid sequence at least 80% identical (e.g., at least 85%, at least 90%, at least 95%, at least 98%, 80-100%, 85-100%, 90-100%, 95-100%, or 100% identical) to PS1259 (SEQ ID NO: 605). In some embodiments, the amino acid recognition factor comprises a polypeptide having an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, 80-100%, 85-100%, 90-100%, 95-100%, or 100% identical to PS1599 (SEQ ID NO: 835).

[0123] In some aspects, the present application provides nucleic acids encoding a single polypeptide having tandem copies of two or more amino acid binding proteins. In some embodiments, the nucleic acid is an expression construct encoding a fusion polypeptide of the present application. In some embodiments, the expression construct encodes a fusion polypeptide having at least two and up to 10 amino acid binding proteins (e.g., at least two and up to 3, 4, 5, 6, 7, 8, 9, or 10 amino acid binding proteins). In some embodiments, the expression construct encodes a fusion polypeptide having five or fewer amino acid binding proteins (e.g., two, three, four, or five amino acid binding proteins).

[0124] F. Shielded Recognition Factors According to the embodiments described herein, single-molecule polypeptide sequencing can be performed by irradiating a surface-immobilized polypeptide with excitation light and detecting the emission generated by a label attached to the amino acid recognition factor. In some cases, radioactive and / or non-radioactive decay generated by the label can cause photodamage to the polypeptide. The inventors of the present application have found that incorporating a shielding element into the amino acid recognition factor can reduce photodamage and extend the recognition time. See, for example, PCT International Publication No. WO2020102741A1, filed November 15, 2019, and PCT International Publication No. WO2021236983A2, filed May 20, 2021, which describe shielded recognition molecules in detail (the relevant contents of which are incorporated by reference in their entirety).

[0125] Thus, in some aspects, the present disclosure provides shielded recognition agents that include at least one amino acid recognition agent (e.g., an amino acid binding protein) described herein, at least one detectable label, and a shielding element (e.g., a "shield") that forms a covalent or non-covalent bond between the recognition agent and the label. In some embodiments, the shield forms a covalent or non-covalent bond between the one or more amino acid binding proteins and the one or more labels.

[0126] In some embodiments, the shielded recognition factor comprises a fusion polypeptide having an amino acid binding protein of the present disclosure and a protein shield joined end-to-end (e.g., C-terminal to N-terminal). In some embodiments, the protein shield comprises a label protein, such as a fluorescent protein or a non-fluorescent protein comprising a luminescent label.

[0127] In some embodiments, the amino acid binding protein and the protein shield are joined end-to-end, either by a covalent bond or by a linker that covalently links the C-terminus of one protein to the N-terminus of the other protein. In some embodiments, a linker, in the context of a fusion polypeptide, refers to one or more amino acids within the fusion polypeptide that connect the amino acid binding protein and the protein shield and that do not form part of the polypeptide sequence corresponding to either the amino acid binding protein or the protein shield. In some embodiments, the linker comprises at least two amino acids (e.g., at least 2, 3, 4, 5, 6, 8, 10, 15, 25, 50, 100, or more amino acids). In some embodiments, the linker comprises up to 5, up to 10, up to 15, up to 25, up to 50, or up to 100 amino acids. In some embodiments, the linker comprises from about 2 to about 200 amino acids (e.g., from about 2 to about 100, from about 5 to about 50, from about 2 to about 20, from about 5 to about 20, or from about 2 to about 30 amino acids).

[0128] In some embodiments, the protein shield of the fusion polypeptide is a protein having a molecular weight of at least 10 kDa. For example, in some embodiments, the protein shield is a protein having a molecular weight of at least 10 kDa up to 500 kDa (e.g., about 10 kDa to about 250 kDa, about 10 kDa to about 150 kDa, about 10 kDa to about 100 kDa, about 20 kDa to about 80 kDa, about 15 kDa to about 100 kDa, or about 15 kDa to about 50 kDa). In some embodiments, the protein shield of the fusion polypeptide is a protein comprising at least 25 amino acids. For example, in some embodiments, the protein shield is a protein comprising at least 25 and up to 1,000 amino acids (e.g., about 100 to about 1,000 amino acids, about 100 to about 750 amino acids, about 500 to about 1,000 amino acids, about 250 to about 750 amino acids, about 50 to about 500 amino acids, about 100 to about 400 amino acids, or about 50 to about 250 amino acids).

[0129] In some embodiments, the protein shield is a polypeptide comprising one or more tag proteins. In some embodiments, the protein shield is a polypeptide comprising at least two tag proteins. In some embodiments, at least two tag proteins are the same (e.g., the polypeptide comprises at least two copies of the tag protein sequence). In some embodiments, at least two tag proteins are different (e.g., the polypeptide comprises at least two different tag protein sequences). Examples of tag proteins include, but are not limited to, Fasciola hepatica 8-kDa antigen (Fh8), maltose-binding protein (MBP), N-utilization substance (NusA), thioredoxin (Trx), small ubiquitin-like modifier (SUMO), glutathione-S-transferase (GST), solubility-enhancer peptide sequence (SET), IgG domain B1 of protein G (GB1), IgG repeat domain ZZ of protein A (ZZ), mutant dehalogenase (HaloTag), solubility-enhancing ubiquitous tag (SNUT), 17 kilodalton protein (Skp), phage T7 protein kinase (T7PK), Escherichia coli secreted protein A (EspA), and monomeric bacteriophage T7. Examples of such proteins include the 0.3 protein (Orc protein; Mocr), Escherichia coli trypsin inhibitor (Ecotin), calcium-binding protein (CaBP), stress-responsive arsenate reductase (ArsC), the N-terminal fragment of translation initiation factor IF2 (IF2-domain I), stress-responsive proteins (e.g., RpoA, SlyD, Tsf, RpoS, PotD, Crr), and Escherichia coli acidic proteins (e.g., msyB, yjgD, rpoD).See, for example, Costa, S. et al., "Fusion tags for protein solubility, purification and immunogenicity in Escherichia coli: the novel Fh8 system." Front Microbiol., 2014 Feb 19;5:63, the relevant contents of which are incorporated herein by reference.

[0130] The shielding elements of the present disclosure can advantageously absorb, deflect, or otherwise block radioactive and / or non-radioactive decay emitted by the label of the amino acid recognition factor. Therefore, it should be understood that an appropriate protein shield for a fusion polypeptide can be easily selected by one of ordinary skill in the art. For example, the inventors of the present application have demonstrated the use of various types of protein shields in the context of fusion polypeptides including polypeptides having an amino acid binding protein fused to an enzyme (e.g., DNA polymerase, glutathione S-transferase), a transport protein (e.g., maltose-binding protein), a fluorescent protein (e.g., GFP), and a commercially available tag protein (e.g., SNAP-tag®). The inventors of the present application have further demonstrated the use of fusion polypeptides having multiple copies of a protein shield oriented in tandem. See, for example, PCT International Publication No. WO2021236983A2, filed May 20, 2021.

[0131] Thus, in some embodiments, the present disclosure provides fusion polypeptides having one or more tandemly oriented amino acid binding proteins fused to one or more tandemly oriented protein shields. In some embodiments, when a fusion polypeptide comprises two or more tandemly oriented binding agents and / or two or more tandemly oriented shields, an end of one of the two or more binding agents is joined end-to-end to an end of one of the two or more shields. Fusion polypeptides having tandem copies of two or more binding agents are described elsewhere herein, and in some embodiments, such fusions may further comprise a protein shield joined end-to-end to one of the two or more binding agents.

[0132] Further exemplary configurations of shielded recognition factors and shielding elements (e.g., oligonucleotide shields, avidin protein shields) have been described and are contemplated for use in accordance with the present disclosure. See, e.g., PCT International Publication No. WO2020102741A1, filed November 15, 2019, and PCT International Publication No. WO2021236983A2, filed May 20, 2021, the relevant contents of each of which are incorporated herein by reference.

[0133] G. Signs In some embodiments, the amino acid recognition factors of the present disclosure comprise one or more labels. In some embodiments, the one or more labels comprise a detectable label, such as a luminescent label or a conductivity label. As described herein, in some embodiments, one or more chemical characteristics of a polypeptide can be determined by monitoring a signal for a change in a signal (e.g., a signal pulse) corresponding to a binding event between one or more amino acid recognition factors and the polypeptide. In some embodiments, the amino acid recognition factor comprises a detectable label that produces a change in a signal during a binding event between the amino acid recognition factor and the polypeptide. Thus, as used herein, a detectable label of an amino acid recognition factor can refer to any label that can produce a detectable change in a signal during a binding event between the amino acid recognition factor and the polypeptide.

[0134] In some embodiments, the one or more labels of the amino acid recognition factor comprise a luminescent label. In some embodiments, the luminescent label comprises at least one fluorophore dye molecule (e.g., at least two, at least three, at least four, at least five, 20 or fewer, 15 or fewer, 10 or fewer fluorophore dye molecules). In some embodiments, the luminescent label comprises at least one FRET pair comprising a donor label and an acceptor label. Examples of luminescent labels and their uses according to the present disclosure are described in detail elsewhere herein.

[0135] In some embodiments, one or more labels of the amino acid recognition factor include a conductive label. In some embodiments, the conductive label is a charged label, such as a charged polymer. Examples of charged labels include dendrimers, nanoparticles, nucleic acids, and other polymers with multiple charged groups. In some embodiments, the conductive label is uniquely identifiable by its net charge (e.g., net positive charge or net negative charge), by its charge density, and / or by the number of its charged groups.

[0136] In some embodiments, one or more tags of an amino acid recognition factor comprise a tag sequence. For example, in some embodiments, an amino acid recognition factor comprises a tag sequence that provides one or more functions other than amino acid binding. In some embodiments, the tag sequence comprises at least one biotin ligase recognition sequence that enables biotinylation of the recognition factor (e.g., incorporation of one or more biotin molecules comprising biotin and bisbiotin moieties). In some embodiments, the tag sequence comprises two biotin ligase recognition sequences oriented in tandem. In some embodiments, a biotin ligase recognition sequence refers to an amino acid sequence recognized by a biotin ligase, which catalyzes the covalent bond between the sequence and a biotin molecule. Each biotin ligase recognition sequence of a tag sequence can be covalently linked to a biotin moiety, such that a tag sequence having multiple biotin ligase recognition sequences can be covalently linked to multiple biotin molecules. A region of a tag sequence having one or more biotin ligase recognition sequences can generally be referred to as a biotinylation tag or biotinylation sequence. In some embodiments, bisbiotin or bisbiotin moiety can refer to two biotins attached to two biotin ligase recognition sequences oriented in tandem.

[0137] Further examples of functional sequences in tag sequences include purification tags, cleavage sites, and other moieties useful for purifying and / or modifying the recognition factor. Table 2 provides a non-limiting list of tag sequences, any one or more of which can be used in combination with any one of the amino acid recognition factors of the present application (e.g., in combination with the sequences set forth in Table 1). It should be understood that the tag sequences set forth in Table 2 are meant to be non-limiting, and that a recognition factor in accordance with the present application can include any one or more of the tag sequences (e.g., His tags and / or biotinylation tags) at the N-terminus or C-terminus of the recognition factor polypeptide, or at an internal position split between the N-terminus and C-terminus, or otherwise rearranged as practiced in the art.

[0138] In some embodiments, one or more labels of the amino acid recognition factor comprise a biotin moiety. In some embodiments, the biotin moiety comprises at least one biotin molecule (e.g., 1, 2, 3, 4, or more biotin molecules). In some embodiments, the biotin moiety is a bis-biotin moiety. In some embodiments, the biotin moiety comprises at least one biotin molecule bound to at least one biotin ligase recognition sequence. For example, in some embodiments, the one or more labels comprise a tag sequence comprising two biotin ligase recognition sequences oriented in tandem, each biotin ligase recognition sequence having a biotin molecule attached thereto. In some embodiments, the biotin moiety comprises at least one biotin molecule bound to the amino acid recognition factor via means other than the tag sequence. For example, in some embodiments, at least one biotin molecule is chemically conjugated to an amino acid (e.g., an unnatural amino acid) of the amino acid binding protein.

[0139] In some embodiments, the one or more labels of the amino acid recognition agent comprise one or more polyol moieties (e.g., one or more moieties selected from dextran, polyvinylpyrrolidone, polyethylene glycol, polypropylene glycol, polyoxyethylene glycol, and polyvinyl alcohol). For example, in some embodiments, the amino acid recognition factor is PEGylated. In some embodiments, polyol modification (e.g., PEGylation) can limit the degree of nonspecific adhesion to the surface of a substrate (e.g., a sequencing chip). In some embodiments, polyol modification can limit the degree of aggregation or interaction between the amino acid recognition factor and other recognition factors, cleavage reagents, or other species present in the sequencing reaction mixture. PEGylation can be performed by incubating the recognition factor (e.g., an amino acid binding protein such as a ClpS protein) with mPEG4-NHS ester, which labels primary amines, such as surface-exposed lysine side chains. Other types of PEG and other methods of polyol modification are known in the art.

[0140] It should be understood that in some embodiments, the amino acid recognition factors of the present disclosure can include one or more different types of labels described herein. For example, in some embodiments, the amino acid recognition factors include one or more labels selected from a detectable label (e.g., a luminescent label, a conductive label), a tag sequence (e.g., a purification tag, a cleavage site, a biotinylation sequence), a biotin moiety, and a polyol moiety. In some embodiments, the amino acid recognition factors include a detectable label (e.g., a luminescent label, a conductive label) and one or more labels selected from a tag sequence (e.g., a purification tag, a cleavage site, a biotinylation sequence), a biotin moiety, and a polyol moiety.

[0141] In some embodiments, one or more labels of the amino acid recognition factor include a luminescent label. As used herein, a luminescent label is a molecule that can absorb one or more photons and then emit one or more photons after one or more time durations. In some embodiments, this term is used interchangeably with "label," "detectable label," or "luminescent molecule," depending on the context. A luminescent label according to certain embodiments described herein may refer to a luminescent label of an amino acid recognition factor, a luminescent label of a cleavage reagent (e.g., a peptidase such as an aminopeptidase), or a luminescent label of another labeled composition described herein.

[0142] In some embodiments, the luminescent label comprises a first chromophore and a second chromophore. In some embodiments, the excited state of the first chromophore can be relaxed via energy transfer to the second chromophore. In some embodiments, the energy transfer is Förster resonance energy transfer (FRET). Such FRET pairs can be useful for providing luminescent labels with properties that make it easier to distinguish the label from among multiple luminescent labels in a mixture, or for providing binding-induced fluorescence that limits background fluorescence, as described elsewhere herein. In still other embodiments, the FRET pair comprises a first chromophore of the first luminescent label and a second chromophore of the second luminescent label. In certain embodiments, the FRET pair can absorb excitation energy in a first spectral range and emit emission in a second spectral range.

[0143] In some embodiments, the luminescent label refers to a fluorophore or a dye. Typically, the luminescent label comprises an aromatic or heteroaromatic compound, and may be pyrene, anthracene, naphthalene, naphthylamine, acridine, stilbene, indole, benzindole, oxazole, carbazole, thiazole, benzothiazole, benzoxazole, phenanthridine, phenoxazine, porphyrin, quinoline, ethidium, benzamide, cyanine, carbocyanine, salicylate, anthranilate, coumarin, 45-fluororescein, rhodamine, xanthene, or other similar compounds.

[0144] In some embodiments, the luminescent label comprises a dye selected from one or more of the following: 5 / 6-carboxyrhodamine 6G, 5-carboxyrhodamine 6G, 6-carboxyrhodamine 6G, 6-TAMRA, Abberior® STAR 440SXP, Abberior® STAR 470SXP, Abberior® STAR 488, Abberior® STAR 490SXP ... Abberior® STAR 512, Abberior® STAR 520SXP, Abberior® STAR 580, Abberior® STAR 600, Abberior® STAR 635, Abberior® STAR 635P, Abberior® STAR Red RED, Alexa Fluor® 350, Alexa Fluor® 405, Alexa Fluor® 430, Alexa Fluor® 480, Alexa Fluor® 488, Alexa Fluor® 514, Alexa Fluor® 532, Alexa Fluor® 546, Alexa Fluor® 555, Alexa Fluor® 568, Alexa Fluor® 594, Alexa Fluor® 610-X, Alexa Fluor® 633, Alexa Fluor® Alexa Fluor® 647, Alexa Fluor® 660, Alexa Fluor® 680, Alexa Fluor® 700,Alexa Fluor® 750, Alexa Fluor® 790, AMCA, ATTO 390, ATTO 425, ATTO 465, ATTO 488, ATTO 495, ATTO 514, ATTO 520, ATTO 532, ATTO 542, ATTO 550, ATTO 565, ATTO 590, ATTO 610, ATTO 620, ATTO 633, ATTO 647, ATTO ) 647N, ATTO 655, ATTO 665, ATTO 680, ATTO 700, ATTO 725, ATTO 740, ATTO Oxa12, ATTO Rho101, ATTO Rho11, ATTO Rho12, ATTO Rho13, ATTO Rho14, ATTO Rho3B, ATTO Rho6G, ATTO Thio12, BD Horizon ( Trademark) V450, BODIPY® 493 / 501, BODIPY® 530 / 550, BODIPY® 558 / 568, BODIPY® 564 / 570, BODIPY® 576 / 589, BODIPY® 581 / 591, BODIPY® 630 / 650, BODIPY® 650 / 665, BODIPY® CAL Fluor (registered trademark) FL, BODIPY (registered trademark) FL-X, BODIPY (registered trademark) R6G, BODIPY (registered trademark) TMR, BODIPY (registered trademark) TR, CAL Fluor (registered trademark) Gold 540, CAL Fluor (registered trademark) Green 510, CAL Fluor (registered trademark) Orange 560, CAL Fluor (registered trademark) Red 590, CAL Fluor (registered trademark) Red 610,CAL Fluor® Red 615, CAL Fluor® Red 635, Cascade® Blue, CF™ 350, CF™ 405M, CF™ 405S, CF™ 488A, CF™ 514, CF™ 532, CF™ 543, CF™ 546, CF™ 555, CF™ 568, CF™ 594, CF™ 620R, CF™ 633, CF™ 633-V1, CF™ 640R, CF™ 640R-V1 , CF(TM) 640R-V2, CF(TM) 660C, CF(TM) 660R, CF(TM) 680, CF(TM) 680R, CF(TM) 680R-V1, CF(TM) 750, CF(TM) 770, CF(TM) 790, Chromeo(TM) 642, Chromis 425N, Chromis 500N, Chromis 515N, Chromis 530N, Chromis 550A, Chromis 550C, Chromis 550Z, Chromis 560N, Chromis 570N, Chromis 577N, Chromis 600N, Chromis 630N, Chromis 645A, Chromis 645C, Chromis 645Z, Chromis 678A, Chromis 678C, Chromis 678Z, Chromis 770A, Chromis 770C, Chromis 80 0A, Chromis 800C, Chromis 830A, Chromis 830C, Cy® 3, Cy® 3.5, Cy® 3B, Cy® 5, Cy® 5.5, Cy® 7, DyLight® 350, DyLight® 405, DyLight® 415-Co1, DyLight® 425Q, DyLight® 485-LS,DyLight® 488, DyLight® 504Q, DyLight® 510-LS, DyLight® 515-LS, DyLight® 521-LS, DyLight® 530-R2, DyLight® 543Q, DyLight® 550, DyLight® 554-R0, DyLight® DyLight® 554-R1, DyLight® 590-R2, DyLight® 594, DyLight® 610-B1, DyLight® 615-B2, DyLight® 633, DyLight® 633-B1, DyLight® 633-B2, DyLight® 650, DyLight® 655-B1, DyLight® DyLight (registered trademark) 655-B2, DyLight (registered trademark) 655-B3, DyLight (registered trademark) 655-B4, DyLight (registered trademark) 662Q, DyLight (registered trademark) 675-B1, DyLight (registered trademark) 675-B2, DyLight (registered trademark) 675-B3, DyLight (registered trademark) 675-B4, DyLight (registered trademark) 679-C5, DyLight ) (registered trademark) 680, DyLight (registered trademark) 683Q, DyLight (registered trademark) 690-B1, DyLight (registered trademark) 690-B2, DyLight (registered trademark) 696Q, DyLight (registered trademark) 700-B1, DyLight (registered trademark) 700-B1, DyLight (registered trademark) 730-B1, DyLight (registered trademark) 730-B2, DyLight (registered trademark) 730-B3,DyLight® 730-B4, DyLight® 747, DyLight® 747-B1, DyLight® 747-B2, DyLight® 747-B3, DyLight® 747-B4, DyLight® 755, DyLight® 766Q, DyLight® 775-B2, DyLight® DyLight® 775-B3, DyLight® 775-B4, DyLight® 780-B1, DyLight® 780-B2, DyLight® 780-B3, DyLight® 800, DyLight® 830-B2, Dyomics-350, Dyomics-350XL, Dyomics-360XL, Dy Dyomics-370XL, Dyomics-375XL, Dyomics-380XL, Dyomics-390XL, Dyomics-405, Dyomics-415, Dyomics-430, Dyomics-431, Dyomics-478, Dyomics-480XL, Dyomics-481XL, Dyomics-485XL, Dyomics Dyomics-490, Dyomics-495, Dyomics-505, Dyomics-510XL, Dyomics-511XL, Dyomics-520XL, Dyomics-521XL, Dyomics-530, Dyomics-547, Dyomics-547P1, Dyomics-548, Dyomics-549,Dyomics-549P1, Dyomics-550, Dyomics-554, Dyomics-555, Dyomics-556, Dyomics-560, Dyomics-590, Dyomics-591, Dyomics, Dyomics-594, Dyomics-601XL, Dyomics-605, Dyomics-610, Dyomics-615, Dyomics-630, Dyomics-631, Dyomics-632, Dyomics-633, Dyomics-634, Dyomics-635, Dyomics-636, Dyomics Dyomics-647, Dyomics-647P1, Dyomics-648, Dyomics-648P1, Dyomics-649, Dyomics-649P1, Dyomics-650, Dyomics-651, Dyomics-652, Dyomics-654, Dyomics-675, Dyomics-676, Dyomics Dyomics-677, Dyomics-678, Dyomics-679P1, Dyomics-680, Dyomics-681, Dyomics-682, Dyomics-700, Dyomics-701, Dyomics-703, Dyomics-704, Dyomics-730, Dyomics-731, Dyomics -732, Dyomics-734, Dyomics-749, Dyomics-749P1, Dyomics-750, Dyomics-751, Dyomics-752, Dyomics-754, Dyomics-776, Dyomics-777, Dyomics-778, Dyomics-780, Dyomics-781,Dyomics-782, Dyomics-800, Dyomics-831, eFluor® 450, Eosin, FITC, Fluorescein, HiLyte™ Fluor 405, HiLyte™ Fluor 488, HiLyte™ Fluor 532, HiLyte™ Fluor 555, HiLyte™ Fluor 594, HiLyte™ Fluor 647, HiLyte™ Fluor 5 LightCycler® Red 680, HiLyte™ Fluor 750, IRDye® 680LT, IRDye® 750, IRDye® 800CW, JOE, LightCycler® 640R, LightCycler® Red 610, LightCycler® Red 640, LightCycler® Red 670, LightCycler® Red 705, Lissamine Rhodamine B, Naphthofluorescein, Oregon Green Green® 488, Oregon Green® 514, Pacific Blue™, Pacific Green™, Pacific Orange™, PET, PF350, PF405, PF415, PF488, PF505, PF532, PF546, PF555P, PF568, PF594, PF610, PF633P, PF647P, Quasar® 570, Quasar® 670, Quasar® 705, Rhodamine 123, Rhodamine 6G, Rhodamine B, Rhodamine Green, Rhodamine Green-X, Rhodamine Red, ROX, Seta™ 375,Seta™ 470, Seta™ 555, Seta™ 632, Seta™ 633, Seta™ 650, Seta™ 660, Seta™ 670, Seta™ 680, Seta™ 700, Seta™ 750, Seta™ 780, Seta™ APC-780, Seta™ PerCP -680, Seta™ R-PE-670, Seta™ 646, SeTau 380, SeTau 425, SeTau 647, SeTau 405, Square 635, Square 650, Square 660, Square 672, Square 680, Sulforhodamine 101, TAMRA, TET, Texas Red®, TMR, TRITC, Yakima Yellow™, Zenon®, Zy3, Zy5, Zy5.5, and Zy7.

[0145] In some aspects, the present disclosure provides methods and compositions for polypeptide analysis (e.g., amino acid recognition) based on one or more luminescent properties of luminescent labels. In some embodiments, luminescent labels are identified based on luminescence lifetime, luminescence intensity, brightness, absorption spectrum, emission spectrum, luminescence quantum yield, or a combination of two or more thereof. In some embodiments, multiple types of luminescent labels can be distinguished from each other based on differences in luminescence lifetime, luminescence intensity, brightness, absorption spectrum, emission spectrum, luminescence quantum yield, or a combination of two or more thereof.

[0146] In some embodiments, luminescence is detected by exposing a luminescent label to a series of discrete light pulses and evaluating the timing or other characteristics of each photon emitted from the label. In some embodiments, information about multiple photons sequentially emitted from the label is aggregated and evaluated to identify the label and thereby identify the associated barcode site. In some embodiments, the luminescence lifetime of the label is determined from multiple photons sequentially emitted from the label, and the luminescence lifetime may be used to identify the label. In some embodiments, the luminescence intensity of the label is determined from multiple photons sequentially emitted from the label, and the luminescence intensity may be used to identify the label. In some embodiments, the luminescence lifetime and luminescence intensity of the label are determined from multiple photons sequentially emitted from the label, and the luminescence lifetime and luminescence intensity may be used to identify the label.

[0147] In some aspects of the present disclosure, a single molecule is exposed to multiple separate light pulses, and the sequence of emitted photons is detected and analyzed. In some embodiments, the sequence of emitted photons provides information about a single molecule that is present and does not change in the mixture over the course of the experiment. However, in some embodiments, the sequence of emitted photons provides information about a sequence of different molecules present at different times in the mixture (e.g., as a reaction or process progresses).

[0148] In certain embodiments, a luminescent label absorbs one photon and emits one photon after a duration. In some embodiments, the luminescent lifetime of a label can be determined or estimated by measuring its duration. In some embodiments, the luminescent lifetime of a label can be determined or estimated by measuring multiple durations for multiple pulse and emission events. In some embodiments, the luminescent lifetime of a label can be distinguished among the luminescent lifetimes of multiple types of labels by measuring the duration. In some embodiments, the luminescent lifetime of a label can be distinguished among the luminescent lifetimes of multiple types of labels by measuring multiple durations for multiple pulse and emission events. In certain embodiments, a label is identified or distinguished among multiple types of labels by determining or estimating the luminescent lifetime of the label. In certain embodiments, a label is identified or distinguished among multiple types of labels by distinguishing the luminescent lifetime of the label among the multiple luminescent lifetimes of multiple types of labels.

[0149] Determining the emission lifetime of a luminescent label can be done using any suitable method (e.g., by measuring the lifetime using a suitable technique or by determining the time-dependent characteristics of the emission). In some embodiments, determining the emission lifetime of one label comprises determining the lifetime relative to another label. In some embodiments, determining the emission lifetime of a label comprises determining the lifetime relative to a reference. In some embodiments, determining the emission lifetime of a label comprises measuring a lifetime (e.g., a fluorescence lifetime). In some embodiments, determining the emission lifetime of a label comprises determining one or more temporal characteristics indicative of the lifetime. In some embodiments, the emission lifetime of a label can be determined based on the distribution of a plurality of emission events (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more emission events) occurring over one or more time-gated windows relative to an excitation pulse. For example, based on the distribution of photon arrival times measured for an excitation pulse, the emission lifetime of a label can be distinguished from multiple labels with different emission lifetimes.

[0150] It should be understood that the luminescence lifetime of a luminescent label indicates the timing of photons emitted after the label reaches an excited state, and labels can be distinguished by information indicating the timing of the photons. Some embodiments may include distinguishing a label from multiple labels based on the luminescence lifetime of the label by measuring the time associated with photons emitted by the label. The distribution of times may provide an indication of the luminescence lifetime, which may be determined from the distribution. In some embodiments, a label can be distinguished from multiple labels based on the distribution of times, such as by comparing the distribution of times to a reference distribution corresponding to known labels. In some embodiments, the luminescence lifetime value is determined from the distribution of times.

[0151] As used herein, in some embodiments, luminescence intensity refers to the number of emission photons per unit time emitted by a luminescent label excited by the delivery of pulsed excitation energy, hi some embodiments, luminescence intensity refers to the detected number of emission photons per unit time emitted by a label excited by the delivery of pulsed excitation energy and detected by a particular sensor or set of sensors.

[0152] As used herein, in some embodiments, "brightness" refers to a parameter that reports the average luminescence intensity per luminescent label. Thus, in some embodiments, "luminescence intensity" can be used generally to refer to the brightness of a composition comprising one or more labels. In some embodiments, the brightness of a label is equal to the product of its quantum yield and extinction coefficient.

[0153] As used herein, in some embodiments, luminescence quantum yield refers to the proportion of excitation events at a given wavelength or within a given spectral range that result in a luminescence event, and is typically less than 1. In some embodiments, the luminescence quantum yield of a luminescent label described herein is 0 to about 0.001, about 0.001 to about 0.01, about 0.01 to about 0.1, about 0.1 to about 0.5, about 0.5 to 0.9, or about 0.9 to 1. In some embodiments, a label is identified by determining or estimating the luminescence quantum yield.

[0154] As used herein, in some embodiments, the excitation energy is a pulse of light from a light source. In some embodiments, the excitation energy is in the visible spectrum. In some embodiments, the excitation energy is in the ultraviolet spectrum. In some embodiments, the excitation energy is in the infrared spectrum. In some embodiments, the excitation energy is at or near the absorption maximum of a luminescent label from which multiple emitted photons are to be detected. In certain embodiments, the excitation energy is between about 500 nm and about 700 nm (e.g., between about 500 nm and about 600 nm, between about 600 nm and about 700 nm, between about 500 nm and about 550 nm, between about 550 nm and about 600 nm, between about 600 nm and about 650 nm, or between about 650 nm and about 700 nm). In certain embodiments, the excitation energy may be monochromatic or may be limited to a spectral range. In some embodiments, the spectral range has a range of between about 0.1 nm and about 1 nm, between about 1 nm and about 2 nm, or between about 2 nm and about 5 nm. In some embodiments, the spectral range has a range of about 5 nm to about 10 nm, about 10 nm to about 50 nm, or about 50 nm to about 100 nm.

[0155] II. Polypeptide Analysis In some embodiments, the present application provides a method for determining at least one chemical characteristic of a polypeptide by monitoring a signal for a signal pulse corresponding to an interaction between the polypeptide and at least one amino acid recognition factor described herein, and determining at least one chemical characteristic of the polypeptide based on a characteristic pattern in the signal.

[0156] A non-limiting example of polypeptide structural analysis by detecting single-molecule binding interactions during the polypeptide degradation process is shown in FIG. 1. Exemplary signal traces are shown showing different association (e.g., binding) events at times corresponding to signal changes. As shown, association events between an amino acid recognition factor and the terminus of a polypeptide result in a change in signal magnitude that persists for a period of time. Different association events are shown for different amino acids exposed at the terminus of a polypeptide. As described herein, an amino acid that is "exposed" at the terminus of a polypeptide is an amino acid that is still bound to the polypeptide and that becomes the terminal amino acid upon removal of the previous terminal amino acid (e.g., alone or together with one or more additional amino acids) during degradation.

[0157] As generally indicated, association events between amino acid recognition factors and different types of amino acids at the termini of a polypeptide result in characteristic changes in signals, referred to herein as signature patterns, which can be used to determine the chemical characteristics of the polypeptide. In some embodiments, a signature pattern corresponding to one type of terminal amino acid can be used to determine structural information about the terminal amino acid and one or more amino acids adjacent to the terminal amino acid. Thus, in some embodiments, a signature pattern corresponding to one type of terminal amino acid can be used to determine structural information about at least two (e.g., at least three, at least four, at least five, two, three, four, or two to five) amino acids of a polypeptide.

[0158] In some embodiments, a transition from one characteristic pattern to another indicates an amino acid cleavage. As used herein, in some embodiments, amino acid cleavage refers to the removal of at least one amino acid from the end of a polypeptide (e.g., the removal of at least one terminal amino acid from a polypeptide). In some embodiments, amino acid cleavage is determined by inference based on the duration between characteristic patterns. In some embodiments, amino acid cleavage is determined by detecting a change in signal resulting from the association of a labeled cleavage reagent with an amino acid at the end of the polypeptide. As amino acids are sequentially cleaved from the end of the polypeptide during degradation, a series of changes in size, or a series of signal pulses, is detected.

[0159] In some embodiments, the signal data may be analyzed to extract signal pulse information by applying a threshold level to one or more parameters of the signal data. For example, in some embodiments, a threshold magnitude level may be applied to the signal data of a signal trace. In some embodiments, the threshold magnitude level is the minimum difference between the signal detected at a given time point and a baseline determined for a given data set. In some embodiments, a signal pulse that exhibits a change in magnitude exceeding the threshold magnitude level and lasts for a certain duration is assigned to each portion of the data. In some embodiments, a threshold duration may be applied to portions of the data that meet the threshold magnitude level to determine whether a signal pulse is assigned to the portion of the data that meets the threshold magnitude level. For example, experimental artifacts may produce a change in magnitude that exceeds the threshold magnitude level, but this may not last for a sufficient duration to assign a signal pulse with a desired degree of confidence (e.g., a transient related event that may be non-discriminatory for amino acid type, e.g., a non-specific detection event such as diffusion into the observation region or reagent adhesion within the observation region). Thus, in some embodiments, signal pulses are extracted from the signal data based on a threshold magnitude level and a threshold duration.

[0160] In some embodiments, the signal pulse magnitude peak is determined by averaging the detected magnitude over a duration that persists above a threshold magnitude level. It should be understood that in some embodiments, a "signal pulse" as used herein can refer to a change in signal data (e.g., raw signal data) that persists for a duration above the baseline, or signal pulse information extracted therefrom (e.g., processed signal data).

[0161] In some embodiments, the signal pulse information can be analyzed to identify different types of amino acids in a polypeptide based on different characteristic patterns in a series of signal pulses. For example, as shown in Figure 1, the signal pulse information indicates different types of amino acids (e.g., arginine, leucine, isoleucine, phenylalanine) at the end of the polypeptide. For example, the signal pulse detected at the earliest time point provides information indicating (at least) arginine at the end of the polypeptide based on a first characteristic pattern, and the signal pulse detected at the latest time point provides information indicating at least phenylalanine at the end of the polypeptide based on a second characteristic pattern.

[0162] In some embodiments, each signal pulse of the characteristic pattern comprises a pulse duration corresponding to an association event between the amino acid recognition factor and the amino acid ligand. In some embodiments, the pulse duration is characteristic of the dissociation rate of binding. In some embodiments, each signal pulse of the characteristic pattern is separated from another signal pulse of the characteristic pattern by an inter-pulse duration. In some embodiments, the inter-pulse duration is characteristic of the association rate of binding. In some embodiments, a change in signal magnitude can be determined for the same signal pulse based on the difference between the baseline and peak of the signal pulse. In some embodiments, the characteristic pattern is determined based on the pulse duration. In some embodiments, the characteristic pattern is determined based on the pulse duration and the inter-pulse duration. In some embodiments, the characteristic pattern is determined based on any one or more of the pulse duration, the inter-pulse duration, and the change in magnitude.

[0163] Thus, as illustrated by Figure 1, in some embodiments, polypeptide analysis is performed by detecting a series of signal pulses that indicate the association of one or more amino acid recognition factors with consecutive amino acids exposed at the termini of the polypeptide in an ongoing degradation reaction. The series of signal pulses can be analyzed to determine a characteristic pattern in the series of signal pulses, and the time course of the characteristic pattern can be used to determine the chemical signature across the amino acid sequence of the polypeptide.

[0164] As described herein, signal pulse information can be used to identify amino acids based on a characteristic pattern in a series of signal pulses. In some embodiments, the characteristic pattern includes a plurality of signal pulses, each signal pulse including a pulse duration. In some embodiments, the plurality of signal pulses can be characterized by a summary statistic (e.g., mean, median, time decay constant) of the distribution of pulse durations in the characteristic pattern. In some embodiments, the average pulse duration of the characteristic pattern is about 1 millisecond to about 10 seconds (e.g., about 1 ms to about 1 s, about 1 ms to about 100 ms, about 1 ms to about 10 ms, about 10 ms to about 10 s, about 100 ms to about 10 s, about 1 s to about 10 s, about 10 ms to about 100 ms, or about 100 ms to about 500 ms). In some embodiments, the average pulse duration is about 50 milliseconds to about 2 seconds, about 50 milliseconds to about 500 milliseconds, or about 500 milliseconds to about 2 seconds.

[0165] In some embodiments, different characteristic patterns corresponding to different types of amino acids in a single polypeptide can be distinguished from one another based on statistically significant differences in summary statistics. For example, in some embodiments, one characteristic pattern can be distinguished from another characteristic pattern based on a difference in mean pulse duration of at least 10 milliseconds (e.g., about 10 ms to about 10 s, about 10 ms to about 1 s, about 10 ms to about 100 ms, about 100 ms to about 10 s, about 1 s to about 10 s, or about 100 ms to about 1 s). In some embodiments, the difference in mean pulse duration is at least 50 ms, at least 100 ms, at least 250 ms, at least 500 ms, or more. In some embodiments, the difference in mean pulse duration is about 50 ms to about 1 s, about 50 ms to about 500 ms, about 50 ms to about 250 ms, about 100 ms to about 500 ms, about 250 ms to about 500 ms, or about 500 ms to about 1 s. In some embodiments, the mean pulse duration of one characteristic pattern differs from the mean pulse duration of another characteristic pattern by about 10-25%, 25-50%, 50-75%, 75-100%, or more than 100%, e.g., about 2-fold, 3-fold, 4-fold, 5-fold, or more. It should be understood that in some embodiments, smaller differences in mean pulse duration between different characteristic patterns may require more pulse durations within each characteristic pattern to be distinguished from one another with statistical reliability.

[0166] In some embodiments, a characteristic pattern generally refers to a plurality of association events between amino acids of a polypeptide and a means for binding amino acids (e.g., amino acid recognition molecules). In some embodiments, a characteristic pattern comprises at least 10 association events (e.g., at least 25, at least 50, at least 75, at least 100, at least 250, at least 500, at least 1,000, or more association events). In some embodiments, a characteristic pattern comprises about 10 to about 1,000 association events (e.g., about 10 to about 500 association events, about 10 to about 250 association events, about 10 to about 100 association events, or about 50 to about 500 association events). In some embodiments, a plurality of association events is detected as a plurality of signal pulses.

[0167] In some embodiments, a characteristic pattern refers to a plurality of signal pulses that can be characterized by summary statistics described herein. In some embodiments, a characteristic pattern includes at least 10 signal pulses (e.g., at least 25, at least 50, at least 75, at least 100, at least 250, at least 500, at least 1,000, or more signal pulses). In some embodiments, a characteristic pattern includes about 10 to about 1,000 signal pulses (e.g., about 10 to about 500 signal pulses, about 10 to about 250 signal pulses, about 10 to about 100 signal pulses, or about 50 to about 500 signal pulses).

[0168] In some embodiments, a characteristic pattern refers to multiple association events between an amino acid recognition molecule and amino acids of a polypeptide that occur over a time interval prior to removal of an amino acid (e.g., a cleavage event). In some embodiments, a characteristic pattern refers to multiple association events that occur over a time interval between two cleavage events (e.g., before removal of an amino acid and after removal of a terminally previously exposed amino acid). In some embodiments, the time interval of a characteristic pattern is about 1 minute to about 30 minutes (e.g., about 1 minute to about 20 minutes, about 1 minute to about 10 minutes, about 5 minutes to about 20 minutes, about 5 minutes to about 15 minutes, or about 5 minutes to about 10 minutes).

[0169] In some embodiments, the series of signal pulses comprises a series of changes in the magnitude of the light signal over time. In some embodiments, the series of changes in the light signal comprises a series of changes in the luminescence generated during the associated event. In some embodiments, the luminescence is generated by a detectable label associated with one or more reagents of the sequencing reaction. For example, in some embodiments, each of the one or more amino acid recognition factors comprises a luminescent label. In some embodiments, the cleavage reagent comprises a luminescent label. Examples of luminescent labels and their uses in accordance with the present application are provided herein.

[0170] In some embodiments, the series of signal pulses comprises a series of changes in the magnitude of the electrical signal over time. In some embodiments, the series of changes in the electrical signal comprises a series of changes in conductance generated during the relevant events. In some embodiments, the conductivity is provided by a detectable label associated with one or more reagents of the sequencing reaction. For example, in some embodiments, each of the one or more amino acid recognition factors comprises a conductive label. Examples of conductive labels and their uses in accordance with the present application are provided elsewhere herein. Methods for identifying single molecules using conductive labels have been described (see, e.g., U.S. Patent Application Publication No. 2017 / 0037462).

[0171] In some embodiments, the series of conductance changes comprises a series of changes in conductance through a nanopore. For example, methods for assessing receptor-ligand interactions using nanopores have been described (see, e.g., Thakur, A.K. and Movileanu, L., (2019) Nature Biotechnology 37(1)). The inventors of the present application have recognized and appreciated that such nanopores can be used to monitor polypeptide sequencing reactions in accordance with the present application. Accordingly, in some embodiments, the present disclosure provides methods of polypeptide analysis comprising contacting a single polypeptide molecule with one or more amino acid recognition factors described herein, wherein the single polypeptide molecule is immobilized in a nanopore. In some embodiments, the present method further comprises detecting a series of changes in conductance through the nanopore, indicative of association of the one or more amino acid recognition factors with consecutive amino acids exposed at a terminus of the single polypeptide, while the single polypeptide is being degraded.

[0172] As described herein, in some embodiments, the amino acid recognition elements of the present disclosure can be used to determine at least one chemical characteristic of a polypeptide. In some embodiments, determining at least one chemical characteristic includes determining the type of amino acid present at the terminal end of the polypeptide and / or the type of amino acid present at one or more positions adjacent to the terminal amino acid. In some embodiments, determining the type of amino acid includes determining the actual amino acid identity, for example, by determining which of the 20 naturally occurring amino acids is present. In some embodiments, the type of amino acid is selected from alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, selenocysteine, serine, threonine, tryptophan, tyrosine, and valine.

[0173] In some embodiments, determining at least one chemical characteristic of the polypeptide includes determining a subset of potential amino acids that may be present in the polypeptide. In some embodiments, this can be achieved by determining that an amino acid is not one or more specific amino acids (and therefore may be any other amino acid). In some embodiments, this can be achieved by determining which of a specific subset of amino acids may be present in the polypeptide (e.g., based on size, charge, hydrophobicity, post-translational modification, binding properties) (e.g., using a recognition element that binds to a specific subset of two or more amino acids).

[0174] In some embodiments, determining at least one chemical characteristic of the polypeptide includes determining that the amino acid comprises a post-translational modification. Non-limiting examples of post-translational modifications include acetylation (e.g., acetylated lysine), ADP-ribosylation, caspase cleavage, citrullination, formylation, N-linked glycosylation (e.g., glycosylated asparagine), O-linked glycosylation (e.g., glycosylated serine, glycosylated threonine), hydroxylation, methylation (e.g., methylated lysine, methylated arginine), myristoylation (e.g., myristoylated glycine), neddylation, nitration ( For example, these include nitrated tyrosine, chlorination (e.g., chlorinated tyrosine), oxidation / reduction (e.g., oxidized cysteine, oxidized methionine), palmitoylation (e.g., palmitoylated cysteine), phosphorylation, prenylation (e.g., prenylated cysteine), S-nitrosylation (e.g., S-nitrosylated cysteine, S-nitrosylated methionine), sulfation, sumoylation (e.g., sumoylated lysine), and ubiquitination (e.g., ubiquitinated lysine).

[0175] In some embodiments, determining at least one chemical characteristic of the polypeptide includes determining that the amino acid comprises an arginine post-translational modification. For example, as described herein, the amino acid recognition element of the present disclosure can distinguish between different arginine modifications, including symmetric dimethylarginine (SDMA), asymmetric dimethylarginine (ADMA), and citrullinated arginine.

[0176] In some embodiments, determining at least one chemical characteristic of the polypeptide comprises determining that the amino acid comprises a phosphorylated side chain. For example, in some embodiments, determining at least one chemical characteristic of the polypeptide comprises determining that the amino acid comprises a phosphorylated threonine (e.g., phospho-threonine). In some embodiments, determining at least one chemical characteristic of the polypeptide comprises determining that the amino acid comprises a phosphorylated tyrosine (e.g., phosphotyrosine). In some embodiments, determining at least one chemical characteristic of the polypeptide comprises determining that the amino acid comprises a phosphorylated serine (e.g., phosphoserine).

[0177] In some embodiments, determining at least one chemical characteristic of the polypeptide includes determining that the amino acids include chemically modified variants, unnatural amino acids, or proteinogenic amino acids such as selenocysteine and pyrrolysine. Examples of unnatural amino acids include, but are not limited to, 2-naphthyl-alanine, statins, homoalanine, α-amino acids, β2-amino acids, β3-amino acids, γ-amino acids, 3-pyridyl-alanine, 4-fluorophenyl-alanine, cyclohexyl-alanine, N-alkyl amino acids, peptoid amino acids, homo-cysteine, penicillamine, 3-nitro-tyrosine, homo-phenyl-alanine, t-leucine, hydroxy-proline, 3-Abz, 5-F-tryptophan, and azabicyclo-[2.2.1]heptane.

[0178] In some embodiments, determining at least one chemical characteristic of the polypeptide includes determining that the amino acid comprises an oxidative modification. For example, as described herein, the amino acid recognition elements of the present disclosure can distinguish between oxidized methionine and its unmodified variant. In some embodiments, the oxidative modification comprises an oxidatively damaged side chain of the amino acid. In some embodiments, the oxidatively damaged side chain is selected from the group consisting of cysteine-derived products (e.g., disulfides, sulfinic acids, sulfonic acids, sulfenic acids, S-nitrosocysteine), tyrosine-derived products (e.g., di-tyrosine, 3,4-dihydroxyphenylalanine, 3-chlorotyrosine, 3-nitrotyrosine), histidine-derived products (e.g., 2-oxohistidine, 4-hydroxy-2-oxohistidine, di-histidine, asparagine, aspartic acid, urea), methionine-derived products, and the like. Oxidatively damaged amino acids include oxidatively damaged amino acids (e.g., sulfoxides, sulfones), tryptophan-derived products (e.g., di-tryptophan, N-formylkynurenine, kynurenine, 2-oxo-tryptophan, oxindolylalanine, 6-nitrotryptophan, hydroxytryptophan), phenylalanine-derived products (e.g., meta-tyrosine, ortho-tyrosine), or common side chain products (e.g., alcohols, hydroperoxides, aldehyde / ketone carbonyls). Examples of oxidatively damaged amino acids are known in the art; see, e.g., Hawkins, CL, Davies, MJ, Detection, identification, and quantification of oxidative protein modifications., J Biol Chem., 2019 Dec. 20, 294(51), 19683-19708.

[0179] In some embodiments, determining at least one chemical characteristic of the polypeptide includes determining that the amino acid comprises a side chain characterized by one or more biochemical properties. For example, the amino acid may comprise a nonpolar aliphatic side chain, a positively charged side chain, a negatively charged side chain, a nonpolar aromatic side chain, or a polar, uncharged side chain. Non-limiting examples of amino acids comprising nonpolar aliphatic side chains include alanine, glycine, valine, leucine, methionine, and isoleucine. Non-limiting examples of amino acids comprising positively charged side chains include lysine, arginine, and histidine. Non-limiting examples of amino acids comprising negatively charged side chains include aspartic acid and glutamic acid. Non-limiting examples of amino acids comprising nonpolar aromatic side chains include phenylalanine, tyrosine, and tryptophan. Non-limiting examples of amino acids comprising polar, uncharged side chains include serine, threonine, cysteine, proline, asparagine, and glutamine.

[0180] In some embodiments, a protein or polypeptide may be digested into multiple smaller polypeptides and chemical characteristics may be determined for one or more of these smaller polypeptides. In some embodiments, a first end (e.g., the N- or C-terminus) of the polypeptide is immobilized and the other end (e.g., the C- or N-terminus) is analyzed as described herein.

[0181] As used herein, sequencing a polypeptide refers to determining sequence information of a polypeptide. In some embodiments, this may include determining the identity of each consecutive amino acid for part (or all) of a polypeptide. However, in some embodiments, this may include assessing the identity of a subset of amino acids within a polypeptide (e.g., determining the relative positions of one or more amino acid types without determining the identity of each amino acid in the polypeptide). However, in some embodiments, amino acid content information can be obtained from a polypeptide without directly determining the relative positions of different types of amino acids in the polypeptide. Amino acid content alone can be used to infer the identity of existing polypeptides (e.g., by comparing amino acid content to a database of polypeptide information and determining which polypeptides have the same amino acid content).

[0182] In some embodiments, sequence information for multiple polypeptide products obtained from a longer polypeptide or protein (e.g., via enzymatic and / or chemical cleavage) can be analyzed to reconstruct or infer the sequence of the longer polypeptide or protein.

[0183] In some embodiments, the polypeptide analysis methods described herein generate data showing how a polypeptide interacts with a binding means while the polypeptide is being degraded by the cleavage means. As described above, the data can include a series of characteristic patterns corresponding to association events at the ends of the polypeptide between the cleavage events at the ends. In some embodiments, the methods of polypeptide analysis described herein include contacting a single polypeptide molecule with a binding means and a cleavage means, wherein the binding means and the cleavage means are configured to achieve at least 10 association events before the cleavage event. In some embodiments, the means are configured to achieve at least 10 association events between two cleavage events.

[0184] In some embodiments, multiple single molecule sequencing reactions are performed in parallel in an array of sample wells. In some embodiments, the array comprises about 10,000 to about 1,000,000 sample wells. The volume of a sample well, in some implementations, is about 10 -21 liters ~ approx. 10 -15 liters. Because the sample wells have small volumes, only about one polypeptide may be present in the sample well at any given time, allowing for detection of single molecule events. Statistically, some sample wells may not contain a single molecule sequencing reaction, and some may contain two or more single polypeptide molecules. However, a significant number of sample wells may each contain a single molecule reaction (e.g., in some embodiments, at least 30%), and thus single molecule analysis may be performed in parallel on a large number of sample wells. In some embodiments, the binding and cleavage means are configured to achieve at least 10 association events before a cleavage event in at least 10% (e.g., 10-50%, more than 50%, 25-75%, at least 80%, or more) of the sample wells in which a single molecule reaction is occurring. In some embodiments, the binding and cleavage means are configured to achieve at least 10 association events before a cleavage event in at least 50% (e.g., more than 50%, 50-75%, at least 80%, or more) of the amino acids of the polypeptide in the single molecule reaction.

[0185] III. Compositions and Reaction Mixtures In some aspects, the present disclosure provides a composition comprising two or more amino acid recognition factors, wherein at least one amino acid recognition factor comprises an amino acid binding protein described herein. In some embodiments, the composition comprises at least one ClpS homologous protein described herein. In some embodiments, the composition comprises at least one UBR homologous protein described herein. In some embodiments, the composition comprises at least one Ntaq1 homologous protein described herein. In some embodiments, the composition comprises two or more of a ClpS homologous protein, a UBR homologous protein, and an Ntaq1 homologous protein. In some embodiments, the composition comprises at least one ClpS homologous protein, at least one UBR homologous protein, and at least one Ntaq1 homologous protein.

[0186] In some embodiments, the composition further comprises at least one type of cleavage reagent. A composition comprising an amino acid recognition factor and a cleavage reagent may be referred to herein as a reaction mixture (e.g., a polypeptide sequencing reaction mixture). Peptidases, also called proteases or proteinases, are enzymes that catalyze the hydrolysis of peptide bonds. Peptidases digest polypeptides into shorter fragments and can generally be classified as endopeptidases and exopeptidases, which cleave polypeptide chains internally and at the terminals, respectively. In some embodiments, the cleavage reagent comprises an exopeptidase (e.g., an aminopeptidase). Examples of suitable peptidases have been described and are contemplated for use in accordance with the present disclosure. See, for example, PCT International Publication No. WO2020102741A1, filed November 15, 2019, and PCT International Publication No. WO2021236983A2, filed May 20, 2021, the relevant contents of each of which are incorporated herein by reference.

[0187] As described herein, the compositions of the present disclosure can be used to determine at least one chemical characteristic of a polypeptide based on a characteristic pattern. In some embodiments, polypeptide sequencing reaction conditions can be configured to achieve a time interval that allows for sufficient related events to provide a desired level of confidence that the characteristic pattern has been identified. This can be achieved by configuring the reaction conditions based on various characteristics, including, for example, reagent concentration, molar ratio of one reagent to another (e.g., ratio of amino acid recognition molecule to cleavage reagent, ratio of one recognition factor to another recognition factor, ratio of one cleavage reagent to another cleavage reagent), number of different reagent types (e.g., number of different types of recognition factors and / or cleavage reagents, number of recognition factor types relative to number of cleavage reagent types), cleavage activity (e.g., peptidase activity), binding characteristics (e.g., kinetic and / or thermodynamic binding parameters for recognition molecule binding), reagent modifications (e.g., polyols and other recognition factor modifications that can alter interaction kinetics), reaction mixture components (e.g., one or more components such as pH, buffers, salts, divalent cations, surfactants, and other reaction mixture components described herein), temperature of the reaction, and various other parameters, as well as combinations thereof, that will be apparent to one of skill in the art. Reaction conditions can be configured based on one or more aspects described herein, including, for example, signal pulse information (e.g., pulse duration, inter-pulse duration, magnitude change), labeling strategy (e.g., number and / or type of fluorophores, linkers with or without shielding elements), surface modification (e.g., modification of sample well surface, including polypeptide immobilization), sample preparation (e.g., polypeptide fragment size, polypeptide modification for immobilization), and other aspects described herein.

[0188] In some embodiments, polypeptide sequencing reactions in accordance with the present application are performed under conditions that allow amino acid recognition and cleavage to occur simultaneously in a single reaction mixture. For example, in some embodiments, polypeptide sequencing reactions are performed in a reaction mixture having a pH that allows association and cleavage events to occur. Thus, in some embodiments, the reaction mixture has a pH of about 6.5 to about 9.0. In some embodiments, the reaction mixture has a pH of about 7.0 to about 8.5 (e.g., about 7.0 to about 8.0, about 7.5 to about 8.5, about 7.5 to about 8.0, or about 8.0 to about 8.5).

[0189] In some embodiments, the polypeptide sequencing reaction is carried out in a reaction mixture comprising one or more buffers. In some embodiments, the reaction mixture comprises a buffer at a concentration of at least 10 mM (e.g., at least 20 mM and up to 250 mM, at least 50 mM, 10-250 mM, 10-100 mM, 20-100 mM, 50-100 mM, or 100-200 mM). In some embodiments, the reaction mixture comprises a buffer at a concentration of about 10 mM to about 50 mM (e.g., about 10 mM to about 25 mM, about 25 mM to about 50 mM, or about 20 mM to about 40 mM). Examples of buffering agents include, but are not limited to, HEPES (4-(2-hydroxyethyl)-1-piperazineethanesulfonic acid), Tris (tris(hydroxymethyl)aminomethane), and MOPS (3-(N-morpholino)propanesulfonic acid).

[0190] In some embodiments, the polypeptide sequencing reaction is carried out in a reaction mixture comprising a salt concentration of at least 10 mM. In some embodiments, the reaction mixture comprises a salt concentration of at least 10 mM (e.g., at least 20 mM, at least 50 mM, at least 100 mM, or more). In some embodiments, the reaction mixture comprises a salt concentration of about 10 mM to about 250 mM (e.g., about 20 mM to about 200 mM, about 50 mM to about 150 mM, about 10 mM to about 50 mM, or about 10 mM to about 100 mM). Examples of salts include, but are not limited to, sodium salts, potassium salts, and acetate salts, such as sodium chloride (NaCl), sodium acetate (NaOAc), and potassium acetate (KOAc).

[0191] Further examples of components for use in the reaction mixture include divalent cations (e.g., Mg 2+ , Co 2+ ) and surfactants (e.g., polysorbate 20). In some embodiments, the reaction mixture contains a divalent cation at a concentration of about 0.1 mM to about 50 mM (e.g., about 10 mM to about 50 mM, about 0.1 mM to about 10 mM, or about 1 mM to about 20 mM). In some embodiments, the reaction mixture contains a surfactant at a concentration of at least 0.01% (e.g., about 0.01% to about 0.10%). In some embodiments, the reaction mixture contains one or more components useful in single molecule analysis, such as an oxygen scavenging system (e.g., a PCA / PCD system or a pyranose oxidase / catalase / glucose system) and / or one or more triplet state quenchers (e.g., Trolox®, COT, and NBA).

[0192] In some embodiments, the polypeptide sequencing reaction is performed at a temperature at which association and cleavage events can occur. In some embodiments, the polypeptide sequencing reaction is performed at a temperature of at least 10°C. In some embodiments, the polypeptide sequencing reaction is performed at a temperature of about 10°C to about 50°C (e.g., 15-45°C, 20-40°C, 25°C or near 25°C, 30°C or near 30°C, 35°C or near 35°C, 37°C or near 37°C). In some embodiments, the polypeptide sequencing reaction is performed at or near room temperature.

[0193] As detailed above, a real-time sequencing process such as that illustrated by FIG. 1 may generally involve cycles of amino acid recognition and terminal amino acid cleavage. In some embodiments, the relative occurrence of recognition and cleavage can be controlled by the concentration difference between one or more amino acid recognition factors and at least one cleavage reagent. In some embodiments, the concentration difference may be optimized so that the number of signal pulses detected during recognition of individual amino acids provides a desired confidence interval for discrimination. For example, if an initial sequencing reaction provides signal data with too few signal pulses between cleavage events to allow for the determination of a characteristic pattern with a desired confidence interval, the sequencing reaction may be repeated using a reduced concentration of non-specific exopeptidase relative to the recognition molecule.

[0194] In some embodiments, polypeptide analysis according to the present disclosure can be performed by contacting a polypeptide with a reaction mixture comprising one or more amino acid recognition factors and one or more cleavage reagents (e.g., peptidases). In some embodiments, the reaction mixture comprises the amino acid recognition factors at a concentration of about 10 nM to about 10 μM. In some embodiments, the reaction mixture comprises the cleavage reagent at a concentration of about 500 nM to about 500 μM.

[0195] In some embodiments, the reaction mixture contains an amino acid recognition factor at a concentration of about 100 nM to about 10 μM, about 250 nM to about 10 μM, about 100 nM to about 1 μM, about 250 nM to about 1 μM, about 250 nM to about 750 nM, or about 500 nM to about 1 μM. In some embodiments, the reaction mixture contains an amino acid recognition factor at a concentration of about 100 nM, about 250 nM, about 500 nM, about 750 nM, or about 1 μM. In some embodiments, the reaction mixture contains a cleavage reagent at a concentration of about 500 nM to about 250 μM, about 500 nM to about 100 μM, about 1 μM to about 100 μM, about 500 nM to about 50 μM, about 1 μM to about 100 μM, about 10 μM to about 200 μM, or about 10 μM to about 100 μM. In some embodiments, the reaction mixture comprises the cleavage reagent at a concentration of about 1 μM, about 5 μM, about 10 μM, about 30 μM, about 50 μM, about 70 μM, or about 100 μM.

[0196] In some embodiments, the reaction mixture comprises an amino acid recognition factor at a concentration of about 10 nM to about 10 μM and a cleavage reagent at a concentration of about 500 nM to about 500 μM. In some embodiments, the reaction mixture comprises an amino acid recognition factor at a concentration of about 100 nM to about 1 μM and a cleavage reagent at a concentration of about 1 μM to about 100 μM. In some embodiments, the reaction mixture comprises an amino acid recognition factor at a concentration of about 250 nM to about 1 μM and a cleavage reagent at a concentration of about 10 μM to about 100 μM. In some embodiments, the reaction mixture comprises an amino acid recognition factor at a concentration of about 500 nM and a cleavage reagent at a concentration of about 25 μM to about 75 μM. In some embodiments, the concentrations of the amino acid recognition factor and / or the cleavage reagent in the reaction mixture are as described elsewhere herein.

[0197] In some embodiments, the reaction mixture comprises the amino acid recognition factor and the cleavage reagent in a molar ratio of about 500:1, about 400:1, about 300:1, about 200:1, about 100:1, about 75:1, about 50:1, about 25:1, about 10:1, about 5:1, about 2:1, or about 1:1. In some embodiments, the reaction mixture comprises the amino acid recognition factor and the cleavage reagent in a molar ratio of about 10:1 to about 200:1. In some embodiments, the reaction mixture comprises the amino acid recognition factor and the cleavage reagent in a molar ratio of about 50:1 to about 150:1. In some embodiments, the molar ratio of amino acid recognition factor to cleavage reagent in the reaction mixture is about 1:1,000 to about 1:1 or about 1:1 to about 100:1 (e.g., 1:1,000, about 1:500, about 1:200, about 1:100, about 1:10, about 1:5, about 1:2, about 1:1, about 5:1, about 10:1, about 50:1, about 100:1). In some embodiments, the molar ratio of amino acid recognition factor to cleavage reagent in the reaction mixture is about 1:100 to about 1:1 or about 1:1 to about 10:1. In some embodiments, the molar ratio of amino acid recognition factor to cleavage reagent in the reaction mixture is as described elsewhere herein.

[0198] In some embodiments, the reaction mixture comprises one or more amino acid recognition factors and one or more cleavage reagents. In some embodiments, the reaction mixture comprises at least three amino acid recognition factors and at least one cleavage reagent. In some embodiments, the reaction mixture comprises two or more cleavage reagents. In some embodiments, the reaction mixture comprises at least one and up to 10 cleavage reagents (e.g., 1-3 cleavage reagents, 2-10 cleavage reagents, 1-5 cleavage reagents, 3-10 cleavage reagents). In some embodiments, the reaction mixture comprises at least 3 and up to 30 amino acid recognition factors (e.g., 3-25, 3-20, 3-10, 3-5, 5-30, 5-20, 5-10, or 10-20 amino acid recognition factors). In some embodiments, the one or more amino acid recognition factors comprise at least one amino acid binding protein selected from Table 1.

[0199] In some embodiments, the reaction mixture comprises two or more amino acid recognition factors and / or two or more cleavage reagents. In some embodiments, a reaction mixture described as comprising two or more amino acid recognition factors (or cleavage reagents) refers to a mixture having two or more types of amino acid recognition factors (or cleavage reagents). For example, in some embodiments, the reaction mixture comprises two or more amino acid binding proteins, and two or more amino acid binding proteins refer to two or more types of amino acid binding proteins. In some embodiments, one type of amino acid binding protein has an amino acid sequence that differs from that of another type of amino acid binding protein in the reaction mixture. In some embodiments, one type of amino acid binding protein associates with (e.g., binds to) amino acids that differ from the amino acids associated with another type of amino acid binding protein in the reaction mixture. In some embodiments, one type of amino acid binding protein associates with (e.g., binds to) a subset of amino acids that differs from the subset of amino acids associated with another type of amino acid binding protein in the reaction mixture.

[0200] IV. Devices and Systems In some embodiments, methods according to the present disclosure can be implemented using a system that enables single-molecule analysis. The system can include an integrated device and an instrument configured to interface with the integrated device. The integrated device can include an array of pixels, each pixel including a sample well and at least one photodetector. The sample wells of the integrated device can be formed on or through a surface of the integrated device and configured to receive a sample disposed on the surface of the integrated device. Collectively, the sample wells can be considered an array of sample wells. The multiple sample wells can have a size and shape suitable for at least a portion of the sample wells to receive a single sample (e.g., a single molecule such as a polypeptide). In some embodiments, the number of samples in the sample wells can be distributed among the sample wells of the integrated device such that some sample wells contain one sample while other sample wells contain zero, two, or more samples.

[0201] Excitation light is provided to the integrated device from one or more light sources external to the integrated device. Optical components of the integrated device can receive the excitation light from the light source and direct the light toward the array of sample wells in the integrated device, illuminating an illumination region within the sample well. In some embodiments, the sample wells can have a configuration that allows a sample to be held in close proximity to the surface of the sample well, which can facilitate delivery of excitation light to the sample and detection of emission light from the sample. A sample positioned within the illumination region can emit emission light in response to being illuminated by the excitation light. For example, the sample can be labeled with a fluorescent label that emits light in response to achieving an excited state through illumination with the excitation light. The emission light emitted by the sample can then be detected by one or more photodetectors in pixels corresponding to the sample well with the sample being analyzed. According to some embodiments, multiple samples can be analyzed in parallel when performed across the entire array of sample wells, which can range in number from approximately 10,000 pixels to 1,000,000 pixels.

[0202] The integrated device may include an optical system for receiving excitation light and directing the excitation light among the sample well array. The optical system may include one or more grating couplers configured to couple the excitation light to other optical components of the integrated device and direct the excitation light to other optical components. For example, the optical system may include optical components that direct the excitation light from the grating coupler toward the sample well array. Such optical components may include an optical splitter, an optical combiner, and a waveguide. In some embodiments, one or more optical splitters can couple the excitation light from the grating coupler and deliver the excitation light to at least one waveguide. According to some embodiments, the optical splitter can have a configuration that enables substantially uniform delivery of excitation light across all waveguides, such that each of the waveguides receives substantially the same amount of excitation light. Such embodiments may improve the performance of the integrated device by improving the uniformity of the excitation light received by the sample wells of the integrated device. For example, examples of suitable components for coupling excitation light into a sample well and / or directing emitted light to a photodetector and for inclusion within an integrated device are described in U.S. patent application Ser. No. 14 / 821,688, filed Aug. 7, 2015, entitled "INTEGRATED DEVICE FOR PROBING, DETECTING AND ANALYZING MOLECULES," and U.S. patent application Ser. No. 14 / 543,865, filed Nov. 17, 2014, entitled "INTEGRATED DEVICE WITH EXTERNAL LIGHT SOURCE FOR PROBING, DETECTING, AND ANALYZING MOLECULES," both of which are incorporated by reference in their entireties. Examples of suitable grating couplers and waveguides that can be implemented in an integrated device are described in U.S. patent application Ser. No. 15 / 844,403, filed Dec. 15, 2017, entitled "OPTICAL COUPLER AND WAVEGUIDE SYSTEM," which is incorporated by reference in its entirety.

[0203] Additional photonic structures may be disposed between the sample well and the photodetector and configured to reduce or prevent excitation light from reaching the photodetector, which might otherwise contribute to signal noise when detecting the emission light. In some embodiments, the metal layer, which may act as a circuit for the integrated device, may also act as a spatial filter. Examples of suitable photonic structures may include spectral filters, polarization filters, and spatial filters, and are described in U.S. Patent Application No. 16 / 042,968, filed July 23, 2018, entitled "OPTICAL REJECTION PHOTONIC STRUCTURES," and U.S. Provisional Patent Application No. 63 / 124,655, filed December 11, 2020, entitled "INTEGRATED CIRCUIT WITH IMPROVED CHARGE TRANSFER EFFICIENCY AND ASSOCIATED TECHNIQUES," both of which are incorporated by reference in their entireties.

[0204] Components located remotely from the integrated device can be used to position and align the excitation source relative to the integrated device. Such components can include optical components, including lenses, mirrors, prisms, windows, apertures, attenuators, and / or optical fibers. Additional mechanical components can be included in the instrument to enable control of one or more alignment components. Such mechanical components can include actuators, stepper motors, and / or knobs. Examples of suitable excitation sources and alignment mechanisms are described in U.S. patent application Ser. No. 15 / 161,088, filed May 20, 2016, entitled "PULSED LASER AND SYSTEM," which is incorporated by reference in its entirety. Another example of a beam steering module is described in U.S. patent application Ser. No. 15 / 842,720, filed December 14, 2017, entitled "COMPACT BEAM SHAPING AND STEERING ASSEMBLY," which is incorporated by reference herein. Further examples of suitable excitation sources are described in U.S. Patent Application No. 14 / 821,688, filed August 7, 2015, entitled "INTEGRATED DEVICE FOR PROBING, DETECTING AND ANALYZING MOLECULES," which is incorporated by reference in its entirety.

[0205] The photodetector(s) associated with each pixel of the integrated device may be configured and arranged to detect light emission from the pixel's corresponding sample well. Examples of suitable photodetectors are described in U.S. Patent Application No. 14 / 821,656, filed August 7, 2015, entitled "INTEGRATED DEVICE FOR TEMPORAL BINNING OF RECEIVED PHOTONS," which is incorporated by reference in its entirety. In some embodiments, the sample wells and their respective photodetectors may be aligned along a common axis. In this manner, the photodetectors may overlap the sample wells within the pixel.

[0206] Characteristics of the detected emitted light can provide an indication for identifying a label associated with the emitted light. Such characteristics can include any suitable type of characteristic, including the arrival time of a photon detected by a photodetector, the amount of photons accumulated over time by a photodetector, and / or the distribution of photons across two or more photodetectors. In some embodiments, such characteristics can be any one or a combination of two or more of the following: luminescence lifetime, luminescence intensity, luminance, absorption spectrum, emission spectrum, luminescence quantum yield, wavelength (e.g., peak wavelength), and signal characteristics (e.g., pulse duration, inter-pulse duration, change in signal intensity).

[0207] In some embodiments, the photodetector may have a configuration that allows for detection of one or more timing characteristics (e.g., luminescence lifetime) associated with the sample's emission. The photodetector may detect a distribution of photon arrival times after a pulse of excitation light propagates through the integrated device, and the distribution of arrival times may provide an indication of the timing characteristics of the sample's emission light (e.g., a proxy for the luminescence lifetime). In some embodiments, one or more photodetectors provide an indication of the probability (e.g., luminescence intensity) of emission light emitted by a label. In some embodiments, multiple photodetectors may be sized and positioned to capture the spatial distribution of emission light. The output signal from the one or more photodetectors may then be used to distinguish one label from multiple labels, and multiple labels may be used to identify a sample within a sample. In some embodiments, the sample may be excited by multiple excitation energies, and the emission light and / or timing characteristics of the emission light emitted by the sample in response to the multiple excitation energies can distinguish one label from multiple labels.

[0208] In operation, parallel analysis of samples in the sample wells is performed by exciting some or all of the samples in the wells using excitation light and detecting signals from the sample emissions using photodetectors. The emitted light from the samples can be detected by corresponding photodetectors and converted into at least one electrical signal. The electrical signal can be transmitted along conductive lines in the circuitry of the integrated device, which can be connected to an instrument interfaced with the integrated device. The electrical signal can then be processed and / or analyzed. The processing or analysis of the electrical signal can be performed on a suitable computing device located either on or off the instrument.

[0209] The device may include a user interface for controlling the operation of the device and / or the integrated device. The user interface may be configured to allow a user to input information into the device, such as commands and / or settings used to control the device's functions. In some embodiments, the user interface may include buttons, switches, dials, and a microphone for voice commands. The user interface may allow a user to receive feedback regarding the device and / or the performance of the integrated device, such as information obtained by proper alignment and / or readout signals from a photodetector on the integrated device. In some embodiments, the user interface may provide feedback using a speaker to provide audible feedback. In some embodiments, the user interface may include indicator lights and / or a display screen to provide visual feedback to the user.

[0210] In some embodiments, the instrument may include a computer interface configured to connect to a computing device. The computer interface may be a USB interface, a FireWire interface, or any other suitable computer interface. The computing device may be any general-purpose computer, such as a laptop or desktop computer. In some embodiments, the computing device may be a server (e.g., a cloud-based server) accessible over a wireless network via a suitable computer interface. The computer interface may facilitate communication of information between the instrument and the computing device. Input information for controlling and / or configuring the instrument may be provided to the computing device and transmitted to the instrument via the computer interface. Output information generated by the instrument may be received by the computing device via the computer interface. The output information may include feedback regarding instrument performance, integrated device performance, and / or data generated from the photodetector readout signal.

[0211] In some embodiments, the instrument may include a processing device configured to analyze data received from one or more photodetectors of the integrated device and / or transmit control signals to the excitation source. In some embodiments, the processing device may comprise a general-purpose processor, a specially adapted processor (e.g., one or more central processing units (CPUs), such as microprocessors or microcontroller cores, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), custom integrated circuits, digital signal processors (DSPs), or combinations thereof). In some embodiments, processing of data from the one or more photodetectors may be performed by both the instrument's processing device and an external computing device. In other embodiments, the external computing device may be omitted, and processing of data from the one or more photodetectors may be performed solely by the integrated device's processing device.

[0212] According to some embodiments, an instrument configured to analyze a sample based on its luminescence emission characteristics may detect differences in luminescence lifetimes and / or intensities between different luminescent molecules and / or differences in lifetimes and / or intensities of the same luminescent molecule in different environments. The inventors have recognized and appreciated that differences in luminescence emission lifetimes can be used to distinguish between the presence or absence of different luminescent molecules and / or to distinguish between different environments or conditions to which the luminescent molecules are exposed. In some cases, aspects of the system can be simplified by distinguishing luminescent molecules based on lifetime (e.g., rather than emission wavelength). As an example, wavelength-discriminating optics (e.g., wavelength filters, dedicated detectors for each wavelength, dedicated pulsed light sources at different wavelengths, and / or diffractive optics) may be reduced in number or eliminated when distinguishing luminescent molecules based on lifetime. In some cases, a single pulsed light source operating at a single characteristic wavelength can be used to excite different luminescent molecules that emit within the same wavelength region of the optical spectrum but have measurably different lifetimes. Analysis systems that use a single pulsed light source, rather than multiple light sources operating at different wavelengths, to excite and distinguish different luminescent molecules that emit in the same wavelength range can be less complex to operate and maintain, more compact, and can be manufactured at lower cost.

[0213] While analytical systems based on luminescence lifetime analysis may have certain advantages, the amount of information obtained by the analytical system and / or the detection accuracy may be increased by enabling additional detection techniques. For example, some embodiments of the system may be further configured to identify one or more characteristics of a sample based on emission wavelength and / or emission intensity. In some implementations, emission intensity may additionally or alternatively be used to distinguish between different luminescent labels. For example, some luminescent labels may emit at significantly different intensities or have significant differences in excitation probability (e.g., at least about 35% difference), even if their decay rates are similar. By referencing binned signals to the measured excitation light, it may be possible to distinguish between different luminescent labels based on intensity levels.

[0214] According to some embodiments, different luminescence lifetimes can be distinguished by a photodetector configured to time-bin luminescence emission events following excitation of the luminescent labels. Time binning can occur during a single charge accumulation cycle for the photodetector. A charge accumulation cycle is the interval between readout events during which photogenerated carriers accumulate in bins of the time-binning photodetector. Examples of time-binning photodetectors are described in U.S. Patent Application No. 14 / 821,656, filed August 7, 2015, entitled "INTEGRATED DEVICE FOR TEMPORAL BINNING OF RECEIVED PHOTONS," which is incorporated herein by reference. In some embodiments, the time-binning photodetector can generate charge carriers in a photon absorption / carrier generation region and directly transfer the charge carriers to charge carrier storage bins in a charge carrier storage region. In such embodiments, the time-binning photodetector may not include a carrier transfer / capture region. Such time-binning photodetectors are sometimes referred to as "direct binning pixels." An example of a time-binning photodetector including a direct binning pixel is described in U.S. Patent Application No. 15 / 852,571, filed December 22, 2017, entitled "INTEGRATED PHOTODETECTOR WITH DIRECT BINNING PIXEL," which is incorporated herein by reference.

[0215] In some embodiments, different numbers of fluorophores of the same type can be attached to different reagents in a sample, so that each reagent can be identified based on its emission intensity. For example, two fluorophores can be attached to a first labeled recognition molecule, and four or more fluorophores can be attached to a second labeled recognition molecule. Due to the different numbers of fluorophores, there can be different excitation and emission probabilities associated with different recognition molecules. For example, there can be more emission events for the second labeled recognition molecule during a signal accumulation interval, so that the apparent intensity of the bin is significantly higher than that of the first labeled recognition molecule.

[0216] The inventors of the present application have recognized and appreciated that distinguishing biological or chemical samples based on fluorophore decay rates and / or fluorophore intensities can allow for simplification of optical excitation and detection systems. For example, optical excitation can be performed using a single wavelength source (e.g., a light source that produces one characteristic wavelength rather than multiple light sources, or a light source that operates at multiple different characteristic wavelengths). Furthermore, wavelength-discriminating optics and filters may not be required in the detection system. Also, a single photodetector may be used for each sample well to detect emissions from different fluorophores. The phrase "characteristic wavelength" or "wavelength" is used to refer to the central or dominant wavelength within a limited emission bandwidth (e.g., the central or peak wavelength within a 20 nm bandwidth output by a pulsed light source). In some cases, "characteristic wavelength" or "wavelength" can be used to refer to the peak wavelength within the full bandwidth of the emission output by a source.

[0217] According to one aspect of the present disclosure, an exemplary integrated device may be configured to perform single molecule analysis in combination with the above-described instruments. It should be understood that the exemplary integrated device described herein is intended to be exemplary, and that other integrated device configurations may be configured to perform any or all of the techniques described herein.

[0218] 29 shows a cross-sectional view of pixel 1-112 of integrated device 1-102. Pixel 1-112 includes a photodetection region, which may be a pinned photodiode (PPD), and a charge storage region, which may be a storage diode (SDO). In some embodiments, the photodetection region and the charge storage region may be formed within the semiconductor material of the pixel by doping regions of the semiconductor material. For example, the photodetection region and the charge storage region may be formed using the same conductivity type (e.g., n-type doping or p-type doping).

[0219] During operation of the pixel 1-112, excitation light can illuminate the sample well 1-108, causing incident photons, including fluorescent emission from the sample, to flow along the optical axis to the photodetection region PPD. As shown in FIG. 29, the pixel 1-112 can include a waveguide 1-220 configured to optically couple (e.g., by evanescent coupling) excitation light from a grating coupler of an integrated device (not shown) into the sample well 1-108. In response, the sample in the sample well 1-108 can emit fluorescent light toward the photodetection region PPD. In some embodiments, the pixel 1-112 can also include one or more photonic structures 1-230, which can include one or more optical rejection structures, such as a spectral filter, a polarizing filter, and / or a spatial filter. For example, the photonic structures 1-230 can be configured to reduce the amount of excitation light reaching the photodetection region PPD and / or increase the amount of fluorescent emission reaching the photodetection region PPD. Also, as shown in pixel 1-112, pixel 1-112 may include one or more metal layers 1-240, which may be configured as filters and / or carry control signals from control circuitry configured to control the transfer gates, as further described herein.

[0220] In some embodiments, pixel 1-112 may include one or more transfer gates configured to control operation of pixel 1-112 by applying an electrical bias to one or more semiconductor regions of pixel 1-112 in response to one or more control signals. For example, when transfer gate ST0 induces a first electrical bias in the semiconductor region between photodetection region PPD and storage region SD0, a transfer path (e.g., a charge transfer channel) may be formed in the semiconductor region. Charge carriers (e.g., photoelectrons) generated in photodetection region PPD by incident photons flow along the transfer path to storage region SD0. In some embodiments, the first electrical bias may be applied during a collection period during which charge carriers from the sample are selectively directed to storage region SD0. Alternatively, when transfer gate ST0 imparts a second electrical bias to the semiconductor region between photodetection region PPD and storage region SD0, charge carriers from photodetection region PPD may be blocked from reaching storage region SD0 along the transfer path. In some embodiments, the drain gate REJ can provide a channel to the drain D to draw noise charge carriers generated in the photodetection region PPD by the excitation light away from the photodetection region PPD and the storage region SD0, such as during a rejection period before fluorescent emission photons from the sample reach the photodetection region PPD. In some embodiments, during a readout period, the transfer gate ST0 can provide a second electrical bias, and the transfer gate TX0 can provide an electrical bias to flow the charge carriers stored in the storage region SD0 to the readout region, which may be a floating diffusion (FD) region, for processing.

[0221] It should be understood that, according to various embodiments, the transfer gates described herein may comprise semiconductor materials and / or metals, and may include the gate of a field effect transistor (FET), the base of a bipolar junction transistor (BJT), etc.

[0222] In some embodiments, operation of the pixels 1-112 may include one or more collection sequences, each of which includes one or more rejection (e.g., drain) periods and one or more collection periods. In one example, a collection sequence performed in accordance with one or more pulses of an excitation light source can begin with a rejection period, such as to discard charge carriers generated within the pixels 1-112 (e.g., within the photodetection region PD) in response to excitation photons from the light source. For example, excitation photons may arrive at the pixels 1-112 prior to the arrival of fluorescent emission photons from the sample well. A transfer gate for the charge storage region can be biased to have low conductivity in a charge transfer channel connecting the charge storage region to the photodetection region, preventing the transfer and accumulation of charge carriers in the charge storage region. A drain gate for the drain region can be biased to have high conductivity in a drain channel between the photodetection region and the drain region, facilitating the draining of charge carriers from the photodetection region to the drain region. The transfer gate for any charge storage region coupled to the photodetector region may be biased to have low conductivity between the photodetector region and the charge storage region, thereby preventing charge carriers from being transferred or stored in the charge storage region during the rejection period.

[0223] The rejection period may be followed by a collection period, during which charge carriers generated in response to incident photons are transferred to one or more charge accumulation regions. During the collection period, incident photons may include fluorescent emission photons, resulting in the accumulation of fluorescent emission charge carriers in the charge accumulation region(s). For example, a transfer gate for one of the charge accumulation regions may be biased to have high conductivity between the photodetection region and the charge accumulation region, facilitating the accumulation of charge carriers in the charge accumulation region. Any drain gate coupled to the photodetection region may be biased to have low conductivity between the photodetection region and the drain region to prevent charge carriers from being discarded during the collection period.

[0224] Some embodiments may include multiple rejection and / or collection periods in a collection sequence, such as a first rejection and collection period followed by a second rejection and collection period, with each pair of rejection and collection periods occurring in response to a pulse of excitation light. In one example, charge carriers generated in the photodetection region during each collection period of a collection sequence (e.g., in response to multiple pulses of excitation light) may be collected in a single charge storage region. In some embodiments, the charge carriers collected in the charge storage region may be read out for processing before the next collection sequence. Alternatively, or additionally, in some embodiments, the charge carriers collected in the first charge storage region during the first collection sequence may be transferred to a second charge storage region sequentially coupled to the first charge storage region and read out simultaneously with the next collection sequence. In some embodiments, the processing circuitry configured to read out the charge carriers from one or more pixels may be configured to determine one or more of luminescence intensity information, luminescence lifetime information, luminescence spectrum information, and / or any other mode of luminescence information associated with performing the techniques described herein.

[0225] In some embodiments, the first collection sequence may include transferring charge carriers generated in the light detection response to the excitation pulse to the charge accumulation region at a first time following each excitation pulse, and the second collection sequence may include transferring charge carriers generated in the light detection response to the excitation pulse to the charge accumulation region at a second time following each excitation pulse. For example, the number of charge carriers collected after the first and second times can indicate luminance lifetime information of the received light.

[0226] As described further herein, the pixels of the integrated circuit can be controlled to perform one or more collection sequences using one or more control signals from the integrated circuit's control circuitry, for example, by providing control signals to the drains and / or transfer gates of the pixels of the integrated circuit. In some embodiments, charge carriers can be read out from the FD region of each pixel for processing during a readout pixel associated with each pixel and / or row or column of pixels. In some embodiments, the FD region of the pixel can be read out using a correlated double sampling (CDS) technique.

[0227] V. Sequence Information Table 1: Non-limiting example sequences of amino acid binding proteins

[0228] [Table 1-1]

[0229] [Table 1-2]

[0230] [Table 1-3]

[0231] [Table 1-4]

[0232] [Table 1-5]

[0233] [Table 1-6]

[0234] [Table 1-7]

[0235]

Table 1-8

[0236]

Table 1-9

[0237]

Table 1-10

[0238]

Table 1-11

[0239]

Table 1-12

[0240]

Table 1-13

[0241]

Table 1-14

[0242]

Table 1-15

[0243]

Table 1-16

[0244]

Table 1-17

[0245]

Table 1-18

[0246]

Table 1-19

[0247]

Table 1-20

[0248]

Table 1-21

[0249]

Table 1-22

[0250]

Table 1-23

[0251]

Table 1-24

[0252]

Table 1-25

[0253]

Table 1-26

[0254]

Table 1-27

[0255]

Table 1-28

[0256]

Table 1-29

[0257]

Table 1-30

[0258]

Table 1-31

[0259]

Table 1-32

[0260]

Table 1-33

[0261]

Table 1-34

[0262]

Table 1-35

[0263]

Table 1-36

[0264]

Table 1-37

[0265]

Table 1-38

[0266]

Table 1-39

[0267]

Table 1-40

[0268]

Table 1-41

[0269]

Table 1-42

[0270]

Table 1-43

[0271]

Table 1-44

[0272]

Table 1-45

[0273]

Table 1-46

[0274]

Table 1-47

[0275] [Table 1-48]

[0276] Table 2: Non-limiting example sequences of tag sequences

[0277] [Table 2] [Example]

[0278] Example 1. Real-time dynamic single-molecule protein sequencing on an integrated semiconductor device In this example, we demonstrate a novel method consisting of a dynamic sequencing-by-degradation approach in which single surface-immobilized peptide molecules are probed in real time with a mixture of dye-labeled N-terminal amino acid recognition factors. By measuring the fluorescence intensity, lifetime, and intermolecular dynamics of the recognition factors on a novel semiconductor chip, we demonstrate the ability to annotate amino acids and collectively identify peptide sequences. By exploiting binding kinetics, each recognition factor can uniquely identify multiple amino acids. Principles and processes for expanding the number of recognizable amino acids are also described herein. Furthermore, we demonstrate that this method is compatible with both synthetic peptides and native peptides isolated from recombinant human proteins, capable of detecting single amino acid changes and post-translational modifications. The results demonstrate a robust core technology that can serve as a highly accurate, sensitive, and scalable next-generation sequencing platform for proteins.

[0279] Measurement of the proteome provides profound and valuable insights into biological processes. However, to fully understand the complex and dynamic state of the proteome in cells and the proteomic changes that occur in disease states, and to make this information more accessible, more sensitive methods are needed. The complex nature of the proteome and the chemical characteristics of proteins present several fundamental challenges to achieving the same comprehensive sensitivity, throughput, and adoption as DNA sequencing technologies. These challenges include the large number of different proteins (>10,000) and even larger number of proteoforms per cell; the extremely wide dynamic range of protein abundance in cells and biofluids and the lack of correlation with transcript levels; and the limited availability of current mass spectrometry methods. 2 These limitations include the high cost and high detection limits of sequencing single protein molecules; and the inability to copy or amplify proteins. Methods that directly sequence single protein molecules have the potential to offer the greatest possible detection sensitivity, enable single-cell input, digital quantification based on read counts, detection of post-translational modifications (PTMs) and low-abundance or aberrant proteoforms, and allow for cost and throughput levels favorable for widespread adoption.

[0280] Here, we demonstrate a novel single-molecule protein sequencing method and integrated system for massively parallel proteomics studies. In this method, peptides are immobilized in nanoscale reaction chambers on a semiconductor chip, and N-terminal amino acids (NAAs) are detected in real time with dye-labeled NAA recognition factors. Aminopeptidases sequentially remove individual NAAs, exposing subsequent amino acids for recognition, eliminating the need for complex chemical and fluidic engineering (Figure 1). We constructed a benchtop device with a 532 nm pulsed laser source for fluorescence excitation and electronics for signal processing (Figure 6A). The semiconductor chip uses intensity and fluorescence lifetime, rather than emission wavelength, to distinguish dye labels. The recognition factors detect one or more types of NAAs, and the temporal order of NAA recognition and on-off binding kinetics provide information for peptide identification.

[0281] Using CMOS fabrication technology, we constructed a custom time-domain-sensitive semiconductor chip with nanosecond precision, featuring fully integrated components for single-molecule detection, including a photosensor, optical waveguide circuitry, and a reaction chamber for biomolecule immobilization (Figure 1). An observation volume of less than 5 attoliters was achieved through evanescent illumination at the bottom of the reaction chamber from a nearby waveguide, enabling highly sensitive single-molecule detection in the presence of high freely diffusing dye concentrations (>1 μM).

[0282] The semiconductor chip uses a novel filterless system that rejects excitation light based on photon arrival time, achieving over 10,000-fold attenuation of the incident excitation light. Eliminating the need for an integrated optical filter layer improves the efficiency of fluorescence collection and enables scalable manufacturing of the chip. To enable discrimination of fluorescent dye labels attached to NAA recognition agents by their fluorescence lifetime and intensity, the chip rapidly alternates between early and late signal collection windows associated with each laser pulse, thereby collecting different portions of the exponential fluorescence lifetime decay curve. The relative signals in these collection windows (called "bin ratios") provide a reliable indicator of fluorescence lifetime (Figures 6B-6F and Materials and Methods).

[0283] For an NAA-binding protein to function as a recognition factor in this approach, the mean lifetime of the bound recognition factor-peptide complex must be long enough (typically >120 ms) to generate detectable single-molecule binding events. We evaluated proteins from the N-end rule adaptor family ClpS, which naturally bind to N-terminal phenylalanine, tyrosine, and tryptophan. Using PS610, a recognition factor derived from ClpS2 from A. tumefaciens, we confirmed that this recognition factor detectably binds to immobilized peptides bearing these NAAs. Importantly, we also determined that the binding kinetics differed for each NAA. To demonstrate these properties, immobilized peptides containing the initial N-terminal sequence FAA, YAA, or WAA were incubated with PS610 on separate chips, and data were collected for 10 hours (Methods). NAA recognition, observed by PS610, was characterized by continuous on-off binding during the incubation period, with different pulse durations (PDs) for each peptide (Figure 2A). The median PD values were 2.51, 0.73, and 0.31 seconds for FAA, YAA, and WAA, respectively. These values were consistent with the mean PD values for each type of protein-NAA interaction. 7 This reflects differences in binding affinity driven by different dissociation rates for α, β, and β (Figures 7A-7B).

[0284] To expand the set of recognizable NAAs, we investigated N-end rule pathway proteins as a source of additional recognition factors. In a comprehensive screen of diverse ClpS family proteins, we discovered a novel group of ClpS proteins from the bacterial phylum Planctomycetes that natively bind to N-terminal leucine, isoleucine, and valine. Applying directed evolution techniques, we generated a Planctomycetes ClpS mutant, PS961, with submicromolar affinity for N-terminal leucine, isoleucine, and valine, demonstrating recognition of these NAAs (Figure 2B). The median PDs for binding to peptides with N-terminal LAA, IAA, and VAA were 1.21, 0.28, and 0.21 s, respectively, consistent with bulk characterization (Figure 7C).

[0285] In another screen, we investigated a diverse set of UBR box domains from the UBR family of ubiquitin ligases, which naturally bind to N-terminal arginine, lysine, and histidine. The UBR box domain from the yeast K. lactis UBR1 protein showed the highest affinity for N-terminal arginine, and we used this protein to generate the arginine recognition factor PS691. PS691 recognized arginine in peptides with an N-terminal RLA, with a median PD of 0.23 seconds (Figure 2C). Its low affinity binding to N-terminal lysine and histidine (Figures 7D-7E) was insufficient for single-molecule detection.

[0286] To demonstrate that amino acids in a single peptide molecule can be sequentially exposed by aminopeptidases and recognized in real time with distinguishable kinetics, an immobilized peptide containing the initial sequence FAAWAAYAA (SEQ ID NO: 1073) was incubated with PS610 for 15 min, followed by the addition of PhTET3 (an aminopeptidase from P. horikoshii). The collected trace consisted of distinct regions of pulsing, termed recognition segments (RS), separated by regions lacking recognition pulsing (non-recognition segments, NRS). Analysis software was developed to automatically identify pulsing regions and transition points within the trace (Methods). The trace began with recognition of phenylalanine, with a median PD of 2.36 s (Figure 2D), consistent with the PD observed for FAA in the recognition-only assay. This pattern terminated after aminopeptidase addition (average 11 min after addition) and was followed by the orderly appearance of two RSs with median PDs of 0.25 s and 0.49 s, corresponding to the short and medium PDs obtained in the YAA and WAA recognition-only assays (Figure 2D). Thus, the introduction of aminopeptidase activity into the reaction resulted in the sequential appearance of distinct RSs with the expected kinetic properties in the correct order.

[0287] To demonstrate dynamic sequencing using two NAA recognition factors, PS610 and PS961 were labeled with the distinguishable dyes atto-Rho6G and Cy3, respectively, and an immobilized peptide with the sequence LAQFASIAAYASDDD (SEQ ID NO: 1035) was exposed to a solution containing both recognition factors. After 15 min, two P. horikoshii aminopeptidases, PhTET2 and PhTET3, with complementary activities covering all 20 amino acids, were added. The collected traces showed distinct segments alternating between PS961 and PS610 pulsing according to the order of the recognizable amino acids in the peptide sequence (Figure 2E). The average bin ratios and average PDs associated with each RS readily distinguished the two dye labels and the four types of recognized NAAs (Figure 2F). The median PDs were 2.70, 1.43, 0.25, and 0.66 s for the N-terminal LAQ, FAS, IAA, and YAS, respectively (Figure 2G).

[0288] Previous studies have shown that NAA-binding ClpS and UBR proteins also contact residues at positions two (P2) and three (P3) from the N-terminus, which influence binding affinity. These effects are reflected in modulation of PD, dependent on downstream P2 and P3 residues, as observed above for LAA (1.21 s) compared with LAQ (2.70 s). We found that these effects on PD vary within informatively favorable ranges and can be empirically determined or approximated in silico to a priori model peptide sequencing behavior (Figures 7F–H). A powerful feature of this recognition behavior for peptide identification is that each RS contains information about potential downstream P2 and P3 residues or PTMs, regardless of whether these positions are targets of the NAA recognition factor.

[0289] To evaluate the kinetic principles of dynamic sequencing when applied to diverse sequences, we first characterized the synthetic peptide DQQRLIFAG (SEQ ID NO: 1036), corresponding to a segment of human ubiquitin (Figure 3A-D). Sequencing reactions were performed using three differentially labeled recognition factors, PS610, PS961, and PS691, in combination with two aminopeptidases, PhTET2 and PhTET3 (Materials and Methods). The example trace in Figure 3A begins with an NRS corresponding to the time interval in which residues in the first DQQ motif are present at the N-terminus. The first RS begins at 120 min, when the N-terminal arginine is exposed for recognition by PS691. Subsequent cleavage events sequentially expose the N-terminal leucine, isoleucine, and phenylalanine to their corresponding recognition factors, with rapid transitions (average <10 s) from one RS to the next. The transition from leucine to isoleucine recognition by PS961 is easily identified as an abrupt change in the average PD. This overall pattern is reproduced across many instances of sequencing the same peptide, with similar PD statistics across traces, as each peptide molecule follows the same reaction pathway over the course of a sequencing run (Figure 3B-3C). Due to the stochastic timing of cleavage events, each trace exhibits distinct onset and duration times for each RS (Figure 3C).

[0290] This technique reports the binding kinetics at each recognizable amino acid position and the kinetics of aminopeptidase cleavage along the peptide sequence. Each RS typically contains tens to hundreds of on-off binding events, resulting in a distribution of PD and inter-pulse duration (IPD) measurements that can be statistically analyzed, allowing highly accurate binding kinetic information to be obtained from a single trace. Because calls are not based on the error-prone detection of a single event associated with a single fluorophore molecule, repeated probing of each NAA also provides highly accurate recognition factor calls (Figure 6F). The recognition factor concentration determines the IPD of each RS. Higher recognition factor concentrations result in shorter average IPDs and faster pulse rates (Figures 8A-8B). However, higher recognition factor concentrations can increase fluorescence background from freely diffusing recognition factors, resulting in lower pulse signal-to-noise and potentially competing with aminopeptidases for N-terminal access. In practice, an IPD in the range of approximately 2-10 seconds provides a favorable balance between these factors.

[0291] The distribution of RS durations across the ensemble of replicate traces defines the cleavage rate of each recognizable NAA. For the DQQRLIFAG (SEQ ID NO: 1036) peptide, mean cleavage times of 31, 54, 39, and 86 min were observed for the N-terminal arginine, leucine, isoleucine, and phenylalanine, respectively, with approximate monoexponential decay statistics for each position (Figure 3D, Figure 8C). The distribution of NRS durations reports the cleavage rate of one or more unrecognized NAA runs. The mean NRS duration for the first DQQ motif was 153 min (Figure 3D). The mean cleavage rate is a critical parameter, controlled by the aminopeptidase concentration in the assay (Figure 8D-8E). Considering the exponential behavior, a mean RS duration of 10–40 min was targeted to provide sufficient time for pulse data collection, avoid RS loss due to rapid cleavage, and minimize excessively long RS durations. We found it useful to visualize peptide sequencing profiles as kinetic signature plots, simplified trace-like representations of the complete peptide sequencing time course, including the median PD for each RS and the mean duration of each RS and NRS (Figure 3E). These highly distinctive features provide rich sequence-dependent information for mapping peptide-derived traces to their proteins of origin.

[0292] To demonstrate that this core methodology and its kinetic principles apply to a wide range of peptide sequences, the synthetic peptides DQQIASSRLAASFAAQQYPDDD (SEQ ID NO: 1037), RLAFSALGAADDD (SEQ ID NO: 1038), and EFIAWLV (SEQ ID NO: 1039) (a segment of human GLP-1) were sequenced under the same sequencing conditions used for DQQRLIFAG (SEQ ID NO: 1036) (Figure 3F). Each peptide generated a distinctive kinetic signature according to its sequence (Figure 3G). Readouts were obtained up to position 18 (the furthest recognizable amino acid residue) of peptide DQQIASSRLAASFAAQQYPDDD (SEQ ID NO: 1037), demonstrating that this method is compatible with long peptides and can deeply access sequence information for peptides of the length found in typical protein digests.

[0293] To demonstrate the sensitivity of the kinetic parameters obtained from sequencing to changes in sequence composition, we performed sequencing using a set of three peptides: RLAFAYPDDD (SEQ ID NO: 1040), RLIFAYPDDD (SEQ ID NO: 1041), and RLVFAYPDDD (SEQ ID NO: 1042), which differ only at a single position located directly downstream of the PS961 N-terminal target leucine. Each type of amino acid at this position had a different effect on the PD obtained during recognition of the N-terminal leucine by PS961. Median PDs of 1.29 s, 2.22 s, and 4.21 s were observed for LAF, LIF, and LVF, respectively (Figure 4B). In addition to the differences in PD for leucine, each peptide exhibited a characteristic RS or NRS in the interval between leucine and phenylalanine recognition (Figure 4A, Figure 9A). These results demonstrate the sensitivity of sequencing readouts to variations at a single position and indicate that both the directly recognized NAA and neighboring residues can affect the complete kinetic signature obtained from sequencing.

[0294] Because the aminoacyl-proline bond of the YP motif in peptides such as RLIFAYPDDD (SEQ ID NO: 1041) cannot be cleaved by PhTET aminopeptidase, the observation of YP pulsing at the end of the trace confirms that cleavage proceeded completely from the first to the last recognizable amino acid. Thus, the sequencing output from RLIFAYPDDD (SEQ ID NO: 1041) provided a convenient data set for investigating biochemical sources of non-ideal behavior that may lead to errors in peptide identification. The primary sources of incomplete information in the trace were the loss of the expected RS due to the stochastic occurrence of rapid sequential cleavage events (Figure 9B) and premature termination of the lead resulting from photodamage or surface detachment (Figure 9C).

[0295] In addition to changes in amino acid sequence composition, the sequencing readout is sensitive to changes due to PTMs. As an example, methionine oxidation was examined. The thioether moiety of the methionine side chain is susceptible to oxidation during peptide synthesis and sequencing. PS961 has a K of 947 nM. DIt was determined that PS961 binds to peptides with an N-terminal methionine at P2 (Figure 9D), and oxidation resulting in a polar methionine sulfoxide side chain was hypothesized to eliminate binding and reduce NAA binding affinity when located at P2. It was computationally determined that methionine sulfoxide is highly unfavorable in the PS961 NAA binding pocket, and that nonpolar residues are preferred at P2 (Figure 9E). The synthetic peptide RLMFAYPDDD (SEQ ID NO: 1043) was sequenced, and two populations of traces with distinct kinetic signatures were observed: the first population, which included leucine recognition with a median PD of 0.86 s, and the second population, which had a median PD of 0.35 s (Figure 4C). Traces from the first population also showed methionine recognition with a short PD in the time interval between leucine and phenylalanine recognition (Figure 4E). Methionine recognition was absent in traces from the second population (Figure 4D), indicating that the methionine side chain in these peptides could not be recognized by PS961. When methionine was fully oxidized by preincubation with hydrogen peroxide (Materials and Methods), as expected, elimination of both the methionine-recognition and leucine-recognition clusters was observed, with a long median PD (Figure 4E). These results demonstrate the ability for highly sensitive detection of PTMs due to their kinetic effect on recognition.

[0296] Proteomics applications require the identification of peptides in mixtures derived from biological sources. To extend our results to peptide mixtures and biologically derived peptides, we performed two experiments. First, we mixed the DQQRLIFAG (SEQ ID NO: 1036) and RLAFSALGAADDD (SEQ ID NO: 1038) peptides, immobilized them on the same chip, and performed sequencing experiments. Data analysis (Materials and Methods) identified two populations of traces corresponding to each peptide, and the kinetic signatures were nearly identical to those identified in runs using individual peptides (Figures 5A and 9F). Second, to demonstrate that our method can be extended to biologically derived peptides, we performed sequencing experiments using peptide libraries generated using a simple workflow from recombinant human ubiquitin (76 amino acids) and GLP-1 (37 amino acids) proteins digested with AspN / LysC and trypsin, respectively (Methods). For both libraries, data analysis readily identified traces matching the expected recognition patterns for the protease cleavage products DQQRLIFAGK (SEQ ID NO: 1045) and EFIAWLVK (SEQ ID NO: 1046) of ubiquitin and GLP-1, respectively, and generated kinetic signatures consistent with synthetic versions of these peptides (Figure 5B, Figure 9G). Matches to the kinetic signature of the ubiquitin peptide DQQRLIFAGK (SEQ ID NO: 1045) were identified across the human proteome using simple sequence constraints provided by the kinetic information (Materials and Methods). Only one protein other than ubiquitin was found to contain a peptide that could potentially match this signature. Thus, even a short signature could be obtained by 10 4 These results demonstrate the potential of full kinetic output from sequencing to enable digital mapping of peptides to their protein of origin.

[0297] Consideration This simple, real-time kinetic approach differs significantly from other recently described single-molecule approaches, which rely on stepwise Edman chemistry or complex iterative methods involving hundreds of cycles of epitope probing. While nanopore approaches offer the potential for real-time readout and simplicity, they face significant challenges related to the size and biophysical complexity of polypeptides. The sequencing technology described herein is easily expandable in its capabilities, and multiple areas for improvement exist. Expanded proteome coverage can be achieved through directed evolution and engineering of recognition factors. The NAA targets demonstrated here comprise approximately 35.6% of the human proteome, although lower affinity NAA targets require longer PDs to enable detection in all sequence contexts.

[0298] New amino acid or PTM recognition factors can be evolved from current recognition factors or identified in screens of other scaffolds, such as other types of NAAs or PTM-binding proteins or aptamers. Generally, expansion to detection of all 20 natural amino acids and multiple PTMs is feasible for de novo sequencing. However, for most proteomics applications that rely on mapping candidate proteins to a predefined set, partial sequences are sufficient. Aminopeptidases can be engineered to optimize cleavage rates and minimize RS loss from rapid sequential cleavage. The dynamic range of the sample and the most suitable application for the system tend to scale with the number of reaction chambers on the chip, and it is anticipated that dynamic range compression may be necessary for certain applications.

[0299] The sequencing technologies demonstrated herein promise to increase the availability of proteomic testing, enable new discoveries in biological and clinical research, and help power a new generation of precision medicine.

[0300] Materials and Methods Semiconductor device operation and bin ratio calculation Experiments were conducted on a pre-fabricated semiconductor chip with 296K active wells, taking into account some losses due to flow cell blockage in the sensor array. The dual-chamber flow cell allows two independent samples to be sequenced in parallel, each utilizing 148K active wells. The initial production device will have 2M active wells, scaling to tens of millions of active wells using standard CMOS processing for the first product line. Pulsed 532 nm excitation light from a 67 MHz mode-locked laser is coupled into a grating coupler at the edge of the semiconductor chip. The use of a single laser wavelength combined with fluorophore discrimination by fluorescence intensity and lifetime reduces size, cost, and complexity, contributing to the scalability of the platform. A network of optical waveguides splits the excitation light and routes it to the sensor array, illuminating each reaction chamber. Each CMOS pixel contains a single photosensitive photodiode with two high-speed global shutters (rejection gate and collection gate) that discard and collect photoelectrons (the chip's photonic architecture reduces pixel-to-pixel crosstalk to less than 2%). Control waveforms are applied to the collection and reject gates in synchronization with the incident pulsed light source (Figure 6B). Approximately 1 ns before the excitation pulse, the reject gate is charged to >3 volts, and the collection gate is discharged to <1 volt. Scattered 532 nm excitation photons generate photoelectrons within the photodiode. The photoelectrons are rapidly transferred to a high-voltage drain by the built-in potential field within the photodiode and the reject gate potential. 1–3 ns after excitation, the collection gate is charged to >3 volts, and the reject gate is discharged to <1 volt. Photoelectrons generated from emitted photons that reach the photodiode after the collection gate is opened are transferred to a storage node within each pixel. Photoelectrons are configurably accumulated within each pixel for 7.5–30 ms over approximately 500,000–2,000,000 laser pulses (Figure 6B). The charge stored in the storage node is measured using standard transfer gates, floating diffusions, source followers, row selects, and on-chip analog-to-digital converters common to all CMOS image sensors, allowing scalability to large array sizes with small pixels.Fluorescence lifetime information is obtained by alternating the timing of the collection and rejection gating waveforms between subsequent measurements. In the first measurement, only emission photoelectrons arriving >3 ns after the excitation pulse are collected (bin 0). In the second measurement, emission photoelectrons arriving >1 ns after the excitation pulse are collected (bin 1). The signal measured from the pixel as the phase relationship between the excitation source and the gating waveform is adjusted throughout the entire excitation cycle demonstrates that the pixel transitions from 100% collection of photons during the collection phase to >99.99% extinction of photons during the rejection phase in less than 1 ns (Figure 6C). The ratio of these two measurements (the bin ratio) provides an estimate of the fluorescence lifetime (Figure 6D). We demonstrated the ability to distinguish between multiple dyes based solely on the bin ratio (Figure 6E).

[0301] Peptide synthesis and labeling Peptides were synthesized on rink amide resin using standard Fmoc chemistry on a PurePrep® Chorus solid-phase peptide synthesizer (Gyros Protein Technology). All synthesized peptides contained a C-terminal Fmoc-azidolysine. The resin was deprotected in a mixture of TFA / TIPS / HO (2.5% / 2.5% / 95%) at room temperature for 1.5 h. The deprotection mixture was concentrated under a stream of argon. Peptides were precipitated from cold diethyl ether, resuspended in 1:1 water-acetonitrile, and purified by reverse-phase HPLC (X-bridge C18, Waters) using a gradient of 10–70% acetonitrile (0.05% TFA) over 20 min. The residue was dried under high vacuum to produce a white pellet. To a solution of DBCO-DNA-biotin (2 nmol in 100 μL PBS) was added 4 μL of the peptide stock solution (5 mM) at room temperature. The reaction progress was monitored by LC-MS (Thermo UlTiMate 3000 Executive Plus). After the reaction was complete, the mixture was conjugated to excess streptavidin. The peptide-DNA-streptavidin complex was purified by ion-exchange HPLC (DNAPac 200, Thermo). A gradient was used: Buffer A, 20 mM sodium phosphate buffer, pH 8.5; Buffer B, 1 M NaBr, 20 mM sodium phosphate buffer, pH 8.5, 20–60% B over 15 min. The purified complex was buffer-exchanged onto a 30K MWCO spin filter into a solution containing 50 mM MOPS (pH 8.0) and 60 mM potassium acetate before use. Fully oxidized methionine-containing peptides were prepared by mixing 3% hydrogen peroxide with the methionine peptide in 1:1 water-methanol at room temperature for 20 min. The product was immediately purified by reverse-phase HPLC using the same peptide purification method as above, and the purity was confirmed by reverse-phase HPLC (Thermo UlTiMate 3000) on an analytical column (Zorbax® SB-Aq, 5 μm, 4.6 × 250 mm), and the exact mass of the oxidation product was confirmed by LC-MS (Agilent LC-MSD-iQ, positive mode).

[0302] Protein digestion and labeling GLP-1 7-37, GLP-2, and ubiquitin (1-76) recombinant proteins were purchased as lyophilized powders from RnD Systems. Each protein was reconstituted to a final concentration of 200 μM in 100 mM HEPES, pH 8.0 (20% acetonitrile). Where necessary, cysteines were reduced and alkylated using TCEP (2 mM) and iodoacetamide (10 mM). GLP1 and GLP2 were digested overnight at 37°C using 1 μg of trypsin (LCMS grade, Pierce). Ubiquitin was digested using 1 μg of LysC (LCMS grade, Pierce) and 1 μg of rAspN (LCMS grade, Promega). After protease digestion, the pH of the peptide mixture was adjusted to pH 10.5 using potassium carbonate (57 mM), and lysines were converted to azidolysines using imidazole-1-sulfonyl azide (ISA, 2 mM) and copper sulfate catalyst (0.5 mM). ISA was quenched using amine-functionalized polyurethane beads (Oligo Factory). The mixture was then filtered and adjusted to pH 7–8 using 1 M acetic acid. The solution was diluted with 50% (v / v) 10 mM MOPS, 10 mM KOAc, pH 7.5, added to the DNA-streptavidin-DBCO complex, and incubated at 37°C for 12–16 h. When necessary, the detergent cetrimonium bromide was added to the reaction to a final concentration of 0.25 mM.

[0303] Recognition factor purification, labeling, and characterization The expression vectors for the recognition factors (with a pET30a+ backbone) and biotin ligase were cotransformed into BL21(DE3) chemically competent E. coli cells. Transformed cells were plated on Luria agar plates containing carbenicillin (50 μg / mL) and kanamycin (25 μg / mL) and incubated overnight at 37°C to obtain single colonies. A starter liquid culture inoculated with the colony was grown in Luria broth containing ampicillin (50 μg / mL) and kanamycin (25 μg / mL) and used to inoculate a large-scale culture at an initial optical density (OD600) of approximately 0.01. The expression culture was incubated at 37°C and 230 rpm until the OD600 approached approximately 0.7. The culture was then induced with 4 mM IPTG. The expressed recognition factors were biotinylated in vivo by adding 8 mM biotin simultaneously with IPTG. Approximately 12 hours after expression, cells were harvested by centrifugation at 10,000 g at 4°C, and the cell pellet was washed with 1x PBS buffer, pH 7.4. The cells were resuspended in Bugbuster® HT (Thermo Fisher Scientific) and incubated on a magnetic stirrer at room temperature for 30 minutes. The cell suspension was then diluted with an equal volume of 2x lysis buffer (100 mM Tris-HCl pH 7.5, 10% glycerol, 0.5 M NaCl) and incubated on a magnetic stirrer at room temperature for 30 minutes. The lysate was centrifuged at 21,000 g at 4°C to remove cell debris. The supernatant was collected and loaded onto a nickel-NTA resin (Cytiva) affinity column pre-equilibrated with buffer A (50 mM Tris-HCl pH 7.5, 10% glycerol, 0.5 M NaCl) on an AKTA Pure (Cytiva) system. The column was washed with at least 10 column volumes of buffer containing 10 mM imidazole. Elution was performed using a 10-300 mM imidazole gradient. The eluted fraction was dialyzed overnight at 4°C in a 10 kDa cassette against 4 L of dialysis buffer (50 mM Tris-HCl pH 7.5, 0.2 M NaCl, 50% glycerol).

[0304] For labeling of the recognition factor, equal volumes of the recognition factor and DNA-dye-streptavidin conjugate were mixed at a molar ratio of 5:1 (recognition factor:DNA-dye-SV). The mixture was incubated on ice for 30 minutes and dialyzed overnight against SEC buffer (25 mM HEPES pH 8.0, 150 mM KCl). The recognition factor-dye conjugate was recovered from dialysis and centrifuged at 10,000 g and 4°C. The supernatant was collected and concentrated using a 10 kDa cutoff concentrator. The concentrated conjugate was purified using a size exclusion column (BioSEC-3 300 Å, 3 μm) on an Agilent 1260 Infinity HPLC system.

[0305] Binding affinity was measured by polarization using labeled peptides. Polarization response and total intensity measurements were performed at 20°C in a microplate fluorometer using 480 nm excitation and 530 nm emission. Interactions between the recognition factor and a labeled peptide containing the target N-terminal residue (XAKLDEESILKQK-FITC (SEQ ID NO: 1074)) were performed in PBS buffer at pH 7.4, with readings collected after 30 minutes. To obtain titration curves, multiple analyses were performed with increasing concentrations of recognition factor at a fixed concentration of target peptide. The equilibrium polarization response at each concentration was plotted and fitted to determine the K D was calculated.

[0306] Using a stopped-flow apparatus, the off-rate (k) of PS610 was measured for various peptides. off ) was measured. Labeled peptide (50 nM) was mixed with PS610 in PBS buffer (pH 7.4) containing 0.01% Tween-20 and incubated at 30 °C. After 30 min of incubation, the recognition factor:peptide complex was rapidly mixed with a 10- to 20-fold molar excess of unlabeled trap peptide, and the reaction was tracked in real time by measuring the fluorescence intensity. At least three time-lapse traces were averaged and fitted to an exponential equation.

[0307] Aminopeptidase purification Expression vectors for the aminopeptidases PhTET2 and PhTET3 (with a pET30a+ backbone) were transformed into BL21(DE3) chemically competent E. coli cells. Transformed cells were plated on Luria agar plates containing kanamycin (25 μg / mL) and incubated overnight at 37°C to obtain single colonies. A starter liquid culture inoculated with the colony was grown in Luria broth (LB) containing kanamycin (25 μg / mL) and used to inoculate a large-scale culture at an initial optical density (OD600) of approximately 0.01. The expression culture was incubated at 37°C and 230 rpm until the OD600 approached approximately 0.7. The culture was then induced with 0.4 mM IPTG. The expressed aminopeptidases were purified as described above for the recognition factors. For conditioning, the aminopeptidase protein was dialyzed against 50 mM MOPS (pH 8.0) / 60 mM potassium acetate and then exposed to a final concentration of 400 μM cobalt acetate at 65°C for 1–1.5 h to form the active dodecamer complex. The conditioned aminopeptidase preparation was further dialyzed against 50 mM MOPS pH 8.0 / 60 mM potassium acetate, aliquoted, and flash-frozen.

[0308] Peptide loading, recognition, and dynamic sequencing The semiconductor chip was placed in the sequencing device, and a chip check was performed to test electronic circuit function and optimize laser coupling alignment. The chip was then removed from the device socket and washed twice with 50 μL of 70% isopropanol, followed by four washes with 30 μL of wash buffer (50 mM MOPS pH 8.0, 60 mM potassium acetate, 50 mM glucose, 20 mM magnesium acetate, and a surfactant mixture) through a flow cell attached to the chip. A second chip check was then performed. The laser was then shut off via an integrated software-controlled shutter, and the peptide complex was added to a final concentration of 1–10 nM, mixed thoroughly, and the chip was incubated for 15 min. The chip was then washed six times with wash buffer, followed by the addition of imaging solution (wash buffer containing 5 mM Trolox and an oxygen scavenging and scavenging system). The laser was then unblocked, and the occupancy percentage (target 10-30%, Poisson distribution) was recorded by acquiring the photobleaching signal from the fluorophore bound to the peptide complex during 5 min of laser irradiation. For NAA recognition-only assays, after peptide loading, labeled recognition factors were added to a final concentration of 50 nM PS610, 100 nM PS691, or 250 nM PS961 (as indicated in the experiment), and data were recorded for 10 h. For kinetic sequencing assays, after peptide loading, a mixture of labeled recognition factors was added to obtain final concentrations of 50 nM PS610, 100 nM PS691, and 250 nM PS961. Data were recorded for 15 min. The laser was then briefly blocked, and aminopeptidase was added to the sequencing reaction via the flow cell and thoroughly mixed (final concentrations of 2-8 µM PhTET2 and / or 20-80 µM PhTET3, as indicated in the experiment). The laser was then unblocked and data was recorded for 10 hours. For all runs, 30 μL of mineral oil was added to the fluid reservoirs of each port of the flow cell to prevent evaporation during the run.

[0309] Signal Processing and Trace Segmentation The signal measured on-chip contains various noise components, the most dominant of which is due to fluorescence emission from the diffusing recognition agent within the reaction chamber. The pulse-caller algorithm for a given reaction chamber begins by estimating the statistical properties of this background noise component. Once an estimate within certain error limits is established, the algorithm operates in an online manner, observing new frames of data as they are generated. At each point in time, the algorithm maintains a state indicating whether the signal is due to the background component alone or whether a pulse from the recognition agent-NAA interaction is being observed. The background-to-pulse state transition is triggered using an edge-detection test, where a shift in the signal is expected to be significant relative to the statistical distribution of the background component. The pulse-to-background state transition is triggered when a small window of recent frames of signal appears to again follow the distribution of the background component. The algorithm maintains an updated model of the background component as new background frames are observed. This, along with a feedback control loop that maintains stable optical coupling of the laser to the chip based on any such detected drift, provides robustness against drifts in signal intensity. Because detected pulses can be due to true recognition factor-dipeptide interaction events as well as other incidental transient noise spikes, a downstream filter layer is used to test the significance of pulse events based on their duration, intensity, and noise pattern within the context of the entire timeline of the run and the entire data set of the reaction chamber.

[0310] Initial regions are determined by performing a sliding-window calculation of pulse rates along the time dimension of the series of pulses. Regions with a mean pulse rate >1 pulse / min are then subdivided according to a greedy bisection approach. Here, the left and right pulses of each potential division are evaluated for statistically significant deviations in any of four distinct pulse characteristics (intensity, time bin ratio, pulse duration, and interpulse duration) using a Mann-Whitney U test. To define the RS, the region is subdivided using the division point with the lowest p-value for any of the four characteristics, with a p-value <10 for any comparison. -5 The process continues until there are no more regions with candidate division points with a . In this way, the transition from one RS to the next in a region of continuous pulsing is determined a priori based on the change in fluorescence characteristics of the pulsing kinetics. The resulting region is called a recognition segment (RS).

[0311] Recognition Segment Classification RS classification for reactions containing a single synthetic peptide was performed using an unsupervised clustering algorithm. A Gaussian mixture model (GMM) was pretrained using a subset of RSs containing those with an average signal-to-noise ratio of constituent pulses ≥3 to identify approximate centroids for each of the N classes of recognition. N is equal to the number of predicted recognizable peptide states with F, Y, W, L, I, V, or R at the N-terminus. Identified clusters were assigned to recognizable peptide states by matching the predominant order of observed cluster sequences to the predicted amino acid sequence and using prior knowledge of dye properties to identify binding factors active in each RS. Subsequent rounds of GMM fitting were performed on all RSs that matched the predicted order of these events to refine the GMM model until no further sequences appeared in the predicted order. The final model was then applied to all RSs in a given reaction.

[0312] RS classification for reactions containing library-prepared peptides and mixtures of peptides was performed using a random forest classifier pre-trained on annotated RS pulse features from previous synthetic peptide experiments. Unless otherwise noted, figures and statistics generated from classified RS are derived from reaction chambers with the expected sequence of RS.

[0313] Molecular dynamics and binding energy calculations A homology model of PS961 complexed to the peptide was generated using the internal crystal structure, mutations were applied, and optimized using protCAD prior to molecular dynamics. AMBER20 implicit solvent molecular dynamics simulations with a general Born solvation potential were performed using the ff19SB force field with no interatomic distance cutoff. Minimization was performed using steepest descent, followed by conjugate gradient minimization. Langevin dynamics and 3ps -1 The system was heated from 0 to 300 K using a collision frequency of 100 kJ / s. Molecular dynamics simulations of the equilibrated recognition factor-peptide complex, the free recognition factor, and the free peptide were performed independently at 300 K for 5 ns, and binding energy calculations were performed using MMPBSA. Here, 125 frames, each containing 10,000 2-femtosecond steps, were used for calculations from three simulations. Binding energies and decompositions of all residues contributing to the binding energy were calculated at 0.15 M salt concentration.

[0314] Example 2. Peptide identification using modeled proteome-wide kinetic signatures Using sequencing and biochemical data, we determined predicted pulse durations for recognition factors that bind to all possible tripeptide targets. Figures 10A–10C show heat maps of predicted pulse durations for PS961-bound tripeptide targets with leucine (Figure 10A), isoleucine (Figure 10B), or valine (Figure 10C) at the N-terminus. Figures 10D–10F show heat maps of predicted pulse durations for PS610-bound tripeptide targets with phenylalanine (Figure 10D), tyrosine (Figure 10E), or tryptophan (Figure 10F) at the N-terminus. Figure 10G shows a heat map of predicted pulse durations for PS1122-bound tripeptide targets with arginine at the N-terminus. The predicted pulse durations correlated highly with the actual pulse durations from on-chip experiments for PS961 (Figure 10H, left plot) and PS610 (Figure 10H, right plot).

[0315] This database of predicted tripeptide pulse durations can be used to model the expected kinetic signatures of all peptides in the human proteome, which may provide improved understanding and utilization of the ability to identify proteins from sequencing output. Kinetic signatures are average representations of peptide-on-chip sequencing behavior, as detailed in Example 1 above. The information in kinetic signatures derived from single-molecule traces dramatically improves the ability to map sequencing data to proteomes (e.g., compared to methods based on text string alignment, such as in DNA sequencing). Kinetic information can include, for example, pulse duration, inter-pulse duration, and recognition segment (RS) duration.

[0316] Because recognition factors bind to peptides, they contact not only the N-terminal residue but also (at least) two adjacent downstream residues, and this kinetic information enhances proteome mapping data. In this way, recognition factors indirectly sense all 20 amino acids, and this information is encoded in the average pulse duration (and potentially in the IPD and RS duration). Furthermore, adjacent visible residues in a peptide are, on average, represented by their immediate neighboring RSs (i.e., a consensus gap exists between two RSs only if there is at least one invisible amino acid between them).

[0317] To prepare a model demonstrating the ability to uniquely map peptides to the human proteome (using the recognition factors PS961, PS610, and PS1122), we performed an in silico AspN / LysC digest of the proteome, followed by the selection of all peptides ending in lysine (used for on-chip immobilization) and greater than seven amino acids in length. The results are shown below.

[0318] [Table 3]

[0319] A predicted pulse duration was assigned to every visible amino acid in the set of 273,112 peptides (positions with a predicted average PD less than 0.18 seconds were treated as invisible). The distribution of predicted RSs in the first 15 residues is shown in Figure 10I (left plot). 82,068 peptides contained four or more RSs (and were therefore considered potentially informative). A kinetic signature was generated for each of these peptides.

[0320] The kinetic signature includes the predicted binding factor and average PD at each visible position, as well as gaps representing runs of one or more invisible amino acids. Next, for each peptide, the number of peptides with the same kinetic signature was determined. (Signatures were considered identical if they had the same order of RS and gaps, and the predicted PDs at each RS were somewhat similar (in any pairwise comparison, the shorter PD was more than half of the longer PD).) This analysis revealed that 38,849 of the 82,068 peptides yielded unique kinetic signatures with no other matches in the human proteome. An additional 10,571 peptides had only one other match. The distribution of kinetic matches per peptide is shown in Figure 10I (center plot). 14,167 proteins (69% of all proteins) contained at least one uniquely mappable peptide. On average, there were 2.5 uniquely mappable peptides per protein. The distribution of uniquely mappable peptides per protein is shown in Figure 10I (right plot).

[0321] To further illustrate this data and how it can be used to model protein behavior, results for the IL6 protein are shown in Figure 10J (for simplicity, the residue immediately preceding the C-terminal lysine was treated as invisible and the XP motif was treated as cleavable). As shown in Figure 10J, two peptides contain at least four RSs. As shown in Figure 10K, one of these peptides uniquely maps to IL6, and the other peptide matches the kinetic signatures of eight different peptides from eight proteins.

[0322] To provide an illustrative example using a smaller proteome, the E. coli proteome (containing only 4392 proteins) was analyzed as described above for the human proteome. The results are shown below.

[0323] [Table 4]

[0324] The distribution of predicted RSs for the first 15 residues is shown in Figure 10L (left plot). 9925 peptides contained four or more RSs (and therefore were considered potentially informative). A kinetic signature was generated for each of these peptides. For each peptide, the number of peptides with the same kinetic signature was determined. This analysis revealed that 7740 of the 9925 peptides yielded unique kinetic signatures with no other matches in the E. coli proteome. The distribution of kinetic matches per peptide is shown in Figure 10L (center plot). 3187 proteins contained at least one uniquely mappable peptide. On average, there were 2.4 uniquely mappable peptides per protein. The distribution of uniquely mappable peptides per protein is shown in Figure 10L (right plot). To illustrate this data and how it can be used to model protein behavior, results using a protein from E. coli containing six uniquely mappable peptides are shown in Figure 10M.

[0325] These results demonstrate the utility of a kinetics-centric view of peptide identification, which also provides the ability to accurately model the informatics impact of changes to reaction conditions such as adding new recognition factors, increasing recognition factor pulse duration, changing frame rates, and adding new dye labels.

[0326] Example 3. Selection of N-terminal alanine and valine binding mutants by yeast display The gene encoding PS557 was used as a template for error-prone PCR, in which multiple nucleotide changes were introduced to generate mutant PS557 proteins. The mutant protein library was transformed into yeast and used for yeast display and flow cytometry, selecting for peptides of interest. Briefly, proteins can be labeled with a tag, such as a myc tag, and a fluorescently labeled antibody against this tag identifies cells expressing the protein. The peptide of interest is biotinylated, and the peptide is labeled with a streptavidin-conjugated fluorophore to identify yeast cells displaying proteins bound to the peptide. The yeast cells, peptide, and fluorophore mixture are incubated at room temperature for 1 hour, followed by two-color FACS to obtain double-positive cells.

[0327] In this example, selection was performed using peptides with amino acid residues V or A at the N-terminus. After three rounds of selection for each peptide, samples were sent for next-generation sequencing, and the results were used to rationally design additional libraries for further iterations of directed evolution or individual proteins to be tested in biochemical assays. Figure 11A shows the results from FACS selection for three rounds of selection, with flow cytometry plots showing cells expressing the protein (y-axis) relative to cells binding to the AVP peptide (x-axis). Figure 11B shows the results from one round of selection of the error-prone PCR library mixture, with flow cytometry plots showing cells expressing the protein (y-axis) relative to cells binding to 1 μM AV or AI peptide (x-axis). In each of the flow cytometry plots shown in Figures 11A-B, signals in the upper right quadrant indicate more binding. A selection of hits obtained from sequencing the alanine and valine libraries is shown in Table 3.

[0328] [Table 5-1]

[0329] [Table 5-2]

[0330] The mutation N41D was selected as the top hit, which can be rationalized via computer modeling. When residue N41 is mutated to D in silico, the Rosetta algorithm yields a more favorable binding energy with the valine peptide (-1.8) compared to wild-type PS557. This predicted binding energy difference is approximately that of one to two hydrogen bonds, and therefore, the experimental K D This could result in a 5- to 10-fold difference in binding. Figure 11C illustrates the predicted binding of an alanine peptide to a mutant protein. The left panel of Figure 11C shows a model of the entire PS557 protein with the peptide shown at the binding site forming a hydrogen-bonding interaction with the N41D mutation. The right panel of Figure 11C shows the superimposed hydrogen-bonding network with negatively charged (light-shaded) and positively charged (associated) surface regions.

[0331] The V72M mutation also appears to be enriched by selection, which can also be rationalized by computer modeling. Therefore, these two mutations were selected and combined together, and the double mutant was expressed in E. coli, purified, and tested for binding activity. Similarly, other mutations were selected from the NGS dataset and combined using a reasonable design to obtain a panel of potential hits to be tested in a high-throughput assay on the Octet platform.

[0332] In the high-throughput assay, Octet sensors are coated with the peptide of interest and immersed in a buffer containing the purified protein. Comparing the response after 200 seconds for each protein provides an approximation of relative binding. Because each protein has approximately the same molecular weight and is used at the same concentration, ranking proteins by their response level in this assay provides an approximation of which proteins have improved binding. Examining on-rates and off-rates can also provide insight into the mechanism of binding. Response values for different peptides with constructs selected from this selection round are shown in Tables 4 and 5, which indicate the binding affinity of the selected candidates compared to the wild-type protein (PS557) in the high-throughput Octet assay. The responses (nm) for each peptide with the two N-terminal residues of the listed peptides after 200 seconds of incubation with the mutant protein candidates are shown (Table 4: AA, VA, LA; Table 5: MA, IA, FA, WA, YA).

[0333] [Table 6]

[0334] [Table 7]

[0335] The results show that binding was improved for both valine and alanine binding in some candidates, such as PS824, which was selected for further characterization. Upon more quantitative measurement by fluorescence polarization, binding was improved five-fold for the amino acid valine. The wild-type protein had a measured K of 1174 nM for peptides with an N-terminal valine. D whereas the PS824 mutant containing the mutations I12F, N41D, Q55R, and V72M had a K of 205 nM. DSimilarly, clone PS852, containing R31H, N41D, Q55R, and V72M, has a K of 142 nM for the valine peptide as measured by fluorescence polarization. D Figure 11D shows exemplary results from a fluorescence polarization study comparing the kinetics of N-terminal alanine peptide binding for selected PS557 variants. The results shown in Figure 11D were obtained using mixtures containing PBS buffer with 0.01% Tween® 20, peptide (AAKLDEESILKQ{LYS(FITC)} (SEQ ID NO: 1075)) at a concentration of 100 nM, and protein concentrations ranging from 100 to 5000 nM.

[0336] [Table 8]

[0337] In parallel, the NGS dataset was also used to design a second-generation library in which additional mutations were layered on top of the clones selected from the first round of directed evolution. Both error-prone PCR and targeted mutagenesis libraries were generated and further rounds of selection were performed as described for the second iteration of directed evolution. The top hits are summarized in Tables 7 and 8, which provide a snapshot of the best clones obtained from these methods.

[0338] [Table 9]

[0339] [Table 10-1]

[0340] [Table 10-2]

[0341] Example 4. Improved N-terminal amino acid binding proteins using SNAP display The gene encoding PS557 was cloned into a vector to be used for SNAP display, with the SNAP tag fused to its N-terminus. During SNAP display, the SNAP protein tag reacts and covalently binds to a benzylguanine (BG) molecule added to the end of the DNA template encoding the protein. This allows for correlation between the protein's phenotype and genotype, as long as the protein is expressed in a droplet emulsion using an in vitro transcription-translation system. The protein-DNA complex is exposed to the desired peptide bound to magnetic beads, unbound complexes are washed away, and high-affinity binding factors are eluted from the beads.

[0342] In these studies, a library of mutant PS557 genes was generated using error-prone PCR and targeted mutagenesis, with a rational design based on computer modeling and previous results. SNAP display was used to select the library for binding to alanine, valine, and methionine peptides. The naive and selected libraries were each sequenced, and the NGS data was used to determine the enrichment of clones from different selection rounds.

[0343] By comparing the frequency with which a given protein sequence appears in the library before and after selection, clones can be ranked by their potential affinity for the peptide. Most sequences exhibit low affinity and are selected. The results indicate that defective clones (e.g., clones with stop codons) are selected, lending credibility to the method. Figure 12A is a heat map showing the enrichment of mutations in the PS557 protein. The amino acid at which the residue was changed is listed on the left, and the residue number is listed below. Each rectangle represents the enrichment of that mutation in the selected library compared to the naive library (dark indicates no enrichment, light indicates enrichment). The rectangle represents the wild-type residue at this position. For simplicity, a subset of mutations, the stop codon, cysteine, or alanine residue, are shown.

[0344] As shown in Figure 12A, stop codons are largely unenriched at all positions in the sequences listed by residue number, however, alanine mutations, for example, are enriched at many positions indicated by light rectangles, more so than, for example, cysteines.

[0345] Four rounds of selection were performed on the library of mutant PS557 proteins. The most enriched sequences from this round of directed evolution on alanine peptides are shown in Table 9, which shows the mutations in PS557 sequences that were found to be enriched after four rounds of SNAP selection on N-terminal alanine peptides. Enrichment was calculated by dividing the percent abundance of clones in the NGS data from the fourth round of sequencing by their abundance in the naive library.

[0346] [Table 11]

[0347] As shown by the results in Table 9, many of the enriched sequences contained similar mutations. The library used for this iteration of directed evolution was designed based on hits from three previous rounds of directed evolution and selection, which could be explained by tracing the evolution of one particular clone (e.g., N41D, Q55R, E63S, L68M, V72M, Y100R). Mutation N41D was first identified as enriched after selection of an error-prone PCR library in yeast display (Table 10, directed evolution round 1).

[0348] [Table 12]

[0349] In the same selection, N41D combined with Q55R (double mutant) was also identified. At the same time, V72M was selected as enriched. However, clones containing combinations of these mutations, such as the triple mutant N41D, Q55R, and V72M, were not selected, likely due to the absence of all possible triple mutation combinations in the naive library. Many mutations, such as V72M, have reliable computational data that rationalize their significance. All of this data was analyzed, and a second library was created in which the combination N41D, Q55R, and V72M, as well as many other reasonably designed combinations, were intentionally included in the library. In the next round of selection using yeast display, this mutant was selected as one of the most enriched clones (Table 10, directed evolution round 2).

[0350] In parallel, a library designed with targeted mutations based on computational design and some of the hits already seen in previous rounds of directed evolution was subjected to one round of directed evolution using SNAP display. Each position was mutated to all 20 amino acids in different combinations. Several similar positions were identified in this selection, as in other previous selections, and in some cases, this residue was mutated to a different amino acid than previously seen. For example, E63 was identified as mutated to lysine (K) in round 1 of directed evolution, but was mutated to alanine (A) or serine (S) in round 3 (Table 10, directed evolution round 3). In the fourth round of directed evolution, a library was generated to test many of these combinations, including N41D, Q55R, E63S, L68M, and V72M, and additional residues were randomly mutated to all 20 amino acids (e.g., Y100). From this fourth round of directed evolution, enriched clones were identified, as listed in Table 9, and clones such as N41D, Q55R, E63S, L68M, V72M, and Y100R were selected for expression in E. coli, purified, and tested for binding activity in high-throughput assays on the Octet platform.

[0351] Candidates selected from the fourth round of directed evolution using SNAP display selection, compared to candidates from other previous selections, are shown in Figure 12B. In a high-throughput assay, an Octet sensor is coated with the peptide of interest and immersed in a buffer containing purified protein. A trace for an alanine peptide is shown in Figure 12B. Improved binding is indicated by an increase in response based on a wavelength shift in nm over time (association curve: 0-200 s, dissociation curve: 200-500 s). By comparing the response after 200 s for each protein, an approximation of relative binding can be obtained. Each protein has approximately the same molecular weight and is used at the same concentration. Ranking the proteins by their response level in this assay provides an approximation of which proteins have improved binding. Examining on- and off-rates can also provide insight into the mechanism of binding. The response values for the different peptides with the constructs selected from this selection round are shown in Table 11, which shows the response (in nm) for each peptide with the two N-terminal residues of the listed peptides (AA, VA, LA, etc.) after 200 seconds of incubation during the association phase of incubation with the mutant protein candidates.

[0352] [Table 13]

[0353] Example 5. Development of arginine recognition factor PS1122 This example describes the development of PS1122, an engineered variant of the UBR protein (PS621) from Kluyveromyces marxianus that exhibits improved arginine recognition on-chip, with increased affinity for arginine and histidine. Based on binding kinetics analysis and on-chip results, PS1122 has approximately 7-fold higher binding affinity for N-terminal arginine than PS621, resulting in a favorable increase in the pulse duration of the RX dipeptide and faster binding. These properties combine to improve the recognition range of arginine tripeptides and the accuracy of ROI detection. PS1122 is estimated to be able to unambiguously detect approximately 52% of all arginine positions in the human proteome, equivalent to 2.9% of the total proteome (an increase from approximately 1.4% in PS621).

[0354] Directed evolution approach PS621 (and its tandem form, PS691) binds arginine (R), histidine (H), and lysine (K). Observable on-chip binding of PS621 to R is limited to approximately 25% of arginine positions in the proteome due to the influence of downstream residues on pulse duration. Directed evolution was used to select PS621 variants with stronger arginine binding. Multiple types of mutant libraries were subjected to many cycles of selection and mutational evolution, resulting in a panel of candidate recognition factor variants (PS1101–PS1122) that were advanced to biochemical and single-molecule investigations.

[0355] Octet analysis Mutant binders (PS1101-PS1122) and controls were expressed in E. coli, purified in a high-throughput workflow, and evaluated for binding to the N-terminal amino acid on the Octet platform. The peptides used in the assays mostly contained a penultimate alanine and consisted of the sequence XAKLDEESILKQK (SEQ ID NO: 1074). A peptide with the sequence RXKLDEESILKQK (SEQ ID NO: 1076) was also used to evaluate the effect of the penultimate residue. The set of Octet response measurements for RX (various R dipeptides), HA, and KA are summarized in Table 12.

[0356] [Table 14]

[0357] Binding affinity by polarization Fluorescence polarization assays were performed with all candidates to measure single-point binding responses at a fixed concentration of binder (Figure 13A). This assay measures the strength of interaction between the binder and a labeled peptide (XAKLDEESILKQK-FITC (SEQ ID NO: 1074)). Based on the binding responses, the inventors further investigated selected candidates by measuring their Kd. The inventors performed multiple polarization binding titration analyses with increasing concentrations of binder protein and determined the Kd from the titration curves (Figure 13B).

[0358] PS1122, PS1115, PS1106, PS1114, PS1121, and PS1104 showed the greatest improvement in binding to the RA peptide. Several mutants also showed improved HA binding compared to PS621. RA-binding affinity determination titration curves for PS621, PS691, and PS1122 are shown in Figure 13B. Compared to PS621 and PS691, PS1122 showed a 5- to 7-fold increase in binding affinity to the RA peptide and an approximately 2-fold increase in binding affinity to the HA peptide.

[0359] k on and k offStopped-flow high-speed kinetic analysis of Using stopped-flow assays, the on-rate constants (k on ) and dissociation rate (k off ) were measured (results are summarized in Table 13). These measurements were performed to predict the relative improvement of pulse duration and inter-pulse duration on-chip.

[0360] [Table 15]

[0361] In these assays, the RX peptide gave better signals and higher accuracy of measurements due to its stronger binding than HA. on The rate constants were comparable to those of the tandem binder PS691. The k of PS1122 for the RA and RL peptides off The rate was approximately 3-3.5 times slower than PS621 or PS691, predicting longer pulse durations on-chip. PS1122 also dissociated from the HA peptide approximately 2-fold slower than PS691. These measurements identified PS1122 as a mutant for evaluation in single-molecule assays.

[0362] In parallel, we used a next-generation on-chip recognition factor screening method to evaluate multiple PS621 mutants on-chip. The PS621 mutants were purified in biotinylated form and labeled on a microscale using a streptavidin-tetra-Cy3B modification protocol, and recognition runs were performed on selected candidates. Various penultimate peptides of RX were used on-chip to evaluate the improved R recognition range against PS1122.

[0363] On-chip recognition of RA dipeptides For detailed on-chip characterization, PS1122 was purified and extensively labeled using streptavidin-tetracytosine Cy3 dye conjugates. Ensemble and screening on-chip assays demonstrated that PS1122 binding to arginines was sufficiently improved to recognize the RA dipeptide on-chip. Analysis of recognition runs using the QP304-RAIFAG peptide confirmed the presence of visible binding of PS1122 to the N-terminal RA at pulse durations longer than those of PS691 (0.16 s) (Figure 13C).

[0364] Sequencing performance and arginine tripeptide coverage by PS1122 The sequencing performance of PS1122 and PS691 was compared using QP433 (RLIFAYP (SEQ ID NO: 1087)) and other peptides. To further evaluate the scope of arginine recognition for the RXA tripeptide, multiple multiplexed kinetic assays were performed using PS1122 (Figure 13D).

[0365] Using the arginine tripeptide pulse duration data determined from the multiplexed run with PS1122, predicted pulse durations were determined for all 400 RXX tripeptides to estimate arginine proteome coverage for PS1122 (Figure 10G). The results predict 52% coverage of arginine positions, corresponding to 2.9% of the human proteome.

[0366] Example 6. PS961 modified peptide binding enhancement PS961 features six point mutations (N41D, Q55R, E63S, L68M, V72M, and Y100R) at the top of the PS557 precursor and binds to L / I / V / A / M / P N-terminal peptides with better affinity or on-chip performance than PS557. Each point mutation was analyzed in the context of the protein complexed with an alanine tripeptide, providing a structure-based rationale for the selection of these substitutions in directed evolution screens against native PS557 amino acid identity.

[0367] Directed enhancement of the binding pocket Exchange of asparagine in the outer ring of the binding pocket at position 41 for aspartic acid enhances both long- and short-range interactions of PS961 with N-terminal peptide ligands (FIG. 14A).

[0368] As shown in the electrostatic surface charge distribution (Figure 14B), long-range electrostatic interactions increase the probability that the positively charged peptide N-terminus will interact with the protein, and the additional negative charge from the aspartic acid side chain enhances electrostatic steering to the binding site compared to the asparagine side chain.

[0369] This amino acid position also forms a short-range electrostatic interaction with the peptide's N-terminal amino group, directly within the triad of residues essential for binding. The aspartic acid at position 41, with two electronegative atoms on either side of its side chain, may interact more favorably with the peptide by lowering the conformational entropy of significant electrostatic interactions in the binding pocket. The aspartic acid side chain is also predicted to have a more electronegative polarization of the orbitals, which may enable stronger hydrogen bonding between PS961 and the peptide N-terminus compared to the partially charged asparagine side chain. Rosetta all-atom energy functions confirm the existence of this stronger interaction by quantitatively identifying that the Coulomb energy (the "fa_elec" term in the scoring function) involving the N-terminal amino acid in PS961 is lower and more favorable in the presence of the triad with asparagine than in the presence of the asparagine triad (Figure 14C).

[0370] The substitution of valine with methionine at position 72 introduces a longer nonpolar side chain into the PS961 binding pocket (Figure 15). These extra nonpolar atoms can further intercalate into the cavity, thereby decreasing the pocket volume and increasing hydrophobic interactions with the smaller N-terminal amino acid.

[0371] Optimization of non-pocket hydrophobic packing Because methionine has a longer nonpolar side chain than leucine, this mutation at position 68 allows it to interact favorably with amino acids in the adjacent β-sheet, resulting in a structural cavity that is often filled in the simulations (Figure 16). Densely packed nonpolar amino acid side chains within the protein core reduce the unfavorable conformational entropy of the interior cavity and increase favorable hydrophobic interactions, both of which likely increase the overall folding stability and reduce the energetic cost of forming the binding pocket prior to binding.

[0372] Reducing potential alternative binding sites and increasing surface and net charge Because the N-terminus of the peptide ligand bears a permanent positive charge, any negatively charged pocket naturally present on the surface of PS557 may interfere with the probability of effective protein-peptide interactions at the binding site by acting as an alternative, lower-affinity, competitive binding site. The Y100R mutation alleviates potential off-pocket interactions by reducing the negative surface charge present in the absence of arginine mutations (Figure 17). Furthermore, in conjunction with Q55R and E63S, the total surface and net charges of the protein increase, more directly supporting the intended binding event for proper on-chip recognition.

[0373] Increased probability of loop conformations that positively interact with the peptide Molecular dynamics simulations of PS961 and PS557 bound to an AAA tripeptide reveal enhanced hydrogen-bonding potential of the arginine side chain at position 100 to adjacent loop residues compared with the tyrosine in PS557. This arginine often participates in a complex hydrogen-bonding network with R106 and the main-chain carbonyl of the penultimate residue of the peptide, providing further stabilization to the bound form of the peptide (Figure 18A). The average occupancy of the R100:R106 hydrogen bond, calculated as the percentage of 1,000 simulation frames that achieve a specific interaction, is approximately six-fold higher in simulations of PS961 bound to an AAA tripeptide compared with PS557 (Figure 18B). Similarly, the penultimate hydrogen bond of R106 is more likely to occur in PS961 than in PS557.

[0374] Figures 19A-C show the secondary structure, sequence, and binding pocket characteristics of PS961. Figure 19A shows the classification of the protein into secondary structure groups. Figure 19B shows a Poisson-Boltzmann electrostatic potential surface map of the binding pocket with the residues that form the binding pocket labeled and the corresponding pocket properties listed. Figure 19C shows the sequences of the native parent protein and engineered variants, highlighting the mutations, pocket positions, and secondary structure assignments for each position.

[0375] PS961 Crystallography The crystal structure of PS961 complexed with a target peptide bearing an N-terminal methionine (Met-Ala-Lys-Leu (MAKL) (SEQ ID NO: 1047)) was solved. The protein:peptide complex was generated and purified to obtain diffraction-grade crystals of PS961:MAKL (Figure 19D). A complete X-ray data set from these crystals was collected, and the final structure was solved and refined at 1.9 Å resolution (Figure 19E). Figures 19F-19K show how PS961 binds to the target peptide MAKL (SEQ ID NO: 1047).

[0376] Figure 19F shows that Met in the target peptide contacts Asp10 and Asp11 in the recognition factor PS961. The crystal structure also shows interactions with two water molecules (spheres), one of which is held in the proper orientation by Asp42 in the recognition factor. Substitution of Asp42 can alter the binding affinity of this N-terminal residue in the peptide. In addition to the ionic interactions shown in Figure 19F, Met1 in the peptide also makes nonpolar contacts with at least six residues in PS961 (Figure 19G, residues labeled other than "MET-1"). These residues in PS961 are arranged around the side chain of Met. With appropriate substitution of some of these residues, other N-terminal amino acids other than Met could potentially be recognized and bound by PS961.

[0377] The next residue in the peptide, Ala2, contacts the main chain of His14 in PS961 because Ala2 also H-bonds with a water molecule (Figure 19H, sphere), which in turn contacts Tyr16 and an additional water molecule. Substitution of Tyr16 could disrupt the presence of the water molecule and alter the binding of this penultimate residue in the peptide. The crystal structure also shows that the oily side chain of Ala2 in the peptide (with three methyl protons) is located between residues Thr15 and Leu73 in the recognition factor (Figure 19I), which opens the possibility of replacing these residues in the recognition factor with smaller side chains to create space for larger side chains in the peptide. Lys3 in the peptide forms a salt bridge with Asp42 in the recognition factor (Figure 19J, ionic interactions shown with dashed lines). Lys3 in the peptide is nearly removed from the binding site, while Tyr16 in the recognition factor is at a very short distance, as indicated by their surface (FIG. 19K).

[0378] The crystal structure of PS961 complexed with a target peptide bearing an N-terminal alanine (Ala-Ala-Lys-Leu (AAKL) (SEQ ID NO: 1048)) was solved. The protein:peptide complex was generated and purified to obtain diffraction-grade crystals of PS961:AAKL. A complete X-ray data set from these crystals was collected, and the final structure was solved and refined at 1.39 Å resolution. Figure 19L shows a panoramic view of the PS961:AAKL complex structure, with PS961 shown in ribbon representation, the AAKL (SEQ ID NO: 1048) peptide shown as sticks, water molecules shown as spheres, and PEG molecules shown in stick-and-ball representation.

[0379] Because of its high resolution, the PS961:AAKL crystal structure shown in Figure 19L shows many water molecules that can be modeled. It also shows alternate conformations for the three residues in the recognition factor. The recognition factor in the AAKL (SEQ ID NO: 1048) complex has the same overall fold as observed in the PS961:MAKL complex. Thus, the main chains of the recognition factors in both structures can be superimposed with an average deviation of only 0.23 Å. Figure 19M shows the superposition of the main chains of P...

Claims

1. 1. A recombinant or synthetic amino acid binding protein having an amino acid sequence that is at least 80% identical to SEQ ID NO:1, wherein the amino acid sequence contains amino acid substitutions at one or more positions corresponding to E22, R31, L39, N41, D42, D43, D44, H45, T46, Y47, V50, Q55, P62, E63, L68, A69, V72, D73, Q75, Y100, and M111 of SEQ ID NO:

1.

2. 2. The amino acid binding protein of claim 1, wherein the amino acid sequence comprises an amino acid substitution at a position corresponding to N41 and one or more positions corresponding to E22, R31, L39, D42, H45, V50, Q55, P62, E63, L68, V72, Q75, Y100, and M111.

3. 3. The amino acid binding protein of claim 1, wherein the amino acid sequence comprises an amino acid substitution at a position corresponding to N41 and one or more positions corresponding to Q55, E63, L68, V72, and Y100.

4. 4. The amino acid binding protein of claim 1, wherein the amino acid substitutions are selected from E22V, R31H, L39M, N41D, D42L / P, H45C / F, V50A / F / Y, Q55H / R, P62R, E63A / G / K / S, L68M, V72M, Q75L, Y100R, and M111A / S.

5. 5. The amino acid binding protein of claim 1, wherein the amino acid substitution is selected from N41D, Q55R, E63S, L68M, V72M, and Y100R.

6. 6. The amino acid binding protein of any one of claims 1 to 5, wherein the amino acid sequence is at least 85%, at least 90%, at least 95%, at least 98%, 80-98%, 80-95%, 80-90%, 85-95%, or 90-98% identical to SEQ ID NO:

1.

7. The amino acid sequences are PS635-645, PS731-732, PS759-766, PS769, PS795-870, PS896-912, PS918-1043, PS1048-1100, PS1124-1137, PS1141-1161, PS1175-1199, PS1203-1217, PS1222-1245, PS1277-1305, PS1321-1350, and PS1425. 1448 (SEQ ID NOs: 22-27, 87-88, 115-122, 125, 151-226, 249-265, 271-390, 395-446, 470-483, 487-507, 521-545, 549-563, 568-591, 622-650, 664-693, and 768-791). The amino acid binding protein of any one of claims 1 to 6.

8. A recombinant or synthetic amino acid binding protein having the structure of formula (I) or a structural equivalent thereof: β1-α1-α2-β2-α3-β3 (I) (In the formula, each of β1, β2, and β3 is a β chain; each of α1, α2, and α3 is an α-helix; Each instance of "-" is a loop, At least a portion of each of α1, α2, the loop between β1 and α1, and the loop between α3 and β3 form a binding pocket for an amino acid ligand, said binding pocket comprising: i) approximately 170 Å 3 The volume of ii) -3.0RTe c -1 The electrostatic potential, iii) negatively charged side chains on at least 35% of the amino acids forming the binding pocket; iv) a plurality of hydrogen bond acceptors configured to form one or more hydrogen bonds in the presence of the amino acid ligand; and v) a plurality of van der Waals contact locations configured to form van der Waals interactions in the presence of said amino acid ligand; (including one or more of the following) an amino acid binding protein,

9. The amino acid binding protein of claim 8 , wherein the amino acid ligand is a polypeptide comprising at least three amino acids.

10. 10. The amino acid binding protein of claim 8 or 9, wherein the amino acid ligand comprises the N-terminal amino acid of a polypeptide.

11. 11. The amino acid binding protein of claim 10, wherein the N-terminal amino acid is selected from leucine, isoleucine, valine, methionine, and alanine.

12. The amino acid binding protein has a dissociation constant (K) of less than 2,000 nM, less than 1,500 nM, less than 1,000 nM, less than 750 nM, less than 500 nM, less than 250 nM, less than 150 nM, less than 100 nM, less than 80 nM, 10-2,000 nM, 25-1,000 nM, 50-500 nM, 10-150 nM, 25-75 nM, or 50-60 nM. D 12. The amino acid binding protein of claim 11, wherein the N-terminal leucine is bound by a nucleotide.

13. the amino acid binding protein has a K of less than 2,000 nM, less than 1,500 nM, less than 1,000 nM, less than 750 nM, less than 500 nM, less than 250 nM, less than 150 nM, less than 100 nM, less than 80 nM, 10-2,000 nM, 25-1,000 nM, 50-500 nM, 10-150 nM, 30-80 nM, or 60-75 nM D The amino acid binding protein of claim 11 or 12, which binds to the N-terminal isoleucine at

14. the amino acid binding protein has a K of less than 2,000 nM, less than 1,500 nM, less than 1,000 nM, less than 750 nM, less than 500 nM, less than 300 nM, less than 250 nM, less than 200 nM, 10-2,000 nM, 25-1,000 nM, 50-500 nM, 50-300 nM, or 100-200 nM D The amino acid binding protein according to any one of claims 11 to 13, which binds to the N-terminal valine at

15. 15. The amino acid binding protein of any one of claims 8 to 14, wherein the amino acid binding protein is at least 50 amino acids in length, at least 75 amino acids in length, at least 100 amino acids in length, 50 to 250 amino acids in length, 50 to 150 amino acids in length, or 100 to 200 amino acids in length.

16. The amino acid binding protein according to any one of claims 8 to 15, wherein the loop between β1 and α1 contains three or more negatively charged amino acids.

17. The amino acid binding protein according to any one of claims 8 to 16, wherein the loop between β1 and α1 contains four or more negatively charged amino acids.

18. 18. The amino acid binding protein of claim 16 or 17, wherein at least two negatively charged amino acids in the loop between β1 and α1 form hydrogen bonds with the amino acid ligand.

19. 19. The amino acid binding protein of claim 16, wherein at least one negatively charged amino acid in the loop between β1 and α1 forms a bifurcated hydrogen bond with the amino acid ligand.

20. 20. The amino acid binding protein of claim 16, wherein the negatively charged amino acid is selected from aspartic acid and glutamic acid.

21. 21. The amino acid binding protein of any one of claims 8 to 20, wherein β1-α1 comprises an amino acid sequence that is at least 80% identical to the sequence of amino acids 35 to 58 of SEQ ID NO:

1.

22. 22. The amino acid binding protein of claim 21, wherein the binding pocket is formed by amino acids at one or more positions corresponding to amino acids 41 to 47 and 50 of SEQ ID NO:

1.

23. 23. The amino acid binding protein of claim 21 or 22, wherein the amino acid sequence comprises amino acid substitutions at one or more positions corresponding to L39, N41, D42, H45, V50, and Q55 of SEQ ID NO:

1.

24. 24. The amino acid binding protein of claim 23, wherein at least one amino acid substitution is at a position corresponding to N41 of SEQ ID NO:

1.

25. 25. The amino acid binding protein of claim 23 or 24, wherein the amino acid substitution is selected from N41D and Q55R.

26. 26. The amino acid binding protein of any one of claims 8 to 25, wherein α2 comprises an amino acid sequence that is at least 80% identical to the sequence of amino acids 62 to 73 of SEQ ID NO:

1.

27. 27. The amino acid binding protein of claim 26, wherein the binding pocket is formed by amino acids at one or more positions corresponding to amino acids 69, 72, and 73 of SEQ ID NO:

1.

28. 28. The amino acid binding protein of claim 26 or 27, wherein the amino acid sequence comprises amino acid substitutions at one or more positions corresponding to P62, E63, L68, and V72 of SEQ ID NO:

1.

29. 29. The amino acid binding protein of claim 28, wherein at least one amino acid substitution is at a position corresponding to V72 of SEQ ID NO:

1.

30. 30. The amino acid binding protein of claim 28 or 29, wherein the amino acid substitution is selected from E63S, L68M, and V72M.

31. 31. The amino acid binding protein of claim 8, wherein the loop between α3 and β3 comprises an amino acid sequence that is at least 80% identical to the sequence of amino acids 99 to 112 of SEQ ID NO:

1.

32. 32. The amino acid binding protein of claim 31 , wherein the binding pocket is formed by an amino acid at a position corresponding to amino acid 111 of SEQ ID NO:

1.

33. 33. The amino acid binding protein of claim 31 or 32, wherein the amino acid sequence comprises an amino acid substitution at one or more positions corresponding to Y100 and M111 of SEQ ID NO:

1.

34. 34. The amino acid binding protein of claim 33, wherein the amino acid substitution is Y100R.

35. 35. The amino acid binding protein of any one of claims 8 to 34, wherein the plurality of hydrogen bond acceptors in the binding pocket are configured to form at least two, at least three, at least four, or at least five hydrogen bonds in the presence of the amino acid ligand.

36. 36. The amino acid binding protein of any one of claims 8 to 35, wherein the binding pocket comprises two or more of (i), (ii), (iii), (iv), and (v).

37. 37. The amino acid binding protein of any one of claims 8 to 36, wherein the binding pocket comprises three or more of (i), (ii), (iii), (iv), and (v).

38. 38. The amino acid binding protein of any one of claims 8 to 37, wherein the binding pocket comprises four or more of (i), (ii), (iii), (iv), and (v).

39. 39. The amino acid binding protein of any one of claims 8 to 38, wherein the binding pocket comprises (i), (ii), (iii), (iv), and (v).

40. The amino acid binding protein of any one of claims 8 to 39, wherein the structural equivalent is a structure in which at least 80% of the secondary structural α carbon atoms are aligned with the structure of formula (I) with a root mean square difference of 5 Å or less.

41. 41. The amino acid binding protein of any one of claims 8 to 40, wherein the structural equivalent is a structure in which at least 80% of the secondary structural alpha carbon atoms are aligned with the structure of formula (I) with a root mean square difference of 4 Å or less, 3 Å or less, 2 Å or less, or 1 Å or less.

42. The amino acid binding proteins are selected from the group consisting of PS635-645, PS731-732, PS759-766, PS769, PS795-870, PS896-912, PS918-1043, PS1048-1100, PS1124-1137, PS1141-1161, PS1175-1199, PS1203-1217, PS1222-1245, PS1277-1305, PS1321-1350, and PS1425-1448 (SEQ ID NOs: 22-27, 87-88, 115-122, 125, 151-226, 249-265, 271-390, 395-446, 470-483, 487-507, 521-545, 549-563, 568-591, 622-650, 664-693, and 768-791).

43. 1. A recombinant or synthetic amino acid binding protein having an amino acid sequence at least 80% identical to SEQ ID NO:2, wherein the amino acid sequence contains amino acid substitutions at one or more positions corresponding to G19, K26, S29, F30, D31, D32, T33, C34, V35, T47, G48, T53, T54, T57, E58, F59, N61, 163, D65, D68, E70, A71, H74, and T75 of SEQ ID NO:

2.

44. 44. The amino acid binding protein of claim 43, wherein the amino acid sequence comprises amino acid substitutions at positions corresponding to I63 and E70, and at one or more positions corresponding to G19, K26, S29, D32, T47, G48, T53, T54, T57, E58, F59, N61, H74, and T75.

45. 45. The amino acid binding protein of claim 43 or 44, wherein the amino acid sequence comprises amino acid substitutions at one or more positions corresponding to K26, D32, T47, I63, and E70.

46. 46. The amino acid binding protein of any one of claims 43 to 45, wherein the amino acid substitutions are selected from G19R, K26R, S29Q, D32R / Y, T47K / L / R, G48R / Y, T53V, T54K, T57K / R, E58K, F59R, N61K, I63E, E70S / T, H74K, and T75E.

47. 47. The amino acid binding protein of any one of claims 43 to 46, wherein the amino acid substitution is selected from K26R, D32R, T47L, I63E, and E70T.

48. 48. The amino acid binding protein of any one of claims 43 to 47, wherein the amino acid sequence is at least 85%, at least 90%, at least 95%, at least 98%, 80-98%, 80-95%, 80-90%, 85-95%, or 90-98% identical to SEQ ID NO:

2.

49. The amino acid sequence of any one of claims 43 to 48, wherein the amino acid sequence is selected from any one of PS1101 to 1122, PS1218 to 1221, and PS1351 to 1398 (SEQ ID NOs: 447 to 468, 564 to 567, and 694 to 741).

50. A recombinant or synthetic amino acid binding protein having the structure of formula (II) or a structural equivalent thereof: β1-α1-β2-α2-α3 (II) (In the formula, each of β1 and β2 is a β strand; each of α1, α2, and α3 is an α-helix; Each instance of "-" is a loop, At least a portion of each of α2, the loop between β1 and α1, and the loop between β2 and α2 form a binding pocket for an amino acid ligand, said binding pocket comprising: i) Approximately 200 Å 3 The volume of ii) -3.0RTe c -1 The electrostatic potential, iii) a plurality of hydrogen bond acceptors configured to form one or more hydrogen bonds in the presence of the amino acid ligand; and iv) a plurality of van der Waals contact locations configured to form van der Waals interactions in the presence of said amino acid ligand; (including one or more of the following) an amino acid binding protein,

51. 51. The amino acid binding protein of claim 50, wherein the amino acid ligand is a polypeptide comprising at least three amino acids.

52. 52. The amino acid binding protein of claim 50 or 51, wherein the amino acid ligand comprises the N-terminal amino acid of a polypeptide.

53. 53. The amino acid binding protein of claim 52, wherein the N-terminal amino acid is arginine.

54. The amino acid binding protein has a dissociation constant (K) of less than 2,000 nM, less than 1,500 nM, less than 1,000 nM, less than 750 nM, less than 500 nM, less than 400 nM, less than 200 nM, less than 100 nM, 10-2,000 nM, 25-1,000 nM, 50-500 nM, 10-400 nM, 25-75 nM, or 50-80 nM. D 54. The amino acid binding protein of claim 53, which binds to an N-terminal arginine at

55. 55. The amino acid binding protein of any one of claims 50 to 54, wherein the amino acid binding protein is at least 50 amino acids in length, at least 75 amino acids in length, at least 100 amino acids in length, 50 to 250 amino acids in length, 50 to 150 amino acids in length, or 100 to 200 amino acids in length.

56. 56. The amino acid binding protein of any one of claims 50 to 55, wherein the loop between β2 and α2 contains three or more negatively charged amino acids.

57. 57. The amino acid binding protein of any one of claims 50 to 56, wherein the loop between β2 and α2 contains four or more negatively charged amino acids.

58. 58. The amino acid binding protein of claim 56 or 57, wherein at least three negatively charged amino acids in the loop between β2 and α2 form hydrogen bonds with the amino acid ligand.

59. 59. The amino acid binding protein of any one of claims 56 to 58, wherein at least one negatively charged amino acid in the loop between β2 and α2 forms a hydrogen bond with the amino terminus of the amino acid ligand.

60. 60. The amino acid binding protein of any one of claims 56 to 59, wherein the negatively charged amino acid is selected from aspartic acid and glutamic acid.

61. 61. The amino acid binding protein of any one of claims 50 to 60, wherein at least one amino acid of α2 forms a hydrogen bond with the amino acid ligand.

62. 62. The amino acid binding protein of any one of claims 50 to 61, wherein α2 comprises one or more polar uncharged amino acids.

63. 63. The amino acid binding protein of claim 62, wherein at least one polar, uncharged amino acid of α2 forms a hydrogen bond with a side chain of the amino acid ligand.

64. 64. The amino acid binding protein of any one of claims 50 to 63, wherein the loop between β1 and α1 comprises an amino acid sequence that is at least 80% identical to the sequence of amino acids 27 to 42 of SEQ ID NO:

2.

65. 65. The amino acid binding protein of claim 64, wherein the binding pocket is formed by amino acids at one or more positions corresponding to amino acids 31, 32, and 34-36 of SEQ ID NO:

2.

66. 66. The amino acid binding protein of any one of claims 50 to 65, wherein the loop between α1 and β2 comprises an amino acid sequence that is at least 50% identical to the sequence of amino acids 47 to 50 of SEQ ID NO:

2.

67. 67. The amino acid binding protein of claim 66, wherein the amino acid sequence comprises an amino acid substitution at a position corresponding to T47 in SEQ ID NO:

2.

68. 68. The amino acid binding protein of claim 67, wherein the amino acid substitution is T47L.

69. 69. The amino acid binding protein of any one of claims 50 to 68, wherein β2-α2 comprises an amino acid sequence that is at least 80% identical to the sequence of amino acids 51 to 71 of SEQ ID NO:

2.

70. 70. The amino acid binding protein of claim 69, wherein the binding pocket is formed by amino acids at one or more positions corresponding to amino acids 63, 65, 68, 70, and 71 of SEQ ID NO:

2.

71. 71. The amino acid binding protein of claim 69 or 70, wherein the amino acid sequence comprises an amino acid substitution at one or more positions corresponding to I63 and E70 of SEQ ID NO:

2.

72. 72. The amino acid binding protein of claim 71, wherein the amino acid substitution is selected from I63E and E70T.

73. 73. The amino acid binding protein of any one of claims 50 to 72, wherein the plurality of hydrogen bond acceptors in the binding pocket are configured to form at least two, at least three, at least four, or at least five hydrogen bonds in the presence of the amino acid ligand.

74. 74. The amino acid binding protein of any one of claims 50 to 73, wherein the binding pocket comprises two or more of (i), (ii), (iii), and (iv).

75. 75. The amino acid binding protein of any one of claims 50 to 74, wherein the binding pocket comprises three or more of (i), (ii), (iii), and (iv).

76. 76. The amino acid binding protein of any one of claims 50 to 75, wherein the binding pocket comprises (i), (ii), (iii), and (iv).

77. 77. The amino acid binding protein of any one of claims 50 to 76, wherein the structural equivalent is a structure in which at least 80% of the secondary structural alpha carbon atoms are aligned with the structure of formula (II) with a root mean square difference of 5 Å or less.

78. 78. The amino acid binding protein of any one of claims 50 to 77, wherein the structural equivalent is a structure in which at least 80% of the secondary structural alpha carbon atoms are aligned with the structure of formula (II) with a root mean square difference of 4 Å or less, 3 Å or less, 2 Å or less, or 1 Å or less.

79. 79. The amino acid binding protein of any one of claims 50 to 78, wherein the amino acid binding protein has an amino acid sequence that is at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, 40-100%, 50-100%, 60-100%, 70-100%, 80-100%, 90-100%, or 95-100% identical to a sequence selected from any one of PS1101-1122, PS1218-1221, and PS1351-1398 (SEQ ID NOs: 447-468, 564-567, and 694-741).

80. 1. A recombinant or synthetic amino acid binding protein having an amino acid sequence that is at least 80% identical to SEQ ID NO:3, wherein the amino acid sequence contains amino acid substitutions at one or more positions corresponding to S22, C23, Y24, C25, E26, S39, W75, D76, Y77, H78, C85, N120, H145, and M146 of SEQ ID NO:

3.

81. 81. The amino acid binding protein of claim 80, wherein the amino acid sequence comprises amino acid substitutions at positions corresponding to S22, C25, H78, C85, and N120.

82. 82. The amino acid binding protein of claim 80 or 81, wherein the amino acid substitution is selected from S22E, C25S, H78Q, H78K, C85T, N120R, and M146E.

83. 83. The amino acid binding protein of any one of claims 80 to 82, wherein the amino acid substitution is selected from S22E, C25S, H78Q, H78K, C85T, and N120R.

84. 84. The amino acid binding protein of any one of claims 80 to 83, wherein the amino acid sequence is at least 85%, at least 90%, at least 95%, at least 98%, 80-98%, 80-95%, 80-90%, 85-95%, or 90-98% identical to SEQ ID NO:

3.

85. The amino acid sequence of any one of claims 80 to 84, wherein the amino acid sequence is selected from any one of PS1258-1260, PS1315-1318, PS1457-1478, PS1480-1499, PS1633-1656, PS1737-1758, PS1821-1898, PS2014-2057, and PS2116-2137 (SEQ ID NOs: 604-606, 660-663, 792-833, and 836-1025).

86. A recombinant or synthetic amino acid binding protein having the structure of formula (III) or a structural equivalent thereof: α1-α2-α3β1-β2-β3-β4-β5-α4-β6α5-α6 (III) (In the formula, each of α1, α2, α3, α4, α5, and α6 is an α-helix; each of β1, β2, β3, β4, β5, and β6 is a β chain; Each instance of "-" is a loop, At least a portion of each of α2, β3, β4, α5, the loop between α1 and α2, and the loop between β3 and β4 form a binding pocket for an amino acid ligand, said binding pocket comprising: i) approximately 160 Å 3 The volume of ii) -2.0RTe c -1 The electrostatic potential, iii) a plurality of hydrogen bond acceptors or donors configured to form one or more hydrogen bonds in the presence of said amino acid ligand; iv) a plurality of van der Waals contact locations configured to form van der Waals interactions in the presence of the amino acid ligand; and v) at least one negatively charged amino acid and at least one positively charged amino acid; (including one or more of the following) an amino acid binding protein,

87. 87. The amino acid binding protein of claim 86, wherein the amino acid ligand is a polypeptide comprising at least three amino acids.

88. 88. The amino acid binding protein of claim 86 or 87, wherein the amino acid ligand comprises the N-terminal amino acid of a polypeptide.

89. 89. The amino acid binding protein of claim 88, wherein the N-terminal amino acid is selected from glutamine, asparagine, glutamic acid, aspartic acid, and cysteine-S-acetamide.

90. The amino acid binding protein has a dissociation constant (K) of less than 2,000 nM, less than 1,500 nM, less than 1,000 nM, less than 750 nM, less than 500 nM, less than 250 nM, less than 150 nM, less than 100 nM, less than 50 nM, 10-2,000 nM, 25-1,000 nM, 50-500 nM, 10-100 nM, 25-250 nM, or 50-150 nM. D 90. The amino acid binding protein of claim 89, which binds to the N-terminal glutamine at

91. 91. The amino acid binding protein of any one of claims 86 to 90, wherein the amino acid binding protein is at least 50 amino acids in length, at least 75 amino acids in length, at least 100 amino acids in length, 50 to 250 amino acids in length, 50 to 150 amino acids in length, or 100 to 200 amino acids in length.

92. 92. The amino acid binding protein of any one of claims 86 to 91, wherein each of α2 and β4 comprises at least one polar uncharged amino acid that forms a hydrogen bond with the amino acid ligand.

93. 93. The amino acid binding protein of claim 92, wherein the at least one polar, uncharged amino acid of α2 is serine.

94. 94. The amino acid binding protein of claim 92 or 93, wherein the at least one polar, uncharged amino acid of β4 is glutamine.

95. 95. The amino acid binding protein of any one of claims 86 to 94, wherein α2 and the loop between α1 and α2 comprise an amino acid sequence that is at least 80% identical to the sequence of amino acids 18 to 40 of SEQ ID NO:

3.

96. 96. The amino acid binding protein of claim 95, wherein the binding pocket is formed by amino acids at one or more positions corresponding to amino acids 23-26 of SEQ ID NO:

3.

97. 97. The amino acid binding protein of claim 95 or 96, wherein the amino acid sequence comprises an amino acid substitution at a position corresponding to C25 of SEQ ID NO:

3.

98. 98. The amino acid binding protein of claim 97, wherein the amino acid substitution is C25S.

99. 99. The amino acid binding protein of any one of claims 86 to 98, wherein β3-β4 comprises an amino acid sequence that is at least 80% identical to the sequence of amino acids 73 to 85 of SEQ ID NO:

3.

100. 100. The amino acid binding protein of claim 99, wherein the binding pocket is formed by amino acids at one or more positions corresponding to amino acids 75-78 of SEQ ID NO:

3.

101. 101. The amino acid binding protein of claim 99 or 100, wherein the amino acid sequence comprises an amino acid substitution at a position corresponding to H78 in SEQ ID NO:

3.

102. 102. The amino acid binding protein of claim 101, wherein the amino acid substitution is H78Q or H78K.

103. 103. The amino acid binding protein of any one of claims 86 to 102, wherein α6 comprises an amino acid sequence that is at least 66% identical to the sequence of amino acids 144 to 146 of SEQ ID NO:

3.

104. 104. The amino acid binding protein of claim 103, wherein the binding pocket is formed by amino acids at one or more positions corresponding to amino acids 145-146 of SEQ ID NO:

3.

105. 105. The amino acid binding protein of any one of claims 86 to 104, wherein the plurality of hydrogen bond acceptors or donors in the binding pocket are configured to form at least two, at least three, at least four, or at least five hydrogen bonds in the presence of the amino acid ligand.

106. 106. The amino acid binding protein of any one of claims 86 to 105, wherein the binding pocket comprises two or more of (i), (ii), (iii), (iv), and (v).

107. 107. The amino acid binding protein of any one of claims 86 to 106, wherein the binding pocket comprises three or more of (i), (ii), (iii), (iv), and (v).

108. 108. The amino acid binding protein of any one of claims 86 to 107, wherein the binding pocket comprises four or more of (i), (ii), (iii), (iv), and (v).

109. 109. The amino acid binding protein of any one of claims 86 to 108, wherein the binding pocket comprises (i), (ii), (iii), and (iv).

110. 110. The amino acid binding protein of claim 109, wherein the amino acid ligand is a polypeptide comprising an N-terminal glutamine or asparagine.

111. 109. The amino acid binding protein of any one of claims 86 to 108, wherein the binding pocket comprises (i), (ii), (iii), (iv), and (v).

112. 112. The amino acid binding protein of claim 111, wherein the amino acid ligand is a polypeptide comprising an N-terminal glutamic acid.

113. Structure of formula (III-A) or a structural equivalent thereof: α1-α2-α3β1-β2-β3-β4-β5-α4-β6α5-α6-α7-β7α8 (III-A) (In the formula, each of α7 and α8 is an α-helix; β7 is the β chain) The amino acid binding protein of any one of claims 86 to 112, comprising:

114. A recombinant or synthetic amino acid binding protein having the structure of formula (III-B) or a structural equivalent thereof: α1-α2-α3β1-β2-β3-β4 (III-B) (In the formula, each of α1, α2, and α3 is an α-helix; each of β1, β2, β3, and β4 is a β chain; Each instance of "-" is a loop, At least a portion of each of α2, β3, β4, the loop between α1 and α2, and the loop between β3 and β4 form a binding pocket for an amino acid ligand, said binding pocket comprising: i) at least one negatively charged amino acid configured to form a hydrogen bond with said amino acid ligand; and ii) at least one positively charged amino acid configured to form a hydrogen bond with said amino acid ligand; (including an amino acid binding protein,

115. 115. The amino acid binding protein of claim 114, wherein the amino acid ligand is a polypeptide comprising at least three amino acids.

116. 116. The amino acid binding protein of claim 114 or 115, wherein the amino acid ligand comprises the N-terminal amino acid of a polypeptide.

117. 117. The amino acid binding protein of claim 116, wherein the N-terminal amino acid is glutamic acid.

118. said at least one negatively charged amino acid forms said hydrogen bond with a main chain atom of said amino acid ligand; the at least one positively charged amino acid forms the hydrogen bond with a side chain atom of the amino acid ligand; An amino acid binding protein according to any one of claims 114 to 117.

119. a side chain atom of said at least one negatively charged amino acid forms said hydrogen bond with said main chain atom of said amino acid ligand; a side chain atom of the at least one positively charged amino acid forms the hydrogen bond with a side chain atom of the amino acid ligand; An amino acid binding protein according to any one of claims 114 to 118.

120. α2 comprises said at least one negatively charged amino acid; β4 comprises said at least one positively charged amino acid; An amino acid binding protein according to any one of claims 114 to 119.

121. the at least one negatively charged amino acid comprises glutamic acid; the at least one positively charged amino acid comprises lysine; An amino acid binding protein according to any one of claims 114 to 120.

122. the at least one negatively charged amino acid corresponds to E26 of SEQ ID NO:3; the at least one positively charged amino acid is a lysine substitution at a position corresponding to H78 in SEQ ID NO: 3; An amino acid binding protein according to any one of claims 114 to 121.

123. 123. The amino acid binding protein of any one of claims 114 to 122, wherein α1-α2 comprises an amino acid sequence that is at least 80% identical to the sequence of amino acids 15 to 39 of SEQ ID NO:

3.

124. 124. The amino acid binding protein of claim 123, wherein the amino acid sequence comprises amino acid substitutions at positions corresponding to S22 and C25 of SEQ ID NO:

3.

125. 125. The amino acid binding protein of claim 124, wherein the amino acid substitutions are S22E and C25S.

126. 126. The amino acid binding protein of any one of claims 114 to 125, wherein β3-β4 comprises an amino acid sequence that is at least 80% identical to the sequence of amino acids 73 to 85 of SEQ ID NO:

3.

127. 127. The amino acid binding protein of claim 126, wherein the amino acid sequence comprises amino acid substitutions at positions corresponding to H78 and C85 of SEQ ID NO:

3.

128. 128. The amino acid binding protein of claim 127, wherein the amino acid substitutions are H78K and C85T.

129. The structural equivalent is a structure in which at least 80% of the secondary structural α carbon atoms are aligned with the structure of formula (III), (III-A) or (III-B) with a root mean square difference of 5 Å or less.

130. 130. The amino acid binding protein of any one of claims 86 to 129, wherein the structural equivalent is a structure in which at least 80% of the secondary structural alpha carbon atoms align with the structure of formula (III), (III-A) or (III-B) with a root mean square difference of 4 Å or less, 3 Å or less, 2 Å or less, or 1 Å or less.

131. The amino acid binding protein is selected from any one of PS1258-1260, PS1315-1318, PS1457-1478, PS1480-1499, PS1633-1656, PS1737-1758, PS1821-1898, PS2014-2057, and PS2116-2137 (SEQ ID NOs: 604-606, 660-663, 792-833, and 836-1025).

131. The amino acid binding protein of any one of claims 86 to 130, having an amino acid sequence that is at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, 40-100%, 50-100%, 60-100%, 70-100%, 80-100%, 90-100%, or 95-100% identical to a sequence

132. 132. The amino acid binding protein of any one of claims 114 to 131, further comprising one or more labels.

133. 133. The amino acid binding protein of claim 132, wherein the one or more labels comprise a luminescent label or a conductive label.

134. 134. The amino acid binding protein of claim 133, wherein the luminescent label comprises at least one fluorophore dye molecule.

135. 135. The amino acid binding protein of claim 133 or 134, wherein the luminescent label comprises 20 or fewer fluorophore dye molecules.

136. 136. The amino acid binding protein of any one of claims 133 to 135, wherein the luminescent label comprises at least one FRET pair comprising a donor label and an acceptor label.

137. The amino acid binding protein of any one of claims 133 to 136, wherein the conductive label comprises a charged polymer.

138. 138. The amino acid binding protein of any one of claims 132 to 137, wherein the one or more labels comprise a tag sequence.

139. 139. The amino acid binding protein of claim 138, wherein the tag sequence comprises one or more of a purification tag, a cleavage site, and a biotinylation sequence.

140. 140. The amino acid binding protein of claim 139, wherein the biotinylation sequence comprises at least one biotin ligase recognition sequence.

141. 141. The amino acid binding protein of claim 139 or 140, wherein the biotinylation sequence comprises two biotin ligase recognition sequences oriented in tandem.

142. 142. The amino acid binding protein of any one of claims 132 to 141, wherein the one or more labels comprise a biotin moiety.

143. 143. The amino acid binding protein of claim 142, wherein the biotin moiety comprises at least one biotin molecule.

144. 144. The amino acid binding protein of claim 142 or 143, wherein the biotin moiety is a bisbiotin moiety.

145. 145. The amino acid binding protein of claim 143 or 144, wherein the label comprises at least one biotin ligase recognition sequence to which the at least one biotin molecule is attached.

146. 146. The amino acid binding protein of any one of claims 132 to 145, wherein the one or more labels comprise one or more polyol moieties.

147. 147. The amino acid binding protein of claim 146, wherein the one or more polyol moieties comprise dextran, polyvinylpyrrolidone, polyethylene glycol, polypropylene glycol, polyoxyethylene glycol, polyvinyl alcohol, or a combination or variation thereof.

148. 148. The amino acid binding protein of any one of claims 132 to 147, wherein the amino acid binding protein comprises one or more unnatural amino acids having the one or more labels attached thereto.

149. 149. An amino acid recognition factor comprising a polypeptide having at least a first amino acid binding protein and a second amino acid binding protein linked end-to-end, wherein the first amino acid binding protein and the second amino acid binding protein are separated by a linker comprising at least two amino acids, and at least one of the first amino acid binding protein and the second amino acid binding protein is an amino acid binding protein described in any one of claims 1 to 148.

150. 150. The amino acid recognition factor of claim 149, wherein the first amino acid binding protein and the second amino acid binding protein are the same.

151. 150. The amino acid recognition factor of claim 149, wherein the first amino acid binding protein and the second amino acid binding protein are different.

152. 152. An amino acid recognition factor according to any one of claims 149 to 151, wherein each of the first and second amino acid binding proteins is independently an amino acid binding protein according to any one of claims 1 to 148.

153. 153. The amino acid recognition element of any one of claims 149 to 152, wherein the linker comprises up to 100 amino acids, up to 80 amino acids, up to 60 amino acids, up to 50 amino acids, from about 5 to about 100 amino acids, or from about 5 to about 50 amino acids.

154. the first amino acid binding protein has an amino acid sequence at least 80% identical to PS961 (SEQ ID NO: 314); The amino acid recognition factor of any one of claims 149 to 153, wherein the second amino acid binding protein has an amino acid sequence that is at least 80% identical to PS961 (SEQ ID NO: 314).

155. The amino acid recognition element of claim 154, comprising an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, 80-100%, 85-100%, 90-100%, 95-100%, or 100% identical to a sequence selected from any one of PS1038, PS1222, and PS1223 (SEQ ID NOs: 389, 568, and 569).

156. the first amino acid binding protein has an amino acid sequence at least 80% identical to PS1122 (SEQ ID NO:468); 154. The amino acid recognition element of any one of claims 149 to 153, wherein the second amino acid binding protein has an amino acid sequence that is at least 80% identical to PS1122 (SEQ ID NO: 468).

157. The amino acid recognition element of claim 156, comprising an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, 80-100%, 85-100%, 90-100%, 95-100%, or 100% identical to a sequence selected from any one of PS1219 to PS1221 (SEQ ID NOs: 565-567).

158. the first amino acid binding protein has an amino acid sequence at least 80% identical to PS1259 (SEQ ID NO: 605); 154. The amino acid recognition factor of any one of claims 149 to 153, wherein the second amino acid binding protein has an amino acid sequence that is at least 80% identical to PS1259 (SEQ ID NO: 605).

159. The amino acid recognition element of claim 158, comprising an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, 80-100%, 85-100%, 90-100%, 95-100%, or 100% identical to PS1599 (SEQ ID NO: 835).

160. An amino acid recognition factor comprising a polypeptide having an amino acid binding protein and a labeled protein linked end-to-end, wherein the amino acid binding protein and the labeled protein are separated by a linker comprising at least two amino acids, and the amino acid binding protein is an amino acid binding protein described in any one of claims 1 to 148.

161. The amino acid recognizer of claim 160, wherein the labeled protein has a molecular weight of at least 10 kDa, about 10 kDa to about 150 kDa, or about 15 kDa to about 100 kDa.

162. 162. The amino acid recognition element of claim 160 or 161, wherein the labeled protein comprises at least 50 amino acids, from about 50 to about 1,000 amino acids, or from about 100 to about 750 amino acids.

163. The amino acid recognition factor of any one of claims 160 to 162, wherein the labeled protein comprises a fluorescent protein.

164. The amino acid recognition factor of any one of claims 160 to 163, wherein the labeled protein comprises a luminescent label.

165. The amino acid recognition factor of any one of claims 160 to 164, wherein the labeled protein comprises a protein selected from maltose binding protein, glutathione S-transferase, green fluorescent protein, SNAP-tag, and DNA polymerase.

166. 166. The amino acid recognition element of any one of claims 160 to 165, wherein the linker comprises up to 100 amino acids, up to 80 amino acids, up to 60 amino acids, up to 50 amino acids, from about 5 to about 100 amino acids, or from about 5 to about 50 amino acids.

167. A composition comprising two or more amino acid recognition factors, wherein at least one amino acid recognition factor is an amino acid binding protein according to any one of claims 1 to 166.

168. The composition of claim 167, wherein the composition comprises at least one amino acid binding protein according to any one of claims 1 to 42.

169. 169. The composition of claim 167 or 168, wherein the composition comprises at least one amino acid binding protein of any one of claims 43 to 79.

170. 170. The composition of any one of claims 167 to 169, wherein said composition comprises at least one amino acid binding protein of any one of claims 80 to 131.

171. The composition comprising: A first amino acid binding protein according to any one of claims 1 to 42; A second amino acid binding protein according to any one of claims 43 to 79; A third amino acid binding protein according to any one of claims 80 to 131; The composition of any one of claims 167 to 170, comprising:

172. 172. The composition of any one of claims 167 to 171, wherein the composition comprises at least one type of cleavage reagent.

173. 173. The composition of claim 172, wherein the cleavage reagent comprises an exopeptidase.

174. 174. The composition of claim 172 or 173, wherein the cleavage reagent comprises an aminopeptidase.

175. 175. The composition of any one of claims 172-174, wherein the molar ratio of amino acid recognition factor to cleavage reagent in the composition is about 1:1,000 to about 1:1, about 1:1 to about 100:1, about 1:100 to about 1:1, about 1:1 to about 10:1, about 1:1,000, about 1:500, about 1:200, about 1:100, about 1:10, about 1:5, about 1:2, about 1:1, about 5:1, about 10:1, about 50:1, or about 100:

1.

176. 1. A method for determining at least one chemical characteristic of a polypeptide, said method comprising: contacting a polypeptide with a composition according to any one of claims 167 to 175; monitoring a signal for a signal pulse corresponding to an interaction between one or more amino acid recognition factors and the polypeptide; determining at least one chemical characteristic of the polypeptide based on a characteristic pattern in the signal; A method comprising:

177. 177. The method of claim 176, wherein determining the at least one chemical characteristic comprises identifying at least one amino acid in the polypeptide as a naturally occurring amino acid, a non-natural amino acid, or a modified variant thereof.

178. 178. The method of claim 176 or 177, wherein determining the at least one chemical characteristic comprises identifying at least an amino acid in the polypeptide as having a side chain that is negatively charged, positively charged, uncharged, polar, non-polar, hydrophobic, aromatic, or a combination thereof.

179. 179. The method of any one of claims 176 to 178, wherein determining the at least one chemical characteristic comprises identifying at least one amino acid in the polypeptide as one type selected from alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, selenocysteine, serine, threonine, tryptophan, tyrosine, and valine.

180. 180. The method of any one of claims 176 to 179, wherein determining the at least one chemical characteristic comprises identifying at least one amino acid in the polypeptide as having a post-translational modification.

181. 181. The method of any one of claims 176 to 180, wherein the post-translational modification is selected from acetylation, ADP-ribosylation, caspase cleavage, citrullination, formylation, N-linked glycosylation, O-linked glycosylation, hydroxylation, methylation, myristoylation, NEDDylation, nitration, oxidation, palmitoylation, phosphorylation, prenylation, S-nitrosylation, sulfation, sumoylation, and ubiquitination.

182. The monitoring 182. The method of any one of claims 176 to 181, comprising detecting a series of signal pulses, wherein a characteristic pattern in the series of signal pulses is indicative of said at least one chemical characteristic of said polypeptide.

183. 183. The method of claim 182, wherein the series of signal pulses corresponds to a respective series of binding events between the one or more amino acid recognition factors and the polypeptide.

184. 184. The method of claim 182 or 183, wherein the characteristic pattern of signal pulses comprises an average pulse duration of about 1 millisecond to about 10 seconds, about 50 milliseconds to about 2 seconds, about 50 milliseconds to about 500 milliseconds, or about 500 milliseconds to about 2 seconds.

185. 185. The method of any one of claims 182 to 184, wherein the characteristic pattern in the series of signal pulses comprises at least 10 signal pulses, about 50 to about 200 signal pulses, or about 25 to about 100 signal pulses.

186. 1. A system comprising: at least one hardware processor; at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by said at least one hardware processor, cause said at least one hardware processor to perform the method of any one of claims 176 to 185; and Including, the system.

187. At least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by at least one hardware processor, cause the at least one hardware processor to perform the method of any one of claims 176 to 185.