Compositions and methods for polypeptide analysis

By developing recombinant or synthetic amino acid binding proteins with improved binding characteristics, the difficulties in the application of sequencing technology in proteomics are solved, and the accuracy and efficiency of peptide structure analysis are improved.

CN119998660APending Publication Date: 2025-05-13QUANTUM SI INC
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202380071246.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-08-04
Filing Date
2023-08-03
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In proteomics, due to the size, dynamic range and inability to amplify the source, it is difficult for the existing technology to effectively apply sequencing technology, which in turn affects the understanding of complex human diseases.

Method used

A recombinant or synthetic amino acid binding protein is provided that has an amino acid sequence at least 80% identical to a particular amino acid sequence and comprises a specific structure such as β1-α1-α2-β2-α3-β3 or β1-α1-β2-α2-α3 with improved binding properties.

Benefits of technology

Through these improved amino acid binding proteins, it is possible to more effectively identify and bind specific amino acid ligands, improving the accuracy and efficiency of peptide structural analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119998660A_ABST
    Figure CN119998660A_ABST
Patent Text Reader

Abstract

Provided is an amino acid recognition agent having improved binding properties, which is capable of obtaining more structural information from a polypeptide based on the kinetics of the on-off binding between the recognition agent and the polypeptide. The amino acid recognition agent may comprise an amino acid binding protein having an engineered binding pocket with one or more modifications relative to a homologous protein. Compared with an unmodified homologous protein binding pocket, the modified binding pocket can increase the number of interactions formed between the binding pocket and an amino acid ligand, and can increase the number of types of amino acid ligands capable of detecting binding compared with the unmodified homologous protein binding pocket. According to the present invention, the amino acid ligands of one or more types can be bound to one or more types of amino acid ligands, and can improve kinetics (e.g., KD, koff, kon) binding to one or more types of amino acid ligands, which advantageously increases the amount or confidence of structural information obtained from polypeptide analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application claims priority under 35 USC §119(e) to U.S. Provisional Patent Application No. 63 / 395,328, filed on August 4, 2022, which is incorporated herein by reference in its entirety. Background Art

[0002] Proteins represent the fundamental building blocks of life, driving critical biological and cellular processes. The function of a protein is driven by its structure, including its sequence. In adjacent fields such as genomics, advances in sequencing technology have proven invaluable in improving our understanding of the progression of complex human diseases. Applying similar approaches to proteomics has been difficult due to the size and dynamic range of the proteome and the inability to amplify sources. Summary of the invention

[0003] In some embodiments, a recombinant or synthetic amino acid binding protein is provided, which has an amino acid sequence that is at least 80% identical to SEQ ID NO:1, wherein the amino acid sequence comprises amino acid substitutions at one or more positions corresponding to E22, R31, L39, N41, D42, D43, D44, H45, T46, Y47, V50, Q55, P62, E63, L68, A69, V72, D73, Q75, Y100 and M111 of SEQ ID NO:1.

[0004] In some embodiments, a recombinant or synthetic amino acid binding protein is provided, comprising a structure of formula (I) or a structural equivalent thereof: β1-α1-α2-β2-α3-β3 (I), wherein: each of β1, β2 and β3 is a β strand; each of α1, α2 and α3 is an α helix; each instance of “–” is a loop; and at least a portion of each of α1, α2, the loop between β1 and α1, and the loop between α3 and β3 forms a binding pocket for an amino acid ligand, wherein the binding pocket comprises one or more of the following: (i) a volume of about (ii) Electrostatic potential is -3.0RTe c -1 or less, (iii) at least 35% of the amino acids forming the binding pocket have negatively charged side chains, (iv) a plurality of hydrogen bond acceptors that are configured to form one or more hydrogen bonds in the presence of the amino acid ligand, and (v) a plurality of van der Waals contact positions that are configured to form van der Waals interactions in the presence of the amino acid ligand.

[0005] In some embodiments, a recombinant or synthetic amino acid binding protein is provided, which has an amino acid sequence that is at least 80% identical to SEQ ID NO:2, wherein the amino acid sequence comprises amino acid substitutions at one or more positions corresponding to G19, K26, S29, F30, D31, D32, T33, C34, V35, T47, G48, T53, T54, T57, E58, F59, N61, I63, D65, D68, E70, A71, H74 and T75 of SEQ ID NO:2.

[0006] In some embodiments, a recombinant or synthetic amino acid binding protein is provided, comprising a structure of formula (II) or a structural equivalent thereof: β1-α1-β2-α2-α3 (II), wherein: each of β1 and β2 is a β strand; each of α1, α2 and α3 is an α helix; each instance of “–” is a loop; and at least a portion of each of α2, the loop between β1 and α1, and the loop between β2 and α2 forms a binding pocket for an amino acid ligand, wherein the binding pocket comprises one or more of the following: (i) a volume of about (ii) Electrostatic potential is -3.0RTe c -1 or less, (iii) a plurality of hydrogen bond acceptors configured to form one or more hydrogen bonds in the presence of the amino acid ligand, and (iv) a plurality of van der Waals contact sites configured to form van der Waals interactions in the presence of the amino acid ligand.

[0007] In some embodiments, a recombinant or synthetic amino acid binding protein is provided, which has an amino acid sequence that is at least 80% identical to SEQ ID NO:3, wherein the amino acid sequence comprises amino acid substitutions at one or more positions corresponding to S22, C23, Y24, C25, E26, S39, W75, D76, Y77, H78, C85, N120, H145 and M146 of SEQ ID NO:3.

[0008] In some embodiments, a recombinant or synthetic amino acid binding protein is provided, comprising a structure of formula (III) or a structural equivalent thereof: α1–α2–α3β1–β2–β3–β4–β5–α4–β6α5–α6 (III), wherein: each of α1, α2, α3, α4, α5, and α6 is an α helix; each of β1, β2, β3, β4, β5, and β6 is a β strand; each instance of “–” is a loop; and at least a portion of each of α2, β3, β4, α5, the loop between α1 and α2, and the loop between β3 and β4 forms a binding pocket for an amino acid ligand, wherein the binding pocket comprises one or more of the following: (i) a volume of about (ii) Electrostatic potential is -2.0RTe c -1 or less, (iii) a plurality of hydrogen bond acceptors or donors configured to form one or more hydrogen bonds in the presence of the amino acid ligand, (iv) a plurality of van der Waals contact sites configured to form van der Waals interactions in the presence of the amino acid ligand, and (v) at least one negatively charged amino acid and at least one positively charged amino acid.

[0009] In some embodiments, an amino acid recognition agent is provided, comprising a polypeptide having at least a first amino acid binding protein and a second amino acid binding protein connected end-to-end, wherein the first amino acid binding protein and the second amino acid binding protein are separated by a linker comprising at least two amino acids, wherein at least one of the first amino acid binding protein and the second amino acid binding protein is an amino acid binding protein according to any aspect of the technology described herein.

[0010] In some embodiments, an amino acid recognition agent is provided, comprising a polypeptide having an amino acid binding protein and a marker protein linked end-to-end, wherein the amino acid binding protein and the marker protein are separated by a linker comprising at least two amino acids, wherein the amino acid binding protein is an amino acid binding protein according to any aspect of the technology described herein.

[0011] In some embodiments, a composition is provided comprising two or more amino acid recognition agents, wherein at least one amino acid recognition agent is an amino acid binding protein according to any aspect of the technology described herein.

[0012] According to some embodiments, a method for determining at least one chemical characteristic of a polypeptide is provided, the method comprising: contacting the polypeptide with a composition according to any aspect of the technology described herein; and monitoring a signal corresponding to a signal pulse of an interaction between one or more amino acid recognition agents and the polypeptide; and determining at least one chemical characteristic of the polypeptide based on a characteristic pattern in the signal.

[0013] In some embodiments, a system is provided, comprising: at least one hardware processor; and at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by the at least one hardware processor, cause the at least one hardware processor to perform a method according to any aspect of the technology described herein.

[0014] In some embodiments, at least one non-transitory computer-readable storage medium is provided, which stores processor-executable instructions that, when executed by at least one hardware processor, cause the at least one hardware processor to perform a method according to any aspect of the technology described herein.

[0015] The details of certain embodiments of the present disclosure are set forth in the detailed description. Other features, objects, and advantages of the present disclosure will be apparent from the examples, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate several embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0017] Figure 1 An example overview of real-time dynamic protein sequencing is shown. Protein samples are digested into peptide fragments, fixed in a nanoscale reaction chamber, and incubated with a mixture of freely diffusing N-terminal amino acid (NAA) recognition agents and aminopeptidases for sequencing. When one of its homologous NAAs is exposed at the N-terminus, the labeled recognition agent will bind or dissociate with the peptide, resulting in a characteristic pulse pattern. The NAA is cleaved by the aminopeptidase, exposing the next amino acid for recognition. The temporal order and binding kinetics of NAA recognition can identify peptides and are sensitive to features that regulate binding kinetics such as post-translational modifications (PTMs). SEQ ID NO: 1090 (RLIFA) is shown.

[0018] Figure 2A-2G An example of NAA recognition and dynamic sequencing is shown. Figure 2A-2C Shown PS610 ( Figure 2A )、PS961( Figure 2B ) and PS691( Figure 2C ) Example trace of single-molecule N-terminal recognition; Figures 2A-2C A scatter plot of the number of pulses per recognition segment (RS) and the average RS pulse duration (PD) for each peptide is shown in the figure, and the median PD is marked. Figure 2D An example trace of kinetic sequencing of the synthetic peptide FAAWAAYAAAADDD (SEQ ID NO: 1034) is shown; the median PD is indicated above each RS. Figure 2E-2G Shown is the dynamic sequencing of the synthetic peptide LAQFASIAAYASDDD (SEQ ID NO: 1035) using PS610 and PS961. Figure 2E Example traces are shown. Figure 2F Scatter plots of RS mean PD versus bin ratios are shown, illustrating the discrimination of bin ratios for discriminants and pulse duration for NAA. Figure 2G Shown are scatter plots of the number of pulses per RS ​​versus the mean PD of the RS, grouped by the amino acid label assigned to the RS.

[0019] Figures 3A-3GAn example of kinetic sequencing of multiple peptides with high-precision kinetic output is shown. Figures 3A-3E Dynamic sequencing of the peptide DQQRLIFAG (SEQ ID NO: 1036) is shown. Figure 3A Example traces are shown. Figure 3B Scatter plots of RS mean PD versus bin ratios are shown. Figure 3C Additional example traces of kinetic sequencing of DQQRLIFAG (SEQ ID NO: 1036) are shown. Figure 3D The distribution of durations of each RS and non-recognized segment (NRS) obtained during sequencing is shown, and the average duration is indicated. Figure 3E A kinetic profile is shown summarizing the characteristic sequencing behavior of the DQQRLIFAG (SEQ ID NO: 1036) peptide. Figure 3F-3G Dynamic sequencing of the synthetic peptides DQQIASSRLAASFAAQQYPDDD (SEQ ID NO: 1037) (top), RLAFSALGAADDD (SEQ ID NO: 1038) (middle), and EFIAWLV (SEQ ID NO: 1039) (bottom) are shown. Figure 3F An example trace for each peptide is shown. Figure 3G The corresponding kinetic profiles are shown.

[0020] Figures 4A-4E Examples of detection of single amino acid changes and PTMs are shown. Figures 4A-4B Shown are dynamic sequencing of synthetic peptides that differ by a single amino acid: RLAFAYPDDD (SEQ ID NO: 1040) (top), RLIFAYPDDD (SEQ ID NO: 1041) (middle), RLVFAYPDDD (SEQ ID NO: 1042) (bottom). Figure 4A Example traces are shown. Figure 4B Scatter plots of RS mean PD versus bin ratios are shown. Figure 4C-4D Detection of oxidized methionine using the peptide RLMFAYPDDD (SEQ ID NO: 1043) is shown. Figure 4C The average PD distribution of leucine is shown; labels indicate populations where leucine is followed by methionine (LM) or methionine sulfoxide (LMo). Figure 4D Example traces are shown where methionine is recognized by PS961 and leucine exhibits a long PD (top), or where methionine is not recognized due to oxidation (RLMoFAYPDDD (SEQ ID NO: 1044), where "Mo" is methionine sulfoxide) and leucine exhibits a short PD (bottom). Figure 4EShown are scatter plots of RS mean PD versus bin ratios for runs in which oxidation was not controlled (top) or methionine was fully oxidized (bottom).

[0021] Figures 5A-5C Examples of distinguishing peptides in a mixture and mapping the peptides to the human proteome are shown. Figure 5A An example trace is shown for sequencing a mixture of peptides DQQRLIFAG (SEQ ID NO: 1036) and RLAFSALGAADDD (SEQ ID NO: 1038) on the same chip; the chip windows indicate the location of the reaction chambers that generated the sequencing readout for each peptide. Figure 5B Shown are example traces of kinetic sequencing of two peptides DQQRLIFAGK (SEQ ID NO: 1045) (top) and EFIAWLVK (SEQ ID NO: 1046) (bottom) isolated from recombinant human proteins ubiquitin and GLP-1, respectively. Figure 5C A schematic diagram showing the identification of protein ubiquitin in an in silico digest of the human proteome matching the kinetic profile of the DQQRLIFAGK (SEQ ID NO: 1045) peptide based on kinetic information is shown. Also shown are IVNFSRLIFHHLK (SEQ ID NO: 1095), DIRLIFSNAK (SEQ ID NO: 1096), GQSRLIFTYGLTNSGK (SEQ ID NO: 1097), DQQRLIFAGK (SEQ ID NO: 1045) and DEHCLRLIFLK (SEQ ID NO: 1098).

[0022] Figures 6A-6F An example of chip operation is shown. Fig. 6A Shown is an exploded view of a compact benchtop instrument designed to support custom semiconductor chips and protein sequencing assays. Figure 6B The chip is shown to achieve electronic rejection by discarding photoelectrons from the pulsed laser and then turning to collect fluorescent photoelectrons from bound NAA recognition agents; the timing of the rejection and collection windows cycles in alternating frames between two modes (Bin 1 and Bin 0, example waveforms shown) to provide a bin ratio estimate of the dye fluorescence lifetime. Figure 6C It shows that the chip achieves more than 10,000 times of incident laser attenuation within 1ns after starting the knockout mode. Fig.6D Example pulses of long and short fluorescence lifetime dyes are shown, illustrating the difference in Bin 0 and Bin 1 signal collection. Fig. 6E Shown are the average RS bin ratio distributions collected for three dyes with different fluorescence lifetimes. Fig. 6F It is shown that the dye channel identification accuracy improves with the number of pulses captured per RS.

[0023] Figures 7A-7H Examples of recognition agent properties are shown. Figures 7A-7E Kinetic characterization of recognition agents using polarization assays is shown (Example 1, Methods). Figure 7A-7B The affinity (K) of PS610 for peptides with N-terminal phenylalanine, tyrosine and tryptophan is shown. D )( Fig. 7A ) and dissociation rate (k off )( Figure 7B ).exist Figure 7B In the figure, FAKLK(FITC)DEESILKQ (SEQ ID NO: 1099), YAKLK(FITC)DEESILKQ (SEQ ID NO: 1100) and WAKLK(FITC)DEESILKQ (SEQ ID NO: 1101) are shown. Figure 7C The affinity of PS961 for peptides with N-terminal leucine, isoleucine and valine is shown. Figure 7D-7E The affinity of PS691 for peptides with N-terminal arginine is shown ( Fig.7D ), as well as single-point polarization data measured for peptides with N-terminal arginine, lysine, and histidine ( Fig. 7E ). Figure 7F The binding energy of peptides with initial sequences of LAX and LXA calculated using the computational model (Example 1, Methods) is shown, where X = all 20 amino acids; the box plot shows the fraction of the total binding energy contributed by the amino acids at position 1 (P1), position 2 (P2), and position 3 (P3), which decreases exponentially from P1 to P3 (R 2 >0.97). Figure 7G Shown are the RS average PDs for LXA and LAX peptides determined using PS961 and for FXA and FAX peptides determined using PS610 in single-molecule assays. Figure 7HThe nonpolar solvation energy terms in the computational binding model using PS961 are shown to have a high correlation with the actual RS average PD values ​​observed in single molecule assays for peptides containing an N-terminal leucine and different amino acids at the P2 position. LVFA (SEQ ID NO: 1102), LIFA (SEQ ID NO: 1103), LVAR (SEQ ID NO: 1104), LAFA (SEQ ID NO: 1105), LQAR (SEQ ID NO: 1106), LDAA (SEQ ID NO: 1107), LCAR (SEQ ID NO: 1108), LGAA (SEQ ID NO: 1109), LMFA (SEQ ID NO: 1110), LSAR (SEQ ID NO: 1111), and LEFA (SEQ ID NO: 1112) are shown.

[0024] Figures 8A-8E Examples of binding and cleavage rates are shown. Figures 8A-8B The results show that PS961 and LIF ( Fig. 8A ) and IFA( Figure 8B ) Mean inter-pulse duration (IPD) of RS combined with ; median IPD values ​​are indicated. Figure 8C Shown are single exponential decay curves fitted to the RS duration distributions of arginine, leucine, isoleucine, and phenylalanine obtained from kinetic sequencing of the synthetic peptide DQQRLIFAG (SEQ ID NO: 1036). Figures 8D-8E It was shown that increasing aminopeptidase concentration in a kinetic sequencing run of the synthetic peptide DQQRLIFAG (SEQ ID NO: 1036) resulted in NRS ( Fig.8D ) and RS( Fig. 8E ) duration decreased; median RS duration values ​​are indicated.

[0025] Figures 9A-9G Examples of single amino acid changes and kinetic features of PTMs are shown. Fig. 9A The kinetic profiles of three peptides are shown: RLAFAYPDDD (SEQ ID NO: 1040) (top), RLIFAYPDDD (SEQ ID NO: 1041) (middle), and RLVFAYPDDD (SEQ ID NO: 1042) (bottom). Figure 9B-9C Shown is the incomplete RS information observed when dynamically sequencing the RLIFAYPDDD (SEQ ID NO: 1041) peptide. Fig. 9BThe percentage of reads for each type and example traces in one or more RS deletions observed in traces starting with arginine recognition and ending with tyrosine recognition are shown. RLIFY (SEQ ID NO: 1115) is shown. Fig. 9C Shown are the percentage of reads and example traces for each type of RS truncation observed in traces starting with arginine. Fig.9D Shown is the affinity of PS961 for peptides with an N-terminal methionine as measured by polarization assay (Example 1, Methods). Fig.9E Shown are binding energy predictions for peptides with N-terminal methionine (MFAY (SEQ ID NO: 1113) and methionine sulfoxide (Mo) (MoFAY (SEQ ID NO: 1114)) to PS961 obtained from computational modeling (Example 1, Methods). Fig.9F Shown are the kinetic profiles of DQQRLIFAG (SEQ ID NO: 1036) and RLAFSALGAADDD (SEQ ID NO: 1038) peptides mixed and run on the same chip. Figure 9G Shown are kinetic profiles of DQQRLIFAGK (SEQ ID NO: 1045) and EFIAWLVK (SEQ ID NO: 1046) peptides obtained from digestion of recombinant human ubiquitin and GLP-1.

[0026] Figures 10A-10M An example of peptide identification using modeled proteome-wide kinetic signatures is shown. Figures 10A-10C PS961 binding to the N-terminal position of leucine ( Fig. 10A ), isoleucine ( Fig. 10B ) or valine ( Fig. 10C ) heatmap of predicted pulse durations for the tripeptide targets. Figures 10D-10F PS610 binding to phenylalanine at the N-terminal position ( Fig. 10D ), tyrosine( Fig.10E ) or tryptophan ( Fig.10F ) heatmap of predicted pulse durations for the tripeptide targets.

[0027] Figure 10G A heat map showing the predicted pulse durations for PS1122 binding to a tripeptide target with arginine at the N-terminal position. Fig. 10H Shown are graphs demonstrating the high correlation of predicted pulse durations for PS961 (left panel) and PS610 (right panel) with actual pulse durations from on-chip experiments. Figures 10I-10K The results of the analysis of the human proteome are shown. Fig.10JIn, the sequences of IL6_HUMAN (SEQ ID NO: 1116), DGISALRK (SEQ ID NO: 1117), SNMCESSK (SEQ ID NO: 1118), EALAENNLNLPK (SEQ ID NO: 1119), DGCFQSGFNEETCLVK (SEQ ID NO: 1120), IITGLLEFEVYLEYLQNRFESSEEQARAVQMSTK (SEQ ID NO: 1121), VLIQFLQK (SEQ ID NO: 1122), DPTTNASLLTK (SEQ ID NO: 1123) and DMTTHLILRSFK (SEQ ID NO: 1124) are shown. Figure 10K In the figure, ACLILRSIEELK (SEQ ID NO: 1125), DMTTHLILRSFK (SEQ ID NO: 1124), SDSRNTLILRCK (SEQ ID NO: 1126), DSSHQISALVLRAQASEILLEELQQGLSQAK (SEQ ID NO: 1127), ARTVGIEELILRIqESK (SEQ ID NO: 1128), STLVLRCHRRRK (SEQ ID NO: 1129), DSPQEPLVLRLK (SEQ ID NO: 1130) and DLVLRATK (SEQ ID NO: 1131) are shown. Figure 10L-10M Results from proteomic analysis of E. coli are shown. Figure 10M In the figure, the sequences of SSUA_ECOLI (SEQ ID NO: 1132), LALAGLLSVSTFAVAAESSPEALRIGYQK (SEQ ID NO: 1133), GSSSHNLLLRALRQAGLK (SEQ ID NO: 1134), DPYYSAALLQGGVRVLK (SEQ ID NO: 1135), DLNQTGSFYLAARPYAEK (SEQ ID NO: 1136), DLFYENRLVPK (SEQ ID NO: 1137) and DIRQRIWQPLEGK (SEQ ID NO: 1138) are shown.

[0028] Figures 11A-11D Example results are shown for the selection and analysis of N-terminal alanine and valine binding variants. Fig.11A Results after 3 rounds of FACS selection are shown. Fig. 11B Results from one round of error-prone PCR library mixture selection are shown. Fig. 11CExample depictions of putative binding of alanine peptides to PS557 variants are shown. Fig.11D Shown are the results of fluorescence polarization studies comparing the kinetics of N-terminal alanine peptide binding of selected PS557 variants.

[0029] Figures 12A-12B Example results of rational design of PS557 recognition agent variants are shown. Fig. 12A A heat map showing the enrichment of mutations in the PS557 protein is shown. Fig. 12B Selected candidate binding assay traces from the Octet platform are shown.

[0030] Figures 13A-13D Example results for the development of arginine recognition agents are shown. Fig.13A Shown are the polarization responses of PS621 variants binding to RA, KA and HA peptides (upper panel) and the binding affinities (Kd) of selected PS621 variants for RA and HA at 20°C as determined by polarization (lower table). Fig. 13B Kd determination titration curves for RA binding of PS621, PS691 and PS1122 are shown. Fig. 13C Shown is the on-chip recognition of RA dipeptide by PS1122 in an on-chip recognition assay using QP304-RAIFAG. Fig.13D An example of multiplexed dynamic chip analysis of PS1122 is shown to demonstrate its enhanced arginine tripeptide coverage. Three peptides containing different RXA motifs (RLQFQALMAADDD (SEQ ID NO: 1139), LAQRQAFDAADDD (SEQ ID NO: 1140), FAQLQARFAADDD (SEQ ID NO: 1141)) were evaluated simultaneously in one sequencing run.

[0031] Figures 14A-14C Example results of computational modeling of PS961 are shown, demonstrating that N41D enhances the electrostatic interaction with the N-terminal amino group of the polypeptide. Fig.14A The amino acid triplets D73, D43 and N41 (left panel) or D41 (right panel) are shown bound to the N-terminal amino group of the AAA-tripeptide (sticks, top middle) via hydrogen bonds (dashed lines). Fig. 14B The binding pocket of PS961 was shown to be more electronegative than that of PS557, increasing the potential for protein-peptide interactions. Fig. 14C It is shown that the average fa_elec energy term of the peptide N-terminus of 50 representative structures in molecular dynamics simulations is lower (more favorable) in PS961 than in PS557.

[0032] Fig.15Example results of computational modeling of PS961 are shown, showing that V72M increases hydrophobic interactions with the peptide and further fills the protein core. The longer methionine side chain can pack tightly onto neighboring residues (van der Waals radii shown as spheres) and interact with the peptide N-terminal side chain (sticks and spheres, top middle) more tightly than its valine precursor.

[0033] Fig.16 Example results of computational modeling of PS961 are shown, showing that L68M further optimizes the packing of the protein core. Similar to V72M, the longer methionine side chain penetrates deeper into the non-optimal protein cavity than leucine (van der Waals radii shown as spheres), providing stabilizing hydrophobic interactions.

[0034] Fig.17 Example results of computational modeling of PS961 are shown, showing that the Y100R mutation neutralizes and increases the charge of a small negative surface pocket distal to the binding pocket, which may compete with the expected recognition pocket.

[0035] Figures 18A-18B Example results of computational modeling of PS961 are shown, showing that Y100R allows for a unique loop structure that favors the penultimate interaction. Fig.18A R100 is shown to stabilize an extensive hydrogen bond network between the penultimate backbone (AP; sticks, left), R106, and other members of the ring. Fig.18B Shows Fig.18A The average occupancies of the R100:R106 and R106:AP hydrogen bonds depicted in Figure 3 are higher in the PS961 simulations than in the PS557 simulations (50 ns trajectories, n=3).

[0036] Figures 19A-19C The secondary structure, sequence and binding pocket properties of PS961 are shown. Fig.19A The classification of the secondary structural groups of the proteins is shown. Fig.19B A Poisson–Boltzmann electrostatic potential surface plot of the binding pocket is shown, the residues forming the binding pocket are labeled, and the corresponding pocket properties are listed. Fig.19C The sequence of residues 32-116 of the native parent protein (PS557 (SEQ ID NO: 1)) and the sequence of residues 32-116 of the engineered variant (PS961 (SEQ ID NO: 314)) are shown, highlighting the mutations, pocket locations, and secondary structure assignments for each position.

[0037] Figures 19D-19S PS961 and N-terminal methionine peptide ( Figures 19D-19K ) or N-terminal alanine peptide ( Figures 19L-19S ) Example results of the crystal structure analysis of the composite. Fig.19DCrystals of the PS961:MAKL complex are shown. Fig.19E Shown is the crystal structure of the recognition agent PS961 (surface) in complex with the target peptide MAKL (SEQ ID NO: 1047) (stick). Figures 19F-19K It shows how PS961 binds to the target peptide MAKL (SEQ ID NO: 1047). Figure 19L Shown is the crystal structure of the recognition agent PS961 (cartoon) in complex with the target peptide AAKL (SEQ ID NO: 1048) (stick). Fig.19M Backbone superposition of PS961 binding to MAKL (SEQ ID NO: 1047) and AAKL (SEQ ID NO: 1048) is shown. Fig.19N Shown is the displacement of the Asp42 side chain in the recognition agent when comparing binding to AAKL (SEQ ID NO: 1048) and MAKL (SEQ ID NO: 1047). Fig.19O Comparison of AAKL (SEQ ID NO: 1048) and MAKL (SEQ ID NO: 1047) peptide binding to PS961 is shown. Figure 19P A comparison of the interactions of AAKL (SEQ ID NO: 1048) (left panel) and MAKL (SEQ ID NO: 1047) (right panel) peptides with residues in PS961 is shown. Figure 19Q Reorientation of residue Aspl2 in PS961 is shown to bind either MAKL (SEQ ID NO: 1047) (via a water molecule as a mediator) or AAKL (SEQ ID NO: 1048) peptides. Figure 19R Reorientation of residue Asp42 in PS961 is shown to bind either MAKL (SEQ ID NO: 1047) or AAKL (SEQ ID NO: 1048) peptides. Figure 19S Shown are superpositions of AAKL (SEQ ID NO: 1048) and MAKL (SEQ ID NO: 1047) peptides, illustrating the 180° flip of the LYS side chain at the third position (left), and the different orientations of the Lys3 side chain when the MAKL (SEQ ID NO: 1047) and AAKL (SEQ ID NO: 1048) peptides are bound to PS961 (right).

[0038] Fig. 20AExample results of computational modeling of PS1122 are shown, showing the PS621 crystal structure electrostatic surface with the modeled mutations. Top right: The E70T mutation creates a new hydrogen bond with the N-terminal arginine side chain. Middle right: The I63E mutation creates a new hydrogen bond with the amino terminus. Bottom right: The T47L mutation was found to be a common mutation in binding selection and may improve protein stability by stabilizing the helical turn near the metal binding site.

[0039] Figures 20B-20F Shown are example results of crystal structure analysis of PS1122 in complex with an N-terminal arginine peptide (RAKL (SEQ ID NO: 1049)). Fig. 20B Crystals of the PS1122:RAKL complex are shown. Fig. 20C Shown is the crystal structure of the recognition agent PS1122 in complex with the target peptide RAKL (SEQ ID NO: 1049) (stick). Fig.20D A portion of PS1122 is shown to be different from a similar portion observed in PS621, the predecessor of PS1122. Fig.20E The NH3 group of the first amino acid of Arg-1 that binds the peptide is shown interacting with amino acid Glu-63, Asp-65, and a water molecule held in place by Glu-63 (interactions are depicted as dashed lines). Fig.20F The side chains of both Glu-63 and Thr-70 of PS1122 were shown to interact with Arg-1 of the bound peptide.

[0040] Fig.21 A-21C shows the secondary structure, sequence, and binding pocket characteristics of PS1122. Fig.21 A shows the classification of proteins into secondary structural groups. Fig.21 B shows the Poisson-Boltzmann electrostatic potential surface plot of the binding pocket, with the residues forming the binding pocket labeled and the corresponding pocket properties listed. Fig.21 C shows the sequence of residues 1-82 of the native parent protein (PS621 (SEQ ID NO: 2)) and the sequence of residues 1-82 of the engineered variant (PS1122 (SEQ ID NO: 468)), highlighting the mutations, pocket locations, and secondary structure assignments for each position.

[0041] Fig. 22 An AlphaFold model of PS1259 is shown with a hydrogen bond network enabled by two mutations: C25S (below the N-terminal glutamine, shown as sticks) and H78Q (above the N-terminal glutamine, shown as sticks).

[0042] Figures 23A-23C The secondary structure, sequence, and binding pocket properties of PS1259 are shown. Fig.23A The classification of the secondary structural groups of the proteins is shown. Fig. 23B A Poisson–Boltzmann electrostatic potential surface plot of the binding pocket is shown, the residues forming the binding pocket are labeled, and the corresponding pocket properties are listed. Fig.23C The sequences of the native parent protein (Ntaq1(sf) (SEQ ID NO:3)) and the engineered variant PS1259 (SEQ ID NO:605) are shown, highlighting the mutations, pocket locations, and secondary structure assignments for each position.

[0043] Figures 24A-24D Example results showing direct identification of an arginine PTM are shown. Fig.24A Different arginine PTMs (YRELRLLK (SEQ ID NO: 1077), YR ADMA ELRLLK (SEQ ID NO: 1078) and YR SDMA ELRLLK (SEQ ID NO: 1079)), including symmetric dimethylarginine (SDMA), asymmetric dimethylarginine (ADMA) and citrullinated arginine. Fig. 24B An exemplary workflow for collecting samples, preparing digested peptide libraries, loading onto the chip, and performing on-chip sequencing and data analysis is shown. Fig.24C Sequencing data are shown, demonstrating that kinetic signatures can distinguish peptides containing arginine, ADMA, and SDMA. Fig.24C -A shows example protein sequencing traces of three synthetic P38 MAPKα-derived peptides containing arginine, ADMA or SDMA at position 2. The full-length peptide sequence of each example trace is indicated. Fig.24C -B shows the distribution of the average pulse duration (PD) of the recognition segment (RS) of the RS corresponding to the initial 4-residue sequence of each peptide: YREL (SEQ ID NO: 1050) (left), YR ADMA EL (SEQ ID NO: 1051) (middle) and YR SDMA EL (SEQ ID NO: 1052) (right). The median of each distribution is indicated. Fig.24C -C shows the inter-pulse duration (IPD) of PS621 detecting arginine and ADMA. Fig.24D Sequencing data are shown, demonstrating that the kinetic signature can distinguish between arginine- and citrulline-containing peptides. Fig.24D-A shows example protein sequencing traces of two synthetic peptides containing arginine or citrulline at position 2: the peptide sequence LRLAFAYPDDDK (SEQ ID NO: 1053) (QP707) and the citrullinated peptide sequence LRcitLAFAYPDDDK (SEQ ID NO: 1054) (QP789). The full-length peptide sequences of each example trace are marked. Fig.24D -B shows the distribution of the average PD of RS corresponding to the initial 5-residue sequences of each peptide: LRLAF (SEQ ID NO: 1055) (left) and LCitLAF (SEQ ID NO: 1056) (right). The median value of each distribution is marked.

[0044] Figures 25A-25C Shows example results demonstrating the identification of threonine PTM. Figures 25A-25B Shows the results of sequencing reactions of the following peptides using the recognition agents PS691, PS610, and PS961: RLTFIAYPDDD (SEQ ID NO: 1057) ( Fig.25A ); and RLpTFIAYPDDD (SEQ ID NO: 1058), where pT is phosphorylated threonine ( Fig.25B ). Fig.25C Shows Fig.25A (left panel) and Fig.25B the duration of the recognition segment (RS) for leucine recognition in the sequencing reaction of (right panel).

[0045] Figures 26A-26B Shows example results demonstrating the identification of tyrosine PTM of the following peptides using the recognition agents PS691, PS610, PS961, and PS1165 in a sequencing reaction: RLYFIAYPDDD (SEQ ID NO: 1059) ( Fig.26A ); and RLpYFIAYPDDD (SEQ ID NO: 1060), where pY is phosphorylated tyrosine ( Fig.26B ).

[0046] Figures 27A-27B Shows example results demonstrating the identification of lysine PTM of the following peptides using the recognition agents PS691, PS610, PS961, and PS1165 in a sequencing reaction: RLYFKAYPDDD (SEQ ID NO: 1061) ( Fig.27A ); and RLK{acetyl}FIAYPDDD (SEQ ID NO: 1062), where K{acetyl} is acetylated lysine ( Fig.27B ).

[0047] Figures 28A-28G Illustrates aspects of an example application of the technique for identifying β-amyloid variants. Fig.28A Examples of beta-amyloid variants are illustrated. Fig.28B An example workflow for β-amyloid variant detection is described. Figures 28C-28G Examples of pulse patterns for beta-amyloid wild-type LVFFAE (SEQ ID NO: 1063) and variants (LVFFAK (SEQ ID NO: 1064), LVFFGK (SEQ ID NO: 1065), LVFFAG (SEQ ID NO: 1066), LVPFAE (SEQ ID NO: 1067)) are illustrated.

[0048] Fig.29 An example schematic diagram of an integrated device pixel is shown.

[0049] Figures 30A-30F Computational modeling ( Figures 30A-30E ) and a model peptide (RLAF (SEQ ID NO: 1142) evaluated with PS621 and PS1122 ( Fig. 30B )、ADMA-LAF (SEQ ID NO: 1143) ( Fig. 30C )、SDMA-LAF (SEQ ID NO: 1144) ( Fig.30D ) and Cit-LAF (SEQ ID NO: 1145) ( Fig.30E )) instance result of the structure ( Fig.30F ).

[0050] Fig.31 The frequencies of amino acids in the human proteome are shown.

[0051] Fig.32A High-throughput expression, purification and conjugation to streptavidin of homologous variants of the hNTAQ protein are shown.

[0052] Fig.32B Example traces of sequencing reactions performed using six recognition agents, including the glutamate recognition agent PS1875, are shown.

[0053] Fig.32C The expression and purification of doubly biotinylated PS2132 at 2L scale is shown.

[0054] Fig.32D Shown are example results of size exclusion chromatography (SEC) of PS2132 labeled with a streptavidin-linked long-lived BODIPY dye.

[0055] Fig.32E Shown are example results of quality control SDS PAGE gel analysis of PS2132 labeled with the long-lived BODIPY dye before and after SEC column purification.

[0056] Figures 33A-33B PS1875 labeled with the long-lived BODIPY dye ( Fig.33A ) and PS2132( Fig.33B ) on-chip identification of E.

[0057] Figures 35A-35B PS1875 ( Fig.35A ) and PS2123( Fig.35B ) on-chip identification of E.

[0058] Fig.36 Shown are pulse durations (top) and inter-pulse durations (bottom) corresponding to RS recognized by E in aligned reads from sequencing runs using the QP1165 (EIAFLKQRVWK (SEQ ID NO: 1084) peptide, using a cocktail of six recognition agents including PS1875, PS2132, PS2121, or PS2123 as glutamate (E) recognition agents.

[0059] Figures 37A-37C The results show that reagent A containing five recognition agents ( Fig.37A ) or reagent A in combination with E recognition agent PS2132 ( Fig.37B ) is an example result of a CDNF library sequencing run performed. Fig.37C Depicted are the sequence of CDNF (SEQ ID NO: 1146) and example traces of three peptides: EFLNRFYK (SEQ ID NO: 1068), ELISFCLDTK (SEQ ID NO: 1069), and ENRLCYYLGATK (SEQ ID NO: 1070).

[0060] Figures 38A-38C Shows the use of ( Fig.38A ) or not use ( Fig.38B ) Example results of a GFAP peptide library sequencing run using the glutamate (E) recognition agent PS2132 in combination with five recognition agents. Fig.38C Depicted are the sequence of GFAP (SEQ ID NO: 1147) and example traces of two peptides: DEMARHLQEYQDLLNVK (SEQ ID NO: 1071) and LALDIEIATYRK (SEQ ID NO: 1072).

[0061] Figures 39A-39B Example results from computational modeling of PS2132 are shown. Fig.39A The binding pocket of PS2132 in complex with the N-terminal glutamate peptide is depicted. Fig.39BSurface charge modeling of PS1259 and PS2132 is depicted. DETAILED DESCRIPTION

[0062] Aspects of the present disclosure relate to compositions and methods for determining the chemical characteristics of a polypeptide based on single-molecule binding interactions between the polypeptide and one or more reagents described herein. In some embodiments, the present disclosure provides a method for performing polypeptide structural analysis based on kinetic information derived from single-molecule binding interactions between a polypeptide and one or more amino acid recognition agents described herein.

[0063] Figure 1 An example of a dynamic peptide sequencing reaction is shown, in which a single on-off binding event generates a signal pulse of a signal output. As shown in the left figure, a protein sample can be fragmented into peptides, which are fixed in the reaction chamber of the array, and the fixed peptides are exposed to one or more amino acid recognition agents and one or more cleavage reagents (e.g., aminopeptidases) in the reaction chamber. As shown in the right figure, the amino acid recognition agent reversibly binds to the end of the peptide, and a pulse is generated in the signal output when the recognition agent binds to the peptide.

[0064] Since the on-off binding of the recognition agent generally occurs at a faster rate than the amino acid cleavage, the binding event before each cleavage event results in a series of changes in the signal (e.g., a signal pulse), which can be used to determine structural information about the amino acids at or near the end of the peptide. Compositions and methods for performing dynamic polypeptide sequencing and analyzing data obtained therefrom are more fully described in PCT International Publication No. WO2020102741A1 filed on November 15, 2019 and PCT International Publication No. WO2021236983A2 filed on May 20, 2021, both of which are incorporated herein by reference in their entirety.

[0065] In some aspects, the present disclosure provides amino acid recognition agents with improved binding properties, which can obtain more structural information from polypeptides based on the kinetics of on-off binding between the recognition agent and the polypeptide. In some embodiments, the amino acid recognition agent comprises an amino acid binding protein with an engineered binding pocket, and the binding pocket has one or more modifications relative to a homologous protein. In some embodiments, the modified binding pocket increases the number of interactions (e.g., hydrogen bond interactions, van der Waals interactions) formed between the binding pocket and the amino acid ligand compared to the unmodified homologous protein binding pocket. In some embodiments, the modified binding pocket increases the number of amino acid ligand types that can be detected for binding compared to the unmodified homologous protein binding pocket. In some embodiments, the modified binding pocket improves the kinetics (e.g., K D , k off, k on ), which advantageously increases the amount or confidence of structural information obtained from the polypeptide analysis described herein.

[0066] I. Amino Acid Recognizer

[0067] In some aspects, the present disclosure provides amino acid recognition agents comprising amino acid binding proteins having an amino acid sequence selected from Table 1. Table 1 herein provides an example sequence list of amino acid binding proteins. It should be understood that these sequences and other examples described herein are meant to be non-limiting, and amino acid recognition agents according to the present disclosure may include any homologues, variants or fragments thereof, which minimally contain a domain or subdomain responsible for amino acid recognition.

[0068] In some embodiments, the disclosure provides an amino acid binding protein having an amino acid sequence that is at least 80% identical to an amino acid sequence selected from Table 1. In some embodiments, the amino acid binding protein has at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98% or more amino acid sequence identity to an amino acid sequence selected from Table 1. In some embodiments, the amino acid binding protein has 25-50%, 50-60%, 60-70%, 70-80%, 80-90%, 90-95%, 95-99%, 40-100%, 50-100%, 60-100%, 70-100%, 80-100%, 90-100% or 95-100% amino acid sequence identity to an amino acid sequence selected from Table 1.

[0069] For comparison of two or more amino acid sequences, the percentage of "sequence identity" (also referred to herein as "amino acid identity") between a first amino acid sequence and a second amino acid sequence can be calculated by dividing [the number of amino acid residues in the first amino acid sequence that are identical to the amino acid residues at corresponding positions in the second amino acid sequence] by [the total number of amino acid residues in the first amino acid sequence], and multiplying by

[100] , wherein each deletion, insertion, substitution or addition of an amino acid residue in the second amino acid sequence compared to the first amino acid sequence is considered to be a difference in a single amino acid residue (position). Alternatively, the degree of sequence identity between two amino acid sequences can be calculated using a known computer algorithm (e.g., by the local homology algorithm of Smith and Waterman (1970) Adv. Appl. Math. 2: 482c, by the homology alignment algorithm of Needleman and Wunsch, J. Mol. Biol. (1970) 48: 443, by the search similarity method of Pearson and Lipman. Proc. Natl. Acad. Sci. USA (1998) 85: 2444, or by a computer implementation of Blast, Clustal Omega or other sequence alignment algorithms, etc., available), for example, using standard settings. Usually, in order to determine the "sequence identity" percentage between two amino acid sequences according to the above calculation method, the amino acid sequence with the most amino acid residues will be used as the "first" amino acid sequence, and the other amino acid sequence will be used as the "second" amino acid sequence.

[0070] Additionally or alternatively, the identity between two or more sequences can also be evaluated.In the context of two or more nucleotides or amino acid sequences, the term "identical" or "identity" percentage refers to two or more identical sequences or subsequences.When comparing and comparing to obtain maximum corresponding relationship on a comparison window or a specified region, as measured by using one of the above-mentioned sequence comparison algorithms or by manual comparison and visual inspection, if two sequences have a specific percentage of identical amino acid residues or nucleotides (for example, at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8% or 99.9% identical) on a specified region or the entire sequence, then the two sequences are "substantially identical".Optionally, identity is present in a region of at least about 25, 50, 75 or 100 amino acids in length, or a region of 100 to 150, 150 to 200, 100 to 200 or more than 200 amino acids in length.

[0071] Additionally or alternatively, the comparison between two or more sequences can also be evaluated.In the context of two or more nucleotide or amino acid sequences, the term "comparison" or "comparison" percentage refers to two or more identical sequences or subsequences.When comparing and comparing to obtain maximum corresponding relationship on a comparison window or a specified region, as measured by using one of the above-mentioned sequence comparison algorithms or by manual comparison and visual inspection, if two sequences have the same amino acid residue or nucleotide of a specified percentage on a specified region or the entire sequence (for example, at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8% or 99.9% identical), then the two sequences are "basic comparisons". Optionally, the alignment exists over a region that is at least about 25, 50, 75, or 100 amino acids in length, or over a region that is 100 to 150, 150 to 200, 100 to 200, or more than 200 amino acids in length.

[0072] In some embodiments, the amino acid recognition agents of the present disclosure include modified amino acid binding proteins and include one or more amino acid deletions, additions or mutations relative to the sequences listed in Table 1. In some embodiments, the modified amino acid binding proteins include deletions, additions or mutations of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acids (which may or may not be consecutive amino acids) relative to the sequences listed in Table 1.

[0073] A. ClpS homologous recognition agent

[0074] In some embodiments, the amino acid recognition agent of the present disclosure binds to an amino acid ligand (e.g., a polypeptide) comprising an N-terminal amino acid selected from leucine, isoleucine, valine, methionine, alanine, or a modified variant thereof (e.g., a post-translational modified variant thereof, an oxidized variant thereof). In some embodiments, the amino acid recognition agent comprises an amino acid binding protein derived from a ClpS protein (e.g., a Planctomycetia bacterium ClpS protein). For example, in some embodiments, the amino acid binding protein is an engineered variant comprising one or more modifications relative to SEQ ID NO: 1 described herein.

[0075] In some embodiments, the amino acid binding protein has a dissociation constant (K) of less than 2,000 nM, less than 1,500 nM, less than 1,000 nM, less than 750 nM, less than 500 nM, less than 250 nM, less than 150 nM, less than 100 nM, less than 80 nM, 10-2,000 nM, 25-1,000 nM, 50-500 nM, 10-150 nM, 25-75 nM, or 50-60 nM. D ) binds to the N-terminal leucine. In some embodiments, the amino acid binding protein has a K of less than 2,000 nM, less than 1,500 nM, less than 1,000 nM, less than 750 nM, less than 500 nM, less than 250 nM, less than 150 nM, less than 100 nM, less than 80 nM, 10-2,000 nM, 25-1,000 nM, 50-500 nM, 10-150 nM, 30-80 nM, or 60-75 nM. D In some embodiments, the amino acid binding protein has a K of less than 2,000 nM, less than 1,500 nM, less than 1,000 nM, less than 750 nM, less than 500 nM, less than 300 nM, less than 250 nM, less than 200 nM, 10-2,000 nM, 25-1,000 nM, 50-500 nM, 50-300 nM, or 100-200 nM. D Binds to the N-terminal valine.

[0076] In some embodiments, the amino acid binding protein binds one or more types of N-terminal amino acids (e.g., leucine, isoleucine, valine, methionine, and / or alanine), wherein each type of binding interaction is characterized by an off-rate (k off ) is at least 0.1s -1 In some embodiments, the dissociation rate is about 0.1 s -1 To about 1,000s -1 between (e.g., between about 0.5s -1 To about 500s -1 Between, about 0.1s -1 To about 100s -1 Between, in about 1s -1 To about 100s -1 Between or about 0.5s -1 To about 50s -1 In some embodiments, the dissociation rate is between about 0.5 s -1 About 20 seconds -1 In some embodiments, the dissociation rate is between about 2 s -1 About 20 seconds -1In some embodiments, the dissociation rate is between about 0.5 s -1 To about 2s -1 between.

[0077] In some aspects, the present disclosure provides a recombinant or synthetic amino acid binding protein having a sequence selected from the group consisting of PS635-645, PS731-732, PS759-766, PS769, PS795-870, PS896-912, PS918-1043, PS1048-1100, PS1124-1137, PS1141-1161, PS1175-1199, PS1203-1217, PS1222-1245, PS1277-1305, PS1321-1350, and PS1425-1448 (SEQ ID NO:22-27, 87-88, 115-122, 125, 151-226, 249-265, 271-390, 395-446, 470-483, 487-507, 521-545, 549-563, 568-591, 622-650, 664-693 and 768-791) are at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98% or 100% identical to the sequence of any one of the sequences. In some embodiments, the amino acid sequence is selected from the group consisting of PS635-645, PS731-732, PS759-766, PS769, PS795-870, PS896-912, PS918-1043, PS1048-1100, PS1124-1137, PS1141-1161, PS1175-1199, PS1203-1217, PS1222-1245, PS1277-1305, PS1321-1350 and PS1425-1448 (SEQ ID NO:22-27, 87-88, 115-122, 125, 151-226, 249-265, 271-390, 395-446, 470-483, 487-507, 521-545, 549-563, 568-591, 622-650, 664-693 and 768-791) are about 40% to about 100% (e.g., 80-98%, 80-95%, 80-90%, 85-95%, 90-98%, 50-100%, 60-100%, 70-100%, 80-100%, 90-100% or 95-100%) identical to any one of the sequences. In some embodiments, the amino acid sequence is at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or 100% identical to PS961 (SEQ ID NO:314).In some embodiments, the amino acid sequence is about 40% to about 100% (e.g., 80-98%, 80-95%, 80-90%, 85-95%, 90-98%, 50-100%, 60-100%, 70-100%, 80-100%, 90-100%, or 95-100%) identical to PS961 (SEQ ID NO:314).

[0078] In some aspects, the present disclosure provides a recombinant or synthetic amino acid binding protein having an amino acid sequence that is at least 80% identical (e.g., at least 85%, at least 90%, at least 95%, at least 98%, 80-98%, 80-95%, 80-90%, 85-95%, or 90-98% identical) to SEQ ID NO: 1, wherein the amino acid sequence comprises an amino acid substitution at one or more positions corresponding to E22, R31, L39, N41, D42, D43, D44, H45, T46, Y47, V50, Q55, P62, E63, L68, A69, V72, D73, Q75, Y100, and M111 of SEQ ID NO: 1.

[0079] In some embodiments, the amino acid sequence comprises an amino acid substitution at a position corresponding to N41 and one or more positions corresponding to E22, R31, L39, D42, H45, V50, Q55, P62, E63, L68, V72, Q75, Y100, and M111. In some embodiments, the amino acid sequence comprises an amino acid substitution at a position corresponding to N41 and one or more positions corresponding to Q55, E63, L68, V72, and Y100. In some embodiments, the amino acid sequence comprises an amino acid substitution selected from the group consisting of E22V, R31H, L39M, N41D, D42L, D42P, H45C, H45F, V50A, V50F, V50Y, Q55H, Q55R, P62R, E63A, E63G, E63K, E63S, L68M, V72M, Q75L, Y100R, M111A, and M111S. In some embodiments, the amino acid substitution is selected from the group consisting of N41D, Q55R, E63S, L68M, V72M, and Y100R.

[0080] In some aspects, the present disclosure provides a recombinant or synthetic amino acid binding protein comprising a structure of formula (I) or a structural equivalent thereof:

[0081] β1–α1–α2–β2–α3–β3

[0082] (I)

[0083] wherein: each of β1, β2 and β3 is a β strand; each of α1, α2 and α3 is an α helix; each instance of “–” is a loop; and at least a portion of each of α1, α2, the loop between β1 and α1, and the loop between α3 and β3 forms a binding pocket for an amino acid ligand.

[0084] In some embodiments, the binding pocket comprises one or more of the following: (i) a volume of about (ii) Electrostatic potential is -3.0RTe c -1 or less, (iii) at least 35% of the amino acids forming the binding pocket have negatively charged side chains, (iv) a plurality of hydrogen bond acceptors configured to form one or more hydrogen bonds (e.g., at least two, at least three, at least four, or at least five hydrogen bonds) in the presence of an amino acid ligand, and (v) a plurality of van der Waals contact positions configured to form van der Waals interactions in the presence of an amino acid ligand. In some embodiments, the binding pocket comprises one, two, three, or four of (i), (ii), (iii), (iv), and (v). In some embodiments, the binding pocket comprises (i), (ii), (iii), (iv), and (v).

[0085] In some embodiments, the binding pocket comprises about In some embodiments, the volume of the binding pocket is at least At least At least At least or at least In some embodiments, the volume of the binding pocket does not exceed No more than No more than No more than or not more than In some embodiments, the volume of the binding pocket is to to to to to to or to Methods for determining the volume of a binding pocket are known in the art and will be apparent to the skilled artisan based on this disclosure. For example, in some embodiments, the volume of a binding pocket is determined using software configured to measure the geometric and topological properties of a protein. In some embodiments, the software can scan the surface of a protein using a specified probe radius to measure the volume of any cavity that overlaps the binding site directly or that overlaps the binding site indirectly through adjacent cavities within van der Waals contacts between each other. A non-limiting example of a suitable probe radius is a solvent probe radius (e.g., about ). A non-limiting example of suitable software is Computational Atlas of Surface Topography of Proteins (CASTp). See, e.g., W. Tian et al., CASTp 3.0: Computed atlas of surface topography of proteins. Nucleic Acids Res. 46, W363-W367 (2018), the relevant contents of which are incorporated herein by reference.

[0086] In some embodiments, the binding pocket has a -3.0 RTe c -1 In some embodiments, the electrostatic potential of the binding pocket is at least -4RTe. c -1 , at least -3RTe c -1 or at least -2RTe c -1 In some embodiments, the electrostatic potential of the binding pocket is -2RTe c -1 or lower, -3RTe c -1 or lower, or -4RTe c -1 Or lower. In some embodiments, the electrostatic potential of the binding pocket is -2RTe c -1 To -3RTe c -1 、-2RTe c -1 To -4RTe c -1 or -3RTe c -1 To -4RTe c -1Methods for determining the electrostatic potential of a binding pocket are known in the art and will be apparent to the skilled person based on this disclosure. For example, in some embodiments, the Adaptive Poisson-Boltzmann Solver (APBS) tool in PyMOL (PyMOL Molecular Graphics System, Version 2.0 LLC) to calculate the electrostatic surface potential of the binding pocket, using pdb2pqr and AMBER force fields to assign protonation states. See, for example, TJ Dolinsky et al., PDB2PQR: An automated pipeline for the setup of Poisson-Boltzmanne electrostatics calculations, Nucleic Acids Res., 32: W665-7 (2004); MG Lerner et al., APBS plugin for PyMOL–Version 2.4 (University of Michigan, Ann Arbor, MI, 2006); JW Ponder et al., Force fields for protein simulations, Adv. Protein Chem. 66: 27–85 (2003). In some cases, this solvent accessible surface area (SASA) can be considered as accessible to the peptide ligand.

[0087] In some embodiments, the binding pocket comprises a plurality of hydrogen bond receptors, which are configured to form one or more hydrogen bonds in the presence of an amino acid ligand. In some embodiments, the binding pocket forms at least two (e.g., at least three, at least four, at least five, 2-10, 4-10, 5-15, 5-10) hydrogen bonds with the amino acid ligand. Methods for determining hydrogen bond interactions between a binding pocket and a ligand are known in the art and will be apparent to the technician on the basis of the present disclosure. For example, in some embodiments, hydrogen bond interactions are determined by computational modeling, as described in the examples herein (e.g., using atomic coordinates of protein-ligand structural data obtained from X-ray crystallography, or predicting protein-ligand three-dimensional structures by computational modeling).

[0088] In some embodiments, the multiple hydrogen bond acceptors include one or more atoms of the side chains of amino acid residues in the binding pocket. For example, in some embodiments, the binding pocket comprises at least one negatively charged amino acid side chain (e.g., aspartic acid, glutamic acid) that forms a bifurcated hydrogen bond with an amino acid ligand (e.g., the N-terminal amino acid of a polypeptide). In some embodiments, at least four hydrogen bonds are formed between an amino acid ligand (e.g., the N-terminal amino acid of a polypeptide) and an amino acid side chain in the binding pocket. In some embodiments, the multiple hydrogen bond acceptors include one or more atoms of the polypeptide backbone in the binding pocket (e.g., backbone carbonyl).

[0089] In some embodiments, the binding pocket forms one or more hydrogen bonds with the side chain of the amino acid ligand. For example, in some embodiments, the binding pocket forms one or more hydrogen bonds with the terminal amino acid side chain of the amino acid ligand (e.g., polypeptide). In some embodiments, the binding pocket forms one or more hydrogen bonds with the polypeptide backbone of the amino acid ligand (e.g., polypeptide). In some embodiments, the binding pocket forms one or more hydrogen bonds with the terminal amino acid in the polypeptide and one or more amino acids adjacent to the terminal amino acid (e.g., amino acids at position 1 and position 2, 3, 4, and / or 5 relative to the end of the polypeptide).

[0090] In some embodiments, the binding pocket comprises a plurality of van der Waals contact positions that are configured to form van der Waals interactions in the presence of an amino acid ligand. Methods for determining van der Waals interactions between a binding pocket and a ligand are known in the art and will be apparent to the skilled person based on the present disclosure. For example, in some embodiments, van der Waals interactions are determined by computational modeling, as described in the Examples herein (e.g., using atomic coordinates of protein-ligand structural data obtained from X-ray crystallography, or by computational modeling to predict protein-ligand three-dimensional structures).

[0091] In some embodiments, the van der Waals contact position comprises a plurality of atoms (e.g., 2-30, 5-25, 10-20, 2-10, 5-10) that are configured to form hydrophobic interactions with the amino acid ligand. In some embodiments, one or more of the plurality of atoms are non-polar atoms. In some embodiments, the van der Waals contact position is formed by a methionine side chain in the binding pocket.

[0092] In some embodiments, the amino acid ligand is a polypeptide comprising at least three amino acids. In some embodiments, the amino acid ligand comprises the N-terminal amino acid of the polypeptide. In some embodiments, the N-terminal amino acid is selected from leucine, isoleucine, valine, methionine and alanine. In some embodiments, the length of the amino acid binding protein is at least 50 amino acids, at least 75 amino acids, at least 100 amino acids, 50-250 amino acids, 50-150 amino acids or 100-200 amino acids.

[0093] In some embodiments, the loop between β1 and α1 comprises three or more negatively charged amino acids. In some embodiments, the loop between β1 and α1 comprises four or more negatively charged amino acids. In some embodiments, at least two negatively charged amino acids in the loop between β1 and α1 form hydrogen bonds with amino acid ligands. In some embodiments, at least one negatively charged amino acid in the loop between β1 and α1 forms a bifurcated hydrogen bond with an amino acid ligand. In some embodiments, the negatively charged amino acids are selected from aspartic acid and glutamic acid.

[0094] In some embodiments, β1–α1 comprises an amino acid sequence that is at least 80% identical to the sequence of amino acids 35-58 of SEQ ID NO: 1. In some embodiments, the binding pocket is formed by amino acids at one or more positions corresponding to amino acids 41-47 and 50 of SEQ ID NO: 1. In some embodiments, the amino acid sequence comprises amino acid substitutions at one or more positions corresponding to L39, N41, D42, H45, V50, and Q55 of SEQ ID NO: 1. In some embodiments, at least one amino acid substitution is at a position corresponding to N41 of SEQ ID NO: 1. In some embodiments, the amino acid substitution is selected from N41D and Q55R.

[0095] In some embodiments, α2 comprises an amino acid sequence that is at least 80% identical to the sequence of amino acids 62-73 of SEQ ID NO: 1. In some embodiments, the binding pocket is formed by amino acids at one or more positions corresponding to amino acids 69, 72, and 73 of SEQ ID NO: 1. In some embodiments, the amino acid sequence comprises amino acid substitutions at one or more positions corresponding to P62, E63, L68, and V72 of SEQ ID NO: 1. In some embodiments, at least one amino acid substitution is at a position corresponding to V72 of SEQ ID NO: 1. In some embodiments, the amino acid substitution is selected from E63S, L68M, and V72M.

[0096] In some embodiments, the loop between α3 and β3 comprises an amino acid sequence that is at least 80% identical to the sequence of amino acids 99-112 of SEQ ID NO: 1. In some embodiments, the binding pocket is formed by amino acids at positions corresponding to amino acids 111 of SEQ ID NO: 1. In some embodiments, the amino acid sequence comprises amino acid substitutions at one or more positions corresponding to Y100 and M111 of SEQ ID NO: 1. In some embodiments, the amino acid substitution is Y100R.

[0097] In some embodiments, structural equivalents are those that differ by no more than wherein at least 80% of the secondary structure α-carbon atoms are aligned with the structure of formula (I). In some embodiments, structural equivalents are structures having a root mean square difference of no more than No more than No more than or not more than A structure wherein at least 80% of the secondary structure α-carbon atoms are aligned with the structure of formula (I).

[0098] Methods for identifying structural equivalents of the structure of formula (I) are known in the art and will be apparent to the skilled artisan based on this disclosure. For example, in some embodiments, structural equivalents of formula (I) are those having a root mean square difference of no more than wherein at least 80% of the secondary structure α-carbon atoms are aligned with a residue selected from the group consisting of PS635-645, PS731-732, PS759-766, PS769, PS795-870, PS896-912, PS918-1043, PS1048-1100, PS1124-1137, PS1141-1161, PS1175-1199, PS1203-1217, PS1222-1245, PS1277-1305, PS1321-1350, and PS1425-1448 (SEQ ID The three-dimensional protein structure alignment of the amino acid sequence of any one of NO:22-27, 87-88, 115-122, 125, 151-226, 249-265, 271-390, 395-446, 470-483, 487-507, 521-545, 549-563, 568-591, 622-650, 664-693 and 768-791). Protein structure comparison can be performed by determining the protein structure of the protein selected from the group consisting of PS635-645, PS731-732, PS759-766, PS769, PS795-870, PS896-912, PS918-1043, PS1048-1100, PS1124-1137, PS1141-1161, PS1175-1199, PS1203-1217, PS1222-1245, PS1277-1305, PS1321-1350, and PS1425-1448 (SEQ ID NO:22-27, 87-88, 115-122, 125, 151-226, 249-265, 271-390, 395-446, 470-483, 487-507, 521-545, 549-563, 568-591, 622-650, 664-693 and 768-791), and comparing the three-dimensional structure of the protein with the structure of the candidate protein to determine whether the candidate protein is a structural equivalent.The three-dimensional protein structure can be predicted by, for example, using atomic coordinates obtained from X-ray crystallography or by computational modeling to have a structure selected from the group consisting of PS635-645, PS731-732, PS759-766, PS769, PS795-870, PS896-912, PS918-1043, PS1048-1100, PS1124-1137, PS1141-1161, PS1175-1199, PS1203-1217, PS1222-1245, PS1277-1305, PS1321-1350, and PS1425-1448 (SEQ ID NO:22-27, 87-88, 115-122, 125, 151-226, 249-265, 271-390, 395-446, 470-483, 487-507, 521-545, 549-563, 568-591, 622-650, 664-693 and 768-791) in any one of the three-dimensional structure of the protein is determined. Methods for comparing protein structures and determining root mean square differences are known in the art (see, e.g., Kufareva I, Abagyan R. Methods of protein structure comparison. Methods Mol Biol. 2012; 857: 231-57).

[0099] B. UBR homologous recognition agent

[0100] In some embodiments, the amino acid recognition agent of the present disclosure binds to an amino acid ligand (e.g., a polypeptide) comprising an N-terminal amino acid selected from arginine or a modified variant thereof (e.g., a post-translational modified variant thereof, an oxidized variant thereof). In some embodiments, the amino acid recognition agent comprises an amino acid binding protein derived from a UBR protein (e.g., a Kluyveromyces marxianus protein). For example, in some embodiments, the amino acid binding protein is an engineered variant comprising one or more modifications relative to SEQ ID NO: 2 described herein.

[0101] In some embodiments, the amino acid binding protein has a dissociation constant (K) of less than 2,000 nM, less than 1,500 nM, less than 1,000 nM, less than 750 nM, less than 500 nM, less than 400 nM, less than 200 nM, less than 100 nM, 10-2,000 nM, 25-1,000 nM, 50-500 nM, 10-400 nM, 25-75 nM, or 50-80 nM. D ) binds to the N-terminal arginine.

[0102] In some embodiments, the amino acid binding protein binds one or more types of N-terminal amino acids (e.g., arginine or modified variants thereof), wherein each type of binding interaction is characterized by a dissociation rate (k off ) is at least 0.1s -1 In some embodiments, the dissociation rate is about 0.1 s -1 About 1,000s -1 between (e.g., between about 0.5s -1 To about 500s -1 Between, about 0.1s -1 To about 100s -1 Between, in about 1s -1 To about 100s -1 Between or about 0.5s -1 To about 50s -1 In some embodiments, the dissociation rate is between about 0.5 s -1 About 20 seconds -1 In some embodiments, the dissociation rate is between about 2 s -1 About 20 seconds -1 In some embodiments, the dissociation rate is between about 0.5 s -1 To about 2s -1 between.

[0103] In some aspects, the present disclosure provides a recombinant or synthetic amino acid binding protein having an amino acid sequence that is at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98% or 100% identical to any one of PS1101-1122, PS1218-1221 and PS1351-1398 (SEQ ID NOs: 447-468, 564-567 and 694-741). In some embodiments, the amino acid sequence is about 40% to about 100% (e.g., 80-98%, 80-95%, 80-90%, 85-95%, 90-98%, 50-100%, 60-100%, 70-100%, 80-100%, 90-100%, or 95-100%) identical to the sequence of any one selected from PS1101-1122, PS1218-1221, and PS1351-1398 (SEQ ID NOs: 447-468, 564-567, and 694-741). In some embodiments, the amino acid sequence is at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or 100% identical to PS1122 (SEQ ID NO: 468) or PS1381 (SEQ ID NO: 724). In some embodiments, the amino acid sequence is about 40% to about 100% (e.g., 80-98%, 80-95%, 80-90%, 85-95%, 90-98%, 50-100%, 60-100%, 70-100%, 80-100%, 90-100%, or 95-100%) identical to PS1122 (SEQ ID NO: 468) or PS1381 (SEQ ID NO: 724).

[0104] In some aspects, the present disclosure provides a recombinant or synthetic amino acid binding protein having an amino acid sequence that is at least 80% identical (e.g., at least 85%, at least 90%, at least 95%, at least 98%, 80-98%, 80-95%, 80-90%, 85-95%, or 90-98% identical) to SEQ ID NO:2, wherein the amino acid sequence comprises an amino acid substitution at one or more positions corresponding to G19, K26, S29, F30, D31, D32, T33, C34, V35, T47, G48, T53, T54, T57, E58, F59, N61, I63, D65, D68, E70, A71, H74, and T75 of SEQ ID NO:2.

[0105] In some embodiments, the amino acid sequence comprises amino acid substitutions at positions corresponding to I63 and E70 and at one or more positions corresponding to G19, K26, S29, D32, T47, G48, T53, T54, T57, E58, F59, N61, H74, and T75. In some embodiments, the amino acid sequence comprises amino acid substitutions at one or more positions corresponding to K26, D32, T47, I63, and E70. In some embodiments, the amino acid sequence comprises amino acid substitutions selected from G19R, K26R, S29Q, D32R, D32Y, T47K, T47L, T47R, G48R, G48Y, T53V, T54K, T57K, T57R, E58K, F59R, N61K, I63E, E70S, E70T, H74K, and T75E. In some embodiments, the amino acid substitution is selected from T47L, I63E and E70T. In some embodiments, the amino acid substitution is selected from K26R and D32R.

[0106] In some aspects, the present disclosure provides a recombinant or synthetic amino acid binding protein comprising a structure of formula (II) or a structural equivalent thereof:

[0107] β1–α1–β2–α2–α3

[0108] (II), wherein: each of β1 and β2 is a β strand; each of α1, α2 and α3 is an α helix; each instance of “–” is a loop; and at least a portion of each of α2, the loop between β1 and α1, and the loop between β2 and α2 forms a binding pocket for an amino acid ligand.

[0109] In some embodiments, the binding pocket comprises one or more of the following: (i) a volume of about (ii) Electrostatic potential is -3.0RTe c -1 or less, (iii) a plurality of hydrogen bond acceptors configured to form one or more hydrogen bonds in the presence of an amino acid ligand, and (iv) a plurality of van der Waals contact positions configured to form van der Waals interactions in the presence of an amino acid ligand. In some embodiments, the binding pocket comprises one, two, or three of (i), (ii), (iii), and (iv). In some embodiments, the binding pocket comprises (i), (ii), (iii), and (iv).

[0110] In some embodiments, the binding pocket comprises about In some embodiments, the volume of the binding pocket is at least At least At least At least or at least In some embodiments, the volume of the binding pocket does not exceed No more than No more than No more than or not more than In some embodiments, the volume of the binding pocket is to to to to to to or to Methods for determining the volume of a binding pocket are known in the art and will be apparent to the skilled artisan based on this disclosure. For example, in some embodiments, the volume of a binding pocket is determined using software configured to measure the geometric and topological properties of a protein. In some embodiments, the software can scan the surface of a protein using a specified probe radius to measure the volume of any cavity that overlaps the binding site directly or that overlaps the binding site indirectly through adjacent cavities within van der Waals contacts between each other. A non-limiting example of a suitable probe radius is a solvent probe radius (e.g., about ). A non-limiting example of suitable software is Computational Atlas of Surface Topography of Proteins (CASTp). See, e.g., W. Tian et al., CASTp 3.0: Computed atlas of surface topography of proteins. Nucleic Acids Res. 46, W363-W367 (2018), the relevant contents of which are incorporated herein by reference.

[0111] In some embodiments, the binding pocket has a -3.0 RTe c -1 In some embodiments, the electrostatic potential of the binding pocket is at least -4RTe. c -1 , at least -3RTe c -1 or at least -2RTe c -1 In some embodiments, the electrostatic potential of the binding pocket is -2RTe c -1 or lower, -3RTe c -1 or lower, or -4RTe c -1 Or lower. In some embodiments, the electrostatic potential of the binding pocket is -2RTec -1 To -3RTe c -1 、-2RTe c -1 To -4RTe c -1 or -3RTe c -1 To -4RTe c -1 Methods for determining the electrostatic potential of a binding pocket are known in the art and will be apparent to the skilled person based on this disclosure. For example, in some embodiments, the Adaptive Poisson-Boltzmann Solver (APBS) tool in PyMOL (PyMOL Molecular Graphics System, Version 2.0 LLC) to calculate the electrostatic surface potential of the binding pocket, using pdb2pqr and AMBER force fields to assign protonation states. See, for example, TJ Dolinsky et al., PDB2PQR: An automated pipeline for the setup of Poisson-Boltzmanne electrostatics calculations, Nucleic Acids Res., 32: W665-7 (2004); MG Lerner et al., APBS plugin for PyMOL–Version 2.4 (University of Michigan, Ann Arbor, MI, 2006); JW Ponder et al., Force fields for protein simulations, Adv. Protein Chem. 66: 27–85 (2003). In some cases, this solvent accessible surface area (SASA) can be considered as accessible to the peptide ligand.

[0112] In some embodiments, the binding pocket comprises a plurality of hydrogen bond receptors, which are configured to form one or more hydrogen bonds in the presence of an amino acid ligand. In some embodiments, the binding pocket forms at least two (e.g., at least three, at least four, at least five, 2-10, 4-10, 5-15, 5-10 hydrogen bonds) with the amino acid ligand. Methods for determining hydrogen bond interactions between a binding pocket and a ligand are known in the art and will be apparent to the technician on the basis of the present disclosure. For example, in some embodiments, hydrogen bond interactions are determined by computational modeling, as described in the examples herein (e.g., using atomic coordinates of protein-ligand structural data obtained from X-ray crystallography, or predicting protein-ligand three-dimensional structures by computational modeling).

[0113] In some embodiments, the multiple hydrogen bond acceptors include one or more atoms of the side chains of amino acid residues in the binding pocket. For example, in some embodiments, the binding pocket comprises at least three negatively charged amino acid side chains, each of which (e.g., aspartic acid, glutamic acid) forms a hydrogen bond with an amino acid ligand. In some embodiments, at least one negatively charged amino acid side chain forms a hydrogen bond with the amino terminus of an amino acid ligand. In some embodiments, the binding pocket comprises at least one polar uncharged amino acid side chain that forms a hydrogen bond with an amino acid ligand. In some embodiments, the multiple hydrogen bond acceptors include one or more atoms of the polypeptide backbone in the binding pocket (e.g., backbone carbonyl).

[0114] In some embodiments, the binding pocket forms one or more hydrogen bonds with the side chain of the amino acid ligand. For example, in some embodiments, the binding pocket forms one or more hydrogen bonds with the terminal amino acid side chain of the amino acid ligand (e.g., polypeptide). In some embodiments, the binding pocket forms one or more hydrogen bonds with the polypeptide backbone of the amino acid ligand (e.g., polypeptide). In some embodiments, the binding pocket forms one or more hydrogen bonds with the terminal amino acid in the polypeptide and one or more amino acids adjacent to the terminal amino acid (e.g., amino acids at position 1 and positions 2, 3, 4, and / or 5 relative to the end of the polypeptide).

[0115] In some embodiments, the binding pocket comprises a plurality of van der Waals contact positions that are configured to form van der Waals interactions in the presence of an amino acid ligand. Methods for determining van der Waals interactions between a binding pocket and a ligand are known in the art and will be apparent to the skilled person based on the present disclosure. For example, in some embodiments, van der Waals interactions are determined by computational modeling, as described in the Examples herein (e.g., using atomic coordinates of protein-ligand structural data obtained from X-ray crystallography, or by computational modeling to predict protein-ligand three-dimensional structures).

[0116] In some embodiments, the van der Waals contact positions include a plurality of atoms (e.g., 2-30, 5-25, 10-20, 2-10, 5-10) that are configured to form hydrophobic interactions with the amino acid ligand. In some embodiments, one or more of the plurality of atoms are non-polar atoms.

[0117] In some embodiments, the amino acid ligand is a polypeptide comprising at least three amino acids. In some embodiments, the amino acid ligand comprises the N-terminal amino acid of the polypeptide. In some embodiments, the N-terminal amino acid is arginine. In some embodiments, the length of the amino acid binding protein is at least 50 amino acids, at least 75 amino acids, at least 100 amino acids, 50-250 amino acids, 50-150 amino acids, or 100-200 amino acids.

[0118] In some embodiments, the loop between β2 and α2 comprises three or more negatively charged amino acids. In some embodiments, the loop between β2 and α2 comprises four or more negatively charged amino acids. In some embodiments, at least three negatively charged amino acids in the loop between β2 and α2 form hydrogen bonds with amino acid ligands. In some embodiments, at least one negatively charged amino acid in the loop between β2 and α2 forms a hydrogen bond with the amino terminus of the amino acid ligand. In some embodiments, the negatively charged amino acids are selected from aspartic acid and glutamic acid.

[0119] In some embodiments, at least one amino acid of α2 forms a hydrogen bond with an amino acid ligand. In some embodiments, α2 comprises one or more polar uncharged amino acids. In some embodiments, at least one polar uncharged amino acid of α2 forms a hydrogen bond with a side chain of an amino acid ligand.

[0120] In some embodiments, the loop between β1 and α1 comprises an amino acid sequence at least 80% identical to amino acids 27-42 of SEQ ID NO: 2. In some embodiments, the binding pocket is formed by amino acids at one or more positions corresponding to amino acids 31, 32, and 34-36 of SEQ ID NO:2.

[0121] In some embodiments, the loop between α1 and β2 comprises an amino acid sequence that is at least 50% identical to the sequence of amino acids 47-50 of SEQ ID NO: 2. In some embodiments, the amino acid sequence comprises an amino acid substitution at a position corresponding to T47 of SEQ ID NO: 2. In some embodiments, the amino acid substitution is T47L.

[0122] In some embodiments, β2-α2 comprises an amino acid sequence that is at least 80% identical to the sequence of amino acids 51-71 of SEQ ID NO: 2. In some embodiments, the binding pocket is formed by amino acids at one or more positions corresponding to amino acids 63, 65, 68, 70, and 71 of SEQ ID NO: 2. In some embodiments, the amino acid sequence comprises amino acid substitutions at one or more positions corresponding to I63 and E70 of SEQ ID NO: 2. In some embodiments, the amino acid substitutions are selected from I63E and E70T.

[0123] In some embodiments, structural equivalents are those that differ by no more than wherein at least 80% of the secondary structure α-carbon atoms are aligned with the structure of formula (II). In some embodiments, structural equivalents are structures with a root mean square difference of no more than No more than No more than or not more than A structure wherein at least 80% of the secondary structure α-carbon atoms are aligned with the structure of formula (II).

[0124] Methods for identifying structural equivalents of the structure of formula (II) are known in the art and will be apparent to the skilled artisan based on this disclosure. For example, in some embodiments, structural equivalents of formula (II) are those having a root mean square difference of no more than A structure wherein at least 80% of the secondary structure α-carbon atoms are aligned with a three-dimensional protein structure having an amino acid sequence selected from any one of PS1101-1122, PS1218-1221 and PS1351-1398 (SEQ ID NOs: 447-468, 564-567 and 694-741) (e.g., PS1122 (SEQ ID NO: 468), PS1381 (SEQ ID NO: 724)). Protein structure comparison can be performed in the following manner: determine the three-dimensional structure of a protein having an amino acid sequence selected from any one of PS1101-1122, PS1218-1221 and PS1351-1398 (SEQ ID NOs: 447-468, 564-567 and 694-741) (e.g., PS1122 (SEQ ID NO: 468), PS1381 (SEQ ID NO: 724)), and compare the three-dimensional structure of the protein with the structure of the candidate protein to determine whether the candidate protein is a structural equivalent. The three-dimensional protein structure can be determined by, for example, using atomic coordinates obtained from X-ray crystallography or by predicting the three-dimensional structure of a protein having an amino acid sequence selected from any one of PS1101-1122, PS1218-1221, and PS1351-1398 (SEQ ID NOs: 447-468, 564-567, and 694-741) (e.g., PS1122 (SEQ ID NO: 468), PS1381 (SEQ ID NO: 724)) using computational modeling. Methods for comparing protein structures and determining root mean square differences are known in the art (see, e.g., Kufareva I, Abagyan R. Methods of protein structure comparison. Methods Mol Biol. 2012; 857: 231-57).

[0125] C. Ntaq1 homology recognition agent

[0126] In some embodiments, the amino acid recognition agent of the present disclosure binds to an amino acid ligand (e.g., a polypeptide) comprising an N-terminal amino acid selected from glutamine, asparagine, glutamic acid, aspartic acid, cysteine-S-acetamide, or a modified variant thereof (e.g., a post-translational modified variant thereof, an oxidized variant thereof). In some embodiments, the amino acid recognition agent comprises an amino acid binding protein derived from an Ntaq1 protein (e.g., a Scleropages formosus Ntaq1 protein). For example, in some embodiments, the amino acid binding protein is an engineered variant comprising one or more modifications relative to SEQ ID NO: 3 described herein.

[0127] In some embodiments, the amino acid binding protein has a dissociation constant (K) of less than 2,000 nM, less than 1,500 nM, less than 1,000 nM, less than 750 nM, less than 500 nM, less than 250 nM, less than 150 nM, less than 100 nM, less than 50 nM, 10-2,000 nM, 25-1,000 nM, 50-500 nM, 10-100 nM, 25-250 nM, or 50-150 nM. D ) binds to the N-terminal glutamine.

[0128] In some embodiments, the amino acid binding protein binds one or more types of N-terminal amino acids (e.g., glutamine, asparagine, glutamic acid, aspartic acid, cysteine-S-acetamide, or modified variants thereof), wherein each type of binding interaction is characterized by a dissociation rate (k off ) is at least 0.1s -1 In some embodiments, the dissociation rate is about 0.1 s -1 About 1,000s -1 between (e.g., between about 0.5s -1 To about 500s -1 Between, about 0.1s -1 To about 100s -1 Between, in about 1s -1 To about 100s -1 Between or about 0.5s -1 To about 50s -1 In some embodiments, the dissociation rate is between about 0.5 s -1 About 20 seconds -1 In some embodiments, the dissociation rate is between about 2 s -1 About 20 seconds -1 In some embodiments, the dissociation rate is between about 0.5 s -1 To about 2s -1 between.

[0129] In some aspects, the present disclosure provides a recombinant or synthetic amino acid binding protein having an amino acid sequence that is at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98% or 100% identical to any one of the sequences selected from PS1258-1260, PS1315-1318, PS1457-1478, PS1480-1499, PS1633-1656, PS1737-1758, PS1821-1898, PS2014-2057 and PS2116-2137 (SEQ ID NOs: 604-606, 660-663, 792-833 and 836-1025). In some embodiments, the amino acid sequence is about 40% to about 100% (e.g., 80-98%, 80-95%, 80-90%, 85-95%, 90-98%, 50-100%, 60-100%, 70-100%, 80-100%, 90-100%, or 95-100%) identical to the sequence of any one selected from PS1258-1260, PS1315-1318, PS1457-1478, PS1480-1499, PS1633-1656, PS1737-1758, PS1821-1898, PS2014-2057, and PS2116-2137 (SEQ ID NOs: 604-606, 660-663, 792-833, and 836-1025). In some embodiments, the amino acid sequence is at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or 100% identical to PS1259 (SEQ ID NO: 605). In some embodiments, the amino acid sequence is about 40% to about 100% (e.g., 80-98%, 80-95%, 80-90%, 85-95%, 90-98%, 50-100%, 60-100%, 70-100%, 80-100%, 90-100%, or 95-100%) identical to PS1259 (SEQ ID NO: 605). In some embodiments, the amino acid sequence is at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or 100% identical to PS2132 (SEQ ID NO: 1020).In some embodiments, the amino acid sequence is about 40% to about 100% (e.g., 80-98%, 80-95%, 80-90%, 85-95%, 90-98%, 50-100%, 60-100%, 70-100%, 80-100%, 90-100%, or 95-100%) identical to PS2132 (SEQ ID NO: 1020).

[0130] In some aspects, the present disclosure provides a recombinant or synthetic amino acid binding protein having an amino acid sequence that is at least 80% identical (e.g., at least 85%, at least 90%, at least 95%, at least 98%, 80-98%, 80-95%, 80-90%, 85-95%, or 90-98% identical) to SEQ ID NO:3, wherein the amino acid sequence comprises an amino acid substitution at one or more positions corresponding to S22, C23, Y24, C25, E26, S39, W75, D76, Y77, H78, C85, N120, H145, and M146 of SEQ ID NO:3.

[0131] In some embodiments, the amino acid sequence comprises amino acid substitutions at positions corresponding to C25 and H78. In some embodiments, the amino acid sequence comprises amino acid substitutions at positions corresponding to S22, C25, H78, C85 and N120. In some embodiments, the amino acid substitutions are selected from S22E, C25S, H78Q, H78K, C85T, N120R and M146E. In some embodiments, the amino acid substitutions are selected from C25S, H78Q and M146E. In some embodiments, the amino acid substitutions are selected from C25S and H78Q. In some embodiments, the amino acid sequence comprises amino acid substitutions at positions corresponding to H78, wherein the amino acid substitutions are H78Q. In some embodiments, the amino acid sequence comprises amino acid substitutions at positions corresponding to H78, wherein the amino acid substitutions are H78K.

[0132] In some aspects, the present disclosure provides a recombinant or synthetic amino acid binding protein comprising a structure of formula (III) or a structural equivalent thereof:

[0133] α1–α2–α3β1–β2–β3–β4–β5–α4–β6α5–α6

[0134] (III)

[0135] wherein: each of α1, α2, α3, α4, α5 and α6 is an alpha helix; each of β1, β2, β3, β4, β5 and β6 is a beta strand; each instance of “–” is a loop; and at least a portion of each of α2, β3, β4, α5, the loop between α1 and α2, and the loop between β3 and β4 forms a binding pocket for an amino acid ligand.

[0136] In some embodiments, the binding pocket comprises one or more of the following: (i) a volume of about (ii) Electrostatic potential is -2.0RTe c -1 or less, (iii) a plurality of hydrogen bond acceptors or donors configured to form one or more hydrogen bonds in the presence of an amino acid ligand, (iv) a plurality of van der Waals contact sites configured to form van der Waals interactions in the presence of an amino acid ligand, and (v) at least one negatively charged amino acid and at least one positively charged amino acid. In some embodiments, the binding pocket comprises one, two, three or four of (i), (ii), (iii), (iv) and (v). In some embodiments, the binding pocket comprises (i), (ii), (iii), (iv) and (v).

[0137] In some embodiments, the binding pocket comprises about In some embodiments, the volume of the binding pocket is at least At least At least At least or at least In some embodiments, the volume of the binding pocket does not exceed No more than No more than No more than or not more than In some embodiments, the volume of the binding pocket is to to to to to to or to Methods for determining the volume of a binding pocket are known in the art and will be apparent to the skilled artisan based on this disclosure. For example, in some embodiments, the volume of a binding pocket is determined using software configured to measure the geometric and topological properties of a protein. In some embodiments, the software can scan the surface of a protein using a specified probe radius to measure the volume of any cavity that overlaps the binding site directly or that overlaps the binding site indirectly through adjacent cavities within van der Waals contacts between each other. A non-limiting example of a suitable probe radius is a solvent probe radius (e.g., about ). A non-limiting example of suitable software is Computational Atlas of Surface Topography of Proteins (CASTp). See, e.g., W. Tian et al., CASTp 3.0: Computed atlas of surface topography of proteins. Nucleic Acids Res. 46, W363-W367 (2018), the relevant contents of which are incorporated herein by reference.

[0138] In some embodiments, the binding pocket has -2.0 RTe c -1 In some embodiments, the electrostatic potential of the binding pocket is at least -3RTe. c -1 , at least -2RTe c -1 or at least -1RTe c -1 In some embodiments, the electrostatic potential of the binding pocket is -1RTe c -1 or lower, -2RTe c -1 or lower, or -3RTe c -1 Or lower. In some embodiments, the electrostatic potential of the binding pocket is -1RTe c -1 To -2RTe c -1 、-1RTe c -1 To -3RTe c -1 or -2RTe c -1 To -3RTe c -1Methods for determining the electrostatic potential of a binding pocket are known in the art and will be apparent to the skilled person based on this disclosure. For example, in some embodiments, the Adaptive Poisson-Boltzmann Solver (APBS) tool in PyMOL (PyMOL Molecular Graphics System, Version 2.0 LLC) to calculate the electrostatic surface potential of the binding pocket, using pdb2pqr and AMBER force fields to assign protonation states. See, for example, TJ Dolinsky et al., PDB2PQR: An automated pipeline for the setup of Poisson-Boltzmanne electrostatics calculations, Nucleic Acids Res., 32: W665-7 (2004); MG Lerner et al., APBS plugin for PyMOL–Version 2.4 (University of Michigan, Ann Arbor, MI, 2006); JW Ponder et al., Force fields for protein simulations, Adv. Protein Chem. 66: 27–85 (2003). In some cases, this solvent accessible surface area (SASA) can be considered as accessible to the peptide ligand.

[0139] In some embodiments, the binding pocket comprises a plurality of hydrogen bond acceptors or donors, which are configured to form one or more hydrogen bonds in the presence of an amino acid ligand. In some embodiments, the binding pocket forms at least two (e.g., at least three, at least four, at least five, 2-10, 4-10, 5-15, 5-10) hydrogen bonds with the amino acid ligand. Methods for determining hydrogen bond interactions between a binding pocket and a ligand are known in the art, and will be apparent to the technician on the basis of the present disclosure. For example, in some embodiments, hydrogen bond interactions are determined by computational modeling, as described in the examples herein (e.g., using atomic coordinates of protein-ligand structural data obtained from X-ray crystallography, or predicting protein-ligand three-dimensional structures by computational modeling).

[0140] In some embodiments, the multiple hydrogen bond acceptors or donors include one or more atoms of the side chains of amino acid residues in the binding pocket. For example, in some embodiments, the binding pocket comprises at least three negatively charged amino acid side chains, each of which (e.g., aspartic acid, glutamic acid) forms a hydrogen bond with an amino acid ligand. In some embodiments, at least one negatively charged amino acid side chain forms a hydrogen bond with the amino terminal of an amino acid ligand. In some embodiments, the binding pocket comprises at least one polar uncharged amino acid side chain (e.g., serine, glutamine), which forms a hydrogen bond with an amino acid ligand. In some embodiments, the multiple hydrogen bond acceptors or donors include one or more atoms of the polypeptide backbone in the binding pocket (e.g., backbone carbonyl). In some embodiments, the binding pocket comprises at least one negatively charged amino acid side chain and at least one positively charged amino acid side chain, each of which forms a hydrogen bond with an amino acid ligand. In some embodiments, at least one negatively charged amino acid side chain (e.g., aspartic acid, glutamic acid) forms a hydrogen bond with a backbone atom (e.g., nitrogen) of an amino acid ligand. In some embodiments, at least one positively charged amino acid side chain (eg, lysine) forms a hydrogen bond with a side chain atom of the amino acid ligand.

[0141] In some embodiments, the binding pocket forms one or more hydrogen bonds with the side chain of the amino acid ligand. For example, in some embodiments, the binding pocket forms one or more hydrogen bonds with the terminal amino acid side chain of the amino acid ligand (e.g., polypeptide). In some embodiments, the binding pocket forms one or more hydrogen bonds with the polypeptide backbone of the amino acid ligand (e.g., polypeptide). In some embodiments, the binding pocket forms one or more hydrogen bonds with the terminal amino acid in the polypeptide and one or more amino acids adjacent to the terminal amino acid (e.g., amino acids at position 1 and positions 2, 3, 4, and / or 5 relative to the end of the polypeptide).

[0142] In some embodiments, the binding pocket comprises a plurality of van der Waals contact positions that are configured to form van der Waals interactions in the presence of an amino acid ligand. Methods for determining van der Waals interactions between a binding pocket and a ligand are known in the art and will be apparent to the skilled person based on the present disclosure. For example, in some embodiments, van der Waals interactions are determined by computational modeling, as described in the Examples herein (e.g., using atomic coordinates of protein-ligand structural data obtained from X-ray crystallography, or by computational modeling to predict protein-ligand three-dimensional structures).

[0143] In some embodiments, the van der Waals contact positions include a plurality of atoms (e.g., 2-30, 5-25, 10-20, 2-10, 5-10) that are configured to form hydrophobic interactions with the amino acid ligand. In some embodiments, one or more of the plurality of atoms are non-polar atoms.

[0144] In some embodiments, the amino acid ligand is a polypeptide comprising at least three amino acids. In some embodiments, the amino acid ligand comprises the N-terminal amino acid of the polypeptide. In some embodiments, the N-terminal amino acid is selected from glutamine, asparagine, glutamic acid, aspartic acid and cysteine-S-acetamide. In some embodiments, the amino acid ligand is a polypeptide comprising N-terminal glutamine or asparagine. In some embodiments, the binding pocket comprises (i), (ii), (iii) and (iv), and the amino acid ligand is a polypeptide comprising N-terminal glutamine or asparagine. In some embodiments, the amino acid ligand is a polypeptide comprising N-terminal glutamic acid. In some embodiments, the binding pocket comprises (i), (ii), (iii), (iv) and (v), and the amino acid ligand is a polypeptide comprising N-terminal glutamic acid. In some embodiments, the length of the amino acid binding protein is at least 50 amino acids, at least 75 amino acids, at least 100 amino acids, 50-250 amino acids, 50-150 amino acids or 100-200 amino acids.

[0145] In some embodiments, each of α2 and β4 comprises at least one polar uncharged amino acid that forms a hydrogen bond with an amino acid ligand. In some embodiments, at least one polar uncharged amino acid of α2 is serine. In some embodiments, at least one polar uncharged amino acid of β4 is glutamine.

[0146] In some embodiments, α2 and the loop between α1 and α2 comprise an amino acid sequence that is at least 80% identical to the sequence of amino acids 18-40 of SEQ ID NO: 3. In some embodiments, the binding pocket is formed by amino acids at one or more positions corresponding to amino acids 23-26 of SEQ ID NO: 3. In some embodiments, the amino acid sequence comprises an amino acid substitution at a position corresponding to C25 of SEQ ID NO: 3. In some embodiments, the amino acid substitution is C25S.

[0147] In some embodiments, β3-β4 comprises an amino acid sequence that is at least 80% identical to the sequence of amino acids 73-85 of SEQ ID NO: 3. In some embodiments, the binding pocket is formed by amino acids at one or more positions corresponding to amino acids 75-78 of SEQ ID NO: 3. In some embodiments, the amino acid sequence comprises an amino acid substitution at a position corresponding to H78 of SEQ ID NO: 3. In some embodiments, the amino acid substitution is H78Q.

[0148] In some embodiments, α6 comprises an amino acid sequence that is at least 66% identical to amino acids 144-146 of SEQ ID NO: 3. In some embodiments, the binding pocket is formed by amino acids at one or more positions corresponding to amino acids 145-146 of SEQ ID NO:3.

[0149] In some embodiments, the amino acid binding protein comprises the structure of formula (III-A) or a structural equivalent thereof:

[0150] α1–α2–α3β1–β2–β3–β4–β5–α4–β6α5–α6–α7–β7α8

[0151] (III-A),

[0152] wherein: α7 and α8 are each an α helix; and β7 is a β strand.

[0153] In some embodiments, the amino acid binding protein comprises the structure of formula (III-B) or a structural equivalent thereof:

[0154] α1–α2–α3β1–β2–β3–β4

[0155] (III-B),

[0156] wherein: each of α1, α2 and α3 is an α helix; each of β1, β2, β3 and β4 is a β strand; each instance of “–” is a loop; and at least a portion of each of α2, β3, β4, the loop between α1 and α2, and the loop between β3 and β4 forms a binding pocket for an amino acid ligand.

[0157] In some embodiments, the binding pocket comprises: (i) at least one negatively charged amino acid that is configured to form a hydrogen bond with the amino acid ligand, and (ii) at least one positively charged amino acid that is configured to form a hydrogen bond with the amino acid ligand. In some embodiments, the amino acid ligand is a polypeptide comprising at least three amino acids. In some embodiments, the amino acid ligand comprises the N-terminal amino acid of the polypeptide. In some embodiments, the N-terminal amino acid is glutamic acid.

[0158] In some embodiments, at least one negatively charged amino acid forms hydrogen bonds with backbone atoms of an amino acid ligand. In some embodiments, at least one positively charged amino acid forms hydrogen bonds with side chain atoms of an amino acid ligand. In some embodiments, at least one negatively charged amino acid forms hydrogen bonds with backbone atoms of an amino acid ligand, and at least one positively charged amino acid forms hydrogen bonds with side chain atoms of an amino acid ligand. In some embodiments, at least one negatively charged amino acid forms hydrogen bonds with backbone atoms of an amino acid ligand. In some embodiments, at least one positively charged amino acid forms hydrogen bonds with side chain atoms of an amino acid ligand. In some embodiments, at least one negatively charged amino acid forms hydrogen bonds with backbone atoms of an amino acid ligand, and at least one positively charged amino acid forms hydrogen bonds with side chain atoms of an amino acid ligand.

[0159] In some embodiments, α2 comprises at least one negatively charged amino acid. In some embodiments, β4 comprises at least one positively charged amino acid. In some embodiments, α2 comprises at least one negatively charged amino acid, and β4 comprises at least one positively charged amino acid. In some embodiments, at least one negatively charged amino acid comprises glutamic acid. In some embodiments, at least one positively charged amino acid comprises lysine. In some embodiments, at least one negatively charged amino acid comprises glutamic acid, and at least one positively charged amino acid comprises lysine. In some embodiments, at least one negatively charged amino acid corresponds to E26 of SEQ ID NO:3. In some embodiments, at least one positively charged amino acid is a lysine substitution at a position corresponding to H78 of SEQ ID NO:3. In some embodiments, at least one negatively charged amino acid corresponds to E26 of SEQ ID NO:3, and at least one positively charged amino acid is a lysine substitution at a position corresponding to H78 of SEQ ID NO:3.

[0160] In some embodiments, α1-α2 comprises an amino acid sequence that is at least 80% identical to the sequence of amino acids 15-39 of SEQ ID NO:3. In some embodiments, the amino acid sequence comprises amino acid substitutions at positions corresponding to S22 and C25 of SEQ ID NO:3. In some embodiments, the amino acid substitutions are S22E and C25S. In some embodiments, β3-β4 comprises an amino acid sequence that is at least 80% identical to the sequence of amino acids 73-85 of SEQ ID NO:3. In some embodiments, the amino acid sequence comprises amino acid substitutions at positions corresponding to H78 and C85 of SEQ ID NO:3. In some embodiments, the amino acid substitutions are H78K and C85T.

[0161] In some embodiments, structural equivalents are those that differ by no more than wherein at least 80% of the secondary structure α-carbon atoms are aligned with the structure of formula (III), (III-A) or (III-B). In some embodiments, structural equivalents are structures having a root mean square difference of no more than No more than No more than or not more than A structure wherein at least 80% of the secondary structure α-carbon atoms are aligned with the structure of formula (III), (III-A) or (III-B).

[0162] Methods for identifying structural equivalents of structures of Formula (III), (III-A), or (III-B) are known in the art and will be apparent to the skilled artisan based on this disclosure. For example, in some embodiments, structural equivalents of Formula (III), (III-A), or (III-B) are those having a root mean square difference of no more than A structure wherein at least 80% of the secondary structure α-carbon atoms are aligned with a three-dimensional protein structure having an amino acid sequence selected from any one of PS1258-1260, PS1315-1318, PS1457-1478, PS1480-1499, PS1633-1656, PS1737-1758, PS1821-1898, PS2014-2057 and PS2116-2137 (SEQ ID NOs: 604-606, 660-663, 792-833 and 836-1025). Protein structure comparison can be performed by determining the three-dimensional structure of a protein having an amino acid sequence selected from any one of PS1258-1260, PS1315-1318, PS1457-1478, PS1480-1499, PS1633-1656, PS1737-1758, PS1821-1898, PS2014-2057 and PS2116-2137 (SEQ ID NOs: 604-606, 660-663, 792-833 and 836-1025), and comparing the three-dimensional structure of the protein with the structure of the candidate protein to determine whether the candidate protein is a structural equivalent. The three-dimensional protein structure can be determined by, for example, using atomic coordinates obtained from X-ray crystallography or by predicting the three-dimensional structure of a protein having an amino acid sequence selected from any one of PS1258-1260, PS1315-1318, PS1457-1478, PS1480-1499, PS1633-1656, PS1737-1758, PS1821-1898, PS2014-2057, and PS2116-2137 (SEQ ID NOs: 604-606, 660-663, 792-833, and 836-1025). Methods for comparing protein structures and determining root mean square differences are known in the art (see, e.g., Kufareva I, Abagyan R. Methods of protein structure comparison. Methods Mol Biol. 2012; 857: 231-57).

[0163] D. BIR domain homologous recognition agent

[0164] In some embodiments, the amino acid recognition agent of the present disclosure binds to an amino acid ligand (e.g., a polypeptide) comprising an N-terminal alanine. In some embodiments, the amino acid recognition agent comprises an amino acid binding protein derived from a protein containing a Baculoviral IAP repeat sequence (BIR) (e.g., a Homo sapiens BIR3 domain protein). For example, in some embodiments, the amino acid binding protein is an engineered variant comprising one or more modifications relative to SEQ ID NO: 511 described herein.

[0165] In some embodiments, the amino acid binding protein has a dissociation constant (K) of less than 2,000 nM, less than 1,500 nM, less than 1,000 nM, less than 750 nM, less than 500 nM, less than 250 nM, less than 150 nM, less than 100 nM, less than 80 nM, 10-2,000 nM, 25-1,000 nM, 50-500 nM, 10-150 nM, 25-75 nM, or 50-60 nM. D ) binds to the N-terminal alanine.

[0166] In some embodiments, the amino acid binding protein binds to the N-terminal alanine, wherein the binding interaction is characterized by a dissociation rate (k off ) is at least 0.1s -1 In some embodiments, the dissociation rate is about 0.1 s -1 About 1,000s -1 between (e.g., between about 0.5s -1 To about 500s -1 Between, about 0.1s -1 To about 100s -1 Between, in about 1s -1 To about 100s -1 Between or about 0.5s -1 To about 50s -1 In some embodiments, the dissociation rate is between about 0.5 s -1 About 20 seconds -1 In some embodiments, the dissociation rate is between about 2 s -1 About 20 seconds -1 In some embodiments, the dissociation rate is between about 0.5 s -1 To about 2s -1 between.

[0167] In some aspects, the present disclosure provides a recombinant or synthetic amino acid binding protein having an amino acid sequence that is at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98% or 100% identical to any one of PS1165-1166 (SEQ ID NOs: 511-512), PS1267 (SEQ ID NO: 613) and PS1399-1424 (SEQ ID NOs: 742-767). In some embodiments, the amino acid sequence is about 40% to about 100% (e.g., 80-98%, 80-95%, 80-90%, 85-95%, 90-98%, 50-100%, 60-100%, 70-100%, 80-100%, 90-100%, or 95-100%) identical to any one of PS1165-1166 (SEQ ID NOs: 511-512), PS1267 (SEQ ID NO: 613), and PS1399-1424 (SEQ ID NOs: 742-767). In some embodiments, the amino acid sequence is at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or 100% identical to PS1165 (SEQ ID NO: 511). In some embodiments, the amino acid sequence is about 40% to about 100% (e.g., 80-98%, 80-95%, 80-90%, 85-95%, 90-98%, 50-100%, 60-100%, 70-100%, 80-100%, 90-100%, or 95-100%) identical to PS1165 (SEQ ID NO:511).

[0168] In some aspects, the present disclosure provides a recombinant or synthetic amino acid binding protein having an amino acid sequence at least 80% identical to PS1165 (SEQ ID NO: 511) and comprising one or more tags described herein. In some embodiments, the amino acid sequence is at least 85%, at least 90%, at least 95%, at least 98%, or 100% identical to PS1165 (SEQ ID NO: 511). In some embodiments, the amino acid sequence is about 80% to about 100% (e.g., 80-98%, 80-95%, 80-90%, 85-95%, 90-98%, 90-100%, or 95-100%) identical to PS1165 (SEQ ID NO: 511).

[0169] E. Tandem recognition agent

[0170] In some embodiments, the amino acid recognition agent comprises a single polypeptide with a tandem copy of two or more amino acid binding proteins, wherein at least one of the two or more amino acid binding proteins is an amino acid binding protein of the present disclosure. As used herein, in some embodiments, the tandem arrangement or orientation of elements in a molecule refers to each element being connected end-to-end with the next element in a linear manner, so that the elements are fused in a tandem manner. For example, in some embodiments, a polypeptide with a tandem copy of two amino acid binding proteins refers to a fusion polypeptide, wherein the C-terminus of one protein is fused to the N-terminus of another protein. Similarly, a polypeptide with a tandem copy of two or more amino acid binding proteins refers to a fusion polypeptide, wherein the C-terminus of the first protein is fused to the N-terminus of the second protein, and the C-terminus of the second protein is fused to the N-terminus of the third protein, and so on. Such fusion polypeptides can include multiple copies of the same amino acid binding protein or multiple copies of different amino acid binding proteins. In some embodiments, the fusion polypeptide of the present application has at least two and up to ten amino acid binding proteins (e.g., at least 2 binders and up to eight, six, five, four or three binders). In some embodiments, the fusion polypeptides of the present application have five or fewer amino acid binding proteins (eg, two, three, four, or five amino acid binding proteins).

[0171] In some embodiments, the fusion polypeptide is provided by expressing a single coding sequence containing segments encoding monomeric amino acid binding protein subunits, which are separated by segments encoding flexible linkers, wherein the expression of the single coding sequence produces a single full-length polypeptide with two or more independent binding sites. In some embodiments, one or more monomer subunits are ClpS homologous proteins, UBR homologous proteins, or Ntaq1 homologous proteins. In some embodiments, the monomer subunits may be the same or different. When different, the monomer subunits may be different variants of the same parent homologous proteins, or they may be derived from different parent homologous proteins. In some embodiments, the fusion polypeptide comprises two or more ClpS homologous monomers, two or more UBR homologous monomers, or two or more Ntaq1 homologous monomers.

[0172] In some embodiments, at least one amino acid binding protein of the fusion polypeptide has an amino acid sequence selected from Table 1 (or has an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, 80-90%, 90-95%, 95-99% or more amino acid sequence identity with an amino acid sequence selected from Table 1). In some embodiments, each amino acid binding protein of the fusion polypeptide has an amino acid sequence that is at least 80% (e.g., 80-90%, 90-95%, 95-99% or more) identical to an amino acid sequence selected from Table 1 (or has an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, 80-90%, 90-95%, 95-99% or more amino acid sequence identity with an amino acid sequence selected from Table 1). In some embodiments, the amino acid binding proteins of the fusion polypeptide are modified and include one or more amino acid deletions, additions, or mutations relative to the sequences listed in Table 1. In some embodiments, the amino acid binding protein of the fusion polypeptide comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acids (which may or may not be consecutive amino acids) deleted, added or mutated relative to the sequence listed in Table 1.

[0173] In some embodiments, the amino acid binding proteins of the fusion polypeptide recognize the same set of one or more amino acids. In some embodiments, the amino acid binding proteins of the fusion polypeptide recognize different sets of one or more amino acids. In some embodiments, the amino acid binding proteins of the fusion polypeptide recognize overlapping sets of amino acids. In some embodiments, when the amino acid binding proteins of the fusion polypeptide recognize the same amino acids, they may recognize the amino acids with the same characteristic pulse pattern or with different characteristic pulse patterns.

[0174] In some embodiments, the amino acid binding protein of the fusion polypeptide is connected end-to-end by a covalent bond or a joint, and the joint covalently connects the C-terminus of a protein to the N-terminus of another protein. In the context of the fusion polypeptide of the present application, a joint refers to one or more amino acids that connect two amino acid binding proteins in the fusion polypeptide and do not form a part of the polypeptide sequence corresponding to any one of the two proteins. In some embodiments, the joint comprises at least two amino acids (for example, at least 2, 3, 4, 5, 6, 8, 10, 15, 25, 50, 100 or more amino acids). In some embodiments, the joint comprises up to 5, up to 10, up to 15, up to 25, up to 50 or up to 100 amino acids. In some embodiments, the joint comprises about 2 to about 200 amino acids (for example, about 2 to about 100, about 5 to about 50, about 2 to about 20, about 5 to about 20 or about 2 to about 30 amino acids).

[0175] Thus, in some aspects, the present disclosure provides an amino acid recognition agent comprising a polypeptide having a first amino acid binding protein and a second amino acid binding protein connected end-to-end, wherein the first amino acid binding protein and the second amino acid binding protein are separated by a linker comprising at least two amino acids.

[0176] In some embodiments, each of the first amino acid binding protein and the second amino acid binding protein independently has an amino acid sequence that is at least 80% identical (e.g., at least 85%, at least 90%, at least 95%, at least 98%, 80-100%, 85-100%, 90-100%, 95-100% or 100% identical) to PS961 (SEQ ID NO: 314). In some embodiments, the amino acid recognition agent comprises a polypeptide having an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, 80-100%, 85-100%, 90-100%, 95-100% or 100% identical to any one of PS1038, PS1222 and PS1223 (SEQ ID NO: 389, 568 and 569).

[0177] In some embodiments, each of the first amino acid binding protein and the second amino acid binding protein independently has an amino acid sequence that is at least 80% identical (e.g., at least 85%, at least 90%, at least 95%, at least 98%, 80-100%, 85-100%, 90-100%, 95-100% or 100% identical) to PS1122 (SEQ ID NO: 468). In some embodiments, the amino acid recognition agent comprises a polypeptide having an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, 80-100%, 85-100%, 90-100%, 95-100% or 100% identical to any one of PS1219-PS1221 (SEQ ID NO: 565-567).

[0178] In some embodiments, each of the first amino acid binding protein and the second amino acid binding protein independently has an amino acid sequence that is at least 80% identical (e.g., at least 85%, at least 90%, at least 95%, at least 98%, 80-100%, 85-100%, 90-100%, 95-100% or 100% identical) to PS1259 (SEQ ID NO: 605). In some embodiments, the amino acid recognition agent comprises a polypeptide having an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, 80-100%, 85-100%, 90-100%, 95-100% or 100% identical to PS1599 (SEQ ID NO: 835).

[0179] In some aspects, the application provides a nucleic acid encoding a single polypeptide having a tandem copy of two or more amino acid binding proteins. In some embodiments, the nucleic acid is an expression construct encoding the fusion polypeptide of the application. In some embodiments, the expression construct encodes a fusion polypeptide having at least two and up to ten amino acid binding proteins (e.g., at least two and up to three, four, five, six, seven, eight, nine or ten amino acid binding proteins). In some embodiments, the expression construct encodes a fusion polypeptide having five or less amino acid binding proteins (e.g., two, three, four or five amino acid binding proteins).

[0180] F. Shielding identification agent

[0181] According to the embodiments described herein, the single-molecule polypeptide sequencing method can be performed by irradiating the surface-fixed polypeptide with excitation light and detecting the luminescence generated by the label connected to the amino acid recognition agent. In some cases, the radiation and / or non-radiative attenuation generated by the label may cause photodamage to the polypeptide. The inventors have found that by incorporating shielding elements into the amino acid recognition agent, photodamage can be reduced and recognition time can be extended. See, for example, PCT International Publication No. WO2020102741A1 filed on November 15, 2019 and PCT International Publication No. WO2021236983A2 filed on May 20, 2021, which describe shielded recognition molecules in detail, and the relevant contents are incorporated herein by reference in their entirety.

[0182] Therefore, in some aspects, the present disclosure provides a shielded recognition agent, which comprises at least one amino acid recognition agent (e.g., amino acid binding protein) described herein, at least one detectable tag, and a shielding element (e.g., "shield") that forms a covalent or non-covalent linking group between the recognition agent and the tag. In some embodiments, the shield forms a covalent or non-covalent linking group between one or more amino acid binding proteins and one or more tags.

[0183] In some embodiments, the shield recognition agent comprises a fusion polypeptide having an amino acid binding protein of the present disclosure and a protein shield connected end-to-end (e.g., in a C-terminal to N-terminal manner). In some embodiments, the protein shield comprises a marker protein, such as a fluorescent protein or a non-fluorescent protein comprising a luminescent tag.

[0184] In some embodiments, the amino acid binding protein and the protein shield are connected end-to-end by a covalent bond or a linker, and the linker covalently connects the C-terminus of one protein to the N-terminus of another protein. In some embodiments, the linker in the fusion polypeptide refers to one or more amino acids in the fusion polypeptide that connect the amino acid binding protein and the protein shield and do not form a part of the polypeptide sequence corresponding to the amino acid binding protein or the protein shield. In some embodiments, the linker comprises at least two amino acids (e.g., at least 2, 3, 4, 5, 6, 8, 10, 15, 25, 50, 100 or more amino acids). In some embodiments, the linker comprises up to 5, up to 10, up to 15, up to 25, up to 50 or up to 100 amino acids. In some embodiments, the linker comprises about 2 to about 200 amino acids (e.g., about 2 to about 100, about 5 to about 50, about 2 to about 20, about 5 to about 20, or about 2 to about 30 amino acids).

[0185] In some embodiments, the protein shield of the fusion polypeptide is a protein with a molecular weight of at least 10 kDa. For example, in some embodiments, the protein shield is a protein with a molecular weight of at least 10 kDa and up to 500 kDa (e.g., about 10 kDa to about 250 kDa, about 10 kDa to about 150 kDa, about 10 kDa to about 100 kDa, about 20 kDa to about 80 kDa, about 15 kDa to about 100 kDa, or about 15 kDa to about 50 kDa). In some embodiments, the protein shield of the fusion polypeptide is a protein comprising at least 25 amino acids. For example, in some embodiments, the protein shield is a protein comprising at least 25 and up to 1,000 amino acids (e.g., about 100 to about 1,000 amino acids, about 100 to about 750 amino acids, about 500 to about 1,000 amino acids, about 250 to about 750 amino acids, about 50 to about 500 amino acids, about 100 to about 400 amino acids, or about 50 to about 250 amino acids).

[0186] In some embodiments, the protein shield is a polypeptide comprising one or more tag proteins. In some embodiments, the protein shield is a polypeptide comprising at least two tag proteins. In some embodiments, the at least two tag proteins are identical (e.g., the polypeptide comprises at least two copies of the tag protein sequence). In some embodiments, the at least two tag proteins are different (e.g., the polypeptide comprises at least two different tag protein sequences). Examples of tag proteins include, but are not limited to, Fasciola hepatica 8-kDa antigen (Fh8), maltose binding protein (MBP), N utilization substance (NusA), thioredoxin (Trx), small ubiquitin-like modifier (SUMO), glutathione-S-transferase (GST), solubility enhancing peptide sequence (SET), IgG domain B1 of protein G (GB1), IgG repeat domain ZZ of protein A (ZZ), mutant dehalogenase (HaloTag), solubility enhancing universal tag (SNUT), cell chaperone protein (Seventeen kilodalton protein, Skp), bacteriophage T7 protein kinase (T7PK), Escherichia coli secretory protein A (EspA), monomeric bacteriophage T7 0.3 protein (Orc protein; Mocr), Escherichia coli trypsin inhibitor (Ecotin), calcium binding protein (CaBP), stress response arsenate reductase (ArsC), N-terminal fragment of translation initiation factor IF2 (IF2-domain I), stress response protein (e.g., RpoA, SlyD, Tsf, RpoS, PotD, Crr) and Escherichia coli acidic protein (e.g., msyB, yjgD, rpoD). See, e.g., Costa, S., et al. "Fusion tags for protein solubility, purification and immunogenicity in Escherichia coli: the novel Fh8 system." Front Microbiol. 2014 Feb 19; 5: 63, the relevant contents of which are incorporated herein by reference.

[0187] The shielding elements of the present disclosure can advantageously absorb, deflect or otherwise block radiation and / or non-radiative attenuation emitted by the tag of the amino acid recognition agent. Therefore, one skilled in the art can easily select a protein shield for a suitable fusion polypeptide. For example, the inventors have demonstrated the use of various types of protein shields in fusion polypeptides, including those with binding to enzymes (e.g., DNA polymerase, glutathione S-transferase), transport proteins (e.g., maltose binding protein), fluorescent proteins (e.g., GFP), and commercially available tag proteins (e.g., ) fused amino acid binding protein polypeptide. The inventors further demonstrated the use of fusion polypeptides with multiple copies of tandemly directed protein shields. For example, see PCT International Publication No. WO2021236983A2 filed on May 20, 2021.

[0188] Thus, in some embodiments, the disclosure provides a fusion polypeptide having one or more tandemly oriented amino acid binding proteins fused to one or more tandemly oriented protein shields. In some embodiments, when the fusion polypeptide comprises two or more tandemly oriented binders and / or two or more tandemly oriented shields, the end of one of the two or more binders is end-to-end connected to the end of one of the two or more shields. Fusion polypeptides having tandem copies of two or more binders are described elsewhere herein, and in some embodiments, such fusions may further comprise a protein shield connected end-to-end to one of the two or more binders.

[0189] Other example configurations of shielding recognition agents and shielding elements (e.g., oligonucleotide shields, avidin shields) have been described and are contemplated for use according to the present disclosure. For example, see PCT International Publication No. WO2020102741A1 filed on November 15, 2019 and PCT International Publication No. WO2021236983A2 filed on May 20, 2021, the relevant contents of which are incorporated herein by reference.

[0190] G. Labels

[0191] In some embodiments, the amino acid identification agent of the present disclosure comprises one or more labels. In some embodiments, the one or more labels include detectable labels, such as luminescent labels or conductivity labels. As described herein, in some embodiments, one or more chemical features of a polypeptide can be determined by monitoring the change of a signal (such as a signal pulse), and these signal changes correspond to the binding events between one or more amino acid identification agents and the polypeptide. In some embodiments, the amino acid identification agent is included in a detectable label that produces a signal change during the binding event between the amino acid identification agent and the polypeptide. Therefore, as used herein, the detectable label of an amino acid identification agent can refer to any label that can produce a detectable signal change during the binding event between the amino acid identification agent and the polypeptide.

[0192] In some embodiments, one or more labels of amino acid recognition agents include luminescent labels. In some embodiments, the luminescent label includes at least one fluorophore dye molecule (e.g., at least 2, at least 3, at least 4, at least 5, 20 or less, 15 or less, 10 or less fluorophore dye molecules). In some embodiments, the luminescent label includes at least one FRET pair, and the FRET pair includes a donor label and an acceptor label. Luminescent labels and examples of their use in the present disclosure are described in detail elsewhere herein.

[0193] In some embodiments, one or more labels of the amino acid recognition agent include a conductivity label. In some embodiments, the conductivity label is a charge label, such as a charged polymer. Examples of charge labels include dendrimers, nanoparticles, nucleic acids, and other polymers with multiple charged groups. In some embodiments, the conductivity label can be uniquely identified by its net charge (e.g., net positive charge or net negative charge), its charge density, and / or the number of its charged groups.

[0194] In some embodiments, one or more tags of an amino acid recognition agent include a tag sequence. For example, in some embodiments, an amino acid recognition agent includes a tag sequence that provides one or more functions other than amino acid binding. In some embodiments, the tag sequence includes at least one biotin ligase recognition sequence, and the biotin ligase recognition sequence allows the biotinylation of the recognition agent (e.g., incorporation of one or more biotin molecules, including biotin and a double biotin moiety). In some embodiments, the tag sequence includes two biotin ligase recognition sequences directed in series. In some embodiments, the biotin ligase recognition sequence refers to an amino acid sequence recognized by a biotin ligase, and the biotin ligase catalyzes the covalent connection between the sequence and the biotin molecule. Each biotin ligase recognition sequence of the tag sequence can be covalently linked to the biotin moiety, so the tag sequence with multiple biotin ligase recognition sequences can be covalently linked to multiple biotin molecules. The region of the tag sequence with one or more biotin ligase recognition sequences can generally be referred to as a biotinylation tag or a biotinylation sequence. In some embodiments, a double biotin or double biotin moiety can refer to two biotins bound to two biotin ligase recognition sequences directed in series.

[0195] Other examples of functional sequences in tag sequences include purification tags, cleavage sites, and other parts for purifying and / or modifying the recognition agent. Table 2 provides a list of non-limiting sequences of tag sequences, any one or more of which can be used in combination with any of the amino acid recognition agents of the present application (e.g., in combination with the sequences listed in Table 1). It should be understood that the tag sequences shown in Table 2 are intended to be non-limiting, and the recognition agent according to the present application may include any one or more of the tag sequences (e.g., His tags and / or biotinylation tags) rearranged in other ways at the N-terminus or C-terminus of the recognition agent polypeptide, or at an internal position, between the N-terminus and the C-terminus, or as practiced in the art.

[0196] In some embodiments, one or more tags of an amino acid recognition agent include a biotin moiety. In some embodiments, the biotin moiety includes at least one biotin molecule (e.g., 1, 2, 3, 4 or more biotin molecules). In some embodiments, the biotin moiety is a double biotin moiety. In some embodiments, the biotin moiety includes at least one biotin molecule connected to at least one biotin ligase recognition sequence. For example, in some embodiments, the one or more tags include a tag sequence, the tag sequence includes two tandemly directed biotin ligase recognition sequences, each biotin ligase recognition sequence is connected thereto with a biotin molecule. In some embodiments, the biotin moiety includes at least one biotin molecule connected to an amino acid recognition agent by a method other than the tag sequence. For example, in some embodiments, the at least one biotin molecule is chemically conjugated to an amino acid (e.g., an unnatural amino acid) of an amino acid binding protein.

[0197] In some embodiments, one or more tags of the amino acid recognition agent include one or more polyol moieties (e.g., one or more moieties selected from dextran, polyvinyl pyrrolidone, polyethylene glycol, polypropylene glycol, polyoxyethylene glycol, and polyvinyl alcohol). For example, in some embodiments, the amino acid recognition agent is pegylated. In some embodiments, polyol modification (e.g., pegylation) can limit the extent of nonspecific adhesion to the surface of a substrate (e.g., a sequencing chip). In some embodiments, polyol modification can limit the extent of aggregation or interaction between amino acid recognition agents and other species present in other recognition agents, cleavage reagents, or sequencing reaction mixtures. Pegylation can be performed by incubating the recognition agent (e.g., amino acid binding protein, such as ClpS protein) with mPEG4-NHS ester, the ester marking primary amines, such as surface-exposed lysine side chains. Other types of polyethylene glycol and other polyol modification methods are known in the art.

[0198] It should be understood that in some embodiments, the amino acid recognition agent of the present disclosure may include one or more different types of labels described herein. For example, in some embodiments, the amino acid recognition agent includes one or more labels selected from a detectable label (e.g., a luminescent label, a conductivity label), a label sequence (e.g., a purification label, a cleavage site, a biotinylation sequence), a biotin moiety, and a polyol moiety. In some embodiments, the amino acid recognition agent includes a detectable label (e.g., a luminescent label, a conductivity label) and one or more labels selected from a label sequence (e.g., a purification label, a cleavage site, a biotinylation sequence), a biotin moiety, and a polyol moiety.

[0199] In some embodiments, one or more labels of amino acid identification agents include luminescent labels. As used herein, luminescent labels are molecules that absorb one or more photons and can subsequently emit one or more photons after one or more time periods. In some embodiments, the term can be used interchangeably with "label", "detectable label" or "luminescent molecule", depending on the context. Luminescent labels according to certain embodiments described herein can refer to the luminescent labels of amino acid identification agents, the luminescent labels of cleavage reagents (e.g., peptidases, such as aminopeptidases) or the luminescent labels of another labeled composition described herein.

[0200] In some embodiments, the luminescent label comprises a first chromophore and a second chromophore. In some embodiments, the excited state of the first chromophore can be relaxed by energy transfer to the second chromophore. In some embodiments, the energy transfer is Förster resonance energy transfer (FRET). Such FRET pairs can be used to provide luminescent labels with characteristics that make labels easier to distinguish from multiple luminescent labels in a mixture, or for providing fluorescence induced by binding to limit background fluorescence, as described elsewhere herein. In other embodiments, the FRET pair comprises a first chromophore of the first luminescent label and a second chromophore of the second luminescent label. In certain embodiments, the FRET pair can absorb excitation energy within a first spectral range and emit luminescence within a second spectral range.

[0201] In some embodiments, the luminescent label refers to a fluorophore or dye. Typically, the luminescent label comprises an aromatic or heteroaromatic compound and can be pyrene, anthracene, naphthalene, naphthylamine, acridine, stilbene, indole, benzindole, oxazole, carbazole, thiazole, benzothiazole, benzoxazole, phenanthridine, phenoxazine, porphyrin, quinoline, ethidium, benzamide, cyanine, carbocyanine, salicylate, anthranilate, coumarin, fluorescein, rhodamine, xanthene or other similar compounds.

[0202] In some embodiments, the luminescent label comprises a dye selected from one or more of the following: 5 / 6-carboxyrhodamine 6G, 5-carboxyrhodamine 6G, 6-carboxyrhodamine 6G, 6-TAMRA, STAR 440SXP、 STAR 470SXP、 STAR 488、 STAR 512、 STAR 520SXP、 STAR 580、 STAR 600、 STAR 635、 STAR 635P、 STAR RED、Alexa 350、Alexa 405、Alexa 430、Alexa 480、Alexa 488、Alexa 514、Alexa 532、Alexa 546、Alexa 555、Alexa 568、Alexa 594、Alexa 610-X、Alexa 633、Alexa 647、Alexa 660、Alexa 680、Alexa 700、Alexa 750、Alexa 790、AMCA、ACT 390、ACT 425、ACT 465、ACT 488、ACT 495、ACT 514、ACT520、ACT 532、ACT 542、ACT 550、ACT 565、ACT 590、ACT 610、ACT 620、ACT 633、ACT 647、ACT 647N、ACT 655、ACT 665、ACT 680、ACT 700、ACT 725、ACT 740、ACTOxa12、ACT Rho101、ACT Rho11、ACT Rho12、ACT Rho13、ACT Rho14、ACT Rho3B、ATTORho6G、ATTO Thio12、BD Horizon TM V450、 493 / 501、 530 / 550、 558 / 568、 564 / 570、 576 / 589、 581 / 591、 630 / 650、 650 / 665、 FL、 FL-X、 R6G、 TMR、 TR、CAL Gold 540、CAL Green 510、CAL Orange 560、CAL Red 590、CAL Red 610、CAL Red 615、CAL Red635、 Blue、CF TM 350、CF TM 405M、CF TM 405S、CF TM 488A、CF TM 514、CF TM 532、CF TM 543、CF TM 546、CF TM 555、CF TM 568、CF TM 594、CF TM 620R、CF TM 633、CF TM 633-V1、CF TM 640R、CF TM 640R-V1、CF TM 640R-V2、CF TM 660C、CF TM 660R、CF TM 680、CF TM 680R、CF TM 680R-V1、CF TM 750、CF TM 770、CF TM 790、Chromeo TM642、Chromis 425N、Chromis 500N、Chromis 515N、Chromis 530N、Chromis 550A、Chromis 550C、Chromis 550Z、Chromis 560N、Chromis 570N、Chromis 577N、Chromis600N、Chromis 630N、Chromis 645A、Chromis 645C、Chromis 645Z、Chromis 678A、Chromis678C、Chromis 678Z、Chromis 770A、Chromis 770C、Chromis 800A、Chromis 800C、Chromis830A、Chromis 830C、 3、 3.5、 3B、 5、 5.5、 7、 350、 405、 415-Co1、 425Q、 485-LS、 488、 504Q、 510-LS、 515-LS、 521-LS、 530-R2、 543Q、 550、 554-R0、 554-R1、 590-R2、 594、 610-B1、 615-B2、 633、 633-B1、 633-B2、 650、 655-B1、 655-B2、 655-B3、 655-B4、 662Q、 675-B1、 675-B2、 675-B3、<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 675-B4<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 679-C5<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 680<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 683Q<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 690-B1<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 690-B2<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 696Q<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 700-B1<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 700-B1<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 730-B1<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 730-B2<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 730-B3<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 730-B4<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 747<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 747-B1<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 747-B2<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 747-B3<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 747-B4<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 755<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 766Q<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 775-B2<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 775-B3<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 775-B4<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 780-B1<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 780-B2<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 780-B3<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 800<h2 style=";text-align:left;direction:ltr"> 830-B2、Dyomics-350、Dyomics-350XL、Dyomics-360XL、Dyomics-370XL、Dyomics-375XL、Dyomics-380XL、Dyomics-390XL、Dyomics-405、Dyomics-415、Dyomics-430、Dyomics-431、Dyomics-478、Dyomics-480XL、Dyomics-481XL、Dyomics-485XL、Dyomics-490、Dyomics-495、Dyomics-505、Dyomics-510XL、Dyomics-511XL、Dyomics-520XL、Dyomics-521XL、Dyomics-530、Dyomics-547、Dyomics-547P1、Dyomics-548、Dyomics-549、Dyomics-549P1、Dyomics-550、Dyomics-554、Dyomics-555、Dyomics-556、Dyomics-560、Dyomics-590、Dyomics-591、Dyomics-594、Dyomics-601XL、Dyomics-605、Dyomics-610、Dyomics-615、Dyomics-630、Dyomics-631、Dyomics-632、Dyomics-633、Dyomics-634、Dyomics-635、Dyomics-636、Dyomics-647、Dyomics-647P1、Dyomics-648、Dyomics-648P1、Dyomics-649、Dyomics-649P1、Dyomics-650、Dyomics-651、Dyomics-652、Dyomics-654、Dyomics-675、Dyomics-676、Dyomics-677、Dyomics-678、Dyomics-679P1、Dyomics-680、Dyomics-681、Dyomics-682、Dyomics-700、Dyomics-701、Dyomics-703、Dyomics-704、Dyomics-730、Dyomics-731、Dyomics-732、Dyomics-734、Dyomics-749、Dyomics-749P1、Dyomics-750、Dyomics-751、Dyomics-752、Dyomics-754、Dyomics-776, Dyomics-777, Dyomics-778, Dyomics-780, Dyomics-781, Dyomics-782, Dyomics-800, Dyomics-831, 450, Eosin, FITC, Fluorescein, HiLyte TM Fluor 405, HiLyte TM Fluor 488, HiLyte TM Fluor 532, HiLyte TM Fluor 555, HiLyte TM Fluor 594, HiLyte TM Fluor 647, HiLyte TM Fluor 680, HiLyte TM Fluor 750, 680LT, 750, 800CW、JOE、 640R, Red 610, Red 640, Red 670, Red 705, Lissamine Rhodamine B, Napthofluorescein, Oregon 488. Oregon 514、Pacific Blue TM 、Pacific Green TM 、Pacific Orange TM , PET, PF350, PF405, PF415, PF488, PF505, PF532, PF546, PF555P, PF568, PF594, PF610, PF633P, PF647P, 570, 670, 705, Rhodamine 123, Rhodamine 6G, Rhodamine B, Rhodamine Green, Rhodamine Green-X, Rhodamine Red, ROX, Seta TM 375. Seta TM 470、Seta TM 555. Seta TM 632. Seta TM 633. Seta TM 650、Seta TM660、Seta TM 670、Seta TM 680、Seta TM 700、Seta TM 750、Seta TM 780、Seta TM APC-780, Seta TM PerCP-680, Seta TM R-PE-670、Seta TM 646, SeTau 380, SeTau 425, SeTau 647, SeTau 405, Square 635, Square 650, Square 660, Square 672, Square 680, Sulforhodamine 101, TAMRA, TET, Texas TMR, TRITC, Yakima Yellow TM , Zy3, Zy5, Zy5.5 and Zy7.

[0203] In some aspects, the disclosure provides methods and compositions for performing polypeptide analysis (e.g., amino acid identification) based on one or more luminescent properties of luminescent tags. In some embodiments, luminescent tags are identified based on luminescence lifetime, luminescence intensity, brightness, absorption spectrum, emission spectrum, luminescence quantum yield, or a combination of two or more thereof. In some embodiments, various types of luminescent tags can be distinguished from each other based on differences in luminescence lifetime, luminescence intensity, brightness, absorption spectrum, emission spectrum, luminescence quantum yield, or a combination of two or more thereof.

[0204] In some embodiments, luminescence is detected by exposing a luminescent tag to a series of separate light pulses and evaluating the timing or other characteristics of each photon emitted from the tag. In some embodiments, information from multiple photons emitted sequentially from a tag is gathered and evaluated to identify the tag and thereby identify the associated barcode site. In some embodiments, the luminescence lifetime of a tag is determined by multiple photons emitted sequentially from the tag, and the luminescence lifetime can be used to identify the tag. In some embodiments, the luminescence intensity of a tag is determined by multiple photons emitted sequentially from the tag, and the luminescence intensity can be used to identify the tag. In some embodiments, the luminescence lifetime and luminescence intensity of a tag are determined by multiple photons emitted sequentially from the tag, and the luminescence lifetime and luminescence intensity can be used to identify the tag.

[0205] In some aspects of the present disclosure, a single molecule is exposed to multiple individual light pulses, and a series of emitted photons are detected and analyzed. In some embodiments, the series of emitted photons provides information about single molecules that are present and do not change in the mixture during the experiment. However, in some embodiments, the series of emitted photons provides information about a series of different molecules that exist in the mixture at different times (e.g., as a reaction or process proceeds).

[0206] In some embodiments, the luminescent tag absorbs a photon and emits a photon after a period of time. In some embodiments, the luminescence lifetime of the tag can be determined or estimated by measuring the period of time. In some embodiments, the luminescence lifetime of the tag can be determined or estimated by measuring multiple periods of multiple pulse events and emission events. In some embodiments, the luminescence lifetime of the tag can be distinguished among the luminescence lifetimes of multiple types of tags by measuring the period of time. In some embodiments, the luminescence lifetime of the tag can be distinguished among the luminescence lifetimes of multiple types of tags by measuring multiple periods of multiple pulse events and emission events. In some embodiments, the tags are identified or distinguished among multiple types of tags by determining or estimating the luminescence lifetime of the tags. In some embodiments, the tags are identified or distinguished among multiple types of tags by distinguishing the luminescence lifetime of the tags among multiple luminescence lifetimes of multiple types of tags.

[0207] The luminescence lifetime of a luminescent tag can be determined using any suitable method (e.g., by measuring the lifetime using a suitable technique or by determining the time-related characteristics of the emission). In some embodiments, determining the luminescence lifetime of a tag includes determining the lifetime relative to another tag. In some embodiments, determining the luminescence lifetime of a tag includes determining the lifetime relative to a reference. In some embodiments, determining the luminescence lifetime of a tag includes measuring the lifetime (e.g., fluorescence lifetime). In some embodiments, determining the luminescence lifetime of a tag includes determining the time characteristics of one or more indication lifetimes. In some embodiments, the luminescence lifetime of a tag can be determined based on the distribution of multiple emission events (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more emission events) occurring in one or more time-gated windows relative to an excitation pulse. For example, the luminescence lifetime of a tag can be distinguished from multiple tags with different luminescence lifetimes based on the distribution of the photon arrival time measured about the excitation pulse.

[0208] It should be understood that the luminescence lifetime of a luminescent tag indicates the timing of photons emitted after the tag reaches an excited state, and the tag can be distinguished by information indicating the timing of the photons. Some embodiments may include distinguishing the tag from multiple tags based on the luminescence lifetime of the tag by measuring the time associated with the photons emitted by the tag. The time distribution can provide an indication of the luminescence lifetime, which can be determined from the distribution. In some embodiments, the tag can be distinguished from multiple tags based on the time distribution, for example, by comparing the time distribution with a reference distribution corresponding to a known tag. In some embodiments, the value of the luminescence lifetime is determined by the time distribution.

[0209] As used herein, in some embodiments, luminescence intensity refers to the number of emission photons emitted per unit time by a luminescent tag that is excited by delivering pulsed excitation energy. In some embodiments, luminescence intensity refers to the number of emission photons detected per unit time that are emitted by a tag excited by the delivery of pulsed excitation energy and detected by a particular sensor or sensor group.

[0210] As used herein, in some embodiments, brightness refers to a parameter that reports the average emission intensity of each luminescent tag. Thus, in some embodiments, "emission intensity" can be used to generally refer to the brightness of a composition comprising one or more tags. In some embodiments, the brightness of a tag is equal to the product of its quantum yield and extinction coefficient.

[0211] As used herein, in some embodiments, luminescence quantum yield refers to the fraction of excitation events that result in emission events at a given wavelength or within a given spectral range, and is generally less than 1. In some embodiments, the luminescence quantum yield of the luminescent tags described herein is between 0 and about 0.001, between about 0.001 and about 0.01, between about 0.01 and about 0.1, between about 0.1 and about 0.5, between about 0.5 and 0.9, or between about 0.9 and 1. In some embodiments, the tag is identified by determining or estimating the luminescence quantum yield.

[0212] As used herein, in some embodiments, the excitation energy is a light pulse from a light source. In some embodiments, the excitation energy is in the visible spectrum. In some embodiments, the excitation energy is in the ultraviolet spectrum. In some embodiments, the excitation energy is in the infrared spectrum. In some embodiments, the excitation energy is in or near the absorption maximum of the luminescent label, and multiple emission photons are detected from the luminescent label. In certain embodiments, the excitation energy is between about 500nm and about 700nm (for example, between about 500nm and about 600nm, between about 600nm and about 700nm, between about 500nm and about 550nm, between about 550nm and about 600nm, between about 600nm and about 650nm, or between about 650nm and about 700nm). In certain embodiments, the excitation energy can be monochromatic or limited to a spectral range. In some embodiments, the spectral range has a range of about 0.1nm to about 1nm, about 1nm to about 2nm, or about 2nm to about 5nm. In some embodiments, the spectral range has a range of about 5 nm to about 10 nm, about 10 nm to about 50 nm, or about 50 nm to about 100 nm.

[0213] II. Peptide Analysis

[0214] In some aspects, the present application provides a method for determining at least one chemical characteristic of a polypeptide by monitoring a signal of a signal pulse corresponding to an interaction between the polypeptide and at least one amino acid recognition agent described herein, and determining at least one chemical characteristic of the polypeptide based on a characteristic pattern in the signal.

[0215] Figure 1 The non-limiting example of analyzing polypeptide structure by the unimolecular binding interaction in the detection polypeptide degradation process is illustrated. Example signal traces show different association (e.g., combination) events corresponding to signal changes at different time points. As shown in the figure, the binding event between the amino acid recognition agent and the polypeptide end can produce a change in signal amplitude, and this change lasts for a period of time. The different binding events of different amino acids exposed at the polypeptide end are illustrated. As described herein, the amino acid "exposed" at the polypeptide end is still connected on the polypeptide and becomes the amino acid (e.g., alone or with one or more additional amino acids) of the terminal amino acid in the degradation process along with the removal of the previous terminal amino acid.

[0216] As generally described, the binding event between the amino acid recognition agent and the different types of amino acids at the end of the polypeptide produces a unique change in the signal, referred to herein as a characteristic pattern, which can be used to determine the chemical characteristics of the polypeptide. In some embodiments, the characteristic pattern corresponding to the terminal amino acid of a type can be used to determine the structural information of the terminal amino acid and one or more amino acids adjacent to the terminal amino acid. Therefore, in some embodiments, the characteristic pattern corresponding to the terminal amino acid of a type can be used to determine the structural information of at least two (e.g., at least three, at least four, at least five, two, three, four or two to five) amino acids in the polypeptide.

[0217] In some embodiments, the transition from one characteristic pattern to another characteristic pattern indicates the cleavage of amino acids. As used herein, in some embodiments, amino acid cleavage refers to the removal of at least one amino acid from the end of a polypeptide (e.g., removing at least one terminal amino acid from a polypeptide). In some embodiments, amino acid cleavage is determined by inferring based on the time period between characteristic patterns. In some embodiments, amino acid cleavage is determined by detecting the signal change generated when a labeled cleavage reagent is combined with the terminal amino acid of the polypeptide. As the amino acids are sequentially cleaved from the end of the polypeptide during the polypeptide degradation process, a series of changes in amplitude or a series of signal pulses will be detected.

[0218] In some embodiments, signal pulse information can be extracted by applying threshold levels to one or more parameters of signal data, thereby analyzing signal data. For example, in some embodiments, a threshold amplitude level (threshold magnitude level) can be applied to the signal data of the signal trace. In some embodiments, the threshold amplitude level is the minimum difference between the signal detected at a certain time point and the baseline determined by a given data set. In some embodiments, a signal pulse is assigned to each part of the data indicating that the amplitude changes exceed the threshold amplitude level and lasts for a period of time. In some embodiments, a threshold duration (thresholdtime duration) can be applied to the data portion that meets the threshold amplitude level to determine whether the portion is assigned as a signal pulse. For example, experimental artifacts may cause the amplitude to change more than the threshold amplitude level, but the duration is not enough to assign signal pulses with the required confidence (for example, transient binding events that may not distinguish between amino acid types, non-specific detection events such as diffusion to the observation area or reagents adhere to the observation area). Therefore, in some embodiments, signal pulses are extracted from the signal data based on the threshold amplitude level and the threshold duration.

[0219] In some embodiments, the peak amplitude of a signal pulse is determined by averaging the amplitudes detected over a period of time that is sustained above a threshold amplitude level. It should be understood that in some embodiments, as used herein, "signal pulse" may refer to a change in signal data that is sustained above a baseline (e.g., raw signal data) for a period of time, or signal pulse information extracted therefrom (e.g., processed signal data).

[0220] In some embodiments, the signal pulse information can be analyzed to identify different types of amino acids in a polypeptide based on different characteristic patterns in a series of signal pulses. Figure 1 As shown, the signal pulse information indicates different types of amino acids (e.g., arginine, leucine, isoleucine, phenylalanine) at the end of the polypeptide. For example, the signal pulse detected at the earliest time point provides information indicating the presence of (at least) arginine at the end of the polypeptide based on the first characteristic pattern, while the signal pulse detected at the latest time point provides information indicating the presence of at least phenylalanine at the end of the polypeptide based on the second characteristic pattern.

[0221] In some embodiments, each signal pulse of the characteristic pattern comprises a pulse duration corresponding to the binding event between the amino acid recognition agent and the amino acid ligand. In some embodiments, the pulse duration is a characteristic of the dissociation rate of the combination. In some embodiments, each signal pulse of the characteristic pattern is separated from another signal pulse of the characteristic pattern by the pulse duration. In some embodiments, the pulse duration is a characteristic of the binding rate of the combination. In some embodiments, the amplitude change in the signal of the signal pulse can be determined based on the difference between the baseline and the signal pulse peak. In some embodiments, the characteristic pattern is determined based on the pulse duration. In some embodiments, the characteristic pattern is determined based on the pulse duration and the pulse duration. In some embodiments, the characteristic pattern is determined based on any one or more of the pulse duration, the pulse duration and the amplitude change.

[0222] Therefore, if Figure 1 As explained, in some embodiments, polypeptide analysis is performed by detecting a series of signal pulses, which indicate the binding of one or more amino acid recognition agents to the continuous amino acids exposed at the end of the polypeptide in a continuous degradation reaction. A series of signal pulses can be analyzed to determine a characteristic pattern in the series of signal pulses, and the time course of the characteristic pattern can be used to determine the chemical features in the polypeptide amino acid sequence.

[0223] As described herein, signal pulse information can be used to identify amino acids based on characteristic patterns in a series of signal pulses. In some embodiments, the characteristic pattern comprises a plurality of signal pulses, and each signal pulse comprises a pulse duration. In some embodiments, a plurality of signal pulses can be characterized by summary statistics (e.g., mean, median, time decay constant) of the pulse duration distribution in the characteristic pattern. In some embodiments, the average pulse duration of the characteristic pattern is between about 1 millisecond to about 10 seconds (e.g., between about 1ms to about 1s, between about 1ms to about 100ms, between about 1ms to about 10ms, between about 10ms to about 10s, between about 100ms to about 10s, between about 1 second to about 10 seconds, between about 10ms to about 100ms, or between about 100ms to about 500ms). In some embodiments, the average pulse duration is between about 50 milliseconds to about 2 seconds, about 50 milliseconds to about 500 milliseconds, or about 500 milliseconds to about 2 seconds.

[0224] In some embodiments, different characteristic patterns corresponding to different types of amino acids in a single polypeptide can be distinguished from each other based on the statistically significant differences of summary statistics. For example, in some embodiments, a characteristic pattern can be distinguished from another characteristic pattern based on the difference of the average pulse duration of at least 10 milliseconds (for example, between about 10ms to about 10s, between about 10ms to about 1s, between about 10ms to about 100ms, between about 100ms to about 10s, between about 1s to about 10s, or between about 100ms to about 1 second). In some embodiments, the difference in average pulse duration is at least 50ms, at least 100ms, at least 250ms, at least 500ms or more. In some embodiments, the difference in average pulse duration is between about 50ms to about 1 second, about 50ms to about 500ms, about 50ms to about 250ms, about 100ms to about 500ms, about 250ms to about 500ms, or about 500ms to about 1 second. In some embodiments, the average pulse duration of one characteristic pattern differs from the average pulse duration of another characteristic pattern by about 10-25%, 25-50%, 50-75%, 75-100%, or more than 100%, such as by a factor of about 2, 3, 4, 5, or more. It should be understood that in some embodiments, smaller differences in average pulse duration between different characteristic patterns may require more pulse durations in each characteristic pattern to be distinguished from each other with statistical confidence.

[0225] In some embodiments, a characteristic pattern generally refers to a plurality of binding events between an amino acid of a polypeptide and a means for binding the amino acid (e.g., an amino acid recognition molecule). In some embodiments, a characteristic pattern comprises at least 10 binding events (e.g., at least 25, at least 50, at least 75, at least 100, at least 250, at least 500, at least 1,000 or more binding events). In some embodiments, a characteristic pattern comprises about 10 to about 1,000 binding events (e.g., about 10 to about 500 binding events, about 10 to about 250 binding events, about 10 to about 100 binding events, or about 50 to about 500 binding events). In some embodiments, a plurality of binding events are detected as a plurality of signal pulses.

[0226] In some embodiments, a characteristic pattern refers to a plurality of signal pulses, which can be characterized by summary statistics as described herein. In some embodiments, a characteristic pattern comprises at least 10 signal pulses (e.g., at least 25, at least 50, at least 75, at least 100, at least 250, at least 500, at least 1,000 or more signal pulses). In some embodiments, a characteristic pattern comprises about 10 to about 1,000 signal pulses (e.g., about 10 to about 500 signal pulses, about 10 to about 250 signal pulses, about 10 to about 100 signal pulses, or about 50 to about 500 signal pulses).

[0227] In some embodiments, characteristic pattern refers to multiple binding events between amino acid recognition molecules and polypeptide amino acids that occur in the time interval before amino acid removal (e.g., cleavage event). In some embodiments, characteristic pattern refers to multiple binding events that occur in the time interval between two cleavage events (e.g., before amino acid removal and after the amino acid removal previously exposed at the end). In some embodiments, the time interval of characteristic pattern is between about 1 minute to about 30 minutes (e.g., between about 1 minute to about 20 minutes, between about 1 minute to 10 minutes, between about 5 minutes to about 20 minutes, between about 5 minutes to about 15 minutes, or between about 5 minutes to about 10 minutes).

[0228] In some embodiments, a series of signal pulses comprises a series of changes in optical signal amplitude over time. In some embodiments, a series of changes in optical signals comprises a series of luminescence changes produced during the binding event. In some embodiments, luminescence is produced by a detectable label associated with one or more reagents of the sequencing reaction. For example, in some embodiments, each of the one or more amino acid recognition agents comprises a luminescent label. In some embodiments, the cleavage reagent comprises a luminescent label. Examples of luminescent labels and uses thereof according to the present application are provided herein.

[0229] In some embodiments, a series of signal pulses comprises a series of changes in the amplitude of the electrical signal over time. In some embodiments, a series of changes in the electrical signal comprises a series of changes in conductance produced during the binding event. In some embodiments, conductance is produced by a detectable label associated with one or more reagents of the sequencing reaction. For example, in some embodiments, each of the one or more amino acid recognition agents comprises a conductivity label. Examples of conductivity labels and uses thereof according to the present application are provided elsewhere herein. Methods for identifying single molecules using conductivity labels have been described (see, e.g., U.S. Patent Publication No. 2017 / 0037462).

[0230] In some embodiments, a series of changes in conductance include a series of changes in conductance through a nanopore. For example, methods for evaluating receptor-ligand interactions using nanopores have been described (Thakur, AK & Movileanu, L. (2019) Nature Biotechnology 37 (1)). The inventors have recognized and appreciated that such nanopores can be used to monitor polypeptide sequencing reactions according to the present application. Therefore, in some embodiments, the present disclosure provides a method for polypeptide analysis, comprising contacting a single polypeptide molecule with one or more amino acid recognition agents described herein, wherein the single polypeptide molecule is fixed to a nanopore. In some embodiments, the method further comprises detecting a series of changes in conductance through a nanopore, which changes indicate the binding of one or more amino acid recognition agents to continuous amino acids exposed at the end of a single polypeptide during degradation.

[0231] As described herein, in some embodiments, amino acid recognition agents of the present disclosure can be used for determining at least one chemical signature of polypeptide.In some embodiments, determining at least one chemical signature comprises determining the amino acid type present at the polypeptide end and / or the amino acid type present at one or more positions adjacent to the terminal amino acid.In some embodiments, determining the amino acid type comprises determining actual amino acid identity, for example, by determining which one of 20 kinds of naturally occurring amino acids exists.In some embodiments, the amino acid type is selected from alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, selenocysteine, serine, threonine, tryptophan, tyrosine and valine.

[0232] In some embodiments, determining at least one chemical feature of a polypeptide comprises determining a possible subset of amino acids that may be present in the polypeptide. In some embodiments, this can be achieved by determining that an amino acid is not one or more specific amino acids (and therefore may be any of the other amino acids). In some embodiments, this can be achieved by determining which amino acids in a specified subset of amino acids (e.g., based on size, charge, hydrophobicity, post-translational modification, binding properties) may be present in the polypeptide (e.g., using a recognition agent that binds to a specified subset of two or more amino acids).

[0233] In some embodiments, determining at least one chemical characteristic of the polypeptide comprises determining that an amino acid comprises a post-translational modification. Non-limiting examples of post-translational modifications include acetylation (e.g., acetylated lysine), ADP-ribosylation, caspase cleavage, citrullination, formylation, N-linked glycosylation (e.g., glycosylated asparagine), O-linked glycosylation (e.g., glycosylated serine, glycosylated threonine), hydroxylation, methylation (e.g., methylated lysine, methylated arginine), myristoylation (e.g., myristoylated glycine), neddylation, nitration (e.g., nitrated tyrosine), chlorination (e.g., chlorinated tyrosine), oxidation / reduction (e.g., oxidized cysteine, oxidized methionine), palmitoylation (e.g., palmitoylated cysteine), phosphorylation, prenylation (e.g., prenylated cysteine), S-nitrosylation (e.g., S-nitrosylated cysteine, S-nitrosylated methionine), sulfation, sumoylation (e.g., sumoylated lysine), and ubiquitination (e.g., ubiquitinated lysine).

[0234] In some embodiments, determining at least one chemical feature of a polypeptide comprises determining that an amino acid comprises an arginine post-translational modification. For example, as described herein, amino acid recognition agents of the present disclosure are capable of distinguishing different arginine modifications, including symmetric dimethylarginine (SDMA), asymmetric dimethylarginine (ADMA) and citrullinated arginine.

[0235] In some embodiments, determining at least one chemical characteristic of a polypeptide comprises determining that an amino acid comprises a phosphorylated side chain. For example, in some embodiments, determining at least one chemical characteristic of a polypeptide comprises determining that an amino acid comprises a phosphorylated threonine (e.g., phospho-threonine). In some embodiments, determining at least one chemical characteristic of a polypeptide comprises determining that an amino acid comprises a phosphorylated tyrosine (e.g., phospho-tyrosine). In some embodiments, determining at least one chemical characteristic of a polypeptide comprises determining that an amino acid comprises a phosphorylated serine (e.g., phospho-serine).

[0236] In some embodiments, determining at least one chemical feature of the polypeptide comprises determining that the amino acid comprises a chemically modified variant, an unnatural amino acid, or a proteinogenic amino acid, such as selenocysteine ​​and pyrrolysine. Examples of unnatural amino acids include, but are not limited to, 2-naphthyl-alanine, statine, homoalanine, α-amino acids, β2-amino acids, β3-amino acids, γ-amino acids, 3-pyridyl-alanine, 4-fluorophenyl-alanine, cyclohexyl-alanine, N-alkyl amino acids, peptoid amino acids, homocysteine, penicillamine, 3-nitro-tyrosine, homophenyl-alanine, t-leucine, hydroxy-proline, 3-Abz, 5-F-tryptophan, and azabicyclo[2.2.1]heptane.

[0237] In some embodiments, determining at least one chemical feature of a polypeptide comprises determining that an amino acid comprises an oxidative modification. For example, as described herein, the amino acid recognition agents of the present disclosure are capable of distinguishing oxidized methionine and its unmodified variant. In some embodiments, the oxidative modification comprises an oxidatively damaged side chain of an amino acid. In some embodiments, the oxidatively damaged side chain comprises a cysteine-derived product (e.g., disulfide, sulfinic acid, sulfonic acid, sulfenic acid, S-nitrosocysteine), a tyrosine-derived product (e.g., dityrosine, 3,4-dihydroxyphenylalanine, 3-chlorotyrosine, 3-nitrotyrosine), a histidine-derived product (e.g., 2-oxohistidine, 4-hydroxy-2-oxohistidine, dihistidine, asparagine, aspartic acid, urea), a methionine-derived product (e.g., sulfoxide, sulfone), a tryptophan-derived product (e.g., ditryptophan, N-formylkynurenine, kynurenine, 2-oxo-tryptophan oxoindoleylalanine, 6-nitrotryptophan, hydroxytryptophan), a phenylalanine-derived product (e.g., m-tyrosine, o-tyrosine), or a universal side chain product (e.g., alcohol, hydroperoxide, aldehyde / ketone carbonyl). Examples of oxidatively damaged amino acids are known in the art, see, for example, Hawkins, CL, Davies, MJ Detection, identification, and quantification of oxidative protein modifications. J Biol Chem. 2019 Dec 20; 294(51): 19683-19708.

[0238] In some embodiments, determining at least one chemical feature of a polypeptide includes determining that an amino acid comprises a side chain characterized by one or more biochemical properties. For example, an amino acid may comprise a nonpolar aliphatic side chain, a positively charged side chain, a negatively charged side chain, a nonpolar aromatic side chain, or a polar uncharged side chain. Non-limiting examples of amino acids comprising nonpolar aliphatic side chains include alanine, glycine, valine, leucine, methionine, and isoleucine. Non-limiting examples of amino acids comprising positively charged side chains include lysine, arginine, and histidine. Non-limiting examples of amino acids comprising negatively charged side chains include aspartic acid and glutamic acid. Non-limiting examples of amino acids comprising nonpolar aromatic side chains include phenylalanine, tyrosine, and tryptophan. Non-limiting examples of amino acids comprising polar uncharged side chains include serine, threonine, cysteine, proline, asparagine, and glutamine.

[0239] In some embodiments, a protein or polypeptide can be digested into multiple smaller polypeptides, and the chemical characteristics of one or more of these smaller polypeptides can be determined. In some embodiments, a first end (e.g., N-terminus or C-terminus) of a polypeptide is fixed and the other end (e.g., C-terminus or N-terminus) is analyzed as described herein.

[0240] As used herein, sequencing a polypeptide refers to determining the sequence information of the polypeptide. In some embodiments, this may involve determining the identity of each contiguous amino acid of a portion (or all) of the polypeptide. However, in some embodiments, this may involve evaluating the identity of a subset of amino acids in the polypeptide (e.g., and determining the relative position of one or more amino acid types without determining the identity of each amino acid in the polypeptide). However, in some embodiments, amino acid content information can be obtained from the polypeptide without directly determining the relative positions of different types of amino acids in the polypeptide. Only the amino acid content can be used to infer the identity of the polypeptide present (e.g., by comparing the amino acid content with a polypeptide information database and determining which polypeptides have the same amino acid content).

[0241] In some embodiments, sequence information of multiple polypeptide products obtained from a longer polypeptide or protein (eg, by enzymatic and / or chemical cleavage) can be analyzed to reconstruct or infer the sequence of the longer polypeptide or protein.

[0242] In some respects, polypeptide analysis as herein described generates data, indicating how the polypeptide interacts with the binding means when being degraded by the cleavage means.As discussed above, data can include a series of characteristic patterns, corresponding to the binding events between the cleavage events at the ends of the polypeptide.In some embodiments, polypeptide analysis methods as herein described include contacting a single polypeptide molecule with the binding means and the cleavage means, wherein the binding means and the cleavage means are configured to realize at least 10 binding events before the cleavage event.In some embodiments, these means are configured to realize at least 10 binding events between two cleavage events.

[0243] In some embodiments, multiple single molecule sequencing reactions are performed in parallel in a sample well array. In some embodiments, the array comprises about 10,000 to about 1,000,000 sample wells. In some embodiments, the volume of the sample wells can be about 10 -21 Increase to about 10 -15 liters. Due to the small volume of the sample well, the detection of single molecule events may be feasible because only about one polypeptide may be in the sample well at any given time. Statistically, some sample wells may not contain single molecule sequencing reactions, while some may contain more than one single polypeptide molecule. However, a considerable number of sample wells may each contain single molecule reactions (e.g., at least 30% in some embodiments), so single molecule analysis can be performed in parallel on a large number of sample wells. In some embodiments, the binding means and the cleavage means are configured to achieve at least 10 binding events before the cleavage events in at least 10% (e.g., 10-50%, more than 50%, 25-75%, at least 80% or more) of the sample wells, wherein the single molecule reaction is performed in the sample well. In some embodiments, the binding means and the cleavage means are configured to achieve at least 10 binding events before the cleavage events of at least 50% (e.g., more than 50%, 50-75%, at least 80% or more) of the polypeptide amino acids in the single molecule reaction.

[0244] III. Compositions and Reaction Mixtures

[0245] In some aspects, the present disclosure provides a composition comprising two or more amino acid recognition agents, wherein at least one amino acid recognition agent comprises an amino acid binding protein as described herein. In some embodiments, the composition comprises at least one ClpS homologous protein as described herein. In some embodiments, the composition comprises at least one UBR homologous protein as described herein. In some embodiments, the composition comprises at least one Ntaq1 homologous protein as described herein. In some embodiments, the composition comprises two or more of ClpS homologous proteins, UBR homologous proteins, and Ntaq1 homologous proteins. In some embodiments, the composition comprises at least one ClpS homologous protein, at least one UBR homologous protein, and at least one Ntaq1 homologous protein.

[0246] In some embodiments, the composition further comprises at least one type of cleavage reagent. The composition comprising an amino acid recognition agent and a cleavage reagent may be referred to herein as a reaction mixture (e.g., a polypeptide sequencing reaction mixture). Peptidase, also referred to as a protease or proteolytic enzyme, is an enzyme that catalyzes the hydrolysis of peptide bonds. Peptidases digest polypeptides into shorter fragments, which can generally be divided into endopeptidases and exopeptidases, which cleave from the inside and end of a polypeptide chain, respectively. In some embodiments, a cleavage reagent comprises an exopeptidase (e.g., an aminopeptidase). Examples of suitable peptidases have been described and are contemplated for use according to the present disclosure. See, for example, PCT International Publication No. WO2020102741A1 filed on November 15, 2019 and PCT International Publication No. WO2021236983A2 filed on May 20, 2021, the relevant contents of which are incorporated herein by reference.

[0247] As described herein, the compositions of the present disclosure can be used to determine at least one chemical feature of a polypeptide based on a characteristic pattern. In some embodiments, polypeptide sequencing reaction conditions can be set to achieve a time interval allowing enough binding events, thereby providing a characteristic pattern with a desired confidence level. This can be achieved, for example, by setting reaction conditions based on various characteristics, including: reagent concentration, a molar ratio of a reagent to another reagent (e.g., the ratio of an amino acid recognition molecule to a cleavage reagent, a ratio of a recognition agent to another recognition agent, a ratio of a cleavage reagent to another cleavage reagent), the quantity of different reagent types (e.g., the quantity of different types of recognition agents and / or cleavage reagents, the number of recognition agent types relative to the number of cleavage reagent types), cleavage activity (e.g., peptidase activity), binding properties (e.g., the kinetics and / or thermodynamic binding parameters of recognition molecule binding), reagent modification (e.g., polyols and other recognition agents that can change the kinetics of interaction), reaction mixture components (e.g., one or more components, such as pH, buffers, salts, divalent cations, surfactants, and other reaction mixture components described herein), reaction temperature, and various other parameters and combinations thereof apparent to those skilled in the art. Reaction conditions can be set based on one or more aspects described herein, including, for example, signal pulse information (e.g., pulse duration, inter-pulse duration, amplitude variation), labeling strategy (e.g., number and / or type of fluorophores, linkers with or without shielding elements), surface modification (e.g., modification of sample well surfaces, including polypeptide immobilization), sample preparation (e.g., polypeptide fragment size, polypeptide modification for immobilization), and other aspects described herein.

[0248] In some embodiments, the peptide sequencing reaction according to the application is carried out under conditions of simultaneous amino acid recognition and cleavage in a single reaction mixture. For example, in some embodiments, the peptide sequencing reaction is carried out in a reaction mixture with a pH that allows binding events and cleavage events to occur. Therefore, in some embodiments, the pH of the reaction mixture is between about 6.5 and about 9.0. In some embodiments, the pH of the reaction mixture is between about 7.0 and about 8.5 (for example, between about 7.0 and about 8.0, between about 7.5 and about 8.5, between about 7.5 and about 8.0, or between about 8.0 and about 8.5).

[0249] In some embodiments, the polypeptide sequencing reaction is carried out in a reaction mixture comprising one or more buffers. In some embodiments, the reaction mixture comprises a buffer having a concentration of at least 10mM (e.g., at least 20mM and up to 250mM, at least 50mM, 10-250mM, 10-100mM, 20-100mM, 50-100mM or 100-200mM). In some embodiments, the reaction mixture comprises a buffer having a concentration of about 10mM to about 50mM (e.g., about 10mM to about 25mM, about 25mM to about 50mM or about 20mM to about 40mM). The example of buffer includes but is not limited to HEPES (4-(2-hydroxyethyl)-1-piperazineethanesulfonic acid), Tris (tris(hydroxymethyl)aminomethane) and MOPS (3-(N-morpholino)propanesulfonic acid).

[0250] In some embodiments, the polypeptide sequencing reaction is carried out in a reaction mixture comprising a salt having a concentration of at least 10 mM. In some embodiments, the reaction mixture comprises a salt having a concentration of at least 10 mM (e.g., at least 20 mM, at least 50 mM, at least 100 mM or more). In some embodiments, the reaction mixture comprises a salt having a concentration of about 10 mM to about 250 mM (e.g., about 20 mM to about 200 mM, about 50 mM to about 150 mM, about 10 mM to about 50 mM, or about 10 mM to about 100 mM). Examples of salts include, but are not limited to, sodium salts, potassium salts, and acetates, such as sodium chloride (NaCl), sodium acetate (NaOAc), and potassium acetate (KOAc).

[0251] Other examples of components used in the reaction mixture include divalent cations (e.g., Mg 2+ 、Co 2+ ) and a surfactant (e.g., polysorbate 20). In some embodiments, the reaction mixture comprises a divalent cation at a concentration of about 0.1 mM to about 50 mM (e.g., about 10 mM to about 50 mM, about 0.1 mM to about 10 mM, or about 1 mM to about 20 mM). In some embodiments, the reaction mixture comprises a surfactant at a concentration of at least 0.01% (e.g., about 0.01% to about 0.10%). In some embodiments, the reaction mixture comprises one or more components useful in single molecule analysis, such as an oxygen scavenging system (e.g., a PCA / PCD system or a pyranose oxidase / catalase / glucose system) and / or one or more triplet quenchers (e.g., trolox, COT, and NBA).

[0252] In some embodiments, the polypeptide sequencing reaction is carried out at a temperature that allows binding events and cleavage events to occur. In some embodiments, the polypeptide sequencing reaction is carried out at a temperature of at least 10°C. In some embodiments, the polypeptide sequencing reaction is carried out at a temperature of about 10°C to about 50°C (e.g., 15-45°C, 20-40°C, 25°C or about 25°C, 30°C or about 30°C, 35°C or about 35°C, 37°C or about 37°C). In some embodiments, the polypeptide sequencing reaction is carried out at room temperature or near room temperature.

[0253] As detailed above, Figure 1 The illustrated real-time sequencing process generally involves a cycle of amino acid recognition and terminal amino acid cleavage. In some embodiments, the relative incidence of recognition and cleavage can be controlled by the concentration difference between one or more amino acid recognition agents and at least one cleavage reagent. In some embodiments, the concentration difference can be optimized so that the number of signal pulses detected during the recognition of a single amino acid provides the required identification confidence interval. For example, if the signal data provided by the initial sequencing reaction has too few signal pulses between cleavage events to determine the characteristic pattern with the required confidence interval, the sequencing reaction can be repeated using a non-specific exopeptidase with a reduced concentration relative to the recognition molecule.

[0254] In some embodiments, polypeptide analysis according to the present disclosure can be carried out by contacting the polypeptide with a reaction mixture comprising one or more amino acid recognition agents and one or more cleavage reagents (e.g., peptidases). In some embodiments, the reaction mixture comprises an amino acid recognition agent at a concentration of about 10 nM to about 10 μM. In some embodiments, the reaction mixture comprises a cleavage reagent at a concentration of about 500 nM to about 500 μM.

[0255] In some embodiments, the reaction mixture comprises an amino acid recognition agent at a concentration of about 100nM to about 10μM, about 250nM to about 10μM, about 100nM to about 1μM, about 250nM to about 1μM, about 250nM to about 750nM, or about 500nM to about 1μM. In some embodiments, the reaction mixture comprises an amino acid recognition agent at a concentration of about 100nM, about 250nM, about 500nM, about 750nM, or about 1μM. In some embodiments, the reaction mixture comprises a lysis reagent at a concentration of about 500nM to about 250μM, about 500nM to about 100μM, about 1μM to about 100μM, about 500nM to about 50μM, about 1μM to about 100μM, about 10μM to about 200μM, or about 10μM to about 100μM. In some embodiments, the reaction mixture comprises a lysis reagent at a concentration of about 1 μM, about 5 μM, about 10 μM, about 30 μM, about 50 μM, about 70 μM, or about 100 μM.

[0256] In some embodiments, the reaction mixture comprises an amino acid recognition agent at a concentration of about 10nM to about 10μM and a cleavage reagent at a concentration of about 500nM to about 500μM. In some embodiments, the reaction mixture comprises an amino acid recognition agent at a concentration of about 100nM to about 1μM and a cleavage reagent at a concentration of about 1μM to about 100μM. In some embodiments, the reaction mixture comprises an amino acid recognition agent at a concentration of about 250nM to about 1μM and a cleavage reagent at a concentration of about 10μM to about 100μM. In some embodiments, the reaction mixture comprises an amino acid recognition agent at a concentration of about 500nM and a cleavage reagent at a concentration of about 25μM to about 75μM. In some embodiments, the concentration of the amino acid recognition agent in the reaction mixture and / or the concentration of the cleavage reagent are as described elsewhere herein.

[0257] In some embodiments, the reaction mixture comprises an amino acid recognition agent and a cleavage agent in a molar ratio of about 500: 1, about 400: 1, about 300: 1, about 200: 1, about 100: 1, about 75: 1, about 50: 1, about 25: 1, about 10: 1, about 5: 1, about 2: 1, or about 1: 1. In some embodiments, the reaction mixture comprises an amino acid recognition agent and a cleavage agent in a molar ratio of about 10: 1 to about 200: 1. In some embodiments, the reaction mixture comprises an amino acid recognition agent and a cleavage agent in a molar ratio of about 50: 1 to about 150: 1. In some embodiments, the molar ratio of the amino acid recognition agent to the cleavage agent in the reaction mixture is between about 1:1,000 and about 1:1 or between about 1:1 and about 100:1 (e.g., 1:1,000, about 1:500, about 1:200, about 1:100, about 1:10, about 1:5, about 1:2, about 1:1, about 5:1, about 10:1, about 50:1, about 100:1). In some embodiments, the molar ratio of the amino acid recognition agent to the cleavage agent in the reaction mixture is between about 1:100 and about 1:1 or between about 1:1 and about 10:1. In some embodiments, the molar ratio of the amino acid recognition agent to the cleavage agent in the reaction mixture is as described elsewhere herein.

[0258] In some embodiments, the reaction mixture comprises one or more amino acid recognition agents and one or more cleavage agents. In some embodiments, the reaction mixture comprises at least three amino acid recognition agents and at least one cleavage agent. In some embodiments, the reaction mixture comprises two or more cleavage agents. In some embodiments, the reaction mixture comprises at least one and up to ten cleavage agents (for example, 1-3 cleavage agents, 2-10 cleavage agents, 1-5 cleavage agents, 3-10 cleavage agents). In some embodiments, the reaction mixture comprises at least three and up to thirty kinds of amino acid recognition agents (for example, 3-25 kinds, 3-20 kinds, 3-10 kinds, 3-5 kinds, 5-30 kinds, 5-20 kinds, 5-10 kinds or 10-20 kinds of amino acid recognition agents). In some embodiments, the one or more amino acid recognition agents include at least one amino acid binding protein selected from Table 1.

[0259] In some embodiments, the reaction mixture comprises more than one amino acid recognition agent and / or more than one cleavage agent. In some embodiments, the reaction mixture described as comprising more than one amino acid recognition agent (or cleavage agent) means that the mixture has more than one type of amino acid recognition agent (or cleavage agent). For example, in some embodiments, the reaction mixture comprises two or more amino acid binding proteins, wherein two or more amino acid binding proteins refer to two or more types of amino acid binding proteins. In some embodiments, one type of amino acid binding protein has an amino acid sequence different from another type of amino acid binding protein in the reaction mixture. In some embodiments, one type of amino acid binding protein has a label different from another type of amino acid binding protein in the reaction mixture. In some embodiments, one type of amino acid binding protein is different from another type of amino acid binding protein in the reaction mixture. The amino acid associated (e.g., combined) of one type of amino acid binding protein is different. In some embodiments, one type of amino acid binding protein is different from another type of amino acid binding protein in the reaction mixture. The amino acid subset associated (e.g., combined) of one type of amino acid binding protein is different from another type of amino acid binding protein in the reaction mixture.

[0260] IV. Devices and Systems

[0261] In some aspects, the method according to the present disclosure can be performed using a system that allows single molecule analysis. The system may include an integrated device and an instrument configured to interface with the integrated device. The integrated device may include a pixel array, wherein each pixel includes a sample hole and at least one photodetector. The sample hole of the integrated device may be formed on the surface of the integrated device or through the surface of the integrated device, and is configured to receive a sample placed on the surface of the integrated device. In general, the sample hole can be considered as an array of sample holes. Multiple sample holes can have a suitable size and shape so that at least a portion of the sample holes receive a single sample (e.g., a single molecule, such as a polypeptide). In some embodiments, the number of samples in the sample hole can be distributed in the sample holes of the integrated device so that some sample holes contain one sample, while other sample holes contain zero, two or more samples.

[0262] Excitation light is provided to the integrated device from one or more light sources outside the integrated device. The optical components of the integrated device can receive the excitation light from the light source and guide the light to the sample hole array of the integrated device and illuminate the illumination area within the sample hole. In some embodiments, the sample hole can have a structure that allows the sample to remain near the surface of the sample hole, which can easily deliver the excitation light to the sample and detect the emission light from the sample. The sample located in the illumination area can emit light in response to being illuminated by the excitation light. For example, the sample can be labeled with a fluorescent label that emits light in response to achieving an excited state by irradiation with the excitation light. The emission light emitted by the sample can then be detected by one or more photodetectors within the pixel corresponding to the sample hole, where the sample is analyzed. According to some embodiments, when performed on a sample hole array that can range in number from about 10,000 pixels to 1,000,000 pixels, multiple samples can be analyzed in parallel.

[0263] The integrated device may include an optical system for receiving excitation light and directing the excitation light between the sample well array. The optical system may include one or more grating couplers configured to couple the excitation light to other optical components of the integrated device and direct the excitation light to the other optical components. For example, the optical system may include an optical component that directs the excitation light from the grating coupler to the sample well array. Such optical components may include a beam splitter, an optical combiner, and a waveguide. In some embodiments, one or more beam splitters may couple the excitation light from the grating coupler and deliver the excitation light to at least one waveguide. According to some embodiments, the beam splitter may have a structure that allows the excitation light to be transmitted substantially uniformly on all waveguides, so that each waveguide receives a substantially similar amount of excitation light. Such an embodiment may improve the performance of the integrated device by improving the uniformity of the excitation light received by the sample wells of the integrated device. For example, examples of suitable components for coupling excitation light to a sample hole and / or directing emission light to a light detector for inclusion in an integrated device are described in U.S. patent application Ser. No. 14 / 821,688, filed on Aug. 7, 2015, and entitled “INTEGRATED DEVICE FOR PROBING, DETECTING AND ANALYZING MOLECULES,” and U.S. patent application Ser. No. 14 / 543,865, filed on Nov. 17, 2014, and entitled “INTEGRATED DEVICE WITH EXTERNAL LIGHT SOURCE FOR PROBING, DETECTING, AND ANALYZING MOLECULES,” both of which are incorporated herein by reference in their entireties. Examples of suitable grating couplers and waveguides that may be implemented in an integrated device are described in U.S. patent application Ser. No. 15 / 844,403, filed on Dec. 15, 2017, and entitled “OPTICAL COUPLER AND WAVEGUIDE SYSTEM,” which is incorporated herein by reference in its entirety.

[0264] Additional photoexcitable structures can be positioned between the sample well and the photodetector and configured to reduce or prevent the excitation light from reaching the photodetector, which could otherwise result in signal noise when detecting the emitted light. In some embodiments, a metal layer that can act as a circuit for an integrated device can also act as a spatial filter. Examples of suitable photoexcitable structures can include spectral filters, polarization filters, and spatial filters, and are described in U.S. Patent Application No. 16 / 042,968, entitled “OPTICALREJECTION PHOTONIC STRUCTURES,” filed on July 23, 2018, and U.S. Provisional Patent Application No. 63 / 124,655, entitled “INTEGRATED CIRCUIT WITH IMPROVED CHARGE TRANSFER EFFICIENCY ANDASSOCIATED TECHNIQUES,” filed on December 11, 2020, the entire contents of which are incorporated herein by reference.

[0265] Components located outside the integrated device can be used to position and align the excitation source to the integrated device. Such components can include optical components, including lenses, mirrors, prisms, windows, apertures, attenuators and / or optical fibers. Additional mechanical components may be included in the instrument to allow control of one or more alignment components. Such mechanical components may include actuators, stepper motors and / or knobs. Examples of suitable excitation sources and alignment mechanisms are described in U.S. patent application No. 15 / 161,088, entitled "PULSEDLASER AND SYSTEM", filed on May 20, 2016, the entire contents of which are incorporated herein by reference. Another example of a beam control module is described in U.S. patent application No. 15 / 842,720, entitled "COMPACT BEAMSHAPING AND STEERING ASSEMBLY", filed on December 14, 2017, which is incorporated herein by reference. Additional examples of suitable excitation sources are described in U.S. Patent Application No. 14 / 821,688, filed on August 7, 2015, entitled “INTEGRATED DEVICE FOR PROBING, DETECTING AND ANALYZING MOLECULES,” which is incorporated herein by reference in its entirety.

[0266] A photodetector positioned with a single pixel of an integrated device can be arranged and positioned to detect emitted light from a corresponding sample well of the pixel. Examples of suitable photodetectors are described in U.S. Patent Application No. 14 / 821,656, filed on August 7, 2015, entitled "INTEGRATED DEVICE FOR TEMPORAL BINNING OF RECEIVED PHOTONS," the entire contents of which are incorporated herein by reference. In some embodiments, a sample well and its corresponding photodetector can be aligned along a common axis. In this way, the photodetector can overlap a sample well within a pixel.

[0267] The characteristics of the detected emitted light can provide an indication for identifying a tag associated with the emitted light. Such characteristics can include any suitable type of characteristics, including the arrival time of the photons detected by the photodetector, the amount of photons accumulated by the photodetector over time, and / or the distribution of photons across two or more photodetectors. In some embodiments, such characteristics can be any one of luminescence lifetime, luminescence intensity, brightness, absorption spectrum, emission spectrum, luminescence quantum yield, wavelength (e.g., peak wavelength), and signal characteristics (e.g., pulse duration, inter-pulse duration, signal amplitude variation), or a combination of two or more.

[0268] In some embodiments, the photodetector may have a structure that allows detection of one or more timing features associated with the emitted light of the sample (e.g., luminescence lifetime). After the excitation light pulse propagates through the integrated device, the photodetector can detect the distribution of the arrival time of the photons, and the distribution of the arrival time can provide an indication of the timing features of the sample emitted light (e.g., a representative of the luminescence lifetime). In some embodiments, one or more photodetectors provide an indication of the probability (e.g., luminescence intensity) of the emitted light emitted by the tag. In some embodiments, the size and arrangement of multiple photodetectors can be set to capture the spatial distribution of the emitted light. The output signal from one or more photodetectors can then be used to distinguish the label from multiple labels, wherein the multiple labels can be used to identify the sample within the sample. In some embodiments, the sample can be excited by a variety of excitation energies, and the emitted light emitted by the sample in response to a variety of excitation energies and / or the timing features of the emitted light can distinguish the label from multiple labels.

[0269] In operation, parallel analysis of samples in sample wells is performed by exciting some or all of the samples in the wells with excitation light and detecting signals emitted from the samples with photodetectors. The emitted light from the samples can be detected by corresponding photodetectors and converted into at least one electrical signal. The electrical signal can be transmitted along a wire in a circuit of the integrated device, which can be connected to an instrument interfaced with the integrated device. The electrical signal can then be processed and / or analyzed. The processing or analysis of the electrical signal can be performed on a suitable computing device located on or outside the instrument.

[0270] The instrument may include a user interface for controlling the operation of the instrument and / or integrated device. The user interface may be configured to allow the user to input information into the instrument, such as commands and / or settings for controlling instrument functions. In some embodiments, the user interface may include buttons, switches, dials, and microphones for voice commands. The user interface may allow the user to receive feedback on the performance of the instrument and / or integrated device, such as proper alignment and / or information obtained by reading signals from a light detector on the integrated device. In some embodiments, the user interface may provide feedback using a speaker to provide auditory feedback. In some embodiments, the user interface may include an indicator light and / or a display screen for providing visual feedback to the user.

[0271] In some embodiments, the instrument may include a computer interface configured to be connected to a computing device. The computer interface may be a USB interface, a FireWire interface or any other suitable computer interface. The computing device may be any general purpose computer, such as a laptop computer or a desktop computer. In some embodiments, the computing device may be a server (e.g., a cloud-based server) accessible on a wireless network via a suitable computer interface. The computer interface may facilitate the information communication between the instrument and the computing device. The input information for controlling and / or arranging the instrument may be provided to the computing device and transmitted to the instrument via the computer interface. The output information generated by the instrument may be received by the computing device via the computer interface. The output information may include feedback on the data generated by the instrument performance, the integrated device performance and / or the readout signal from the photodetector.

[0272] In some embodiments, the instrument may include a processing device configured to analyze data received from one or more light detectors of the integrated device and / or transmit a control signal to an excitation source. In some embodiments, the processing device may include a general-purpose processor, a specially adapted processor (e.g., a central processing unit (CPU), such as one or more microprocessor or microcontroller cores, a field programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a custom integrated circuit, a digital signal processor (DSP), or a combination thereof). In some embodiments, the processing of data from one or more light detectors may be performed by both the processing device of the instrument and an external computing device. In other embodiments, the external computing device may be omitted, and the processing of data from one or more light detectors may be performed only by the processing device of the integrated device.

[0273] According to some embodiments, an instrument configured to analyze a sample based on a luminescent emission feature can detect differences in luminescent lifetimes and / or intensities between different luminescent molecules, and / or differences between lifetimes and / or intensities of the same luminescent molecule in different environments. The inventors have recognized and understood that differences in luminescent emission lifetimes can be used to distinguish the presence or absence of different luminescent molecules and / or to distinguish different environments or conditions to which luminescent molecules are subjected. In some cases, distinguishing luminescent molecules based on lifetimes (e.g., rather than emission wavelengths) can simplify aspects of the system. As an example, when distinguishing luminescent molecules based on lifetimes, wavelength-distinguishing optical devices (e.g., wavelength filters, dedicated detectors for each wavelength, dedicated pulsed light sources for different wavelengths, and / or diffractive optical devices) can be reduced in number or eliminated. In some cases, a single pulsed light source operating with a single characteristic wavelength can be used to excite different luminescent molecules that are emitted in the same wavelength region of the spectrum but have measurable different lifetimes. The analysis system that uses a single pulsed light source instead of multiple light sources operating at different wavelengths to excite and distinguish different luminescent molecules emitted within the same wavelength range has lower complexity in operation and maintenance, is more compact, and can be manufactured at a lower cost.

[0274] Although an analysis system based on luminescence lifetime analysis may have certain benefits, the amount of information obtained by the analysis system and / or the accuracy of detection may be increased by allowing additional detection techniques. For example, some embodiments of the system may be additionally configured to distinguish one or more characteristics of a sample based on luminescence wavelength and / or luminescence intensity. In some embodiments, luminescence intensity may be used additionally or alternatively to distinguish different luminescent labels. For example, some luminescent labels may emit at significantly different intensities or have significant differences in their excitation probabilities (e.g., at least about 35% difference), even though their decay rates may be similar. By referencing the binned signal to the measured excitation light, different luminescent labels can be distinguished based on intensity levels.

[0275] According to some embodiments, different luminescence lifetimes can be distinguished by a photodetector that is configured to time-bin the luminescence emission event after the luminescent tag is excited. Time binning can occur during a single charge accumulation period of the photodetector. The charge accumulation period is the interval between readout events during which photogenerated carriers are accumulated in the bin of the time-binned photodetector. Examples of time-binned photodetectors are described in U.S. patent application No. 14 / 821,656, entitled "INTEGRATED DEVICE FOR TEMPORAL BINNING OF RECEIVED PHOTONS," filed on August 7, 2015, which is incorporated herein by reference. In some embodiments, the time-binned photodetector can generate charge carriers in the photon absorption / carrier generation region and transfer the charge carriers directly to the charge carrier storage bin in the charge carrier storage region. In such an embodiment, the time-binned photodetector may not include a carrier travel / capture region. Such a time-binned photodetector may be referred to as a "direct binning pixel." Examples of time-binned photodetectors including direct binned pixels are described in U.S. patent application Ser. No. 15 / 852,571, filed on Dec. 22, 2017, and entitled “INTEGRATED PHOTODETECTOR WITH DIRECT BINNING PIXEL,” which is incorporated herein by reference.

[0276] In some embodiments, different numbers of fluorophores of the same type can be connected to different reagents in a sample, so that each reagent can be identified based on luminescence intensity. For example, two fluorophores can be connected to a recognition molecule of a first label, and four or more fluorophores can be connected to a recognition molecule of a second label. Due to different numbers of fluorophores, there may be different excitation and fluorophore emission probabilities associated with different recognition molecules. For example, during the signal accumulation interval, the recognition molecule of the second label may have more emission events, so the apparent intensity of the bin is significantly higher than the recognition molecule of the first label.

[0277] The inventors have recognized and appreciated that differentiating biological or chemical samples based on fluorophore decay rates and / or fluorophore intensities can simplify optical excitation and detection systems. For example, optical excitation can be performed with a single wavelength source (e.g., a source that produces one characteristic wavelength instead of multiple sources or a source that operates at multiple different characteristic wavelengths). In addition, wavelength recognition optics and filters may not be required in the detection system. In addition, each sample well can use a single photodetector to detect emissions from different fluorophores. The phrase "characteristic wavelength" or "wavelength" is used to refer to a central or dominant wavelength within a limited radiation bandwidth (e.g., a central or peak wavelength within a 20 nm bandwidth output by a pulsed light source). In some cases, "characteristic wavelength" or "wavelength" may be used to refer to a peak wavelength within the total bandwidth of the source radiation output.

[0278] According to one aspect of the present disclosure, an exemplary integrated device can be configured to perform single molecule analysis in conjunction with the above-described instrument. It should be understood that the exemplary integrated device described herein is intended to be illustrative, and other integrated device configurations can be configured to perform any or all of the techniques described herein.

[0279] Fig.29 A cross-sectional view of a pixel 1-112 of an integrated device 1-102 is shown. The pixel 1-112 includes a light detection region, which may be a pinned photodiode (PPD), and a charge storage region, which may be a storage diode (SDO). In some embodiments, the light detection region and the charge storage region may be formed in a semiconductor material of the pixel by doping regions of the semiconductor material. For example, the light detection region and the charge storage region may be formed using the same conductivity type (e.g., n-type doping or p-type doping).

[0280] During operation of the pixel 1-112, the excitation light may illuminate the sample aperture 1-108, causing incident photons (including fluorescent emissions from the sample) to flow along the optical axis toward the photodetection region PPD. Fig.29As shown, the pixel 1-112 may include a waveguide 1-220 that is configured to optically (e.g., briefly) couple excitation light from a grating coupler (not shown) of the integrated device to the sample hole 1-108. In response, the sample in the sample hole 1-108 may emit fluorescence to the light detection area PPD. In some embodiments, the pixel 1-112 may also include one or more photoexcitatory structures 1-230, which may include one or more optical rejection structures, such as spectral filters, polarization filters, and spatial filters. For example, the photoexcitatory structure 1-230 may be configured to reduce the amount of excitation light reaching the light detection area PPD and / or increase the amount of fluorescence emission reaching the light detection area PPD. Also shown in the pixel 1-112, the pixel 1-112 may include one or more metal layers 1-240, which may be configured as filters and / or may carry control signals from a control circuit configured to control a transfer gate, as further described herein.

[0281] In some embodiments, the pixel 1-112 may include one or more transfer gates that are configured to control the operation of the pixel 1-112 by applying an electrical bias to one or more semiconductor regions of the pixel 1-112 in response to one or more control signals. For example, when the transfer gate ST0 causes a first electrical bias in the semiconductor region between the light detection region PPD and the storage region SD0, a transfer path (e.g., a charge transfer channel) may be formed in the semiconductor region. Charge carriers (e.g., photoelectrons) generated by incident photons in the light detection region PPD may flow along the transfer path to the storage region SD0. In some embodiments, the first electrical bias may be applied during a collection period, during which charge carriers from the sample are selectively directed toward the storage region SD0. Alternatively, when the transfer gate ST0 provides a second electrical bias in the semiconductor region between the light detection region PPD and the storage region SD0, charge carriers from the light detection region PPD may be prevented from reaching the storage region SD0 along the transfer path. In some embodiments, the drain gate REJ can provide a channel for the drain D to absorb noise charge carriers generated by the excitation light in the light detection region PPD away from the light detection region PPD and the storage region SD0, such as during a rejection period before the fluorescence emission photons of the sample reach the light detection region PPD. In some embodiments, during a readout period, the transfer gate ST0 can provide a second electrical bias and the transfer gate TX0 can provide an electrical bias to allow the charge carriers stored in the storage region SD0 to flow to the readout region (which can be a floating diffusion (FD) region) for processing.

[0282] It should be understood that according to various embodiments, the transfer gate described herein may include semiconductor materials and / or metals and may include the gate of a field effect transistor (FET), the base of a bipolar junction transistor (BJT), and / or the like.

[0283] In some embodiments, the operation of the pixel 1-112 may include one or more collection sequences, each collection sequence including one or more rejection (e.g., drain) periods and one or more collection periods. In one example, a collection sequence performed in accordance with one or more pulses of an excitation light source may begin with a rejection period, for example, discarding charge carriers generated in the pixel 1-112 (e.g., in the photodetection region PD) in response to excitation photons from the light source. For example, the excitation photons may arrive at the pixel 1-112 before the fluorescent emission photons from the sample hole arrive. The transfer gate of the charge storage region may be biased to have a low conductivity in the charge transfer channel coupling the charge storage region with the photodetection region to prevent the transfer and accumulation of charge carriers in the charge storage region. The drain gate for the drain region may be biased to have a high conductivity in the drain channel between the photodetection region and the drain region to promote the discharge of charge carriers from the photodetection region to the drain region. The transfer gate of any charge storage region coupled to the photodetection region may be biased to have low conductivity between the photodetection region and the charge storage region so that charge carriers are not transferred to or accumulated in the charge storage region during the knockout period.

[0284] After the knockout period, there may be a collection period during which charge carriers generated in response to incident photons are transferred to one or more charge storage regions. During the collection period, the incident photons may include fluorescent emission photons, resulting in accumulation of fluorescent emission charge carriers in the charge storage regions. For example, a transfer gate of one of the charge storage regions may be biased to have a high conductivity between the photodetection region and the charge storage region to promote accumulation of charge carriers in the charge storage region. Any drain gate coupled to the photodetection region may be biased to have a low conductivity between the photodetection region and the drain region so as not to discard charge carriers during the collection period.

[0285] Some embodiments may include multiple knockout periods and / or collection periods in a collection sequence, such as a second knockout period and a second collection period following a first knockout period and collection period, wherein each pair of knockout periods and collection periods is performed in response to an excitation light pulse. In one example, charge carriers generated in a light detection region during each collection period of a collection sequence (e.g., in response to multiple excitation light pulses) may be accumulated in a single charge storage region. In some embodiments, the charge carriers accumulated in the charge storage region may be read out for processing prior to a next collection sequence. Alternatively or additionally, in some embodiments, charge carriers accumulated in a first charge storage region during a first collection sequence may be transferred to a second charge storage region that is sequentially coupled to the first charge storage region and read out simultaneously with the next collection sequence. In some embodiments, a processing circuit configured to read out charge carriers from one or more pixels may be configured to determine one or more of luminous intensity information, luminous lifetime information, luminous spectrum information, and / or any other mode of luminous information associated with performing the techniques described herein.

[0286] In some embodiments, the first collection sequence may include transferring charge carriers generated in the light detection reaction in response to the excitation pulse to a charge storage region at a first time after each excitation pulse, and the second collection sequence may include transferring charge carriers generated in the light detection reaction in response to the excitation pulse to the charge storage region at a second time after each excitation pulse. For example, the number of charge carriers collected after the first and second times may indicate brightness lifetime information of the received light.

[0287] As further described herein, the pixels of the integrated device can be controlled using one or more control signals from control circuitry of the integrated circuit, such as by providing control signals to a drain and / or transfer gate of the integrated circuit pixel, to perform one or more collection sequences. In some embodiments, charge carriers can be read out from the FD region of each pixel during a readout pixel associated with each pixel and / or a row or column of pixels for processing. In some embodiments, the FD region of the pixel can be read out using a correlated double sampling (CDS) technique.

[0288] V. Sequence Information

[0289] Table 1. Non-limiting example sequences of amino acid binding proteins.

[0290]

[0291]

[0292]

[0293]

[0294]

[0295]

[0296]

[0297]

[0298]

[0299]

[0300]

[0301]

[0302]

[0303]

[0304]

[0305]

[0306]

[0307]

[0308]

[0309]

[0310]

[0311]

[0312]

[0313]

[0314]

[0315]

[0316]

[0317]

[0318]

[0319]

[0320]

[0321]

[0322]

[0323]

[0324]

[0325]

[0326]

[0327]

[0328]

[0329]

[0330]

[0331]

[0332]

[0333]

[0334]

[0335]

[0336]

[0337]

[0338]

[0339]

[0340]

[0341]

[0342]

[0343]

[0344]

[0345]

[0346]

[0347]

[0348]

[0349]

[0350]

[0351]

[0352]

[0353]

[0354]

[0355]

[0356] Table 2. Non-limiting examples of tag sequences.

[0357]

[0358] Example

[0359] Example 1. Real-time dynamic single-molecule protein sequencing on an integrated semiconductor device

[0360] In this example, a new method is demonstrated, including a dynamic degradation sequencing method, in which a single surface-immobilized peptide molecule is detected in real time by a mixture of dye-labeled N-terminal amino acid recognition agents. The ability to annotate amino acids and collectively identify peptide sequences is demonstrated by measuring the fluorescence intensity, lifetime, and intermolecular dynamics of the recognition agents on the new semiconductor chip. Using binding kinetics, each recognition agent can uniquely identify multiple amino acids. The principles and processes for expanding the number of identifiable amino acids are also described herein. In addition, it is shown that this method is compatible with synthetic peptides and natural peptides isolated from recombinant human proteins, and can detect single amino acid changes and post-translational modifications. These results demonstrate a powerful core technology that can serve as an accurate, sensitive, and scalable next-generation protein sequencing platform.

[0361] Measurements of the proteome provide deep and valuable insights into biological processes. However, to fully understand the complex and dynamic state of the proteome in cells and the changes that occur in the proteome in disease states, and to make this information more accessible, more sensitive methods are needed. The complexity of the proteome and the chemical properties of proteins pose several fundamental challenges to achieving the same comprehensive sensitivity, throughput, and adoption as DNA sequencing technologies. These challenges include: the large number of different proteins (more than 10,000) and even more proteoforms in each cell; the very large dynamic range of protein abundance in cells and biological fluids and the lack of correlation with transcript levels; the high cost and high detection limits of current mass spectrometry methods. 2 ; and the inability to replicate or amplify proteins. Methods that directly sequence individual protein molecules offer the greatest possible detection sensitivity, with the potential to enable single-cell input, digital quantification based on read counts, detection of post-translational modifications (PTMs) and low-abundance or aberrant protein forms, and at cost and throughput levels conducive to widespread adoption.

[0362] This paper demonstrates a new single-molecule protein sequencing method and an integrated system for massively parallel proteomic studies. In this method, peptides are immobilized in nanoscale reaction chambers on a semiconductor chip, and the N-terminal amino acid (NAA) is detected in real time using a dye-labeled NAA recognition agent. Aminopeptidases sequentially remove individual NAAs, exposing subsequent amino acids for recognition without the need for complex chemistry and fluidics ( Figure 1 A benchtop device was fabricated with a 532 nm pulsed laser source for fluorescence excitation and electronics for signal processing ( Fig. 6A ). The semiconductor chip uses intensity and fluorescence lifetime rather than emission wavelength to distinguish dye tags. The recognition agent detects one or more types of NAA and provides information for peptide identification based on the temporal order of NAA recognition and on-off binding kinetics.

[0363] CMOS manufacturing technology is used to create custom time-domain sensitive semiconductor chips with nanosecond precision that contain fully integrated components for single-molecule detection, including light sensors, optical waveguide circuits, and reaction chambers for immobilizing biomolecules ( Figure 1 Brief illumination of the reaction chamber bottom via a nearby waveguide enables an observation volume of less than 5 attoliter, enabling sensitive single-molecule detection at high free-diffusing dye concentrations (greater than 1 μM).

[0364] The semiconductor chip uses a new filter-free system to exclude excitation light based on the arrival time of photons, achieving greater than 10,000-fold attenuation of the incident excitation light. No integrated optical filter layer is required, which improves the fluorescence collection efficiency and enables scalable manufacturing of the chip. In order to be able to distinguish fluorescent dye tags attached to NAA identifiers by fluorescence lifetime and intensity, the chip rapidly alternates between early and late signal collection windows associated with each laser pulse, thereby collecting different parts of the exponential fluorescence lifetime decay curve. The relative signal in these collection windows (called the "bin ratio") provides a reliable indicator of fluorescence lifetime ( Figure 6B -F, and ‘Materials and methods’).

[0365] In order for NAA binding proteins to function as recognizers in this approach, the average lifetime of the bound recognizer-peptide complex must be long enough (typically greater than 120 ms) to produce detectable single-molecule binding events. Proteins from the N-end rule aptamer family ClpS that naturally bind to N-terminal phenylalanine, tyrosine, and tryptophan were evaluated. Using PS610, a recognizer for ClpS2 from A. tumefaciens, it was determined that this recognizer was able to detectably bind to immobilized peptides with these NAAs. Importantly, it was also determined that the binding kinetics varied for each NAA. To demonstrate these properties, immobilized peptides containing the initial N-terminal sequence FAA, YAA, or WAA were incubated with PS610 on separate chips and data collected for 10 h (Methods). NAA recognition by PS610 was observed, characterized by continuous on-off binding during incubation, with a different pulse duration (PD) for each peptide ( Figure 2A ). The median PDs for FAA, YAA, and WAA were 2.51, 0.73, and 0.31 s, respectively. These values ​​reflect the differences in binding affinity resulting from the different dissociation rates of each type of protein-NAA interaction. 7 ( Fig. 7A -B).

[0366] To expand the set of recognizable NAAs, proteins of the N-end rule pathway were investigated as a source of additional recognizers. In a comprehensive screen of multiple ClpS family proteins, a new set of ClpS proteins from the bacterial phylum Planctomycetes were discovered that naturally bind to N-terminal leucine, isoleucine, and valine. Directed evolution was applied to generate a Planctomycetes ClpS variant—PS961—with submicromolar affinities for N-terminal leucine, isoleucine, and valine, and recognition of these NAAs was demonstrated ( Figure 2BThe median PDs for binding to peptides with N-terminal LAA, IAA, and VAA were 1.21, 0.28, and 0.21 s, respectively, consistent with bulk characterization ( Figure 7C ).

[0367] In another screen, a panel of different UBR-box domains from the UBR family of ubiquitin ligases were investigated, which naturally bind N-terminal arginine, lysine, and histidine. The UBR-box domain from the yeast K. lactis UBR1 protein had the highest affinity for N-terminal arginine and was used to generate the arginine recognition agent PS691. PS691 recognized arginine in peptides with an N-terminal RLA with a median PD of 0.23s( Figure 2C ). Binds with lower affinity to N-terminal lysine and histidine ( Fig.7D -E), which is insufficient for single-molecule detection.

[0368] To demonstrate that amino acids within a single peptide molecule can be sequentially exposed by an aminopeptidase and recognized in real time with distinguishable kinetics, an immobilized peptide containing the initial sequence FAAWAAYAA (SEQ ID NO: 1073) was incubated with PS610 for 15 min, followed by the addition of the aminopeptidase PhTET3 from P. horikoshii. The collected trace consisted of distinct pulse regions, referred to as recognition segments (RS), separated from regions lacking recognition pulses (non-recognition segments, NRS). Analysis software was developed to automatically identify pulse regions and transition points within the trace (Methods). The trace began with recognition of phenylalanine with a median PD of 2.36 s ( Figure 2D ), consistent with the PD observed in the FAA-only recognition assay. This pattern terminated after the addition of aminopeptidase (average 11 min after addition), followed by the sequential appearance of two RSs with median PDs of 0.25 s and 0.49 s ( Figure 2D ), which corresponds to the short and medium PDs obtained in the YAA and WAA recognition assays alone. Thus, upon introduction of aminopeptidase activity into the reaction, discrete RSs with the expected kinetics appeared sequentially in the correct order.

[0369] To demonstrate the dynamic sequencing of two NAA recognition agents, PS610 and PS961 were labeled with distinguishable dyes atto-Rho6G and Cy3, respectively, and an immobilized peptide with the sequence LAQFASIAAYASDDD (SEQ ID NO: 1035) was exposed to a solution containing the two recognition agents. After 15 min, two P. horikoshii aminopeptidases PhTET2 and PhTET3 with complementary activities were added, covering all 20 amino acids. The collected traces showed discrete pulse segments alternating between PS961 and PS610 according to the order of the recognizable amino acids in the peptide sequence ( Figure 2E The average bin ratios and average PD associated with each RS easily distinguish the two dye labels and the four types of identified NAAs ( Figure 2F The median PD of N-terminal LAQ, FAS, IAA and YAS were 2.70, 1.43, 0.25 and 0.66 s, respectively ( Figure 2G ).

[0370] Previous studies have shown that NAA-bound ClpS and UBR proteins also make contacts at residues 2 (P2) and 3 (P3) from the N-terminus, influencing binding affinity. These effects are reflected in modulation of PD depending on the downstream P2 and P3 residues, as observed above for LAA (1.21 s) relative to LAQ (2.70 s). These effects on PD were found to vary within a range of information advantages and can be determined empirically or approximated in silico to pre-establish models of peptide sequencing behavior ( Figure 7F A powerful feature of this recognition behavior in terms of peptide identification is that each RS contains information about potential downstream P2 and P3 residues or PTMs, regardless of whether these positions are targets of NAA recognizers.

[0371] To evaluate the kinetics of the dynamic sequencing approach when applied to different sequences, a synthetic peptide corresponding to a segment of human ubiquitin, DQQRLIFAG (SEQ ID NO: 1036), was first characterized ( Figure 3A -D). Sequencing reactions were performed using a combination of three differently labeled recognition agents, PS610, PS961 and PS691, and two aminopeptidases, PhTET2 and PhTET3 (Materials and Methods). Figure 3AThe example traces in begin with NRSs, corresponding to the time intervals at which residues from the initial DQQ motif appear at the N-terminus. The first RS begins at 120 min, after the N-terminal arginine is recognized by PS691. Subsequent cleavage events expose the N-terminal leucine, isoleucine, and phenylalanine to their corresponding recognition agents in sequence, with the transition from one RS to the next being rapid (less than 10 s on average). The transition from leucine to isoleucine recognition by PS961 is easily identified as a sharp change in the average PD. Because each peptide molecule follows the same reaction pathway over the course of the sequencing run, this overall pattern is replicated in multiple sequencing instances of the same peptide, with similar PD statistics for each trace ( Figure 3B Due to the random timing of cleavage events, each trace shows a different onset and duration of each RS ( Figure 3C ).

[0372] This approach reports the binding kinetics for each identifiable amino acid position as well as the kinetics of aminopeptidase cleavage along the peptide sequence. Since each RS typically contains tens to hundreds of on-off binding events, highly accurate binding kinetic information can be obtained from a single trace, resulting in a distribution of PD and interpulse duration (IPD) measurements that can be statistically analyzed. Repetitive probing of each NAA also provides accurate identifier calls because calls are not based on the error-prone detection of a single event associated with one fluorescent molecule ( Fig. 6F The recognition agent concentration determines the IPD of each RS; the higher the recognition agent concentration, the shorter the average IPD and the faster the pulse rate ( Fig. 8A -B). However, higher recognition agent concentrations increase the fluorescence background from freely diffusing recognition agents, resulting in a decrease in the pulse signal-to-noise ratio, and compete with aminopeptidases for N-terminal access. In practice, IPDs in the range of about 2 to 10 s strike a favorable balance between these factors.

[0373] The distribution of RS durations in a series of repeated traces defines the cleavage rate of each identifiable NAA. For the DQQRLIFAG (SEQ ID NO: 1036) peptide, the average cleavage times observed for the N-terminal arginine, leucine, isoleucine, and phenylalanine were 31, 54, 39, and 86 min, respectively, with approximately single exponential decay statistics for each position ( Figure 3D , Figure 8C The distribution of NRS durations reports the cleavage rate of one or more unidentified NAA runs. The average NRS duration of the initial DQQ motif was 153 min ( Figure 3D The average cleavage rate is a key parameter that is controlled by the concentration of aminopeptidase in the assay ( Fig.8D-E). Considering the exponential behavior, an average RS duration of 10 to 40 min was determined to provide sufficient time for pulse data collection, avoid missing RS due to rapid cleavage, and minimize excessive RS duration. It is found helpful to visualize the sequencing profile of the peptide as a kinetic profile, which is a simplified trace-like representation of the time course of the complete peptide sequencing, containing the median PD of each RS and the average duration of each RS and NRS ( Figure 3E ). These highly characteristic features provide a wealth of sequence-dependent information for mapping peptide traces to their source proteins.

[0374] To demonstrate that the core method and its kinetic principles are applicable to a wide range of peptide sequences, the synthetic peptides DQQIASSRLAASFAAQQYPDDD (SEQ ID NO: 1037), RLAFSALGAADDD (SEQ ID NO: 1038), and EFIAWLV (SEQ ID NO: 1039), a segment of human GLP-1, were sequenced under the same sequencing conditions as DQQRLIFAG (SEQ ID NO: 1036). Figure 3F Each peptide produces characteristic kinetics based on its sequence ( Figure 3G In the peptide DQQIASSRLAASFAAQQYPDDD (SEQ ID NO: 1037), the furthest accessible read was obtained for position 18 (the furthest identifiable amino acid), demonstrating that the method is compatible with long peptides and is able to obtain in-depth sequence information for peptides of length found in typical protein digests.

[0375] To illustrate how sequencing-derived kinetic parameters are sensitive to changes in sequence composition, a set of three peptides—RLAFAYPDDD (SEQ ID NO: 1040), RLIFAYPDDD (SEQ ID NO: 1041), and RLVFAYPDDD (SEQ ID NO: 1042)—were sequenced, differing only at a single position, directly downstream of the PS961 N-terminal target leucine. Each type of amino acid at this position had a different effect on the PD obtained during PS961 recognition of the N-terminal leucine. The median PD observed for LAF, LIF, and LVF were 1.29 s, 2.22 s, and 4.21 s ( Figure 4B In addition to differences in the PD for leucine, each peptide displayed a characteristic RS or NRS in the interval between leucine and phenylalanine recognition ( Figure 4A , Fig. 9A ). These results demonstrate the sensitivity of the sequencing readout to changes in a single position and suggest that both the directly identified NAA and neighboring residues can influence the overall kinetic signature obtained from sequencing.

[0376] Since the aminoacyl-proline bond of the YP motif in peptides such as RLIFAYPDDD (SEQ ID NO: 1041) cannot be cleaved by the PhTET aminopeptidase, observing the YP pulse at the end of the trace ensures that cleavage has been completed from the first identifiable amino acid to the last identifiable amino acid. Therefore, the sequencing output of RLIFAYPDDD (SEQ ID NO: 1041) provides a convenient data set for examining biochemical sources of non-ideal behavior that may lead to peptide identification errors. The main source of incomplete information in the trace is the deletion of the expected RS due to the random occurrence of cleavage events in rapid succession ( Fig. 9B ), and premature termination of the read due to photodamage or surface detachment ( Fig. 9C ).

[0377] In addition to changes in amino acid sequence composition, sequencing readouts are also sensitive to changes caused by PTMs. As an example, methionine oxidation was examined. The thioether moiety of the methionine side chain is easily oxidized during peptide synthesis and sequencing. The K of PS961 binding to peptides with an N-terminal methionine was determined. D 947nM( Fig.9D ), assuming that oxidation would generate a polar methionine sulfoxide side chain, thereby abolishing binding and reducing NAA binding affinity at P2. It was determined computationally that methionine sulfoxide is highly unfavorable in the PS961 NAA binding pocket, whereas nonpolar residues are more preferred at P2 ( Fig.9E ). The synthetic peptide RLMFAYPDDD (SEQ ID NO: 1043) was sequenced and two trace populations with different kinetic characteristics were observed - the first population contained leucine recognition with a median PD of 0.86 s; the second population had a median PD of 0.35 s ( Figure 4C The trace for the first population also shows methionine recognition, with a short PD in the time interval between leucine and phenylalanine recognition ( Figure 4E ). There is no methionine recognition in the trace of the second population ( Figure 4D ), indicating that the methionine side chains in these peptides cannot be recognized by PS961. When methionine was fully oxidized by preincubation with hydrogen peroxide (Materials and Methods), as expected, both the methionine recognition and leucine recognition clusters were eliminated, and the median PD was longer ( Figure 4E ). These results demonstrate the ability to perform extremely sensitive detection of PTMs due to their kinetic effects on recognition.

[0378] Proteomics applications require the identification of peptides in mixtures derived from biological sources. To extend the results to peptide mixtures and peptides from biological sources, two experiments were performed. First, the DQQRLIFAG (SEQ ID NO: 1036) and RLAFSALGAADDD (SEQ ID NO: 1038) peptides were mixed, immobilized on the same chip, and sequenced. Data analysis (Materials and Methods) identified two populations of traces corresponding to each peptide, with kinetic features that were very consistent with those identified in the individual peptide runs ( Figure 5A , Fig.9F ). Second, to demonstrate the applicability of the approach to biologically derived peptides, peptide libraries generated from recombinant human ubiquitin (76 amino acids) and GLP-1 (37 amino acids) proteins digested with AspN / LysC and trypsin, respectively, were sequenced using a simple workflow (Methods). For both libraries, data analysis readily identified traces that matched the expected recognition patterns of the protease cleavage products of ubiquitin and GLP-1, DQQRLIFAGK (SEQ ID NO: 1045) and EFIAWLVK (SEQ ID NO: 1046), and produced kinetic signatures consistent with synthetic versions of these peptides ( Figure 5B , Figure 9G Using simple sequence constraints provided by the kinetic information, matches to the kinetic signature of the ubiquitin peptide DQQRLIFAGK (SEQ ID NO: 1045) were identified throughout the human proteome (Materials and Methods). In addition to ubiquitin, only one protein was found to contain a peptide that could potentially match this signature; thus, even short signatures can be shown to have a low affinity for the peptide. 4 These results demonstrate that the full kinetic output of sequencing has the potential to enable digital mapping of peptides to their source proteins.

[0379] discuss

[0380] This simple, real-time, dynamic approach is distinctly different from other recently described single-molecule approaches that rely on complex iterative approaches involving stepwise Edman chemistry or hundreds of cycles of epitope probing; nanopore approaches have the potential for real-time readout and simplicity but face significant challenges related to the size and biophysical complexity of peptides. The sequencing technology described here is readily available to expand its capabilities, and there are multiple areas for improvement. Expanded proteome coverage could be achieved through directed evolution and engineering of recognition agents. The NAA targets demonstrated here represent approximately 35.6% of the human proteome, but lower affinity NAA targets require longer PDs for detection in all sequence contexts.

[0381] Recognizers for novel amino acids or PTMs can be evolved from current recognizers or identified in screens of other scaffolds, such as other types of NAA or PTM binding proteins or aptamers. In general, expansion to detect all 20 natural amino acids and multiple PTMs is feasible for de novo sequencing; however, partial sequences will suffice for most proteomics applications, which rely on mapping to a predetermined set of candidate proteins. Aminopeptidases can be engineered to optimize cleavage rates and minimize RS starvation caused by rapid succession of cleavages. It is envisioned that the dynamic range of samples and the applications best suited to the system will tend to scale with the number of reaction chambers on the chip, and some applications will require a compressed dynamic range.

[0382] It is expected that the sequencing technology demonstrated here will increase the accessibility of proteomic studies, enable new discoveries in biological and clinical research, and power a new generation of precision medicine.

[0383] Materials and Methods

[0384] Semiconductor device operation and bin ratio calculation

[0385] The experiments were performed on a pre-production semiconductor chip with 296K active wells, accounting for some losses due to flow cell blockage of the sensor array. The dual-chamber flow cell allows for parallel sequencing of two independent samples, each using 148K active wells. The initial production unit has 2M active wells, and the first product line is scalable to tens of millions of active wells using standard CMOS processing. Pulsed 532nm excitation light from a 67MHz mode-locked laser is coupled into grating couplers at the edge of the semiconductor chip. The use of a single laser wavelength, combined with the ability to differentiate fluorescent dyes by fluorescence intensity and lifetime, reduces size, cost, and complexity, aiding the scalability of the platform. A network of optical waveguides splits the excitation light and delivers it to the sensor array to illuminate each reaction chamber. Each CMOS pixel contains a single light-sensitive photodiode with two high-speed global shutters (rejection gate and collection gate) to discard and collect photoelectrons (the chip's photoexcitation structure reduces pixel-to-pixel crosstalk to less than 2%). Control waveforms are applied to the collection and reject gates (in sync with the incident pulse light source). Figure 6B). Approximately 1 ns before the excitation pulse, the reject gate is charged to greater than 3 volts and the collection gate is discharged to less than 1 volt. Scattered 532nm excitation photons generate photoelectrons in the photodiode. The photoelectrons are rapidly transferred to the high voltage drain via the built-in potential field within the photodiode and the reject gate potential. Between 1 and 3 ns after excitation, the collection gate is charged to greater than 3 volts and the reject gate is discharged to less than 1 volt. Photoelectrons generated by emission photons that reach the photodiode after the collection gate is opened are transferred to a storage node within each pixel. Photoelectrons within each pixel are accumulated for 7.5 to 30 ms (configurable) over approximately 500,000 to 2,000,000 laser pulses ( Figure 6B ). The accumulated charge in the storage node is measured via a standard transfer gate, floating diffusion, source follower, row select, and on-chip analog-to-digital converter common to all CMOS image sensors, enabling scaling to large array sizes with small pixels. Fluorescence lifetime information is obtained by alternating the timing of the collection gate and reject gate waveforms between subsequent measurements. In the first measurement, only emitted photoelectrons that arrive greater than 3ns after the excitation pulse (bin 0) are collected. In the second measurement, emitted photoelectrons that arrive greater than 1ns after the excitation pulse (bin 1) are collected. Throughout the excitation cycle, as the phase relationship between the excitation source and the gate waveform is adjusted, the signal measured from the pixel shows that the pixel transitions from 100% photon collection during the collection phase to greater than 99.99% photon extinction during the reject phase in less than 1ns ( Figure 6C ). The ratio of these two measurements (bin ratio) can provide an estimate of the fluorescence lifetime ( Fig.6D We have demonstrated the ability to distinguish multiple dyes based on bin ratios alone ( Fig. 6E ).

[0386] Peptide synthesis and labeling

[0387] Peptides were synthesized on Rink Amide Resin of PurePrep Chorus solid phase peptide synthesizer (Gyros ProteinTechnology) using standard Fmoc chemistry. All synthetic peptides contain C-terminal Fmoc-azidolysine. The resin was deprotected in a mixture of TFA / TIPS / H2O (2.5% / 2.5% / 95%) at room temperature for 1.5h. The deprotection mixture was concentrated under an argon stream. The peptide was precipitated from cold diethyl ether, resuspended in 1:1 water-acetonitrile, and purified on a reversed phase HPLC (X-bridge C18, Waters) with a 10-70% acetonitrile (0.05% TFA) gradient for 20min. The residue was dried under high vacuum to generate a white precipitate. At room temperature, a peptide stock solution (4 μL, 5mM) was added to a DBCO-DNA biotin solution (2nmol in 100 μL PBS). The progress of the reaction was monitored by LC-MS (Thermo UltiMate 3000 Executive Plus). After the reaction was completed, the mixture was conjugated with an excess of streptavidin. The peptide-DNA-streptavidin complex was purified on an ion exchange HPLC (DNAPac 200, Thermo). Gradient: buffer A, 20 mM sodium phosphate buffer, pH 8.5; buffer B, 1 M NaBr, 20 mM sodium phosphate buffer, pH 8.5; 20-60% B, 15 min. Before use, the purified complex was buffer exchanged on a 30K MWCO spin filter to a solution containing 50 mM MOPS (pH 8.0) and 60 mM potassium acetate. Peptides containing fully oxidized methionine were prepared by mixing 3% hydrogen peroxide with the methionine peptide in 1:1 water-methanol at room temperature for 20 min. The product was immediately purified on reverse phase HPLC using the same peptide purification method described above, the purity was verified by reverse phase HPLC (Thermo UltiMate 3000) on an analytical column (Zorbax SB-Aq, 5 μm, 4.6×250 mm), and the correct mass of the oxidation product was verified by LC-MS (Agilent LC-MSD-iQ, positive mode).

[0388] Protein digestion and labeling

[0389] GLP-1 7-37, GLP-2, and ubiquitin (1-76) recombinant proteins were purchased from RnD Systems as lyophilized powders. Each protein was reconstituted in 100 mM HEPES (20% acetonitrile) at pH 8.0 to a final concentration of 200 μM. Cysteine ​​was reduced and alkylated using TCEP (2 mM) and iodoacetamide (10 mM) when necessary. GLP1 and GLP2 were digested overnight at 37°C using 1 μg trypsin (LCMS grade, Pierce). Ubiquitin was digested using 1 μg LysC (LCMS grade, Pierce) and 1 μg rAspN (LCMS grade, Promega). After protease digestion, the pH of the peptide mixture was adjusted to 10.5 with potassium carbonate (57 mM), and lysine was converted to azidolysine using imidazole-1-sulfonyl azide (ISA, 2 mM) and copper sulfate catalyst (0.5 mM). The ISA was quenched using amine-functionalized polyurethane beads (Oligo Factory). The mixture was then filtered and the pH was adjusted to 7-8 with 1 M acetic acid. The solution was diluted in 50% (v / v) 10 mM MOPS, 10 mM KOAc, pH 7.5 and added to the DNA-streptavidin-DBCO complex and incubated at 37°C for 12-16 h. When necessary, the detergent cetrimide was added to the reaction at a final concentration of 0.25 mM.

[0390] Purification, labeling and characterization of recognition agents

[0391] The expression vectors (with pET30 a+ backbone) for the recognition agent and biotin ligase were co-transfected into BL21 (DE3) E. coli chemically competent cells. The transformed cells were plated on Luria agar plates containing carbenicillin (50 μg / mL) and kanamycin (25 μg / mL) and incubated overnight at 37°C to obtain single colonies. The starter culture inoculated with the colony was grown in Luria broth containing ampicillin (50 μg / mL) and kanamycin (25 μg / mL) and inoculated into large cultures at a starting optical density (OD600) of about 0.01. The expression culture was incubated at 37°C and 230 rpm until the OD600 was close to about 0.7. The culture was then induced with 4 mM IPTG. 8 mM biotin was added at the same time as IPTG to biotinylate the expressed recognition agent. After about 12 hours of expression, the cells were harvested by centrifugation at 4°C and 10,000g, and the cell pellet was washed with 1xPBS buffer at pH 7.4. The cells were resuspended in Bug buster HT (Thermo Fisher Scientific) and incubated on a magnetic stirrer for 30 minutes at room temperature. The cell suspension was then diluted with an equal volume of 2x lysis buffer (100mM Tris-HCl pH7.5, 10% glycerol, 0.5M NaCl) and incubated on a magnetic stirrer for 30 minutes at room temperature. The lysate was centrifuged at 4°C and 21,000g to remove cell debris. The supernatant was collected and loaded onto a nickel NTA resin (Cytiva) affinity column on an AKTAPure (Cytiva) system pre-equilibrated with buffer A (50mM Tris-HCl pH 7.5, 10% glycerol, 0.5M NaCl). The column was washed with at least 10 column volumes of buffer containing 10mM imidazole. Elution was performed with an imidazole gradient of 10-300 mM. The eluted fraction was dialyzed in a 10 kDa cassette against 4 L of dialysis buffer (50 mM Tris-HCl pH 7.5, 0.2 M NaCl, 50% glycerol) at 4°C overnight.

[0392] To label the recognition agent, equal volumes of recognition agent and DNA-dye-streptavidin complex were mixed at a molar ratio of 5:1 (recognizer: DNA-dye-streptavidin). The mixture was incubated on ice for 30 min and dialyzed overnight with SEC buffer (25 mM HEPES pH 8.0, 150 mM KCl). The recognition agent-dye conjugate was harvested from the dialysis and centrifuged at 4 ° C, 10,000 g. The supernatant was collected and concentrated using a 10 kDa cutoff concentrator. A size exclusion column (BioSEC-3 The conjugate was purified and concentrated using a 3 μm column.

[0393] Binding affinity was measured by using polarization of the labeled peptide. Polarization response and total intensity measurements were performed on a microplate fluorimeter at 20°C with excitation at 480 nm and emission at 530 nm. The recognition agent interacted with the labeled peptide containing the target N-terminal residue (XAKLDEESILKQK-FITC (SEQ ID NO: 1074)) in PBS buffer at pH 7.4 and readings were collected after 30 min. At a fixed concentration of the target peptide, multiple analyses were performed with increasing concentrations of the recognition agent to obtain titration curves. The equilibrium polarization response at each concentration was plotted and fitted to calculate K D .

[0394] The dissociation rates (k) of PS610 of various peptides were measured using a stopped-flow instrument. off ). The labeled peptide (50 nM) was mixed with PS610 in PBS buffer at pH 7.4 containing 0.01% Tween-20 and incubated at 30°C. After 30 min of incubation, the recognition agent:peptide complex was quickly mixed with a 10-20-fold molar excess of unlabeled trap peptide, and the reaction was tracked in real time by measuring the fluorescence intensity. The traces of at least three time courses were averaged and fitted to an exponential equation.

[0395] Purification of aminopeptidase

[0396] The expression vectors (pET30 a+ backbone) for aminopeptidases PhTET2 and PhTET3 were transformed into BL21 (DE3) E. coli chemically competent cells. The transformed cells were plated on Luria agar plates containing kanamycin (25 μg / mL) and incubated overnight at 37°C to obtain single colonies. The starter culture inoculated with the colonies was grown in Luria broth (LB) containing kanamycin (25 μg / mL) and inoculated into large cultures at a starting optical density (OD600) of about 0.01. The expression culture was incubated at 37°C and 230 rpm until the OD600 was close to about 0.7. The culture was then induced with 0.4 mM IPTG. The expressed aminopeptidases were purified as described above for the identification agent. During conditioning, the aminopeptidase protein was dialyzed against 50 mM MOPS pH 8.0 / 60 mM potassium acetate and then exposed to cobalt acetate at 65°C for 1-1.5 h at a final concentration of 400 μM to form an active dodecameric complex. The conditioned aminopeptidase preparation was further dialyzed against 50 mM MOPS pH 8.0 / 60 mM potassium acetate, aliquoted and snap frozen.

[0397] Peptide loading, identification and dynamic sequencing

[0398] The semiconductor chip was placed in the sequencing device and a chip check was performed to test the electronic circuit function and optimize the laser coupling alignment. The chip was then removed from the device socket, washed twice with 50 μL 70% isopropanol, and then washed four times with 30 μL wash buffer (50 mM MOPS pH 8.0, 60 mM potassium acetate, 50 mM glucose, 20 mM magnesium acetate and surfactant mixture) through a flow cell connected to the chip. A second chip check was then performed. The laser was then blocked by an integrated software-controlled shutter, the peptide complex was added to a final concentration of 1-10 nM and mixed thoroughly, and the chip was incubated for 15 min. The chip was then washed six times with wash buffer, and imaging solution (wash buffer containing 5 mM Trolox and oxygen scavenging system) was added. The laser was unblocked, and the occupancy percentage (target 10-30%, Poisson distribution) was recorded by obtaining the photobleaching signal of the fluorophore attached to the peptide complex during 5 min of laser irradiation. For the determination of recognition of NAA only, after peptide loading, the labeled recognition agent was added to a final concentration of 50nM PS610, 100nM PS691 or 250nM PS961 (depending on the experiment) and the data was recorded for 10h. For the dynamic sequencing determination, after peptide loading, a mixture of labeled recognition agents was added to obtain a final concentration of 50nM PS610, 100nM PS691 and 250nM PS961. The data was recorded for 15min. The laser was then briefly blocked, and aminopeptidase was added to the sequencing reaction through the flow cell and mixed thoroughly (the final concentration was 2-8μMPhTET2 and / or 20-80μM PhTET3, depending on the experiment). The laser was then unblocked and the data was recorded for 10h. For all runs, 30μL of mineral oil was added to the reservoir of each port of the flow cell to prevent evaporation during the run.

[0399] Signal processing and trace segmentation

[0400] The measured signal on the chip includes various noise components, the most important of which is the fluorescence emission generated by the diffused recognition agent in the reaction chamber. The pulse calling algorithm for a given reaction chamber first estimates the statistical properties of this background noise component. Once the estimate is determined within a certain error range, the algorithm observes newly generated data frames in an online manner. At each time point, the algorithm maintains state showing whether the signal is caused by the background component alone or by a pulse generated by the observed recognition agent-NAA interaction. The state transition from background to pulse is triggered by an edge detection test, in which the transition in the expected signal is significant relative to the statistical distribution of the background component. The state transition from pulse to background is triggered when a small window of the most recent frame of signal again matches the distribution of the background component. As new background frames are observed, the algorithm maintains the updated model of the background component. This effectively prevents drift in signal intensity, while also maintaining stable optical coupling of the laser to the chip based on any such drift detected through a feedback control loop. Since detected spikes could originate from both true recognition agent-dipeptide interaction events as well as other occasional transient noise spikes, a downstream filtering layer was employed to test the significance of spike events in the context of the entire run timeline and the entire reaction chamber dataset based on their duration, intensity, and noise pattern.

[0401] Initial regions were determined by performing a sliding window calculation of the pulse rate along the temporal dimension of a series of pulses. Regions with an average pulse rate greater than 1 pulse / minute were then subdivided according to a greedy bisection approach. In this paper, the Mann-Whitney U test was used to assess whether there was a statistically significant deviation between the pulses to the left and right of each potential split point in any of the four independent pulse characteristics: intensity, time bin ratio, pulse duration, and inter-pulse duration. When defining RS, the split point with the lowest p-value for any of the four characteristics was used to subdivide the region, and this process was continued until no candidate split point for the region maintained a p-value less than 10 in any comparison. -5 In this way, the transition from one RS to the next in the region of consecutive pulses can be predetermined based on the changes in fluorescence properties of the pulse dynamics. The resulting region is called the recognition segment (RS).

[0402] Identify segment classification

[0403] Reactions containing a single synthetic peptide were classified into RS using an unsupervised clustering algorithm. A Gaussian mixture model (GMM) was pre-trained using a subset of RS (including those whose constituent pulses had an average signal-to-noise ratio ≥ 3) to determine the approximate centroids for each of N recognition categories, where N equals the number of expected recognizable peptide states with F, Y, W, L, I, V, or R at the N-terminus. The identified clusters were assigned to recognizable peptide states by matching the predominant order of the observed cluster sequences to the expected amino acid sequence and using prior knowledge of the dye properties to identify the binders active during each RS. Multiple rounds of GMM fitting were then performed for all RS that matched the expected order of these events to refine the GMM model until no more sequences appeared in the expected order. The final model was then applied to all RS in a given reaction.

[0404] Reactions containing library preparation peptides and peptide mixtures were classified using a random forest classifier. Unless otherwise stated, plots and statistics generated from classified RS are from reaction chambers containing the expected RS sequence.

[0405] Molecular dynamics and binding energy calculations

[0406] A homology model of PS961 in complex with the peptide was generated using an in-house crystal structure, and mutations were applied and optimized using protCAD prior to molecular dynamics. AMBER20 implicit solvent molecular dynamics simulations using generalized Born solvation potentials were performed without cutting off interatomic distances using the ff19SB force field. Minimization was performed using the steepest descent method followed by conjugate gradient minimization. Langevin dynamics and 3ps -1 The system was thermally treated from 0 to 300 K with a collision frequency of . Molecular dynamics simulations of the equilibrated recognizer-peptide complex, free recognizer, and free peptide were independently run at 300 K for 5 ns, and binding energy calculations were performed using MMPBSA. A total of 125 frames were used for three simulations, each containing 10,000 2-fs steps. The binding energy and the decomposition of all residues contributing to the binding energy were calculated at 0.15 M salt concentration.

[0407] Example 2. Peptide identification using modeled proteome-wide dynamics

[0408] Sequencing and biochemical data were used to determine the predicted pulse duration of recognition agents binding to all possible tripeptide targets. Figures 10A-10C PS961 binding to the N-terminal position of leucine ( Fig. 10A ), isoleucine ( Fig. 10B ) or valine ( Fig. 10C ) heatmap of predicted pulse durations for the tripeptide targets. Figures 10D-10F PS610 binding to phenylalanine at the N-terminal position ( Fig. 10D ), tyrosine( Fig.10E ) or tryptophan ( Fig.10F ) heatmap of predicted pulse durations for the tripeptide targets. Figure 10G A heat map showing the predicted pulse durations for PS1122 binding to a tripeptide target with arginine at the N-terminal position. Fig. 10H The predicted pulse durations for PS961 (left) and PS610 (right) show a high correlation with the actual pulse durations from on-chip experiments. Fig. 10H , left) and PS610( Fig. 10H , right).

[0409] With this predicted tripeptide pulse duration database, the expected kinetic characteristics of each peptide in the human proteome can be modeled, thereby better understanding and utilizing the ability to identify proteins from sequencing results. As described in detail in Example 1 above, the kinetic characteristics are an average representation of the sequencing behavior of peptides on the chip. The kinetic characteristic information obtained from single molecule traces can significantly improve the ability to map sequencing data to the proteome (for example, compared to methods based on text string alignment in DNA sequencing). Kinetic information can include, for example, pulse duration, inter-pulse duration, and recognition segment (RS) duration.

[0410] Kinetic information can improve the mapping data of the proteome, because the recognition agent contacts (at least) two adjacent downstream residues when binding to the peptide, rather than just the N-terminal residue. In this way, they indirectly sense all 20 amino acids and encode this information in the average pulse duration (and possibly in the IPD and RS duration). In addition, adjacent visible residues in the peptide are represented by the directly adjacent RS on average (i.e., if there is at least one invisible amino acid between two RSs, there is only a consensus gap between them).

[0411] To build a model to demonstrate the ability to uniquely map peptides to the human proteome (using identifiers PS961, PS610, and PS1122), the proteome was digested in silico using AspN / LysC, followed by selection of all peptides ending in lysine (for on-chip immobilization) and greater than 7 amino acids in length. The results are shown below.

[0412] Human protein (SWISS-Prot): 20,595 proteins Peptides from AspN / LysC digestion: 1,148,192 Peptides ending in lysine: 652,225 Peptides greater than 7 amino acids 273,112

[0413] A predicted pulse duration was assigned to each visible amino acid in the set of 273,112 peptides (positions with predicted average PD less than 0.18 s were considered invisible). The distribution of predicted RS in the first 15 residues is shown in Fig.10I (left panel). 82,068 peptides contained 4 or more RSs (and were therefore considered potentially informative). Kinetic profiles were created for each of these peptides.

[0414] The kinetic signature contains the expected binders and average PD for each visible position, as well as gaps representing one or more invisible amino acid runs. Next, for each peptide, the number of peptides with the same kinetic signature was determined (signatures were considered identical if they had the same RS and gap order and the predicted PD for each RS was somewhat similar (the shorter PD was not less than half the longer PD in any pairwise comparison)). Based on this analysis, 38,849 of the 82,068 peptides produced unique kinetic signatures with no other matching peptides in the human proteome. An additional 10,571 peptides had only 1 other match. Fig.10I (Middle) shows the distribution of kinetic matches for each peptide. 14,167 proteins (69% of all proteins) have at least one uniquely matchable peptide. On average, each protein has 2.5 uniquely mappable peptides. Fig.10I (Right) The distribution of uniquely mappable peptides per protein is shown.

[0415] To further illustrate these data and how they can be used to model protein behavior, Fig.10J Results for the IL6 protein are shown (for simplicity, the residues immediately before the C-terminal lysine are considered invisible and the XP motif is considered cleavable). Fig.10J As shown in Figure 2, both peptides contain at least 4 RS. Figure 10K As shown, one of these peptides mapped uniquely to IL6 and the other peptide matched the kinetic signature of eight different peptides from eight proteins.

[0416] To provide an illustrative example using a smaller proteome, the human proteome analysis was performed on the E. coli proteome (containing only 4,392 proteins), as described above. The results are shown below.

[0417] E. coli protein: 4,392 Peptides from AspN / LysC digestion: 126,439 Peptides ending in lysine: 59,697 Peptides greater than 7 amino acids: 28,046 Peptides with 4+ visible RS in the first 15 residues: 9,925 Peptides with unique kinetic characteristics: 7,740(78%) Proteins with at least one peptide containing 4+ RS in the first 15 residues: 3,527 Proteins with at least one uniquely mappable peptide (2.4 peptides on average): 3,187 / 3527

[0418] Fig.10LThe distribution of predicted RSs in the first 15 residues is shown (left panel). 9,925 peptides contained 4 or more RSs (and were therefore considered potentially informative). A kinetic signature was created for each of these peptides. For each peptide, the number of peptides with the same kinetic signature was determined. Based on this analysis, 7,740 of the 9,925 peptides produced unique kinetic signatures with no other matches in the E. coli proteome. Fig.10L The distribution of kinetic matches for each peptide is shown (middle panel). 3,187 proteins contained at least one uniquely matchable peptide. On average, each protein had 2.4 uniquely mappable peptides. Fig.10L The distribution of uniquely mappable peptides per protein is shown (right). To illustrate this data and how it can be used to model protein behavior, Figure 10M Results are shown for an E. coli protein that contained 6 uniquely mappable peptides.

[0419] These results demonstrate the utility of a kinetic-centric view of peptide identification. This view also provides the ability to accurately model the informative effects of changes in reaction conditions, such as the addition of new recognition agents, increases in recognition agent pulse duration, changes in frame rate, and the addition of new dye labels.

[0420] Example 3: Selection of N-terminal alanine and valine binding variants by yeast display

[0421] The gene encoding PS557 is used as a template to carry out error-prone PCR, wherein multiple nucleotide changes are introduced to produce the PS557 protein of mutation. The protein library of mutation is transformed into yeast, used for yeast display and flow cytometry, wherein selection is carried out for the target peptide. In short, the protein can be marked with a tag (such as a myc tag), and the cells expressing the protein are identified with a fluorescently labeled antibody for the tag. The target peptide is biotinylated, and can be marked with a streptavidinized fluorophore to identify the yeast cells of the protein combined with the peptide. The mixture of yeast cells, peptides and fluorophores is incubated at room temperature for 1 hour, and then double-color FACS is performed to obtain double positive cells.

[0422] In this example, selections were performed using peptides with an N-terminal amino acid residue of either V or A. After 3 rounds of selection for each peptide, samples were sent for next generation sequencing and the results were used to rationally design additional libraries, perform another round of directed evolution or test individual proteins in biochemical assays. Fig.11A Results from three rounds of FACS selection are shown, with flow cytometry plots showing cells expressing protein (y-axis) versus cells binding AVP-peptide (x-axis). Fig. 11BResults from one round of error-prone PCR library pool selection are shown, with flow cytometry plots showing cells expressing protein (y-axis) versus cells binding 1 μM AV or AI peptide (x-axis). Figures 11A-11B In each flow cytometry plot shown, the signal in the upper right quadrant indicates more binding. Table 3 gives a selection of hits obtained from sequencing of the alanine and valine libraries.

[0423] Table 3. Hits obtained from sequencing of alanine and valine libraries.

[0424]

[0425]

[0426] Mutation N41D was selected as the top hit, which can be rationalized by computational modeling. When residue N41 was mutated to D in computer simulation, the Rosetta algorithm showed that its binding energy to the valine peptide (-1.8) was more favorable than that of wild-type PS557. This predicted difference in binding energy is about 1-2 hydrogen bonds and thus may result in a K of 1.5. D 5-10 times difference. Fig. 11C The putative binding of the alanine peptide to the mutant protein was elucidated. Fig. 11C The left panel shows a model of the entire PS557 protein with the peptide forming hydrogen bond interactions with the N41D mutation in the binding site. Fig. 11C The right image shows the hydrogen bond network with negatively charged (light shaded) and positively charged (correlated) surface regions superimposed.

[0427] Mutation V72M also appeared to be enriched by selection and could also be rationalized by computational modeling. Therefore, it was chosen to combine these two mutations together and express the double mutant in E. coli and purify it to test its binding activity. Similarly, other mutations were selected from the NGS dataset and combined by rational design to generate a set of potential hits for testing in high-throughput assays on the Octet platform.

[0428] In the high-throughput assay, the Octet sensor is coated with the target peptide and immersed in a buffer containing the purified protein. By comparing the response after 200 seconds, an approximate estimate of the relative binding of each protein can be obtained. Each protein has approximately the same molecular weight and is used at the same concentration, so ranking the proteins by their response level in this assay can give an approximate estimate of which proteins have improved binding. Studying the association and dissociation rates can also provide insight into the binding mechanism. Tables 4 and 5 give the response values ​​of different peptides to the constructs selected from this round of selection, showing the binding affinity of the selected candidates compared to the wild-type protein (PS557) in the high-throughput Octet assay. The response (in nm) after 200 seconds of incubation with the mutant protein candidate during the binding period of incubation is given for each peptide, which lists the two N-terminal residues of the peptide (Table 4: AA, VA, LA; Table 5: MA, IA, FA, WA, YA).

[0429] Table 4. Binding affinity of selected candidates to AA, VA and LA peptides.

[0430]

[0431]

[0432] Table 5. Binding affinity of selected candidates to MA, IA, FA, WA and YA peptides.

[0433]

[0434]

[0435] The results showed that some candidates, such as PS824, showed improved binding to valine and alanine and were therefore selected for further characterization. Using more quantitative measurements by fluorescence polarization, binding of the amino acid valine was improved 5-fold. The measured K of the wild-type protein D The PS824 mutants containing mutations I12F, N41D, Q55R, and V72M had a K of 1174 nM for peptides with an N-terminal valine. D The K of clone PS852 containing R31H, N41D, Q55R and V72M for valine peptide was 205 nM (Table 6). D was 142 nM as measured by fluorescence polarization. Fig.11D Example results of fluorescence polarization studies comparing the kinetics of binding of selected PS557 variants to an N-terminal alanine peptide are shown. Fig.11DThe results shown were obtained using a mixture containing PBS buffer (containing 0.01% Tween 20), peptide (AAKLDEESILKQ{LYS(FITC)} (SEQ ID NO: 1075) at a concentration of 100 nM, and protein at concentrations spanning 100-5000 nM.

[0436] Table 6. Binding affinity of selected candidates measured by fluorescence polarization.

[0437]

[0438] In parallel, the NGS dataset was also used to design a second generation library in which additional mutations were layered on top of the top ranked clones selected in the first round of directed evolution. Error-prone PCR and targeted mutation libraries were created and further rounds of selection were performed as described for the second round of directed evolution. Tables 7-8 summarize the top ranked hits and illustrate a snapshot of the best clones obtained from these approaches.

[0439] Table 7. Candidate amino acids for mutation in PS557.

[0440]

[0441]

[0442] Table 8. Candidate amino acid combinations for mutation in PS557.

[0443]

[0444]

[0445]

[0446]

[0447] Example 4: Display of improved N-terminal amino acid binding proteins using SNAP

[0448] The gene encoding PS557 is cloned into a vector for SNAP display such that the SNAP tag is fused to the N-terminus. During SNAP display, the SNAP protein tag reacts and covalently binds to a phenylguanine (BG) molecule that is added to the end of the DNA template encoding itself. This allows for a link between the phenotype and genotype of the protein, provided that the protein is expressed in a droplet emulsion using an in vitro transcription-translation system. The protein-DNA complex is exposed to the desired peptide bound to magnetic beads, unbound complexes are washed away, and high affinity binders are eluted from the beads.

[0449] In these studies, libraries of mutant PS557 genes were created using error-prone PCR and targeted mutagenesis, rationally designed based on computational modeling and previous results. Variants in the libraries were selected for peptide binding to alanine, valine, and methionine using SNAP display. Initial and selected libraries were sequenced separately, and NGS data were used to determine the enrichment of clones in different rounds of selection.

[0450] By comparing the frequency of a given protein sequence in the library before and after selection, clones can be ranked by potential affinity for the peptide. Most sequences exhibit low affinity and are screened out. The results show that defective clones (such as those with stop codons) are screened out, demonstrating the reliability of this method. Fig. 12A is a heat map showing the enrichment of mutations in the PS557 protein. The amino acids whose residues were changed are listed on the left, and the residue numbers are listed at the bottom. Each rectangle represents the enrichment of that mutation in the selected library compared to the initial library (dark is not enriched, light is enriched). The rectangle represents the wild-type residue at that position. For simplicity, a subset of mutations is shown, including stop codons, cysteine, or alanine residues.

[0451] like Fig. 12A As shown, stop codons are mostly not enriched at all positions in the sequence. However, for example, alanine mutations are enriched at many positions, represented by light-colored rectangles, and are more enriched than, for example, cysteine.

[0452] Four rounds of selection were performed on the mutant PS557 protein library. The most enriched sequences in this round of directed evolution against alanine peptides are given in Table 9, showing mutations in the PS557 sequence that were found to be enriched after four rounds of SNAP selection against N-terminal alanine peptides. Enrichment was calculated by dividing the percent abundance of clones in the NGS data from the fourth round of sequencing by the abundance in the initial library.

[0453] Table 9. Mutations enriched in the PS557 sequence.

[0454] mutation Enrichment N41D,Q55R,H60I,V72M 319.81 N41D,Q55R,E63S,L68M,V72M,Y100R 165.55 N41D,Q55R,P62Y,E63T,V72M,Y100R 158.03 N41D,Q55R,E63A,V72M,Y100R 154.26 N41D,Q55R,E63W,V72M,R106H 132.96 N41D,E63S,L68M,V72M,R106H 124.16

[0455] As shown in the results in Table 9, many of the enriched sequences contain similar mutations. The library for this round of directed evolution was designed based on the hits obtained in the first three rounds of directed evolution and selection, which can be explained by following the evolution of a specific clone, for example: N41D, Q55R, E63S, L68M, V72M, Y100R. Mutation N41D was first identified as enriched in yeast display after error-prone PCR library selection (Table 10, directed evolution round 1).

[0456] Table 10. PS557 mutations identified among hits enriched from selections of alanine-bearing peptides.

[0457]

[0458]

[0459] In the same selection, the combination of N41D and Q55R, i.e. double mutant, was also identified. Meanwhile, V72M was selected as enrichment. However, clones containing these mutation combinations, such as triple mutants N41D, Q55R, V72M, were not selected, which was considered to be due to the lack of all possible triple mutation combinations in the initial library. Many mutations (such as V72M) have convincing calculation data to rationalize their importance. All these data were analyzed, and a second library was created, wherein the combination of combination N41D, Q55R, V72M and many other rational designs were included in the library on purpose. In the next round of selection using yeast display, this mutant was selected as one of the highest clones in enrichment (table 10, directed evolution the 2nd round).

[0460] At the same time, a round of directed evolution was performed using SNAP display to design libraries with targeted mutations based on computational design and some hits that have been seen in previous rounds of directed evolution. Each position is allowed to mutate into different combinations of all 20 amino acids. Some positions similar to those in other previous selections were identified in this selection, and in some cases, residues mutated into amino acids different from those previously. For example, E63 was identified as mutated into lysine (K) in the first round of directed evolution, but mutated into alanine (A) or serine (S) in the third round (Table 10, directed evolution round 3). In the fourth round of directed evolution, a library was created, in which many of these combinations were tested, including N41D, Q55R, E63S, L68M, V72M, and other residues were randomly mutated into all 20 amino acids (such as Y100). From this fourth round of directed evolution, the enriched clones listed in Table 9 were identified, and clones such as N41D, Q55R, E63S, L68M, V72M, Y100R were selected, expressed in E. coli, purified, and tested for their binding activity in a high-throughput assay on the Octet platform.

[0461] Fig. 12B Candidates selected from the fourth round of directed evolution using SNAP display selection are shown, compared to other previously selected candidates. In a high-throughput assay, the Octet sensor is coated with the target peptide and immersed in a buffer containing the purified protein. Fig. 12BThe trace of alanine peptide is shown. The response increase of wavelength shift (in nm) over time (binding curve between 0 and 200 seconds, dissociation curve between 200 and 500 seconds) is used to illustrate the improvement in binding. By comparing the response of each protein after 200 seconds, an approximate estimate of relative binding can be obtained. The molecular weight of each protein is roughly the same and used at the same concentration. In this assay, ranking proteins by response level can provide an approximate estimate of which proteins have improved binding. Studying the binding and dissociation rates can also provide a deep understanding of the binding mechanism. Table 11 gives the response values ​​of the constructs selected from this round of selection and different peptides, which shows the response (in nm) of each peptide and the mutant protein candidate after 200 seconds of incubation in the binding phase of incubation, and each peptide has two N-terminal residues (AA, VA, LA, etc.) of the listed peptides.

[0462] Table 11. Binding affinity of selected candidates compared to wild-type protein (PS557) in the high-throughput Octet assay.

[0463]

[0464] Example 5: Development of arginine recognition agent PS1122

[0465] This example describes the development of PS1122, an engineered variant of the UBR protein (PS621) from Kluyveromyces marxianus that has higher affinity for arginine and histidine, showing improved arginine recognition on the chip. Based on analysis of binding kinetics and on-chip results, PS1122 has a binding affinity for N-terminal arginine that is approximately 7 times higher than PS621, resulting in a favorable increase in pulse duration and faster binding speed for RX dipeptides. These properties combine to improve the recognition range of arginine tripeptides and the accuracy of ROI detection. It is estimated that PS1122 can clearly detect approximately 52% of the total arginine positions in the human proteome, accounting for 2.9% of the total proteome (an increase compared to approximately 1.4% for PS691).

[0466] Directed evolution method

[0467] PS621 (and its tandem version PS691) binds to arginine (R), histidine (H), and lysine (K). Due to the influence of downstream residues on pulse duration, observable binding of PS621 to R on the chip is limited to approximately 25% of arginine positions in the proteome. PS621 variants with stronger arginine binding were selected using directed evolution. Multiple types of variant libraries were subjected to multiple rounds of selection and mutational evolution to obtain a set of candidate recognizer variants (PS1101-PS1122), which were used for biochemical and single-molecule studies.

[0468] Octet analysis

[0469] Variant binders (PS1101-PS1122) and controls were expressed in E. coli, purified by a high-throughput workflow, and evaluated for binding to the N-terminal amino acid on the Octet platform. The peptides used in the assay mostly contained the penultimate alanine and consisted of the sequence XAKLDEESILKQK (SEQ ID NO: 1074). A peptide of the sequence RXKLDEESILKQK (SEQ ID NO: 1076) was also used to evaluate the impact of the penultimate residue. Table 12 summarizes a set of Octet response measurements for RX (various R dipeptides), HA, and KA.

[0470] Table 12. Octet responses of PS621 variant binders measured with RX (various R dipeptides), HA and KA peptides at 30°C.

[0471] Conjugate RA RL HA KA RE RQ RS RR RF PSll01 6.5 7.7 0.1 1.4 1.9 4.7 4.3 2.8 6.0 PSll02 8.6 10.4 2.5 6.6 6.6 9.3 8.5 7.0 9.3 PSll03 6.2 8.4 0.1 2.9 PSll04 9.4 11.3 3.4 7.3 7.6 9.7 8.9 8.0 10.5 PSll05 6.5 7.5 0.2 2.6 2.9 5.9 5.5 4.4 7.2 PSll06 7.6 9.0 0.6 4.7 PSll07 7.8 8.1 0.6 3.9 PSll08 8.2 10.3 1.2 5.0 PSll09 8.2 10.5 0.6 4.7 P1110 7.0 9.8 2.4 6.6 P.S.llll 7.9 10.0 0.1 2.7 PS1112 8.6 9.9 0.1 23 PS1113 7.7 9.3 0.2 3.8 PS1114 9.2 9.7 1.2 5.3 PS1115 10.8 11.8 3.0 6.5 7.6 9.1 8.5 6.7 10.0 PS1116 7.1 8.5 0.1 1.2 P S1117 8.6 9.0 0.5 3.5 Ps1118 9.8 10.6 1.8 5.0 P S1119 8.6 9.6 0.7 3.3 PSll20 9.9 11.7 1.7 5.5 8.8 9.5 8.3 5.9 10.2 PSll21 9.5 10.6 2.2 5.4 PSll22 9.9 11.2 1.3 4.6 PS621 9.3 10.2 2.0 5.2 5.8 7.4 7.1 6.3 8.4

[0472] Measuring binding affinity by polarization

[0473] Fluorescence polarization assays were performed with all candidates and single-site binding responses were measured at fixed binder concentrations ( Fig.13A The assay measures the strength of the interaction between the binder and a labeled peptide (XAKLDEESILKQK-FITC (SEQ ID NO: 1074). Based on the binding response, we further investigated selected candidates by measuring their Kd. Multiple polarization binding titrations were performed with increasing concentrations of binder protein and the Kd was determined from the titration curves ( Fig. 13B ).

[0474] PS1122, PS1115, PS1106, PS1114, PS1121, and PS1104 showed the greatest improvement in RA peptide binding. Several variants also showed improvement in HA binding over PS621. Fig. 13B The RA binding affinity determination titration curves of PS621, PS691 and PS1122 are shown. Compared with PS621 and PS691, PS1122 has a 5- to 7-fold increase in binding affinity for RA peptides and an approximately 2-fold increase in binding affinity for HA peptides.

[0475] k on and k off Stopped-flow rapid kinetic analysis

[0476] The binding rate constants (k) of PS621 and variants for RA and HA peptides were measured using a stopped-flow assay. on ) and dissociation rate (koff ) (The results are summarized in Table 13). These measurements were performed to predict the relative improvement of pulse duration and inter-pulse duration on the chip.

[0477] Table 13. k values ​​of PS621 variants obtained by stopped-flow instrumentation and fast kinetics. on Rate constant and k off Rate (dashes indicate not measured).

[0478]

[0479] In these assays, the RX peptide provided better signals and more accurate measurements than the HA peptide because the RX peptide bound more tightly. on The rate constants were comparable to those of the tandem conjugate PS691. The k values ​​of PS1122 for RA and RL peptides off The rate was about 3-3.5 times slower than PS621 or PS691, which predicts longer pulse durations on the chip. The dissociation of PS1122 from the HA peptide was also about 2 times slower than PS691. These measurements identified PS1122 as a variant to evaluate in the single molecule assay.

[0480] In parallel, a next-generation on-chip recognition agent screening approach was employed to evaluate multiple PS621 variants on-chip. PS621 variants were purified in biotinylated form, micro-labeled using a modified protocol of streptavidin-tetraCy3B, and recognition runs were performed on selected candidates. Improvements in coverage of PS1122 for R recognition were evaluated on-chip using various Rx penultimate peptides.

[0481] Identification of RA dipeptide on chip

[0482] For in-depth on-chip characterization, PS1122 was purified on a large scale and labeled with a streptavidin-tetraCy3 dye complex. Pooled and on-chip screening assays showed that PS1122 binding to arginine was improved enough to recognize the RA dipeptide on-chip. Recognition run analysis using the QP304-RAIFAG peptide confirmed the presence of visible binding of PS1122 to N-terminal RA with a longer pulse duration (0.29 s) than PS691 (0.16 s). Fig. 13C ).

[0483] Sequencing performance and arginine tripeptide coverage of PS1122

[0484] The sequencing performance of PS1122 and PS691 was compared using QP433 (RLIFAYP (SEQ ID NO: 1087) and other peptides. Multiple multiplexed dynamic assays were performed using PS1122 to further evaluate its arginine recognition range for the RXA tripeptide ( Fig.13D ).

[0485] Using the arginine tripeptide pulse duration data determined from multiplex runs of PS1122, we determined the predicted pulse durations for all 400 RXX tripeptides and estimated the arginine proteome coverage of PS1122 ( Figure 10G ). These results predicted coverage of 52% of arginine positions, representing 2.9% of the human proteome.

[0486] Example 6: PS961 engineered peptide binding enhancement

[0487] PS961 was based on the PS557 precursor with six point mutations (N41D, Q55R, E63S, L68M, V72M, Y100R) and outperformed PS557 in binding affinity or on-chip performance to the L / I / V / A / M / P N-terminal peptides. Each point mutation was analyzed in the context of the protein in complex with an alanine tripeptide, and a structure-based rationale was provided for the selection of these substitutions relative to the native PS557 amino acid identity in the directed evolution screen.

[0488] Direct enhancement of the binding pocket

[0489] Substituting asparagine for aspartic acid at position 41 of the external loop of the binding pocket enhanced both the long-range and short-range interactions of PS961 with the N-terminal peptide ligand ( Fig.14A ).

[0490] Long-range electrostatic interactions increase the probability of the positively charged peptide N-terminus to interact with the protein, and the additional negative charge of the aspartic acid side chain compared to the asparagine side chain enhances electrostatic steering into the binding site, as shown by the electrostatic surface charge distribution ( Fig. 14B ).

[0491] This amino acid position is also directly located in a triplet of residues that is essential for binding, as these form short-range electrostatic interactions with the N-terminal amino group of the peptide. The aspartic acid at position 41 has two negatively charged atoms on either side of the side chain and can interact more favorably with the peptide by reducing the conformational entropy of significant electrostatic interactions in the binding pocket. In addition, the aspartate side chain is expected to carry a stronger polarization of orbital electronegativity, which should allow for stronger hydrogen bonding between PS961 and the N-terminus of the peptide compared to the partially charged asparagine side chain. The Rosetta all-atom energy function confirms the presence of this stronger interaction by quantitatively identifying lower and more favorable Coulomb energies (the “fa_elec” term in the scoring function) involving the N-terminal amino acid in PS961, relative to the presence of a triplet carrying an asparagine ( Fig. 14C ).

[0492] The substitution of valine to methionine at position 72 introduced a longer nonpolar side chain into the PS961 binding pocket ( Fig.15 ). These additional nonpolar atoms can further insert into the cavity, thereby reducing the pocket volume and increasing hydrophobic interactions with the smaller N-terminal amino acid.

[0493] Optimization of non-pocket hydrophobic stacking

[0494] Since the nonpolar side chain of methionine is longer than that of leucine, this mutation at position 68 can interact favorably with the amino acids on the adjacent β-sheet, so that the structural cavity is often filled in the simulations ( Fig.16 The tightly packed nonpolar amino acid side chains inside the protein core reduce the unfavorable conformational entropy of the internal cavity and increase favorable hydrophobic interactions, which together increase the stability of the overall fold and may reduce the energetic cost of forming the binding pocket before binding.

[0495] Reduction in potential alternative binding sites and increase in surface and net charge

[0496] Since the N-terminus of the peptide ligand carries a permanent positive charge, any negatively charged pocket naturally present on the PS557 surface may interfere with the probability of effective protein-peptide interaction in the binding site by acting as an alternative low-affinity competing binding site. The mutation Y100R mitigates potential off-pocket interactions by reducing the negative surface charge that is present in the absence of the arginine mutation. Fig.17 ). In addition, by binding to Q55R and E63S, the overall surface and net charge of the protein increased, more directly favoring the expected binding events for correct on-chip identification.

[0497] The probability of a loop conformation that positively interacts with the peptide increases

[0498] Molecular dynamics simulations of PS961 and PS557 binding to the AAA-tripeptide showed that the arginine side chain at position 100 has an enhanced potential to form hydrogen bonds with adjacent ring residues compared to the tyrosine in PS557. This arginine is often involved in a complex hydrogen bond network involving R106 and the backbone carbonyl of the penultimate residue of the peptide, providing additional stability to the bound form of the peptide ( Fig.18A In simulations of PS961 binding to the AAA-tripeptide, the average occupancy of the R100:R106 hydrogen bond (calculated as the percentage of 1000 simulation frames that achieve a specific interaction) was approximately six times higher than that of PS557 ( Fig.18B ). Accordingly, the third-to-last hydrogen bond of R106 is more likely to occur in PS961 than in PS557.

[0499] Figures 19A-19C The secondary structure, sequence and binding pocket properties of PS961 are shown. Fig.19A The classification of the secondary structural groups of the proteins is shown. Fig.19B A Poisson–Boltzmann electrostatic potential surface plot of the binding pocket is shown, the residues forming the binding pocket are labeled, and the corresponding pocket properties are listed. Fig.19C The sequences of the native parent protein and the engineered variants are shown, highlighting the mutations, pocket locations, and secondary structure assignments for each position.

[0500] PS961 Crystallography

[0501] The crystal structure of PS961 in complex with a target peptide with an N-terminal methionine (Met-Ala-Lys-Leu (MAKL) (SEQ ID NO: 1047) was solved. The protein:peptide complex was generated and purified, and diffracting crystals of PS961:MAKL were obtained ( Fig.19D ). A complete X-ray data set was collected from these crystals, and the final structure was Resolution analysis and refinement ( Fig.19E ). Figures 19F-19K It shows how PS961 binds to the target peptide MAKL (SEQ ID NO: 1047).

[0502] Fig.19F The Met in the target peptide is shown to contact Asp10 and Asp11 in the recognizer PS961. The crystal structure also shows interactions with two water molecules (spheres), one of which is held in the proper orientation by Asp42 of the recognizer. Substitution of Asp42 may alter the binding affinity of this N-terminal residue in the peptide. Fig.19F In addition to the ionic interactions shown in , Met1 in the peptide also makes nonpolar contacts with at least six residues in PS961 ( Figure 19G, marked as residues other than "MET-1"). These residues in PS961 are located around the side chain of Met. By appropriately replacing some of these residues, PS961 may recognize and bind to other N-terminal amino acids different from Met.

[0503] The next residue in the peptide, Ala2, contacts the backbone of His14 in PS961 and also forms hydrogen bonds with water molecules ( Fig.19H , spheres), which in turn contacts Tyr16 and another water molecule. Therefore, replacing Tyr16 could affect the presence of said water molecule, thereby altering the binding of the penultimate residue in the peptide. The crystal structure also shows that the oily side chain of Ala2 in the peptide (with three methyl protons) is located between residues Thr15 and Leu73 in the recognition agent ( Fig.19I ), which opens the possibility of replacing these residues in the recognition agent with smaller side chains to make room for the larger side chains in the peptide. Lys3 in the peptide forms a salt bridge with Asp42 in the recognition agent ( Fig.19J , dashed lines indicate ionic interactions). Lys3 in the peptide is mostly removed from the binding site; however, Tyr16 in the recognizer is in close proximity, as shown on their surfaces ( Figure 19K ).

[0504] The crystal structure of PS961 in complex with a target peptide with an N-terminal alanine (Ala-Ala-Lys-Leu (AAKL) (SEQ ID NO: 1048)) was solved. The protein:peptide complex was generated and purified, and diffracting crystals of PS961:AAKL were obtained. A complete X-ray data set was collected from these crystals, and the final structure was presented in Resolution analysis and refinement. Figure 19L A panoramic view of the PS961:AAKL complex structure is shown, with PS961 shown in a cartoon representation, AAKL (SEQ ID NO: 1048) peptide shown in stick form, water molecules shown in sphere form, and PEG molecules shown in a ball-and-stick representation.

[0505] Due to the high resolution, Figure 19L The PS961:AAKL crystal structure shown shows a number of water molecules that can be modeled. It also shows alternative conformations for three residues in the recognition agent. The overall fold of the recognition agent in the AAKL (SEQ ID NO: 1048) complex is identical to that observed in the PS961:MAKL complex. Thus, the backbones of the recognition agents in the two structures can be superimposed with an average deviation of only Fig.19M Backbone superposition of PS961 bound to Met and Ala peptides is shown.

[0506] Based on the comparison of different PS961 complexes, some side chain changes in PS961 are obvious. Since the AAKL (SEQ ID NO: 1048) and MAKL (SEQ ID NO: 1047) peptides differ only in the first residue, the binding pocket of these residues was evaluated. Since the side chain of Ala1 in the AAKL (SEQ ID NO: 1048) peptide is smaller than Met1 in MAKL (SEQ ID NO: 1047), the recognition agent rearranged some of its residues to make room for Met ( Fig.19N When bound to AAKL (SEQ ID NO: 1048), the side chain of Asp42 points toward the binding pocket, but when bound to the MAKL (SEQ ID NO: 1047) peptide, the same side chain is displaced. Fig.19N The displacement of the Asp42 side chain in the recognition agent when comparing binding to AAKL (SEQ ID NO: 1048) and MAKL (SEQ ID NO: 1047) is shown. The arrow indicates how the Asp42 side chain moves to engage with Met1 of the MAKL peptide.

[0507] In general, both peptides (AAKL (SEQ ID NO: 1048) and MAKL (SEQ ID NO: 1047) have the same orientation and, as expected, the terminal amino group (NH2) in both peptides adopts approximately the same configuration. Fig.19O A comparison of AAKL (SEQ ID NO: 1048) (bottom stick) and MAKL (SEQ ID NO: 1047) (top stick) peptides when bound to PS961 is shown. The backbones of both peptides follow the same trajectory, but overall displacement from each other is observed. It is noteworthy that the main difference between the peptides is the configuration of the Lys3 side chain. In addition, the fourth residue of the peptide (Leu4) could not be observed, which is also the case in the PS961:MAKL complex. Most binding interactions are concentrated at the first residue of the peptide, partially at the second residue, and almost not at the third position, suggesting that some downstream residues are more free and disordered.

[0508] The binding of the N-terminal residues in each structure was evaluated. The terminal amino group ...

Claims

1. A recombinant or synthetic amino acid binding protein having an amino acid sequence at least 80% identical to SEQ ID NO: 1, wherein the amino acid sequence comprises an amino acid substitution at one or more positions corresponding to E22, R31, L39, N41, D42, D43, D44, H45, T46, Y47, V50, Q55, P62, E63, L68, A69, V72, D73, Q75, Y100, and M111 of SEQ ID NO:

1.

2. The amino acid binding protein of claim 1, wherein the amino acid sequence comprises an amino acid substitution at a position corresponding to N41 and at one or more positions corresponding to E22, R31, L39, D42, H45, V50, Q55, P62, E63, L68, V72, Q75, Y100, and M111.

3. The amino acid binding protein according to claim 1 or 2, wherein the amino acid sequence comprises an amino acid substitution at the position corresponding to N41 and at one or more positions corresponding to Q55, E63, L68, V72 and Y100.

4. The amino acid binding protein of any one of claims 1-3, wherein the amino acid substitutions are selected from E22V, R31H, L39M, N41D, D42L / P, H45C / F, V50A / F / Y, Q55H / R, P62R, E63A / G / K / S, L68M, V72M, Q75L, Y100R, and M111A / S.

5. The amino acid binding protein according to any one of claims 1 to 4, wherein the amino acid substitution is selected from N41D, Q55R, E63S, L68M, V72M and Y100R.

6. The amino acid binding protein of any one of claims 1-5, wherein the amino acid sequence is at least 85%, at least 90%, at least 95%, at least 98%, 80-98%, 80-95%, 80-90%, 85-95% or 90-98% identical to SEQ ID NO:

1.

7. The amino acid binding protein according to any one of claims 1 to 6, wherein the amino acid sequence is selected from PS635-645, PS731-732, PS759-766, PS769, PS795-870, PS896-912, PS918-1043, PS1048-1100, PS1124-1137, PS1141-1161, PS1175-1199, PS1203-1217, PS1222-1245, PS1277-1305, PS1321-1350 and PS1425-1448 (SEQ ID NO:22-27, 87-88, 115-122, 125, 151-226, 249-265, 271-390, 395-446, 470-483, 487-507, 521-545, 549-563, 568-591, 622-650, 664-693 and 768-791).

8. A recombinant or synthetic amino acid binding protein comprising the structure of formula (I) or a structural equivalent thereof: β1–α1–α2–β2–α3–β3 (I), in: Each of β1, β2, and β3 is a β strand; Each of α1, α2, and α3 is an α-helix; Every instance of "–" is a ring; and At least a portion of each of α1, α2, the loop between β1 and α1, and the loop between α3 and α3 forms a binding pocket for an amino acid ligand, wherein the binding pocket comprises one or more of the following: i) approx. The volume of ii)-3.0RTe c -1 or lower electrostatic potential, iii) at least 35% of the amino acids forming the binding pocket have negatively charged side chains, iv) a plurality of hydrogen bond acceptors arranged to form one or more hydrogen bonds in the presence of the amino acid ligands, and v) A plurality of van der Waals contact positions, which are configured to form van der Waals interactions in the presence of amino acid ligands.

9. The amino acid binding protein according to claim 8, wherein the amino acid ligand is a polypeptide comprising at least three amino acids.

10. The amino acid binding protein according to claim 8 or 9, wherein the amino acid ligand comprises the N-terminal amino acid of the polypeptide.

11. The amino acid binding protein according to claim 10, wherein the N-terminal amino acid is selected from the group consisting of leucine, isoleucine, valine, methionine and alanine.

12. The amino acid binding protein of claim 11, wherein the amino acid binding protein has a dissociation constant (K) of less than 2,000 nM, less than 1,500 nM, less than 1,000 nM, less than 750 nM, less than 500 nM, less than 250 nM, less than 150 nM, less than 100 nM, less than 80 nM, 10-2,000 nM, 25-1,000 nM, 50-500 nM, 10-150 nM, 25-75 nM, or 50-60 nM. D ) binds to the N-terminal leucine.

13. The amino acid binding protein of claim 11 or 12, wherein the amino acid binding protein has a K of less than 2000 nM, less than 1500 nM, less than 1000 nM, less than 750 nM, less than 500 nM, less than 250 nM, less than 150 nM, less than 100 nM, less than 80 nM, 10-2000 nM, 25-1000 nM, 50-500 nM, 10-150 nM, 30-80 nM or 60-75 nM. D Binds to N-terminal isoleucine.

14. The amino acid binding protein of any one of claims 11-13, wherein the amino acid binding protein has a K of less than 2000 nM, less than 1500 nM, less than 1000 nM, less than 750 nM, less than 500 nM, less than 300 nM, less than 250 nM, less than 200 nM, 10-2000 nM, 25-1000 nM, 50-500 nM, 50-300 nM, or 100-200 nM. D Binds to the N-terminal valine.

15. The amino acid binding protein of any one of claims 8-14, wherein the amino acid binding protein has a length of at least 50 amino acids, at least 75 amino acids, at least 100 amino acids, 50-250 amino acids, 50-150 amino acids, or 100-200 amino acids.

16. The amino acid binding protein according to any one of claims 8 to 15, wherein the loop between β1 and α1 comprises three or more negatively charged amino acids.

17. The amino acid binding protein according to any one of claims 8 to 16, wherein the loop between β1 and α1 comprises four or more negatively charged amino acids.

18. The amino acid binding protein according to claim 16 or 17, wherein at least two negatively charged amino acids in the loop between β1 and α1 form hydrogen bonds with the amino acid ligand.

19. The amino acid binding protein according to any one of claims 16 to 18, wherein at least one negatively charged amino acid in the loop between β1 and α1 forms a bifurcated hydrogen bond with an amino acid ligand.

20. The amino acid binding protein according to any one of claims 16-19, wherein the negatively charged amino acid is selected from aspartic acid and glutamic acid.

21. The amino acid binding protein of any one of claims 8-20, wherein β1-α1 comprises an amino acid sequence that is at least 80% identical to the sequence of amino acids 35-58 of SEQ ID NO:

1.

22. The amino acid binding protein of claim 21, wherein the binding pocket is formed by amino acids at one or more positions corresponding to amino acids 41-47 and 50 of SEQ ID NO:

1.

23. The amino acid binding protein according to claim 21 or 22, wherein the amino acid sequence comprises an amino acid substitution at one or more positions corresponding to L39, N41, D42, H45, V50 and Q55 of SEQ ID NO:

1.

24. The amino acid binding protein of claim 23, wherein at least one amino acid substitution is at the position corresponding to N41 of SEQ ID NO:

1.

25. The amino acid binding protein of claim 23 or 24, wherein the amino acid substitution is selected from N41D and Q55R.

26. The amino acid binding protein of any one of claims 8-25, wherein α2 comprises an amino acid sequence that is at least 80% identical to the sequence of amino acids 62-73 of SEQ ID NO:

1.

27. The amino acid binding protein of claim 26, wherein the binding pocket is formed by amino acids at one or more positions corresponding to amino acids 69, 72 and 73 of SEQ ID NO:

1.

28. The amino acid binding protein according to claim 26 or 27, wherein the amino acid sequence comprises an amino acid substitution at one or more positions corresponding to P62, E63, L68 and V72 of SEQ ID NO:

1.

29. The amino acid binding protein of claim 28, wherein at least one amino acid substitution is at the position corresponding to V72 of SEQ ID NO:

1.

30. The amino acid binding protein of claim 28 or 29, wherein the amino acid substitution is selected from E63S, L68M and V72M.

31. The amino acid binding protein of any one of claims 8-30, wherein the loop between α3 and β3 comprises an amino acid sequence that is at least 80% identical to the sequence of amino acids 99-112 of SEQ ID NO:

1.

32. The amino acid binding protein of claim 31, wherein the binding pocket is formed by the amino acid at the position corresponding to amino acid 111 of SEQ ID NO:

1.

33. The amino acid binding protein according to claim 31 or 32, wherein the amino acid sequence comprises an amino acid substitution at one or more positions corresponding to Y100 and M111 of SEQ ID NO:

1.

34. The amino acid binding protein of claim 33, wherein the amino acid substitution is Y100R.

35. The amino acid binding protein of any one of claims 8-34, wherein the plurality of hydrogen bond acceptors of the binding pocket are configured to form at least two, at least three, at least four or at least five hydrogen bonds in the presence of an amino acid ligand.

36. The amino acid binding protein of any one of claims 8-35, wherein the binding pocket comprises two or more of (i), (ii), (iii), (iv) and (v).

37. The amino acid binding protein of any one of claims 8-36, wherein the binding pocket comprises three or more of (i), (ii), (iii), (iv) and (v).

38. The amino acid binding protein of any one of claims 8-37, wherein the binding pocket comprises four or more of (i), (ii), (iii), (iv) and (v).

39. The amino acid binding protein of any one of claims 8-38, wherein the binding pocket comprises (i), (ii), (iii), (iv) and (v).

40. according to any one of claims 8-39 amino acid binding proteins, wherein the structural equivalent is a root mean square difference of no more than A structure wherein at least 80% of the secondary structure α-carbon atoms are aligned with the structure of formula (I).

41. according to any one of claims 8-40 amino acid binding proteins, wherein the structural equivalent is a root mean square difference of no more than No more than No more than or not more than A structure wherein at least 80% of the secondary structure α-carbon atoms are aligned with the structure of formula (I).

42. The amino acid binding protein of any one of claims 8-41, wherein the amino acid binding protein has a residue selected from the group consisting of PS635-645, PS731-732, PS759-766, PS769, PS795-870, PS896-912, PS918-1043, PS1048-1100, PS1124-1137, PS1141-1161, PS1175-1199, PS1203-1217, PS1222-1245, PS1277-1305, PS1321-1350, and PS1425-1448 (SEQ ID NO:22-27, 87-88, 115-122, 125, 151-226, 249-265, 271-390, 395-446, 470-483, 487-507, 521-545, 549-563, 568-591, 622-650, 664-693 and 768-791) is at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, 40-100%, 50-100%, 60-100%, 70-100%, 80-100%, 90-100% or 95-100% identical amino acid sequence to any one of the sequences.

43. A recombinant or synthetic amino acid binding protein having an amino acid sequence at least 80% identical to SEQ ID NO:2, wherein the amino acid sequence comprises an amino acid substitution at one or more positions corresponding to G19, K26, S29, F30, D31, D32, T33, C34, V35, T47, G48, T53, T54, T57, E58, F59, N61, I63, D65, D68, E70, A71, H74 and T75 of SEQ ID NO:

2.

44. The amino acid binding protein of claim 43, wherein the amino acid sequence comprises amino acid substitutions at positions corresponding to I63 and E70 and at one or more positions corresponding to G19, K26, S29, D32, T47, G48, T53, T54, T57, E58, F59, N61, H74, and T75.

45. The amino acid binding protein of claim 43 or 44, wherein the amino acid sequence comprises an amino acid substitution at one or more positions corresponding to K26, D32, T47, I63 and E70.

46. ​​The amino acid binding protein of any one of claims 43-45, wherein the amino acid substitutions are selected from G19R, K26R, S29Q, D32R / Y, T47K / L / R, G48R / Y, T53V, T54K, T57K / R, E58K, F59R, N61K, I63E, E70S / T, H74K and T75E.

47. The amino acid binding protein of any one of claims 43-46, wherein the amino acid substitution is selected from K26R, D32R, T47L, I63E and E70T.

48. The amino acid binding protein of any one of claims 43-47, wherein the amino acid sequence is at least 85%, at least 90%, at least 95%, at least 98%, 80-98%, 80-95%, 80-90%, 85-95% or 90-98% identical to SEQ ID NO:

2.

49. The amino acid binding protein according to any one of claims 43-48, wherein the amino acid sequence is a sequence selected from any one of PS1101-1122, PS1218-1221 and PS1351-1398 (SEQ ID NOs: 447-468, 564-567 and 694-741).

50. A recombinant or synthetic amino acid binding protein comprising a structure of formula (II) or a structural equivalent thereof: β1–α1–β2–α2–α3 (II) in: Each of β1 and β2 is a β strand; Each of α1, α2, and α3 is an α-helix; Every instance of "–" is a ring; as well as At least a portion of each of α2, the loop between β1 and α1, and the loop between β2 and α2 forms a binding pocket for an amino acid ligand, wherein the binding pocket comprises one or more of the following: i) approx. The volume of ii)-3.0RTe c -1 or lower electrostatic potential, iii) a plurality of hydrogen bond acceptors configured to form one or more hydrogen bonds in the presence of the amino acid ligands, and iv) a plurality of van der Waals contact positions which are configured to form van der Waals interactions in the presence of amino acid ligands.

51. The amino acid binding protein of claim 50, wherein the amino acid ligand is a polypeptide comprising at least three amino acids.

52. The amino acid binding protein of claim 50 or 51, wherein the amino acid ligand comprises the N-terminal amino acid of a polypeptide.

53. The amino acid binding protein of claim 52, wherein the N-terminal amino acid is arginine.

54. The amino acid binding protein of claim 53, wherein the amino acid binding protein has a dissociation constant (K) of less than 2,000 nM, less than 1,500 nM, less than 1,000 nM, less than 750 nM, less than 500 nM, less than 400 nM, less than 200 nM, less than 100 nM, 10-2,000 nM, 25-1,000 nM, 50-500 nM, 10-400 nM, 25-75 nM, or 50-80 nM. D ) binds to the N-terminal arginine.

55. The amino acid binding protein of any one of claims 50-54, wherein the amino acid binding protein is at least 50 amino acids, at least 75 amino acids, at least 100 amino acids, 50-250 amino acids, 50-150 amino acids, or 100-200 amino acids in length.

56. The amino acid binding protein of any one of claims 50-55, wherein the loop between β2 and α2 comprises three or more negatively charged amino acids.

57. The amino acid binding protein of any one of claims 50-56, wherein the loop between β2 and α2 comprises four or more negatively charged amino acids.

58. The amino acid binding protein of claim 56 or 57, wherein at least three negatively charged amino acids of the loop between β2 and α2 form hydrogen bonds with the amino acid ligand.

59. The amino acid binding protein of any one of claims 56-58, wherein at least one negatively charged amino acid of the loop between β2 and α2 forms a hydrogen bond with the amino terminus of the amino acid ligand.

60. The amino acid binding protein of any one of claims 56-59, wherein the negatively charged amino acid is selected from aspartic acid and glutamic acid.

61. The amino acid binding protein of any one of claims 50-60, wherein at least one amino acid of α2 forms a hydrogen bond with an amino acid ligand.

62. The amino acid binding protein of any one of claims 50-61, wherein α2 comprises one or more polar uncharged amino acids.

63. The amino acid binding protein of claim 62, wherein at least one polar uncharged amino acid of α2 forms a hydrogen bond with a side chain of an amino acid ligand.

64. The amino acid binding protein of any one of claims 50-63, wherein the loop between β1 and α1 comprises an amino acid sequence that is at least 80% identical to the sequence of amino acids 27-42 of SEQ ID NO:

2.

65. The amino acid binding protein of claim 64, wherein the binding pocket is formed by amino acids at one or more positions corresponding to amino acids 31, 32, and 34-36 of SEQ ID NO:

2.

66. The amino acid binding protein of any one of claims 50-65, wherein the loop between α1 and β2 comprises an amino acid sequence that is at least 50% identical to the sequence of amino acids 47-50 of SEQ ID NO:

2.

67. The amino acid binding protein of claim 66, wherein the amino acid sequence comprises an amino acid substitution at the position corresponding to T47 of SEQ ID NO:

2.

68. The amino acid binding protein of claim 67, wherein the amino acid substitution is T47L.

69. The amino acid binding protein of any one of claims 50-68, wherein β2-α2 comprises an amino acid sequence that is at least 80% identical to amino acids 51-71 of SEQ ID NO:

2.

70. The amino acid binding protein of claim 69, wherein the binding pocket is formed by amino acids at one or more positions corresponding to amino acids 63, 65, 68, 70 and 71 of SEQ ID NO:

2.

71. The amino acid binding protein of claim 69 or 70, wherein the amino acid sequence comprises an amino acid substitution at one or more positions corresponding to I63 and E70 of SEQ ID NO:

2.

72. The amino acid binding protein of claim 71, wherein the amino acid substitution is selected from I63E and E70T.

73. The amino acid binding protein of any one of claims 50-72, wherein the plurality of hydrogen bond acceptors of the binding pocket are configured to form at least two, at least three, at least four, or at least five hydrogen bonds in the presence of an amino acid ligand.

74. The amino acid binding protein of any one of claims 50-73, wherein the binding pocket comprises two or more of (i), (ii), (iii) and (iv).

75. The amino acid binding protein of any one of claims 50-74, wherein the binding pocket comprises three or more of (i), (ii), (iii) and (iv).

76. The amino acid binding protein of any one of claims 50-75, wherein the binding pocket comprises (i), (ii), (iii) and (iv).

77. The amino acid binding protein of any one of claims 50-76, wherein the structural equivalents are RMS differences of no more than A structure wherein at least 80% of the secondary structure α-carbon atoms are aligned with the structure of formula (II).

78. The amino acid binding protein of any one of claims 50-77, wherein the structural equivalents are RMS differences of no more than No more than No more than or not more than A structure wherein at least 80% of the secondary structure α-carbon atoms are aligned with the structure of formula (II).

79. The amino acid binding protein of any one of claims 50-78, wherein the amino acid binding protein has an amino acid sequence that is at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, 40-100%, 50-100%, 60-100%, 70-100%, 80-100%, 90-100% or 95-100% identical to a sequence selected from any one of PS1101-1122, PS1218-1221 and PS1351-1398 (SEQ ID NOs: 447-468, 564-567 and 694-741).

80. A recombinant or synthetic amino acid binding protein having an amino acid sequence at least 80% identical to SEQ ID NO:3, wherein the amino acid sequence comprises amino acid substitutions at one or more positions corresponding to S22, C23, Y24, C25, E26, S39, W75, D76, Y77, H78, C85, N120, H145 and M146 of SEQ ID NO:

3.

81. The amino acid binding protein of claim 80, wherein the amino acid sequence comprises amino acid substitutions at positions corresponding to S22, C25, H78, C85 and N120.

82. The amino acid binding protein of claim 80 or 81, wherein the amino acid substitution is selected from S22E, C25S, H78Q, H78K, C85T, N120R and M146E.

83. The amino acid binding protein of any one of claims 80-82, wherein the amino acid substitution is selected from S22E, C25S, H78Q, H78K, C85T and N120R.

84. The amino acid binding protein of any one of claims 80-83, wherein the amino acid sequence is at least 85%, at least 90%, at least 95%, at least 98%, 80-98%, 80-95%, 80-90%, 85-95% or 90-98% identical to SEQ ID NO:

3.

85. The amino acid binding protein of any one of claims 80-84, wherein the amino acid sequence is a sequence selected from any one of PS1258-1260, PS1315-1318, PS1457-1478, PS1480-1499, PS1633-1656, PS1737-1758, PS1821-1898, PS2014-2057, and PS2116-2137 (SEQ ID NOs: 604-606, 660-663, 792-833, and 836-1025).

86. A recombinant or synthetic amino acid binding protein comprising a structure of formula (III) or a structural equivalent thereof: α1–α2–α3β1–β2–β3–β4–β5–α4–β6α5–α6 (III) in: Each of α1, α2, α3, α4, α5, and α6 is an α-helix; Each of β1, β2, β3, β4, β5, and β6 is a β strand; Every instance of "–" is a ring; and At least a portion of each of α2, β3, β4, α5, the loop between α1 and α2, and the loop between β3 and β4 forms a binding pocket for an amino acid ligand, wherein the binding pocket comprises one or more of the following: i) approx. The volume of ii) 2.0RTe c -1 or lower electrostatic potential, iii) a plurality of hydrogen bond acceptors or donors configured to form one or more hydrogen bonds in the presence of the amino acid ligands, iv) a plurality of van der Waals contact positions configured to form van der Waals interactions in the presence of an amino acid ligand, and v) at least one negatively charged amino acid and at least one positively charged amino acid.

87. The amino acid binding protein of claim 86, wherein the amino acid ligand is a polypeptide comprising at least three amino acids.

88. The amino acid binding protein of claim 86 or 87, wherein the amino acid ligand comprises the N-terminal amino acid of a polypeptide.

89. The amino acid binding protein of claim 88, wherein the N-terminal amino acid is selected from the group consisting of glutamine, asparagine, glutamic acid, aspartic acid, and cysteine-S-acetamide.

90. The amino acid binding protein of claim 89, wherein the amino acid binding protein has a dissociation constant (K) of less than 2,000 nM, less than 1,500 nM, less than 1,000 nM, less than 750 nM, less than 500 nM, less than 250 nM, less than 150 nM, less than 100 nM, less than 50 nM, 10-2,000 nM, 25-1,000 nM, 50-500 nM, 10-100 nM, 25-250 nM, or 50-150 nM. D ) binds to the N-terminal glutamine.

91. The amino acid binding protein of any one of claims 86-90, wherein the amino acid binding protein is at least 50 amino acids, at least 75 amino acids, at least 100 amino acids, 50-250 amino acids, 50-150 amino acids, or 100-200 amino acids in length.

92. The amino acid binding protein of any one of claims 86-91, wherein each of α2 and β4 comprises at least one polar uncharged amino acid that forms a hydrogen bond with an amino acid ligand.

93. The amino acid binding protein of claim 92, wherein at least one polar uncharged amino acid of α2 is serine.

94. The amino acid binding protein of claim 92 or 93, wherein at least one polar uncharged amino acid of β4 is glutamine.

95. The amino acid binding protein of any one of claims 86-94, wherein α2 and the loop between α1 and α2 comprise an amino acid sequence that is at least 80% identical to the sequence of amino acids 18-40 of SEQ ID NO:

3.

96. The amino acid binding protein of claim 95, wherein the binding pocket is formed by amino acids at one or more positions corresponding to amino acids 23-26 of SEQ ID NO:

3.

97. The amino acid binding protein of claim 95 or 96, wherein the amino acid sequence comprises an amino acid substitution at the position corresponding to C25 of SEQ ID NO:

3.

98. The amino acid binding protein of claim 97, wherein the amino acid substitution is C25S.

99. The amino acid binding protein of any one of claims 86-98, wherein β3-β4 comprises an amino acid sequence that is at least 80% identical to the sequence of amino acids 73-85 of SEQ ID NO:

3.

100. The amino acid binding protein of claim 99, wherein the binding pocket is formed by amino acids at one or more positions corresponding to amino acids 75-78 of SEQ ID NO:

3.

101. The amino acid binding protein of claim 99 or 100, wherein the amino acid sequence comprises an amino acid substitution at the position corresponding to H78 of SEQ ID NO:

3.

102. The amino acid binding protein of claim 101, wherein the amino acid substitution is H78Q or H78K.

103. The amino acid binding protein of any one of claims 86-102, wherein α6 comprises an amino acid sequence that is at least 66% identical to the sequence of amino acids 144-146 of SEQ ID NO:

3.

104. The amino acid binding protein of claim 103, wherein the binding pocket is formed by amino acids at one or more positions corresponding to amino acids 145-146 of SEQ ID NO:

3.

105. The amino acid binding protein of any one of claims 86-104, wherein the plurality of hydrogen bond acceptors or donors of the binding pocket are configured to form at least two, at least three, at least four, or at least five hydrogen bonds in the presence of an amino acid ligand.

106. The amino acid binding protein of any one of claims 86-105, wherein the binding pocket comprises two or more of (i), (ii), (iii), (iv) and (v).

107. The amino acid binding protein of any one of claims 86-106, wherein the binding pocket comprises three or more of (i), (ii), (iii), (iv) and (v).

108. The amino acid binding protein of any one of claims 86-107, wherein the binding pocket comprises four or more of (i), (ii), (iii), (iv) and (v).

109. The amino acid binding protein of any one of claims 86-108, wherein the binding pocket comprises (i), (ii), (iii) and (iv).

110. The amino acid binding protein of claim 109, wherein the amino acid ligand is a polypeptide comprising an N-terminal glutamine or asparagine.

111. The amino acid binding protein of any one of claims 86-108, wherein the binding pocket comprises (i), (ii), (iii), (iv) and (v).

112. The amino acid binding protein of claim 111, wherein the amino acid ligand is a polypeptide comprising an N-terminal glutamic acid.

113. The amino acid binding protein according to any one of claims 86-112, comprising the structure of formula (III-A) or a structural equivalent thereof: α1–α2–α3β1–β2–β3–β4–β5–α4–β6α5–α6–α7–β7α8 (III-A), in: Each of α7 and α8 is an α-helix; and β7 is the β strand.

114. A recombinant or synthetic amino acid binding protein comprising a structure of formula (III-B) or a structural equivalent thereof: α1–α2–α3β1–β2–β3–β4 (III-B), in: Each of α1, α2, and α3 is an α-helix; Each of β1, β2, β3, and β4 is a β strand; Every instance of "–" is a ring; and At least a portion of each of α2, β3, β4, the loop between α1 and α2, and the loop between β3 and β4 forms a binding pocket for an amino acid ligand, wherein the binding pocket comprises: i) at least one negatively charged amino acid configured to form a hydrogen bond with an amino acid ligand, and ii) at least one positively charged amino acid arranged to form a hydrogen bond with the amino acid ligand.

115. The amino acid binding protein of claim 114, wherein the amino acid ligand is a polypeptide comprising at least three amino acids.

116. The amino acid binding protein of claim 114 or 115, wherein the amino acid ligand comprises the N-terminal amino acid of a polypeptide.

117. The amino acid binding protein of claim 116, wherein the N-terminal amino acid is glutamic acid.

118. The amino acid binding protein of any one of claims 114-117, wherein the at least one negatively charged amino acid forms hydrogen bonds with backbone atoms of the amino acid ligand, and The at least one positively charged amino acid forms a hydrogen bond with a side chain atom of the amino acid ligand.

119. The amino acid binding protein of any one of claims 114-118, wherein the side chain atoms of the at least one negatively charged amino acid form hydrogen bonds with backbone atoms of the amino acid ligand, and The side chain atom of the at least one positively charged amino acid forms a hydrogen bond with the side chain atom of the amino acid ligand.

120. The amino acid binding protein of any one of claims 114-119, wherein α2 comprises at least one negatively charged amino acid, and Wherein β4 contains at least one positively charged amino acid.

121. The amino acid binding protein of any one of claims 114-120, wherein the at least one negatively charged amino acid comprises glutamic acid, and Wherein the at least one positively charged amino acid comprises lysine.

122. The amino acid binding protein of any one of claims 114-121, wherein the at least one negatively charged amino acid corresponds to E26 of SEQ ID NO: 3, and wherein the at least one positively charged amino acid is a lysine substitution at the position corresponding to H78 of SEQ ID NO:

3.

123. The amino acid binding protein of any one of claims 114-122, wherein α1-α2 comprises an amino acid sequence that is at least 80% identical to the sequence of amino acids 15-39 of SEQ ID NO:

3.

124. The amino acid binding protein of claim 123, wherein the amino acid sequence comprises amino acid substitutions at positions corresponding to S22 and C25 of SEQ ID NO:

3.

125. The amino acid binding protein of claim 124, wherein the amino acid substitutions are S22E and C25S.

126. The amino acid binding protein of any one of claims 114-125, wherein β3-β4 comprises an amino acid sequence that is at least 80% identical to the sequence of amino acids 73-85 of SEQ ID NO:

3.

127. The amino acid binding protein of claim 126, wherein the amino acid sequence comprises amino acid substitutions at positions corresponding to H78 and C85 of SEQ ID NO:

3.

128. The amino acid binding protein of claim 127, wherein the amino acid substitutions are H78K and C85T.

129. according to any one of claims 86-128 amino acid binding proteins, wherein said structural equivalents are root mean square differences of no more than A structure wherein at least 80% of the secondary structure α-carbon atoms are aligned with the structure of formula (III), (III-A) or (III-B).

130. The amino acid binding proteins of any one of claims 86-129, wherein the structural equivalents are RMS differences of no more than No more than No more than or not more than A structure wherein at least 80% of the secondary structure α-carbon atoms are aligned with the structure of formula (III), (III-A) or (III-B).

131. The amino acid binding protein of any one of claims 86-130, wherein the amino acid binding protein has a residue selected from the group consisting of PS1258-1260, PS1315-1318, PS1457-1478, PS1480-1499, PS1633-1656, PS1737-1758, PS1821-1898, PS2014-2057, and PS2116-2137 (SEQ ID NO:604-606, 660-663, 792-833 and 836-1025) is at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, 40-100%, 50-100%, 60-100%, 70-100%, 80-100%, 90-100% or 95-100% identical amino acid sequence to any one of the sequences.

132. The amino acid binding protein of any preceding claim, further comprising one or more tags.

133. The amino acid binding protein of claim 132, wherein the one or more tags comprise a luminescent tag or a conductivity tag.

134. The amino acid binding protein of claim 133, wherein the luminescent tag comprises at least one fluorophore dye molecule.

135. The amino acid binding protein of claim 133 or 134, wherein the luminescent tag comprises 20 or fewer fluorophore dye molecules.

136. The amino acid binding protein of any one of claims 133-135, wherein the luminescent tag comprises at least one FRET pair comprising a donor tag and an acceptor tag.

137. The amino acid binding protein of any one of claims 133-136, wherein the conductivity tag comprises a charged polymer.

138. The amino acid binding protein of any one of claims 132-137, wherein the one or more tags comprise a tag sequence.

139. The amino acid binding protein of claim 138, wherein the tag sequence comprises one or more of a purification tag, a cleavage site, and a biotinylation sequence.

140. The amino acid binding protein of claim 139, wherein the biotinylation sequence comprises at least one biotin ligase recognition sequence.

141. The amino acid binding protein of claim 139 or 140, wherein the biotinylation sequence comprises two tandem biotin ligase recognition sequences.

142. The amino acid binding protein of any one of claims 132-141, wherein the one or more tags comprise a biotin moiety.

143. The amino acid binding protein of claim 142, wherein the biotin moiety comprises at least one biotin molecule.

144. The amino acid binding protein of claim 142 or 143, wherein the biotin moiety is a bi-biotin moiety.

145. The amino acid binding protein of claim 143 or 144, wherein the tag comprises at least one biotin ligase recognition sequence having the at least one biotin molecule attached thereto.

146. The amino acid binding protein of any one of claims 132-145, wherein the one or more tags comprise one or more polyol moieties.

147. The amino acid binding protein of claim 146, wherein the one or more polyol moieties comprise dextran, polyvinyl pyrrolidone, polyethylene glycol, polypropylene glycol, polyoxyethylene glycol, polyvinyl alcohol, or a combination or variant thereof.

148. The amino acid binding protein of any one of claims 132-147, wherein the amino acid binding protein comprises one or more unnatural amino acids having the one or more tags attached thereto.

149. An amino acid recognition agent comprising a polypeptide having at least a first amino acid binding protein and a second amino acid binding protein connected end-to-end, wherein the first amino acid binding protein and the second amino acid binding protein are separated by a linker comprising at least two amino acids, wherein at least one of the first amino acid binding protein and the second amino acid binding protein is an amino acid binding protein according to any of the preceding claims.

150. The amino acid recognition agent of claim 149, wherein the first amino acid binding protein and the second amino acid binding protein are the same.

151. The amino acid recognition agent of claim 149, wherein the first amino acid binding protein and the second amino acid binding protein are different.

152. The amino acid recognition agent of any one of claims 149-151, wherein each of the first amino acid binding protein and the second amino acid binding protein is independently an amino acid binding protein according to any preceding claim.

153. The amino acid recognition agent of any one of claims 149-152, wherein the linker comprises up to 100 amino acids, up to 80 amino acids, up to 60 amino acids, up to 50 amino acids, about 5 to about 100 amino acids, or about 5 to about 50 amino acids.

154. The amino acid recognition agent according to any one of claims 149-153, wherein: The first amino acid binding protein has an amino acid sequence that is at least 80% identical to PS961 (SEQ ID NO: 314); and The second amino acid binding protein has an amino acid sequence that is at least 80% identical to PS961 (SEQ ID NO: 314).

155. An amino acid recognition agent according to claim 154, comprising an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, 80-100%, 85-100%, 90-100%, 95-100% or 100% identical to any one of PS1038, PS1222 and PS1223 (SEQ ID NOs: 389, 568 and 569).

156. An amino acid recognition agent according to any one of claims 149-153, wherein: The first amino acid binding protein has an amino acid sequence that is at least 80% identical to PS1122 (SEQ ID NO:468); and The second amino acid binding protein has an amino acid sequence that is at least 80% identical to PS1122 (SEQ ID NO: 468).

157. The amino acid recognition agent of claim 156, comprising an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, 80-100%, 85-100%, 90-100%, 95-100% or 100% identical to a sequence selected from any one of PS1219-PS1221 (SEQ ID NOs: 565-567).

158. An amino acid recognition agent according to any one of claims 149-153, wherein: The first amino acid binding protein has an amino acid sequence that is at least 80% identical to PS1259 (SEQ ID NO: 605); and The second amino acid binding protein has an amino acid sequence that is at least 80% identical to PS1259 (SEQ ID NO: 605).

159. The amino acid recognition agent of claim 158, comprising an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, 80-100%, 85-100%, 90-100%, 95-100% or 100% identical to PS1599 (SEQ ID NO: 835).

160. An amino acid recognition agent comprising a polypeptide having an amino acid binding protein and a marker protein connected end-to-end, wherein the amino acid binding protein and the marker protein are separated by a linker comprising at least two amino acids, wherein the amino acid binding protein is the amino acid binding protein according to any preceding claim.

161. The amino acid recognition agent of claim 160, wherein the marker protein has a molecular weight of at least 10 kDa, about 10 kDa to about 150 kDa, or about 15 kDa to about 100 kDa.

162. The amino acid recognition agent of claim 160 or 161, wherein the marker protein comprises at least 50 amino acids, about 50 to about 1000 amino acids, or about 100 to about 750 amino acids.

163. An amino acid recognition agent according to any one of claims 160-162, wherein the marker protein comprises a fluorescent protein.

164. The amino acid recognition agent of any one of claims 160-163, wherein the marker protein comprises a luminescent tag.

165. The amino acid recognition agent according to any one of claims 160-164, wherein the marker protein comprises a protein selected from the group consisting of maltose binding protein, glutathione S-transferase, green fluorescent protein, SNAP-tag and DNA polymerase.

166. The amino acid recognition agent of any one of claims 160-165, wherein the linker comprises up to 100 amino acids, up to 80 amino acids, up to 60 amino acids, up to 50 amino acids, about 5 to about 100 amino acids, or about 5 to about 50 amino acids.

167. A composition comprising two or more amino acid recognition agents, wherein at least one amino acid recognition agent is an amino acid binding protein according to any preceding claim.

168. The composition of claim 167, wherein the composition comprises at least one amino acid binding protein according to any one of claims 1-42.

169. The composition of claim 167 or 168, wherein the composition comprises at least one amino acid binding protein according to any one of claims 43-79.

170. The composition of any one of claims 167-169, wherein the composition comprises at least one amino acid binding protein of any one of claims 80-131.

171. The composition of any one of claims 167-170, wherein the composition comprises: The first amino acid binding protein according to any one of claims 1 to 42, The second amino acid binding protein according to any one of claims 43-79, and The third amino acid binding protein according to any one of claims 80-131.

172. The composition of any one of claims 167-171, wherein the composition comprises at least one type of lysis reagent.

173. The composition of claim 172, wherein the cleavage reagent comprises an exopeptidase.

174. The composition of claim 172 or 173, wherein the cleavage reagent comprises an aminopeptidase.

175. The composition of any one of claims 172-174, wherein the molar ratio of the amino acid recognition agent to the cleavage reagent in the composition is about 1:1,000 to about 1:1, about 1:1 to about 100:1, about 1:100 to about 1:1, about 1:1 to about 10:1, about 1:1,000, about 1:500, about 1:200, about 1:100, about 1:10, about 1:5, about 1:2, about 1:1, about 5:1, about 10:1, about 50:1 or about 100:

1.

176. A method of determining at least one chemical characteristic of a polypeptide, the method comprising: Contacting the polypeptide with a composition according to any one of claims 167-175; and monitoring a signal corresponding to a signal pulse of interaction between the one or more amino acid recognition agents and the polypeptide; and At least one chemical characteristic of the polypeptide is determined based on the characteristic pattern in the signal.

177. The method of claim 176, wherein determining at least one chemical feature comprises identifying at least one amino acid in the polypeptide as a naturally occurring amino acid, a non-natural amino acid, or a modified variant thereof.

178. The method of claim 176 or 177, wherein determining at least one chemical characteristic comprises identifying at least one amino acid in the polypeptide as having a side chain that is negatively charged, positively charged, uncharged, polar, non-polar, hydrophobic, aromatic, or a combination thereof.

179. The method of any one of claims 176-178, wherein determining at least one chemical feature comprises identifying at least one amino acid in the polypeptide as a type selected from alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, selenocysteine, serine, threonine, tryptophan, tyrosine, and valine.

180. The method of any one of claims 176-179, wherein determining at least one chemical feature comprises identifying at least one amino acid in the polypeptide as having a post-translational modification.

181. The method of any one of claims 176-180, wherein the post-translational modification is selected from the group consisting of acetylation, ADP-ribosylation, caspase cleavage, citrullination, formylation, N-linked glycosylation, O-linked glycosylation, hydroxylation, methylation, myristoylation, ubiquitination, nitration, oxidation, palmitoylation, phosphorylation, prenylation, S-nitrosylation, sulfation, sumoylation, and ubiquitination.

182. The method of any one of claims 176-181, wherein monitoring comprises: A series of signal pulses is detected, wherein a characteristic pattern in the series of signal pulses is indicative of at least one chemical characteristic of the polypeptide.

183. The method of claim 182, wherein the series of signal pulses corresponds to a series of binding events between one or more amino acid recognition agents and a polypeptide, respectively.

184. A method according to claim 182 or 183, wherein the signal pulses of the characteristic mode include an average pulse duration of about 1 millisecond to about 10 seconds, about 50 milliseconds to about 2 seconds, about 50 milliseconds to about 500 milliseconds, or about 500 milliseconds to about 2 seconds.

185. The method of any one of claims 182-184, wherein the characteristic pattern in the series of signal pulses comprises at least 10 signal pulses, about 50 to about 200 signal pulses, or about 25 to about 100 signal pulses.

186. A system comprising: at least one hardware processor; as well as At least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by at least one hardware processor, cause the at least one hardware processor to perform a method according to any one of claims 176-185.

187. At least one non-transitory computer-readable storage medium storing processor-executable instructions which, when executed by at least one hardware processor, cause the at least one hardware processor to perform a method according to any one of claims 176-185.

Citation Information

Patent Citations

  • Optical rejection photonic structures using two spatial filters

    US11237326B2

  • Integrated device with external light source for probing detecting and analyzing molecules

    US20150141267A1

  • Integrated device for temporal binning of received photons

    US20160133668A1

  • Pulsed laser and bioanalytic system

    US20160344156A1

  • Single-molecule nanofet sequencing systems and methods

    US20170037462A1