MODIFIED DIPEPTIDE CLEAVASEN, USES THEREOF AND ASSOCIATED KITS

DE602021047771T2Active Publication Date: 2026-02-11ENCODIA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE602021047771
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-09-30
Filing Date
2021-03-19
Publication Date
2026-02-11
Estimated Expiration
2041-03-19

AI Technical Summary

Technical Problem

Existing methods for peptide and protein degradation, such as Edman degradation, are limited by harsh chemical conditions and lack substrate-specific enzymes for efficient amino acid removal, which can be incompatible with sensitive analysis methods like nucleic acid analysis.

Method used

Development of modified dipeptide cleavases with specific amino acid substitutions in the substrate binding site of dipeptidyl aminopeptidases, allowing for the selective removal of labeled terminal amino acids or dipeptides from polypeptides, while maintaining compatibility with sensitive analysis methods.

Benefits of technology

The modified dipeptide cleavases enable efficient and selective removal of labeled amino acids, facilitating protein sequencing and analysis without the use of harsh chemicals, and can be used in kits for targeted polypeptide treatment.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This disclosure claims priority to U.S. Provisional Patent Application No. 62 / 994,216, filed on March 24, 2020 and U.S. Provisional Patent Application No. 63 / 085,977, filed on September 30, 2020.STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH

[0002] This invention was made with Government support awarded by National Institute of General Medical Sciences of the National Institutes of Health under Grant No. 1R43GM130185-01 and Grant No. 5R44GM123836-03. The United States Government has certain rights in this invention pursuant to these grants.SEQUENCE LISTING ON ASCII TEXT

[0003] This patent or application file contains a Sequence Listing submitted in computer readable ASCII text format (file name: 4614-2002240_SeqList_ST25.txt, date recorded: March 17, 2021, size: 211,194 bytes). The content of the Sequence Listing file is incorporated herein by reference in its entirety.TECHNICAL FIELD

[0004] The present disclosure relates to modified dipeptide cleavases for cleaving amino acids from peptides, polypeptides, and proteins, including modified peptides, polypeptides, and proteins. Also provided are methods of using the modified dipeptide cleavases for treating polypeptides, and kits comprising the modified dipeptide cleavase. The methods and the kits also include other components for macromolecule sequencing and / or analysis.BACKGROUND

[0005] Enzymes that are involved in degradation of peptides and proteins, e.g., aminopeptidases, dipeptidyl peptidases, carboxypeptidases, endopeptidases, and others, hydrolyze peptide bonds (Sanderink et al., J. Clin. Chem. Clin. Biochem. (1988) 26:795-807). Various peptidases have been isolated and discovered in a number of organisms and from various tissues. Aminopeptidases naturally occur as monomeric and multimeric enzymes, and may be metal or ATP-dependent. Some substrate-specific peptidases specifically remove one or two amino acid residues at a time from the amino-terminus of the peptide while others remove from the carboxy-terminus of the protein or peptide. Natural aminopeptidases generally have limited specificity and eliminate amino acids in a processive manner, eliminating one amino acid one after another.

[0006] In some embodiments, methods for peptide degradation are useful for applications in protein analysis and / or sequencing. For example, peptide sequencing may involve Edman degradation to achieve stepwise degradation of the N-terminal amino acid (NTAA) on a peptide through a series of chemical modifications and downstream HPLC analysis or mass spectrometry analysis. However, in general, Edman degradation peptide sequencing may be limited, for example, typical Edman degradation requires deployment of high temperature and harsh chemical conditions (e.g., strong acids; anhydrous TFA) for long incubation times. In some cases, Edman degradation may not be compatible with processes for protein analysis methods which may be sensitive to harsh chemical conditions, such as analysis methods which employ nucleic acids (e.g., DNA).

[0007] Suzuki, Yoshiyuki, et al. (in "Identification of the catalytic triad of family S46 exopeptidases, closely related to clan PA endopeptidases." Scientific Reports 4.1 (2014): 4292.) disclosed catalytic residues in the dipeptidyl aminopeptidase DAP BII.

[0008] US 2002 / 164759 disclosed isolated polypeptides, dipeptidylpeptidases, active analogs, active fragments, or active modifications thereof, having amidolytic activity for cleavage of a peptide bond between the second and third amino acids from the N-terminal end of a target polypeptide.

[0009] WO 2017 / 192633 disclosed a method for analyzing macromolecules, including peptides, polypeptides, and proteins, employing nucleic acid encoding.

[0010] US 2019 / 246664 disclosed a site-specific mutagenesis modified yeast dipeptidyl peptidase III.

[0011] Accordingly, there remains a need for improved reagents for degradation of amino acids. For example, enzymatic methods for removing, eliminating, or cleaving amino acids from polypeptides may be desired. Furthermore, the availability of additional substrate-specific enzymes which can bind and remove desired amino acids from a polypeptide is also desired. In some cases, such improved reagents for removing amino acids are useful for protein sequencing and / or analysis. The present disclosure fulfills these and other related needs.

[0012] These and other aspects of the invention will be apparent upon reference to the following detailed description. To this end, various references are set forth herein which describe in more detail certain background information, procedures, compounds and / or compositions.BRIEF SUMMARY

[0013] The present invention is disclosed in independent claims 1, 8 13 and 14. Preferred embodiments are disclosed in the dependent claims.

[0014] In one aspect, disclosed herein is a modified dipeptide cleavase comprising three or more amino acid substitutions in a substrate binding site of a dipeptidyl aminopeptidase, wherein: (i) the dipeptidyl aminopeptidase removes or is configured to remove two terminal amino acids from a polypeptide; and (ii) the modified dipeptide cleavase removes or is configured to remove from the polypeptide having a terminal amino acid residue labeled with a chemical reagent (a) the single labeled terminal amino acid residue or (b) a labeled terminal dipeptide, wherein the dipeptidyl aminopeptidase comprises an amino acid sequence having at least 30 % sequence identity to the amino acid sequence of SEQ ID NO: 31 and also comprises an asparagine residue at a position corresponding to position 191 of SEQ ID NO: 31, a tryptophan or phenylalanine residue at a position corresponding to position 192 of SEQ ID NO: 31, an arginine residue at a position corresponding to position 196 of SEQ ID NO: 31, an asparagine residue at a position corresponding to position 306 of SEQ ID NO: 31, an aspartate residue at a position corresponding to position 650 of SEQ ID NO: 31; and wherein the modified dipeptide cleavase comprises three or more amino acid substitutions in residues corresponding to positions 191, 192, 196, 306, 650 of SEQ ID NO: 31.

[0015] According to an embodiment, the modified dipeptide cleavase does not remove an unlabeled terminal dipeptide from the polypeptide.

[0016] In yet another embodiment, the modified dipeptide cleavase comprises at least four amino acid substitutions in residues corresponding to positions 191, 192, 196, 306, 650 of SEQ ID NO: 31.

[0017] In yet another embodiment, the modified dipeptide cleavase removes or is configured to remove a single N-terminally labeled amino acid of the polypeptide.

[0018] In yet another embodiment, the modified dipeptide cleavase further comprises one or more amino acid substitutions in residues corresponding to positions 310, 628, 648, 651, 659, 669.

[0019] In yet another embodiment, the dipeptidyl aminopeptidase is a protein classified in MEROPS S46, or a functional homolog or fragment thereof.

[0020] In yet another embodiment, the single terminal amino acid or terminal dipeptide is labeled with a N-terminal modification that comprises a N-terminal blocking group (NTM blk ) and, optionally, a natural or unnatural amino acid portion (NTMaa), wherein the NTMaa comprises a compound selected from the group consisting of: a naturally-occurring amino acid residue, 3-(3'-pyridyl)-L-alanine, L-cyclohexylglycine, α-aminoisobutyric acid, 3-(4'-pyridyl)-L-alanine, L-azetidine-2-carboxylic acid, isonipecotic acid, L-phenylglycine, β-(2-thienyl)-L-alanine, 3-(4-thiazolyl)-L-alanine, 1-aminocyclopentane-1-carboxylic acid, (2-trifluoromethyl)-L-Phenylalanine, L-cyclopropylalanine, 3-(2'-pyridyl)-L-alanine, beta-cyano-L-alanine, α-methyl-L-4-Fluorophenylalanine, α-methyl-D-4-fluorophenylalanine, 3-amino-2,2-difluoro-propionic acid, O-sulfo-L-tyrosine sodium salt, L-2-furylalanine, 1-aminocyclopropane-1-carboxylic acid, 3,5-dinitro-L-tyrosine, pentafluoro-L-phenylalanine, 3,5-difluoro-L-phenylalanine, 3-fluoro-L-phenylalanine, N-cyclopentylglycine, 1-(amino)cyclohexanecarboxylic acid, N-methylalanine, 4-amino-tetrahydropyran-4-carboxylic acid, 4-amino-1,1-dioxothiane-4-carboxylic acid, 4-amino-1-methyl-4-piperidinecarboxylic acid, 2-amino-N-(2,4-dimethoxybenzyl)acetamido)acetic acid, or N-alkylated derivatives; and the NTM blk comprises a compound selected from the group consisting of: 4-methylbenzoic acid, 4-(dimethylamio)benzoic acid, nicotinic acid, 3-aminonicotinic acid, 2-pyrazinecarbooxylic acid, 5-amino-2-fluoro-isonicotinic acid, 2,3-pyrazinedicarboxylic acid, 4,7-Difluoroisobenzofuran-1,3-dicarboxylic acid, 4-chloro-2-aminobenzoic acid, 4-nitro-2-aminobenzoic acid, 7-methoxy-1h-benzo[d][1,3]oxazine-2,4-dione, 4-carboxy-2-aminobenzoic acid, 6-(Trifluoromethyl)-2,4-dihydro-1h-3, 1-benzoxazine-2,4-dione, 7-(Trifluoromethyl)-1h-benzo[d][1,3]oxazine-2,4-dione, 6-fluoro-2-aminobenzoic acid, 4-fluoro-2-aminobenzoic acid, 5-methoxy-2-aminobenzoic acid, 4-fluorobenzoic acid, 4-(trifluoromethyl)benzoic acid, 2-ethynyl-6-fluorobenzaldehyde, 2-aminobenzoic acid, Succinic anhydride, 3,6-Difluoropyridine-2-carboxylic acid, 2-Fluoronicotinic acid, 5-Bromo-2-hydroxynicotinic acid, 4-(Trifluoromethyl)pyrimidine-5-carboxylic acid, 2-Oxo-1,2-dihydropyridine-3-carboxylic acid, 5-Methyl-2-aminobenzoic acid, 6-Fluoropicolinic acid, 3-Methyl-2-aminobenzoic acid, 4-Methyl-2-aminobenzoic acid, 2-Amino-6-methylbenzoic acid, 2-Amino-6-fluorobenzoic acid, 2-Amino-5-fluorobenzoic acid, 2-Amino-3-fluorobenzoic acid, 2-Amino-4-fluorobenzoic acid, 2-Aminonicotinic acid, 4-Aminonicotinic acid, 3-Aminopicolinic acid, 2-Amino-4,5-difluorobenzoic acid, 3,4-difluorobenzoic acid, 3,4,5-difluorobenzoic acid, 3-(Methoxycarbonyl)bicyclo[1.1.1]pentane-1-carboxylic acid, 3,3-Difluorocyclobutane-1-carboxylic acid, 1-Methyl-2-oxo-piperidine-4-carboxylic acid, Tetrahydropyran-4-carboxylic acid, 5-Fluoroorotic acid, 3-Fluoro-4-nitrobenzoic acid, 3-(Difluoromethyl)-1-methyl-1H-pyrazole-4-carboxylic acid, 4-(Difluoromethoxy)benzoic acid, 1-(Difluoromethyl)-1h-pyrazole-3-carboxylic acid, 4-(Methanesulfonylamino)benzoic acid, 5-Fluoro-6-methoxynicotinic acid, Tetrahydro-2H-thiopyran-4-carboxylic acid 1,1-dioxide, 4-(1H-Tetrazol-5-yl)benzoic acid, 1,2,3-Thiadiazole-4-carboxylic acid, 1,3-Benzodioxole-4-carboxylic acid, 2,1,3-Benzoxadiazole-5-carboxylic acid, 1-Benzyl-3-methyl-1h-pyrazole-5-carboxylic acid, 1-Cyclopropyl-6,7-difluoro-1,4-dihydro-4-oxoquinoline-3-carboxylic acid, 3,4-Dichlorobenzoic acid, 5-Fluoro-6-methylpyridine-2-carboxylic acid, 4,5-Dimethyl-2-(1h-pyrrol-1-yl)thiophene-3-carboxylic acid, 1,3-Dimethyl-1h-thieno[2,3-c]pyrazole-5-carboxylic acid, 1-[(4-Fluorobenzene)sulfonyl]piperidine-3-carboxylic acid, 1-(4-Fluorobenzyl)-5-oxopyrrolidine-3-carboxylic acid, 3-Fluoro-4-methoxybenzoic acid, 4-Fluoro-3-nitrobenzoic acid, 6-Fluoro-4-oxochromene-2-carboxylic acid, 3-Fluorophenylacetic acid, 4-Fluoro-3-(trifluoromethyl)benzoic acid, 5-Furan-2-yl-isoxazole-3-carboxylic acid, 1-Isopropyl-2-(trifluoromethyl)-1h-benzimidazole-5-carboxylic acid, Levofloxacin carboxylic acid, 3,5,7-Trifluoroadamantane-1-carboxylic acid, 3,4,5-Trimethoxybenzoic acid, 2-Oxo-2,3-dihydro-1h-benzo[d]imidazole-4-carboxylic acid, 1-Methyl-3-(trifluoromethyl)-1h-pyrazole-5-carboxylic acid, 2-Morpholin-4-yl-isonicotinic acid, 1,3-Oxazole-4-carboxylic acid, 4-Carboxybenzenesulfonamide, 3,4-difluorobenzenesulfonyl chloride.

[0021] In another aspect, disclosed herein is a method of treating a polypeptide, comprising the following steps: labeling a terminal amino acid of the polypeptide with a chemical reagent to produce a labeled polypeptide; and contacting the labeled polypeptide with a modified dipeptide cleavase, wherein the modified dipeptide cleavase comprises three or more amino acid substitutions in a substrate binding site of a dipeptidyl aminopeptidase, wherein: (i) the dipeptidyl aminopeptidase removes or is configured to remove two terminal amino acids from the polypeptide upon contacting; and (ii) the modified dipeptide cleavase removes or is configured to remove from the polypeptide having a terminal amino acid labeled with a chemical reagent (a) the single labeled terminal amino acid or (b) a labeled terminal dipeptide, wherein the dipeptidyl aminopeptidase comprises an amino acid sequence having at least 30 % sequence identity to the amino acid sequence of SEQ ID NO: 31 and also comprises an asparagine residue at a position corresponding to position 191 of SEQ ID NO: 31, a tryptophan or phenylalanine residue at a position corresponding to position 192 of SEQ ID NO: 31, an arginine residue at a position corresponding to position 196 of SEQ ID NO: 31, an asparagine residue at a position corresponding to position 306 of SEQ ID NO: 31, an aspartate residue at a position corresponding to position 650 of SEQ ID NO: 31; and wherein the modified dipeptide cleavase comprises three or more amino acid substitutions in residues corresponding to positions 191, 192, 196, 306, 650 of SEQ ID NO: 31.

[0022] In an embodiment of the method, the modified dipeptide cleavase does not remove an unlabeled terminal dipeptide from the polypeptide.

[0023] In yet another embodiment of the method, the modified dipeptide cleavase comprises at least four amino acid substitutions in residues corresponding to positions 191, 192, 196, 306, 650 of SEQ ID NO: 31.

[0024] In yet another embodiment of the method, the single terminal amino acid or terminal dipeptide is labeled with a N-terminal modification that comprises a N-terminal blocking group (NTM blk ) and, optionally, a natural or unnatural amino acid portion (NTMaa), wherein the NTMaa comprises a compound selected from the group consisting of: a naturally-occurring amino acid residue, 3-(3'-pyridyl)-L-alanine, L-cyclohexylglycine, α-aminoisobutyric acid, 3-(4'-pyridyl)-L-alanine, L-azetidine-2-carboxylic acid, isonipecotic acid, L-phenylglycine, β-(2-thienyl)-L-alanine, 3-(4-thiazolyl)-L-alanine, 1-aminocyclopentane-1-carboxylic acid, (2-trifluoromethyl)-L-Phenylalanine, L-cyclopropylalanine, 3-(2'-pyridyl)-L-alanine, beta-cyano-L-alanine, α-methyl-L-4-Fluorophenylalanine, α-methyl-D-4-fluorophenylalanine, 3-amino-2,2-difluoro-propionic acid, O-sulfo-L-tyrosine sodium salt, L-2-furylalanine, 1-aminocyclopropane-1-carboxylic acid, 3,5-dinitro-L-tyrosine, pentafluoro-L-phenylalanine, 3,5-difluoro-L-phenylalanine, 3-fluoro-L-phenylalanine, N-cyclopentylglycine, 1-(amino)cyclohexanecarboxylic acid, N-methylalanine, 4-amino-tetrahydropyran-4-carboxylic acid, 4-amino-1,1-dioxothiane-4-carboxylic acid, 4-amino-1-methyl-4-piperidinecarboxylic acid, 2-amino-N-(2,4-dimethoxybenzyl)acetamido)acetic acid, or N-alkylated derivatives; and the NTM blk comprises a compound selected from the group consisting of: 4-methylbenzoic acid, 4-(dimethylamio)benzoic acid, nicotinic acid, 3-aminonicotinic acid, 2-pyrazinecarbooxylic acid, 5-amino-2-fluoro-isonicotinic acid, 2,3-pyrazinedicarboxylic acid, 4,7-Difluoroisobenzofuran-1,3-dicarboxylic acid, 4-chloro-2-aminobenzoic acid, 4-nitro-2-aminobenzoic acid, 7-methoxy-1h-benzo[d][1,3]oxazine-2,4-dione, 4-carboxy-2-aminobenzoic acid, 6-(Trifluoromethyl)-2,4-dihydro-1h-3, 1-benzoxazine-2,4-dione, 7-(Trifluoromethyl)-1h-benzo[d][1,3]oxazine-2,4-dione, 6-fluoro-2-aminobenzoic acid, 4-fluoro-2-aminobenzoic acid, 5-methoxy-2-aminobenzoic acid, 4-fluorobenzoic acid, 4-(trifluoromethyl)benzoic acid, 2-ethynyl-6-fluorobenzaldehyde, 2-aminobenzoic acid, Succinic anhydride, 3,6-Difluoropyridine-2-carboxylic acid, 2-Fluoronicotinic acid, 5-Bromo-2-hydroxynicotinic acid, 4-(Trifluoromethyl)pyrimidine-5-carboxylic acid, 2-Oxo-1,2-dihydropyridine-3-carboxylic acid, 5-Methyl-2-aminobenzoic acid, 6-Fluoropicolinic acid, 3-Methyl-2-aminobenzoic acid, 4-Methyl-2-aminobenzoic acid, 2-Amino-6-methylbenzoic acid, 2-Amino-6-fluorobenzoic acid, 2-Amino-5-fluorobenzoic acid, 2-Amino-3-fluorobenzoic acid, 2-Amino-4-fluorobenzoic acid, 2-Aminonicotinic acid, 4-Aminonicotinic acid, 3-Aminopicolinic acid, 2-Amino-4,5-difluorobenzoic acid, 3,4-difluorobenzoic acid, 3,4,5-difluorobenzoic acid, 3-(Methoxycarbonyl)bicyclo[1.1.1]pentane-1-carboxylic acid, 3,3-Difluorocyclobutane-1-carboxylic acid, 1-Methyl-2-oxo-piperidine-4-carboxylic acid, Tetrahydropyran-4-carboxylic acid, 5-Fluoroorotic acid, 3-Fluoro-4-nitrobenzoic acid, 3-(Difluoromethyl)-1-methyl-1H-pyrazole-4-carboxylic acid, 4-(Difluoromethoxy)benzoic acid, 1-(Difluoromethyl)-1h-pyrazole-3-carboxylic acid, 4-(Methanesulfonylamino)benzoic acid, 5-Fluoro-6-methoxynicotinic acid, Tetrahydro-2H-thiopyran-4-carboxylic acid 1,1-dioxide, 4-(1H-Tetrazol-5-yl)benzoic acid, 1,2,3-Thiadiazole-4-carboxylic acid, 1,3-Benzodioxole-4-carboxylic acid, 2,1,3-Benzoxadiazole-5-carboxylic acid, 1-Benzyl-3-methyl-1h-pyrazole-5-carboxylic acid, 1-Cyclopropyl-6,7-difluoro-1,4-dihydro-4-oxoquinoline-3-carboxylic acid, 3,4-Dichlorobenzoic acid, 5-Fluoro-6-methylpyridine-2-carboxylic acid, 4,5-Dimethyl-2-(1h-pyrrol-1-yl)thiophene-3-carboxylic acid, 1,3-Dimethyl-1h-thieno[2,3-c]pyrazole-5-carboxylic acid, 1-[(4-Fluorobenzene)sulfonyl]piperidine-3-carboxylic acid, 1-(4-Fluorobenzyl)-5-oxopyrrolidine-3-carboxylic acid, 3-Fluoro-4-methoxybenzoic acid, 4-Fluoro-3-nitrobenzoic acid, 6-Fluoro-4-oxochromene-2-carboxylic acid, 3-Fluorophenylacetic acid, 4-Fluoro-3-(trifluoromethyl)benzoic acid, 5-Furan-2-yl-isoxazole-3-carboxylic acid, 1-Isopropyl-2-(trifluoromethyl)-1h-benzimidazole-5-carboxylic acid, Levofloxacin carboxylic acid, 3,5,7-Trifluoroadamantane-1-carboxylic acid, 3,4,5-Trimethoxybenzoic acid, 2-Oxo-2,3-dihydro-1h-benzo[d]imidazole-4-carboxylic acid, 1-Methyl-3-(trifluoromethyl)-1h-pyrazole-5-carboxylic acid, 2-Morpholin-4-yl-isonicotinic acid, 1,3-Oxazole-4-carboxylic acid, 4-Carboxybenzenesulfonamide, 3,4-difluorobenzenesulfonyl chloride.

[0025] In yet another embodiment, the method further comprises a step of contacting the polypeptide with a binding agent configured to bind to the single labeled terminal amino acid or to the labeled terminal dipeptide, wherein the step of labeling the terminal amino acid of the polypeptide is before the step of contacting the polypeptide with the binding agent; and the step of contacting the polypeptide with the binding agent is before the step of contacting the polypeptide with the modified dipeptide cleavase.

[0026] In yet another aspect, disclosed herein is a set of dipeptide cleavase enzymes, comprising at least two different modified dipeptide cleavases, wherein: (i) each of the modified dipeptide cleavases from the set of dipeptide cleavase enzymes comprises three or more amino acid substitutions in a substrate binding site of a dipeptidyl aminopeptidase, wherein the dipeptidyl aminopeptidase is configured to remove two terminal amino acids from an unlabeled polypeptide; (ii) each of the modified dipeptide cleavases from the set of dipeptide cleavase enzymes is configured to remove a single labeled terminal amino acid from a polypeptide having the terminal amino acid labeled with a chemical reagent, wherein the dipeptidyl aminopeptidase comprises an amino acid sequence having at least 30 % sequence identity to the amino acid sequence of SEQ ID NO: 31 and also comprises an asparagine residue at a position corresponding to position 191 of SEQ ID NO: 31, a tryptophan or phenylalanine residue at a position corresponding to position 192 of SEQ ID NO: 31, an arginine residue at a position corresponding to position 196 of SEQ ID NO: 31, an asparagine residue at a position corresponding to position 306 of SEQ ID NO: 31, an aspartate residue at a position corresponding to position 650 of SEQ ID NO: 31; and wherein the modified dipeptide cleavase comprises three or more amino acid substitutions in residues corresponding to positions 191, 192, 196, 306, 650 of SEQ ID NO: 31; and (iii) the modified dipeptide cleavases from the set of dipeptide cleavase enzymes have different specificities for the labeled terminal amino acids, which the modified dipeptide cleavases are configured to remove.

[0027] In yet another aspect, disclosed herein is a kit for treating a polypeptide, comprising: (a) a chemical reagent for labeling a terminal amino acid of the polypeptide; and (b) a set of dipeptide cleavase enzymes, comprising at least two different modified dipeptide cleavases, wherein: (i) each of the modified dipeptide cleavases from the set of dipeptide cleavase enzymes comprises three or more amino acid substitutions in a substrate binding site of a dipeptidyl aminopeptidase, wherein the dipeptidyl aminopeptidase is configured to remove two terminal amino acids from an unlabeled polypeptide; (ii) each of the modified dipeptide cleavases from the set of dipeptide cleavase enzymes is configured to remove a single labeled terminal amino acid from a polypeptide having the terminal amino acid labeled with a chemical reagent, wherein the dipeptidyl aminopeptidase comprises an amino acid sequence having at least 30 % sequence identity to the amino acid sequence of SEQ ID NO: 31 and also comprises an asparagine residue at a position corresponding to position 191 of SEQ ID NO: 31, a tryptophan or phenylalanine residue at a position corresponding to position 192 of SEQ ID NO: 31, an arginine residue at a position corresponding to position 196 of SEQ ID NO: 31, an asparagine residue at a position corresponding to position 306 of SEQ ID NO: 31, an aspartate residue at a position corresponding to position 650 of SEQ ID NO: 31; and wherein the modified dipeptide cleavase comprises three or more amino acid substitutions in residues corresponding to positions 191, 192, 196, 306, 650 of SEQ ID NO: 31; and (iii) the modified dipeptide cleavases from the set of dipeptide cleavase enzymes have different specificities for the labeled terminal amino acids, which the modified dipeptide cleavases are configured to remove.

[0028] In an embodiment of the kit, (i) the chemical reagent is configured to attach a N-terminal modification to the terminal amino acid of the polypeptide; (ii) the N-terminal modification comprises a N-terminal blocking group (NTM blk ) and, optionally, a natural or unnatural amino acid portion (NTMaa); (iii) the NTMaa comprises a compound selected from the group consisting of: a naturally-occurring amino acid residue, 3-(3'-pyridyl)-L-alanine, L-cyclohexylglycine, α-aminoisobutyric acid, 3-(4'-pyridyl)-L-alanine, L-azetidine-2-carboxylic acid, isonipecotic acid, L-phenylglycine, β-(2-thienyl)-L-alanine, 3-(4-thiazolyl)-L-alanine, 1-aminocyclopentane-1-carboxylic acid, (2-trifluoromethyl)-L-Phenylalanine, L-cyclopropylalanine, 3-(2'-pyridyl)-L-alanine, beta-cyano-L-alanine, α-methyl-L-4-Fluorophenylalanine, α-methyl-D-4-fluorophenylalanine, 3-amino-2,2-difluoro-propionic acid, O-sulfo-L-tyrosine sodium salt, L-2-furylalanine, 1-aminocyclopropane-1-carboxylic acid, 3,5-dinitro-L-tyrosine, pentafluoro-L-phenylalanine, 3,5-difluoro-L-phenylalanine, 3-fluoro-L-phenylalanine, N-cyclopentylglycine, 1-(amino)cyclohexanecarboxylic acid, N-methylalanine, 4-amino-tetrahydropyran-4-carboxylic acid, 4-amino-1,1-dioxothiane-4-carboxylic acid, 4-amino-1-methyl-4-piperidinecarboxylic acid, 2-amino-N-(2,4-dimethoxybenzyl)acetamido)acetic acid, or N-alkylated derivatives; and (iv) the NTM blk comprises a compound selected from the group consisting of: 4-methylbenzoic acid, 4-(dimethylamio)benzoic acid, nicotinic acid, 3-aminonicotinic acid, 2-pyrazinecarbooxylic acid, 5-amino-2-fluoro-isonicotinic acid, 2,3-pyrazinedicarboxylic acid, 4,7-Difluoroisobenzofuran-1,3-dicarboxylic acid, 4-chloro-2-aminobenzoic acid, 4-nitro-2-aminobenzoic acid, 7-methoxy-1h-benzo[d][1,3]oxazine-2,4-dione, 4-carboxy-2-aminobenzoic acid, 6-(Trifluoromethyl)-2,4-dihydro-1h-3, 1-benzoxazine-2,4-dione, 7-(Trifluoromethyl)-1h-benzo[d][1,3]oxazine-2,4-dione, 6-fluoro-2-aminobenzoic acid, 4-fluoro-2-aminobenzoic acid, 5-methoxy-2-aminobenzoic acid, 4-fluorobenzoic acid, 4-(trifluoromethyl)benzoic acid, 2-ethynyl-6-fluorobenzaldehyde, 2-aminobenzoic acid, Succinic anhydride, 3,6-Difluoropyridine-2-carboxylic acid, 2-Fluoronicotinic acid, 5-Bromo-2-hydroxynicotinic acid, 4-(Trifluoromethyl)pyrimidine-5-carboxylic acid, 2-Oxo-1,2-dihydropyridine-3-carboxylic acid, 5-Methyl-2-aminobenzoic acid, 6-Fluoropicolinic acid, 3-Methyl-2-aminobenzoic acid, 4-Methyl-2-aminobenzoic acid, 2-Amino-6-methylbenzoic acid, 2-Amino-6-fluorobenzoic acid, 2-Amino-5-fluorobenzoic acid, 2-Amino-3-fluorobenzoic acid, 2-Amino-4-fluorobenzoic acid, 2-Aminonicotinic acid, 4-Aminonicotinic acid, 3-Aminopicolinic acid, 2-Amino-4,5-difluorobenzoic acid, 3,4-difluorobenzoic acid, 3,4,5-difluorobenzoic acid, 3-(Methoxycarbonyl)bicyclo[1.1.1]pentane-1-carboxylic acid, 3,3-Difluorocyclobutane-1-carboxylic acid, 1-Methyl-2-oxo-piperidine-4-carboxylic acid, Tetrahydropyran-4-carboxylic acid, 5-Fluoroorotic acid, 3-Fluoro-4-nitrobenzoic acid, 3-(Difluoromethyl)-1-methyl-1H-pyrazole-4-carboxylic acid, 4-(Difluoromethoxy)benzoic acid, 1-(Difluoromethyl)-1h-pyrazole-3-carboxylic acid, 4-(Methanesulfonylamino)benzoic acid, 5-Fluoro-6-methoxynicotinic acid, Tetrahydro-2H-thiopyran-4-carboxylic acid 1,1-dioxide, 4-(1H-Tetrazol-5-yl)benzoic acid, 1,2,3-Thiadiazole-4-carboxylic acid, 1,3-Benzodioxole-4-carboxylic acid, 2,1,3-Benzoxadiazole-5-carboxylic acid, 1-Benzyl-3-methyl-1h-pyrazole-5-carboxylic acid, 1-Cyclopropyl-6,7-difluoro-1,4-dihydro-4-oxoquinoline-3-carboxylic acid, 3,4-Dichlorobenzoic acid, 5-Fluoro-6-methylpyridine-2-carboxylic acid, 4,5-Dimethyl-2-(1h-pyrrol-1-yl)thiophene-3-carboxylic acid, 1,3-Dimethyl-1h-thieno[2,3-c]pyrazole-5-carboxylic acid, 1-[(4-Fluorobenzene)sulfonyl]piperidine-3-carboxylic acid, 1-(4-Fluorobenzyl)-5-oxopyrrolidine-3-carboxylic acid, 3-Fluoro-4-methoxybenzoic acid, 4-Fluoro-3-nitrobenzoic acid, 6-Fluoro-4-oxochromene-2-carboxylic acid, 3-Fluorophenylacetic acid, 4-Fluoro-3-(trifluoromethyl)benzoic acid, 5-Furan-2-yl-isoxazole-3-carboxylic acid, 1-Isopropyl-2-(trifluoromethyl)-1h-benzimidazole-5-carboxylic acid, Levofloxacin carboxylic acid, 3,5,7-Trifluoroadamantane-1-carboxylic acid, 3,4,5-Trimethoxybenzoic acid, 2-Oxo-2,3-dihydro-1h-benzo[d]imidazole-4-carboxylic acid, 1-Methyl-3-(trifluoromethyl)-1h-pyrazole-5-carboxylic acid, 2-Morpholin-4-yl-isonicotinic acid, 1,3-Oxazole-4-carboxylic acid, 4-Carboxybenzenesulfonamide, 3,4-difluorobenzenesulfonyl chloride.BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Accompanying figures are intended to support the invention; the figures are schematic and are not intended to be drawn to scale. For purposes of illustration, not every component is labeled in every figure, nor is every component of each embodiment of the invention shown where illustration is not necessary to allow those of ordinary skill in the art to understand the invention. FIG. 1 is a schematic depicting the removal of a single modified amino acid by exemplary modified dipeptide cleavases as provided herein. In FIG. 1 on the left, an exemplary unmodified dipeptide cleavase removes two amino acids as a dipeptide from the N-terminus of the polypeptide, cleaving the bond between the penultimate (P2) and antepenultimate amino acid (P3) residues. On the right, an exemplary modified dipeptide cleavase removes a labeled dipeptide including the terminal labeled amino acid from the N-terminus of the polypeptide, cleaving the bond between the penultimate terminal amino acid residue (P2) and the antepenultimate amino acid residue (P3). FIG. 2A-2C is a schematic depicting a cycle of terminal amino acid removal using the modified dipeptide cleavase and terminal amino acid labeling. In FIG. 2A-2B, a polypeptide with a labeled N-terminal amino acid residue is cleaved at the bond between the penultimate amino acid and antepenultimate amino acid by the modified dipeptide cleavase and the terminal dipeptide, including the label or modification (diamond), is released. In FIG. 2C, the new terminal amino acid is labeled and the modified penultimate cleavase is able to recognize the new labeled terminal amino acid for further cleavage and release of the terminal dipeptide following the next cleavage step. FIG. 3. depicts N-terminal amino acid (NTAA) conversion efficiency with different exemplary reagents for labeling the N-terminal amino acid. Two different peptides were tested in solution: N-Terminal G (NT-G) = GRFSGIY (SEQ ID NO:29); N-Terminal W (NT-W) = WTQIFGA (SEQ ID NO:30). LC-MS was used to quantitate conversion efficiency. FIG. 4A-4C depicts results from a WebLogo analysis of sequence conservation of DAP BII homologs with 60% sequence similarity or identity. The height of each stack indicates the sequence conservation at that position (measured in bits), and the height of symbols within the stack reflects the relative frequency of the corresponding amino acid at the indicated position (in reference to SEQ ID NO: 20). FIG. 5 shows a graph of Michaelis-Menten kinetics for two modified dipeptide cleavases (containing the amino acid sequences as set forth in SEQ ID NO: 18 and SEQ ID NO: 27) tested at various concentrations. Fig. 6 depicts a model of an exemplary anticalin scaffold bound with N-terminal modified amino acid. The modification is shown in orange spheres the preferably occupy part of a surface accessible pocket. The P1 sidechain (i.e., Leucine in magenta) is surrounded by amino acids (shown in blue stick) that can be mutated to provide specificity. Fig. 7A-7H illustrates exemplary Luminex-based binding affinity profile of anticalin clones chosen from a phage display screen against M15-L-P1 peptides. Eight exemplary engineered anticalin binders are shown to have mostly mono-specificity for P1 residues except for the I / L binder. The anticalin clones are isolated from phage library panning. Clones with specificity to different P1 residues, such as E, F, G, H, I, L, P, W, as well as clones with specificity to two different P1 residues, such as T / S, A / T / S, T / V / I / A, F / L, were successfully isolated. Fig. 8A-Billustrates exemplary analysis of P2 dependence via ProteoCode ™< encoding assay. Fig. 8A shows Encoding versus Luminex binding signal for M15-L-G clone shown in Fig. 4. Fig. 8B shows P2 dependence determined by ProteoCode ™< encoding assay using the M15-L-G binder clone on various M15-L-G-P2 peptides. Fig. 9 shows ProteoCode ™< encoding assay with modified NTAA binders and modified cleavases used for high-throughput polypeptide sequencing. Polypeptide molecules are each labeled with a DNA recording tag and attached to a solid support (beads) at a low molecular density, a sparsity that permits only intramolecular information transfer to occur. (1) At the beginning of a sequencing cycle, the polypeptide N-terminal amino acid (NTAA) is functionalized with a N-terminal modification (NTM) or label. (2) Next, an engineered NTAA binding agent labelled with a DNA coding tag binds to the labeled NTAA residue. After binding and washing, the coding tag information is transferred enzymatically to the recording tag (by extension or ligation). (3) Removal of the NTM-labeled N-terminal residue is accomplished by using a modified Cleavase enzyme that specifically cleaves the NTM-labeled N-terminal residue. After n cycles, a DNA library element representing the n amino acids of the polypeptide sequence is formed as a part of extended recording tag and can be sequenced by a next-generation sequencing (NGS) method. A representative structure of an NGS library element after 7 cycles is shown. Fig. 10A illustrates exemplary cleavage of M15-L-modified NTAAs of a model polypeptide (M15-L-P1-AR) with Cleavase enzymes. A compilation of seven different modified Cleavase clones was used to generate the spectrum of cleavage profile across the M15-L-modified NTAAs as shown. Data were generated by HPLC analysis (UV absorbance) of cleaved versus intact peptides after cleavase assay. Fig. 10B shows the same cleavage events using SDS-PAGE analysis. Fig. 10C shows a cleavage profile for an exemplary set of two selected modified Cleavase clones, M15-L_Z001, having specificity towards A, I, L, M, Q, V in the P1 position (cleavage efficiency of M15-L_Z001 is shown by the left columns for each amino acid), and M15-L_Z002, having specificity towards D and E in the P1 position (cleavage efficiency of M15-L_Z002 is shown by the right columns for each amino acid). Fig. 11 illustrates cleavage of an exemplary polypeptide by unmodified dipeptide cleavase (dipeptidyl aminopeptidase DAP BII, SEQ ID NO: 13). 1-5 correspond to cleavage results at the following time points: 0 min, 5 min, 30 min, 45 min, 60 min. Fig. 12A-B. Exemplary N-terminal modifications (NTMs) to enable NTM-NTAA cleavage at P1 residue by modified dipeptide cleavases. Fig. 12A. Structures of a bipartite NTM comprised of an amino acid-like portion (NTMaa) and a N-terminal blocking group (NTM blk ) connected by an amide bond (upper) and other possible NTMs that would accommodate modified substrate binding pockets of cleavases. NTM can also be a small chemical entity (NTM B ) with a similar bipartite shape configuration as NTM A , or a differently shaped NTM C . Fig. 12B. NTMs are activated using standard methods (activated ester) and are coupled to the N-terminal amine on the P1 residue of a polypeptide. The arrow indicates a cleavage site of the modified dipeptide cleavase enzyme. Fig. 13. The cleavage efficiency of the NTM-labeled NTAA of a target polypeptide depends on a particular NTM. Fig. 14A-B. Time course of cleavage reactions of two labeled peptides (M15-LAAR and M19-LAAR, cleavage efficiencies are shown in left and right columns, respectively, for each time point) by two modified dipeptidyl cleavases selected using M15 NTM ( Fig. 14A) and M19 NTM ( Fig. 14B). Fig. 15 depicts results from a WebLogo analysis of sequence conservation of DAP BII homologs. The height of each stack indicates the sequence conservation at that position (measured in bits), and the height of symbols within the stack reflects the relative frequency of the corresponding amino acid at the indicated position (in reference to positions of SEQ ID NO: 20). Conservation of N215, W216(F), R220, N330, D674 positions is highlighted. Fig. 16. Similarity distribution relative to DapBII (Pseudoxanthomonas mexicana) of 2125 sequences clustered at 80% sequence identity. DETAILED DESCRIPTION

[0030] Provided herein are modified dipeptide cleavases comprising a mutation (e.g., one or more modifications in an unmodified dipeptide cleavase) and related methods of selecting, engineering, and using the modified dipeptide cleavases. Also provided are kits comprising the modified dipeptide cleavases. In some embodiments, the kits comprising the modified dipeptide cleavase is used for treating peptides, polypeptides, and proteins, such as for sequencing and / or analysis. In some embodiments, protein analysis using the modified dipeptide cleavase employs barcoding and nucleic acid encoding of molecular recognition events, and / or detectable labels. In some embodiments, the kits also include other components for treating the polypeptides, including tags (e.g., DNA tag or DNA recording tag), solid supports, and other reagents for preparing the polypeptides and other reagents for polypeptide analysis.

[0031] Various enzymes that degrade peptides and proteins by hydrolyzing peptide bonds, (e.g., aminopeptidases, dipeptidyl peptidases, carboxypeptidases, endopeptidases) have been isolated and discovered in a number of organisms and from various tissues. However, natural aminopeptidases may have limited specificity, and generically eliminate N-terminal amino acids in a processive manner, eliminating one amino acid off after another. Some substrate-specific peptidases specifically remove one or two amino acid residues at a time from the amino-terminus or carboxy-terminus of peptides.

[0032] In some embodiments, methods for peptide degradation are useful for applications in protein analysis and / or sequencing. For example, peptide sequencing may involve Edman degradation to achieve stepwise degradation of the N-terminal amino acid on a peptide through a series of chemical modifications and downstream HPLC analysis or mass spectrometry analysis. However, in general, Edman degradation peptide sequencing may be limited, for example, typical Edman degradation requires deployment of high temperature and harsh chemical conditions (e.g., strong acids; anhydrous TFA) for long incubation times. In some cases, Edman degradation may not be compatible with processes for protein analysis methods which may be sensitive to harsh chemical conditions, such as analysis methods which employ nucleic acids (e.g., DNA).

[0033] Accordingly, there remains a need for improved reagents and techniques for degradation of amino acids from a polypeptide. For example, enzymatic methods for removing, eliminating, or cleaving amino acids from polypeptides may be desired. Provided herein are modified dipeptide cleavases that meet such needs. In some embodiments, provided herein are enzymatic methods and reagents for removing amino acids as dipeptides or as single labeled amino acids from polypeptides. In some cases, the removal of amino acids by the provided modified enzymes (e.g., dipeptide cleavases) are used for stepwise degradation of amino acids from polypeptides. In some embodiments, the removal of amino acids by the provided modified dipeptide cleavase are suitable for cyclic removal of dipeptides or single amino acids from the polypeptide. In some embodiments, the modified dipeptide cleavase removes or is configured to remove a labeled terminal dipeptide from a polypeptide. In some embodiments, the modified dipeptide cleavase removes or is configured to remove a single N-terminally modified amino acid from a target polypeptide. In some embodiments, the modified dipeptide cleavase removes a labeled dipeptide (the terminal and penultimate terminal amino acids) from the C-terminus or N-terminus of a polypeptide. In some embodiments, the modified dipeptide cleavase is derived from a wild-type or unmodified dipeptide cleavase. For example, the unmodified dipeptide cleavase is a protein classified in EC 3.4.14, EC 3.4.15, MEROPS S9, MEROPS S46, MEROPS M49, or a functional homolog or fragment thereof. For example, the modified dipeptide cleavase is derived from a wild-type or unmodified dipeptide cleavase (e.g., a dipeptidyl peptidase, a dipeptidyl aminopeptidase, a peptidyl-dipeptidase, or a dipeptidyl carboxypeptidase).

[0034] The present disclosure also relates to a binder that specifically binds to an N-terminally modified polypeptide and modified or an engineered cleavase that removes or is configured to remove a single N-terminally modified amino acid from a polypeptide. Also provided herein is a method and related kits for treating a polypeptide using or comprising the binder and / or modified cleavase.

[0035] In some embodiments, peptidases may be engineered to possess specific binding or catalytic activity to specific terminal amino acids only when modified with a label. For example, a cleavase may be engineered or modified, compared to a wild-type or unmodified dipeptide cleavase, such than it only eliminates a terminal amino acid if it is labeled by a chemical label. Using this exemplary approach, the modified dipeptide cleavase eliminates only terminal single labeled amino acids or dipeptides containing a labeled amino acid from the terminus of the polypeptide, and allows control of degradation in a desired manner. In some embodiments, the modified dipeptide cleavase is configured to remove a labeled terminal dipeptide (including the terminal and penultimate terminal amino acids) from the C-terminus or N-terminus of a polypeptide. In some embodiments, the modified dipeptide cleavase is non-selective as to amino acid residue identity while being selective for the label (e.g., will remove any single labeled amino acids or terminal dipeptides containing any two amino acids associated with a label or modification). In some other embodiments, the modified dipeptide cleavase exhibits some preference for certain amino acid residues or classes of amino acids (e.g. at the P1 and / or P2 terminal positions of the polypeptide). In some cases, two or more modified dipeptide cleavases with different preferences for certain amino acids (or classes of amino acids) may be used in combination. In some embodiments, the modified dipeptide cleavase binds and removes single labeled amino acids dipeptides from the N-terminus of the polypeptide. In some embodiments, the modified dipeptide cleavase binds and removes dipeptides from the C-terminus of the polypeptide.

[0036] In some embodiments, known peptidases may be modified to achieve specific characteristics for binding and / or cleaving. An example of a model of modifying the specificity of enzymatic N-terminal amino acid (NTAA) degradation involves a methionine aminopeptidase converted into a leucine aminopeptidase (Borgo et al., Protein Sci. (2014) 23(3):312-320). In another example, aminopeptidase mutants were engineered to bind to and eliminate individual or small groups of labelled (biotinylated) NTAAs (see, PCT Publication No. WO2010 / 065322). Provided herein are modified dipeptide cleavases which are selected or modified to remove terminal dipeptides that are labeled, such as dipeptides containing a chemically-modified terminal amino acid on a polypeptide. In some embodiments, a wild-type cleavase is engineered (e.g., using structural-function based-design and / or directed evolution) to cleave or remove only a terminal dipeptide containing an N-terminal amino acid having a chemical group present as the label (e.g., PTC / DNP / acetyl / Cbz).

[0037] The unmodified dipeptide cleavase may be from any suitable organism. In some examples, the wild-type or unmodified dipeptide cleavase is from a mammal, e.g., Homo sapiens, a fungus or yeast, e.g., Saccharomyces cerevisiae, or a bacterium, e.g., Bacteroides thetaiotaomicron, Porphyromonas gingivalis, Pseudomonas sp., Pseudoxanthomonas mexicana or Caldithrix abyssi. In some cases, these enzymes are stable, robust, and active at room temperature and at or around pH 8.0, and thus compatible with mild conditions preferred for peptide analysis. In some embodiments, it is preferred to have a thermophilic cleavase capable of removing labeled single amino acids or terminal dipeptides at elevated temperatures to minimize peptide secondary structure.

[0038] In another embodiment, cyclic elimination or removal of amino acids is attained by engineering the dipeptide cleavase to be active only in the presence of a terminal amino acid label. In some embodiments, the label is a chemical label. Moreover, the dipeptide cleavase may be engineered to be non-specific, such that it does not selectively recognize particular amino acids over another, but recognizes any amino acid at the terminus that has a label. In some embodiments, the modified dipeptide cleavase is selective for one or more, two or more, three or more, four or more, five or more, ten or more, fifteen or more, twenty or more etc. amino acids.

[0039] In some embodiments, the provided modified dipeptide cleavases are used for treating polypeptides obtained from a sample. In some cases, the sample and / or the polypeptide obtained from the sample is treated with other reagents for processing the polypeptides, such as digesting the polypeptides. In some aspects, the polypeptides are treated with a reagent for labeling the terminal amino acid (e.g., a chemical reagent) prior to treating the polypeptide with the modified dipeptide cleavase provided. In some embodiments, the polypeptides comprise a plurality of polypeptides obtained from a sample. In some embodiments, the sample is obtained from a subject.

[0040] In certain embodiments, the modified dipeptide cleavase is derived from a metallopeptidase, a zinc-dependent metallopeptidase, or a zinc-dependent hydrolase. In some cases, the modified dipeptide cleavase is a metallo-peptidase and requires a metal ion for activation. In some embodiments, the use of a monomeric metallo-aminopeptidase provides the advantage of having controllable activity can be turned on / off at will by adding or removing the appropriate metal cation.

[0041] Numerous specific details are set forth in the following description in order to provide a thorough understanding of the present disclosure. These details are provided for the purpose of example and the claimed subject matter may be practiced according to the claims without some or all of these specific details. It is to be understood that other embodiments can be used and structural changes can be made without departing from the scope of the claimed subject matter. It should be understood that the various features and functionality described in one or more of the individual embodiments are not limited in their applicability to the particular embodiment with which they are described. They instead can, be applied, alone or in some combination, to one or more of the other embodiments of the disclosure, whether or not such embodiments are described, and whether or not such features are presented as being a part of a described embodiment. For the purpose of clarity, technical material that is known in the technical fields related to the claimed subject matter has not been described in detail so that the claimed subject matter is not unnecessarily obscured.

[0042] Citation of the publications or documents is not intended as an admission that any of them is pertinent prior art, nor does it constitute any admission as to the contents or date of these publications or documents.

[0043] All headings are for the convenience of the reader and should not be used to limit the meaning of the text that follows the heading, unless so specified.DEFINITIONS

[0044] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as is commonly understood by one of ordinary skill in the art to which the present disclosure belongs. If a definition set forth in this section is contrary to or otherwise inconsistent with a definition set forth in the patents, applications, published applications and other publications that are herein cited, the definition set forth in this section prevails over the definition of the cited reference.

[0045] As used herein, the singular forms "a," "an" and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a peptide" includes one or more peptides, or mixtures of peptides. Also, and unless specifically stated or obvious from context, as used herein, the term "or" is understood to be inclusive and covers both "or" and "and".

[0046] The term "about" as used herein refers to the usual error range for the respective value readily known to the skilled person in this technical field. Reference to "about" a value or parameter herein includes (and describes) embodiments that are directed to that value or parameter per se. For example, description referring to "about X" includes description of "X.

[0047] The term "antibody" herein is used in the broadest sense and includes polyclonal and monoclonal antibodies, including intact antibodies and functional (antigen-binding) antibody fragments, including fragment antigen binding (Fab) fragments, F(ab') 2 fragments, Fab' fragments, Fv fragments, recombinant IgG (rIgG) fragments, single chain antibody fragments, including single chain variable fragments (scFv), and single domain antibodies (e.g., sdAb, sdFv, nanobody) fragments. The term encompasses genetically engineered and / or otherwise modified forms of immunoglobulins, such as intrabodies, peptibodies, chimeric antibodies, fully human antibodies, humanized antibodies, and heteroconjugate antibodies, multispecific, e.g., bispecific, antibodies, diabodies, triabodies, and tetrabodies, tandem di-scFv, tandem tri-scFv. Unless otherwise stated, the term "antibody" should be understood to encompass functional antibody fragments thereof. The term also encompasses intact or full-length antibodies, including antibodies of any class or sub-class, including IgG and sub-classes thereof, IgM, IgE, IgA, and IgD.

[0048] An "individual" or "subject" includes a mammal. Mammals include, but are not limited to, domesticated animals (e.g., cows, sheep, cats, dogs, and horses), primates (e.g., humans and non-human primates such as monkeys), rabbits, and rodents (e.g., mice and rats). An "individual" or "subject" may include birds such as chickens, vertebrates such as fish and mammals such as mice, rats, rabbits, cats, dogs, pigs, cows, ox, sheep, goats, horses, monkeys and other non-human primates. In certain embodiments, the individual or subject is a human.

[0049] As used herein, the term "sample" refers to anything which may contain an analyte for which an analyte assay is desired. As used herein, a "sample" can be a solution, a suspension, liquid, powder, a paste, aqueous, non-aqueous or any combination thereof. The sample may be a biological sample, such as a biological fluid or a biological tissue. Examples of biological fluids include urine, blood, plasma, serum, saliva, semen, stool, sputum, cerebral spinal fluid, tears, mucus, amniotic fluid or the like. Biological tissues are aggregate of cells, usually of a particular kind together with their intercellular substance that form one of the structural materials of a human, animal, plant, bacterial, fungal or viral structure, including connective, epithelium, muscle and nerve tissues. Examples of biological tissues also include organs, tumors, lymph nodes, arteries and individual cell(s).

[0050] In some embodiments, the sample is a biological sample. A biological sample of the present disclosure encompasses a sample in the form of a solution, a suspension, a liquid, a powder, a paste, an aqueous sample, or a non-aqueous sample. As used herein, a "biological sample" includes any sample obtained from a living or viral (or prion) source or other source of macromolecules and biomolecules, and includes any cell type or tissue of a subject from which nucleic acid, protein and / or other macromolecule can be obtained. The biological sample can be a sample obtained directly from a biological source or a sample that is processed. For example, isolated nucleic acids that are amplified constitute a biological sample. Biological samples include, but are not limited to, body fluids, such as blood, plasma, serum, cerebrospinal fluid, synovial fluid, urine and sweat, tissue and organ samples from animals and plants and processed samples derived therefrom. In some embodiments, the sample can be derived from a tissue or a body fluid, for example, a connective, epithelium, muscle or nerve tissue; a tissue selected from the group consisting of brain, lung, liver, spleen, bone marrow, thymus, heart, lymph, blood, bone, cartilage, pancreas, kidney, gall bladder, stomach, intestine, testis, ovary, uterus, rectum, nervous system, gland, and internal blood vessels; or a body fluid selected from the group consisting of blood, urine, saliva, bone marrow, sperm, an ascitic fluid, and subfractions thereof, e.g., serum or plasma.

[0051] The terms "level" or "levels" are used to refer to the presence and / or amount of a target, e.g., a substance or an organism that is part of the etiology of a disease or disorder, and can be determined qualitatively or quantitatively. A "qualitative" change in the target level refers to the appearance or disappearance of a target that is not detectable or is present in samples obtained from normal controls. A "quantitative" change in the levels of one or more targets refers to a measurable increase or decrease in the target levels when compared to a healthy control.

[0052] As used herein, the term "polypeptide" encompasses peptides and proteins, and refers to a molecule comprising a chain of two or more amino acids joined by peptide bonds. In some embodiments, a polypeptide comprises 2 to 50 amino acids, e.g., having more than 20-30 amino acids. In some embodiments, a peptide does not comprise a secondary, tertiary, or higher structure. In some embodiments, the polypeptide is a protein. In some embodiments, a protein comprises 30 or more amino acids, e.g. having more than 50 amino acids. In some embodiments, in addition to a primary structure, a protein comprises a secondary, tertiary, or higher structure. The amino acids of the polypeptides are most typically L-amino acids, but may also be D-amino acids, modified amino acids, amino acid analogs, amino acid mimetics, or any combination thereof. Polypeptides may be naturally occurring, synthetically produced, or recombinantly expressed. Polypeptides may be synthetically produced, isolated, recombinantly expressed, or be produced by a combination of methodologies as described above. Polypeptides may also comprise additional groups modifying the amino acid chain, for example, functional groups added via post-translational modification. The polymer may be linear or branched, it may comprise modified amino acids, and it may be interrupted by non-amino acids. The term also encompasses an amino acid polymer that has been modified naturally or by intervention; for example, disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation or modification, such as conjugation with a labeling component.

[0053] As used herein, the term "amino acid" refers to an organic compound comprising an amine group, a carboxylic acid group, and a side-chain specific to each amino acid, which serve as a monomeric subunit of a peptide. An amino acid includes the 20 standard, naturally occurring or canonical amino acids as well as non-standard amino acids. The standard, naturally-occurring amino acids include Alanine (A or Ala), Cysteine (C or Cys), Aspartic Acid (D or Asp), Glutamic Acid (E or Glu), Phenylalanine (F or Phe), Glycine (G or Gly), Histidine (H or His), Isoleucine (I or Ile), Lysine (K or Lys), Leucine (L or Leu), Methionine (M or Met), Asparagine (N or Asn), Proline (P or Pro), Glutamine (Q or Gln), Arginine (R or Arg), Serine (S or Ser), Threonine (T or Thr), Valine (V or Val), Tryptophan (W or Trp), and Tyrosine (Y or Tyr). An amino acid may be an L-amino acid or a D-amino acid. Non-standard amino acids may be modified amino acids, amino acid analogs, amino acid mimetics, non-standard proteinogenic amino acids, or non-proteinogenic amino acids that occur naturally or are chemically synthesized. Examples of non-standard amino acids include, but are not limited to, selenocysteine, pyrrolysine, and N-formylmethionine, β-amino acids, Homo-amino acids, Proline and Pyruvic acid derivatives, 3-substituted alanine derivatives, glycine derivatives, ring-substituted phenylalanine and tyrosine derivatives, linear core amino acids, N-methyl amino acids.

[0054] As used herein, the term "post-translational modification" refers to modifications that occur on a peptide after its translation, e.g., translation by ribosomes, is complete. A post-translational modification may be a covalent chemical modification or enzymatic modification. Examples of post-translation modifications include, but are not limited to, acylation, acetylation, alkylation (including methylation), biotinylation, butyrylation, carbamylation, carbonylation, deamidation, deiminiation, diphthamide formation, disulfide bridge formation, eliminylation, flavin attachment, formylation, gamma-carboxylation, glutamylation, glycylation, glycosylation, glypiation, heme C attachment, hydroxylation, hypusine formation, iodination, isoprenylation, lipidation, lipoylation, malonylation, methylation, myristolylation, oxidation, palmitoylation, pegylation, phosphopantetheinylation, phosphorylation, prenylation, propionylation, retinylidene Schiff base formation, S-glutathionylation, S-nitrosylation, S-sulfenylation, selenation, succinylation, sulfination, ubiquitination, and C-terminal amidation. A post-translational modification includes modifications of the amino terminus and / or the carboxyl terminus of a peptide. Modifications of the terminal amino group include, but are not limited to, des-amino, N-lower alkyl, N-di-lower alkyl, and N-acyl modifications. Modifications of the terminal carboxy group include, but are not limited to, amide, lower alkyl amide, dialkyl amide, and lower alkyl ester modifications (e.g., wherein lower alkyl is C 1 -C 4 alkyl). A post-translational modification also includes modifications, such as but not limited to those described above, of amino acids falling between the amino and carboxy termini. The term post-translational modification can also include peptide modifications that include one or more detectable labels.

[0055] As used herein, the term "binding agent" or "binder" refers to a nucleic acid molecule, a peptide, a polypeptide, a protein, carbohydrate, or a small molecule that binds to, associates, unites with, recognizes, or combines with a binding target, e.g., a polypeptide or a component or feature of a polypeptide. A binding agent may form a covalent association or non-covalent association with the polypeptide or component or feature of a polypeptide. A binding agent may also be a chimeric binding agent, composed of two or more types of molecules, such as a nucleic acid molecule-peptide chimeric binding agent or a carbohydrate-peptide chimeric binding agent. A binding agent may be a naturally occurring, synthetically produced, or recombinantly expressed molecule. A binding agent may bind to a single monomer or subunit of a polypeptide (e.g., a single amino acid of a polypeptide) or bind to a plurality of linked subunits of a polypeptide (e.g., a di-peptide, tri-peptide, or higher order peptide of a longer peptide, polypeptide, or protein molecule). A binding agent may bind to a linear molecule or a molecule having a three-dimensional structure (also referred to as conformation). For example, an antibody binding agent may bind to linear peptide, polypeptide, or protein, or bind to a conformational peptide, polypeptide, or protein. A binding agent may bind to an N-terminal peptide, a C-terminal peptide, or an intervening peptide of a peptide, polypeptide, or protein molecule. A binding agent may bind to an N-terminal amino acid, C-terminal amino acid, or an intervening amino acid of a peptide molecule. A binding agent may preferably bind to a chemically modified or labeled amino acid (e.g., an amino acid that has been labeled by a reagent comprising a compound of any one of Formula (I)-(IV) as described herein) over a non-modified or unlabeled amino acid. For example, a binding agent may preferably bind to an amino acid that has been labeled or modified over an amino acid that is unlabeled or unmodified. A binding agent may bind to a post-translational modification of a peptide molecule. A binding agent may exhibit selective binding to a component or feature of a polypeptide (e.g., a binding agent may selectively bind to one of the 20 possible natural amino acid residues and with bind with very low affinity or not at all to the other 19 natural amino acid residues). A binding agent may exhibit less selective binding, where the binding agent is capable of binding or configured to bind to a plurality of components or features of a polypeptide (e.g., a binding agent may bind with similar affinity to two or more different amino acid residues). A binding agent may comprise a coding tag, which may be joined to the binding agent by a linker.

[0056] As used herein, the term "linker" refers to one or more of a nucleotide, a nucleotide analog, an amino acid, a peptide, a polypeptide, a polymer, or a non-nucleotide chemical moiety that is used to join two molecules. A linker may be used to join a binding agent with a coding tag, a recording tag with a polypeptide, a polypeptide with a solid support, a recording tag with a solid support, etc. In certain embodiments, a linker joins two molecules via enzymatic reaction or chemistry reaction (e.g., click chemistry).

[0057] The term "ligand" as used herein refers to any molecule or moiety connected to the compounds described herein. "Ligand" may refer to one or more ligands attached to a compound. In some embodiments, the ligand is a pendant group or binding site (e.g., the site to which the binding agent binds).

[0058] As used herein, the term "proteome" can include the entire set of proteins, polypeptides, or peptides (including conjugates or complexes thereof) expressed by a genome, cell, tissue, or organism at a certain time, of any organism. In one aspect, it is the set of expressed proteins in a given type of cell or organism, at a given time, under defined conditions. Proteomics is the study of the proteome. For example, a "cellular proteome" may include the collection of proteins found in a particular cell type under a particular set of environmental conditions, such as exposure to hormone stimulation. An organism's complete proteome may include the complete set of proteins from all of the various cellular proteomes. A proteome may also include the collection of proteins in certain sub-cellular biological systems. For example, all of the proteins in a virus can be called a viral proteome. As used herein, the term "proteome" include subsets of a proteome, including but not limited to a kinome; a secretome; a receptome (e.g., GPCRome); an immunoproteome; a nutriproteome; a proteome subset defined by a post-translational modification (e.g., phosphorylation, ubiquitination, methylation, acetylation, glycosylation, oxidation, lipidation, and / or nitrosylation), such as a phosphoproteome (e.g., phosphotyrosine-proteome, tyrosine-kinome, and tyrosine-phosphatome), a glycoproteome, etc.; a proteome subset associated with a tissue or organ, a developmental stage, or a physiological or pathological condition; a proteome subset associated a cellular process, such as cell cycle, differentiation (or de-differentiation), cell death, senescence, cell migration, transformation, or metastasis; or any combination thereof. As used herein, the term "proteomics" refers to quantitative analysis of the proteome within cells, tissues, and bodily fluids, and the corresponding spatial distribution of the proteome within the cell and within tissues. Additionally, proteomics studies include the dynamic state of the proteome, continually changing in time as a function of biology and defined biological or chemical stimuli.

[0059] The terminal amino acid at one end of a peptide or polypeptide chain that has a free amino group is referred to herein as the "N-terminal amino acid" (NTAA). The terminal amino acid at the other end of the chain that has a free carboxyl group is referred to herein as the "C-terminal amino acid" (CTAA). The amino acids making up a peptide may be numbered in order, with the peptide being "n" amino acids in length. As used herein, NTAA is considered the n th< amino acid (also referred to herein as the "n NTAA"). Using this nomenclature, the next amino acid is the n-1 amino acid, then the n-2 amino acid, and so on down the length of the peptide from the N-terminal end to C-terminal end. In certain embodiments, an NTAA, CTAA, or both may be modified or labeled with a moiety or a chemical moiety.

[0060] As used herein, the term "barcode" refers to a nucleic acid molecule of about 2 to about 30 bases (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 bases) providing a unique identifier tag or origin information for a polypeptide, a binding agent, a set of binding agents from a binding cycle, a sample polypeptides, a set of samples, polypeptides within a compartment (e.g., droplet, bead, or separated location), polypeptides within a set of compartments, a fraction of polypeptides, a set of polypeptide fractions, a spatial region or set of spatial regions, a library of polypeptides, or a library of binding agents. A barcode can be an artificial sequence or a naturally occurring sequence. In certain embodiments, each barcode within a population of barcodes is different. In other embodiments, a portion of barcodes in a population of barcodes is different, e.g., at least about 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, or 99% of the barcodes in a population of barcodes is different. A population of barcodes may be randomly generated or non-randomly generated. In certain embodiments, a population of barcodes are error correcting barcodes. Barcodes can be used to computationally deconvolute the multiplexed sequencing data and identify sequence reads derived from an individual polypeptide, sample, library, etc. A barcode can also be used for deconvolution of a collection of polypeptides that have been distributed into small compartments for enhanced mapping. For example, rather than mapping a peptide back to the proteome, the peptide is mapped back to its originating protein molecule or protein complex.

[0061] As used herein, the term "coding tag" refers to a polynucleotide with any suitable length, e.g., a nucleic acid molecule of about 2 bases to about 100 bases, including any integer including 2 and 100 and in between, that comprises identifying information for its associated binding agent. A "coding tag" may also be made from a "sequenceable polymer" (see, e.g., Niu et al., 2013, Nat. Chem. 5:282-292; Roy et al., 2015, Nat. Commun. 6:7237; Lutz, 2015, Macromolecules 48:4759-4767). A coding tag may comprise an encoder sequence, which is optionally flanked by one spacer on one side or optionally flanked by a spacer on each side. A coding tag may also be comprised of an optional UMI and / or an optional binding cycle-specific barcode. A coding tag may be single stranded or double stranded. A double stranded coding tag may comprise blunt ends, overhanging ends, or both. A coding tag may refer to the coding tag that is directly attached to a binding agent, to a complementary sequence hybridized to the coding tag directly attached to a binding agent (e.g., for double stranded coding tags), or to coding tag information present in an extended recording tag. In certain embodiments, a coding tag may further comprise a binding cycle specific spacer or barcode, a unique molecular identifier, a universal priming site, or any combination thereof.

[0062] As used herein, the term "spacer" (Sp) refers to a nucleic acid molecule of about 1 base to about 20 bases (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 bases) in length that is present on a terminus of a recording tag or coding tag. In certain embodiments, a spacer sequence flanks an encoder sequence of a coding tag on one end or both ends. Following binding of a binding agent to a polypeptide, annealing between complementary spacer sequences on their associated coding tag and recording tag, respectively, allows transfer of binding information through a primer extension reaction or ligation to the recording tag, coding tag, or a di-tag construct. Sp' refers to spacer sequence complementary to Sp. Preferably, spacer sequences within a library of binding agents possess the same number of bases. A common (shared or identical) spacer may be used in a library of binding agents. A spacer sequence may have a "cycle specific" sequence in order to track binding agents used in a particular binding cycle. The spacer sequence (Sp) can be constant across all binding cycles, be specific for a particular class of polypeptides, or be binding cycle number specific. Polypeptide class-specific spacers permit annealing of a cognate binding agent's coding tag information present in an extended recording tag from a completed binding / extension cycle to the coding tag of another binding agent recognizing the same class of polypeptides in a subsequent binding cycle via the class-specific spacers. Only the sequential binding of correct cognate pairs results in interacting spacer elements and effective primer extension. A spacer sequence may comprise sufficient number of bases to anneal to a complementary spacer sequence in a recording tag to initiate a primer extension (also referred to as polymerase extension) reaction, or provide a "splint" for a ligation reaction, or mediate a "sticky end" ligation reaction. A spacer sequence may comprise a fewer number of bases than the encoder sequence within a coding tag.

[0063] As used herein, the term "recording tag" refers to a moiety, e.g., a chemical coupling moiety, a nucleic acid molecule, or a sequenceable polymer molecule (see, e.g., Niu et al., 2013, Nat. Chem. 5:282-292; Roy et al., 2015, Nat. Commun. 6:7237; Lutz, 2015, Macromolecules 48:4759-4767) to which identifying information of a coding tag can be transferred, or from which identifying information about the macromolecule (e.g., UMI information) associated with the recording tag can be transferred to the coding tag. Identifying information can comprise any information characterizing a molecule such as information pertaining to sample, fraction, partition, spatial location, interacting neighboring molecule(s), cycle number, etc. Additionally, the presence of UMI information can also be classified as identifying information. In certain embodiments, after a binding agent binds to a polypeptide, information from a coding tag linked to a binding agent can be transferred to the recording tag associated with the polypeptide while the binding agent is bound to the polypeptide. In other embodiments, after a binding agent binds to a polypeptide, information from a recording tag associated with the polypeptide can be transferred to the coding tag linked to the binding agent while the binding agent is bound to the polypeptide. A recoding tag may be directly linked to a polypeptide, linked to a polypeptide via a multifunctional linker, or associated with a polypeptide by virtue of its proximity (or co-localization) on a solid support. A recording tag may be linked via its 5' end or 3' end or at an internal site, as long as the linkage is compatible with the method used to transfer coding tag information to the recording tag or vice versa. A recording tag may further comprise other functional components, e.g., a universal priming site, unique molecular identifier, a barcode (e.g., a sample barcode, a fraction barcode, spatial barcode, a compartment tag, etc.), a spacer sequence that is complementary to a spacer sequence of a coding tag, or any combination thereof. The spacer sequence of a recording tag is preferably at the 3'-end of the recording tag in embodiments where polymerase extension is used to transfer coding tag information to the recording tag.

[0064] As used herein, the term "primer extension", also referred to as "polymerase extension", refers to a reaction catalyzed by a nucleic acid polymerase (e.g., DNA polymerase) whereby a nucleic acid molecule (e.g., oligonucleotide primer, spacer sequence) that anneals to a complementary strand is extended by the polymerase, using the complementary strand as template.

[0065] As used herein, the term "unique molecular identifier" or "UMI" refers to a nucleic acid molecule of about 3 to about 40 bases (3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 bases) in length providing a unique identifier tag for each macromolecule, polypeptide or binding agent to which the UMI is linked. A polypeptide UMI can be used to computationally deconvolute sequencing data from a plurality of extended recording tags to identify extended recording tags that originated from an individual polypeptide. A polypeptide UMI can be used to accurately count originating polypeptide molecules by collapsing NGS reads to unique UMIs. A binding agent UMI can be used to identify each individual molecular binding agent that binds to a particular polypeptide. For example, a UMI can be used to identify the number of individual binding events for a binding agent specific for a single amino acid that occurs for a particular peptide molecule. It is understood that when UMI and barcode are both referenced in the context of a binding agent or polypeptide, that the barcode refers to identifying information other that the UMI for the individual binding agent or polypeptide (e.g., sample barcode, compartment barcode, binding cycle barcode).

[0066] As used herein, the term "universal priming site" or "universal primer" or "universal priming sequence" refers to a nucleic acid molecule, which may be used for library amplification and / or for sequencing reactions. A universal priming site may include, but is not limited to, a priming site (primer sequence) for PCR amplification, flow cell adaptor sequences that anneal to complementary oligonucleotides on flow cell surfaces enabling bridge amplification in some next generation sequencing platforms, a sequencing priming site, or a combination thereof. Universal priming sites can be used for other types of amplification, including those commonly used in conjunction with next generation digital sequencing. For example, extended recording tag molecules may be circularized and a universal priming site used for rolling circle amplification to form DNA nanoballs that can be used as sequencing templates (Drmanac et al., 2009, Science 327:78-81). Alternatively, recording tag molecules may be circularized and sequenced directly by polymerase extension from universal priming sites (Korlach et al., 2008, Proc. Natl. Acad. Sci. 105:1176-1181). The term "forward" when used in context with a "universal priming site" or "universal primer" may also be referred to as "5'" or "sense". The term "reverse" when used in context with a "universal priming site" or "universal primer" may also be referred to as "3'" or "antisense".

[0067] As used herein, the term "extended recording tag" refers to a recording tag to which information of at least one binding agent's coding tag (or its complementary sequence) has been transferred following binding of the binding agent to a polypeptide. Information of the coding tag may be transferred to the recording tag directly (e.g., ligation) or indirectly (e.g., primer extension). Information of a coding tag may be transferred to the recording tag enzymatically or chemically. An extended recording tag may comprise binding agent information of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200 or more coding tags. The base sequence of an extended recording tag may reflect the temporal and sequential order of binding of the binding agents identified by their coding tags, may reflect a partial sequential order of binding of the binding agents identified by the coding tags, or may not reflect any order of binding of the binding agents identified by the coding tags. In certain embodiments, the coding tag information present in the extended recording tag represents with at least 25%, 30%, 35% , 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity the polypeptide sequence being analyzed. In certain embodiments where the extended recording tag does not represent the polypeptide sequence being analyzed with 100% identity, errors may be due to off-target binding by a binding agent, or to a "missed" binding cycle (e.g., because a binding agent fails to bind to a polypeptide during a binding cycle, because of a failed primer extension reaction), or both.

[0068] As used herein, the term "extended coding tag" refers to a coding tag to which information of at least one recording tag (or its complementary sequence) has been transferred following binding of a binding agent, to which the coding tag is joined, to a polypeptide, to which the recording tag is associated. Information of a recording tag may be transferred to the coding tag directly (e.g., ligation), or indirectly (e.g., primer extension). Information of a recording tag may be transferred enzymatically or chemically. In certain embodiments, an extended coding tag comprises information of one recording tag, reflecting one binding event. As used herein, the term "di-tag" or "di-tag construct" or "di-tag molecule" refers to a nucleic acid molecule to which information of at least one recording tag (or its complementary sequence) and at least one coding tag (or its complementary sequence) has been transferred following binding of a binding agent, to which the coding tag is joined, to a polypeptide, to which the recording tag is associated. Information of a recording tag and coding tag may be transferred to the di-tag indirectly (e.g., primer extension). Information of a recording tag may be transferred enzymatically or chemically. In certain embodiments, a di-tag comprises a UMI of a recording tag, a compartment tag of a recording tag, a universal priming site of a recording tag, a UMI of a coding tag, an encoder sequence of a coding tag, a binding cycle specific barcode, a universal priming site of a coding tag, or any combination thereof.

[0069] As used herein, the term "solid support", "solid surface", or "solid substrate", or "sequencing substrate", or "substrate" refers to any solid material, including porous and nonporous materials, to which a polypeptide can be associated directly or indirectly, by any means known in the art, including covalent and non-covalent interactions, or any combination thereof. A solid support may be two-dimensional (e.g., planar surface) or three-dimensional (e.g., gel matrix or bead). A solid support can be any support surface including, but not limited to, a bead, a microbead, an array, a glass surface, a silicon surface, a plastic surface, a filter, a membrane, a PTFE membrane, a PTFE membrane, a nitrocellulose membrane, a nitrocellulose-based polymer surface, nylon, a silicon wafer chip, a flow through chip, a flow cell, a biochip including signal transducing electronics, a channel, a microtiter well, an ELISA plate, a spinning interferometry disc, a nitrocellulose membrane, a nitrocellulose-based polymer surface, a polymer matrix, a nanoparticle, or a microsphere. Materials for a solid support include but are not limited to acrylamide, agarose, cellulose, dextran, nitrocellulose, glass, gold, quartz, polystyrene, polyethylene vinyl acetate, polypropylene, polyester, polymethacrylate, polyacrylate, polyethylene, polyethylene oxide, polysilicates, polycarbonates, poly vinyl alcohol (PVA), Teflon, fluorocarbons, nylon, silicon rubber, polyanhydrides, polyglycolic acid, polyvinylchloride, polylactic acid, polyorthoesters, functionalized silane, polypropylfumerate, collagen, glycosaminoglycans, polyamino acids, dextran, or any combination thereof. Solid supports further include thin film, membrane, bottles, dishes, fibers, woven fibers, shaped polymers such as tubes, particles, beads, microspheres, microparticles, or any combination thereof. For example, when solid surface is a bead, the bead can include, but is not limited to, a ceramic bead, polystyrene bead, a polymer bead, a polyacrylate bead, a methylstyrene bead, an agarose bead, a cellulose bead, a dextran bead, an acrylamide bead, a solid core bead, a porous bead, a paramagnetic bead, a glass bead, a controlled pore bead, a silica-based bead, or any combinations thereof. A bead may be spherical or an irregularly shaped. A bead or support may be porous. A bead's size may range from nanometers, e.g., 100 nm, to millimeters, e.g., 1 mm. In certain embodiments, beads range in size from about 0.2 micron to about 200 microns, or from about 0.5 micron to about 5 micron. In some embodiments, beads can be about 1, 1.5, 2, 2.5, 2.8, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, 10, 10.5, 15, or 20 µm in diameter. In certain embodiments, "a bead" solid support may refer to an individual bead or a plurality of beads. In some embodiments, the solid surface is a nanoparticle. In certain embodiments, the nanoparticles range in size from about 1 nm to about 500 nm in diameter, for example, between about 1 nm and about 20 nm, between about 1 nm and about 50 nm, between about 1 nm and about 100 nm, between about 10 nm and about 50 nm, between about 10 nm and about 100 nm, between about 10 nm and about 200 nm, between about 50 nm and about 100 nm, between about 50 nm and about 150, between about 50 nm and about 200 nm, between about 100 nm and about 200 nm, or between about 200 nm and about 500 nm in diameter. In some embodiments, the nanoparticles can be about 10 nm, about 50 nm, about 100 nm, about 150 nm, about 200 nm, about 300 nm, or about 500 nm in diameter. In some embodiments, the nanoparticles are less than about 200 nm in diameter.

[0070] As used herein, the term "nucleic acid molecule" or "polynucleotide" refers to a single- or double-stranded polynucleotide containing deoxyribonucleotides or ribonucleotides that are linked by 3'-5' phosphodiester bonds, as well as polynucleotide analogs. A nucleic acid molecule includes, but is not limited to, DNA, RNA, and cDNA. A polynucleotide analog may possess a backbone other than a standard phosphodiester linkage found in natural polynucleotides and, optionally, a modified sugar moiety or moieties other than ribose or deoxyribose. Polynucleotide analogs contain bases capable of hydrogen bonding by Watson-Crick base pairing to standard polynucleotide bases, where the analog backbone presents the bases in a manner to permit such hydrogen bonding in a sequence-specific fashion between the oligonucleotide analog molecule and bases in a standard polynucleotide. Examples of polynucleotide analogs include, but are not limited to xeno nucleic acid (XNA), bridged nucleic acid (BNA), glycol nucleic acid (GNA), peptide nucleic acids (PNAs), yPNAs, morpholino polynucleotides, locked nucleic acids (LNAs), threose nucleic acid (TNA), 2'-O-Methyl polynucleotides, 2'-O-alkyl ribosyl substituted polynucleotides, phosphorothioate polynucleotides, and boronophosphate polynucleotides. A polynucleotide analog may possess purine or pyrimidine analogs, including for example, 7-deaza purine analogs, 8-halopurine analogs, 5-halopyrimidine analogs, or universal base analogs that can pair with any base, including hypoxanthine, nitroazoles, isocarbostyril analogues, azole carboxamides, and aromatic triazole analogues, or base analogs with additional functionality, such as a biotin moiety for affinity binding. In some embodiments, the nucleic acid molecule or oligonucleotide is a modified oligonucleotide. In some embodiments, the nucleic acid molecule or oligonucleotide is a DNA with pseudo-complementary bases, a DNA with protected bases, an RNA molecule, a BNA molecule, an XNA molecule, a LNA molecule, a PNA molecule, a γPNA molecule, or a morpholino DNA, or a combination thereof. In some embodiments, the nucleic acid molecule or oligonucleotide is backbone modified, sugar modified, or nucleobase modified. In some embodiments, the nucleic acid molecule or oligonucleotide has nucleobase protecting groups such as Alloc, electrophilic protecting groups such as thiranes, acetyl protecting groups, nitrobenzyl protecting groups, sulfonate protecting groups, or traditional base-labile protecting groups.

[0071] As used herein, "nucleic acid sequencing" means the determination of the order of nucleotides in a nucleic acid molecule or a sample of nucleic acid molecules.

[0072] As used herein, "next generation sequencing" refers to high-throughput sequencing methods that allow the sequencing of millions to billions of molecules in parallel. Examples of next generation sequencing methods include sequencing by synthesis, sequencing by ligation, sequencing by hybridization, polony sequencing, ion semiconductor sequencing, and pyrosequencing. By attaching primers to a solid substrate and a complementary sequence to a nucleic acid molecule, a nucleic acid molecule can be hybridized to the solid substrate via the primer and then multiple copies can be generated in a discrete area on the solid substrate by using polymerase to amplify (these groupings are sometimes referred to as polymerase colonies or polonies). Consequently, during the sequencing process, a nucleotide at a particular position can be sequenced multiple times (e.g., hundreds or thousands of times) - this depth of coverage is referred to as "deep sequencing." Examples of high throughput nucleic acid sequencing technology include platforms provided by Illumina, BGI, Qiagen, Thermo-Fisher, and Roche, including formats such as parallel bead arrays, sequencing by synthesis, sequencing by ligation, capillary electrophoresis, electronic microchips, "biochips," microarrays, parallel microchips, and single-molecule arrays (See e.g., Service, Science (2006) 311:1544-1546).

[0073] As used herein, "single molecule sequencing" or "third generation sequencing" refers to next-generation sequencing methods wherein reads from single molecule sequencing instruments are generated by sequencing of a single molecule of DNA. Unlike next generation sequencing methods that rely on amplification to clone many DNA molecules in parallel for sequencing in a phased approach, single molecule sequencing interrogates single molecules of DNA and does not require amplification or synchronization. Single molecule sequencing includes methods that need to pause the sequencing reaction after each base incorporation ('wash-and-scan' cycle) and methods which do not need to halt between read steps. Examples of single molecule sequencing methods include single molecule real-time sequencing (Pacific Biosciences), nanopore-based sequencing (Oxford Nanopore), duplex interrupted nanopore sequencing, and direct imaging of DNA using advanced microscopy.

[0074] As used herein, "analyzing" the polypeptide means to identify, detect, quantify, characterize, distinguish, or a combination thereof, all or a portion of the components of the polypeptide. For example, analyzing a peptide, polypeptide, or protein includes determining all or a portion of the amino acid sequence (contiguous or non-continuous) of the peptide. Analyzing a polypeptide also includes partial identification of a component of the polypeptide. For example, partial identification of amino acids in the polypeptide protein sequence can identify an amino acid in the protein as belonging to a subset of possible amino acids. Analysis typically begins with analysis of the n NTAA, and then proceeds to the next amino acid of the peptide (i.e., n-1, n-2, n-3, and so forth). This is accomplished by elimination of the n NTAA, thereby converting the n-1 amino acid of the peptide to an N-terminal amino acid (referred to herein as the "n-1 NTAA"). Analyzing the peptide may also include determining the presence and frequency of post-translational modifications on the peptide, which may or may not include information regarding the sequential order of the post-translational modifications on the peptide. Analyzing the peptide may also include determining the presence and frequency of epitopes in the peptide, which may or may not include information regarding the sequential order or location of the epitopes within the peptide. Analyzing the peptide may include combining different types of analysis, for example obtaining epitope information, amino acid sequence information, post-translational modification information, or any combination thereof.

[0075] The term "unmodified" (also "wild-type" or "native") as used herein is used in connection with biological materials such as nucleic acid molecules and proteins (e.g., cleavase), refers to those which are found in nature and not modified by human intervention.

[0076] The term "modified" or "engineered" (or "variant" or mutant") as used in reference to nucleic acid molecules and protein molecules, e.g., a modified dipeptide cleavase, implies that such molecules are created by human intervention and / or they are non-naturally occurring. The variant, mutant or modified dipeptide cleavase is a polypeptide having an altered amino acid sequence, relative to an unmodified or wild-type dipeptide cleavase. The variant or modified dipeptide cleavase is a polypeptide which differs from a wild-type dipeptide cleavase sequence by one or more amino acid substitutions, deletions, additions, or combinations thereof. A variant, mutant or modified dipeptide cleavase can contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more amino acid differences (e.g., mutations) compared to the wild-type cleavase. A variant or modified dipeptide cleavase polypeptide generally exhibits at least 25%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a corresponding wild-type or unmodified dipeptide cleavase. Non-naturally occurring amino acids as well as naturally occurring amino acids are included within the scope of permissible substitutions or additions. A variant, mutant or modified dipeptide cleavase is not limited to any variant, mutant or modified dipeptide cleavase made or generated by a particular method of making and includes, for example, a variant, mutant or modified dipeptide cleavase made or generated by genetic selection, protein engineering, directed evolution, de novo recombinant DNA techniques, or combinations thereof. A mutant, variant or modified dipeptide cleavase polypeptide is altered in primary amino acid sequence by substitution, addition, or deletion of amino acid residues. The term "variant" in the context of variant or modified dipeptide cleavase is not be construed as imposing any condition for any particular starting composition or method by which the variant or modified dipeptide cleavase is created. Thus, variant or modified dipeptide cleavase denotes a composition and not necessarily a product produced by any given process. A variety of techniques including genetic selection, protein engineering, recombinant methods, chemical synthesis, or combinations thereof, may be employed.

[0077] In some embodiments, variants of a modified dipeptide cleavase displaying only non-substantial or negligible differences in structure can be generated by making conservative amino acid substitutions in the modified dipeptide cleavase. By doing this, modified dipeptide cleavase variants that comprise a sequence having at least 90% (90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99%) sequence identity with the modified dipeptide cleavase sequences provided in the attached Sequence Listing can be generated, retaining at least one functional activity, e g., cleavase activity for a given substrate. Examples of conservative amino acid changes are known in the art. Examples of non-conservative amino acid changes that are likely to cause major changes in protein structure are those that cause substitution of (a) a hydrophilic residue, e.g., serine or threonine, for (or by) a hydrophobic residue, e.g., leucine, isoleucine, phenylalanine, valine or alanine; (b) a cysteine or proline for (or by) any other residue; (c) a residue having an electropositive side chain, e.g., lysine, arginine, or histidine, for (or by) an electronegative residue, e.g., glutamic acid or aspartic acid; or (d) a residue having a bulky side chain, e.g., phenylalanine, for (or by) one not having a side chain, e g., glycine. Methods of making targeted amino acid substitutions, deletions, truncations, and insertions are generally known in the art. For example, amino acid sequence variants can be prepared by mutations in the DNA. Methods for polynucleotide alterations are well known in the art, for example, Kunkel et al. (1987) Methods in Enzymol. 154:367-382; U.S. Pat. No. 4,873,192 and the references cited therein.

[0078] The term "sequence identity" as used herein refers to the sequence identity between genes or proteins at the nucleotide or amino acid level, respectively. "Sequence identity" is a measure of identity between proteins at the amino acid level and a measure of identity between nucleic acids at nucleotide level. The protein sequence identity may be determined by comparing the amino acid sequence in a given position in each sequence when the sequences are aligned. Similarly, the nucleic acid sequence identity may be determined by comparing the nucleotide sequence in a given position in each sequence when the sequences are aligned. "Sequence identity" means the percentage of identical subunits at corresponding positions in two sequences when the two sequences are aligned to maximize subunit matching, i.e., taking into account gaps and insertions. Sequence identity is present when a subunit position in both of the two sequences is occupied by the same nucleotide or amino acid, e.g., if a given position is occupied by an adenine in each of two DNA molecules, then the molecules are identical at that position. For example, if 7 positions in a sequence of 10 nucleotides in length are identical to the corresponding positions in a second 10-nucleotide sequence, then the two sequences have 70% sequence identity. Methods for the alignment of sequences for comparison are well known in the art, such methods include GAP, BESTFIT, BLAST, FASTA and TFASTA. The BLAST algorithm calculates percent sequence identity and performs a statistical analysis of the similarity between the two sequences. The software for performing BLAST analysis is publicly available through the National Center for Biotechnology Information (NCBI) website.

[0079] The terms "corresponding to position(s)" or "position(s) ... with reference to position(s)" of or within a polypeptide or a polynucleotide, such as recitation that nucleotides or amino acid positions "correspond to" nucleotides or amino acid positions of a disclosed sequence, such sequence set forth in the Sequence Listing, refers to nucleotides or amino acid positions identified in the polynucleotide or in the polypeptide upon alignment with the disclosed sequence using a standard alignment algorithm, such as the BLAST algorithm (NCBI). For example, one skilled in the art can identify a residue in a given polypeptide at a position corresponding to position 191 of SEQ ID NO: 13 by making a BLASTP alignment of the polypeptide sequence together with SEQ ID NO: 13, and find the residue in the polypeptide that is aligned with the residue 191 of SEQ ID NO: 13. By aligning the sequences, one skilled in the art can identify corresponding residues in a given polypeptide, for example, by using conserved and identical amino acid residues in the alignment as guides. Similarly, one skilled in the art can identify any given amino acid residue in a given polypeptide at a position corresponding to a particular position of a reference sequence, such as set forth in the Sequence Listing, by performing alignment of the polypeptide sequence with the reference sequence (for example, by BLASTP publicly available through the NCBI website), matching the corresponding position of the reference sequence with the position in polypeptide sequence and thus identifying the amino acid residue within the polypeptide.

[0080] As used herein, domain (such as a sequence of amino acid residues) refers to a portion of a molecule, such as a protein or encoding nucleic acid, that is structurally and / or functionally distinct from other portions of the molecule and is identifiable. For example, domains include those portions of a polypeptide chain that can form an independently folded structure within a protein made up of one or more structural motifs and / or that is recognized by virtue of a functional activity, such as binding activity. A protein can have one, or more than one, distinct domains. For example, a domain can be identified, defined or distinguished by homology of the primary sequence or structure to related family members, such as homology to motifs. In another example, a domain can be distinguished by its function, such as an ability to interact with a biomolecule, such as a cognate binding partner. A domain independently can exhibit a biological function or activity such that the domain independently or fused to another molecule can perform an activity, such as, for example binding. A domain can be a linear sequence of amino acids or a non-linear sequence of amino acids. Many polypeptides contain a plurality of domains. Such domains are known, and can be identified by those of skill in the art. For exemplification herein, definitions are provided, but it is understood that it is well within the skill in the art to recognize particular domains by name. If needed, appropriate software can be employed to identify domains.

[0081] As used herein, the term "alkyl" refers to and includes saturated linear and branched univalent hydrocarbon structures and combination thereof, having the number of carbon atoms designated (i.e., C 1 -C 10 or C 1-10 means one to ten carbons). Particular alkyl groups are those having 1 to 20 carbon atoms (a "C 1 -C 20 alkyl"). More particular alkyl groups are those having 1 to 8 carbon atoms (a "C 1 -C 8 alkyl"), 3 to 8 carbon atoms (a "C 3 -C 8 alkyl"), 1 to 6 carbon atoms (a "C 1 -C 6 alkyl"), 1 to 5 carbon atoms (a "C 1 -C 5 alkyl"), or 1 to 4 carbon atoms (a "C 1 -C 4 alkyl"), unless otherwise specified Examples of alkyl include, but are not limited to, groups such as methyl, ethyl, n-propyl, isopropyl, n-butyl, t-butyl, isobutyl, sec-butyl, homologs and isomers of, for example, n-pentyl, n-hexyl, n-heptyl, n-octyl, and the like.

[0082] As used herein, "alkenyl" as used herein refers to an unsaturated linear or branched univalent hydrocarbon chain or combination thereof, having at least one site of olefinic unsaturation (i.e., having at least one moiety of the formula C=C) and having the number of carbon atoms designated (i.e., C 2 -C 10 means two to ten carbon atoms). The alkenyl group may be in "cis" or "trans" configurations, or alternatively in "E" or "Z" configurations. Particular alkenyl groups are those having 2 to 20 carbon atoms (a "C 2 -C 20 alkenyl"), having 2 to 8 carbon atoms (a "C 2 -C 8 alkenyl"), having 2 to 6 carbon atoms (a "C 2 -C 6 alkenyl"), or having 2 to 4 carbon atoms (a "C 2 -C 4 alkenyl"). Examples of alkenyl include, but are not limited to, groups such as ethenyl (or vinyl), prop-1-enyl, prop-2-enyl (or allyl), 2-methylprop-1-enyl, but-1-enyl, but-2-enyl, but-3-enyl, buta-1,3-dienyl, 2-methylbuta-1,3-dienyl, homologs and isomers thereof, and the like.

[0083] The term "aminoalkyl" refers to an alkyl group that is substituted with one or more -NH 2 groups. In certain embodiments, an aminoalkyl group is substituted with one, two, three, four, five or more -NH 2 groups. An aminoalkyl group may optionally be substituted with one or more additional substituents as described herein.

[0084] As used herein, "aryl" or "Ar" refers to an unsaturated aromatic carbocyclic group having a single ring (e.g., phenyl) or multiple condensed rings (e.g., naphthyl or anthryl) which condensed rings may or may not be aromatic. In one variation, the aryl group contains from 6 to 14 annular carbon atoms. An aryl group having more than one ring where at least one ring is non-aromatic may be connected to the parent structure at either an aromatic ring position or at a non-aromatic ring position. In one variation, an aryl group having more than one ring where at least one ring is non-aromatic is connected to the parent structure at an aromatic ring position. In some embodiments, phenyl is a preferred aryl group.

[0085] As used herein, the term "arylalkyl" refers to an aryl group, as defined herein, appended to the parent molecular moiety through an alkyl group, as defined herein. Representative examples of arylalkyl include, but are not limited to, benzyl, 2- phenylethyl, 3-phenylpropyl, 2-naphth-2-ylethyl, and the like.

[0086] As used herein, the term "cycloalkyl" refers to and includes cyclic univalent hydrocarbon structures, which may be fully saturated, mono- or polyunsaturated, but which are non-aromatic, having the number of carbon atoms designated (e.g., C 1 -C 10 means one to ten carbons). Cycloalkyl can consist of one ring, such as cyclohexyl, or multiple rings, such as adamantly, but excludes aryl groups. A cycloalkyl comprising more than one ring may be fused, spiro or bridged, or combinations thereof. In some embodiments, the cycloalkyl is a cyclic hydrocarbon having from 3 to 13 annular carbon atoms. In some embodiments, the cycloalkyl is a cyclic hydrocarbon having from 3 to 8 annular carbon atoms (a "C 3 -C 8 cycloalkyl"). Examples of cycloalkyl include, but are not limited to, cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl, 1-cyclohexenyl, 3-cyclohexenyl, cycloheptyl, norbornyl, and the like.

[0087] As used herein, the "halogen" represents chlorine, fluorine, bromine, or iodine. The term "halo" represents chloro, fluoro, bromo, or iodo.

[0088] The term "haloalkyl" refers to an alkyl group as described above, wherein one or more hydrogen atoms on the alkyl group have been replaced by a halo group. Examples of such groups include, without limitation, fluoroalkyl groups, such as fluoroethyl, trifluoromethyl, difluoromethyl, trifluoroethyl and the like.

[0089] As used herein, the term "heteroaryl" refers to and includes unsaturated aromatic cyclic groups having from 1 to 10 annular carbon atoms and at least one annular heteroatom, including but not limited to heteroatoms such as nitrogen, oxygen and sulfur, wherein the nitrogen and sulfur atoms are optionally oxidized, and the nitrogen atom(s) are optionally quaternized. It is understood that the selection and order of heteroatoms in a heteroaryl ring must conform to standard valence requirements and provide an aromatic ring character, and also must provide a ring that is sufficiently stable for use in the reactions described herein. Typically, a heteroaryl ring has 5-6 ring atoms and 1-4 heteroatoms, which are selected from N, O and S unless otherwise specified; and a bicyclic heteroaryl group contains two 5-6 membered rings that share one bond and contain at least one heteroatom and up to 5 heteroatoms selected from N, O and S as ring members. A heteroaryl group can be attached to the remainder of the molecule at an annular carbon or at an annular heteroatom, in which case the heteroatom is typically nitrogen. Heteroaryl groups may contain additional fused rings (e.g., from 1 to 3 rings), including additionally fused aryl, heteroaryl, cycloalkyl, and / or heterocyclyl rings. Examples of heteroaryl groups include, but are not limited to, pyrazolyl, imidazolyl, triazolyl, pyrrolyl, pyridyl, pyrimidyl, pyrazinyl, pyridazinyl, triazinyl, thiophenyl, furanyl, thiazolyl, and the like.

[0090] As used herein, the term "heterocycle", "heterocyclic", or "heterocyclyl" refers to a saturated or an unsaturated non-aromatic group having from 1 to 10 annular carbon atoms and from 1 to 4 annular heteroatoms, such as nitrogen, sulfur or oxygen, and the like, wherein the nitrogen and sulfur atoms are optionally oxidized, and the nitrogen atom(s) are optionally quaternized. A heterocyclyl group may have a single ring or multiple condensed rings, but excludes heteroaryl groups. A heterocycle comprising more than one ring may be fused, spiro or bridged, or any combination thereof. In fused ring systems, one or more of the fused rings can be aryl or heteroaryl. Examples of heterocyclyl groups include, but are not limited to, tetrahydropyranyl, dihydropyranyl, piperidinyl, piperazinyl, pyrrolidinyl, thiazolinyl, thiazolidinyl, tetrahydrofuranyl, tetrahydrothiophenyl, 2,3-dihydrobenzo[b]thiophen-2-yl, 4-amino-2-oxopyrimidin-1(2H)-yl, and the like.

[0091] The term "substituted" means that the specified group or moiety bears one or more substituents in place of a hydrogen atom of the unsubstituted group, including, but not limited to, substituents such as alkoxy, acyl, acyloxy, carbonylalkoxy, acylamino, amino, aminoacyl, aminocarbonylamino, aminocarbonyloxy, cycloalkyl, cycloalkenyl, aryl, heteroaryl, aryloxy, cyano, azido, halo, hydroxyl, nitro, carboxyl, thiol, thioalkyl, cycloalkyl, cycloalkenyl, alkyl, alkenyl, alkynyl, heterocyclyl, aralkyl, aminosulfonyl, sulfonylamino, sulfonyl, oxo, carbonylalkylenealkoxy and the like. The term "unsubstituted" means that the specified group bears no substituents. The term "optionally substituted" means that the specified group is unsubstituted or substituted by one or more substituents and thus includes both substituted and unsubstituted versions of the group. Where the term "substituted" is used to describe a structural system, the substitution is meant to occur at any valency-allowed position on the system.

[0092] It is understood that aspects and embodiments of the invention described herein include "consisting" and / or "consisting essentially of" aspects and embodiments.

[0093] Throughout this disclosure, various aspects of this invention are presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the invention. Accordingly, the description of a range should be considered to have specifically disclosed all the possible sub-ranges as well as individual numerical values within that range. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed sub-ranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, for example, 1, 2, 3, 4, 5, and 6. This applies regardless of the breadth of the range.

[0094] Other objects, advantages and features of the present invention will become apparent from the following specification taken in conjunction with the accompanying drawings.I. MODIFIED OR ENGINEERED CLEAVASES

[0095] In another aspect, provided herein is a modified or an engineered cleavase comprising a mutation, e.g., one or more amino acid modification(s), in an unmodified cleavase, wherein: said modified or engineered cleavase is derived from a dipeptidyl peptidase of Thermomonas hydrothermalis or Caldithrix abyssii and removes or is configured to remove a single N-terminally modified amino acid from a target polypeptide. In some embodiments, the present modified or engineered cleavase is configured to cleave the peptide bond between a N-terminally modified amino acid residue and a penultimate terminal amino acid residue of the target polypeptide.

[0096] The present modified or engineered cleavase can comprise any suitable active site. For example, the present modified or engineered cleavase can comprise an active site that interacts with the amide bond between the N-terminally modified amino acid residue and a penultimate terminal amino acid residue of the target polypeptide.

[0097] The present modified or engineered cleavase can be derived from any suitable type of dipeptidyl peptidase. For example, the present modified or engineered cleavase can be derived from a protein or enzyme classified as a S46 dipeptidyl peptidase (see e.g., Shakh M.A. Rouf, Yuko Ohara-Nemoto, Tomonori Hoshino, Taku Fujiwara, Toshio Ono, Takayuki K. Nemoto, Discrimination based on Gly and Arg / Ser at position 673 between dipeptidyl-peptidase (DPP) 7 and DPP11, widely distributed DPPs in pathogenic and environmental gram-negative bacteria, Biochimie, Volume 95, Issue 4, 2013, Pages 824-832, ISSN 0300-9084), or a functional homolog or fragment thereof.

[0098] The present modified or engineered cleavase can remove or can be configured to remove any suitable single N-terminally modified amino acid from a target polypeptide. For example, the present modified or engineered cleavase can remove or can be configured to remove a N-terminal amino acid that is labeled with a chemical or an enzymatic reagent or moiety.

[0099] The present modified or engineered cleavase can remove or can be configured to remove any suitable single N-terminally modified amino acid from a target polypeptide containing any suitable N-terminal modification (NTM), such as synthetic NTM. In another example, NTM can comprise an amino acid moiety and / or has a size, e.g., length axis or volume, shape, and / or configuration similar to or exceeding a natural amino acid. In some embodiments, NTM can be a bipartite N-terminal modification that comprises a natural or unnatural amino acid portion (NTMaa) and a N-terminal blocking group (NTM blk ). The amino acid-like portion (NTMaa) and the N-terminal blocking group (NTM blk ) can be connected or linked by any suitable bond or linkage. For example, the amino acid portion (NTMaa) and the N-terminal blocking group (NTM blk ) can be connected with an amide bond.

[0100] In some embodiments, NTM does not comprise an amino acid moiety. In some embodiments, NTM comprises a N-terminal blocking group (NTM blk ) and does not comprise a NTMaa group. In some embodiments, NTM can be a bipartite N-terminal modification that comprises a small (or small molecule) chemical entity having a size, e.g., length axis or volume, shape, and / or configuration similar to or exceeding a natural amino acid, and a N-terminal blocking group (NTM blk ). The small (or small molecule) chemical entity and the N-terminal blocking group (NTM blk ) can be connected or linked by any suitable bond or linkage. For example, the small (or small molecule) chemical entity and the N-terminal blocking group can be connected with an amide bond. The small (or small molecule) chemical entity can have any suitable size, e.g., length axis or volume. For example, the small (or small molecule) chemical entity can have a size, e.g., length axis of about 5-10 Å and volume of about 100 - 1000 Å 3< . In some embodiments, the small (or small molecule) chemical entity has a length axis of about 5, 6, 7, 8, 9 or 10 Å, or any range thereof. In some embodiments, , the small (or small molecule) chemical entity has a volume of about 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 Å 3< or any range thereof.

[0101] In another example, the N-terminal modification can comprise a chemical label.

[0102] In some embodiments, a chemical reagent for the N-terminal modification is selected from the group consisting of: 2-aminobenzamide, 2-(N-methylamino)-benzamide, 2-(N-acetylamine)-benzamide, 2-(N-benzylamine)-benzamide, 4-methylbenzamide, 4-(dimethylamino)benzamide, nicotinamide, 3-aminonicotinamide, 2-pyrazinecarbonyl, 5-amino-2-fluoro-isonicotinamide, 2-carboxylic acid pyrazinecarbonyl, 3,6-difluoro-2-carboxybenzamide, 4-chloro-2-aminobenzamide, 4-nitro-2-aminobenzamide, 4-methoxy-2-aminobenzamide, 4-carboxylic acid-2-aminobenzamide, 5-(trifluoromethyl-2-aminobenzamide, 4-(trifluoromethyl-2-aminobenzamide, 6-fluoro-2-aminobenzamide, 4-fluoro-2-aminobenzamide, 5-methoxy-2-aminobenzamide, 4-fluorobenzamide, 4-(trifluoromethyl)benzamide, 8-fluoroisoquinolinium, 1-hydroxy-2,3,1-benzodiazaborinine-2(1H)-carbonyl, Succinamide, 3,6-Difluoropyridine-2-carbamide, 2-Fluoronicotinamide, 5-Bromo-2-hydroxynicotinamide, 4-(Trifluoromethyl)pyrimidine-5-carbamide, 2-Oxo-1,2-dihydropyridine-3-carbamide, 5-Methyl-2-aminobenzamide, 6-Fluoropicolinamide, 3-Methyl-2-aminobenzamide, 4-Methyl--2-aminobenzamide, 2-Amino-6-methylbenzamide, 2-Amino-6-fluorobenzamide, 2-Amino-5-fluorobenzoamide, 2-Amino-3-fluorobenzoamide, 2-Amino-4-fluorobenzoamide, 2-Aminonicotinamide, 4-Aminonicotinamide, 3-Aminopicolinamide, or a derivative thereof. In some embodiments, the chemical reagent for the N-terminal modification is an isatoic anhydride, an isonicotinic anhydride, an azaisatoic anhydride, a succinic anhydride, an aryl activated ester, a heteroaryl activated ester, a non-aromatic ring activated ester, or a derivative thereof. In some embodiments, the chemical reagent for the N-terminal modification is selected from the group consisting of wherein the chemical reagent is selected from the group consisting of 4-Nitrophenyl Anthranilate, N-Methyl-isatoic anhydride, N-acetyl-isatoic anhydride, N-benzyl-isatoic anhydride, 4-methylbenzoic acid, 4-(dimethylamino)benzoyl chloride, nicotinic acid-NHS, 3-aminonicotinic acid, 2-pyrazinecarbonyl chloride, 5-amino-2-fluoro-isonicotinic acid, 2,3-pyrazinedicarboxylic anhydride, 3,6-difluorophthalic anhydride, 4-chloroisatoic anhydride, 4-nitroisatoic anhydride, 7-methoxy-1h-benzo[d][1,3]oxazine-2,4-dione, 4-carboxylic acid isatoic anhydride, 6-(Trifluoromethyl)-2,4-dihydro-1h-3,1-benzoxazine-2,4-dione, 7-(Trifluoromethyl)-lh-benzo[d][1,3]oxazine-2,4-dione, 6-fluoroisatoic anhydride, 4-fluoroisatoic anhydride, 5-methoxyisatoic anhydride, 4-fluorobenzoic acid anhydride, 4-(trifluoromethyl)benzoic acid anhydride, 2-ethynyl-6-fluorobenzaldehyde, 1-hydroxy-2,3,1-benzodiazaborinine-2(1H)-carboxylic acid, Isatoic anhydride, Succinic anhydride 3,6-Difluoropyridine-2-carboxylic acid, 2-Fluoronicotinic acid, 5-Bromo-2-hydroxynicotinic acid, 4-(Trifluoromethyl)pyrimidine-5-carboxylic acid, 2-Oxo-1,2-dihydropyridine-3-carboxylic acid, 5-Methylisatoic anhydride, 6-Fluoropicolinic acid, 3-Methylisatoic anhydride, 4-Methyl-isatoic anhydride, 2-Amino-6-methylbenzoic acid, 2-Amino-6-fluorobenzoic acid, 2-Amino-5-fluorobenzoic acid, 2-Amino-3-fluorobenzoic acid, 2-Amino-4-fluorobenzoic acid, 2-Aminonicotinic acid, 4-Aminonicotinic acid, 3-Aminopicolinic acid, or a derivative thereof.

[0103] The present modified or engineered cleavase can comprise any suitable amino acid sequence variation(s) as compared with the amino acid sequence of the unmodified cleavase. For example, the present modified or engineered cleavase can comprise an amino acid sequence that exhibits at least 50 % identity, at least 60 % identity, at least 70 % identity, at least 80 % identity, or at least 90 %, or at least 95 %, or more identity with the unmodified cleavase.

[0104] The present or engineered modified cleavase can comprise any suitable type of mutation(s). For example, wherein the mutation can comprise an amino acid substitution, deletion, addition, or a combination thereof.

[0105] The present modified or engineered cleavase can remove or can be configured to remove a single N-terminally modified amino acid from a target polypeptide with any suitable length. For example, the length of the target polypeptide can be greater than 4 amino acids, greater than 5 amino acids, greater than 6 amino acids, greater than 7 amino acids, greater than 8 amino acids, greater than 9 amino acids, greater than 10 amino acids, greater than 11 amino acids, greater than 12 amino acids, greater than 13 amino acids, greater than 14 amino acids, greater than 15 amino acids, greater than 20 amino acids, greater than 25 amino acids, or greater than 30 amino acids.

[0106] The present modified or engineered cleavase can comprise mutation(s) at any suitable site(s). For example, the present modified or engineered cleavase can comprise a modification within its substrate binding site. Substrate binding site of a cleavase is comprised of amino acid residues that are involved in interaction with the substrate during substrate recognition and cleavage. The specificity for the substrate is due to the favorable binding interaction of the substrate amino acid side chains with residues that form the substrate binding site of the cleavase (also called specificity pocket). For example, the binding site of dipeptidyl aminopeptidases comprises residues that are involved in interaction with the N-terminal amino group of a polypeptide (these residues form an amine binding site that is a part of the substrate binding site), and residues that are involved in interaction with P1 and P2 residues of the polypeptide. For modified dipeptidyl aminopeptidases amino acid residues in the substrate binding site are modified to interact with the NTM of a labeled polypeptide, and also with P1 or P2 residues of the polypeptide. In some cases, amino acid residues in the substrate binding site of a modified dipeptidyl aminopeptidase are modified such that the modified dipeptidyl aminopeptidase would not recognize the N-terminal amino group of a polypeptide, and thus the modified dipeptidyl aminopeptidase would not cleave unlabeled polypeptide. The substrate binding site of a dipeptidyl cleavase can be determined for example using crystal structure of the dipeptidyl cleavase with its substrate or with an inhibitor, mimicking the substrate.

[0107] In another example, the present modified or engineered cleavase can comprise a modification within its catalytic domain. In still another example, the present modified or engineered cleavase can comprise a modification within its chymotrypsin fold. In yet another example, the present modified or engineered cleavase can comprise a modification at an amine binding site. In yet another example, the present modified or engineered cleavase can comprise a modification in its S1 and / or S2 sites. In yet another example, the present modified or engineered cleavase can comprise a modification for improving accessibility to the active site of the modified or engineered cleavase.

[0108] In some embodiments, the present modified or engineered cleavase is derived from a dipeptidyl peptidase of Thermomonas hydrothermalis comprising an amino acid sequence set forth in SEQ ID NO:33 (wild type (WT) sequence with the signal peptide) or SEQ ID NO:31 (WT sequence without the signal peptide).

[0109] The present modified or engineered cleavase can comprise any suitable amino acid sequence variations as compared with the amino acid sequence of the unmodified cleavase. For example, the present modified or engineered cleavase can comprise an amino acid sequence that exhibits at least 30 % identity, at least 40 % identity, at least 50 % identity, at least 60 % identity, at least 70 % identity, at least 80 % identity, at least 90 % or more identity or at least 95 % or more identity to the amino acid sequence set forth in SEQ ID NO:33 or SEQ ID NO:31, or a specific binding fragment thereof.

[0110] In some embodiments, the present modified or engineered cleavase has a mutation, with reference to positions of SEQ ID NO: 33, selected from the group consisting of N214X, W215X, R219X, N329X, N333X, A671X, D673X, G674X, N682X, M692X, I651X, and a combination thereof, X being one of the 20 naturally occurring amino acids other than the amino acid residue of the unmodified dipeptidyl peptidase at the mutated position. In some embodiments, the present modified or engineered cleavase has one or more amino acid modification(s) of N214M, W215G, R219T, N329R, D673A, and / or G674V with reference to positions of SEQ ID NO: 33.

[0111] In some embodiments, the present modified or engineered cleavase exhibits the substrate specificity of the above modified or engineered cleavase. In present some embodiments, the modified or engineered cleavase comprises an amino acid sequence that comprises a catalytic domain, an amine binding site, or S1 and / or S2 sites with at least 30 % identity, at least 40 % identity, at least 50 % identity, at least 60 % identity, at least 70 % identity, at least 80 % identity, or at least 90 %, 95 %, or more identity with the catalytic domain, the amine binding site, or the S1 and / or S2 sites of the above modified or engineered cleavase. By definition, when referring to proteases, the S1 site is defined as the region of the protease that binds to the amino acid just upstream (amino side) of the cleavage position and the S2 site binds to the amino acid two residues upstream (amino side) to the cleavage position. For a native DPP which binds to the N-terminal dipeptide of a peptide, the S2 site binds to the N-terminal amino acid and the S1 site binds to the penultimate amino acid. In the modified or engineered Cleavase, derived from a DPP, the S2 site can be used to participate in binding to the N-terminal modification (NTM), and the S1 site used to bind to the NTAA residue, creating a single modified amino acid cleavage.

[0112] In some embodiments, the present modified or engineered cleavase is derived from a dipeptidyl peptidase of Caldithrix abyssii comprising an amino acid sequence set forth in SEQ ID NO:34 (WT sequence with the signal peptide) or SEQ ID NO:32 (WT sequence without the signal peptide).

[0113] The present modified or engineered cleavase can comprise any suitable amino acid sequence variations as compared with the amino acid sequence of the unmodified cleavase. For example, the present modified or engineered cleavase can comprise an amino acid sequence that exhibits at least 30 % identity, at least 40 % identity, at least 50 % identity, at least 60 % identity, at least 70 % identity, at least 80 % identity, at least 90 % or more identity or at least 95 % or more identity to the amino acid sequence set forth in SEQ ID NO:34 or SEQ ID NO:32, or a specific binding fragment thereof.

[0114] In some embodiments, the present modified or engineered cleavase has a mutation, with reference to positions of SEQ ID NO: 34, selected from the group consisting of N207M, W208X, R212X, N322X, D663X, and a combination thereof, X being one of the 20 naturally occurring amino acids other than the amino acid residue of the unmodified dipeptidyl peptidase at the mutated position. In some embodiments, the present modified or engineered cleavase has one or more amino acid modification(s) of N207M, W208G, R212V, N322I, D663A, or a combination thereof, with reference to positions of SEQ ID NO: 34.

[0115] In some embodiments, the present modified or engineered cleavase exhibits the substrate specificity of the above modified or engineered cleavase. In present some embodiments, the modified or engineered cleavase comprises an amino acid sequence that comprises a catalytic domain, an amine binding site, or S1 and / or S2 sites, with at least 30 % identity, at least 40 % identity, at least 50 % identity, at least 60 % identity, at least 70 % identity, at least 80 % identity, or at least 90 % or more identity with the catalytic domain, the amine binding site, or the S1 and / or S2 sites of the above modified or engineered cleavase.

[0116] A nucleic acid encoding the above modified or engineered cleavase is provided herein. A vector, e.g., an expression vector, comprising the nucleic acid encoding the above modified or engineered cleavase is also provided herein. A host cell comprising the above nucleic acid or the vector is further provided herein. The host cell can be any suitable type of cell. For example, the host cell can be a mammalian or human host cell.

[0117] In another aspect, provided herein is a modified dipeptide cleavase comprising a mutation, e.g., one or more amino acid modification(s) in an unmodified dipeptide cleavase, wherein the modified dipeptide cleavase removes or is configured to remove a labeled terminal dipeptide from a polypeptide. In some embodiments, the modified dipeptide cleavase is configured to remove a single labeled dipeptide (the terminal and penultimate terminal amino acids) from the C-terminus or N-terminus of a polypeptide. In some embodiments, the modified dipeptide cleavase is derived from a wild-type or unmodified dipeptide cleavase. In some cases, the dipeptide removed contains a terminal labeled amino acid residue that is an N-terminal amino acid. In some embodiments, the dipeptide removed contains a terminal labeled amino acid residue that is a C-terminal amino acid. In some aspects, a labeled amino acid (removed as part of the dipeptide) is a terminal amino acid that is modified by treating with a chemical reagent. In some aspects, the modified dipeptide cleavase comprises an active site that interacts with an amide bond (e.g., the amide bond between a penultimate and antepenultimate terminal amino acid residues of the polypeptide). In some embodiments, the modified dipeptide cleavase contains a mutation that is an amino acid substitution, deletion, addition, or any combinations thereof.

[0118] In some embodiments, the modified dipeptide cleavase exhibits activity that is different from the activity of the unmodified or wild-type dipeptide cleavase. "Unmodified dipeptide cleavase" or "wild-type dipeptide cleavase" as used herein refers to any natural or wild-type exopeptidase, or a functional homolog or fragment thereof, that possesses catalytic activity to remove a dipeptide from the terminus of a polypeptide (e.g., from the C-terminus or N-terminus of a polypeptide). The unmodified or wild-type dipeptide cleavase may be an exopeptidase that catalyzes the cleavage of an penultimate peptide bond to release a dipeptide from the peptide chain. The unmodified or wild-type dipeptide cleavase removes unlabeled or unmodified dipeptides from the peptide chain. The unmodified or wild-type dipeptide cleavase may be a proteolytic enzyme such as an aminopeptidase or a carboxypeptidase. The unmodified dipeptide cleavase described herein may be used to refer to a protein classified by the Enzyme Commission (EC) as EC 3.4.14, EC 3.4.15, MEROPS S9, MEROPS S46, MEROPS M49, or a functional homolog or fragment thereof. The unmodified dipeptide cleavase described herein may be used to refer to a dipeptidyl peptidase, a dipeptidyl aminopeptidase, a peptidyl-dipeptidase, or a dipeptidyl carboxypeptidase.

[0119] A "modified dipeptide cleavase" or "variant dipeptide cleavase" refers to any exopeptidase that has been modified from a unmodified or wild-type dipeptide cleavase as described. The modified or variant dipeptide cleavase may be derived from an unmodified or wild-type dipeptide cleavase (e.g. a dipeptidyl peptidase, a dipeptidyl aminopeptidase, a peptidyl-dipeptidase, a dipeptidyl carboxypeptidase). As compared to an unmodified or wild-type dipeptide cleavase which removes an unlabeled P1-P2 terminal amino acids from a polypeptide as a dipeptide at a time, a modified dipeptide cleavase removes or is configured to remove a labeled terminal dipeptide from a polypeptide, or a single labeled terminal amino acid (such as N-terminal amino acid or NTAA). In some embodiments, a modified dipeptide cleavase preferentially removes a labeled P1-P2 terminal dipeptide from the polypeptide as compared to the cleavage of an unlabeled P1-P2 terminal dipeptide from the polypeptide. In some embodiments, a modified dipeptide cleavase removes only a labeled P1 terminal residue from the polypeptide as compared to the cleavage of an unlabeled P1-P2 terminal dipeptide from the polypeptide. In the present disclosure P1 is a terminal amino acid residue of a polypeptide, such as N-terminal amino acid (NTAA), and P2 is a penultimate terminal amino acid residue of a polypeptide.

[0120] In some embodiments, the modified dipeptide cleavase removes a dipeptide comprising a labeled terminal amino acid, e.g., a labeled NTAA. In some cases, the modified dipeptide cleavase removes a dipeptide comprising a labeled C-terminal amino acid (CTAA). In some embodiments, the modified dipeptide cleavase derived from the dipeptide cleavase is configured to cleave the peptide bond between a penultimate terminal labeled amino acid residue and a antepenultimate terminal amino acid residue of the polypeptide.

[0121] In another example, the peptide bond between the P1 and P2 can be cleaved using a modified cleavase. In some embodiments, the peptide bond between the P1 and P2 is cleaved using an above descried modified or engineered cleavase (see e.g., the above Section I.) In some embodiments, the peptide bond between the P1 and P2 is cleaved using a modified or an engineered cleavase described and / or claimed in U.S. provisional application serial Nos. 62 / 823,927, filed March 26, 2019, 62 / 824,157, filed March 26, 2019, and 62 / 931,737, filed November 6, 2019, and in application WO 2020 / 198264 published on October 01, 2020.

[0122] Information regarding various known unmodified or wild-type exopeptidases is available from databases such as MEROPS and / or the BRENDA enzyme information system (See e.g., Schomburg et al., J Biotechnol. (2017) 261:194-206). A protein may be classified using more than one classification system. The Enzyme Commission (EC) sets forth a numbering system for the classification of enzyme based upon specificity using recommendations from Nomenclature Committee of the International Union of Biochemistry and Molecular Biology (IUBMB) for describing each type of characterized enzyme for which an EC (Enzyme Commission) number has been provided (See e.g., Bairoch A., (2000) Nucleic Acids Res. 28:304-305). In some aspects, the unmodified or wild-type dipeptide cleavase is a protein classified in EC 3.4.14, EC 3.4.15, MEROPS S9, MEROPS S46, MEROPS M49, or a homolog thereof. In some aspects, the unmodified or wild-type dipeptide cleavase is a protein provided in Tables 1, 2, 3, 4, 5A-5B , or a homolog thereof. In some embodiments, the modified dipeptide cleavase is derived from a protein classified EC 3.4.14, EC 3.4.15, MEROPS S9, MEROPS S46, MEROPS M49, or a functional homolog or fragment thereof (as provided in Tables 1, 2, 3, 4, 5A-5B ). Table 1. Exemplary Dipeptide Cleavases from EC 3.4.14 Enzyme Commission Number Name 3.4.14.1dipeptidyl-peptidase I3.4.14.2dipeptidyl-peptidase II3.4.14.4dipeptidyl-peptidase III3.4.14.5dipeptidyl-peptidase IV3.4.14.6dipeptidyl-dipeptidase3.4.14.11Xaa-Pro dipeptidyl-peptidase3.4.14.13gamma-D-glutamyl-L-ly sine dipeptidyl-peptidase Table 2. Exemplary Dipeptide Cleavases from EC 3.4.15 Enzvme Commission Number Name 3.4.15.1peptidyl-dipeptidase A3.4.15.3dipeptidyl carboxypeptidase3.4.15.4peptidyl-dipeptidase B3.4.15.5peptidyl-dipeptidase Dcp Table 3. Exemplary Dipeptide Cleavases from MEROPS S9 MEROPS ID Name S09.003dipeptidyl-peptidase IV (eukaryote)S09.009dipeptidyl-peptidase 4 (bacteria-type 1)S09.012dipeptidyl-peptidase V / dipeptidyl-peptidase 5S09.013dipeptidyl-peptidase 4 (bacteria-type 2)S09.018dipeptidyl-peptidase 8S09.019dipeptidyl-peptidase 9S09.056dipeptidyl-peptidase IV, membrane-type (protistan)S09.075dipeptidyl-peptidase 5 (Porphyromonas sp.)Unassignedsubfamily S9B unassigned peptidases Table 4. Exemplary Dipeptide Cleavases from MEROPS S46 MEROPS ID Name S46.001dipeptidyl-peptidase 7S46.002dipeptidyl-peptidase 11S46.003dipeptidyl-peptidase BIIS46.004BF9343_2924 g.p. Table 5A. Exemplary Dipeptide Cleavases from MEROPS M49 MEROPS ID Name M49.001dipeptidyl-peptidase IIIM49.003dipeptidyl-peptidase IIIB (Bacteroides thetaiotaomicron-type)M49.004dipeptidyl-peptidase III (Saccharomyces-type) Table 5B. Other Exemplary Dipeptide Cleavases MEROPS ID Name C01.070dipeptidyl-peptidase IS28.002dipeptidyl-peptidase IIS15.001Xaa-Pro dipeptidyl-peptidaseS9G.084dipeptidyl-peptidase IV beta

[0123] Various peptidases (e.g., cleavases) that sequentially cleave off dipeptides or tripeptides from unsubstituted N-terminals of oligopeptides have been identified. Cleavases as described herein refer to enzymes that are classified under the Enzyme Commission (EC) Class 3 of hydrolases. Dipeptidyl peptidases (DPPs from Enzyme Commission number 3.4.14; Table 1) are a class of exopeptidases which digest dipeptides (two amino acid residues, P1-P2) from the N-terminal end of a peptide, typically in a processive manner. Peptidyl dipeptidases also known as dipeptidyl carboxypeptidases (EC 3.4.15; Table 2) act from the C-terminal end in removing dipeptides in a processive manner. In some embodiments, the unmodified or wild-type dipeptide cleavase is an exoaminopeptidase or an exopeptidase. In some aspects, the unmodified or wild-type dipeptide cleavase is a metallopeptidase, e.g., a zinc-dependent metallopeptidase or a zinc-dependent hydrolase. In some aspects, the unmodified or wild-type dipeptide cleavase is a serine exopeptidase or a serine protease. DPPs typically recognize the N-terminal alpha amine, and cleave the peptide bond between the penultimate and antepenultimate amino acid residues of a polypeptide (P2-P3). See e.g., Sanderink et al., J. Clin. Chem. Clin. Biochem. (1988) 26:795-807) and Baral et al., J Biol Chem (2008) 283(32): 22316-22324.

[0124] In some embodiments, the modified dipeptide cleavase exhibits activity including the removal of a labeled terminal dipeptide from polypeptides or proteins (e.g., from the N-terminus or C-terminus). In a general manner, the peptidase activity is capable of removing the amino acids Xaa 1 and Xaa 2 from the terminus of a peptide, polypeptide, or protein, wherein Xaa may represent any amino acid residue selected from the group consisting of Ala, Arg, Asn, Asp, Cys, Gln, Glu, Gly, His, Ile, Leu, Lys, Met, Phe, Pro, Ser, Thr, Trp, Tyr, and Val. It will be understood that the modified dipeptide cleavase of the present disclosure may be unspecific as to the amino acid sequence of the peptide, polypeptide, or protein to be cleaved. In some embodiments, the modified dipeptide cleavase is partially specific or selective. In some aspects, the modified dipeptide cleavase preferentially cleaves or removes some amino acids at the P1 and / or P2 position of the peptide over others. In some cases, the modified dipeptide cleavase preferentially cleaves or removes a class of amino acids over others, e.g., preferentially removing hydrophobic amino acids over other classes of amino acids. In some aspects, the modified dipeptide cleavase may also have a preference for one or more amino acids at the second, third, fourth, fifth, etc. positions from the terminal amino acid. In some cases, the modified dipeptide cleavase exhibits specificity to subsets of amino acids and preferentially removes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15 or more specific terminal amino acid over others.

[0125] In some embodiments, the modified dipeptide cleavase is a polypeptide having an altered amino acid sequence, relative to an unmodified or wild-type dipeptide cleavase. In some cases, the modified dipeptide cleavase is a polypeptide which differs from a wild-type dipeptide cleavase sequence by one or more amino acid substitutions, deletions, additions, or combinations thereof. A variant or modified dipeptide cleavase can contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more mutations, e.g., amino acid differences, compared to the wild-type cleavase.

[0126] In some embodiments, the variant or modified dipeptide cleavase polypeptide generally exhibits at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a corresponding wild-type or unmodified dipeptide cleavase. In some embodiments, the wild-type or unmodified dipeptide cleavase comprises the amino acid sequence of any one of SEQ ID NO: 5-8, 10-16, 20, 33, 34, a mature sequence thereof that excludes a signal peptide, or a portion thereof containing the active site. In some embodiments, the variant or modified dipeptide cleavase polypeptide generally exhibits at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a wild-type or unmodified dipeptide cleavase set forth in SEQ ID NOs: 5-8, 10-16, 20, 33, 34, or a mature sequence thereof that excludes a signal peptide.

[0127] It is within the level of a skilled artisan to identify the corresponding position of a mutation or modification, e.g., amino acid substitution, in a dipeptide cleavase polypeptide, including a portion thereof, such as by alignment with a reference sequence. In some embodiments, the unmodified or reference dipeptide cleavase polypeptide generally exhibits at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to any of the sequences set forth in SEQ ID NOs: 5-8 and 10-16, 20, 33, 34. For example, corresponding residues can be determined by alignment of a reference sequence with a sequence provided herein (for example, sequences set forth in SEQ ID NOs: 5-8, 10-16, 20, 33, 34, or a functional homolog or fragment thereof) using known alignment methods. By aligning the sequences, one skilled in the art can identify corresponding residues, for example, using conserved and identical amino acid residues as guides. In some cases, while the numbering (positions) of the residues provided herein may differ from a reference sequence, using the alignment method will allow determination of corresponding residues.

[0128] In some embodiments, the modified dipeptide cleavase comprises a mutation, e.g., one or more amino acid modification(s), in an unmodified dipeptide cleavase, wherein the unmodified dipeptide cleavase is a dipeptidyl peptidase 3. Dipeptidyl peptidase 3 (also known as dipeptidyl peptidase III, dipeptidyl aminopeptidase III, dipeptidyl arylamidase III, enkephalinase B, red cell angiotensinase, DPP3, or DPP III) is a metalloproteinase (zinc-dependent) that sequentially removes dipeptides (two amino acid residues) from the N-terminus of short peptides. Wild-type or unmodified DPP3 is classified in the M49 family (MEROPS database identifier M49.001). In some cases, the unmodified dipeptidyl peptidase 3 exhibits at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to UniProt Accession No. Q9NY33 as set forth in SEQ ID NO: 5, UniProt Accession No. Q08225 as set forth in SEQ ID NO: 6, UniProt Accession No. Q8A6N1 as set forth in SEQ ID NO: 7, UniProt Accession No. H1XW48 as set forth in SEQ ID NO: 8, or UniProt Accession No. O55096 as set forth in SEQ ID NO: 15 (See e.g., Prajapati et al., FEBS J. 2011; 278(18):3256-276; Fukasawa et al., Biochem J. 1998 Jan 15; 329(Pt 2): 275-282; Fukasawa et al., J Amino Acids. 2011; 2011: 574816). DPP3 preferentially digests peptides that are 3 to 10 amino acids in length. DPP3 harbors a unique HEXXGH catalytic motif (SEQ ID NO: 1). Both histidines in this motif along with the glutamate residue of a second conserved EEXRAE / D motif are involved in zinc coordination (SEQ ID NO: 2). In some cases, C-terminal peptide modifications do not affect the activity of DPP3 enzymes. See Kumar et al., Sci Rep. (2016) 6:23787. Several substrate-bound structures of DPP3 have been solved, including in complex with peptides that have an N-terminal tyrosine. An N-terminal tyrosine is structurally similar to a phenylisothiocyanate (PITC), nitro-PITC, sulfo-PITC, or a phenylisocyanate version of these modifiers, and these substrate-bound structures may be useful for a targeted active-site design approach. In some embodiments, provided is a modified dipeptide cleavase derived from a dipeptidyl peptidase 3 that cleaves labeled terminal amino dipeptides, e.g., dipeptides containing a labeled N-terminal amino acid residue.

[0129] In some embodiments, the modified dipeptide cleavase comprises a mutation, e.g., one or more amino acid modification(s), in an unmodified dipeptide cleavase, wherein the unmodified dipeptide cleavase is a dipeptidyl peptidase 5. Dipeptidyl peptidase 5 is also known as allergen Tri m 4 (Trichophyton mentagrophytes), allergen Tri r 4 (Trichophyton rubrum), allergen Tri t 4 (Trichophyton tonsurans), dipeptidyl-peptidase V, DPP V, and secreted alanyl dipeptidyl peptidase (Aspergillus oryzae). Wild-type or unmodified dipeptidyl peptidase 5 is classified in the peptidase family S9 (MEROPS database identifier S09.012). Wild-type or unmodified dipeptidyl peptidase 5 has been observed to catalyze the hydrolysis of X-Ala, His-Ser, and Ser-Tyr dipeptides at a neutral pH optimum (See e.g., Beauvais et al., J Biol Chem. 1997; 272(10):6238-44). Wild-type or unmodified dipeptidyl peptidase 5 is described as a secreted dipeptidyl peptidase which contains the consensus sequences of the catalytic site of the nonclassical serine proteases. In some cases, the unmodified dipeptide cleavase exhibits at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to UniProt Accession No. P0C959 as set forth in SEQ ID NO: 10 or UniProt Accession No. B2RIT0 as set forth in SEQ ID NO: 16. In some embodiments, the mutations, e.g., one or more amino acid modifications (e.g., substitutions, deletions, additions) in the modified dipeptide cleavase is in reference to the amino acid sequence set forth in reference to positions of SEQ ID NO: 10 or 16.

[0130] In some embodiments, the modified dipeptide cleavase comprises a mutation, e.g., one or more amino acid modification(s), in an unmodified dipeptide cleavase, wherein the unmodified dipeptide cleavase is a dipeptidyl peptidase 7 (DPP7). Wild-type or unmodified DPP7 is classified in S46 protease family (MEROPS database identifier S46.001). Wild-type or unmodified DPP7 has been observed to catalyze the removal of dipeptides from the N-terminus of oligopeptides, including a broad specificity for both aliphatic and aromatic residues in the P1 position, with glycine or proline being not acceptable in this position (See e.g., Banbula et al., J. Biol. Chem. 2001, 276:6299-6305). DPP7 has been shown to exhibit activity for cleaving the synthetic substrates Met-Leu-methylcoumaryl-7-amide (Met-Leu-MCA), Leu-Arg-MCA, and Lys-Ala-MCA (Rouf et al., FEBS Open Bio. 2013; 3:177-83). In some cases, the unmodified dipeptide cleavase exhibits at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to UniProt Accession No. B2RKV3 as set forth in SEQ ID NO: 11. In some embodiments, the mutations, e.g., one or more amino acid modification(s) (e.g., substitutions, deletions, additions) in the modified dipeptide cleavase is in reference to the amino acid sequence set forth in reference to positions of SEQ ID NO: 11.

[0131] In some embodiments, the modified dipeptide cleavase comprises a mutation, e.g., one or more amino acid modification(s), in an unmodified dipeptide cleavase, wherein the unmodified dipeptide cleavase is a dipeptidyl peptidase 11. Dipeptidyl peptidase 11 is also known as Asp / Glu-specific dipeptidyl-peptidase or DPP11. Wild-type or unmodified dipeptidyl peptidase 11 is classified in S46 protease family (MEROPS database identifier S46.002), and shares 38.7% sequence identity with dipeptidyl peptidase 7. Wild-type or unmodified dipeptidyl peptidase 11 has been observed to catalyze the removal of dipeptides from the N-terminus of oligopeptides, including removing dipeptides from oligopeptides with the penultimate N-terminal Asp and Glu and has a P2-position preference to hydrophobic residues (See e.g., Ohara-Nemoto et al., J Biol Chem. 2011; 286(44):38115-27). In some cases, the unmodified dipeptide cleavase exhibits at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to UniProt Accession No. B2RID1 or F8WQK8 as set forth in SEQ ID NO: 12 and 14, respectively. In some embodiments, the mutations, e.g., one or more amino acid modifications (e.g., substitutions, deletions, additions) in the modified dipeptide cleavase is in reference to the amino acid sequence set forth in reference to positions of SEQ ID NO: 12. In some embodiments, the mutations, e.g., one or more amino acid modifications (e.g., substitutions, deletions, additions) in the modified dipeptide cleavase is in reference to the amino acid sequence set forth in reference to positions of SEQ ID NO: 14.

[0132] In some embodiments, the modified dipeptide cleavase comprises a mutation, e.g., one or more amino acid modification(s), in an unmodified dipeptide cleavase, wherein the unmodified dipeptide cleavase is a dipeptidyl aminopeptidase BII (DAP BII or dipeptidyl peptidase BII). Wild-type or unmodified DAP BII catalyzes the removal of dipeptides from the amino terminus of peptides (See e.g., Ogasawara et al., J. Bacteriol. 1996, 178:6288-6295); Sakamoto et al., Scientific Reports 2014, 4:4977). DAP BII is a serine protease that belongs to the serine peptidase family S46 (MEROPS database identifier S46.003). The amino acid sequence of the catalytic unit of DAP BII exhibits significant similarity to those classified in the clan PA endopeptidases. In some cases, the unmodified dipeptide cleavase exhibits at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to UniProt Accession No. V5YM14 as set forth in SEQ ID NO: 13. In some embodiments, the modified dipeptide cleavase contains one or more amino acids modifications in the catalytic domain of an unmodified DAP BII (e.g., residues 1-252 and residues 550 to 698 of SEQ ID NO: 13). In some embodiments, the mutations, e.g., one or more amino acid modifications (e.g., substitutions, deletions, additions) in the modified dipeptide cleavase is in reference to the amino acid sequence set forth in reference to positions of SEQ ID NO: 13. It has been shown that the unmodified or wild-type DAP BII hydrolyses peptides from the N-terminus of oligopeptides and small proteins, cleaving dipeptide units (NH2-P2-P1-) when the second (P1) residue is Ala, Leu, Ile, Phe, Tyr, Arg, or His (but not Pro) (See e.g., Sakamoto et al., Scientific Reports 2014, 4:4977).

[0133] In some embodiments, the modified dipeptide cleavase is derived from DAP BII and removes or is configured to remove a labeled terminal dipeptide from a polypeptide. In some embodiments, the modified dipeptide cleavase has one or more amino acid modifications (e.g. substitutions, deletions, additions, or combinations thereof) in an unmodified DAP BII cleavase corresponding to any one or more of positions 126, 188, 189, 190, 191, 192, 196, 238, 302, 306, 307, 310, 525, 528, 546, 604, 650, 651, 665, and / or 692, with reference to positions of SEQ ID NO: 13. In some embodiments, the modified dipeptide cleavase comprises one or more amino acid modifications in an unmodified dipeptide cleavase, corresponding to positions 126, 188, 189, 190, 191, 192, 196, 238, 302, 306, 307, 310, 525, 528, 546, 604, 650, 651, 665, and / or 692, with reference to positions of SEQ ID NO: 13, and comprises an amino acid sequence that exhibits at least 30 % identity, at least 40 % identity, at least 50 % identity, at least 60 % identity, at least 70 % identity, at least 80 % identity, or at least 90 % or more identity to any of SEQ ID NOs: 17-19 or 23-28. In some embodiments, the modified dipeptide cleavase contains a functional fragment of any of the provided sequences (e.g., a functional fragment of any of SEQ ID NOs: 17-19 or 23-28).

[0134] In some embodiments, the modified dipeptide cleavase has one or more amino acid modifications (e.g. substitutions, deletions, additions, or combinations thereof) in an unmodified DAP BII cleavase or fragment thereof corresponding to any one or more of positions 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, and / or 202, with reference to positions of SEQ ID NO: 13. In some embodiments, the modified dipeptide cleavase has one or more amino acid modifications (e.g. substitutions, deletions, additions, or combinations thereof) in an unmodified DAP BII cleavase or fragment thereof corresponding to any one or more of positions 188, 189, 190, 191, 192, 302, and / or 310, with reference to positions of SEQ ID NO: 13. In some embodiments, the modified dipeptide cleavase has one or more amino acid modifications (e.g. substitutions, deletions, additions, or combinations thereof) in an unmodified DAP BII cleavase or fragment thereof corresponding to any one or more of positions 191, 192, 196, 306, and / or 650, with reference to positions of SEQ ID NO: 13. In some embodiments, the modified dipeptide cleavase has one or more amino acid modifications (e.g. substitutions, deletions, additions, or combinations thereof) in an unmodified DAP BII cleavase or fragment thereof corresponding to any one or more of positions 323-544 with reference to positions of SEQ ID NO: 13. In some embodiments, the modified dipeptide cleavase has one or more amino acid modifications (e.g. substitutions, deletions, additions, or combinations thereof) in an unmodified DAP BII cleavase or fragment thereof corresponding to any one or more of positions 310, 651, 655, and / or 656 with reference to positions of SEQ ID NO: 13. In some embodiments, the modified dipeptide cleavase has one or more amino acid modifications (e.g. substitutions, deletions, additions, or combinations thereof) in an unmodified DAP BII cleavase or fragment thereof corresponding to any one or more of positions 627, 628, 630, 648, 651, 655, and / or 669, with reference to positions of SEQ ID NO: 13.

[0135] In some embodiments, the modified dipeptide cleavase is derived from DAP BII and removes or is configured to remove a labeled terminal dipeptide from a polypeptide. In some embodiments, the modified dipeptide cleavase has one or more amino acid modifications (e.g. substitutions, deletions, additions, or combinations thereof) in an unmodified DAP BII cleavase corresponding to any one or more of positions 126, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 238, 302, 306, 307, 310, 525, 528, 546, 604, 627, 628, 630, 648, 650, 651, 655, 656, 665, 669, and / or 692, with reference to positions of SEQ ID NO: 13.

[0136] In some embodiments, the modified dipeptide cleavase has one or more amino acid substitutions selected from the group consisting of A126T, D188V, I189A, D190S, N191L, N191M, W192G, R196S, R196T, R196V, G238V, A302W, N306R, T307K, N310K, N525K, A528V, F546L, A604V, D650A, G651V, K6651, and / or K692N, with reference to positions of SEQ ID NO: 13, or a conservative amino acid substitution thereof. In some embodiments, the one or more amino acid modification is N191M / W192G / R196T / N306R / D650A, N191M / W192G / R196V / N306R / D650A, D188V / I189A / D190S / N191L / W192G / R196S / A302W / N310K / D650A, N191M / W192G / R196T / N306R / T307K / D650A, N191M / W192G / R196T / N306R / N525K / A528V / A604V / D650A / K692N, A126T / N191M / W192G / R196T / G238V / N306R / D650A, N191M / W192G / R196T / N306R / F546L / D650A, N191M / W192G / R196T / N306R / D650A / G651V / K665I, or N191M / W192G / R196T / N306R / D650A / G651V.

[0137] In some embodiments, the modified dipeptide cleavase has an amino acid sequence that has at least at least 30 % identity, at least 40 % identity, at least 50 % identity, at least 60 % identity, at least 70 % identity, at least 80 % identity, or at least 90 % or more identity to any of SEQ ID NOs: 17-19, 23-28, or a specific binding fragment thereof. In some embodiments, the specific binding fragment has a length ranging from about 10 amino acids to about 400 amino acids, from about 10 amino acids to about 300 amino acids, from about 10 amino acids to about 200 amino acids, from about 10 amino acids to about 100 amino acids, or from about 10 amino acids to about 50 amino acids. In some specific examples, the modified dipeptide cleavase comprises the sequence of amino acids set forth in any of SEQ ID NOs: 17-19, 23-28, or a sequence of amino acids that exhibits at least 95% sequence identity to any of SEQ ID NOs: 17-19, 23-28, or a specific binding fragment thereof. In some examples, the modified dipeptide cleavase contains one or more of the amino acid substitutions provided in SEQ ID NOs: 17-19 or 23-28. In some aspects, the modified dipeptide cleavase comprises one or more amino acid modifications in an unmodified dipeptide cleavase, corresponding to positions 188, 189, 190, 191, 192, 196, 302, 306, 310, and / or 650, with reference to positions of SEQ ID NO: 13, and has an amino acid sequence that has at least 30 % identity, at least 40 % identity, at least 50 % identity, at least 60 % identity, at least 70 % identity, at least 80 % identity, or at least 90 % or more identity to any of SEQ ID NOs: 17-19 or 23-28. In some aspects, the modified dipeptide cleavase comprises one or more amino acid modifications in an unmodified dipeptide cleavase, corresponding to positions 191, 192, 196, 306, and / or 650, with reference to positions of SEQ ID NO: 13, and has an amino acid sequence that has at least 30 % identity, at least 40 % identity, at least 50 % identity, at least 60 % identity, at least 70 % identity, at least 80 % identity, or at least 90 % or more identity to any of SEQ ID NOs: 17-19 or 23-28. In some embodiments, the modified dipeptide cleavase exhibits the substrate specificity of any one of the sequences in SEQ ID NOs: 17-19 or 23-28. In some embodiments, the modified dipeptide cleavase has the cleaving activity of any one of the sequences in SEQ ID NOs: 17-19 or 23-28.

[0138] In some embodiments, the modified dipeptide cleavase has an amino acid sequence that comprises a catalytic domain with at least at least 30 % identity, at least 40 % identity, at least 50 % identity, at least 60 % identity, at least 70 % identity, at least 80 % identity, or at least 90 % or more identity with the catalytic domain of any of SEQ ID NOs: 17-19 or 23-28. In some embodiments, the modified dipeptide cleavase has an amino acid sequence that comprises an amine binding site with at least at least 30 % identity, at least 40 % identity, at least 50 % identity, at least 60 % identity, at least 70 % identity, at least 80 % identity, or at least 90 % or more identity with the amine binding site of any of SEQ ID NOs: 17-19 or 23-28. In some embodiments, the modified dipeptide cleavase has an amino acid sequence that comprises a loop domain with at least 30 % identity, at least 40 % identity, at least 50 % identity, at least 60 % identity, at least 70 % identity, at least 80 % identity, or at least 90 % or more identity with the loop domain of any of SEQ ID NOs: 17-19 or 23-28.

[0139] In some specific examples, a desired modified dipeptide cleavase may exhibit reduced bias towards specific amino acids in the P1 or P2 position of the polypeptide. In some embodiments, such modified dipeptide cleavases may be obtained by targeting the P1 or P1 pocket of the wildtype or unmodified enzyme in genetic selection. In some cases, the residues N310, G651, S655, and / or V656 with reference to positions of SEQ ID NO: 13 may be targeted to reduce bias. In some cases, the residues N215, W216, R220, N330, and / or D674 with reference to positions of SEQ ID NO: 20 may be targeted to reduce bias.

[0140] Table 6 also provides exemplary amino acid substitutions in sequences of exemplary modified cleavases by reference to positions of the indicated SEQ ID NOs. In some examples, the modified cleavase contains one or more of the amino acid substitutions provided in Table 6. Table 6. Exemplary Modified Cleavases Mutation(s)SEQ ID NOD188V / I189A / D1905 / N191L / W192G / R196S / A302W / N310K / D650A17N191M / W192G / R196T / N306R / D650A18N191M / W192G / R196V / N306R / D650A19N191M1W192G / R196T / N306R / T307K / D650A23N191M1W192G / R196T / N306R / N525K / A528V / A604V / D650A / K692N24A126T / N191M / W192G / R196T / G238V / N306R / D650A25N191M1W192G / R196T / N306R / F546L / D650A26N191M1W192G / R196T / N306R / D650A / G651V / K665I27N191M1W192G / R196T / N306R / D650A / G651V28

[0141] In some embodiments, the removed dipeptide comprises an amino acid that is labeled or modified by a chemical reagent or enzymatic reagent. In some embodiments, selection of the appropriate label with an appropriately engineered or modified dipeptide cleavase enables cleavage of a labeled terminal dipeptide. In some embodiments, the active site and / or amino acid binding site(s) of the unmodified dipeptide cleavase is modified. In some embodiments, the modified dipeptide cleavase comprises a mutation or modification within its substrate binding site, at the boundary of the substrate binding site, or a combination thereof. In some embodiments, the modified dipeptide cleavase is derived from a dipeptide cleavase (e.g., a dipeptidyl peptidase, a dipeptidyl aminopeptidase, a peptidyl-dipeptidase, or a dipeptidyl carboxypeptidase) modified to fit and recognize a label (e.g. a chemical label or a chemical modification).

[0142] In some embodiments, the modified dipeptide cleavase comprises an amino acid mutation (e.g., modifications, substitutions, deletions, additions, or combinations thereof) compared to the wild-type dipeptide cleavase in its substrate binding site, at the boundary of the substrate binding site, in the catalytic domain, in the P1 or P2 pocket, in a chymotrypsin fold, at an amine binding site, in the loop domain, or a combination thereof. In some embodiments, the modified dipeptide cleavase comprises an amino acid mutation (e.g., substitutions, deletions, additions, or combinations thereof) compared to the wild-type dipeptide cleavase polypeptide in the hinge region of the cleavase. In some embodiments, the modified dipeptide cleavase comprises an amino acid mutation compared to the wild-type dipeptide cleavase polypeptide in the binding cleft of the cleavase. In some cases, the modified dipeptide cleavase comprises an amino acid mutation compared to the wild-type dipeptide cleavase polypeptide in the inter-lobe cleft of the cleavase. In some embodiments, the modified dipeptide cleavase comprises an amino acid mutation compared to the wild-type dipeptide cleavase polypeptide in the alpha amine binding region of the cleavase. For example, the modified dipeptide cleavase exhibits reduced alpha amine binding compared to the wild-type cleavase polypeptide. See e.g., Kumar et al., Sci Rep. (2016) 6:23787.

[0143] In some embodiments, the modified dipeptide cleavase comprises an amino acid mutation (e.g., modifications, substitutions, deletions, additions, or combinations thereof) compared to the wild-type dipeptide cleavase polypeptide in the chymotrypsin fold of the cleavase. In some embodiments, the modified dipeptide cleavase comprises an amino acid mutation (e.g., modifications, substitutions, deletions, additions, or combinations thereof) compared to the wild-type dipeptide cleavase polypeptide in at an amine binding site. In some embodiments, the modified dipeptide cleavase comprises an amino acid mutation (e.g., modifications, substitutions, deletions, additions, or combinations thereof) compared to the wild-type dipeptide cleavase polypeptide in the loop domain. In some aspects, the modified dipeptide cleavase comprises an amino acid mutation (e.g., modifications, substitutions, deletions, additions, or combinations thereof) compared to the wild-type dipeptide cleavase polypeptide for improving accessibility to the active site of the modified dipeptide cleavase. In some cases, the modified dipeptide cleavase exhibits greater accessibility of the substrate (e.g., polypeptide) to the active site compared to the unmodified dipeptide cleavase. For example, the modified dipeptide cleavase may allow larger substrates to access the active site.

[0144] In some embodiments, the modified dipeptide cleavase exhibits altered activity, substrate binding capability, or cleavage characteristics compared to the unmodified dipeptide cleavase. In some embodiments, the modified dipeptide cleavase is modified in the catalytic motif or catalytic domain of the unmodified dipeptide cleavase (e.g., the HEXXGH catalytic motif as set forth in SEQ ID NO: 1). In some embodiments, the mutations, e.g., one or more amino acid modifications (e.g., substitutions, deletions, additions) corresponds to positions 316, 391, and / or 394 with reference to positions of SEQ ID NO: 5. In some embodiments, the mutations, e.g., one or more amino acid modifications (e.g., substitutions, deletions, additions) corresponds to amino acid residue positions 419, 420, 421, 422, 423, 424, 425, 426, or a combination thereof, with reference to positions of SEQ ID NO: 5.

[0145] In some embodiments, the unmodified dipeptide cleavase is a metallopeptidase. In some embodiments, the modified dipeptide cleavase is a metallopeptidase. In some embodiments, the modified dipeptide cleavase is a zinc-dependent metallopeptidase or a zinc-dependent hydrolase or derived from such. Some known metallopeptidase are characterized by the presence of a conventional catalytic signature motif HEXXH. In some aspects, the two His residues of the HEXXH motif contribute to coordinate the divalent metal ion (e.g., Zn 2+< , Mn 2+< , Co 2+< , Ni 2+< , Cu 2+< ). For example, the modified dipeptide cleavase requires the presence of or contact with specific metal ions (e.g., zinc ions, chloride ions) for activation. In some embodiments, function of the modified dipeptide cleavase can be modulated or controlled by the presence or absence of metal ions, or by contacting with metal chelating agents.

[0146] In some embodiments, the modified dipeptide cleavase exhibits altered binding affinity and / or specificity to specific substrates compared to the unmodified dipeptide cleavase. For example, the modified dipeptide cleavase exhibits increased binding affinity and / or specificity for labeled terminal amino acids compared to the unmodified dipeptide cleavase. For example, the modified dipeptide cleavase may not remove an unlabeled terminal dipeptide from the polypeptide. In some cases, the modified dipeptide cleavase does not remove a terminal dipeptide that does not contain a labeled amino acid. In comparison, a wild-type or unmodified dipeptide cleavase does not remove dipeptides comprising a labeled amino acid. In some embodiments, the modified dipeptide cleavase exhibits decreased binding affinity and / or specificity for a substrate compared to the unmodified dipeptide cleavase. In some embodiments, the modified dipeptide cleavase exhibits one or more desired characteristics, such as binding rate, rate of hydrolysis, rate of release. In some embodiments, the modified dipeptide cleavase can be removed or released from the polypeptide at a desired rate.

[0147] In some of any such embodiments, a reaction with a modified dipeptide cleavase can be enhanced by recruiting the modified dipeptide cleavase to the labeled terminal amino acid. For example, one or more modified dipeptide cleavases can be recruited to the labeled terminal amino acid of the polypeptide via hybridization of complementary universal priming sequences a DNA tag or sequence associated with the modified dipeptide cleavase and a DNA tag or sequence associated with the polypeptide to be treated with the modified dipeptide cleavase(s). This hybridization step may improve the effective affinity of the modified dipeptide cleavase for the labeled terminal amino acid (e.g., NTAA). In some cases, after the labeled terminal amino acid is removed as a dipeptide, it may diffuse away, and the associated modified dipeptide cleavase can be removed by stripping the hybridized DNA tag.

[0148] In some embodiments, the modified dipeptide cleavase is attached to an anchoring sequence. In some cases, the modified dipeptide cleavase is attached to the anchoring sequence directly or indirectly. In some cases, the anchoring sequence is complementary to a sequence attached to the polypeptides. In some embodiments, the anchoring sequence is a universal sequence or a universal DNA tag. In some embodiments, the polypeptide is also attached to a universal sequence. In some examples, the anchoring sequence on the modified dipeptide cleavase brings the enzyme in proximity to the polypeptide. In some embodiments, the anchoring sequence brings the enzyme in proximity or co-localizes the modified dipeptide cleavase to the polypeptide. In some embodiments, this co-localization of the modified dipeptide cleavase and the polypeptide aids in binding and / or removal of the labeled dipeptide from said polypeptide.

[0149] In any of the embodiments provided herein, recruitment of one or more modified dipeptide cleavases to the terminal amino acid of the polypeptide may be enhanced by utilizing a chimeric modified dipeptide cleavase containing a first tethering moiety and a second tethering moiety associated with the polypeptide, or is colocalized with the polypeptide, wherein the first tethering moiety is configured to form a stable complex with the second tethering moiety upon contact (moieties are capable of a binding reaction with each other). For example, the polypeptide may be immobilized on a solid support, and the second tethering moiety is attached to the solid support in proximity to the polypeptide, or attached directly to the polypeptide. Examples of first and second tethering moieties include biotin-streptavidin pair, two complementary polynucleotide molecules that form a stable double strand complex upon contact, or other known in the art molecules that can strongly interact upon contact under physiological conditions (or under conditions used in a cleavase assay). In one example, a modified dipeptide cleavase is a low affinity enzyme (> µM Kd) and it is recruited to the polypeptide associated with a biotin using a streptavidin-chimeric modified dipeptide cleavase. In some cases, the efficiency of modified dipeptide cleavase to remove labeled terminal amino acid can be improved due to the increase in effective local concentration as a result of the first tethering moiety-second tethering moiety interaction. In some cases, this approach effectively increases the affinity KD of the modified dipeptide cleavase from µM to subpicomolar. A number of different bioconjugation recruitment strategies can also be employed. An azide modified PITC is commercially available (4-Azidophenyl isothiocyanate, Sigma), allowing a number of simple transformations of azide-PITC into other bioconjugates of PITC, such as biotin-PITC via a click chemistry reaction with alkyne-biotin. In some aspects, after the labeled terminal amino acid is removed, it may diffuse away with the associated modified dipeptide cleavase from the polypeptide.

[0150] In some embodiments, the modified dipeptide cleavase can be a single polypeptide chain or a multimer (dimers or higher order multimers) of at least two polypeptide chains. Thus, monomeric, dimeric, and higher order multimeric modified dipeptide cleavase polypeptides are within the scope of the defined term. Multimeric polypeptides can be homomultimeric (of identical polypeptide chains) or heteromultimeric (of non-identical polypeptide chains). In some embodiments, the modified dipeptide cleavase is a monomeric enzyme. In some embodiments, the modified dipeptide cleavase is a fusion molecule or a chimeric molecule. For example, the modified dipeptide cleavase may be attached or associated, directly or indirectly via a linker, to a oligonucleotide. In some specific cases, the modified dipeptide cleavase may be joined to a moiety such as a SpyTag / SpyCatcher or SnoopTag / SnoopCatcher.A. Labeled Terminal Amino Acid

[0151] In some embodiments, the terminal amino acid of a peptide removed by the modified dipeptide cleavase is labeled or modified. A label can comprise any suitable material or moiety. Any suitable molecule or materials may be employed for this purpose, including proteins, amino acids, nucleic acids, carbohydrates, chemical moieties, and small molecules. In some embodiments, a suitable label is capable of fitting in the binding pocket of the modified dipeptide cleavase. In some aspects, the labeling of a terminal amino acid is performed in a manner that is nucleic acid-compatible (e.g., the labeling is performed in a manner that is not damaging to nucleic acids). In some embodiments, a suitable label enables the modified dipeptide cleavase to remove a labeled dipeptide from the polypeptide. The terminal amino acid of the polypeptides may be labeled by any suitable methods. In some examples, the terminal amino acid is labeled chemically or enzymatically. In some embodiments, the terminal amino acid is labeled by a reagent that is or comprises a chemical agent, an enzyme, and / or a biological agent. In some cases, the terminal amino acid is labeled with a chemical label or moiety.

[0152] In some embodiments, a precursor polypeptide (e.g., an unlabeled polypeptide) is contacted with a reagent for labeling the terminal amino acid of the precursor polypeptide to provide a polypeptide prepared for treatment with the modified dipeptide cleavase. In some cases, the contacting of the precursor polypeptide with the reagent for labeling the terminal amino acid is performed prior to contacting the polypeptide with a modified dipeptide cleavase. In some aspects, the modified dipeptide cleavase is contacted with a polypeptide that has been labeled or modified. In some cases, the contacting of the precursor polypeptide with the reagent for labeling the terminal amino acid and contacting the polypeptide with a modified dipeptide cleavase are performed simultaneously or substantially simultaneously.

[0153] In some embodiments, the dipeptide for removal or removed by the modified dipeptide cleavase comprises an amino acid that is labeled with a chemical label. In some examples, the amino acid in the dipeptide for removal by the modified dipeptide cleavase is labeled with a chemical reagent. In some aspects, the labeling of a terminal amino acid by treating with a chemical reagent is performed in a manner that is nucleic acid-compatible (e.g., the labeling is performed under conditions that is not damaging to nucleic acids).

[0154] In some embodiments, the modified dipeptide cleavase removes dipeptides comprising amino acid(s) that are labeled, such as a chemically-modified or labeled (e.g., PTC / DNP / acetyl / Cbz-modified or labeled) amino acids on a polypeptide. In some cases, the labeled amino acid is removed as part of a terminal dipeptide. In some embodiments, the modified dipeptide cleavase removes a dipeptide comprising an N-terminal amino acid having a PTC / DNP / acetyl / Cbz group present as the label.

[0155] In some embodiments, at least one amino acid (as part of a dipeptide) for removal by the modified dipeptide cleavase, which may be the terminal amino acid of a dipeptide to be removed by the dipeptide cleavase, is labeled with a reagent selected from the group consisting of a phenyl isothiocyanate (PITC), a nitro-PITC, a sulfo-PITC, a phenyl isocyanate (PIC), a nitro-PIC, a sulfo-PIC, benzyloxycarbonyl chloride or carbobenzoxy chloride (Cbz-Cl), N-(Benzyloxycarbonyloxy)succinimide (Cbz-OSu or Cbz-O-NHS), a 1-fluoro-2,4-dinitrobenzene (Sanger's reagent, DNFB), dansyl chloride (DNS-Cl, or 1-dimethylaminonaphthalene-5-sulfonyl chloride), 4-sulfonyl-2-nitrofluorobenzene (SNFB), an anhydride, 2-Pyridinecarboxaldehyde, 2-Formylphenylboronic acid, 2-Acetylphenylboronic acid, 1-Fluoro-2,4-dinitrobenzene, 4-Chloro-7-nitrobenzofurazan, Pentafluorophenylisothiocyanate, 4-(Trifluoromethoxy)-phenylisothiocyanate, 4-(Trifluoromethyl)-phenylisothiocyanate, 3-(Carboxylic acid)-phenylisothiocyanate, 3-(Trifluoromethyl)-phenylisothiocyanate, 1-Naphthylisothiocyanate, N-nitroimidazole-1-carboximidamide, N,N'-Bis(pivaloyl)-1H-pyrazole-1-carboxamidine, N,N'-Bis(benzyloxycarbonyl)-1H-pyrazole-1-carboxamidine, an acetylating reagent, a guanidinylation reagent, a thioacylation reagent, a thioacetylation reagent, a thiobenzylation reagent, and a diheterocyclic methanimine reagent, or a derivative thereof.

[0156] In some embodiments, the terminal amino acid for removal by the modified dipeptide cleavase, which may be the terminal amine of a dipeptide to be removed by the dipeptide cleavase, is labeled with an anhydride or derivative thereof. In some embodiments, the reagent for labeling the amino acid for removal by the modified dipeptide cleavase is selected from the group consisting of: S-Acetylmercaptosuccinic anhydride, cis-Aconitic anhydride, 4-Amino-1,8-naphthalic anhydride, endo-Bicyclo[2.2.2]oct-5-ene-2,3-dicarboxylic anhydride, 5-Bromoisatoic anhydride, Bromomaleic anhydride, 4-Bromo-1,8-naphthalic anhydride, Citraconic anhydride, Crotonic anhydride, trans-1,2-Cyclohexanedicarboxylic anhydride, 1-Cyclopentene-1,2-dicarboxylic anhydride, 2,3-Dichloromaleic anhydride, 3,6-Dichlorophthalic anhydride, 3,6-Difluorophthalic anhydride, Diglycolic anhydride, 2,2-Dimethylglutaric anhydride, 3,3-Dimethylglutaric anhydride, 2,3-Dimethylmaleic anhydride, 2,2-Dimethylsuccinic anhydride, (2-Dodecen-1-yl)succinic anhydride, Dodecenylsuccinic anhydride, Glutaric anhydride, Hexafluoroglutaric anhydride, Hexahydro-4-methylphthalic anhydride, Homophthalic anhydride, 3-Hydroxyphthalic anhydride, Itaconic anhydride, Maleic anhydride, 3-Methylglutaric anhydride, N-Methylisatoic anhydride, Methylsuccinic anhydride, 1,8-Naphthalic anhydride, 3-Nitro-1,8-naphthalic anhydride, 4-Nitro-1,8-naphthalic anhydride, 3-Nitrophthalic anhydride, 4-Nitrophthalic anhydride, 2-Octen-1-ylsuccinic anhydride, 2,5-Oxazolidinedione, 2-Phenylglutaric anhydride, Phenylmaleic anhydride, Phenylsuccinic anhydride, N-Phthaloyl-DL-glutamic anhydride, 2,3-Pyrazinedicarboxylic anhydride, 3,4-Pyridinedicarboxylic anhydride, Succinic anhydride, 4-Sulfo-1,8-naphthalic anhydride, Tetrabromophthalic anhydride, Tetrachlorophthalic anhydride, Tetrafluorophthalic anhydride, 3,4,5,6-Tetrahydrophthalic anhydride, 3,3-Tetramethyleneglutaric anhydride, Trimellitic anhydride chloride, and 2-(Triphenylphosphoranylidene)succinic anhydride. See e.g. Staiger et al., J. Org. Chem. 1959, 24, 9, 1214-1219; Jiang et al. J. Org. Chem. 2019, 84, 4, 2022-2031; U.S. Patent No. 9,867,883.

[0157] In some examples, in preparation for treatment with a modified dipeptide cleavase of the invention, a polypeptide is treated with a chemical reagent that comprises an isatoic anhydride, an isonicotinic anhydride, an azaisatoic anhydride, a succinic anhydride, or a derivative of one of these, and the terminal amino acid of the polypeptide is modified, or labeled, by the chemical reagent. Specific examples of labeling of a terminal amino acid of a polypeptide to be treated with a modified dipeptide cleavase of the invention include: wherein: G 1< -G 4< are each independently selected from CH, CX, and N; X at each occurrence is independently selected from C 1 -C 2 alkyl,, NO 2 , C 1 -C 2 haloalkyl, C 1 -C 2 haloalkoxy, halo, -OR 2< , -N(R 2< ) 2 , -SR 2< , SO 2 R 3< , SO 3 R 2< , -B(OR 2< ) 2 , C(=O)R 2< , CN, CON(R 2< ) 2 , -COOR 2< , -C(-O)Ar, and tetrazole; R represents the side chain of an amino acid, e.g. one of the side chains of the 20 common amino acids; R 1< is selected from H, R 3< , C(-O)R 2< , -C(=O)N(R 2< ) 2 , -C(=O)Ar, and -SO 2 N(R 2< ) 2 ; R 2< is independently at each occurrence selected from H and C 1 -C 2 alkyl; R 3< is independently at each occurrence selected from C 1 -C 2 alkyl; Ar is independently selected at each occurrence from phenyl, pyridinyl, pyrimidinyl, pyridazinyl, and pyrazinyl, each of which is optionally substituted by one or two groups selected from halo, CN, NO 2 , C 1 -C 2 alkyl, C 1 -C 2 haloalkyl, C 1 -C 2 haloalkoxy, and -OR 2< ; and PP represents a portion of a polypeptide, particularly the portion of a polypeptide being prepared for treatment with a modified dipeptide cleavase of the invention excluding the N-terminal amino acid. Thus the compound of Formula (C) is typically a polypeptide for use in the methods of the invention, and R represents the side chain of the terminal amino acid of the polypeptide.

[0158] In preferred embodiments, the terminal amino acid shown in Formula (C) is in the L-configuration when R is not H.

[0159] The compounds of Formula (B) are polypeptides, sometimes referred to as labeled polypeptides, that have been prepared for use in the modified dipeptide cleavase reactions described herein.

[0160] In some aspects, the amino acid for removal by the modified dipeptide cleavase is labeled with an exemplary reagent derived from an isatoic anhydride, an isonicotinic anhydride or an azaisatoic anhydride, especially compounds of Formula (A) as described herein. In some embodiments, the amino acid for removal by the modified dipeptide cleavase is labeled with an exemplary reagent selected from the list consisting of N-Methyl-isatoic anhydride, N-acetyl-isatoic anhydride, 4-carboxylic acid isatoic anhydride, 5-methoxy-isatoic anhydride, 5-nitro-isatoic anhydride, 4-chloro-isatoic anhydride, 4-fluoro-isatoic anhydride, 6-fluoro-isatoic anhydride, N-benzyl-isatoic anhydride, 4-trifluoromethyl-isatoic anhydride, 5-trifluoromethyl-isatoic anhydride, 4-nitro-isatoic anhydride, 4-methoxy-isatoic anhydride, and 5-Amino-2-fluoro-isonicotinic anhydride (6-fluoro-1H-pyrido[3,4-d][1,3]oxazine-2,4-dione), or a derivative thereof. In some examples, the labeled amino acid or dipeptide removed by the action of a modified dipeptide cleavase of the invention comprises an optionally substituted benzamide, typically one derived from any of the optionally substituted isatoic anhydrides disclosed herein, including a compound of Formula (B) as described herein.

[0161] In other favored embodiments of the invention, in preparation for treatment with a modified dipeptide cleavase of the invention, the polypeptide is treated with a chemical reagent that comprises a succinic anhydride, a phthalic anhydride, a pyrazinedicarboxylic anhydride, or a derivative of one of these, and the terminal amino acid is modified, or labeled, by the chemical reagent.

[0162] Additional specific examples of reactions for labeling of a terminal amino acid of a polypeptide to be treated with a modified dipeptide cleavase of the invention include: wherein: n is 0 or 1; Ring Cy represents a 5- or 6-membered ring or an 8-10 membered bicyclic ring that may be absent or present; when present, ring Cy may be saturated, unsaturated, or aromatic, and the dashed bond may be a single bond, double bond, or aromatic bond; when Cy is present, it may be a carbocyclic ring, or it may contain one or two heteroatoms selected from N, O and S as ring members; when Ring Cy is present, it is optionally substituted with one to six groups (or with one to four groups when Cy is aromatic) selected from halo, CN, NO 2 , C 1 -C 2 alkyl, C 1 -C 2 haloalkyl, C 1 -C 2 haloalkoxy, and -OR 4< ; when ring Cy is absent, the dashed bond may be a single bond or a double bond, and the dashed bond is optionally substituted by one or two groups selected from halo, CN, C 1 -C 2 alkyl, C 1 -C 2 haloalkyl, C 1 -C 2 haloalkoxy, CO 2 R 4< , and -OR 4< ; R represents the side chain of an amino acid, e.g. one of the side chains of the 20 common amino acids; R 4< is independently selected at each occurrence from H, C 1 -C 2 alkyl, and C 1 -C 2 haloalkyl; R 5< is independently selected at each occurrence from H, halo, C 1 -C 2 alkyl, C 1 -C 2 haloalkyl, C 1 -C 2 alkoxy, and C 1 -C 2 haloalkoxy; PP represents a portion of a polypeptide, particularly the portion of a polypeptide being prepared for treatment with a modified dipeptide cleavase of the invention excluding the N-terminal amino acid. Thus the compound of Formula (C) is typically a polypeptide for use in the methods of the invention, and R represents the side chain of the terminal amino acid of the polypeptide.

[0163] In preferred embodiments of these reagents of Formula (D), ring Cy is absent when n is 1. In additional preferred embodiments of the chemical reagents of Formula (D), ring Cy is present and is a phenyl ring or a 2,3-pyrazine ring, each of which is optionally substituted as described above, and n is 0.

[0164] In preferred embodiments, the terminal amino acid shown in Formula (C) is in the L-configuration when R is not H.

[0165] The compounds of Formula (E) are polypeptides, sometimes referred to as labeled polypeptides, that have been prepared for use in the modified dipeptide cleavase reactions described herein.

[0166] In some embodiments, the amino acid for removal by the modified dipeptide cleavase is labeled with an exemplary reagent derived from succinic anhydride, or a compound of Formula (D) wherein ring Cy is absent and the dashed bond represents a single bond. In some embodiments, the reagent is 3,6, difluorophthalic anhydride, 2,3 pyrazinedicarboxylic anhydride, or succinic anhydride. In some examples, the removed labeled amino acid or dipeptide comprises 4-carboxybutylamide.

[0167] In some embodiments, in preparation for treatment with a modified dipeptide cleavase of the invention, a polypeptide is treated with any suitable chemical reagent that is capable of forming an amide bond with the α-amine of the polypeptide N-terminus. A number of chemical reagents react with terminal amines of the polypeptide to form a modified polypeptide with an amide bond linking the polypeptide to the modification; this N-terminal modified polypeptide can be a substrate for a modified dipeptide cleavase. Chemical reagents that react with amines to form an amide bond are known from the field of peptide coupling, including but not limited to: acyl halides (chlorides, fluorides, bromides), acyl imidazoles, O-acyl isoureas, activated esters [N-hydroxysuccinimide (NHS or HOSu),N-hydroxysulfosuccinimide (sulfo-NHS) p-nitrophenyl (PNP), Pentafluorophenyl (Pfp), 4-sulfo-2,3,5,6,-tetrafluorophenyl, 2,4,5-trichlorophenol, N-hydroxy-5-norbornene-2,3-dicarboximide (HONB), 3-hydroxy-4-oxo-3,4-dihydro-1,2,3-benzotriazine (HODhbt), hydroxybenzotriazole (HOBt), 1-hydroxy-7-azabenzotriazole (HOAt), 1-Hydroxy-1H-1,2,3-triazole-4-carboxylate (HOCt)], Ethyl (2Z)-2-cyano-2-hydroxyiminoacetate (Oxyma)], alkyl esters, carbodiimides, etc. (Hermanson (2013) Bioconjugation Techniques, Academic Press; Montalbetti et al., (2005) Tetrahedron 61: 10827-10852; Montalbetti et al., Wiley Encyclopedia of Chemical Biology: 1-17. de Figueiredo et al., (2016) Chem Rev 116(19): 12029-12122). N-terminal modifications can be installed with an amide bond linking to the polypeptide via enzymatic methods as well (Philpott et al., (2018) Green Chemistry 20(15): 3426-3431). An example of labeling a polypeptide with a PNP ester is provided with 4-Nitrophenyl Anthranilate which can be used to label a polypeptide under the following conditions: 4-Nitrophenol anthranilate (PNPA) is dissolved in DMSO at 100 mM; and PNPA used at 10 mM with 1 mM peptide in 1XPBS (pH 8.5) or 100 mM NaHCO3 carbonate buffer (pH 8.5) in 10% DMSO for 37 °C for 1 hr. The resulting peptide product generated is equivalent to labeling a peptide with isatoic anhydride, and generates a 2-aminobenzamide-modified peptide suitable as a substrate for a modified dipeptide cleavase (e.g., derived from DAP BII) as illustrated in Table 8.

[0168] In some examples, in preparation for treatment with a modified dipeptide cleavase of the invention, a polypeptide is treated with a chemical reagent that comprises an amine-protected activated ester to form an amide bond and the terminal amino acid of the polypeptide is modified, or labeled, by the chemical reagent. This modified polypeptide can then be further appropriately treated to remove the designated protecting group, yielding a modified polypeptide for treatment with a modified dipeptide cleavase.

[0169] Specific examples of labeling of a terminal amino acid of a polypeptide to be treated with a modified dipeptide cleavase of the invention include: wherein: G 1< -G 4< are each independently selected from CH, CX, and N; X at each occurrence is independently selected from H, C 1 -C 2 alkyl, NO 2 , C 1 -C 2 haloalkyl, C 1 -C 2 haloalkoxy, halo, -OR 2< , -N(R 2< ) 2 , -SR 2< , SO 2 R 3< , SO 3 R 2< , -B(OR 2< ) 2 , C(=O)R 2< , CN, CON(R 2< ) 2 , -COOR 2< , -C(-O)Ar, and tetrazole; R represents the side chain of an amino acid, e.g. one of the side chains of the 20 common amino acids; R 1< is selected from H, R 3< , C(-O)R 2< , -C(=O)N(R 2< ) 2 , -C(=O)Ar, and -SO 2 N(R 3< ) 2 ; R 2< is independently at each occurrence selected from H and C 1 -C 2 alkyl; R 3< is independently at each occurrence selected from C 1 -C 2 alkyl; Ar is independently selected at each occurrence from phenyl, pyridinyl, pyrimidinyl, pyridazinyl, and pyrazinyl, each of which is optionally substituted by one or two groups selected from halo, CN, NO 2 , C 1 -C 2 alkyl, C 1 -C 2 haloalkyl, C 1 -C 2 haloalkoxy, and -OR 2< ; L is a leaving group selected from halo, N-hydroxysuccinimide (NHS), N-hydroxybenzotriazole, sulfo N-hydroxysuccinimide (sulfoNHS), 2,3,4,5,6-pentafluorophenol (pFP), 4-sulfo-2,3,5,6-tetrafluoro phenol, chloro, 4-nitrophenol, and -O(C=)-O-(C1-6 alkyl); optionally, -NR 1< -PG can be replaced by -N 3 ; and PG is H or a nitrogen protecting groups which may be selected from tert-butyloxycarbonyl (Boc), 2,2,2-trichloroethoxycarbonyl (Troc), 2-(trimethylsilyl)ethoxycarbonyl (Teoc), carboxylbenzyl (Cbz), para-nitrocarboxylbenzyl (p-NO 2 Cbz), allyloxycarbonyl (Alloc), 9-fluorenylmethoxycarbonyl (Fmoc), para-azidocarboxylbenzyl (p-N 3 Cbz), 2,2,6,6-tetramethylpiperidin-1-yloxycarbonyl (Tempoc), and other N-protecting groups.

[0170] In some examples, in preparation for treatment with a modified dipeptide cleavase of the invention, a polypeptide is treated with 4-Nitrophenyl Anthranilate.

[0171] In some embodiments, the chemical reagent for modifying or labeling the amino acid for removal by the modified dipeptide cleavase is one or more of any of the compounds of Formula (A) or (D), described herein, or a salt or conjugate thereof.

[0172] In some embodiments, the chemical reagent for modifying or labeling the amino acid for removal as a dipeptide by the modified dipeptide cleavase is one or more of any of the compounds of Formula (I), (II), (III), (IV), or (AB), described herein, or a salt or conjugate thereof.

[0173] In some embodiments, the reagent for modifying or labeling the amino acid for removal as a dipeptide by the modified dipeptide cleavase comprises a compound selected from the group consisting of a compound of Formula (I): or a salt or conjugate thereof, wherein R 1< and R 2< are each independently H, C 1-6 alkyl, cycloalkyl, -C(O)R a< , -C(O)OR b< , or -S(O) 2 R c< ; R a< , R b< , and R c< are each independently H, C 1-6 alkyl, C 1-6 haloalkyl, arylalkyl, aryl, or heteroaryl, wherein the C 1-6 alkyl, C 1-6 haloalkyl, arylalkyl, aryl, and heteroaryl are each unsubstituted or substituted; R 3< is heteroaryl, -NR d< C(O)OR e< , or -SR f< , wherein the heteroaryl is unsubstituted or substituted; R d< , R e< , and R f< are each independently H or C 1-6 alkyl.

[0174] In some embodiments, when R 3< is R 1< and R 2< are not both H. In some embodiments of Formula (I), both R 1< and R 2< are H. In some embodiments, neither R 1< nor R 2< are H. In some embodiments, one of R 1< and R 2< is C 1-6 alkyl. In some embodiments, one of R 1< and R 2< is H, and the other is C 1-6 alkyl, cycloalkyl, -C(O)R a< , -C(O)OR b< , or -S(O) 2 R c< . In some embodiments, one or both of R 1< and R 2< is C 1-6 alkyl. In some embodiments, one or both of R 1< and R 2< is cycloalkyl. In some embodiments, one or both of R 1< and R 2< is -C(O)R a< . In some embodiments, one or both of R 1< and R 2< is -C(O)OR b< . In some embodiments, one or both of R 1< and R 2< is -S(O) 2 R c< . In some embodiments, one or both of R 1< and R 2< is -S(O) 2 R c< , wherein R c< is C 1-6 alkyl, C 1-6 haloalkyl, arylalkyl, aryl, or heteroaryl. In some embodiments, R 1< is In some embodiments, R 2< is In some embodiments, both R 1< and R 2< are In some embodiments, R 1< or R 2< is

[0175] In some embodiments of the compound of Formula (I), R 3< is a monocyclic heteroaryl group. In some embodiments of Formula (I), R 3< is a 5- or 6-membered monocyclic heteroaryl group. In some embodiments of Formula (I), R 3< is a 5- or 6-membered monocyclic heteroaryl group containing one or more N. Preferably, R 3< is selected from pyrazole, imidazole, triazole and tetrazole, and is linked to the amidine of Formula (I) via a nitrogen atom of the pyrazole, imidazole, triazole or tetrazole ring, and R 3< is optionally substituted by a group selected from halo, C 1-3 alkyl, C 1-3 haloalkyl, and nitro. In some embodiments, R 3< is wherein G 1 is N, CH, or CX where X is halo, C 1-3 alkyl, C 1-3 haloalkyl, or nitro. In some embodiments, R 3< is or , where X is Me, F, Cl, CF 3 , or NO 2 . In some embodiments, R 3< is wherein G 1 is N or CH. In some embodiments, R 3< is In some embodiments, R 3< is a bicyclic heteroaryl group. In some embodiments, R 3< is a 9- or 10-membered bicyclic heteroaryl group. In some embodiments, R 3< is

[0176] In some embodiments, the compound of Formula (I) is In some embodiments, the compound of Formula (I) is not

[0177] In some embodiments, the compound of Formula (I) is selected from the group consisting of and optionally also including (N-Boc,N'-trifluoroacetyl-pyrazolecarboxamidine, N,N'-bisacetyl-pyrazolecarboxamidine, N-methyl-pyrazolecarboxamidine, N,N'-bisacetyl-N-methyl-pyrazolecarboxamidine, N,N'-bisacetyl-N-methyl-4-nitro-pyrazolecarboxamidine, and N,N'-bisacetyl-N-methyl-4-trifluoromethyl-pyrazolecarboxamidine), or a salt or conjugate of any of these.

[0178] In some embodiments, the chemical reagent additionally comprises Mukaiyama's reagent (2-chloro-1-methylpyridinium iodide). In some embodiments, the reagent comprises at least one compound of Formula (I) and Mukaiyama's reagent.

[0179] In some embodiments, the chemical reagent comprising a cyanamide derivative is used to label one or more amino acids of the polypeptide. (See, e.g., Kwon et al., Org. Lett. 2014, 16, 6048-6051).

[0180] In some embodiments, the chemical reagent comprises a compound selected from the group consisting of a compound of Formula (II): or a salt or conjugate thereof, wherein R 4< is H, C 1-6 alkyl, cycloalkyl, -C(O)R g< , or -C(O)OR g< ; and R g< is H, C 1-6 alkyl, C 2-6 alkenyl, C 1-6 haloalkyl, or arylalkyl, wherein the C 1-6 alkyl, C 2-6 alkenyl, C 1-6 haloalkyl, and arylalkyl are each unsubstituted or substituted.

[0181] In some embodiments, a reagent comprising an isothiocyanate derivative is used to label the terminal amino acid (e.g., NTAA) of a polypeptide. (See, e.g., Martin et al., Organometallics. 2006, 34, 1787-1801).

[0182] In some embodiments, the chemical reagent comprises a compound selected from the group consisting of a compound of Formula (III):         R 5< -N=C=S     (III) or a salt or conjugate thereof, wherein R 5< is C 1-6 alkyl, C 2-6 alkenyl, cycloalkyl, heterocyclyl, aryl or heteroaryl; wherein the C 1-6 alkyl, C 2-6 alkenyl, cycloalkyl, heterocyclyl, aryl or heteroaryl are each unsubstituted or substituted with one or more groups selected from the group consisting of halo, -NR h< R i< , -S(O) 2 R j< , or heterocyclyl; R h< , R i< , and R j< are each independently H, C 1-6 alkyl, C 1-6 haloalkyl, arylalkyl, aryl, or heteroaryl, wherein the C 1-6 alkyl, C 1-6 haloalkyl, arylalkyl, aryl, and heteroaryl are each unsubstituted or substituted.

[0183] In some embodiments of Formula (III), R 5< is substituted phenyl. In some embodiments, R 5< is substituted phenyl substituted with one or more groups selected from halo, -NR h< R i< , -S(O) 2 R j< , or heterocyclyl. In some embodiments, R 5< is unsubstituted C 1-6 alkyl. In some embodiments, R 5< is substituted C 1-6 alkyl. In some embodiments, R 5< is substituted C 1-6 alkyl, substituted with one or more groups selected from halo, -NR h< R i< , -S(O) 2 R j< , or heterocyclyl. In some embodiments, R 5< is unsubstituted C 2-6 alkenyl. In some embodiments, R 5< is C 2-6 alkenyl. In some embodiments, R 5< is substituted C 2-6 alkenyl, substituted with one or more groups selected from halo, -NR h< R i< , -S(O) 2 R j< , or heterocyclyl. In some embodiments, R 5< is unsubstituted aryl. In some embodiments, R 5< is substituted aryl. In some embodiments, R 5< is aryl, substituted with one or more groups selected from halo, -NR h< R i< , -S(O) 2 R j< , or heterocyclyl. In some embodiments, R 5< is unsubstituted cycloalkyl. In some embodiments, R 5< is substituted cycloalkyl. In some embodiments, R 5< is cycloalkyl, substituted with one or more groups selected from halo, -NR h< R i< , -S(O) 2 R j< , or heterocyclyl. In some embodiments, R 5< is unsubstituted heterocyclyl. In some embodiments, R 5< is substituted heterocyclyl. In some embodiments, R 5< is heterocyclyl, substituted with one or more groups selected from halo, -NR h< R i< , -S(O) 2 R j< , or heterocyclyl. In some embodiments, R 5< is unsubstituted heteroaryl. In some embodiments, R 5< is substituted heteroaryl. In some embodiments, R 5< is heteroaryl, substituted with one or more groups selected from halo, -NR h< R i< , -S(O) 2 R j< , or heterocyclyl.

[0184] In some embodiments, the compound of Formula (III) is trimethylsilyl isothiocyanate (TMSITC) or pentafluorophenyl isothiocyanate (PFPITC).

[0185] In some embodiments, the compound is not trifluoromethyl isothiocyanate, allyl isothiocyanate, dimethylaminoazobenzene isothiocyanate, 4-sulfophenyl isothiocyanate, 3-pyridyl isothiocyanate, 2-piperidinoethyl isothiocyanate, 3-(4-morpholino) propyl isothiocyanate, or 3-(diethylamino)propyl isothiocyanate.

[0186] In some embodiments, the reagent is or comprises an alkyl amine. In some embodiments, the reagent additionally comprises DIPEA, trimethylamine, pyridine, and / or N-methylpiperidine. In some embodiments, the reagent additionally comprises pyridine and triethylamine in acetonitrile. In some embodiments, the reagent additionally comprises N-methylpiperidine in water and / or methanol.

[0187] In some embodiments, the polypeptide is also contacted with a carbodiimide compound.

[0188] In some embodiments, the chemical reagent comprises a carbodiimide derivative (See, e.g., Chi et al., 2015, Chem. Eur. J. 2015, 21, 10369-10378).

[0189] In some embodiments, the NTAA of a polypeptide is labeled via acylation. (See, e.g., Protein Science (1992), I, 582-589).

[0190] In some embodiments, the chemical reagent comprises a compound selected from the group consisting of a compound of Formula (IV): or a salt or conjugate thereof, wherein R 8< is halo or -OR m< ; R m< is H, C 1-6 alkyl, or heterocyclyl; and R 9< is hydrogen, halo, or C 1-6 haloalkyl.

[0191] In some embodiments of Formula (IV), R 8< is halo. In some embodiments, R 8< is chloro. In some embodiments, R 8< In some embodiments, R 9< is hydrogen. In some embodiments, R 9< is halo, such as bromo. In some embodiments, the compound of Formula (IV) is selected from acetyl chloride, acetyl anhydride, and acetyl-NHS. In some embodiments, the compound is not acetyl anhydride or acetyl-NHS.

[0192] In some embodiments, the polypeptide is also contacting with a peptide coupling reagent. In some embodiments, the peptide coupling reagent is a carbodiimide compound. In some embodiments, the carbodiimide compound is diisopropylcarbodiimide (DIC) or 1-ethyl-3-(3-dimethylaminopropyl)carbodiimide (EDC). In some embodiments, the method includes contacting with at least one compound of Formula (I) and a carbodiimide compounds, such as DIC or EDC.

[0193] In some embodiments, the chemical reagent comprises a conjugate of Formula (I), Formula (II), Formula (III), or Formula (IV). In some embodiments, the reagent used to modify the terminal amino acid of a polypeptide comprises a compound of Formula (I), Formula (II), Formula (III), or Formula (IV) conjugated to a ligand.

[0194] In some embodiments, the chemical reagent comprises a conjugate of Formula (I)-Q, Formula (II)-Q, Formula (III)-Q, or Formula (IV)-Q, wherein Formula (I)-(IV) are as defined above, and Q is a ligand.

[0195] In some embodiments, the ligand Q is a pendant group or binding site (e.g., the site to which the binding agent binds). In some embodiments, the polypeptide binds covalently to a binding agent. In some embodiments, the polypeptide comprises a terminal amino acid which includes a ligand group that is capable of covalent binding to a binding agent. In certain embodiments, the polypeptide comprises a labeled NTAA with a compound of Formula (I)-Q, Formula (II)-Q, Formula (III)-Q, or Formula (IV)-Q, , wherein the Q binds covalently to a binding agent. In some embodiments, a coupling reaction is carried out to create a covalent linkage between the polypeptide and the binding agent (e.g., a covalent linkage between the ligand Q and a functional group on the binding agent).

[0196] In some embodiments, the chemical reagent comprises a conjugate of Formula (I)-Q wherein R 1< , R 2< , and R 3< are as defined above and Q is a ligand.

[0197] In some embodiments, the chemical reagent comprises a conjugate of Formula (II)-Q wherein R 4< is as defined above, and Q is a ligand.

[0198] In some embodiments, the chemical reagent comprises a conjugate of Formula (III)-Q wherein R 5< is as defined above and Q is a ligand.

[0199] In some embodiments, the chemical reagent comprises a conjugate of Formula (IV)-Q wherein R 8< and R 9< are as defined above and Q is a ligand.

[0200] In some embodiments, Q is selected from the group consisting of -C 1-6 alkyl, -C 2-6 alkenyl, -C 2-6 alkynyl, aryl, heteroaryl, heterocyclyl, -N=C=S, -CN, -C(O)R n< , -C(O)OR o< , --SR p< or -S(O) 2 R q< ; wherein the -C 1-6 alkyl, -C 2-6 alkenyl, -C 2-6 alkynyl, aryl, heteroaryl, and heterocyclyl are each unsubstituted or substituted, and R n< , R o< , R p< , and R q< are each independently selected from the group consisting of -C 1-6 alkyl, -C 1-6 haloalkyl, -C 2-6 alkenyl, -C 2-6 alkynyl, aryl, heteroaryl, and heterocyclyl. In some embodiments, Q is selected from the group consisting of and

[0201] In some embodiments, Q is a fluorophore. In some embodiments, Q is selected from a lanthanide, europium, terbium, XL665, d2, quantum dots, green fluorescent protein, red fluorescent protein, yellow fluorescent protein, fluorescein, rhodamine, eosin, Texas red, cyanine, indocarbocyanine, ocacarbocyanine, thiacarbocyanine, merocyanine, pyridyloxadole, benzoxadiazole, cascade blue, nile red, oxazine 170, acridine orange, proflavin, auramine, malachite green crystal violet, porphine phtalocyanine, and bilirubin.

[0202] Provided in other aspects are reagents used in labeling the terminal amino acid or dipeptide for removal by the modified dipeptide cleavase with more than one label.

[0203] In some embodiments, labeling the terminal amino acid (e.g., NTAA) or amino acid for removal as a dipeptide by the modified dipeptide cleavase includes using a first reagent and a second reagent. In some embodiments, the terminal amino acid is concurrently or sequentially labeled with the first reagent and the second reagent. In some embodiments, the first reagent comprises a compound selected from the group consisting of a compound of Formula (I), (II), (III), (IV), and (IV), or a salt or conjugate thereof, as described herein.

[0204] In some embodiments, the second reagent comprises a compound of Formula (Va) or (Vb): or a salt or conjugate thereof, wherein R 13< is H, C 1-6 alkyl, aryl, heteroaryl, cycloalkyl, or heterocyclyl, wherein the C 1-6 alkyl, aryl, heteroaryl, cycloalkyl, and heterocyclyl are each unsubstituted or substituted; or         R 13< -X     (Vb) wherein R 13< is C 1-6 alkyl, aryl, heteroaryl, cycloalkyl, or heterocyclyl, each of which is unsubstituted or substituted; and X is a halogen.

[0205] In some embodiments of Formula (Va), R 13< is H. In some embodiments, R 13< is methyl. In some embodiments, R 13< is ethyl, propyl, isopropyl, butyl, isobutyl, secbutyl, pentyl, or hexyl. In some embodiments, R 13< is C 1-6 alkyl, which is substituted. In some embodiments, R 13< is C 1-6 alkyl, which is substituted with aryl, heteroaryl, cycloalkyl, or heterocyclyl. In some embodiments, R 13< is C 1-6 alkyl, which is substituted with aryl. In some embodiments, R 13< is - CH 2 CH 2 Ph, -CH 2 Ph, -CH(CH 3 )Ph, or -CH(CH 3 )Ph.

[0206] In some embodiments of Formula (Vb), R 13< is methyl. In some embodiments, R 13< is ethyl, propyl, isopropyl, butyl, isobutyl, secbutyl, pentyl, or hexyl. In some embodiments, R 13< is C 1-6 alkyl, which is substituted. In some embodiments, R 13< is C 1-6 alkyl, which is substituted with aryl, heteroaryl, cycloalkyl, or heterocyclyl. In some embodiments, R 13< is C 1-6 alkyl, which is substituted with aryl. In some embodiments, R 13< is -CH 2 CH 2 Ph, -CH 2 Ph, -CH(CH 3 )Ph, or - CH(CH 3 )Ph.

[0207] In some embodiments, the reagent for modifying or labeling the terminal amino acid for removal as part of a dipeptide by the modified dipeptide cleavase comprises formaldehyde. In some embodiments, the reagent for modifying or labeling the terminal amino acid comprises methyl iodide.

[0208] In some embodiments, the polypeptide is also contacted with a reducing agent. In some embodiments, the reducing agent comprises a borohydride, such as NaBH 4 , KBH 4 , ZnBH 4 , NaBH 3 CN or LiBu 3 BH. In some embodiments, the reducing agent comprises an aluminum or tin compound, such as LiAlH 4 or SnCl. In some embodiments, the reducing agent comprises a borane complex, such as B 2 H 6 and dimethyamine borane. In some embodiments, the reagent additionally comprises NaBH 3 CN.

[0209] In some embodiments, the reagents that may be used to label the terminal amino acid (e.g., NTAA) include: 4-sulfophenyl isothiocyanate (sulfo-PITC), 4-nitrophenyl isothiocyanate (nitro-PITC), 3-pyridyl isothiocyanate (PYITC), a phenyl isocyanate (PIC), a nitro-PIC, a sulfo-PIC, an anhydride (e.g., an isatoic anhydride, an isonicotinic anhydride, an azaisatoic anhydride, a succinic anhydride), 2-piperidinoethyl isothiocyanate (PEITC), 3-(4-morpholino) propyl isothiocyanate (MPITC), 3-(diethylamino)propyl isothiocyanate (DEPTIC) (Wang et al., 2009, Anal Chem 81: 1893-1900), (1-fluoro-2,4-dinitrobenzene (Sanger's reagent, DNFB), dansyl chloride (DNS-Cl, or 1-dimethylaminonaphthalene-5-sulfonyl chloride), 4-sulfonyl-2-nitrofluorobenzene (SNFB), acetylation reagents, amidination (guanidinylation) reagents (including PCA and PCA derivatives), 2-carboxy-4,6-dinitrochlorobenzene, 7-methoxycoumarin acetic acid, a thioacylation reagent, a thioacetylation reagent, and / or a thiobenzylation reagent. Many of these reagents are unreactive or minimally reactive with DNA including PITC, nitro-PITC, sulfo-PITC, PYITC, and guanidinylation reagents (e.g., PCA compounds). If the amino acid is blocked to labeling, there are a number of approaches to unblock the terminus, such as removing N-acetyl blocks with acyl peptide hydrolase (APH) (Farries, Harris et al., 1991, Eur. J. Biochem. 196:679-685). Methods of unblocking the N-terminus of a peptide are known in the art (see, e.g., Krishna et al., 1991, Anal. Biochem. 199:45-50; Leone et al., 2011, Curr. Protoc. Protein Sci., Chapter 11:Unit11.7; Fowler et al., 2001, Curr. Protoc. Protein Sci., Chapter 11: Unit 11.7).

[0210] Dansyl chloride reacts with the free amine group of a peptide to yield a dansyl derivative of the NTAA. DNFB and SNFB react the α-amine groups of a peptide to produce DNP-NTAA, and SNP-NTAA, respectively. Additionally, both DNFB and SNFB also react with the with ε-amine of lysine residues. DNFB also reacts with tyrosine and histidine amino acid residues. In some embodiments, SNFB has better selectivity for amine groups than DNFB (Carty et al., J Biol Chem (1968) 243(20): 5244-5253). In certain embodiments, lysine ε-amines are pre-blocked with an organic anhydride prior to polypeptide protease digestion into peptides.

[0211] Isothiocyanates, in the presence of ionic liquids, have been shown to have enhanced reactivity to primary amines. Ionic liquids are excellent solvents (and serve as a catalyst) in organic chemical reactions and can enhance the reaction of isothiocyanates with amines to form thioureas. Moreover, ionic liquids may act as absorbers of microwave radiation to further enhance reactivity (Martinez-Palou, J. Mex. Chem. Soc (2007) 51(4): 252-264). An example is the use of the ionic liquid 1-butyl-3-methyl-imidazolium tetraflouoraborate [Bmim][BF4] for rapid and efficient functionalization of aromatic and aliphatic amines by phenyl isothiocyanate (PITC) (Le, Chen et al. 2005).

[0212] In some embodiments, the peptide may be labeled by treating with a chemical reagent comprising a compound of Formula (AB) as shown in the scheme below:

[0213] In some embodiments, the peptide treated with a chemical reagent to modify the N-terminal amino acid (NTAA) of peptides is treated with a diheterocyclic methanimine reagent. In some embodiments, the reagent for modifying or labeling the terminal amino acid for removal as part of a dipeptide by the modified dipeptide cleavase comprises a compound of Formula (AB): wherein: R 2< is H, R 4< , OH, OR 4< , NH 2 , or -NHR 4< ; R 4< is C 1-6 alkyl, which is optionally substituted with one or two members selected from halo, C 1-3 alkyl, C 1-3 alkoxy, C 1-3 haloalkyl, phenyl, 5-membered heteroaryl, and 6-membered heteroaryl, wherein each phenyl, 5-membered heteroaryl, and 6-membered heteroaryl is optionally substituted with one or two members selected from halo, -OH, C 1-3 alkyl, C 1-3 alkoxy, C 1-3 haloalkyl, NO 2 , CN, COOR", and CON(R") 2 , where each R" is independently H or C 1-3 alkyl; ring A and ring B are each independently a 5-membered heteroaryl ring containing up to three N atoms as ring members and each is optionally fused to an additional phenyl or a 5-6 membered heteroaryl ring, and wherein the 5-membered heteroaryl ring and optional fused phenyl or 5-6 membered heteroaryl ring are each optionally substituted with one or two groups selected from C 1-4 alkyl, C 1-4 alkoxy, -OH, halo, C 1-4 haloalkyl, NO 2 , COOR, CONR 2 , -SO 2 R*, - NR 2 , phenyl, and 5-6 membered heteroaryl; wherein each R is independently selected from H and C 1-3 alkyl optionally substituted with OH, OR*, -NH 2 , -NHR*, or -NR* 2 ; and each R* is C 1-3 alkyl, optionally substituted with OH, oxo, C 1-2 alkoxy, or CN; wherein two R, or two R", or two R* on the same N can optionally be taken together to form a 4-7 membered heterocyclic ring, optionally containing an additional heteroatom selected from N, O and S as a ring member, and optionally substituted with one or two groups selected from halo, C 1-2 alkyl, OH, oxo, C 1-2 alkoxy, or CN. or a salt thereof. In some embodiments, Ring A and Ring B are not both unsubstituted imidazole, and that Ring A and Ring B are not both unsubstituted benzotriazole;

[0214] In an example of this embodiment, R 2< is H or R 4< . In these embodiments, In these embodiments, the 5-membered heteroaryl group, when present, can be a 5-membered ring comprising one to three heteroatoms selected from N, O and S as ring members, and the 6-membered heteroaryl group when present can be a 6-membered ring comprising one to three nitrogen atoms as ring members. In some of these embodiments, neither ring A nor ring B is unsubstituted imidazole or unsubstituted benzotriazole. In some embodiments, R 2< is H. In some of these embodiments, neither ring A nor ring B is unsubstituted imidazole or unsubstituted benzotriazole.

[0215] In some embodiments, Ring A and Ring B are different. In some embodiments, Ring A and Ring B are the same. Specific compounds of this embodiment include:

[0216] In some aspects, each 5-6 membered heteroaryl ring is independently selected and contains 1 or 2 heteroatoms selected from N, O and S as ring members. In these embodiments, each 5-membered heteroaryl group present can be a 5-membered ring comprising one or two heteroatoms selected from N, O and S as ring members, and the 6-membered heteroaryl group can be a 6-membered ring comprising one to two nitrogen atoms as ring members.

[0217] In some specific embodiments, Ring A and Ring B are selected from: wherein: each R x< , R y< and R z< is independently selected from H, halo, C 1-2 alkyl, C 1-2 haloalkyl, NO 2 , SO 2 (C 1-2 alkyl), COOR #< , C(O)N(R #< ) 2 , and phenyl optionally substituted with one or two groups selected from halo, C 1-2 alkyl, C 1-2 haloalkyl, NO 2 , SO 2 (C 1-2 alkyl), COOR #< , and C(O)N(R #< ) 2 , and two R x< , R y< or R z< on adjacent atoms of a ring can optionally be taken together to form a phenyl group, 5-membered heteroaryl group, or 6-membered heteroaryl group fused to the ring, and the fused phenyl, 5-membered heteroaryl, or 6-membered heteroaryl group can optionally be substituted with one or two groups selected from halo, C 1-2 alkyl, C 1-2 haloalkyl, NO 2 , SO 2 (C 1-2 alkyl), COOR #< , and C(O)N(R #< ) 2 ; wherein each R #< is independently H or C 1-2 alkyl; and wherein two R# on the same nitrogen can optionally be taken together to form a 4-7 membered heterocycle optionally containing an additional heteroatom selected from N, O and S as a ring member, wherein the 4-7 membered heterocycle is optionally substituted with one or two groups selected from halo, OH, OMe, Me, oxo, NH 2 , NHMe and NMe 2 ; or a salt thereof.

[0218] In these embodiments, each 5-membered heteroaryl group present can be a 5-membered ring comprising one to three heteroatoms selected from N, O and S as ring members, and the 6-membered heteroaryl group can be a 6-membered ring comprising one to three nitrogen atoms as ring members.

[0219] In some embodiments, Ring A and Ring B are the same and are selected from:

[0220] The compound of embodiment 30, which is selected from the following: B. Engineering and Genetic Selection

[0221] In some embodiments, the modified dipeptide cleavase provided herein can be made, isolated, engineered, or selected for using any suitable methods. In some cases, the variant or modified dipeptide cleavase polypeptide is altered in primary amino acid sequence compared to the wild-type or unmodified dipeptide cleavase by introducing one or more substitutions, additions, or deletions of amino acid residues. In some embodiments, the modified dipeptide cleavase is derived from a wild-type or unmodified dipeptide cleavase (e.g., a dipeptidyl peptidase, a dipeptidyl aminopeptidase, a peptidyl-dipeptidase, a dipeptidyl carboxypeptidase, or a protein classified in EC 3.4.14, EC 3.4.15, MEROPS S9, MEROPS S46, MEROPS M49, or a functional homolog or fragment thereof) via engineering and genetic selection. A variety of techniques including genetic selection, protein engineering, recombinant methods, chemical synthesis, or combinations thereof, may be employed.

[0222] In some embodiments, the modified dipeptide cleavase is engineered using a rational design approach for select activities, substrate binding capability, or other cleaving characteristics. In some embodiments, a rational design approach is based on crystal structure of the unmodified dipeptide cleavase. In some examples, the rational design approach is based on crystal structure of the unmodified dipeptide cleavase with substrates to identify target amino acid residues for modification. In some cases, the modifications may be targeted at residues of specific domains of the unmodified dipeptide cleavase (See e.g., Sakamoto et al., Scientific Reports 2014, 4:4977). In some embodiments, a rational design is used to engineer a modified dipeptide cleavase with modified amino acids in the substrate binding domain of the unmodified dipeptide cleavase. In some embodiments, the mutations, e.g., one or more amino acid modifications (e.g., substitutions, additions, deletions) corresponds to positions 316, 391, 394, or a combination thereof, with reference to positions of SEQ ID NO: 5 or the sequence of a human dipeptidyl peptidase 3 (DPP3) or a homolog thereof. In some embodiments, a rational design is used to engineer a modified dipeptide cleavase that is able to bind or cleave polypeptides of increased length compared to an unmodified dipeptide cleavase. In some embodiments, the mutations, e.g., one or more amino acid modifications (e.g., substitutions, additions, deletions) corresponds to amino acid residue positions 419, 420, 421, 422, 423, 424, 425, 426, or a combination thereof, with reference to positions of SEQ ID NO: 5. In some embodiments, a rational design is used to engineer a modified dipeptide cleavase with modified amino acids in the hinge region of the unmodified dipeptide cleavase.

[0223] In some embodiments, the genetic selection or other engineering methods are designed to identify modified dipeptide cleavases that are active on labeled polypeptides (e.g. chemically labeled polypeptides). In some embodiments, the genetic selection or other engineering methods are designed to identify modified dipeptide cleavases that are active on modified or labeled polypeptides having a labeled N-terminal amino acid. In some cases, the size or other characteristics of the moiety or label on the labeled polypeptide is considered in the design of the genetic selection or other engineering methods to obtain a desired modified dipeptide cleavase.

[0224] It is understood that references to amino acids, including to specific sequences set forth in the Sequence Listing as SEQ ID NOs used herein to describe domain organization of a wild-type or modified dipeptide cleavase are for illustrative purposes only and are not meant to limit the scope of the embodiments provided. It is understood that polypeptides and the description of domains thereof are theoretically derived based on homology analysis and alignments with similar molecules. Thus, the exact locus can vary, and is not necessarily the same for each protein. Hence, the specific domain, such as specific binding domain, loop domain, or other functional domain), can be identified in a homolog or enzyme derived from another species using known analyses and alignment methods.

[0225] In some examples, amino acids for modification in a wildtype dipeptide cleavase can be chosen using analysis of crystal structure of the wild-type cleavase (e.g. wildtype DAP BII) and its substrate to identify contact residues and other residues at the protein interaction interface. This analysis can be performed for example, using Rosetta software suite for macromolecular modeling (Das et al., Annu Rev Biochem (2008) 77:363-382). In some embodiments, using the selected target residues for modification, an alignment of wildtype cleavase sequences of other organisms can be used to identify conserved residues (Crooks et al., Genome Res (2004) 14(6): 1188-1190). Based on this analysis, conserved target residues or corresponding residues in homologs can be modified. In some embodiments, the identified contact residues or other residues of interest are modified to introduce new functions. As shown in FIG. 4A-4C, a WebLogo analysis of DAP BII homologs with 60% sequence similarity or identity showed sequence conservation across various residues. For example, sequence conservation was observed for residues at the amine binding sites of DAP BII including positions N215, W216, R220, N330, and D674 in reference to the wildtype DAP BII sequence set forth in SEQ ID NO: 20. In another example, sequence conservation was observed for residues at the amine binding sites of DAP BII including positions G207, K208, F209, G210, G211, D212, I213, D214, N215, W216, M217, W218, P219, R220, H221, T222, G223, A224, F225, A226, A326, and N334, in reference to the wildtype DAP BII sequence set forth in SEQ ID NO: 20.

[0226] In some aspects, a rational design approach for engineering DAP BII may be used to target domains or residues such that the resulting modified dipeptide cleavase removes or is configured to remove a labeled N-terminal amino acid (NTAA) using crystal structures of DAP BII in complex with substrates (Sakamoto et al., Scientific Reports 2014, 4:4977). For example, the DAP BII structure in complex with a peptide substrate at the residues N191, W192, R196, N306, and D650 (based on the sequence of the protein set forth in SEQ ID NO: 13; UniProt Accession No. V5YM14) interacts with the peptide N-terminal amine group. Additionally, a loop of approximately 20 residues (residue 183-202 in reference to SEQ ID NO: 13) makes contact with the N-terminal residue and penultimate residue of a bound peptide substrate. These amine binding residues and NTAA and penultimate NTAA binding residues, individually or in combination, may be targeted for modification.

[0227] In some aspects, it may be desired to modify the specificity of the unmodified or wildtype cleavase (See e.g., Sakamoto et al., Scientific Reports 2014, 4:4977). In some examples, residues in the S1 subsite or pocket of DAP BII can be targeted to engineer a modified cleavase with preferred specificity (e.g., reduced specificity for a specific amino acid residue at the P1 position of the polypeptide treated with the modified cleavase). In some embodiments, the modified dipeptide cleavase comprises mutations, e.g., one or more amino acid modifications (e.g., substitutions, additions, deletions) corresponding to positions D627, I628, G630, A648, G651, S655, M669, or a combination thereof, with reference to positions of SEQ ID NO: 13.

[0228] In some embodiments, a modified dipeptide cleavase variant can be identified using a genetic screen. In some cases, the genetic screen uses a cell-based system. In some embodiments, the genetic screen uses prokaryotic cells, such as E. coli strains including E. coli variants or mutants. In some embodiments, the genetic screen uses eukaryotic cells, such as yeast two-hydrid systems. In some embodiments, the genetic selection is designed to select for modified dipeptide cleavases with desired characteristics for binding of substrates, cleaving, and / or removal of labeled terminal amino acids.

[0229] In some embodiments, carrying out a genetic selection screen involves preparing various dipeptide cleavase genes (e.g., a dipeptidyl peptidase, a dipeptidyl aminopeptidase, a peptidyl-dipeptidase, a dipeptidyl carboxypeptidase) for expression. A plasmid or cosmid containing nucleic acid sequences encoding mutated or modified dipeptide cleavase polypeptides is readily constructed using standard techniques well known in the art. In some embodiments, the expression of any of the dipeptide cleavases (e.g., any of SEQ ID NOs: 5-8, 10-16, 20, 31, 32) may further include a signal sequence. In some cases, the use of a signal sequence may be useful for purification purposes. For example, a periplasm targeting sequence such as PelB can be included in the expression construct. Recombinant vectors can be generated using any of the recombinant techniques known in the art.

[0230] In some embodiments, the vectors can include a prokaryotic origin of replication and / or a gene whose expression confers a detectable or selectable marker for propagation and / or selection in prokaryotic systems. Once the vector or DNA sequence containing the constructs has been prepared for expression, the DNA constructs may be introduced into an appropriate host. In some embodiments, prokaryotic hosts can be used including bacteria such as E. coli., Bacillus, Streptomyces, Pseudomonas, Salmonella, Serratia, etc. Various techniques may be employed, such as protoplast fusion, calcium phosphate precipitation, electroporation or other conventional techniques. After the fusion, the cells are grown in media and screened for appropriate activities.

[0231] In some examples, libraries of mutated dipeptide cleavase genes can be generated by error prone PCR or rational mutagenesis using the crystal structure of the cleavase as a guide, or a combination thereof. Other suitable methods for generating mutations or generating a library may also be used. A library of mutated dipeptide cleavase genes can be subsequently cloned into a vector and transformed into an E. coli auxotroph strain (available from CSSC E. coli Genetic Stock Center at Yale - https: / / cgsc2.biology.yale.edu / ). In some embodiments, the screen involves isolating colonies growing on the selection media and extracting and analyzing plasmid DNA to identify modified dipeptide cleavase polypeptides that remove a labeled terminal dipeptide from a polypeptide. In some embodiments, a screen can be performed to identify and isolate a modified dipeptide cleavase that cleaves or is configured to cleave a polypeptide with a labeled amino acid (e.g., a PITC-labeled NTAA or a Cbz-labeled NTAA, etc). In some embodiments, the genetic screen is aimed at selecting for the binding of the label but not for a specific amino acid, therefore, the screen uses polypeptides with various labeled terminal amino acids. In some embodiments, selecting a modified dipeptide cleavase further includes purifying, characterizing, assessing and / or optimizing of the activity of the modified dipeptide cleavase. The modified dipeptide cleavase may be isolated and purified in accordance with conventional methods, such as extraction, precipitation, chromatography, affinity chromatography, electrophoresis, or the like.

[0232] In some embodiments, a genetics screen or other selection methods can also be used to select for and obtain modified dipeptide cleavases with an altered active site of the cleavase or with altered binding pockets of the cleavase. In some embodiments, genetics screen or selection methods can also be used to select for and obtain modified dipeptide cleavases with an altered hinge region of the cleavase. In some embodiments, genetic screen or selection methods can also be used to select for and obtain modified dipeptide cleavases with an altered binding cleft of the cleavase. In some cases, genetic screen or selection methods can also be used to select for and obtain modified dipeptide cleavases with an altered inter-lobe cleft of the cleavase. In some embodiments, genetic screen or selection methods can also be used to select for and obtain modified dipeptide cleavases with an altered alpha amine binding region of the cleavase. For example, the modified dipeptide cleavase exhibits reduced alpha amine binding compared to the wild-type cleavase polypeptide.

[0233] In some embodiments, a genetic screen or other selection methods can also be used to select for and obtain modified dipeptide cleavases configured to remove a dipeptide comprising a labeled terminal amino acid from polypeptides of various lengths. In some cases, porin size in the E. coli outer membrane limits the peptide length that can be uptaken. In some embodiments, this length limitation is overcome by briefly treating E. coli with Tris-EDTA or the small molecule MAC13243 which permeabilizes the E. coli outer membrane (e.g., Leive, L. (1974). Ann N Y Acad Sci 235(0): 109-129; Muheim, C. (2017). Scientific Reports 7(1): 17629) and allows uptake of peptides into the periplasmic space. In some embodiments, the modified dipeptide cleavase is capable of cleaving or is configured to remove amino acids from polypeptides that are greater than 5 amino acids in length, greater than 6 amino acids in length, greater than 7 amino acids in length, greater than 8 amino acids in length, greater than 9 amino acids in length, greater than 10 amino acids in length, greater than 15 amino acids in length, greater than 20 amino acids in length, greater than 25 amino acids in length, or greater than 30 amino acids in length. In some embodiments, the modified dipeptide cleavase is capable of cleaving or is configured to remove amino acids from polypeptides that are less than 30 amino acids in length, less than 40 amino acids in length, less than 50 amino acids in length, less than 75 amino acids in length, less than 100 amino acids in length, less than 200 amino acids in length, less than 300 amino acids in length, less than 400 amino acids in length, less than 500 amino acids in length, less than 600 amino acids in length, less than 700 amino acids in length, less than 800 amino acids in length, less than 900 amino acids in length, or less than 1000 amino acids in length. In some embodiments, the modified dipeptide cleavase is capable of cleaving or is configured to remove amino acids from polypeptides that are between 5 to 100 amino acids in length, between 10 to 100 amino acids in length, between 20 to 100 amino acids in length, between 30 to 100 amino acids in length, between 5 to 50 amino acids in length, between 10 to 50 amino acids in length, between 20 to 50 amino acids in length, between 30 to 50 amino acids in length, between 5 to 30 amino acids in length, between 10 to 30 amino acids in length, between 20 to 30 amino acids in length, between 10 to 20 amino acids in length. In some embodiments, the modified dipeptide cleavase is capable of cleaving or is configured to remove amino acids from polypeptides that are between 50 to 1000 amino acids in length, between 100 to 1000 amino acids in length, between 300 to 1000 amino acids in length, between 500 to 1000 amino acids in length, between 10 to 500 amino acids in length, between 50 to 500 amino acids in length, between 100 to 500 amino acids in length, or between 200 to 500 amino acids in length.

[0234] In some embodiments, the modified dipeptide cleavase is capable or configured to remove dipeptides from partial or digested proteins and polypeptides (e.g., protein or polypeptide fragments). In some embodiments, the modified dipeptide cleavase is capable or configured to remove dipeptides from whole or undigested proteins and polypeptides.

[0235] In some embodiments, the modified dipeptide cleavase removes the terminal dipeptide by contacting the polypeptide with a modified dipeptide cleavase for less than 5 minutes, less than 10 minutes, less than 20 minutes, less than 30 minutes, less than 40 minutes, less than 50 minutes, less than 60 minutes, less than 2 hours, less than 5 hours, less than 8 hours, or less than 10 hours.

[0236] In some embodiments, the modified dipeptide cleavase achieves a yield of polypeptides with the terminal dipeptide removed of >30%, >40%, >50%, >60%, >70%, >80%, >90%, >95%, >99% or more by treating the polypeptide with the modified dipeptide cleavase for about less than 15 minutes. In some embodiments, the modified dipeptide cleavase achieves a yield of polypeptides with the terminal dipeptide removed of >30%, >40%, >50%, >60%, >70%, >80%, >90%, >95%, >99% or more by treating the polypeptide with the modified dipeptide cleavase for about less than 30 minutes. In some embodiments, the modified dipeptide cleavase achieves a yield of polypeptides with the terminal dipeptide removed of >30%, >40%, >50%, >60%, >70%, >80%, >90%, >95%, >99% or more by treating the polypeptide with the modified dipeptide cleavase for about less than 45 minutes. In some embodiments, the modified dipeptide cleavase achieves a yield of polypeptides with the terminal dipeptide removed of >30%, >40%, >50%, >60%, >70%, >80%, >90%, >95%, >99% or more by treating the polypeptide with the modified dipeptide cleavase for about less than 1 hour. In some embodiments, the modified dipeptide cleavase achieves a yield of polypeptides with the terminal dipeptide removed of >30%, >40%, >50%, >60%, >70%, >80%, >90%, >95%, >99% or more by treating the polypeptide with the modified dipeptide cleavase for about less than 2 hours. In some embodiments, the modified dipeptide cleavase achieves a yield of polypeptides with the terminal dipeptide removed of >30%, >40%, >50%, >60%, >70%, >80%, >90%, >95%, >99% or more by treating the polypeptide with the modified dipeptide cleavase for about less than 5 hours.

[0237] In some embodiments, the modified dipeptide cleavase is capable of cleaving dipeptides or functions at a temperature of higher than about 10° C, higher than about 20° C higher than about 30° C, or higher than about 40° C. In some embodiments, the modified dipeptide cleavase is capable of cleaving terminal dipeptides or functions at a temperature of about 10° C to 20° C, about 10° C to 30° C, about 10° C to 40° C, about 10° C to 50° C, about 10° C to 60° C, about 10° C to 70° C, about 10° C to 80° C, about 10° C to 90° C or about 10° C to 100° C; about 20° C to 30° C, about 20° C to 40° C, about 20° C to 50° C, about 20° C to 60° C, about 20° C to 70° C, about 20° C to 80° C, about 20° C to 90° C, or about 20° C to 100° C; about 30° C to 40° C, about 30° C to 50° C, about 30° C to 60° C; about 50° C to 70° C, about 50° C to 80° C, about 50° C to 90° C, or about 50° C to 100° C. In some embodiments, the modified dipeptide cleavase is capable of cleaving terminal dipeptides at a temperature at which the secondary structure of the polypeptide is disrupted. In some embodiments, the modified dipeptide cleavase functions at about 20 to 25° C. In some embodiments, the method includes contacting the modified dipeptide cleavase with the polypeptide while applying heating. In some embodiments, the heating is achieved by applying microwave energy. In some embodiments of any of the methods provided herein, the contacting of the modified dipeptide cleavase with the polypeptide to remove a terminal dipeptide is performed in the presence of microwave energy.

[0238] Provided herein are isolated DNA molecules encoding any of the modified dipeptide cleavases as described in Section I. Also provided are recombinant expression vectors comprising a DNA molecule encoding any of the modified dipeptide cleavases as described in Section I. In some cases, the DNA molecules and recombinant expression vectors are isolated from the genetic engineering and selection methods described. In some cases, a host cell comprising the DNA molecule is also provided. In some embodiments, a fusion protein containing a fragment of a modified dipeptide cleavase is provided.

[0239] In some embodiments, provided herein is a method of producing a modified or variant dipeptide cleavase, comprising introducing the nucleic acid molecule according to any one of the embodiments described herein or vector according to any one of the embodiments described herein into a host cell under conditions to express the protein in the cell. Also provided herein are methods for producing any of the modified dipeptide cleavases provided herein including: cultivating a transformed host cell under conditions suitable for expression of the modified dipeptide cleavase, and separating, purifying and / or recovering the mutant organism expressing the modified dipeptide cleavase. In some embodiments, provided herein is a host cell comprising a DNA molecule encoding a modified dipeptide cleavase. In some embodiments, the host cell comprises a recombinant expression vector for expressing a modified dipeptide cleavase. In some embodiments, the method further includes isolating or purifying the variant or modified dipeptide cleavase from the cell.

[0240] In some embodiments, provided herein is an engineered cell, expressing the variant or modified dipeptide cleavase polypeptide according to any one of the embodiments described herein or the nucleic acid molecule encoding a variant or modified dipeptide cleavase described herein, or the vector according to any one of the embodiments described herein. In some embodiments, the variant or modified dipeptide cleavase polypeptide contains a signal peptide.II. POLYPEPTIDES

[0241] In some embodiments, the present disclosure relates to the treatment of polypeptides with any of the modified dipeptide cleavases provided herein. In some embodiments, the labeled terminal amino acid is removed as part of a dipeptide from a polypeptide (including a partial or fragmented polypeptide).

[0242] In some embodiments, the terminal amino acid is removed as a dipeptide from a polypeptide that has a length of greater than 4 amino acids, greater than 5 amino acids, greater than 6 amino acids, greater than 7 amino acids, greater than 8 amino acids, greater than 9 amino acids, greater than 10 amino acids, greater than 11 amino acids, greater than 12 amino acids, greater than 13 amino acids, greater than 14 amino acids, greater than 15 amino acids, greater than 20 amino acids, greater than 25 amino acids, or greater than 30 amino acids. In some cases, the length of the polypeptide is greater than 10 amino acids. In some embodiments, the terminal amino acid is removed as a dipeptide from a polypeptide that has a length of less than 30 amino acids, less than 40 amino acids, less than 50 amino acids, less than 75 amino acids, less than 100 amino acids, less than 200 amino acids, less than 300 amino acids, less than 400 amino acids, less than 500 amino acids, less than 600 amino acids, less than 700 amino acids, less than 800 amino acids, less than 900 amino acids, or less than 1000 amino acids. In some embodiments, the terminal amino acid is removed as a dipeptide from a polypeptide that has a length of between 5 to 100 amino acids, between 10 to 100 amino acids, between 20 to 100 amino acids, between 30 to 100 amino acids, between 5 to 50 amino acids, between 10 to 50 amino acids, between 20 to 50 amino acids, between 30 to 50 amino acids, between 5 to 30 amino acids, between 10 to 30 amino acids, between 20 to 30 amino acids, between 10 to 20 amino acids. In some embodiments, the terminal amino acid is removed as a dipeptide from a polypeptide that has a length of between 50 to 1000 amino acids, between 100 to 1000 amino acids, between 300 to 1000 amino acids, between 500 to 1000 amino acids, between 10 to 500 amino acids, between 50 to 500 amino acids, between 100 to 500 amino acids, or between 200 to 500 amino acids.

[0243] In some embodiments, the terminal amino acid is removed as a dipeptide from a partial or digested protein and polypeptide (e.g., a polypeptide fragment). In some embodiments, the terminal amino acid is removed as a dipeptide from a whole or undigested protein and polypeptide.

[0244] A polypeptide treated with the modified dipeptide cleavases provided herein and according the methods disclosed herein may be obtained from a suitable source or sample, including but not limited to: biological samples, such as cells (both primary cells and cultured cell lines), cell lysates or extracts, cell organelles or vesicles, including exosomes, tissues and tissue extracts; biopsy; fecal matter; bodily fluids (such as blood, whole blood, serum, plasma, urine, lymph, bile, cerebrospinal fluid, interstitial fluid, aqueous or vitreous humor, colostrum, sputum, amniotic fluid, saliva, anal and vaginal secretions, perspiration and semen, a transudate, an exudate (e.g., fluid obtained from an abscess or any other site of infection or inflammation) or fluid obtained from a joint (normal joint or a joint affected by disease such as rheumatoid arthritis, osteoarthritis, gout or septic arthritis) of virtually any organism, with mammalian-derived samples, including microbiome-containing samples, being preferred and human-derived samples, including microbiome-containing samples, being particularly preferred; environmental samples (such as air, agricultural, water and soil samples); microbial samples including samples derived from microbial biofilms and / or communities, as well as microbial spores; research samples including extracellular fluids, extracellular supernatants from cell cultures, inclusion bodies in bacteria, cellular compartments including mitochondrial compartments, and cellular periplasm.

[0245] In certain embodiments, the polypeptide is a protein or a protein complex. Amino acid sequence information and post-translational modifications of the polypeptide are transduced into a nucleic acid encoded library that can be analyzed via next generation sequencing methods. A polypeptide may comprise L-amino acids, D-amino acids, or both. A polypeptide may comprise a standard, naturally occurring amino acid, a modified amino acid (e.g., post-translational modification), an amino acid analog, an amino acid mimetic, or any combination thereof. In some embodiments, the polypeptide is naturally occurring, synthetically produced, or recombinantly expressed. In any of the aforementioned embodiments, the polypeptide may further comprise a post-translational modification.

[0246] Standard, naturally occurring amino acids include Alanine (A or Ala), Cysteine (C or Cys), Aspartic Acid (D or Asp), Glutamic Acid (E or Glu), Phenylalanine (F or Phe), Glycine (G or Gly), Histidine (H or His), Isoleucine (I or Ile), Lysine (K or Lys), Leucine (L or Leu), Methionine (M or Met), Asparagine (N or Asn), Proline (P or Pro), Glutamine (Q or Gln), Arginine (R or Arg), Serine (S or Ser), Threonine (T or Thr), Valine (V or Val), Tryptophan (W or Trp), and Tyrosine (Y or Tyr). Non-standard amino acids include selenocysteine, pyrrolysine, and N-formylmethionine, β-amino acids, Homo-amino acids, Proline and Pyruvic acid derivatives, 3-substituted Alanine derivatives, Glycine derivatives, Ring-substituted Phenylalanine and Tyrosine Derivatives, Linear core amino acids, and N-methyl amino acids.

[0247] A post-translational modification (PTM) of a polypeptide may be a covalent modification or enzymatic modification. Examples of post-translation modifications include, but are not limited to, acylation, acetylation, alkylation (including methylation), biotinylation, butyrylation, carbamylation, carbonylation, deamidation, deiminiation, diphthamide formation, disulfide bridge formation, eliminylation, flavin attachment, formylation, gamma-carboxylation, glutamylation, glycylation, glycosylation (e.g., N-linked, O-linked, C-linked, phosphoglycosylation), glypiation, heme C attachment, hydroxylation, hypusine formation, iodination, isoprenylation, lipidation, lipoylation, malonylation, methylation, myristolylation, oxidation, palmitoylation, pegylation, phosphopantetheinylation, phosphorylation, prenylation, propionylation, retinylidene Schiff base formation, S-glutathionylation, S-nitrosylation, S-sulfenylation, selenation, succinylation, sulfination, ubiquitination, and C-terminal amidation. A post-translational modification includes modifications of the amino terminus and / or the carboxyl terminus of a peptide, polypeptide, or protein. Modifications of the terminal amino group include, but are not limited to, des-amino, N-lower alkyl, N-di-lower alkyl, and N-acyl modifications. Modifications of the terminal carboxy group include, but are not limited to, amide, lower alkyl amide, dialkyl amide, and lower alkyl ester modifications (e.g., wherein lower alkyl is C 1 -C 4 alkyl). A post-translational modification also includes modifications, such as but not limited to those described above, of amino acids falling between the amino and carboxy termini of a peptide, polypeptide, or protein. Post-translational modification can regulate a protein's "biology" within a cell, e.g., its activity, structure, stability, or localization. Phosphorylation is the most common post-translational modification and plays an important role in regulation of protein, particularly in cell signaling (Prabakaran et al., (2012) Wiley Interdiscip Rev Syst Biol Med 4: 565-583). The addition of sugars to proteins, such as glycosylation, has been shown to promote protein folding, improve stability, and modify regulatory function. The attachment of lipids to proteins enables targeting to the cell membrane. A post-translational modification can also include modifications to include one or more detectable labels.

[0248] In certain embodiments, the polypeptide can be fragmented. For example, the fragmented polypeptide can be obtained by fragmenting a polypeptide, protein or protein complex from a sample, such as a biological sample. The polypeptide, protein or protein complex can be fragmented by any means known in the art, including fragmentation by a protease or endopeptidase. In some embodiments, fragmentation of a polypeptide, protein or protein complex is targeted by use of a specific protease or endopeptidase. A specific protease or endopeptidase binds and cleaves at a specific consensus sequence (e.g., TEV protease which is specific for ENLYFQ\S consensus sequence). In other embodiments, fragmentation of a peptide, polypeptide, or protein is non-targeted or random by use of a non-specific protease or endopeptidase. A non-specific protease may bind and cleave at a specific amino acid residue rather than a consensus sequence (e.g., proteinase K is a non-specific serine protease). Proteinases and endopeptidases are well known in the art, and examples of such that can be used to cleave a protein or polypeptide into smaller peptide fragments include proteinase K, trypsin, chymotrypsin, pepsin, thermolysin, thrombin, Factor Xa, furin, endopeptidase, papain, pepsin, subtilisin, elastase, enterokinase, Genenase ™< I, Endoproteinase LysC, Endoproteinase AspN, Endoproteinase GluC, etc. (Granvogl et al., (2007) Anal Bioanal Chem 389: 991-1002). In certain embodiments, a peptide, polypeptide, or protein is fragmented by proteinase K, or optionally, a thermolabile version of proteinase K to enable rapid inactivation. Proteinase K is quite stable in denaturing reagents, such as urea and SDS, enabling digestion of completely denatured proteins.

[0249] In some embodiments, the polypeptide is contacted with one or more enzymes in addition to a modified dipeptide cleavase to eliminate the NTAA (e.g., a proline aminopeptidase to remove an N-terminal proline, if present). In some embodiments, the additional enzyme eliminates an NTAA from the polypeptide that is a proline. In some specific examples, the enzyme is a proline aminopeptidase, a proline iminopeptidase (PIP), or a pyroglutamate aminopeptidase (pGAP). In some embodiments, one or more modified dipeptide cleavases are used in combination with other enzymes to treat the polypeptides. In some embodiments, the polypeptide is first contacted with a proline aminopeptidase under conditions suitable to remove an N-terminal proline, if present.

[0250] Chemical reagents can also be used to digest proteins into peptide fragments. A chemical reagent may cleave at a specific amino acid residue (e.g., cyanogen bromide hydrolyzes peptide bonds at the C-terminus of methionine residues). Chemical reagents for fragmenting polypeptides or proteins into smaller peptides include cyanogen bromide (CNBr), hydroxylamine, hydrazine, formic acid, BNPS-skatole [2-(2-nitrophenylsulfenyl)-3-methylindole], iodosobenzoic acid, •NTCB +Ni (2-nitro-5-thiocyanobenzoic acid), etc.

[0251] In certain embodiments, some polypeptides can be treated with a reagent for enzymatic or chemical elimination. In certain embodiments, following enzymatic or chemical elimination, the resulting polypeptide fragments are approximately the same desired length, e.g., from about 10 amino acids to about 70 amino acids, from about 10 amino acids to about 60 amino acids, from about 10 amino acids to about 50 amino acids, about 10 to about 40 amino acids, from about 10 to about 30 amino acids, from about 20 amino acids to about 70 amino acids, from about 20 amino acids to about 60 amino acids, from about 20 amino acids to about 50 amino acids, about 20 to about 40 amino acids, from about 20 to about 30 amino acids, from about 30 amino acids to about 70 amino acids, from about 30 amino acids to about 60 amino acids, from about 30 amino acids to about 50 amino acids, or from about 30 amino acids to about 40 amino acids. A elimination reaction may be monitored, preferably in real time, by spiking the protein or polypeptide sample with a short test FRET (fluorescence resonance energy transfer) polypeptide comprising a peptide sequence containing a proteinase or endopeptidase elimination site. In the intact FRET peptide, a fluorescent group and a quencher group are attached to either end of the peptide sequence containing the elimination site, and fluorescence resonance energy transfer between the quencher and the fluorophore leads to low fluorescence. Upon elimination of the test peptide by a protease or endopeptidase, the quencher and fluorophore are separated giving a large increase in fluorescence. An elimination reaction can be stopped when a certain fluorescence intensity is achieved, allowing a reproducible elimination end point to be achieved.A. Providing the Polypeptide Joined to a Support or in Solution

[0252] In some embodiments, polypeptides of the present disclosure are joined to a surface of a solid support (also referred to as "substrate surface"). In some cases, the polypeptides are joined to a solid support prior to contacting with the modified dipeptide cleavase. In some cases, the modified dipeptide cleavase removes a labeled terminal amino acid from a polypeptide that is join (directly or indirectly) to a solid support. In some embodiments, the labeled terminal amino acid is removed as a single amino acid or as part of a dipeptide.

[0253] The solid support can be any porous or non-porous support surface including, but not limited to, a bead, a microbead, an array, a glass surface, a silicon surface, a plastic surface, a filter, a membrane, a PTFE membrane, a PTFE membrane, a nitrocellulose membrane, a nitrocellulose-based polymer surface, nylon, a silicon wafer chip, a flow cell, a flow through chip, a biochip including signal transducing electronics, a microtiter well, an ELISA plate, a spinning interferometry disc, a nitrocellulose membrane, a nitrocellulose-based polymer surface, a nanoparticle, or a microsphere. Materials for a solid support include but are not limited to acrylamide, agarose, cellulose, dextran, nitrocellulose, glass, gold, quartz, polystyrene, polyethylene vinyl acetate, polypropylene, polyester, polymethacrylate, polyacrylate, polyethylene, polyethylene oxide, polysilicates, polycarbonates, poly vinyl alcohol (PVA), Teflon, fluorocarbons, nylon, silicon rubber, polyanhydrides, polyglycolic acid, polyvinylchloride, polylactic acid, polyorthoesters, functionalized silane, polypropylfumerate, collagen, glycosaminoglycans, polyamino acids, or any combination thereof. Solid supports further include thin film, membrane, bottles, dishes, fibers, woven fibers, shaped polymers such as tubes, particles, beads, microparticles, or any combination thereof. For example, when solid surface is a bead, the bead can include, but is not limited to, a polystyrene bead, a polymer bead, a polyacrylate bead, a methylstyrene bead, an agarose bead, a cellulose bead, a dextran bead, an acrylamide bead, a solid core bead, a porous bead, a paramagnetic bead, glass bead, a controlled pore bead, a silica-based bead, or any combinations thereof.

[0254] In certain embodiments, a solid support is a bead, which may refer to an individual bead or a plurality of beads. In some embodiments, the bead is compatible with a selected next generation sequencing platform that will be used for downstream analysis (e.g., SOLiD or 454). In some embodiments, a solid support is an agarose bead, a paramagnetic bead, a polystyrene bead, a polymer bead, an acrylamide bead, a solid core bead, a porous bead, a glass bead, or a controlled pore bead. In further embodiments, a bead may be coated with a binding functionality (e.g., amine group, affinity ligand such as streptavidin for binding to biotin labeled polypeptide, antibody) to facilitate binding to a polypeptide.

[0255] Proteins, polypeptides, or peptides can be joined to the solid support, directly or indirectly, by any means known in the art, including covalent and non-covalent interactions, or any combination thereof (see, e.g., Chan et al., 2007, PLoS One 2:e1164; Cazalis et al., Bioconj. Chem. 15:1005-1009; Soellner et al., 2003, J. Am. Chem. Soc. 125:11790-11791; Sun et al., 2006, Bioconjug. Chem. 17-52-57; Decreau et al., 2007, J. Org. Chem. 72:2794-2802; Camarero et al., 2004, J. Am. Chem. Soc. 126:14730-14731; Girish et al., 2005, Bioorg. Med. Chem. Lett. 15:2447-2451; Kalia et al., 2007, Bioconjug. Chem. 18:1064-1069; Watzke et al., 2006, Angew Chem. Int. Ed. Engl. 45:1408-1412; Parthasarathy et al., 2007, Bioconjugate Chem. 18:469-476; and Bioconjugate Techniques, G. T. Hermanson, Academic Press (2013)). For example, the peptide may be joined to the solid support by a ligation reaction. Alternatively, the solid support can include an agent or coating to facilitate joining, either direct or indirectly, the peptide to the solid support. Any suitable molecule or materials may be employed for this purpose, including proteins, nucleic acids, carbohydrates and small molecules. For example, in one embodiment the agent is an affinity molecule. In another example, the agent is an azide group, which group can react with an alkynyl group in another molecule to facilitate association or binding between the solid support and the other molecule.

[0256] Proteins, polypeptides, or peptides can be joined to the solid support using methods referred to as "click chemistry." For this purpose, any reaction which is rapid and substantially irreversible can be used to attach proteins, polypeptides, or peptides to the solid support. Exemplary reactions include the copper catalyzed reaction of an azide and alkyne to form a triazole (Huisgen 1, 3-dipolar cycloaddition), strain-promoted azide alkyne cycloaddition (SPAAC), reaction of a diene and dienophile (Diels-Alder), strain-promoted alkyne-nitrone cycloaddition, reaction of a strained alkene with an azide, tetrazine or tetrazole, alkene and azide [3+2] cycloaddition, alkene and tetrazine inverse electron demand Diels-Alder (IEDDA) reaction (e.g., m-tetrazine (mTet) or phenyl tetrazine (pTet) and trans-cyclooctene (TCO); or pTet and an alkene), alkene and tetrazole photoreaction, Staudinger ligation of azides and phosphines, and various displacement reactions, such as displacement of a leaving group by nucleophilic attack on an electrophilic atom (Horisawa, Front Physiol (2014). 5: 457; Knall, Hollauf et al., Tetrahedron Lett (2014) 55(34): 4763-4766). Exemplary displacement reactions include reaction of an amine with: an activated ester; an N-hydroxysuccinimide ester; an isocyanate; an isothioscyanate, an aldehyde, an epoxide, or the like.

[0257] In some embodiments, the polypeptide and solid support are joined by a functional group capable of formation by reaction of two complementary reactive groups, for example a functional group which is the product of one of the foregoing "click" reactions. In various embodiments, functional group can be formed by reaction of an aldehyde, oxime, 97erivatiz, hydrazide, alkyne, amine, azide, acylazide, acylhalide, nitrile, nitrone, sulfhydryl, disulfide, sulfonyl halide, isothiocyanate, imidoester, activated ester (e.g., N-hydroxysuccinimide ester, pentynoic acid STP ester), ketone, α,β-unsaturated carbonyl, alkene, maleimide, α-haloimide, epoxide, aziridine, tetrazine, tetrazole, phosphine, biotin or thiirane functional group with a complementary reactive group. An exemplary reaction is a reaction of an amine (e.g., primary amine) with an N-hydroxysuccinimide ester or isothiocyanate.

[0258] In some embodiments, the functional group comprises an alkene, ester, amide, thioester, disulfide, carbocyclic, heterocyclic or heteroaryl group. In further embodiments, the functional group comprises an alkene, ester, amide, thioester, thiourea, disulfide, carbocyclic, heterocyclic or heteroaryl group. In other embodiments, the functional group comprises an amide or thiourea. In some more specific embodiments, functional group is a triazolyl functional group, an amide, or thiourea functional group.

[0259] In some embodiments, iEDDA click chemistry is used for immobilizing polypeptides to a solid support since it is rapid and delivers high yields at low input concentrations. In another embodiment, m-tetrazine rather than tetrazine is used in an iEDDA click chemistry reaction, as m-tetrazine has improved bond stability. In another embodiment, phenyl tetrazine (pTet) is used in an iEDDA click chemistry reaction.

[0260] In some embodiments, the substrate surface is functionalized with TCO, and the recording tag-labeled protein, polypeptide, peptide is immobilized to the TCO coated substrate surface via an attached m-tetrazine moiety.

[0261] In some embodiments, polypeptides are immobilized to a surface of a solid support by its C-terminus, N-terminus, or an internal amino acid, for example, via an amine, carboxyl, or sulfydryl group. Standard activated supports used in coupling to amine groups include CNBr-activated, NHS-activated, aldehyde-activated, azlactone-activated, and CDI-activated supports. Standard activated supports used in carboxyl coupling include carbodiimide-activated carboxyl moieties coupling to amine supports. Cysteine coupling can employ maleimide, idoacetyl, and pyridyl disulfide activated supports. An alternative mode of peptide carboxy terminal immobilization uses anhydrotrypsin, a catalytically inert derivative of trypsin that binds peptides containing lysine or arginine residues at their C-termini without cleaving them.

[0262] In certain embodiments, a polypeptide is immobilized to a solid support via covalent attachment of a solid surface bound linker to a lysine group of the protein, polypeptide, or peptide.

[0263] In certain embodiments, a polypeptide is first labeled with a DNA tag, and the chimeric DNA-polypeptide molecule is immobilized to a solid support via nucleic acid hybridization and ligation to a DNA sequence attached to the solid support. In some embodiments, protein and polypeptide fragmentation into peptides can be performed before or after attachment of a DNA tag or DNA recording tag.B. Optional Processing of Polypeptides

[0264] A sample of polypeptides can undergo protein fractionation methods prior to attachment to a solid support, where proteins or peptides are separated by one or more properties such as cellular location, molecular weight, hydrophobicity, or isoelectric point, or protein enrichment methods. Alternatively, or additionally, protein enrichment methods may be used to select for a specific protein or peptide (see, e.g., Whiteaker et al., (2007) Anal. Biochem. 362:44-54) or to select for a particular post translational modification (see, e.g., Huang et al., (2014) J. Chromatogr. A 1372:1-17). Alternatively, a particular class or classes of proteins such as immunoglobulins, or immunoglobulin (Ig) isotypes such as IgG, can be affinity enriched or selected for analysis. In the case of immunoglobulin molecules, analysis of the sequence and abundance or frequency of hypervariable sequences involved in affinity binding are of particular interest, particularly as they vary in response to disease progression or correlate with healthy, immune, and / or disease phenotypes. Overly abundant proteins can also be subtracted from the sample using standard immunoaffinity methods. Depletion of abundant proteins can be useful for plasma samples where over 80% of the protein constituent is albumin and immunoglobulins. Several commercial products are available for depletion of plasma samples of overly abundant proteins, such as PROTIA and PROT20 (Sigma-Aldrich).

[0265] In some embodiments, the methods provided herein may be performed on polypeptides that have been normalized. In some embodiments, subtraction of certain protein species (e.g., highly abundant proteins) from the sample is performed. This can be accomplished, for example, using commercially available protein depletion reagents such as Sigma's PROT20 immuno-depletion kit, which deplete the top 20 plasma proteins. Additionally, it would be useful to have an approach that greatly reduced the dynamic range even further to a manageable 3-4 orders. In certain embodiments, a protein sample dynamic range can be modulated by fractionating the protein sample using standard fractionation methods, including electrophoresis and liquid chromatography (Zhou et al., Anal Chem (2012) 84(2): 720-734), or partitioning the fractions into compartments (e.g., droplets) loaded with limited capacity protein binding beads / resin (e.g. hydroxylated silica particles) (McCormick, Anal Biochem (1989) 181(1): 66-74) and eluting bound protein. Excess protein in each compartmentalized fraction is washed away.

[0266] Examples of electrophoretic methods include capillary electrophoresis (CE), capillary isoelectric focusing (CIEF), capillary isotachophoresis (CITP), free flow electrophoresis, gel-eluted liquid fraction entrapment electrophoresis (GELFrEE). Examples of liquid chromatography protein separation methods include reverse phase (RP), ion exchange (IE), size exclusion (SE), hydrophilic interaction, etc. Examples of compartment partitions include emulsions, droplets, microwells, physically separated regions on a flat substrate, etc. Exemplary protein binding beads / resins include silica nanoparticles derivatized with phenol groups or hydroxyl groups (e.g., StrataClean Resin from Agilent Technologies, RapidClean from LabTech, etc.). By limiting the binding capacity of the beads / resin, highly-abundant proteins eluting in a given fraction will only be partially bound to the beads, and excess proteins removed.III. EXEMPLARY USE OF MODIFIED DIPEPTIDE CLEAVASE AND RELATED METHODS

[0267] Provided herein is a method of treating one or more polypeptides comprising contacting the polypeptide with a modified dipeptide cleavase. In some embodiments, the modified dipeptide cleavase comprises a mutation, e.g., one or more amino acid modifications in an unmodified dipeptide cleavase, wherein the modified dipeptide cleavase removes a labeled terminal dipeptide from a polypeptide. In some embodiments, polypeptides are contacted with any one or more of the modified dipeptide cleavases as described in Section I. In some embodiments, the method further comprises contacting the polypeptide with a reagent for labeling the terminal amino acid. In some embodiments, the contacting with the reagent for labeling the terminal amino acid is with any one or more of the reagents described in Section I.A. In some embodiments, one or more cycles of contacting the polypeptide with the modified dipeptide cleavase and contacting with a reagent to label the terminal amino acid is performed, such as in a cyclic manner as depicted in FIG. 2A-2C and Fig. 9. In some embodiments, the polypeptide is bound to a support. In some embodiments, the method includes joining the polypeptides to a solid support (e.g., directly or indirectly). In some embodiments, the removal of NTAA as part of a dipeptide from a polypeptide using the provided modified dipeptide cleavases can be combined with a chemical method for removing the NTAA from a peptide, such as described in PCT publication number WO 2019 / 089846.

[0268] In some embodiments, the modified dipeptide cleavases provided herein can be used for treating polypeptides to be analyzed and / or sequenced. In some embodiments, the methods are for determining the sequence of at least a portion of the polypeptide. In some embodiments, the provided methods can be used in the context of a degradation-based polypeptide sequencing assay. In some cases, the method may include performing any of the methods as described in International Patent Publication No. WO 2017 / 192633. In some cases, the sequence of the polypeptide is analyzed by construction of an extended recording tag (e.g., DNA sequence) representing the polypeptide sequence, such as an extended recording tag. In some cases, the methods provided herein apply to or can be used in combination with a ProteoCode ™< assay. In some embodiments employing a cyclic degradation-based polypeptide analysis method, the provided modified dipeptide cleavase provides certain advantages. For example, the recognition and removal of labeled amino acids as dipeptides may provide a pause to amino acid removal as compared to an enzyme which removes unlabeled dipeptides, which may continuously remove amino acids from the polypeptide before other steps of the assay can be performed (e.g., binding of the NTAA by a binding agent and recording information of the NTAA to a recording tag). Thus, in some cases, by recognizing and removing labeled dipeptides, the modified dipeptide cleavase removes the NTAA (as part of a dipeptide) only after a labeling step has occurred. Thereby, the modified dipeptide cleavase provides control over the removal of dipeptides (containing a labeled amino acid) compared to the unmodified dipeptide cleavase which removes unlabeled dipeptides.

[0269] In some embodiments, a method comprising the modified dipeptide cleavase is conducted in the absence of a condition that degrades nucleic acids (e.g., DNA, such as a recording tag). In some embodiments, the method comprising the modified dipeptide cleavase is conducted in the absence of a chemical condition that degrades nucleic acids. In some embodiments, the method comprising the modified dipeptide cleavase is conducted in conditions compatible with a degradation-based polypeptide sequencing assay (e.g., the methods as described in International Patent Publication No. WO 2017 / 192633). In some cases, the method comprising the modified dipeptide cleavase is conducted in the presence of conditions compatible with nucleic acids. In some embodiments, the method comprising the modified dipeptide cleavase is conducted in the absence of a strong acid or a strong base. In some aspects, the strong acid is a strong anhydrous acid. In some examples, the method comprising the modified dipeptide cleavase is conducted in the absence of anhydrous TFA.

[0270] In some embodiments, the method includes contacting the polypeptide with more than one modified dipeptide cleavase. In some cases, various modified dipeptide cleavases may exhibit different characteristics, for example, binding preferences for polypeptides and / or differences in cleaving dipeptides. In some embodiments, different modified dipeptide cleavases may be used in any of the described methods, as a mixture of enzymes or each separately. In some embodiments, the different modified dipeptide cleavases are contacted with polypeptides simultaneously or sequentially.

[0271] In some embodiments, the polypeptide is contacted with one or more additional enzymes to eliminate the NTAA (e.g., a proline aminopeptidase to remove an N-terminal proline, if present). The methods of the invention may include optionally treating the polypeptides with an enzyme to remove one or more NTAAs (e.g., proline aminopeptidase) before, during, or after treatment with any of the provided chemical reagents for labeling the NTAA. The methods of the invention may include optionally treating the polypeptides with an enzyme to remove one or more NTAAs (e.g., proline aminopeptidase) before, during, or after treatment with any of the provided modified dipeptide cleavases. In some embodiments, the enzyme eliminates an NTAA from the polypeptide that is a proline. In some specific examples, the enzyme is a proline aminopeptidase, a proline iminopeptidase (PIP), or a pyroglutamate aminopeptidase (pGAP). In some embodiments, one or more modified dipeptide cleavases are used in combination with other enzymes to treat the polypeptides. In some specific cases, the modified dipeptide cleavase and / or other enzymes are provided as a cocktail.

[0272] In some embodiments, the method further comprises contacting the polypeptide with one or more binding agents capable of binding to the terminal amino acid of the polypeptide, wherein each binding agent comprises a coding tag with identifying information regarding the binding agent. In some cases, the binding agent may bind to a labeled terminal amino acid of the polypeptide. In some further embodiments, the method further comprises transferring the identifying information of the coding tag to a recording tag attached to the polypeptide, thereby generating an extended recording tag on the polypeptide. In some particular embodiments, the method further comprises removing or releasing the one or more binding agents from the polypeptide.

[0273] In some embodiments, one or more steps of contacting the polypeptide with various reagents, including for example, contacting with the modified dipeptide cleavase, with the reagent to label the terminal amino acid, and / or with binding reagent(s), is repeated in a cyclic manner. In some embodiments, provided is a method for analyzing a polypeptide, comprising the steps of: (a) contacting a polypeptide with a binding agent capable of binding to the terminal amino acid of the polypeptide, wherein each binding agent comprises a coding tag with identifying information regarding the binding agent; (b) transferring the identifying information of the coding tag to a recording tag associated with each of the polypeptides to generate an extended recording tag; (c) contacting the polypeptide with a reagent to label the terminal amino acid of the polypeptide; and (d) contacting the polypeptide with a modified dipeptide cleavase comprising a mutation, e.g., one or more amino acid modifications in an unmodified dipeptide cleavase, whereby the modified dipeptide cleavase removes a terminal dipeptide labeled by the reagent in step (c) from the polypeptide. In some embodiments, steps (a)-(d) are repeated for "n" binding cycles, wherein the information of each coding tag of each binding agent that binds to the polypeptide is transferred to the extended recording tag generated from the previous binding cycle to generate an nth order extended recording tag. In some embodiments, the method further comprises (b 1) removing or releasing the one or more binding agents from the plurality of polypeptides. In some examples, the polypeptide is contacted with the reagent to label the terminal amino acid of the polypeptide prior to contacting the polypeptide with the modified dipeptide cleavase. In some embodiments, the polypeptides includes a plurality of polypeptides. In some embodiments, the polypeptide is contacted with a plurality of binding agents. In some embodiments, the polypeptide is contacted with two or more binding agents.

[0274] In some examples, step (a) is performed before step (b); step (a) is performed before step (c); step (a) is performed before step (d); step (b) is performed before step (c); step (b) is performed before step (d); step (c) is performed before step (a); step (c) is performed before step (b); and / or step (c) is performed before step (d). In some particular embodiments, the steps are performed in the order: (a), (b), (c), and (d). In some particular embodiments, the steps are performed in the order: (c), (a), (b), and (d). In some embodiments, the method further comprises (e) analyzing the nth order extended recording tag. In some embodiments, the method further comprises removing the one or more binding agents. In some embodiments, step (b1) is performed after step (a); step (b1) is performed after step (b); step (b1) is performed before step (c); and / or step (b1) is performed before step (d).

[0275] In an exemplary workflow, the treatment and analysis of the polypeptides is as follows: a large collection of polypeptides (e.g., 50 million - 1 billion or more) from a proteolytic digest are immobilized randomly on a single molecule sequencing substrate (e.g., beads) at an appropriate intramolecular spacing. In some cases, the polypeptides are attached to recording tags. In a cyclic manner, the terminal amino acid (e.g., N-terminal amino acid) of each peptide is labeled (e.g., PTC, modified-PTC, Cbz, DNP, SNP, acetyl, guanidinyl, amino guanidinyl, heterocyclic methanimine). In some cases, the labeling of the terminal amino acid can be performed as a later step. The labeled N-terminal amino acid (e.g., PITC-NTAA, Cbz-NTAA, DNP-NTAA, SNP-NTAA, acetyl-NTAA, guanidinylated-NTAA, heterocyclic methanimine-NTAA) of each immobilized peptide is bound by the cognate NTAA binding agent which is attached to a coding tag, and information from the coding tag associated with the bound NTAA binding agent is transferred to the recording tag associated with the immobilized peptide, thereby generating an extended recording tag. In some embodiments, the one or more bindings agents is removed or released from the polypeptides. The labeled NTAA is removed as a dipeptide by contacting with a modified dipeptide cleavase. One or more cycles of the labeling, contacting with the binding agent, transferring identifying information, and removal of the labeled dipeptide can be performed.

[0276] In some examples, the final extended recording tag is optionally flanked by universal priming sites to facilitate downstream amplification and / or DNA sequencing. The forward universal priming site (e.g., Illumina's P5-S1 sequence) can be part of the original recording tag design and the reverse universal priming site (e.g., Illumina's P7-S2' sequence) can be added as a final step in the extension of the recording tag. In some embodiments, the addition of forward and reverse priming sites can be done independently of a binding agent.

[0277] In some embodiments, the order of the steps in the process for a degradation-based peptide or polypeptide sequencing assay can be reversed or be performed in various orders. For example, in some embodiments, the terminal amino acid labeling can be conducted before and / or after the polypeptide is bound to the binding agent. In some embodiments, contacting with the one or more binding agents is before contacting the polypeptide with the reagent for labeling the terminal amino acid. In some cases, contacting with the one or more binding agents is before contacting the polypeptide with the modified dipeptide cleavase to remove the labeled terminal amino acid.

[0278] In some embodiments, the terminal amino acid labeling can be conducted before or after the polypeptide is bound to a support. In some embodiments, the terminal amino acid removal can be conducted before and / or after the polypeptide is bound to the binding agent. In some embodiments, the contacting of the polypeptides with the reagent for labeling the terminal amino acid is before the contacting with the binding agent and the contacting with the one or more binding agents is before the contacting of the polypeptides with the modified dipeptide cleavase. In some embodiments, transferring of the identifying information is performed after the contacting of the polypeptide with the one or more binding agents and before the contacting of the polypeptide with the modified dipeptide cleavase.

[0279] In some of any such embodiments, removing the one or more binding agents is after the transferring of identifying information from the coding tag to a recording tag associated with each of the polypeptides to generate an extended recording tag. In some of any such embodiments, removing the one or more binding agents is before contacting the polypeptides with a reagent to label the terminal amino acid of the polypeptide. In some embodiments, removing the one or more binding agents is before contacting the polypeptide with a modified dipeptide cleavase.

[0280] In some embodiments, the order of any of the steps of the provided methods for treating the proteins or polypeptides can be reversed or be performed in various orders.A. Attaching Recording Tags to Polypeptides

[0281] In some embodiments, the methods provided comprise contacting polypeptides with the modified dipeptide cleavase and optionally other reagents for polypeptide analysis. In one embodiment, the protein or polypeptide is labeled with DNA recording tags through standard amine coupling chemistries. The ε-amino group (e.g., of lysine residues) and the N-terminal amino group are particularly susceptible to labeling with amine-reactive coupling agents, depending on the pH of the reaction (Mendoza et al., Mass Spectrom Rev (2009) 28(5): 785-815). In a particular embodiment, the recording tag is comprised of a reactive moiety (e.g., for conjugation to a solid surface, a multifunctional linker, or a polypeptide), a linker, a universal priming sequence, a barcode (e.g., compartment tag, partition barcode, sample barcode, fraction barcode, or any combination thereof), an optional UMI, and a spacer (Sp) sequence for facilitating information transfer to / from a coding tag. In some cases, wherein ligation is used, the Sp sequence can serve as an overhang of 1-8 bases. In some cases, the recording tag does not include a spacer. In another embodiment, the protein can be first labeled with a universal DNA tag, and the barcode-Sp sequence (representing a sample, a compartment, a physical location on a slide, etc.) are attached to the protein later through and enzymatic or chemical coupling step. A universal DNA tag comprises a short sequence of nucleotides that are used to label a polypeptide and can be used as point of attachment for a barcode (e.g., compartment tag, recording tag, etc.). For example, a recording tag may comprise at its terminus a sequence complementary to the universal DNA tag. In certain embodiments, a universal DNA tag is a universal priming sequence. Upon hybridization of the universal DNA tags on the labeled protein to complementary sequence in recording tags (e.g., bound to beads), the annealed universal DNA tag may be extended via primer extension, transferring the recording tag information to the DNA tagged protein. In a particular embodiment, the protein is labeled with a universal DNA tag prior to proteinase digestion into peptides. The universal DNA tags on the labeled peptides from the digest can then be converted into an informative and effective recording tag. In some embodiments, protein and polypeptide fragmentation into peptides can be performed before or after attachment of a DNA tag or DNA recording tag.

[0282] At least one recording tag is associated or co-localized directly or indirectly with the polypeptide and joined to the solid support. A recording tag may comprise DNA, RNA, or polynucleotide analogs including PNA, gPNA, GNA, HNA, BNA, XNA, TNA, or a combination thereof. A recording tag may be single stranded, or partially or completely double stranded. A recording tag may have a blunt end or overhanging end. In certain embodiments, upon binding of a binding agent to a polypeptide, identifying information of the binding agent's coding tag is transferred to the recording tag to generate an extended recording tag. Further extensions to the extended recording tag can be made in subsequent binding cycles.

[0283] A recording tag can be joined to the solid support, directly or indirectly (e.g., via a linker), by any means known in the art, including covalent and non-covalent interactions, or any combination thereof. For example, the recording tag may be joined to the solid support by a ligation reaction. Alternatively, the solid support can include an agent or coating to facilitate joining, either direct or indirectly, of the recording tag, to the solid support. Strategies for immobilizing nucleic acid molecules to solid supports (e.g., beads) have been described in U.S. Patent 5,900,481; Steinberg et al. (2004) Biopolymers 73:597-605; Lund et al., (1988) Nucleic Acids Res. 16: 10861-10880).

[0284] In certain embodiments, the co-localization of a polypeptide and associated recording tag is achieved by conjugating polypeptide and recording tag to a bifunctional linker attached directly to the solid support surface (Steinberg et al. (2004) Biopolymers 73:597-605). In further embodiments, a trifunctional moiety is used to derivatize the solid support (e.g., beads), and the resulting bifunctional moiety is coupled to both the polypeptide and recording tag. In other embodiments, the co-localization of a polypeptide and associated recording tag is achieved by coupling the polypeptide to the associated DNA recording tag and ligating the chimera to a DNA decorated solid support surface.

[0285] Methods and reagents (e.g., click chemistry reagents and photoaffinity labelling reagents) such as those described for attachment of polypeptides and solid supports, may also be used for attachment of recording tags.

[0286] In a particular embodiment, a single recording tag is attached to a polypeptide, preferably via the attachment to a de-blocked N- or C-terminal amino acid. In another embodiment, multiple recording tags are attached to the polypeptide, preferably to the lysine residues or peptide backbone. In some embodiments, a polypeptide labeled with multiple recording tags is fragmented or digested into smaller peptides, with each peptide labeled on average with one recording tag.

[0287] In certain embodiments, a polypeptide is first labeled with a DNA recording tag, and the chimeric DNA-polypeptide molecule is immobilized to a solid support via nucleic acid hybridization and ligation to a DNA sequence attached to the solid support.

[0288] In certain embodiments, a recording tag comprises an optional, unique molecular identifier (UMI), which provides a unique identifier tag for each polypeptide to which the UMI is associated with. A UMI can be about 3 to about 40 bases, or a subrange thereof, e.g., about 3 to about 30 bases, about 3 to about 20 bases, or about 3 to about 10 bases, or about 3 to about 8 bases. In some embodiments, a UMI is about 3 bases, 4 bases, 5 bases, 6 bases, 7 bases, 8 bases, 9 bases, 10 bases, 11 bases, 12 bases, 13 bases, 14 bases, 15 bases, 16 bases, 17 bases, 18 bases, 19 bases, 20 bases, 25 bases, 30 bases, 35 bases, or 40 bases in length. A UMI can be used to de-convolute sequencing data from a plurality of extended recording tags to identify sequence reads from individual polypeptides. In some embodiments, within a library of polypeptides, each polypeptide is associated with a single recording tag, with each recording tag comprising a unique UMI. In other embodiments, multiple copies of a recording tag are associated with a single polypeptide, with each copy of the recording tag comprising the same UMI. In some embodiments, a UMI has a different base sequence than the spacer or encoder sequences within the binding agents' coding tags to facilitate distinguishing these components during sequence analysis.

[0289] In certain embodiments, a recording tag comprises a barcode, e.g., other than the UMI if present. A barcode is a nucleic acid molecule of about 3 to about 30 bases, or a subrange thereof, e.g., about 3 to about 25 bases, about 3 to about 20 bases, about 3 to about 10 bases, about 3 to about 10 bases, about 3 to about 8 bases in length. In some embodiments, a barcode is about 3 bases, 4 bases, 5 bases, 6 bases, 7 bases, 8 bases, 9 bases, 10 bases, 11 bases, 12 bases, 13 bases, 14 bases, 15 bases, 20 bases, 25 bases, or 30 bases in length. In one embodiment, a barcode allows for multiplex sequencing of a plurality of samples or libraries. A barcode may be used to identify a partition, a fraction, a compartment, a sample, a spatial location, or library from which the polypeptide derived. Barcodes can be used to de-convolute multiplexed sequence data and identify sequence reads from an individual sample or library. For example, a barcoded bead is useful for methods involving emulsions and partitioning of samples, e.g., for purposes of partitioning the proteome.

[0290] A barcode can represent a compartment tag in which a compartment, such as a droplet, microwell, physical region on a solid support, etc. is assigned a unique barcode. The association of a compartment with a specific barcode can be achieved in any number of ways such as by encapsulating a single barcoded bead in a compartment, e.g., by direct merging or adding a barcoded droplet to a compartment, by directly printing or injecting a barcode reagent to a compartment, etc. The barcode reagents within a compartment are used to add compartment-specific barcodes to the polypeptide or fragments thereof within the compartment. Applied to protein partitioning into compartments, the barcodes can be used to map analysed peptides back to their originating protein molecules in the compartment. This can greatly facilitate protein identification. Compartment barcodes can also be used to identify protein complexes.

[0291] In other embodiments, multiple compartments that represent a subset of a population of compartments may be assigned a unique barcode representing the subset.

[0292] Alternatively, a barcode may be a sample identifying barcode. A sample barcode is useful in the multiplexed analysis of a set of samples in a single reaction vessel or immobilized to a single solid substrate or collection of solid substrates (e.g., a planar slide, population of beads contained in a single tube or vessel, etc.). Polypeptides from many different samples can be labeled with recording tags with sample-specific barcodes, and then all the samples pooled together prior to immobilization to a solid support, cyclic binding, and recording tag analysis. Alternatively, the samples can be kept separate until after creation of a DNA-encoded library, and sample barcodes attached during PCR amplification of the DNA-encoded library, and then mixed together prior to sequencing. This approach could be useful when assaying analytes (e.g., proteins) of different abundance classes. For example, the sample can be split and barcoded, and one portion processed using binding agents to low abundance analytes, and the other portion processed using binding agents to higher abundance analytes. In a particular embodiment, this approach helps to adjust the dynamic range of a particular protein analyte assay to lie within the "sweet spot" of standard expression levels of the protein analyte.

[0293] In certain embodiments, polypeptides from multiple different samples are labeled with recording tags containing sample-specific barcodes. The multi-sample barcoded polypeptides can be mixed together prior to a cyclic binding reaction. In this way, a highly-multiplexed alternative to a digital reverse phase protein array (RPPA) is effectively created (Guo et al., Proteome Sci (2012) 10(1): 56; Assadi, Lamerz et al., Mol Cell Proteomics (2013) 12(9): 2615-2622; Akbani et al. 2014; Mol Cell Proteomics (2014) 13(7): 1625-1643; Creighton et al., Drug Des Devel Ther (2015) 9: 3519-3527). The creation of a digital RPPA-like assay has numerous applications in translational research, biomarker validation, drug discovery, clinical, and precision medicine.

[0294] In certain embodiments, a recording tag comprises a universal priming site, e.g., a forward or 5' universal priming site. A universal priming site is a nucleic acid sequence that may be used for priming a library amplification reaction and / or for sequencing. A universal priming site may include, but is not limited to, a priming site for PCR amplification, flow cell adaptor sequences that anneal to complementary oligonucleotides on flow cell surfaces (e.g., Illumina next generation sequencing), a sequencing priming site, or a combination thereof. A universal priming site can be about 10 bases to about 60 bases. In some embodiments, a universal priming site comprises an Illumina P5 primer (5'-AATGATACGGCGACCACCGA-3' - SEQ ID NO:3) or an Illumina P7 primer (5'-CAAGCAGAAGACGGCATACGAGAT - 3' - SEQ ID NO:4).

[0295] In certain embodiments, a recording tag comprises a spacer at its terminus, e.g., 3' end. As used herein reference to a spacer sequence in the context of a recording tag includes a spacer sequence that is identical to the spacer sequence associated with its cognate binding agent, or a spacer sequence that is complementary to the spacer sequence associated with its cognate binding agent. The terminal, e.g., 3', spacer on the recording tag permits transfer of identifying information of a cognate binding agent from its coding tag to the recording tag during the first binding cycle (e.g., via annealing of complementary spacer sequences for primer extension or sticky end ligation).

[0296] In one embodiment, the spacer sequence is about 1-20 bases in length or a subrange thereof, e.g., about 2-12 bases in length, or 5-10 bases in length. The length of the spacer may depend on factors such as the temperature and reaction conditions of the primer extension reaction for transferring coding tag information to the recording tag. In some embodiments, the recording tag does not comprise a spacer.

[0297] In a preferred embodiment, the spacer sequence in the recording is designed to have minimal complementarity to other regions in the recording tag; likewise, the spacer sequence in the coding tag should have minimal complementarity to other regions in the coding tag. In other words, the spacer sequence of the recording tags and coding tags should have minimal sequence complementarity to components such unique molecular identifiers, barcodes (e.g., compartment, partition, sample, spatial location), universal primer sequences, encoder sequences, cycle specific sequences, etc. present in the recording tags or coding tags.

[0298] In some embodiments, the recording tags associated with a library of polypeptides share a common spacer sequence. In other embodiments, the recording tags associated with a library of polypeptides have binding cycle specific spacer sequences that are complementary to the binding cycle specific spacer sequences of their cognate binding agents, which can be useful when using non-concatenated extended recording tags.

[0299] In some cases, the collection of extended recording tags can be concatenated. For example, after the binding cycles are complete, the bead solid supports, each bead comprising on average one or fewer than one polypeptide per bead, each polypeptide having a collection of extended recording tags that are co-localized at the site of the polypeptide, are placed in an emulsion. The emulsion is formed such that each droplet, on average, is occupied by at most 1 bead. An optional assembly PCR reaction is performed in-emulsion to amplify the extended recording tags co-localized with the polypeptide on the bead and assemble them in co-linear order by priming between the different cycle specific sequences on the separate extended recording tags (Xiong et al., FEMS Microbiol Rev (2008) 32(3): 522-540). Afterwards the emulsion is broken and the assembled extended recording tags are sequenced.

[0300] In another embodiment, the DNA recording tag is comprised of a universal priming sequence (U1), one or more barcode sequences (BCs), and a spacer sequence (Sp1) specific to the first binding cycle. In the first binding cycle, binding agents employ DNA coding tags comprised of an Sp1 complementary spacer, an encoder barcode, and optional cycle barcode, and a second spacer element (Sp2). The utility of using at least two different spacer elements is that the first binding cycle selects one of potentially several DNA recording tags and a single DNA recording tag is extended resulting in a new Sp2 spacer element at the end of the extended DNA recording tag. In the second and subsequent binding cycles, binding agents contain just the Sp2' spacer rather than Sp1'. In this way, only the single extended recording tag from the first cycle is extended in subsequent cycles. In another embodiment, the second and subsequent cycles can employ binding agent specific spacers.

[0301] In some embodiments, a recording tag comprises from 5' to 3' direction: a universal forward (or 5') priming sequence, a UMI, and a spacer sequence. In some embodiments, a recording tag comprises from 5' to 3' direction: a universal forward (or 5') priming sequence, an optional UMI, a barcode (e.g., sample barcode, partition barcode, compartment barcode, spatial barcode, or any combination thereof), and a spacer sequence. In some other embodiments, a recording tag comprises from 5' to 3' direction: a universal forward (or 5') priming sequence, a barcode (e.g., sample barcode, partition barcode, compartment barcode, spatial barcode, or any combination thereof), an optional UMI, and a spacer sequence.

[0302] Combinatorial approaches may be used to generate UMIs from modified DNA and PNAs. In one example, a UMI may be constructed by "chemical ligating" together sets of short word sequences (4-15mers), which have been designed to be orthogonal to each other (Spiropulos and Heemstra 2012). A DNA template is used to direct the chemical ligation of the "word" polymers. The DNA template is constructed with hybridizing arms that enable assembly of a combinatorial template structure simply by mixing the sub-components together in solution. In certain embodiments, there are no "spacer" sequences in this design. The size of the word space can vary from 10's of words to 10,000's or more words or a subrange thereof. In certain embodiments, the words are chosen such that they differ from one another to not cross hybridize, yet possess relatively uniform hybridization conditions. In one embodiment, the length of the word will be on the order of 10 bases, with about 1000's words in the subset (this is only 0.1% of the total 10-mer word space ~ 4 10< = 1 million words). Sets of these words (1000 in subset) can be concatenated together to generate a final combinatorial UMI with complexity = 1000 n< power. For 4 words concatenated together, this creates a UMI diversity of 10 12< different elements. These UMI sequences will be appended to the polypeptide at the single molecule level. In one embodiment, the diversity of UMIs exceeds the number of molecules of polypeptides to which the UMIs are attached. In this way, the UMI uniquely identifies the polypeptide of interest. The use of combinatorial word UMI's facilitates readout on high error rate sequencers, (e.g., nanopore sequencers, nanogap tunneling sequencing, etc.) since single base resolution is not required to read words of multiple bases in length. Combinatorial word approaches can also be used to generate other identity-informative components of recording tags or coding tags, such as compartment tags, partition barcodes, spatial barcodes, sample barcodes, encoder sequences, cycle specific sequences, and barcodes. Methods relating to nanopore sequencing and DNA encoding information with error-tolerant words (codes) are known in the art (see, e.g., Kiah et al., 2015, Codes for DNA sequence profiles. IEEE International Symposium on Information Theory (ISIT); Gabrys et al., 2015, Asymmetric Lee distance codes for DNA-based storage. IEEE Symposium on Information Theory (ISIT); Laure et al., 2016, Coding in 2D: Using Intentional Dispersity to Enhance the Information Capacity of Sequence-Coded Polymer Barcodes. Angew. Chem. Int. Ed. doi:10.1002 / anie.201605279; Yazdi et al., 2015, IEEE Transactions on Molecular, Biological and Multi-Scale Communications 1:230-248; and Yazdi et al., 2015, Sci Rep 5:14138). Thus, in certain embodiments, an extended recording tag, an extended coding tag, or a di-tag construct in any of the embodiments described herein is comprised of identifying components (e.g., UMI, encoder sequence, barcode, compartment tag, cycle specific sequence, etc.) that are error correcting codes. In some embodiments, the error correcting code is selected from: Hamming code, Lee distance code, asymmetric Lee distance code, Reed-Solomon code, and Levenshtein-Tenengolts code. For nanopore sequencing, the current or ionic flux profiles and asymmetric base calling errors are intrinsic to the type of nanopore and biochemistry employed, and this information can be used to design more robust DNA codes using the aforementioned error correcting approaches. An alternative to employing robust DNA nanopore sequencing barcodes, one can directly use the current or ionic flux signatures of barcode sequences (U.S. Patent No. 7,060,507), avoiding DNA base calling entirely, and immediately identify the barcode sequence by mapping back to the predicted current / flux signature as described by Laszlo et al. (2014, Nat. Biotechnol. 32:829-833). For example, Laszlo et al. describe the current signatures generated by the biological nanopore, MspA, when passing different word strings through the nanopore, and the ability to map and identify DNA strands by mapping resultant current signatures back to an in silico prediction of possible current signatures from a universe of sequences (Laszlo et al., (2014) Nat. Biotechnol. 32:829-833). Similar concepts can be applied to DNA codes and the electrical signal generated by nanogap tunneling current-based DNA sequencing (Ohshiro et al., 2012, Sci Rep 2: 501).

[0303] Thus, in certain embodiments, the identifying components of a coding tag, recording tag, or both are capable of generating a unique current or ionic flux or optical signature, wherein the analysis step of any of the methods provided herein comprises detection of the unique current or ionic flux or optical signature in order to identify the identifying components. In some embodiments, the identifying components are selected from an encoder sequence, barcode, UMI, compartment tag, cycle specific sequence, or any combination thereof.

[0304] In certain embodiments, all or a substantial amount of the polypeptides (e.g., at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%) within a sample are labeled with a recording tag. Attaching of the recording tag to the polypeptides may occur before or after immobilization of the polypeptides to a solid support.

[0305] In other embodiments, a subset of polypeptides within a sample are labeled with recording tags. In a particular embodiment, a subset of polypeptides from a sample undergo targeted (analyte specific) labeling with recording tags. Targeted recording tag labeling of proteins may be achieved using target protein-specific binding agents (e.g., antibodies, aptamers, etc.) that are linked a short target-specific DNA capture probe, e.g., analyte-specific barcode, which anneal to complementary target-specific bait sequence, e.g., analyte-specific barcode, in recording tags. The recording tags comprise a reactive moiety for a cognate reactive moiety present on the target protein (e.g., click chemistry labeling, photoaffinity labeling). For example, recording tags may comprise an azide moiety for interacting with alkyne-derivatized proteins, or recording tags may comprise a benzophenone for interacting with native proteins, etc. Upon binding of the target protein by the target protein specific binding agent, the recording tag and target protein are coupled via their corresponding reactive. After the target protein is labeled with the recording tag, the target-protein specific binding agent may be removed by digestion of the DNA capture probe linked to the target-protein specific binding agent. For example, the DNA capture probe may be designed to contain uracil bases, which are then targeted for digestion with a uracil-specific excision reagent (e.g., USER ™< ), and the target-protein specific binding agent may be dissociated from the target protein.

[0306] In one example, antibodies specific for a set of target proteins can be labeled with a DNA capture probe that hybridizes with recording tags designed with complementary bait sequence. Sample-specific labeling of proteins can be achieved by employing DNA-capture probe labeled antibodies hybridizing with complementary bait sequence on recording tags comprising of sample-specific barcodes.

[0307] In another example, target protein-specific aptamers are used for targeted recording tag labeling of a subset of proteins within a sample. A target specific-aptamer is linked to a DNA capture probe that anneals with complementary bait sequence in a recording tag. The recording tag comprises a reactive chemical or photo-reactive chemical probes (e.g. benzophenone (BP)) for coupling to the target protein having a corresponding reactive moiety. The aptamer binds to its target protein molecule, bringing the recording tag into close proximity to the target protein, resulting in the coupling of the recording tag to the target protein.

[0308] Photoaffinity (PA) protein labeling using photo-reactive chemical probes attached to small molecule protein affinity ligands has been previously described (Park, Koh et al. 2016). Typical photo-reactive chemical probes include probes based on benzophenone (reactive diradical, 365 nm), phenyldiazirine (reactive carbon, 365 nm), and phenylazide (reactive nitrene free radical, 260 nm), activated under irradiation wavelengths as previously described (Smith et al., Future Med Chem. (2015) 7(2): 159-183). In a preferred embodiment, target proteins within a protein sample are labeled with recording tags comprising sample barcodes using the method disclosed by Li et al., in which a bait sequence in a benzophenone labeled recording tag is hybridized to a DNA capture probe attached to a cognate binding agent (e.g., nucleic acid aptamer (Li et al., Angew Chem Int Ed Engl (2013) 52(36): 9544-9549). For photoaffinity labeled protein targets, the use of DNA / RNA aptamers as target protein-specific binding agents are preferred over antibodies since the photoaffinity moiety can self-label the antibody rather than the target protein. In contrast, photoaffinity labeling is less efficient for nucleic acids than proteins, making aptamers a better vehicle for DNA-directed chemical or photo-labeling. Similar to photo-affinity labeling, one can also employ DNA-directed chemical labeling of reactive lysine's (or other moieties) in the proximity of the aptamer binding site in a manner similar to that described by Rosen et al. (Rosen et al, Nature Chemistry volume (2014) 6:804-809; Kodal et al., ChemBioChem (2016) 17:1338-1342).

[0309] In the aforementioned embodiments, other types of linkages besides hybridization can be used to link the target specific binding agent and the recording tag. For example, the two moieties can be covalently linked, using a linker that is designed to be cleaved and release the binding agent once the captured target protein (or other polypeptide) is covalently linked to the recording tag. A suitable linker can be attached to various positions of the recording tag, such as the 3' end, or within the linker attached to the 5' end of the recording tag.

[0310] Recording tags can be attached to the protein, polypeptide, or peptides pre- or post-immobilization to the solid support. For example, proteins, polypeptides, or peptides can be first labeled with recording tags and then immobilized to a solid surface via a recording tag comprising at two functional moieties for coupling. One functional moiety of the recording tag couples to the protein, and the other functional moiety immobilizes the recording tag-labeled protein to a solid support.

[0311] In other embodiments, polypeptides are immobilized to a solid support prior to labeling of the proteins, polypeptides or peptides with recording tags. For example, proteins can first be derivatized with reactive groups such as click chemistry moieties. The activated protein molecules can then be attached to a suitable solid support and then labeled with recording tags using the complementary click chemistry moiety. As an example, proteins derivatized with alkyne and mTet moieties may be immobilized to beads derivatized with azide and TCO and attached to recording tags labeled with azide and TCO.

[0312] In certain embodiments, the surface of a solid support is passivated (blocked) to minimize non-specific absorption to binding agents. A "passivated" surface refers to a surface that has been treated ...

Claims

1. A modified dipeptide cleavase comprising three or more amino acid substitutions in a substrate binding site of a dipeptidyl aminopeptidase, wherein: (i) the dipeptidyl aminopeptidase removes or is configured to remove two terminal amino acids from a polypeptide; and (ii) the modified dipeptide cleavase removes or is configured to remove from the polypeptide having a terminal amino acid residue labeled with a chemical reagent (a) the single labeled terminal amino acid residue or (b) a labeled terminal dipeptide, wherein the dipeptidyl aminopeptidase comprises an amino acid sequence having at least 30 % sequence identity to the amino acid sequence of SEQ ID NO: 31 and also comprises an asparagine residue at a position corresponding to position 191 of SEQ ID NO: 31, a tryptophan or phenylalanine residue at a position corresponding to position 192 of SEQ ID NO: 31, an arginine residue at a position corresponding to position 196 of SEQ ID NO: 31, an asparagine residue at a position corresponding to position 306 of SEQ ID NO: 31, an aspartate residue at a position corresponding to position 650 of SEQ ID NO: 31; and wherein the modified dipeptide cleavase comprises three or more amino acid substitutions in residues corresponding to positions 191, 192, 196, 306, 650 of SEQ ID NO: 31.

2. The modified dipeptide cleavase of claim 1, wherein the modified dipeptide cleavase does not remove an unlabeled terminal dipeptide from the polypeptide.

3. The modified dipeptide cleavase of claim 1, which comprises at least four amino acid substitutions in residues corresponding to positions 191, 192, 196, 306, 650 of SEQ ID NO: 31.

4. The modified dipeptide cleavase of claim 1, wherein the modified dipeptide cleavase removes or is configured to remove a single N-terminally labeled amino acid of the polypeptide.

5. The modified dipeptide cleavase of claim 1, further comprising one or more amino acid substitutions in residues corresponding to positions 310, 628, 648, 651, 659, 669.

6. The modified dipeptide cleavase of claim 1, wherein the dipeptidyl aminopeptidase is a protein classified in MEROPS S46, or a functional homolog or fragment thereof.

7. The modified dipeptide cleavase of claim 1, wherein the single terminal amino acid or terminal dipeptide is labeled with a N-terminal modification that comprises a N-terminal blocking group (NTMblk) and, optionally, a natural or unnatural amino acid portion (NTMaa), wherein the NTMaa comprises a compound selected from the group consisting of: a naturally-occurring amino acid residue, 3-(3'-pyridyl)-L-alanine, L-cyclohexylglycine, α-aminoisobutyric acid, 3-(4'-pyridyl)-L-alanine, L-azetidine-2-carboxylic acid, isonipecotic acid, L-phenylglycine, β-(2-thienyl)-L-alanine, 3-(4-thiazolyl)-L-alanine, 1-aminocyclopentane-1-carboxylic acid, (2-trifluoromethyl)-L-Phenylalanine, L-cyclopropylalanine, 3-(2'-pyridyl)-L-alanine, beta-cyano-L-alanine, α-methyl-L-4-Fluorophenylalanine, α-methyl-D-4-fluorophenylalanine, 3-amino-2,2-difluoro-propionic acid, O-sulfo-L-tyrosine sodium salt, L-2-furylalanine, 1-aminocyclopropane-1-carboxylic acid, 3,5-dinitro-L-tyrosine, pentafluoro-L-phenylalanine, 3,5-difluoro-L-phenylalanine, 3-fluoro-L-phenylalanine, N-cyclopentylglycine, 1-(amino)cyclohexanecarboxylic acid, N-methylalanine, 4-amino-tetrahydropyran-4-carboxylic acid, 4-amino-1,1-dioxothiane-4-carboxylic acid, 4-amino-1-methyl-4-piperidinecarboxylic acid, 2-amino-N-(2,4-dimethoxybenzyl)acetamido)acetic acid, or N-alkylated derivatives; and the NTMblk comprises a compound selected from the group consisting of: 4-methylbenzoic acid, 4-(dimethylamio)benzoic acid, nicotinic acid, 3-aminonicotinic acid, 2-pyrazinecarbooxylic acid, 5-amino-2-fluoro-isonicotinic acid, 2,3-pyrazinedicarboxylic acid, 4,7-Difluoroisobenzofuran-1,3-dicarboxylic acid, 4-chloro-2-aminobenzoic acid, 4-nitro-2-aminobenzoic acid, 7-methoxy-1h-benzo[d][1,3]oxazine-2,4-dione, 4-carboxy-2-aminobenzoic acid, 6-(Trifluoromethyl)-2,4-dihydro-1h-3,1-benzoxazine-2,4-dione, 7-(Trifluoromethyl)-1h-benzo[d][1,3]oxazine-2,4-dione, 6-fluoro-2-aminobenzoic acid, 4-fluoro-2-aminobenzoic acid, 5-methoxy-2-aminobenzoic acid, 4-fluorobenzoic acid, 4-(trifluoromethyl)benzoic acid, 2-ethynyl-6-fluorobenzaldehyde, 2-aminobenzoic acid, Succinic anhydride, 3,6-Difluoropyridine-2-carboxylic acid, 2-Fluoronicotinic acid, 5-Bromo-2-hydroxynicotinic acid, 4-(Trifluoromethyl)pyrimidine-5-carboxylic acid, 2-Oxo-1,2-dihydropyridine-3-carboxylic acid, 5-Methyl-2-aminobenzoic acid, 6-Fluoropicolinic acid, 3-Methyl-2-aminobenzoic acid, 4-Methyl-2-aminobenzoic acid, 2-Amino-6-methylbenzoic acid, 2-Amino-6-fluorobenzoic acid, 2-Amino-5-fluorobenzoic acid, 2-Amino-3-fluorobenzoic acid, 2-Amino-4-fluorobenzoic acid, 2-Aminonicotinic acid, 4-Aminonicotinic acid, 3-Aminopicolinic acid, 2-Amino-4,5-difluorobenzoic acid, 3,4-difluorobenzoic acid, 3,4,5-difluorobenzoic acid, 3-(Methoxycarbonyl)bicyclo[1.1.1]pentane-1-carboxylic acid, 3,3-Difluorocyclobutane-1-carboxylic acid, 1-Methyl-2-oxo-piperidine-4-carboxylic acid, Tetrahydropyran-4-carboxylic acid, 5-Fluoroorotic acid, 3-Fluoro-4-nitrobenzoic acid, 3-(Difluoromethyl)-1-methyl-1H-pyrazole-4-carboxylic acid, 4-(Difluoromethoxy)benzoic acid, 1-(Difluoromethyl)-1h-pyrazole-3-carboxylic acid, 4-(Methanesulfonylamino)benzoic acid, 5-Fluoro-6-methoxynicotinic acid, Tetrahydro-2H-thiopyran-4-carboxylic acid 1,1-dioxide, 4-(1H-Tetrazol-5-yl)benzoic acid, 1,2,3-Thiadiazole-4-carboxylic acid, 1,3-Benzodioxole-4-carboxylic acid, 2,1,3-Benzoxadiazole-5-carboxylic acid, 1-Benzyl-3-methyl-1h-pyrazole-5-carboxylic acid, 1-Cyclopropyl-6,7-difluoro-1,4-dihydro-4-oxoquinoline-3-carboxylic acid, 3,4-Dichlorobenzoic acid, 5-Fluoro-6-methylpyridine-2-carboxylic acid, 4,5-Dimethyl-2-(1h-pyrrol-1-yl)thiophene-3-carboxylic acid, 1,3-Dimethyl-1h-thieno[2,3-c]pyrazole-5-carboxylic acid, 1-[(4-Fluorobenzene)sulfonyl]piperidine-3-carboxylic acid, 1-(4-Fluorobenzyl)-5-oxopyrrolidine-3-carboxylic acid, 3-Fluoro-4-methoxybenzoic acid, 4-Fluoro-3-nitrobenzoic acid, 6-Fluoro-4-oxochromene-2-carboxylic acid, 3-Fluorophenylacetic acid, 4-Fluoro-3-(trifluoromethyl)benzoic acid, 5-Furan-2-yl-isoxazole-3-carboxylic acid, 1-Isopropyl-2-(trifluoromethyl)-1h-benzimidazole-5-carboxylic acid, Levofloxacin carboxylic acid, 3,5,7-Trifluoroadamantane-1-carboxylic acid, 3,4,5-Trimethoxybenzoic acid, 2-Oxo-2,3-dihydro-1h-benzo[d]imidazole-4-carboxylic acid, 1-Methyl-3-(trifluoromethyl)-1h-pyrazole-5-carboxylic acid, 2-Morpholin-4-yl-isonicotinic acid, 1,3-Oxazole-4-carboxylic acid, 4-Carboxybenzenesulfonamide, 3,4-difluorobenzenesulfonyl chloride.

8. A method of treating a polypeptide, comprising the following steps: labeling a terminal amino acid of the polypeptide with a chemical reagent to produce a labeled polypeptide; and contacting the labeled polypeptide with a modified dipeptide cleavase, wherein the modified dipeptide cleavase comprises three or more amino acid substitutions in a substrate binding site of a dipeptidyl aminopeptidase, wherein: (i) the dipeptidyl aminopeptidase removes or is configured to remove two terminal amino acids from the polypeptide upon contacting; and (ii) the modified dipeptide cleavase removes or is configured to remove from the polypeptide having a terminal amino acid labeled with a chemical reagent (a) the single labeled terminal amino acid or (b) a labeled terminal dipeptide, wherein the dipeptidyl aminopeptidase comprises an amino acid sequence having at least 30 % sequence identity to the amino acid sequence of SEQ ID NO: 31 and also comprises an asparagine residue at a position corresponding to position 191 of SEQ ID NO: 31, a tryptophan or phenylalanine residue at a position corresponding to position 192 of SEQ ID NO: 31, an arginine residue at a position corresponding to position 196 of SEQ ID NO: 31, an asparagine residue at a position corresponding to position 306 of SEQ ID NO: 31, an aspartate residue at a position corresponding to position 650 of SEQ ID NO: 31; and wherein the modified dipeptide cleavase comprises three or more amino acid substitutions in residues corresponding to positions 191, 192, 196, 306, 650 of SEQ ID NO: 31.

9. The method of claim 8, wherein the modified dipeptide cleavase does not remove an unlabeled terminal dipeptide from the polypeptide.

10. The method of claim 8, wherein the modified dipeptide cleavase comprises at least four amino acid substitutions in residues corresponding to positions 191, 192, 196, 306, 650 of SEQ ID NO: 31.

11. The method of claim 8, wherein the single terminal amino acid or terminal dipeptide is labeled with a N-terminal modification that comprises a N-terminal blocking group (NTMblk) and, optionally, a natural or unnatural amino acid portion (NTMaa), wherein the NTMaa comprises a compound selected from the group consisting of: a naturally-occurring amino acid residue, 3-(3'-pyridyl)-L-alanine, L-cyclohexylglycine, α-aminoisobutyric acid, 3-(4'-pyridyl)-L-alanine, L-azetidine-2-carboxylic acid, isonipecotic acid, L-phenylglycine, β-(2-thienyl)-L-alanine, 3-(4-thiazolyl)-L-alanine, 1-aminocyclopentane-1-carboxylic acid, (2-trifluoromethyl)-L-Phenylalanine, L-cyclopropylalanine, 3-(2'-pyridyl)-L-alanine, beta-cyano-L-alanine, α-methyl-L-4-Fluorophenylalanine, α-methyl-D-4-fluorophenylalanine, 3-amino-2,2-difluoro-propionic acid, O-sulfo-L-tyrosine sodium salt, L-2-furylalanine, 1-aminocyclopropane-1-carboxylic acid, 3,5-dinitro-L-tyrosine, pentafluoro-L-phenylalanine, 3,5-difluoro-L-phenylalanine, 3-fluoro-L-phenylalanine, N-cyclopentylglycine, 1-(amino)cyclohexanecarboxylic acid, N-methylalanine, 4-amino-tetrahydropyran-4-carboxylic acid, 4-amino-1,1-dioxothiane-4-carboxylic acid, 4-amino-1-methyl-4-piperidinecarboxylic acid, 2-amino-N-(2,4-dimethoxybenzyl)acetamido)acetic acid, or N-alkylated derivatives; and the NTMblk comprises a compound selected from the group consisting of: 4-methylbenzoic acid, 4-(dimethylamio)benzoic acid, nicotinic acid, 3-aminonicotinic acid, 2-pyrazinecarbooxylic acid, 5-amino-2-fluoro-isonicotinic acid, 2,3-pyrazinedicarboxylic acid, 4,7-Difluoroisobenzofuran-1,3-dicarboxylic acid, 4-chloro-2-aminobenzoic acid, 4-nitro-2-aminobenzoic acid, 7-methoxy-1h-benzo[d][1,3]oxazine-2,4-dione, 4-carboxy-2-aminobenzoic acid, 6-(Trifluoromethyl)-2,4-dihydro-1h-3,1-benzoxazine-2,4-dione, 7-(Trifluoromethyl)-1h-benzo[d][1,3]oxazine-2,4-dione, 6-fluoro-2-aminobenzoic acid, 4-fluoro-2-aminobenzoic acid, 5-methoxy-2-aminobenzoic acid, 4-fluorobenzoic acid, 4-(trifluoromethyl)benzoic acid, 2-ethynyl-6-fluorobenzaldehyde, 2-aminobenzoic acid, Succinic anhydride, 3,6-Difluoropyridine-2-carboxylic acid, 2-Fluoronicotinic acid, 5-Bromo-2-hydroxynicotinic acid, 4-(Trifluoromethyl)pyrimidine-5-carboxylic acid, 2-Oxo-1,2-dihydropyridine-3-carboxylic acid, 5-Methyl-2-aminobenzoic acid, 6-Fluoropicolinic acid, 3-Methyl-2-aminobenzoic acid, 4-Methyl-2-aminobenzoic acid, 2-Amino-6-methylbenzoic acid, 2-Amino-6-fluorobenzoic acid, 2-Amino-5-fluorobenzoic acid, 2-Amino-3-fluorobenzoic acid, 2-Amino-4-fluorobenzoic acid, 2-Aminonicotinic acid, 4-Aminonicotinic acid, 3-Aminopicolinic acid, 2-Amino-4,5-difluorobenzoic acid, 3,4-difluorobenzoic acid, 3,4,5-difluorobenzoic acid, 3-(Methoxycarbonyl)bicyclo[1.1.1]pentane-1-carboxylic acid, 3,3-Difluorocyclobutane-1-carboxylic acid, 1-Methyl-2-oxo-piperidine-4-carboxylic acid, Tetrahydropyran-4-carboxylic acid, 5-Fluoroorotic acid, 3-Fluoro-4-nitrobenzoic acid, 3-(Difluoromethyl)-1-methyl-1H-pyrazole-4-carboxylic acid, 4-(Difluoromethoxy)benzoic acid, 1-(Difluoromethyl)-1h-pyrazole-3-carboxylic acid, 4-(Methanesulfonylamino)benzoic acid, 5-Fluoro-6-methoxynicotinic acid, Tetrahydro-2H-thiopyran-4-carboxylic acid 1,1-dioxide, 4-(1H-Tetrazol-5-yl)benzoic acid, 1,2,3-Thiadiazole-4-carboxylic acid, 1,3-Benzodioxole-4-carboxylic acid, 2,1,3-Benzoxadiazole-5-carboxylic acid, 1-Benzyl-3-methyl-1h-pyrazole-5-carboxylic acid, 1-Cyclopropyl-6,7-difluoro-1,4-dihydro-4-oxoquinoline-3-carboxylic acid, 3,4-Dichlorobenzoic acid, 5-Fluoro-6-methylpyridine-2-carboxylic acid, 4,5-Dimethyl-2-(1h-pyrrol-1-yl)thiophene-3-carboxylic acid, 1,3-Dimethyl-1h-thieno[2,3-c]pyrazole-5-carboxylic acid, 1-[(4-Fluorobenzene)sulfonyl]piperidine-3-carboxylic acid, 1-(4-Fluorobenzyl)-5-oxopyrrolidine-3-carboxylic acid, 3-Fluoro-4-methoxybenzoic acid, 4-Fluoro-3-nitrobenzoic acid, 6-Fluoro-4-oxochromene-2-carboxylic acid, 3-Fluorophenylacetic acid, 4-Fluoro-3-(trifluoromethyl)benzoic acid, 5-Furan-2-yl-isoxazole-3-carboxylic acid, 1-Isopropyl-2-(trifluoromethyl)-1h-benzimidazole-5-carboxylic acid, Levofloxacin carboxylic acid, 3,5,7-Trifluoroadamantane-1-carboxylic acid, 3,4,5-Trimethoxybenzoic acid, 2-Oxo-2,3-dihydro-1h-benzo[d]imidazole-4-carboxylic acid, 1-Methyl-3-(trifluoromethyl)-1h-pyrazole-5-carboxylic acid, 2-Morpholin-4-yl-isonicotinic acid, 1,3-Oxazole-4-carboxylic acid, 4-Carboxybenzenesulfonamide, 3,4-difluorobenzenesulfonyl chloride.

12. The method of claim 8, further comprising a step of contacting the polypeptide with a binding agent configured to bind to the single labeled terminal amino acid or to the labeled terminal dipeptide, wherein the step of labeling the terminal amino acid of the polypeptide is before the step of contacting the polypeptide with the binding agent; and the step of contacting the polypeptide with the binding agent is before the step of contacting the polypeptide with the modified dipeptide cleavase.

13. A set of dipeptide cleavase enzymes, comprising at least two different modified dipeptide cleavases, wherein: (i) each of the modified dipeptide cleavases from the set of dipeptide cleavase enzymes comprises three or more amino acid substitutions in a substrate binding site of a dipeptidyl aminopeptidase, wherein the dipeptidyl aminopeptidase is configured to remove two terminal amino acids from an unlabeled polypeptide; (ii) each of the modified dipeptide cleavases from the set of dipeptide cleavase enzymes is configured to remove a single labeled terminal amino acid from a polypeptide having the terminal amino acid labeled with a chemical reagent, wherein the dipeptidyl aminopeptidase comprises an amino acid sequence having at least 30 % sequence identity to the amino acid sequence of SEQ ID NO: 31 and also comprises an asparagine residue at a position corresponding to position 191 of SEQ ID NO: 31, a tryptophan or phenylalanine residue at a position corresponding to position 192 of SEQ ID NO: 31, an arginine residue at a position corresponding to position 196 of SEQ ID NO: 31, an asparagine residue at a position corresponding to position 306 of SEQ ID NO: 31, an aspartate residue at a position corresponding to position 650 of SEQ ID NO: 31; and wherein the modified dipeptide cleavase comprises three or more amino acid substitutions in residues corresponding to positions 191, 192, 196, 306, 650 of SEQ ID NO: 31; and (iii) the modified dipeptide cleavases from the set of dipeptide cleavase enzymes have different specificities for the labeled terminal amino acids, which the modified dipeptide cleavases are configured to remove.

14. A kit for treating a polypeptide, comprising: (a) a chemical reagent for labeling a terminal amino acid of the polypeptide; and (b) a set of dipeptide cleavase enzymes, comprising at least two different modified dipeptide cleavases, wherein: (i) each of the modified dipeptide cleavases from the set of dipeptide cleavase enzymes comprises three or more amino acid substitutions in a substrate binding site of a dipeptidyl aminopeptidase, wherein the dipeptidyl aminopeptidase is configured to remove two terminal amino acids from an unlabeled polypeptide; (ii) each of the modified dipeptide cleavases from the set of dipeptide cleavase enzymes is configured to remove a single labeled terminal amino acid from a polypeptide having the terminal amino acid labeled with a chemical reagent, wherein the dipeptidyl aminopeptidase comprises an amino acid sequence having at least 30 % sequence identity to the amino acid sequence of SEQ ID NO: 31 and also comprises an asparagine residue at a position corresponding to position 191 of SEQ ID NO: 31, a tryptophan or phenylalanine residue at a position corresponding to position 192 of SEQ ID NO: 31, an arginine residue at a position corresponding to position 196 of SEQ ID NO: 31, an asparagine residue at a position corresponding to position 306 of SEQ ID NO: 31, an aspartate residue at a position corresponding to position 650 of SEQ ID NO: 31; and wherein the modified dipeptide cleavase comprises three or more amino acid substitutions in residues corresponding to positions 191, 192, 196, 306, 650 of SEQ ID NO: 31; and (iii) the modified dipeptide cleavases from the set of dipeptide cleavase enzymes have different specificities for the labeled terminal amino acids, which the modified dipeptide cleavases are configured to remove.

15. The kit of claim 14, wherein (i) the chemical reagent is configured to attach a N-terminal modification to the terminal amino acid of the polypeptide; (ii) the N-terminal modification comprises a N-terminal blocking group (NTMblk) and, optionally, a natural or unnatural amino acid portion (NTMaa); (iii) the NTMaa comprises a compound selected from the group consisting of: a naturally-occurring amino acid residue, 3-(3'-pyridyl)-L-alanine, L-cyclohexylglycine, α-aminoisobutyric acid, 3-(4'-pyridyl)-L-alanine, L-azetidine-2-carboxylic acid, isonipecotic acid, L-phenylglycine, β-(2-thienyl)-L-alanine, 3-(4-thiazolyl)-L-alanine, 1-aminocyclopentane-1-carboxylic acid, (2-trifluoromethyl)-L-Phenylalanine, L-cyclopropylalanine, 3-(2'-pyridyl)-L-alanine, beta-cyano-L-alanine, α-methyl-L-4-Fluorophenylalanine, α-methyl-D-4-fluorophenylalanine, 3-amino-2,2-difluoro-propionic acid, O-sulfo-L-tyrosine sodium salt, L-2-furylalanine, 1-aminocyclopropane-1-carboxylic acid, 3,5-dinitro-L-tyrosine, pentafluoro-L-phenylalanine, 3,5-difluoro-L-phenylalanine, 3-fluoro-L-phenylalanine, N-cyclopentylglycine, 1-(amino)cyclohexanecarboxylic acid, N-methylalanine, 4-amino-tetrahydropyran-4-carboxylic acid, 4-amino-1,1-dioxothiane-4-carboxylic acid, 4-amino-1-methyl-4-piperidinecarboxylic acid, 2-amino-N-(2,4-dimethoxybenzyl)acetamido)acetic acid, or N-alkylated derivatives; and (iv) the NTMblk comprises a compound selected from the group consisting of: 4-methylbenzoic acid, 4-(dimethylamio)benzoic acid, nicotinic acid, 3-aminonicotinic acid, 2-pyrazinecarbooxylic acid, 5-amino-2-fluoro-isonicotinic acid, 2,3-pyrazinedicarboxylic acid, 4,7-Difluoroisobenzofuran-1,3-dicarboxylic acid, 4-chloro-2-aminobenzoic acid, 4-nitro-2-aminobenzoic acid, 7-methoxy-1h-benzo[d][1,3]oxazine-2,4-dione, 4-carboxy-2-aminobenzoic acid, 6-(Trifluoromethyl)-2,4-dihydro-1h-3,1-benzoxazine-2,4-dione, 7-(Trifluoromethyl)-1h-benzo[d][1,3]oxazine-2,4-dione, 6-fluoro-2-aminobenzoic acid, 4-fluoro-2-aminobenzoic acid, 5-methoxy-2-aminobenzoic acid, 4-fluorobenzoic acid, 4-(trifluoromethyl)benzoic acid, 2-ethynyl-6-fluorobenzaldehyde, 2-aminobenzoic acid, Succinic anhydride, 3,6-Difluoropyridine-2-carboxylic acid, 2-Fluoronicotinic acid, 5-Bromo-2-hydroxynicotinic acid, 4-(Trifluoromethyl)pyrimidine-5-carboxylic acid, 2-Oxo-1,2-dihydropyridine-3-carboxylic acid, 5-Methyl-2-aminobenzoic acid, 6-Fluoropicolinic acid, 3-Methyl-2-aminobenzoic acid, 4-Methyl-2-aminobenzoic acid, 2-Amino-6-methylbenzoic acid, 2-Amino-6-fluorobenzoic acid, 2-Amino-5-fluorobenzoic acid, 2-Amino-3-fluorobenzoic acid, 2-Amino-4-fluorobenzoic acid, 2-Aminonicotinic acid, 4-Aminonicotinic acid, 3-Aminopicolinic acid, 2-Amino-4,5-difluorobenzoic acid, 3,4-difluorobenzoic acid, 3,4,5-difluorobenzoic acid, 3-(Methoxycarbonyl)bicyclo[1.1.1]pentane-1-carboxylic acid, 3,3-Difluorocyclobutane-1-carboxylic acid, 1-Methyl-2-oxo-piperidine-4-carboxylic acid, Tetrahydropyran-4-carboxylic acid, 5-Fluoroorotic acid, 3-Fluoro-4-nitrobenzoic acid, 3-(Difluoromethyl)-1-methyl-1H-pyrazole-4-carboxylic acid, 4-(Difluoromethoxy)benzoic acid, 1-(Difluoromethyl)-1h-pyrazole-3-carboxylic acid, 4-(Methanesulfonylamino)benzoic acid, 5-Fluoro-6-methoxynicotinic acid, Tetrahydro-2H-thiopyran-4-carboxylic acid 1,1-dioxide, 4-(1H-Tetrazol-5-yl)benzoic acid, 1,2,3-Thiadiazole-4-carboxylic acid, 1,3-Benzodioxole-4-carboxylic acid, 2,1,3-Benzoxadiazole-5-carboxylic acid, 1-Benzyl-3-methyl-1h-pyrazole-5-carboxylic acid, 1-Cyclopropyl-6,7-difluoro-1,4-dihydro-4-oxoquinoline-3-carboxylic acid, 3,4-Dichlorobenzoic acid, 5-Fluoro-6-methylpyridine-2-carboxylic acid, 4,5-Dimethyl-2-(1h-pyrrol-1-yl)thiophene-3-carboxylic acid, 1,3-Dimethyl-1h-thieno[2,3-c]pyrazole-5-carboxylic acid, 1-[(4-Fluorobenzene)sulfonyl]piperidine-3-carboxylic acid, 1-(4-Fluorobenzyl)-5-oxopyrrolidine-3-carboxylic acid, 3-Fluoro-4-methoxybenzoic acid, 4-Fluoro-3-nitrobenzoic acid, 6-Fluoro-4-oxochromene-2-carboxylic acid, 3-Fluorophenylacetic acid, 4-Fluoro-3-(trifluoromethyl)benzoic acid, 5-Furan-2-yl-isoxazole-3-carboxylic acid, 1-Isopropyl-2-(trifluoromethyl)-1h-benzimidazole-5-carboxylic acid, Levofloxacin carboxylic acid, 3,5,7-Trifluoroadamantane-1-carboxylic acid, 3,4,5-Trimethoxybenzoic acid, 2-Oxo-2,3-dihydro-1h-benzo[d]imidazole-4-carboxylic acid, 1-Methyl-3-(trifluoromethyl)-1h-pyrazole-5-carboxylic acid, 2-Morpholin-4-yl-isonicotinic acid, 1,3-Oxazole-4-carboxylic acid, 4-Carboxybenzenesulfonamide, 3,4-difluorobenzenesulfonyl chloride.