Mirror-image trypsin, the preparation method and use thereof

EP4801936A1Pending Publication Date: 2026-09-09WESTLAKE UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2024809426
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-01
Filing Date
2024-10-30
Publication Date
2026-09-09

AI Technical Summary

Technical Problem

The development of mirror-image biology systems is hindered by the lack of effective methods to sequence mirror-image (D-) proteins, as existing sequencing strategies require digestion by a site-specific D-protease like trypsin, which is not currently chemically synthesized.

Method used

Chemical synthesis of a mirror-image version of trypsin, specifically by synthesizing and in vitro folding its zymogen form, trypsinogen, which is then activated to form a functional mirror-image trypsin, capable of digesting D-peptides and D-proteins.

Benefits of technology

The synthesized mirror-image trypsin enables effective sequencing of long D-peptides and D-proteins, facilitating research in mirror-image biology and related applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2024060671_08052025_PF_FP_ABST
    Figure IB2024060671_08052025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a mirror-image trypsin, the preparation method and use thereof. In particular, the present disclosure provides a novel mirror-image trypsin which digests D-peptides and / or D-proteins, a mirror-image trypsinogen, a kit comprising said mirror-image trypsin, a method for chemically synthesizing a trypsin, a method for sequencing D-peptide and / or D-protein, a method for writing and reading information in a D-peptide and / or D-protein, as well as the use of the mirror-image trypsin in mirror-image biology systems and related applications.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Mirror-image trypsin, the preparation method and use thereof

[0002] CROSS-REFERNECE TO RELATED APPLICATIONS

[0003] This application claims priority from U.S. provisional application number 63 / 546,881, filed on November 1, 2023, which is incorporated herein by reference.

[0004] REFERENCE TO SEQUENCE LISTING

[0005] The content of the electronically submitted sequence listing (Name: FPCH2416038 IP. Sequence Listing_ST26.xml, Size: ~11 KB, and Date of Creation: October 23, 2024) submitted in this application is incorporated herein by reference in its entirety.

[0006] FIELD OF THE INVENTION

[0007] The present invention relates to the field of mirror-image biology. In particular, the present invention relates to a novel mirror-image trypsin, the preparation method and the use of the same in sequencing D-peptides and / or D-proteins.

[0008] BACKGROUND TO THE INVENTION

[0009] Proteins composed of unnatural D-amino acids and the achiral amino acid glycine are mirror-image forms of their native L-protein counterparts. D-Proteins can facilitate structure determination of their native L-forms that are difficult to crystallize (racemic Xray crystallography); D-proteins can serve as the bait for library screening to ultimately yield pharmacologically superior D-peptide / D-protein therapeutics (mirror-image phage display); D-proteins can also be used as a powerful mechanistic tool for probing molecular events in biology, drug discovery, and immunology.

[0010] Mirror-image peptides and proteins composed of D-amino acids and the achiral glycine are widely investigated as potential therapeutic and enzymatic tools because of their exceptional biostability-. However, the development of mirror-image biology systems and related applications are hindered by the lack of effective methods to sequence mirrorimage (D-) proteins. While natural-chirality (L-) proteins can be sequenced by bottom-up liquid chromatography-tandem mass spectrometry (LC-MS / MS), the sequencing of long D-peptides and D-proteins with the same strategy requires digestion by a site-specific D- protease prior to mass analysis. This is because fragmentation is required with site-specific proteases such as trypsin and endopeptidase Lys-C to generate short peptides of typically 300-2000 m / z for tandem mass analysis--.

[0011] Through combining solid-phase peptide synthesis (SPPS)- and native chemical ligation (NCL)-, the total chemical synthesis of mirror-image versions of several enzymes has been realized, including the human immunodeficiency virus type 1 (HIV-1) protease-, 4-hydroxy-tetrahydrodipicolinate synthase (DapA)-, Bacillus amyloliquefaciens ribonuclease (barnase)-, African swine fever virus polymerase X (ASFV pol X)-, Sulfolobus solfataricus P2 DNA polymerase IV (Dpo4)— — , Pyrococcus furiosus (Pfu) DNA polymerase—, and bacteriophage T7 RNA polymerase—.

[0012] Trypsin is a serine protease enriched in the small intestine which cuts peptide chains mainly at the carboxyl side of the amino acid lysine or arginine. Trypsin is a commonly used tool enzyme used in biochemical research and protein / peptide sequencing. However, it is not reported yet in the art that a mirror-image trypsin has been chemically synthesized. Therefore, there is a need in the field of mirror-image biology systems for a mirror-image trypsin which digests D-peptides and / or D-proteins, thereby achieving the effective sequencing of mirror-image (D-) proteins / peptides. Such mirror-image trypsin would provide a powerful tool for the study in mirror-image biology systems and related applications.

[0013] SUMMARY OF THE INVENTION

[0014] The inventors of the present invention for the first time chemically synthesized a mirror-image version of the enzyme trypsin which is able to digest D-peptides and D- proteins, thereby facilitating sequencing long D-peptides and D-proteins.

[0015] Although trypsin is a relatively small enzyme of 223 amino acids (aa), the in vitro folding might be problematic since the protease digests itself——. Attempts to directly chemically synthesize trypsin have proven to be inefficient and thereby difficult to obtain in sufficient quantities. The inventors reasoned that trypsin autolysis can be avoided by chemically synthesizing and in vitro folding its zymogen form, trypsinogen—, which does not exhibit protease activity because of an inhibitory propeptide at its N-terminus until its cleavage during activation122-2. By chemically synthesizing trypsinogen first and then folding and activating the trypsinogen in vitro into its active form, a novel mirror-image trypsin is successfully synthesized with a desired yield.

[0016] In addition, the inventors further unexpectedly found that by incorporating isoacyl dipeptide into trypsinogen, at least one of the trypsinogen segments could be synthesized more easily, and the yield of such peptide segment is greatly improved. In particular, an O- acyl isopeptide bond is incorporated before at least one serine or threonine residue of at least one of the peptide segments instead of a normal A-acyl peptide bond. The incorporated O-acyl isopeptide bond could interrupt the original structure of the peptide, leading to an increased solubility which might accelerate the synthesis and the ligation of the peptide, thereby increasing the yield of trypsinogen. The O-acyl isopeptide bond is further converted into a normal A-acyl peptide bond through a pH-promoted O,N-acyl shift.

[0017] Therefore, in a first aspect, the present invention relates to a novel mirror-image trypsin which digests D-peptides and / or D-proteins.

[0018] In a second aspect, the present invention relates to a mirror-image trypsinogen which is capable of being activated into a mirror-image trypsin.

[0019] In a third aspect, the present invention relates to a method for chemically synthesizing a trypsin.

[0020] In a fourth aspect, the present invention relates to a trypsin which is prepared according to the method of the present invention.

[0021] In a fifth aspect, the present invention relates to a method for sequencing D-peptide and / or D-protein by use of the mirror-image trypsin.

[0022] In a sixth aspect, the present invention relates to a kit for sequencing D-peptide and / or D-protein which comprises the mirror-image trypsin.

[0023] In a seventh aspect, the present invention relates to a method for writing and reading information in a long D-peptide and / or D-protein by use of the mirror-image trypsin.

[0024] Other aspects, embodiments and advantages of the present invention are set forth in part in the description, which may be obvious from the description or may be learned from the practice of the invention.

[0025] DESCRIPTION OF DRAWINGS / FIGURES

[0026] The following drawings / figures constitute part of the present specification and are included to further demonstrate certain aspects of the present invention. The invention may be better understood by reference to one or more of these figures in combination with the detailed description of specific embodiments presented herein.

[0027] Figure 1 illustrates the synthetic natural-chirality (L-) and synthetic mirror-image (D-) trypsin, a, Structures of natural-chirality (L-) trypsin (Protein Data Bank ID: 5XWL) and mirror-image (D-) trypsin (model generated by digital reflection), b, Recombinant L-, synthetic L-, and synthetic D-trypsin analyzed by sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE) and stained by Coomassie brilliant blue. M, protein marker. The experiment was performed three times with similar results, c-e, Kinetic assays of the digestion of substrate peptide L-LYAARLYAVR by the recombinant L- (panel c) and synthetic L-trypsin (panel d), and the digestion of substrate peptide D-LYAARLYAVR (SEQ ID NO: 3) by the synthetic D-trypsin (panel e). Data are presented as best-fit curves, with individual data points (ri = 3) and 95% confidence intervals (Cis).

[0028] Figure 2 illustrates the mirror-image trypsin digestion and sequencing of mirrorimage ribosomal protein L25. a-f, Analytical HPLC chromatograms of the synthetic L-L25 before (panel a) and after (panel b) digestion by the recombinant L-trypsin, and before (panel c) and after (panel d) digestion by the synthetic L-trypsin, and of the synthetic D- L25 before (panel e) and after (panel f) digestion by the synthetic D-trypsin. g-i, Protein sequence coverages of the synthetic L-L25 digested by the recombinant L- (panel g) and synthetic L-trypsin (panel h), and of the synthetic D-L25 digested by the synthetic D- trypsin (panel i). Solid triangles, trypsin cleavage sites. Open triangles, prohibited trypsin cleavage sites with proline at the Pl' site. The experiments were performed twice with similar results.

[0029] Figure 3 illustrates the mirror-image trypsin digestion and sequencing of mirrorimage Dpo4. a-f, Analytical HPLC chromatograms of the recombinant L-Dpo4-5m before (panel a) and after (panel b) digestion by the recombinant L-trypsin, and before (panel c) and after (panel d) digestion by the synthetic L-trypsin, and of the synthetic D-Dpo4-5m before (panel e) and after (panel f) digestion by the synthetic D-trypsin. g-i, Protein sequence coverages of the recombinant L-Dpo4-5m digested by the recombinant L- (panel g) and synthetic L-trypsin (panel h), and of the synthetic D-Dpo4-5m digested by the synthetic D-trypsin (panel i). Solid triangles, trypsin cleavage sites. Open triangles, prohibited trypsin cleavage sites with proline at the Pl' site. The experiments were performed twice with similar results.

[0030] Figure 4 illustrates the writing and reading information in a long D-peptide. a, Design of an information-storing 50-aa D-peptide, chemically synthesized by SPPS and purified by RP-HPLC. Solid triangles, trypsin cleavage sites. Amino acids in bold, N-terminal index amino acids of the 10-aa D-peptide fragments, in the order of A, F, G, H, and L. Amino acids underlined, C-terminal arginines of the 10-aa D-peptide fragments, b, ESI-MS spectrum of the mirror-image trypsin-digested information-storing 50-aa D-peptide. c,d, Tandem mass spectrum (panel c) and de novo sequencing (panel d), with an example of the tandem mass spectrum of the 10-aa D-peptide fragment indexed by alanine (AHWAGGHGHR) shown, e, Sorting of the de novo sequencing results with the sums of ALC of the filtered potential 10-aa D-peptide sequences indexed by alanine. The five 10- aa D-peptide sequences (panel e and Fig. 9) were arranged according to the index amino acids and decoded into the original phrase: “Mirror-image biology” (panel a). The experiments were performed twice with similar results.

[0031] Figure 5 illustrates the design of the synthetic L- / D-trypsinogen. a, Amino acid sequence of the L- / D-trypsinogen (UniProt ID: P00761). The amino acid sequences shown in bold, underlined, italic, double underlined, or boxed, respectively, correspond to the peptide segments used in panel b. b, Synthetic route for the total chemical synthesis of the L- / D-trypsinogen.

[0032] Figure 6 illustrates the chiral specificity of trypsin digestion of substrate peptides. a,b, Analytical HPLC chromatograms of the chemically synthesized L-LYAARLYAVR (panel a) and D-LYAARLYAVR (panel b) before digestion. c,d, Analytical HPLC chromatograms of L-LYAARLYAVR (panel c) and D-LYAARLYAVR (panel d) digested by the recombinant L-trypsin. e,f, Analytical HPLC chromatograms of L-LYAARLYAVR (panel e) and D-LYAARLYAVR (panel f) digested by the synthetic L-trypsin. g,h, Analytical HPLC chromatograms of L-LYAARLYAVR (panel g) and of D- LYAARLYAVR (panel h) digested by the synthetic D-trypsin. The experiments were performed twice with similar results.

[0033] Figure 7 illustrates the chiral specificity of trypsin digestion of ribosomal protein L25 and Dpo4. a-c, Analytical RP-HPLC chromatograms of the synthetic D-L25 before (panel a) and after digestion by the recombinant L- (panel b) and synthetic L-trypsin (panel c). d, e, Analytical RP-HPLC chromatograms of the synthetic L-L25 before (panel d) and after digestion by the synthetic D-trypsin (panel e). f-h, Analytical RP-HPLC chromatograms of the synthetic D-Dpo4-5m before (panel f) and after digestion by the recombinant L- (panel g) and synthetic L-trypsin (panel h). i, j, Analytical RP-HPLC chromatograms of the recombinant L-Dpo4-5m before (panel i) and after digestion by the synthetic D-trypsin (panel j). The experiments were performed twice with similar results.

[0034] Figure 8 illustrates the cleavage site specificity of trypsin digestion of ribosomal protein L25 and Dpo4. a, Observed cleavage frequency at Pl site of L-L25 digested by the recombinant L- and synthetic L-trypsin, and of D-L25 by the synthetic D-trypsin. b, Observed cleavage frequency at Pl site of L-Dpo4-5m digested by the recombinant L- and synthetic L-trypsin , and of D-Dpo4-5m by the synthetic D-trypsin. c, Observed cleavage frequency of L-L25 digested by the recombinant L- and synthetic L-trypsin, and of D-L25 by the synthetic D-trypsin, at trypsin cleavage sites with lysine at Pl site and with or without proline at Pl' site, d, Observed cleavage frequency of L-Dpo4-5m digested by the recombinant L- and synthetic L-trypsin, and of D-Dpo4-5m by the synthetic D-trypsin, at trypsin cleavage sites with lysine at Pl site and with or without proline at Pl' site. Panel c and panel d are displayed on log scales. The experiments were performed twice with similar results.

[0035] Figure 9 illustrates the sorting of the de novo sequencing results, a-e, Sorting of the de novo sequencing results with the sum of ALC of each filtered potential 10-aa D-peptide sequence indexed by alanine (panel a, also shown in Fig. 4e), phenylalanine (panel b), glycine (panel c), histidine (panel d), and leucine (panel e). The experiments were performed twice with similar results.

[0036] Figure 10 illustrates the preparation of L-trypsinogen-1. a, Analytical HPLC chromatogram ( = 214 nm) of the purified L-trypsinogen-1, analyzed using the Ultimate XB-C4 300 A (5 pm, 4.6 x 250 mm) column with a gradient of 20-70% CFLCN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min. b, ESLMS spectrum of L-trypsinogen-1 purified by semi-preparative HPLC, with the MS data obtained across the entire UV- absorbing peak shown in panel b. The observed molecular mass is presented with ± s.d.. Asterisk, -115, aspartate deletion.

[0037] Figure 11 illustrates the preparation of L-trypsinogen-2. a, Analytical HPLC chromatogram ( = 214 nm) of the purified L-trypsinogen-2, analyzed using the Ultimate XB-C4 300 A (5 pm, 4.6 x 250 mm) column with a gradient of 20-70% CH3CN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min. b, ESLMS spectrum of L-trypsinogen-2 purified by semi-preparative HPLC, with the MS data obtained across the entire UV- absorbing peak shown in panel b. The observed molecular mass is presented with ± s.d..

[0038] Figure 12 illustrates the preparation of L-trypsinogen-3. a, Analytical HPLC chromatogram ( = 214 nm) of the purified L-trypsinogen-3, analyzed using the Ultimate XB-C4 300 A (5 pm, 4.6 x 250 mm) column with a gradient of 20-70% CH3CN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min. b, ESLMS spectrum of L-trypsinogen-3 purified by semi-preparative HPLC, with the MS data obtained across the entire UV- absorbing peak shown in panel b. The observed molecular mass is presented with ± s.d..

[0039] Figure 13 illustrates the preparation of L-trypsinogen-4. a, Analytical HPLC chromatogram ( = 214 nm) of the purified L-trypsinogen-4, analyzed using the Ultimate XB-C4 300 A (5 pm, 4.6 x 250 mm) column with a gradient of 20-70% CH3CN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min. b, ESLMS spectrum of L-trypsinogen-4 purified by semi-preparative HPLC, with the MS data obtained across the entire UV- absorbing peak shown in panel b. The observed molecular mass is presented with ± s.d..

[0040] Figure 14 illustrates the preparation of L-trypsinogen-5. a, Analytical HPLC chromatogram ( = 214 nm) of the purified L-trypsinogen-5, analyzed using the Ultimate XB-C4 300 A (5 pm, 4.6 x 250 mm) column with a gradient of 20-70% CH3CN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min. b, ESLMS spectrum of L-trypsinogen-5 purified by semi-preparative HPLC, with the MS data obtained across the entire UV- absorbing peak shown in panel b. The observed molecular mass is presented with ± s.d..

[0041] Figure 15 illustrates the preparation of L-trypsinogen-6. a, 190 mg L-trypsinogen-2 was dissolved in 9.0 ml acidified ligation buffer (6 M Gn HCl, 0.1 M NaH2PO4, pH 3.0). The mixture was cooled in ice-salt bath at -10 °C, and 900 pl 0.5 M NaNCh in acidified ligation buffer was added. The reaction mixture was kept in ice-salt bath under stirring for 25 min, after which 9 ml 0.2 M MPAA in 8 M Gn HCl, 0.1 M Na2HPO4, pH 5.5 was added. After the addition of 197 mg L-trypsinogen-3, the pH of the reaction mixture was adjusted to 6.49 with NaOH solution at room temperature. After 15 h, the pH was adjusted to 9.0 to promote the complete removal of the Tfa group. After 1 h, 270 mg MeONHi HC1 was added to carry out the conversion of Thz into cysteine. TCEP HC1 was added until the pH of the reaction mixture reached 4.0. After 3 h, the reaction mixture was purified by semipreparative HPLC using the Ultimate XB-C4 120 A (5 pm, 10 x 250 mm) columns with a gradient of 25-55% CH3CN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min, and 174 mg L-trypsinogen-6 was obtained with a yield of 49.9%. b, Analytical HPLC chromatogram ( = 214 nm) of the purified L-trypsinogen-6, analyzed using the Ultimate XB-C4 300 A (5 pm, 4.6 x 250 mm) column with a gradient of 20-70% CH3CN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min. c, ESI-MS spectrum of L-trypsinogen-6 purified by semi-preparative HPLC, with the MS data obtained across the entire UV- absorbing peak shown in panel b. The observed molecular mass is presented with ± s.d..

[0042] Figure 16 illustrates the preparation of L-trypsinogen-7. a, 135 mg L-trypsinogen-1 was dissolved in 4.5 ml acidified ligation buffer (6 M Gn HC1, 0.1 M NaJUPCL, pH 3.0). The mixture was cooled in ice-salt bath at -10 °C, and 450 pl 0.5 M NaNCh in acidified ligation buffer was added. The reaction mixture was kept in ice-salt bath under stirring for 25 min, after which 4.5 ml 0.2 M MPAA in 8 M Gn HCl, 0.1 M Na2HPO4, pH 5.5 was added. After the addition of 174 mg L-trypsinogen-6, the pH of the reaction mixture was adjusted to 6.52 with NaOH solution at room temperature. After 15 h, the reaction mixture was reduced by TCEP and purified by semi-preparative HPLC using the Ultimate XB-C4 120 A (5 pm, 10 x 250 mm) column with a gradient of 25-55% CH3CN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min, and 110 mg L-trypsinogen-7 was obtained with a yield of 38.9%. b, Analytical HPLC chromatogram ( = 214 nm) of the purified L- trypsinogen-7, analyzed using the Ultimate XB-C4 300 A (5 pm, 4.6 x 250 mm) column with a gradient of 20-70% CH3CN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min. c, ESLMS spectrum of L-trypsinogen-7 purified by semi-preparative HPLC, with the MS data obtained across the entire UV-absorbing peak shown in panel b. The observed molecular mass is presented with ± s.d.. Asterisk, -115, aspartate deletion.

[0043] Figure 17 illustrates the preparation of L-trypsinogen-8. a, 28.3 mg L-trypsinogen-7 was dissolved in 19 ml 200 mM TCEP solution (6.7 M Gn HCl, 0.1 M Na2HPO4, pH 7.0), containing 20 mM VA-044 and 40 mM reduced L-glutathione. The reaction mixture was under stirring overnight at 37 °C and purified by semi-preparative HPLC using the Ultimate XB-C4 120 A (5 pm, 10 x 250 mm) column with a gradient of 30-60% CH3CN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min, and 16.6 mg L-trypsinogen-8 was obtained with a yield of 59.0%. b, Analytical HPLC chromatogram ( = 214 nm) of the purified L- trypsinogen-8, analyzed using the Ultimate XB-C4 300 A (5 pm, 4.6 x 250 mm) column with a gradient of 20-70% CH3CN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min. c, ESLMS spectrum of L-trypsinogen-8 purified by semi-preparative HPLC, with the MS data obtained across the entire UV-absorbing peak shown in panel b. The observed molecular mass is presented with ± s.d.. Asterisk, -115, aspartate deletion.

[0044] Figure 18 illustrates the preparation of L-trypsinogen-9. a, 70.3 mg L-trypsinogen-8 (from three synthesis batches) was dissolved in 16.5 ml Acm deprotection buffer (8 M Gn HCl, 0.1 M Na2HPO4, 40 mM TCEP, pH 7.0). 60 mg PdCl2dissolved in 700 pl 8 M Gn HCl was added to the peptide solution, under stirring at 25 °C for 15 h, before 16.5 ml 2 M DTT in 8 M Gn HCl was added. The reaction mixture was under stirring for 30 min and purified by semi-preparative HPLC using the Ultimate XB-C4 120 A (5 pm, 10 x 250 mm) column with a gradient of 30-60% CH3CN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min, and 32.7 mg L-trypsinogen-9 was obtained with a yield of 47.3%. b, Analytical HPLC chromatogram ( = 214 nm) of the purified L-trypsinogen-9, analyzed using the Ultimate XB-C4 300 A (5 pm, 4.6 x 250 mm) column with a gradient of 20-70% CH3CN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min. c, ESLMS spectrum of L- trypsinogen-9 purified by semi-preparative HPLC, with the MS data obtained across the entire UV-absorbing peak shown in panel b. The observed molecular mass is presented with ± s.d.. Asterisk, -115, aspartate deletion.

[0045] Figure 19 illustrates the preparation of L-trypsinogen-10. a, 105 mg L-trypsinogen-4 was dissolved in 3.0 ml acidified ligation buffer (6 M Gn HCl, 0.1 M NaH2PO4, pH 3.0). The mixture was cooled in ice-salt bath at -10 °C, and 300 pl 0.5 M NaNO2in acidified ligation buffer was added. The reaction mixture was kept in ice-salt bath under stirring for 25 min, after which 3.0 ml 0.2 M MPAA in 8 M Gn HCl, 0.1 M Na2HPO4, pH 5.5 was added. After the addition of 84.0 mg L-trypsinogen-5, the pH of the reaction mixture was adjusted to 6.47 with NaOH solution at room temperature. After 2.5 h, the pH was adjusted to 9.0 to promote the complete removal of the Tfa group. After 1 h, 93 mg MeONH2HC1 was added to carry out the conversion of Thz into cysteine. TCEP HC1 was added until the pH of the reaction mixture reached 4.0. After 3 h, the reaction mixture was purified by semi -preparative HPLC using the Ultimate XB-C4 120 A (5 pm, 10 x 250 mm) columns with a gradient of 30-60% CH3CN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min, and 17.8 mg L-trypsinogen-10 was obtained with a yield of 9.7%. b, Analytical HPLC chromatogram ( = 214 nm) of the purified L-trypsinogen-10, analyzed using the Ultimate XB-C4 300 A (5 pm, 4.6 x 250 mm) column with a gradient of 20-70% CH3CN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min. c, ESLMS spectrum of L-trypsinogen-10 purified by semi-preparative HPLC, with the MS data obtained across the entire UV- absorbing peak shown in panel b. The observed molecular mass is presented with ± s.d..

[0046] Figure 20 illustrates the preparation of L-trypsinogen-11. a, 22.6 mg L-trypsinogen- 9 was dissolved in 300 pl acidified ligation buffer (6 M Gn HC1, 0.1 M NaFbPCh, pH 3.0). The mixture was cooled in ice-salt bath at -10 °C, and 30 pl 0.5 M NaNCh in acidified ligation buffer was added. The reaction mixture was kept in ice-salt bath under stirring for 25 min, after which 300 pl 0.2 M MPAA in 8 M Gn HCl, 0.1 M Na2HPO4, pH 5.5 was added. After the addition of 17.8 mg L-trypsinogen-10, the pH of the reaction mixture was adjusted to 6.49 with NaOH solution at room temperature. After 5 h, the reaction mixture was reduced by TCEP and purified by semi-preparative HPLC using the Ultimate XB-C4 120 A (5 pm, 10 x 250 mm) column with a gradient of 25-55% CH3CN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min, and 2.5 mg L-trypsinogen-11 was obtained with a yield of 6.8%. b, Analytical HPLC chromatogram ( = 214 nm) of the purified L- trypsinogen-11, analyzed using the Ultimate XB-C4 300 A (5 pm, 4.6 x 250 mm) column with a gradient of 20-70% CH3CN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min. c, HR-ESI-MS spectrum of L-trypsinogen-11 purified by semi -preparative HPLC, with the MS data obtained across the entire UV-absorbing peak shown in panel b. The observed molecular mass is presented with ± s.d.. Asterisk, -115, aspartate deletion, d, Deconvoluted HR-ESI-MS spectrum of L-trypsinogen-11. The purity was estimated to be -7.5% according to the original HR-ESI-MS spectrum without deconvolution. The synthesis was performed twice with similar results.

[0047] Figure 21 illustrates the preparation of D-trypsinogen-1. a, Analytical HPLC chromatogram ( = 214 nm) of the purified D-trypsinogen-1, analyzed using the Ultimate XB-C4 300 A (5 pm, 4.6 x 250 mm) column with a gradient of 20-70% CH3CN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min. b, ESLMS spectrum of D-trypsinogen-1 purified by semi-preparative HPLC, with the MS data obtained across the entire UV- absorbing peak shown in panel b. The observed molecular mass is presented with ± s.d.. Asterisk, -115, aspartate deletion.

[0048] Figure 22 illustrates the preparation of D-trypsinogen-2. a, Analytical HPLC chromatogram ( = 214 nm) of the purified D-trypsinogen-2, analyzed using the Ultimate XB-C4 300 A (5 pm, 4.6 x 250 mm) column with a gradient of 20-70% CH3CN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min. b, ESLMS spectrum of D-trypsinogen-2 purified by semi-preparative HPLC, with the MS data obtained across the entire UV- absorbing peak shown in panel b. The observed molecular mass is presented with ± s.d..

[0049] Figure 23 illustrates the preparation of D-trypsinogen-3. a, Analytical HPLC chromatogram ( = 214 nm) of the purified D-trypsinogen-3, analyzed using the Ultimate XB-C4 300 A (5 pm, 4.6 x 250 mm) column with a gradient of 20-70% CH3CN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min. b, ESI-MS spectrum of D-trypsinogen-3 purified by semi-preparative HPLC, with the MS data obtained across the entire UV- absorbing peak shown in panel b. The observed molecular mass is presented with ± s.d..

[0050] Figure 24 illustrates the preparation of D-trypsinogen-4. a, Analytical HPLC chromatogram ( = 214 nm) of the purified D-trypsinogen-4, analyzed using the Ultimate XB-C4 300 A (5 pm, 4.6 x 250 mm) column with a gradient of 20-70% CH3CN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min. b, ESLMS spectrum of D-trypsinogen-4 purified by semi-preparative HPLC, with the MS data obtained across the entire UV- absorbing peak shown in panel b. The observed molecular mass is presented with ± s.d..

[0051] Figure 25 illustrates the preparation of D-trypsinogen-5. a, Analytical HPLC chromatogram ( = 214 nm) of the purified D-trypsinogen-5, analyzed using the Ultimate XB-C4 300 A (5 pm, 4.6 x 250 mm) column with a gradient of 20-70% CH3CN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min. b, ESLMS spectrum of D-trypsinogen-5 purified by semi-preparative HPLC, with the MS data obtained across the entire UV- absorbing peak shown in panel b. The observed molecular mass is presented with ± s.d..

[0052] Figure 26 illustrates the preparation of D-trypsinogen-6. a, 146 mg D-trypsinogen-2 was dissolved in 7.0 ml acidified ligation buffer (6 M Gn HCl, 0.1 M NaH2PO4, pH 3.0). The mixture was cooled in ice-salt bath at -10 °C, and 700 pl 0.5 M NaNCh in acidified ligation buffer was added. The reaction mixture was kept in ice-salt bath under stirring for 25 min, after which 7 ml 0.2 M MPAA in 8 M Gn HCl, 0.1 M Na2HPO4, pH 5.5 was added. After the addition of 153 mg D-trypsinogen-3, the pH of the reaction mixture was adjusted to 6.51 with NaOH solution at room temperature. After 15 h, the pH was adjusted to 9.0 to promote the complete removal of the Tfa group. After 1 h, 210 mg MeONHi HC1 was added to carry out the conversion of Thz into cysteine. TCEP HC1 was added until the pH of the reaction mixture reached 4.0. After 3 h, the reaction mixture was purified by semipreparative HPLC using the Ultimate XB-C4 120 A (5 pm, 10 x 250 mm) columns with a gradient of 25-55% CH3CN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min, and 200 mg D-trypsinogen-6 was obtained with a yield of 73.8%. b, Analytical HPLC chromatogram (X = 214 nm) of the purified D-trypsinogen-6, analyzed using the Ultimate XB-C4 300 A (5 pm, 4.6 x 250 mm) column with a gradient of 20-70% CH3CN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min. c, ESI-MS spectrum of D-trypsinogen-6 purified by semi-preparative HPLC, with the MS data obtained across the entire UV- absorbing peak shown in panel b. The observed molecular mass is presented with ± s.d..

[0053] Figure 27 illustrates the preparation of D-trypsinogen-7. a, 50.8 mg D-trypsinogen-1 was dissolved in 2.0 ml acidified ligation buffer (6 M Gn HC1, 0.1 M NaJUPCh, pH 3.0). The mixture was cooled in ice-salt bath at -10 °C, and 200 pl 0.5 M NaNCh in acidified ligation buffer was added. The reaction mixture was kept in ice-salt bath under stirring for 25 min, after which 2.0 ml 0.2 M MPAA in 8 M Gn HCl, 0.1 M Na2HPO4, pH 5.5 was added. After the addition of 98.2 mg D-trypsinogen-6, the pH of the reaction mixture was adjusted to 6.49 with NaOH solution at room temperature. After 15 h, the reaction mixture was reduced by TCEP and purified by semi-preparative HPLC using the Ultimate XB-C4 120 A (5 pm, 10 x 250 mm) column with a gradient of 25-55% CH3CN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min, and 43.7 mg D-trypsinogen-7 was obtained with a yield of 33.3%. b, Analytical HPLC chromatogram ( = 214 nm) of the purified D- trypsinogen-7, analyzed using the Ultimate XB-C4 300 A (5 pm, 4.6 x 250 mm) column with a gradient of 20-70% CH3CN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min. c, ESLMS spectrum of D-trypsinogen-7 purified by semi-preparative HPLC, with the MS data obtained across the entire UV-absorbing peak shown in panel b. Asterisk, -115, aspartate deletion. The observed molecular mass is presented with ± s.d..

[0054] Figure 28 illustrates the preparation of D-trypsinogen-8. a, 37.3 mg D-trypsinogen-7 was dissolved in 25 ml 200 mM TCEP solution (6.7 M Gn HCl, 0.1 M Na2HPO4, pH 7.0), containing 20 mM VA-044 and 40 mM reduced L-glutathione. The reaction mixture was under stirring overnight at 37 °C and purified by semi-preparative HPLC using the Ultimate XB-C4 120 A (5 pm, 10 x 250 mm) column with a gradient of 30-60% CH3CN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min, and 21.7 mg D-trypsinogen-8 was obtained with a yield of 58.5%. b, Analytical HPLC chromatogram ( = 214 nm) of the purified D- trypsinogen-8, analyzed using the Ultimate XB-C4 300 A (5 pm, 4.6 x 250 mm) column with a gradient of 20-70% CH3CN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min. c, ESLMS spectrum of D-trypsinogen-8 purified by semi-preparative HPLC, with the MS data obtained across the entire UV-absorbing peak shown in panel b. The observed molecular mass is presented with ± s.d.. Asterisk, -115, aspartate deletion.

[0055] Figure 29 illustrates the preparation of D-trypsinogen-9. a, 37.1 mg D-trypsinogen-8 (from two synthesis batches) was dissolved in 8.5 ml Acm deprotection buffer (8 M Gn HCl, 0.1 M Na2HPO4, 40 mM TCEP, pH 7.0). 30 mg PdCl2dissolved in 350 pl 8 M Gn HCl was added to the peptide solution, under stirring at 25 °C for 15 h, before 8.5 ml 2 M DTT in 8 M Gn HCl was added. The reaction mixture was under stirring for 30 min and purified by semi -preparative HPLC using the Ultimate XB-C4 120 A (5 pm, 10 x 250 mm) column with a gradient of 30-60% CH3CN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min, and 17.6 mg D-trypsinogen-9 was obtained with a yield of 48.2%. b, Analytical HPLC chromatogram ( = 214 nm) of the purified D-trypsinogen-9, analyzed using the Ultimate XB-C4 300 A (5 pm, 4.6 x 250 mm) column with a gradient of 20-70% CH3CN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min. c, ESLMS spectrum of D-trypsinogen-9 purified by semi-preparative HPLC, with the MS data obtained across the entire UV- absorbing peak shown in panel b. The observed molecular mass is presented with ± s.d.. Asterisk, -115, aspartate deletion.

[0056] Figure 30 illustrates the preparation of D-trypsinogen-10. a, 92.2 mg D-trypsinogen- 4 was dissolved in 2.5 ml acidified ligation buffer (6 M Gn HCl, 0.1 M NaH2PO4, pH 3.0). The mixture was cooled in ice-salt bath at -10 °C, and 250 pl 0.5 M NaNO2in acidified ligation buffer was added. The reaction mixture was kept in ice-salt bath under stirring for 25 min, after which 2.5 ml 0.2 M MPAA in 8 M Gn HCl, 0.1 M Na2HPO4, pH 5.5 was added. After the addition of 69.4 mg D-trypsinogen-5, the pH of the reaction mixture was adjusted to 6.48 with NaOH solution at room temperature. After 2.5 h, the pH was adjusted to 9.0 to promote the complete removal of the Tfa group. After 1 h, 77 mg MeONH2HC1 was added to carry out the conversion of Thz into cysteine. TCEP HC1 was added until the pH of the reaction mixture reached 4.0. After 3 h, the reaction mixture was purified by semi -preparative HPLC using the Ultimate XB-C4 120 A (5 pm, 10 x 250 mm) columns with a gradient of 30-60% CH3CN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min, and 11.9 mg D-trypsinogen-10 was obtained with a yield of 7.8%. b, Analytical HPLC chromatogram ( = 214 nm) of the purified D-trypsinogen-10, analyzed using the Ultimate XB-C4 300 A (5 pm, 4.6 x 250 mm) column with a gradient of 20-70% CH3CN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min. c, ESLMS spectrum of D-trypsinogen-10 purified by semi-preparative HPLC, with the MS data obtained across the entire UV- absorbing peak shown in panel b. The observed molecular mass is presented with ± s.d..

[0057] Figure 31 illustrates the preparation of D-trypsinogen-11. a, 17.6 mg D-trypsinogen- 9 was dissolved in 300 pl acidified ligation buffer (6 M Gn HC1, 0.1 M NaFbPCh, pH 3.0). The mixture was cooled in ice-salt bath at -10 °C, and 30 pl 0.5 M NaNCh in acidified ligation buffer was added. The reaction mixture was kept in ice-salt bath under stirring for 25 min, after which 300 pl 0.2 M MPAA in 8 M Gn HCl, 0.1 M Na2HPO4, pH 5.5 was added. After the addition of 22.5 mg D-trypsinogen-10 (from two synthesis batches), the pH of the reaction mixture was adjusted to 6.53 with NaOH solution at room temperature. After 5 h, the reaction mixture was reduced by TCEP and purified by semi-preparative HPLC using the Ultimate XB-C4 120 A (5 pm, 10 x 250 mm) column with a gradient of 25-55% CH3CN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min, and 2.0 mg D- trypsinogen-11 was obtained with a yield of 5.9%. b, Analytical HPLC chromatogram ( = 214 nm) of the purified D-trypsinogen-11, analyzed using the Ultimate XB-C4 300 A (5 pm, 4.6 x 250 mm) column with a gradient of 20-70% CH3CN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min. c, HR-ESI-MS spectrum of D-trypsinogen-11 purified by semipreparative HPLC, with the MS data obtained across the entire UV-absorbing peak shown in panel b. The observed molecular mass is presented with ± s.d.. Asterisk, -115, aspartate deletion, d, Deconvoluted HR-ESI-MS spectrum of D-trypsinogen-11. The purity was estimated to be -7.5% according to the original HR-ESI-MS spectrum without deconvolution. The synthesis was performed once.

[0058] Figure 32 illustrates the direct LC-MS / MS analysis of the undigested informationstoring 50-aa D-peptide. a, Design of an information-storing 50-aa D-peptide, also shown in Fig. 4a. b, c, ESLMS spectrum of the undigested information-storing 50-aa D-peptide (panel b), with an example of the tandem mass spectra of the undigested D-peptide shown (panel c). No 50-aa sequence was present in the de novo sequencing results. The experiment was performed twice with similar results.

[0059] DETAILED DESCRIPTION OF THE INVENTION

[0060] Addressing the need in the field of mirror-image biology systems for a mirror-image trypsin, the present invention provides a novel mirror-image trypsin, a preparation method for chemically synthesizing trypsin, and the use of the mirror-image trypsin in mirrorimage biology. Unless otherwise defined, scientific and technical terms used in connection with the present disclosure shall have the meanings that are commonly understood by those of ordinary skill in the art. Further, unless otherwise required by context, singular terms shall include pluralities and plural terms shall include the singular. Generally, nomenclatures utilized in connection with, and techniques of molecular biology, biochemistry, and protein and oligo- or polynucleotide chemistry described herein are those well-known and commonly used in the art.

[0061] Alpha-amino acids - the basic building blocks of proteins - are chiral molecules that exist in two forms: L-enantiomer (‘L’ for levorotatory or left-handed) and D-enantiomer (‘D’ for dextrorotatory or right-handed), except that glycine is not a chiral molecule. The two non-superimposable forms of amino acid differing in handedness or chirality are mirror images of one another and have otherwise identical physical and chemical properties. The term "chirality" used according to the present invention refers to a property that a thing or an object in 3 dimensions might have, which means it is different from its reflection in the mirror. For example, the left hand and right hand both have chirality.

[0062] The term "mirror-image" used according to the present invention refers to a thing or an object that is “mirror-image” to the other. If one thing or object is mirror symmetrical to the other, that is to say, it looks like the same as the reflection of the other in the mirror, the two things or objects are called "mirror-image". For example, the left hand is mirrorimage to the right hand, or to say, the left hand is the “mirror-image right hand”.

[0063] The term "natural-chirality (L-) protein" used according to the present invention refers to a protein consisting of L-amino acids and / or glycine, which often exists in natural organisms or their products and thus called natural-chirality (L-) protein.

[0064] The term "mirror-image (D-) peptide" or "D-peptide" used according to the present invention refers to a peptide consisting of D-amino acids and / or glycine, which is mirrorimage to natural -chirality (L-) peptide and thus called mirror-image (D-) peptide or D- peptide.

[0065] The term "mirror-image (D-) protein" or "D-protein" used according to the present invention refers to a protein consisting of D-amino acids and / or glycine, which is mirrorimage to natural-chirality (L-) protein and thus called mirror-image (D-) protein or D- protein.

[0066] The terms "mirror-image (D-) peptide / D-peptide" and "mirror-image (D-) protein / D- protein" are used interchangeably herein to refer to a polymer of D-amino acids and / or glycine residue.

[0067] The term "peptide segment" used according to the present invention refers to a peptide that is corresponding to a part of a full-length protein / polypeptide. Peptide segments can be, such as, at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200 or more amino acids in length. Peptide segments can also be, such as, at most 300, 250, 200, 175, 150, 125, 100, 90, 80, 70, 60, 50, 40, 30, 20, 15, 14, 13, 12, 11, 10, or 5 amino acids in length.

[0068] The term "identity" used according to the present invention refers to a relationship between the sequences of two or more polypeptide / protein molecules as determined by aligning and comparing the sequences. "Percent identity / percentage identity" means the percentage of identical residues between the amino acids in the compared molecules and is calculated based on the size of the smallest of the molecules being compared. In calculating percent identity, the sequences being compared are aligned in a way that gives the largest match between the sequences. For these calculations, if there are gaps in alignments, the gaps must be addressed by a particular mathematical model or computer program which is also called an algorithm. The common algorithm used for calculating the percent identity includes, such as, Needleman-Wunsch algorithm and Smith-Waterman algorithm.

[0069] Trypsin is an enzyme produced by pancreas which is released as its zymogen form, trypsinogen, into the small intestine and activated to trypsin by enterokinase or autoactivation, and activated trypsin can further activate other trypsinogen. Trypsin is capable of cleaving the peptide or protein with the same chirality at C-terminus of amino acid Lys or Arg. According to the present invention, trypsin may be vertebrate trypsin, such as mammalian trypsin, for example, trypsin from human, chimpanzee, cynomolgus monkey, rhesus monkey, mouse, rat, cat, dog, rabbit, cow, sheep, goat, horse, pig, or camel, among others. The details of trypsin from various origins can be found, for example, at the website https: / / www.uniprot.org / uniprotkb?query=%28ec%3A3.4.21.4%29&facets=reviewed%3 Atrue.

[0070] According to the present invention, L25 is a protein in the large subunit of ribosome derived from E. coli. Dpo4 (Sulfolobus solfataricus P2 DNA polymerase IV) is a thermostable polymerase which can also synthesize DNA at 37°C. Its mismatch rate is between 8* 10'3to 3 * 10'4. It is a polymerase that can replace Taq for multi-cycle PCR reaction. In addition, its mutant version Y12S capable of catalyzing the polymerization of RNA. Its amino acid sequence length is within the reach of current chemical synthesis techniques.

[0071] Solid-phase peptide synthesis (SPPS) refers to a solid-phase synthesis technique commonly used for peptide synthesis. During SPPS, an N-protected amino acid is bound to a solid phase substrate, forming a covalent bond between the carbonyl group and the substrate, most often an amido or an ester bond. The a-amino group is then deprotected and reacted with the carbonyl group of the next N-protected amino acid, and after that, the solid phase bears a dipeptide. This cycle is repeated to form the desired peptide chain. After all reactions are complete, the synthesized peptide is cleaved from the substrate. A number of amino acids bear functional groups in the side chain which must be protected specifically from reacting with the incoming N-protected amino acids. The protecting groups for the a- amino groups mostly used in the peptide synthesis are 9-fluorenylmethyloxycarbonyl group (Fmoc) and t-butyloxycarbonyl (Boc).

[0072] Native chemical ligation (NCL) is an extension of the chemical ligation field, a concept for constructing a large polypeptide formed by the assembling of two or more unprotected peptides segments. Especially, NCL is a powerful ligation method for synthesizing native backbone proteins or modified proteins of small and moderate size. In native chemical ligation, the thiol group of an N-terminal cysteine residue of an unprotected peptide attacks the C-terminal thioester of a second unprotected peptide. This reversible transthioesterification step is chemoselective and regioselective and leads to form a thioester intermediate. This intermediate rearranges by an intramolecular S,N- acyl shift that results in the formation of a native amide (peptide) bond at the ligation site.

[0073] The term "O,N-acyl shift" used according to the present invention refers to the conversion of an ( -acyl ester bond into an A- acyl amide bond between two adjacent amino acids in a peptide. The phrase "pH-promoted O,N-acyl shift" used according to the present invention means that the O,N-acyl shift is achieved by adjusting pH of a solution, for example, to a pH > 7.

[0074] According to the present invention, the expression "writing and reading information" means that the information of a text or an image is encoded into a D-peptide and / or D- protein by designing the sequence of the D-peptide and / or D-protein according to a text encoding standard, and the encoded information is then decoded from the sequence of D- peptide and / or D-protein. According to the present invention, the text refers to a phrase or a sentence, among others, which comprises certain information. The text encoding standard used in the method according to the present invention includes, such as, the 128 American Standard Code for Information Interchange (ASCII) codes, UTF-8 codes, Unicode codes, or other known text encoding standards.

[0075] The term "American Standard Code for Information Interchange (ASCII) code" used according to the present invention refers to an incipient coding format using 1 Byte to code 128 common characters, including control characters, letters, numbers, and symbols like space, comma, and dot. The ASCII code is now the first section of many modern coding formats, such as UTF-8 and Unicode. Table 1 shows an example of conversion relevance between printable ASCII codes and amino acids.

[0076] Table 1. Conversion between printable ASCII codes and amino acids

[0077] Therefore, in a first aspect, the present disclosure provides a novel mirror-image trypsin which digests D-peptides and / or D-proteins.

[0078] In some embodiments, the mirror-image trypsin is derived from a mirror-image trypsinogen. In some embodiments, the mirror-image trypsinogen is human, chimpanzee, cynomolgus monkey, rhesus monkey, mouse, rat, cat, dog, rabbit, cow, sheep, goat, horse, pig, or camel mirror-image trypsinogen. In preferable embodiments, the mirror-image trypsinogen is a porcine mirror-image trypsinogen. In more preferable embodiments, the mirror-image trypsin is a porcine mirror-image trypsin. In some embodiments, the mirror- image trypsin comprises an amino acid sequence having at least 70% identity, such as at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, or at least 99% identity, to SEQ ID NO: 1. In some embodiments, the mirror-image trypsin comprises the amino acid sequence as set forth in SEQ ID NO: 1. In some embodiments, the mirror- image trypsin essentially consists of the amino acid sequence as set forth in SEQ ID NO: 1. In some embodiments, the mirror-image trypsin consists of the amino acid sequence as set forth in SEQ ID NO: 1.

[0079] The amino acid sequence of porcine trypsin is as set forth in SEQ ID NO: 1. porcine trypsin (SEQ ID NO: 1)

[0080] IVGGYTCAANSIPYQVSLNSGSHFCGGSLINSQWVVSAAHCYKSRIQVRLGE HNIDVLEGNEQFINAAKIITHPNFNGNTLDNDIMLIKLSSPATLNSRVATVSLPRSCA AAGTECLISGWGNTKSSGSSYPSLLQCLKAPVLSDSSCKSSYPGQITGNMICVGFLE GGKDSCQGDSGGPVVCNGQLQGIVSWGYGCAQKNKPGVYTKVCNYVNWIQQTIA AN

[0081] In a second aspect, the present disclosure provides a mirror-image trypsinogen which is capable of being activated into a mirror-image trypsin.

[0082] In some embodiments, the mirror-image trypsinogen is human, chimpanzee, cynomolgus monkey, rhesus monkey, mouse, rat, cat, dog, rabbit, cow, sheep, goat, horse, pig, or camel mirror-image trypsinogen. In preferable embodiments, the mirror-image trypsinogen is a porcine mirror-image trypsinogen. In some embodiments, the mirrorimage trypsinogen comprises an amino acid sequence having at least 70% identity, such as at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, or at least 99% identity, to SEQ ID NO: 2. In some embodiments, the mirror-image trypsinogen comprises the amino acid sequence as set forth in SEQ ID NO: 2. In some embodiments, the mirror-image trypsinogen essentially consists of the amino acid sequence as set forth in SEQ ID NO: 2. In some embodiments, the mirror-image trypsinogen consists of the amino acid sequence as set forth in SEQ ID NO: 2.

[0083] The amino acid sequence of porcine trypsinogen is as set forth in SEQ ID NO: 2. porcine trypsinogen (SEQ ID NO: 2)

[0084] FPTDDDDKIVGGYTCAANSIPYQVSLNSGSHFCGGSLINSQWVVSAAHCYKS RIQVRLGEHNIDVLEGNEQFINAAKIITHPNFNGNTLDNDIMLIKLSSPATLNSRVAT VSLPRSCAAAGTECLISGWGNTKSSGSSYPSLLQCLKAPVLSDSSCKSSYPGQITGN MICVGFLEGGKDSCQGDSGGPVVCNGQLQGIVSWGYGCAQKNKPGVYTKVCNYV NWIQQTIAAN

[0085] In a third aspect, the present disclosure provides a method for chemically synthesizing a trypsin, which comprises the steps of: a) preparing the peptide segments of a trypsinogen by SPPS; b) assembling the peptide segments into trypsinogen by NCL; and c) folding and activating the trypsinogen into active trypsin.

[0086] In some embodiments, in step a) of the method, the trypsinogen is divided into 4, 5, 6, 7, 8, 9, or 10 peptide segments, and then each peptide segment is synthesized respectively. In some embodiments, all these peptide segments are assembled into a whole trypsinogen by NCL. In preferable embodiments, the trypsinogen is divided into 5 peptide segments in step a) of the method. In more preferable embodiments, the peptide segments are corresponding to aa 1-46, aa 47-75, aa 76-116, aa 117-180, and aa 181-231 in SEQ ID NO: 2, respectively.

[0087] In some embodiments, the peptide segments range from 10 to 150 aa, from 10 to 140 aa, from 10 to 130 aa, from 10 to 120 aa, from 10 to 110 aa, from 10 to 100 aa, from 10 to 90 aa, from 10 to 80 aa, from 10 to 70 aa, from 10 to 60 aa, from 10 to 50 aa, from 20 to 150 aa, from 20 to 140 aa, from 20 to 130 aa, from 20 to 120 aa, from 20 to 110 aa, from 20 to 100 aa, from 20 to 90 aa, from 20 to 80 aa, from 20 to 70 aa, from 20 to 60 aa, from 20 to 50 aa, from 30 to 90 aa, from 30 to 80 aa, from 30 to 70 aa, from 30 to 60 aa, from 25 to 95 aa, from 25 to 85 aa, from 25 to 75 aa, from 25 to 65 aa, from 25 to 55 aa, for example, from 29 to 64 aa, in length. In preferable embodiments, each of the peptide segments comprises 46, 29, 41, 64, or 51 amino acids, respectively.

[0088] In some embodiments, during synthesizing the peptide segments in step a), an O-acyl isopeptide bond is incorporated before at least one serine or threonine residue of at least one of the peptide segments instead of a normal 7V-acyl peptide bond. In preferable embodiments, the O-acyl isopeptide bond is incorporated before at least one serine residue of at least one of the peptide segments instead of a normal 7V-acyl peptide bond. In more preferable embodiments, the O-acyl isopeptide bond is incorporated between valine and serine residue of the at least one of the peptide segments. In even more preferable embodiments, the O-acyl isopeptide bond is incorporated between the valine residue at position 199 and the serine residue at position 200 numbered according to SEQ ID NO: 2.

[0089] In some embodiments, the O-acyl isopeptide bond is converted into a normal 7V-acyl peptide bond through a pH-promoted O,N-acyl shift. In preferable embodiments, the pH- promoted O,N-acyl shift is performed by adjusting pH to pH > 7.

[0090] In some embodiments, in step c) of the method, the trypsinogen is auto-activated or activated by an enzyme into active trypsin. In preferable embodiments, the trypsinogen is auto-activated in a buffer. In more preferable embodiments, the buffer contains Ca2+. In more preferable embodiments, the buffer contains Ca2+in a concentration ranging from 1 mM to 100 mM, for example, from 1 mM to 90 mM, from 1 mM to 80 mM, from 1 mM to 70 mM, from 1 mM to 60 mM, from 1 mM to 50 mM, from 1 mM to 40 mM, from 1 mM to 30 mM, from 1 mM to 20 mM, from 1 mM to 10 mM, from 2 mM to 90 mM, from 2 mM to 80 mM, from 2 mM to 70 mM, from 2 mM to 60 mM, from 2 mM to 50 mM, from 2 mM to 40 mM, from 2 mM to 30 mM, from 2 mM to 20 mM, from 2 mM to 10 mM, from 3 mM to 20 mM, from 4 mM to 20 mM, from 5 mM to 20 mM, from 5 mM to 15 mM, from 5 mM to 10 mM. In even more preferable embodiments, the buffer contains Ca2+in a concentration of 5 mM or 10 mM. In preferable embodiments, the trypsinogen is activated by enterokinase, cathepsin B, activated trypsin, or other enzymes which are capable of activating trypsinogen. In more preferable embodiments, the trypsinogen is activated by enterokinase, or activated trypsin.

[0091] In some embodiments, the method further comprises purifying the peptide segments by reversed-phase high-performance liquid chromatography (RP-HPLC) before step b). In some embodiments, the method further comprises a step of performing metal-free radicalbased desulfurization to the trypsinogen of step b) to convert unprotected cysteine to alanine. In some embodiments, the NCL is step b) is hydrazide-based NCL.

[0092] In some embodiments, the trypsin is a natural-chirality trypsin or a mirror-image trypsin. In preferable embodiments, the trypsinogen is a natural-chirality or mirror-image version of porcine trypsinogen. In more preferable embodiments, the trypsinogen is a natural -chirality or mirror-image version of porcine trypsinogen having the amino acid sequence as set forth in SEQ ID NO: 2. In even more preferable embodiments, the trypsin chemically synthesized by the method of the present invention is a natural-chirality trypsin or a mirror-image trypsin having the amino acid sequence as set forth in SEQ ID NO: 1.

[0093] In a fourth aspect, the present disclosure provides a trypsin which is prepared according to the method of the third aspect of the present application. In some embodiments, the trypsin is a natural-chirality trypsin or a mirror-image trypsin. In preferable embodiments, the trypsin comprises, essentially consists of, or consists of the amino acid sequence as set forth in SEQ ID NO: 1.

[0094] In a fifth aspect, the present disclosure provides a method for sequencing D-peptide and / or D-protein, which comprises the steps of: a) digesting the D-peptide and / or D-protein with a mirror-image trypsin; and b) analyzing the sequence of the D-peptide and / or D- protein with LC-MS / MS.

[0095] In some embodiments, the mirror-image trypsin is a mirror-image trypsin chemically synthesized by the method of the present application. In some embodiments, the D-peptide and / or D-protein is denatured in a denature buffer prior to trypsin digestion. In preferable embodiments, method for sequencing D-peptide and / or D-protein further comprises a step of desalting the trypsin digested proteins after step a).

[0096] In some embodiments, the D-peptide and / or D-protein ranges from 10 to 2,000 aa, for example, from 10 to 1,900 aa, from 10 to 1,800 aa, from 10 to 1,700 aa, from 10 to 1,600 aa, from 10 to 1,500 aa, from 10 to 1,400 aa, from 10 to 1,300 aa, from 10 to 1,200 aa, from 10 to 1,100 aa, from 10 to 1,000 aa, from 10 to 900 aa, from 10 to 800 aa, from 10 to 700 aa, from 10 to 600 aa, from 10 to 500 aa, from 10 to 400 aa, from 10 to 300 aa, from 10 to 200 aa, or from 10 to 100 aa, in length.

[0097] In some embodiments, the D-peptide and / or D-protein is a D-peptide / protein comprising trypsin digestion site. In some embodiments, the D-peptide and / or D-protein is a mirror-image ribosomal protein L25. In some embodiments, the D-peptide and / or D- protein is a mirror-image Dpo4-5m. In some embodiments, the D-peptide and / or D-protein is a mirror-image Dpo4-5m-Y12S.

[0098] The amino acid sequence of ribosomal protein L25 is as set forth in SEQ ID NO: 4. The amino acid sequence of protein Dpo4-5m is as set forth in SEQ ID NO: 5. The amino acid sequence of protein Dpo4-5m-Y12S is as set forth in SEQ ID NO: 6. ribosomal protein L25 (SEQ ID NO: 4)

[0099] MFTINAEVRKEQGKGASRRLRAANKFPAIIYGGKEAPLAIELDHDKVMNMQ AKAEFYSEVLTIVVDGKEIKVKAQDVQRHPYKPKLQHIDFVRA

[0100] Dpo4-5m (SEQ ID NO: 5)

[0101] HHHHHHMIVLFVDFDYFYAQVEEVLNPSLKGKPVVVSVFSGRFEDSGAVAT ANYEARKFGVKAGIPIVEAKKILPNAVYLPMRKEVYQQVSCRIMNLLREYSEKIEIA SIDEAYLDISDKVRDYREAYALGLEIKNKILEKEKITVTVGISKNKVFAKIAADMAK PNGIK VIDDEE VKRLIRELDIADVPGIGNITAEKLKKLGINKLVDTLAIEFDKLKGMI GEAKAKYLISLARDEYNEPIRTRVRKSIGRIVTMKRNSRNLEEIKPYLFRAIEESYYK LDKRIPKAIHVVAVTEDLDIVSRGRTFPHGISKETAYAESVKLLQKILEEDERKIRRI GVRFSKFIEAIGLDKFFDT Dpo4-5m-Y12S (SEQ ID NO: 6)

[0102] HHHHHH(Nle)IVLFVDFDYFSAQVEEVLNPSLKGKPVVVSVFSGRFEDSGAVA TANYEARKFGVKAGIPIVEAKKILPNAVYLP(Nle)RKEVYQQVSCRI(Nle)NLLREYS EKIEIASIDEAYLDISDKVRDYREAYALGLEIKNKILEKEKITVTVGISKNKVFAKIA AD(Nle)AKPNGIKVIDDEEVKRLIRELDIADVPGIGNITAEKLKKLGINKLVDTLAIEF DKLKG(Nle)IGEAKAKYLISLARDEYNEPIRTRVRKSIGRIVT(Nle)KRNSRNLEEIKP YLFRAIEESYYKLDKRIPKAIHVVAVTEDLDIVSRGRTFPHGISKETAYAESVKLLQ KILEEDERKIRRIGVRFSKFIEAIGLDKFFDT

[0103] "Nle" represents norleucine residue.

[0104] In some embodiments, the method for sequencing D-peptide and / or D-protein of the present application is performed so as to distinguish between different mutants of D- peptide and / or D-protein, and the method further comprises a step of c) identifying each mutant of D-peptide and / or D-protein.

[0105] In preferable embodiments, the method further comprises a step of desalting the trypsin digested proteins prior to analyzing the sequence of the D-peptide and / or D-protein. In some embodiments, the mutants of D-peptide and / or D-protein differ from each other by at least one amino acid, at least two amino acids, at least three amino acids, at least four amino acids, at least five amino acids, at least six amino acids, at least seven amino acids, at least eight amino acids, at least nine amino acids, or at least ten amino acids.

[0106] In some embodiments, the mutants of D-peptide and / or D-protein differ from each other by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more amino acids.

[0107] In some embodiments, the method of the fifth aspect is performed to distinguish between mutants D-Dpo4-5m and D-Dpo4-5m-Y12S. The difference between the two different mutants of D-Dpo4 is shown in Table 2. Table 2. Distinguishing between the two different mutants of D-Dpo4

[0108] PSMs for distinguishing between the two mutants of D-Dpo4: D-Dpo4-5m and D- Dpo4-5m-Y12S. Y12S corresponds to the mutated 18th residue in the D-proteins with N- terminal D-Hise tags.

[0109] In a sixth aspect, the present disclosure provides a kit for sequencing D-peptide and / or D-protein which comprises a mirror-image trypsin. In some embodiments, the mirror-image trypsin is a mirror-image trypsin chemically synthesized by the method of the present application. In preferable embodiments, the mirror-image trypsin comprises, essentially consists of, or consists of the amino acid sequence as set forth in SEQ ID NO: 1.

[0110] In a seventh aspect, the present disclosure provides a method for writing and reading information in a D-peptide and / or D-protein, which comprises the steps of: a) designing a text and encoding the text information into a D-peptide and / or D-protein; b) chemically synthesizing the D-peptide and / or D-protein; c) sequencing the information-storing D- peptide and / or D-protein by the method for sequencing D-peptide and / or D-protein according to the present application; and d) decoding the sequence of the D-peptide and / or D-protein into the original text and obtaining the text information in the D-peptide and / or D-protein.

[0111] In some embodiments, the text includes a phrase or sentence. In some embodiments, the text information is encoded into and decoded from the D-peptide and / or D-protein according to a text encoding standard. In some embodiments, the text encoding standard includes ASCII codes, UTF-8 codes, Unicode codes, or other known text encoding standards. In preferable embodiments, the text information is encoded and decoded according to the ASCII codes.

[0112] In some embodiments, the D-peptide and / or D-protein comprises arginine or lysine to provide trypsin cleavage sites. In some embodiments, the D-peptide and / or D-protein has from 20 to 1,000 aa, for example, from 20 to 900 aa, from 20 to 800 aa, from 20 to 700 aa, from 20 to 600 aa, from 20 to 500 aa, from 20 to 400 aa, from 20 to 300 aa, from 20 to 200 aa, from 20 to 100 aa in length. In some embodiments, the D-peptide and / or D-protein is a 50-aa, 60-aa, 70-aa, 80-aa, 90-aa, or 100-aa D-peptide and / or D-protein. In some embodiments, the designed text comprises information consisting of at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, or at least 500 characters / codes.

[0113] All publications, patents and patent applications discussed in the specification are indicative of the level of those skilled in the art to which the invention pertains. All publications, patents and patent applications are herein incorporated by reference to the same extent as if each individual publication, patent or patent application was specifically and individually indicated to be incorporated by reference.

[0114] The above description of various illustrated embodiments of the disclosure is not intended to be exhaustive or to limit the scope to the precise form disclosed. While specific embodiments of examples are described herein for illustrative purposes, various equivalent modifications are possible within the scope of the disclosure, as those skilled in the relevant art will recognize. The teachings provided herein can be applied to other purposes, other than the examples described above. Numerous modifications and variations are possible in light of the above teachings and, therefore, are within the scope of the appended claims.

[0115] These and other changes may be made in light of the above detailed description. In general, in the following claims, the terms used should not be construed to limit the scope to the specific embodiments disclosed in the specification and the claims.

[0116] EXAMPLES

[0117] Materials and Methods

[0118] Materials

[0119] D-DNA oligos and the porcine trypsinogen gene for recombinant protein expression were ordered from Genewiz (Beijing, China). Tris base and guanidine hydrochloride (Gn HCl) were purchased from Amresco (PA, U.S.). PBS, L-cystine, and L-cysteine were purchased from Solarbio Life Sciences (Beijing, China). Urea, CaCL, and Amicon Ultra centrifugal filter (0.5 ml, 10,000 MWCO) were purchased from Sigma (MO, U.S.). TransStart FastPfu Fly DNA polymerase, pEASY-Uni seamless cloning and assembly kit, and ProteinRuler I were purchased from TransGen Biotech (Beijing, China). Sep-Pak tC18 cartridge (100 mg sorbent per cartridge) was purchased from Waters Corp (MA, U.S.). 2- Chlorotrityl chloride resin was purchased from Tianjin Nankai Hecheng Science & Technology (Tianjin, China). Wang ChemMatrix resin, Fmoc-L-amino acids, Fmoc-D- amino acids, Boc-L-Ser(Fmoc-L-Val)-OH, Boc-D-Ser(Fmoc-D-Val)-OH, and O-(6- chlorobenzotriazol-l-yl)-A,A, ,M-tetramethyluronium hexafluorophosphate (HCTU) were purchased from GL Biochem (Shanghai, China). A,A-dimethylformamide (DMF), A,A-diisopropylethylamine (DIEA), trifluoroacetic acid (TFA), thioanisole, triisopropyl silane (TIPS), 1,2-ethanedithiol (EDT), palladium chloride (PdCU), and 2,2'- azobis[2-(2-imidazolin-2-yl)propane]dihydrochloride (VA-044) were purchased from J&K Scientific (Beijing, China). 4-Mercaptophenylacetic acid (MPAA) was purchased from Alfa Aesar Chemicals (MA, U.S.). Piperidine, Na2HPO4 I2H2O, NaEEPO^ELO, and NaNO2 were purchased from Sinopharm Chemical Reagent (Shanghai, China). NaCl, NaOH, and HC1 were purchased from Sinopharm Chemical Reagent (Beijing, China). Tris(2-carboxyethyl)phosphine hydrochloride (TCEP HC1), 9-fluorenylmethyl carbazate (Fmoc-hydrazide), ethyl cyanoglyoxylate-2-oxime (Oxyma), A,7V- diisopropylcarbodiimide (DIC), and DL-l,4-dithiothreitol (DTT) were purchased from Adamas Reagent (Shanghai, China). Glutathione reduced (GSH) was purchased from Acros Organics (NJ, U.S.). Anhydrous ether was purchased from Beijing Tongguang Fine Chemicals (Beijing, China). Acetonitrile (HPLC grade) was purchase from J. T. Baker (NJ, U.S.).

[0120] Trypsinogen expression and purification

[0121] The porcine trypsinogen gene was amplified by PCR using the TransStart FastPfu Fly DNA polymerase and cloned into the pET-28c vector by the pEASY-Uni seamless cloning and assembly kit. The recombinant trypsinogen was expressed in E. coli strain BL21(DE3) in lysogeny broth (LB) medium. The induced cells were harvested and resuspended in PBS. The cell lysate was disrupted by sonication at 4 °C for 10 min, and the proteins were subsequently precipitated by centrifugation at 20,000 g at 4 °C for 40 min. The precipitate was further washed three times with PBS. The trypsinogen in inclusion bodies was solubilized in 8M Gn HCl, purified by HPLC, and lyophilized.

[0122] Fmoc-SPPS

[0123] All the peptides were synthesized by Fmoc-SPPS on the Liberty Blue automated microwave peptide synthesizer (CEM Corporation, NC, U.S.). Isoacyl dipeptide— was incorporated at positions Vall99-Ser200 in L- / D-trypsinogen-5. The peptides with C- terminal carboxylate such as L- / D-trypsinogen-5 were synthesized on Wang ChemMatrix resin (0.6 mmol / g) preloaded with the first C-terminal residue, and the other peptides were synthesized on Fmoc-hydrazine 2-chlorotrityl resin (0.53 mmol / g) to prepare peptide hydrazides—, the scale of which was typically 0.25 mmol. The first residue was manually attached to the Wang ChemMatrix resin by a double coupling method: in the first coupling reaction, the amino acid was coupled at 30 °C for 40 min with 1 mmol amino acid, 0.98 mmol HCTU, and 2 mmol DIEA dissolved in 4 ml DMF, after which the resin was washed with DMF. Without deprotection, the second coupling reaction was performed at 25 °C overnight with 1 mmol amino acid, 1 mmol Oxyma, and 1 mmol DIC dissolved in 4 ml DMF. All the resins were swelled in DMF for 5-10 min before use. The Fmoc groups of the resins and coupled amino acids were removed by treatment with 20% piperidine and 0.1 M Oxyma in DMF. The coupling of amino acids except Fmoc-Cys(Trt)-OH, Fmoc- Cys(Acm)-OH, and Fmoc-His(Trt)-OH was performed in 10 ml DMF with 1 mmol amino acid, 1 mmol Oxyma, and 2 mmol DIC at 85 °C for 3 min. The coupling of Fmoc-Cys(Trt)- OH, Fmoc-Cys(Acm)-OH, and Fmoc-His(Trt)-OH was performed at 50 °C for 10 min to reduce side reactions at higher temperatures. The coupling of trifluoroacetyl thiazolidine- 4-caboxylic acid-OH (Tfa-Thz-OH) was performed with Oxyma / DIC activation at room temperature overnight—. The synthesized peptides were cleaved from resin using H2O / thioanisole / TIPS / EDT / TFA (0.5 / 0.5 / 0.5 / 0.25 / 8.25, v / v) under agitation at 27 °C for 2.5 h. Most of the TFA in the mixture was removed by N2 blowing, and cold ether was added to precipitate the crude peptides. After centrifugation, the supernatant was discarded and the precipitate was washed twice with ether. The crude peptides were dissolved in CH3CN / H2O or 6 M Gn HCl, analyzed by HPLC and ESI-MS, and purified by semipreparative HPLC.

[0124] NCL

[0125] The peptide segments were assembled by hydrazide-based NCL with a convergent assembly strategy—. The C-terminal peptide hydrazide segment (4-6 mM) was dissolved in acidified ligation buffer (6 M Gn HCl, 0.1 M NaH2PO4, pH 3.0), with pH monitored by the FiveEasy Plus pH meter and InLab Micro pH electrode (Mettler Toledo, OH, U.S.). The mixture was cooled in ice-salt bath at -10 °C, after which 25 mM NaNO2 in acidified ligation buffer was added. The reaction mixture was kept in ice-salt bath under stirring for 25 min, after which 100 mM MPAA in 6 M Gn HCl, 0.1 M Na2HPO4, pH 5-6 was added. After the addition of the N-terminal cysteine peptide (to 2-3 mM final concentration for both the C-terminal and N-terminal peptide segments), the pH of the reaction mixture was adjusted to 6.5 at room temperature. After overnight reaction, 150 mM TCEP in ligation buffer (pH 7.0) was added to dilute the reaction mixture twice, under stirring at room temperature for 30 min. Next, the ligation product was analyzed by HPLC and ESLMS, and purified by semi -preparative HPLC. The preparations of L- / D-trypsinogen-10 and L- / D-trypsinogen-11 suffered from low ligation yields (Figures 19-20 and 30-31), likely because the solubility of L- / D-trypsinogen-5 and L- / D-trypsinogen-10 decreased after the O-acyl isopeptide bond being converted to the native A-acyl peptide bond.

[0126] Desulfurization

[0127] Metal-free radical-based desulfurization— was performed by dissolving the cysteine- containing peptides (3 mg / ml) in desulfurization buffer (6 M Gn HCl, 0.1 M Na2HPO4, 200 mM TCEP, 40 mM GSH, 20 mM VA-044, pH 6.8), under stirring at 37 °C overnight, and the desulfurization product was analyzed by HPLC and ESLMS, and purified by semipreparative HPLC.

[0128] Acetamidomethyl (Acm) deprotection Pd-assisted Acm deprotection— was performed by dissolving the Acm-protected peptides in Acm deprotection buffer (6 M Gn HCl, 0.1 M Na2HPO4, 40 mM TCEP, pH 7.0), after which 20 mM (final concentration) PdCU was added, under stirring at 25 °C overnight, with 50 mM (final concentration) DTT added to quench the reaction. The reaction mixture was under stirring for 1 h, analyzed by HPLC and ESI-MS, and purified by semi-preparative HPLC.

[0129] RP-HPLC and ESI-MS

[0130] All the RP-HPLC analysis and purification experiments were performed on the Shimadzu Prominence HPLC systems (Shimadzu Corp, Kyoto, Japan) with SPD-20A ultraviolet- visible detectors and LC-20AT solvent delivery units. The Ultimate XB-C4 120 A (5 pm, 21.2 x 250 mm, Welch Materials, Shanghai, China) and C18 120 A (5 pm, 21.2 x 250 mm) columns were used to purify the crude peptides at a flow rate of 8 ml / min. The Ultimate XB-C4 120 A (5 pm, 10 x 250 mm) column was used to separate the ligation products at a flow rate of 4 ml / min. The Ultimate XB-C4 300 A (5 pm, 4.6 x 250 mm) and Inertsil C4 150 A (5 pm, 4.6 x 250 mm, GL Sciences, Tokyo, Japan) columns were used to monitor the ligation reactions and analyze the purity of the ligation products and the trypsin-digested peptides or proteins at a flow rate of 1 ml / min. The molecular mass with standard deviation (s.d.) of each purified peptide segment and ligation product was characterized by ESI-MS on the Shimadzu LC / MS-2020 system (Shimadzu Corp, Kyoto, Japan). The final products (L-trypsinogen-11 and D-trypsinogen-11) were further characterized by high-resolution ESI-MS (HR-ESI-MS) on the Waters SYNAPT G2-Si HDMS Mass Spectrometer (Waters Corp, MA, U.S.), with the purities of the final products estimated by the MassLynx software (Waters Corp, MA, U.S.).

[0131] Folding and activation of trypsinogen in vitro

[0132] Lyophilized recombinant L-, synthetic L-, and synthetic D- versions of trypsinogen were dissolved in denaturation buffer (8 M urea, 20 mM DTT, and 20 mM Tris, pH 8.5) at room temperature for 4 h. Trypsinogen folding was performed by diluting the protein with 39x volumes of renaturation buffer (20 mM Tris, 2 M urea, 1 mM L-cystine, 3 mM L- cysteine, pH 7.5) under stirring at 4 °C for 12 h. After folding, the precipitates were removed by centrifugation at 12,000 g at 4 °C for 30 min, followed by ultrafiltration using the Amicon Ultra centrifugal filter (0.5 ml, 10,000 MWCO). The trypsinogen was autoactivated in activation buffer (40 mM Tris, 0.1 M NaCl, 10 mM CaCU, pH 8.0) at 37 C for 12 h. The autoactivated trypsin was aliquoted and stored at -20 °C.

[0133] Trypsin kinetic assay

[0134] Substrate peptides (L- / D-LYAARLYAVR) chemically synthesized by SPPS and purified by RP-HPLC were used for the trypsin kinetic assays. The substrate peptides were dissolved and diluted into activation buffer (40 mM Tris, 0.1 M NaCl, 10 mM CaCL, pH 8.0), digested at 37 °C for 10 min with a final trypsin concentration at 100 nM, and quenched by adding 1 x volume of 1% TFA in H2O. The trypsin-digested substrate peptides were analyzed by HPLC using the Inertsil C4 150 A (5 pm, 4.6 x 250 mm) column, with the peak areas of UV absorption at 214 nm used to calculate the concentrations of the substrate peptides and trypsin-digested products with a standard curve. The initial reaction rates were calculated using the concentrations of the trypsin-digested products divided by the reaction time, and least-square curve fitting was used to calculate the Kmand kcat in the Michaelis-Menten equation. Because the different activities may result in differences of trypsin-digested peptides and proteins, we carefully adjusted the concentrations of recombinant L-, synthetic L-, and synthetic D-trypsin to compensate for their different activities (enzyme to substrate ratio at 1 / 133, 1 / 80, 1 / 53, w / w, respectively).

[0135] Trypsin digestion of peptides and proteins

[0136] The synthetic L- and synthetic D- versions of E. coli ribosomal protein L25, the synthetic D-Dpo4-5m, and the synthetic D-Dpo4-5m-Y12S were synthesized in our previous work^52. The synthetic L- and synthetic D- versions of L25, and the synthetic D- peptide for information storage were dissolved and diluted into activation buffer (40 mM Tris, 0.1 M NaCl, 10 mM CaCL, pH 8.0) before trypsin digestion. The recombinant L- Dpo4-5m, synthetic D-Dpo4-5m, and synthetic D-Dpo4-5m-Y12S were denatured in a buffer containing 8 M urea and 5 mM DTT for 1 h, after which 12.5 mM iodoacetamide was added to alkylate the cysteines in dark at room temperature for 30 min. lodoacetamide was inactivated by exposure to room light for 15 min, and the protein solution was diluted with 5x volumes of activation buffer before trypsin digestion. The trypsin digestion was performed at 37 °C for 12 h, quenched by adding l x volume of 1% TFA in H2O, and analyzed by analytical HPLC using the Inertsil C4 150 A (5 pm, 4.6 x 250 mm) column with a gradient of 5-95% CH3CN (with 0.1% TFA) in H2O (with 0.1% TFA) over 30 min.

[0137] LC-MS / MS analysis of trypsin-digested peptides and proteins

[0138] The trypsin-digested peptides and proteins were desalted by the Sep-Pak tC18 cartridge, dried using a CV200 vacuum centrifugal concentrator (Beijing JM Technology, Beijing, China), and dissolved in 20 pl 0.1% TFA in H2O, of which 6 pl was used for LC- MS / MS analysis. The samples were separated using a fused silica capillary column (100 pm x 150 mm, packed in-house with ReproSil-Pur C18-AQ 1.9 pm resin, Dr. Maisch, Germany) with a gradient of 5-95% CH3CN (with 0.1% formic acid) in H2O (with 0.1% formic acid) over 60 min at a flow rate of 0.30 pl / min, directly interfaced with an Orbitrap Exploris 480 mass spectrometer (Thermo Fisher Scientific, MA, U.S.). For sequencing the ribosomal protein L25 and different mutants of Dpo4, the tandem mass spectra were searched by the Proteome Discoverer software (version 2.5, Thermo Fisher Scientific, MA, U.S.) against a protein database containing the proteome of E. coli (UniProt Taxonomy ID 83333, 4530 proteins), porcine trypsin, and different mutants of Dpo4 (the norleucines in Dpo4-5m-Y12S were replaced by leucines of the same mass), with the following settings: full trypsin specificity (for calculating protein sequence coverage) or no enzyme specificity (for calculating the ratio of non-specific cleavage), two missed cleavages allowed, oxidation of methionine (+ 16.00 Da) and deamination of asparagine and glutamine (+ 0.98 Da) set to variable modifications, carbamidomethyl of cysteine (+ 57.02 Da) set to static modification for sequencing the mutants of Dpo4, precursor ion mass tolerance set to 10 ppm, and fragment ion mass tolerance set to 0.02 Da. The false discovery rate (FDR) was set to 0.01 based on q-value calculated by Percolator in the Proteome Discoverer software. For the cleavage site-specificity analysis, the intensity of the precursor ion of each PSM was used to calculate the amino acid frequency at Pl site. For information storage in a long D-peptide, the PEAKS Studio software (version 8.5, BioInformatics Solutions, ON, Canada) was used for de novo sequencing, with the following settings: no enzyme specificity, oxidation of methionine (+ 16.00 Da) and deamination of asparagine and glutamine (+ 0.98 Da) set to variable modifications, precursor ion mass tolerance set to 10 ppm, and fragment ion mass tolerance set to 0.02 Da. The local confidence and ALC were calculated by the software with the de novo sequencing results.

[0139] Writing and reading information in a long D-peptide

[0140] To encode the 128 American Standard Code for Information Interchange (ASCII) codes, 12 different amino acids were used as two-letter codes (122> 128). Among the 20 proteinogenic amino acids, lysine and arginine were used as cleavage sites, proline at Pl' site prohibits trypsin cleavage, aspartate and glutamate at P2, P3, Pl', or P2' site inhibit trypsin cleavage—, cysteine forms disulfate bond, and isoleucine cannot be distinguished from leucine by mass spectrometry. Among the remaining 13 amino acids, the hydrophobic amino acid valine was not chosen in order to improve the solubility of the synthetic D- peptide. The remaining 12 amino acids (alanine, phenylalanine, glycine, histidine, leucine, methionine, asparagine, glutamine, serine, threonine, tryptophan, and tyrosine) were used for D-peptide information storage. The printable ASCII codes were converted to duodecimal numbers, and to the corresponding amino acids (Table 1). Arginines were inserted into the C-terminus of each fragment to provide trypsin cleavage sites, and index amino acids were inserted into the N-terminus of each fragment to provide the order of the digested peptide fragments. A 20-character phrase: “Mirror-image biology” was encoded into a 50-aa D-peptide, chemically synthesized by SPPS and purified by RP-HPLC, digestible by the synthetic D-trypsin into five 10-aa D-peptide fragments. The informationstoring 50-aa D-peptide was digested by the synthetic D-trypsin and the desalted digestion products were analyzed by ESI-MS, followed by LC-MS / MS. The de novo sequencing results were filtered by a length of 10 aa, an ALC above 80%, with the ESI-MS m / z error of ± 0.5. The sums of ALC of the potential 10-aa D-peptide sequences were sorted to determine their sequences, which were arranged according to the index amino acids and decoded into the original phrase: “Mirror-image biology”.

[0141] Example 1. Chemical synthesis and biochemical characterization of natural-chirality and mirror-image trypsin.

[0142] To perform their total chemical synthesis, the natural-chirality and mirror-image versions of the porcine trypsinogen were each divided into five peptide segments ranging from 29 to 64 aa in length (Figure 5). All of the peptide segments were prepared by 9- fluorenylmethoxycarbonyl (Fmoc)-SPPS, purified by reversed-phase high-performance liquid chromatography (RP-HPLC), and assembled by hydrazide-based NCL— , followed by metal-free radical-based desulfurization— to convert unprotected cysteine to alanine—. After the synthesis, ligation, purification, and lyophilization (Figures 10-31), the L- and D- versions of trypsinogen were obtained at milligram (mg) scales with the expected molecular mass of 24.4 kDa (Table 3). Meanwhile, the recombinant porcine trypsinogen was expressed and purified from Escherichia coli (E. coli). The recombinant L-, synthetic L-, and synthetic D- versions of trypsinogen were folded and autoactivated in vitro, and analyzed by sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE) (Fig. lb, Table 4).

[0143] Table 3. Synthetic natural-chirality (L-) and mirror-image (D-) trypsinogen Table 4. Folding and activation yields of trypsinogen

[0144] To biochemically characterize the autoactivated recombinant L-, synthetic L-, and synthetic D- versions of trypsin, the L- and D- versions of a short substrate peptide (LYAARLYAVR, SEQ ID NO: 3) were chemically synthesized by SPPS and purified by RP-HPLC (Fig. 6a-b), and RP-HPLC was used to analyze the trypsin-digested products—.

[0145] It was found that the synthetic L-peptide was digestible by the recombinant L- and synthetic L-trypsin, but not by the synthetic D-trypsin; whereas the synthetic D-peptide was digestible by the synthetic D-trypsin, but not by the recombinant L- and synthetic L-trypsin (Fig. 6c-h). The reciprocal chiral specificity was consistent with previous studies on the synthetic mirror-image HIV-1 protease-. The initial reaction rates were determined at different substrate concentrations, and fitted to the Michaelis-Menten equation, suggesting that the Michaelis constants (Km) of the recombinant L-, synthetic L-, and synthetic D- trypsin were similar, while the catalytic constants (&Cat) of the synthetic L- and synthetic D-trypsin were about half of that of the recombinant L-trypsin (Fig. Ic-e), likely owing to the lower purity of the synthetic proteases (Figures 20 and 31).

[0146] Example 2. Mirror-image trypsin digestion and sequencing of mirror-image ribosomal protein L25.

[0147] Next, the synthetic D-trypsin was applied to the bottom-up sequencing of small D- proteins. Ribosomal proteins are generally lysine- and arginine-rich and thus are suitable substrates for trypsin digestion. The enzymatic activities of the recombinant L-, synthetic L-, and synthetic D-trypsin on the synthetic L- and synthetic D-L25, an E. coli ribosomal protein24, were tested. It was found that the synthetic L-L25 was digestible by the recombinant L- and synthetic L-trypsin, but not by the synthetic D-trypsin; whereas the synthetic D-L25 was digestible by the synthetic D-trypsin, but not by the recombinant L- and synthetic L-trypsin (Fig. 2a-f, Fig. 7).

[0148] The trypsin-digested and desalted peptide fragments of L- and D-L25 were analyzed by LC-MS / MS. The peptide-spectrum matches (PSMs) (Fig. 10-31) resulted in a protein sequence coverage of 98.9% for L-L25 after digestion by the recombinant L- (Fig. 2g) and synthetic L-trypsin (Fig. 2h), and 98.9% for D-L25 after digestion by the synthetic D- trypsin (Fig. 2i), hence validating the D-protein sequence of D-L25. Meanwhile, the cleavage site specificity was confirmed with precursor intensity of tryptic and nontryptic peptides—, suggesting that the recombinant L-, synthetic L-, and synthetic D-trypsin displayed similar cleavage preferences at the C-terminus of lysine or arginine (Fig. 8a), with proline at the Pl' site prohibiting trypsin cleavage (Fig. 8c)——. These results suggested that the synthetic D-trypsin was equally effective and site-specific in digesting small D-proteins including the 94-aa D-L25 as the recombinant L- and the synthetic L- trypsin do for protein sequencing.

[0149] Example 3. Mirror-image trypsin digestion and sequencing of mirror-image Dpo4.

[0150] Next, the synthetic D-trypsin on the bottom-up sequencing of a larger D-protein, Dpo4-5m, a mutant version of Dpo4 that facilitated its total chemical synthesis—12’—, was tested. Because the 358-aa Dpo4-5m (with an N-terminal Hise tag) is much larger than the 94-aa ribosomal protein L25, denaturation of Dpo4-5m was performed prior to trypsin digestion (see Materials and Methods). It was found that the recombinant L-Dpo4-5m was digestible by the recombinant L- and synthetic L-trypsin, whereas the synthetic D-Dpo4- 5m was digestible by the synthetic D-trypsin (Fig. 3a-f).

[0151] The trypsin-digested and desalted peptide fragments of L- and D-Dpo4-5m were analyzed by LC-MS / MS. The PSMs (Fig. 10-31) resulted in a protein sequence coverage of 99.2% and 97.8% for L-Dpo4-5m after digestion by the recombinant L- (Fig. 3g) and synthetic L-trypsin (Fig. 3h), respectively, and 97.8% for D-Dpo4-5m after digestion by the synthetic D-trypsin (Fig. 3i), hence validating the D-protein sequence of D-Dpo4-5m. Similarly, the cleavage site specificity was confirmed with precursor intensity of tryptic and nontryptic peptides—, suggesting that the recombinant L-, synthetic L-, and synthetic D-trypsin displayed similar cleavage preferences at the C-terminus of lysine or arginine (Fig. 8b), with proline at the Pl' site prohibiting trypsin cleavage (Fig. 8d)— — . These results suggested that the synthetic D-trypsin was equally effective and site-specific in digesting large D-proteins including the 358-aa D-Dpo4-5m as the recombinant L- and the synthetic L-trypsin do for protein sequencing.

[0152] Example 4. Distinguishing between different mutants of mirror-image Dpo4.

[0153] Encouraged by the high protein sequence coverage of D-Dpo4-5m, the inventors sought to apply the mirror-image trypsin digestion and D-protein sequencing method to distinguishing between two different mutants of D-Dpo4: D-Dpo4-5m and D-Dpo4-5m- Y12S (a mutant version capable of polymerizing L-RNA, with all the 6 methionines replaced by norleucines and a tyrosine by serine compared with D-Dpo4-5m)— .

[0154] The trypsin-digested and desalted peptide fragments of the synthetic D-Dpo4-5m and D-Dpo4-5m-Y12S were analyzed by LC-MS / MS. To distinguish between the two different mutants of D-Dpo4, the inventors focused on the PSMs (Fig. 10-31) containing the aforementioned 7 mutated residues. PSMs of D-Dpo4-5m identified all the 6 methionines and the unmutated tyrosine, whereas PSMs of D-Dpo4-5m-Y12S identified all the 6 norleucines and the mutated serine (Table 2). These results suggested that the mirror-image trypsin digestion and D-protein sequencing method was capable of distinguishing different mutants of a D-protein.

[0155] Example 5. Writing and reading information in a long D-peptide.

[0156] While information storage in short L- and D-peptides with the lengths of 10-18 aa has been reported^1, the retrieval of information from D-peptides longer than ~20 aa remains challenging without a site-specific protease, limiting the amount of information to be stored in D-peptides. It is reasoned that using D-trypsin, information-storing long D- peptides can be digested into short peptides and sequenced by LC-MS / MS, enabling the storage of more information in long D-peptides. Here, a 20-character phrase “Mirror-image biology” was selected and encoded it into a 50-aa D-peptide (SEQ ID NO: 7), chemically synthesized by SPPS and purified by RP- HPLC, with 5 arginines to provide trypsin cleavage sites and 5 index amino acids to provide the order of the digested peptide fragments (Fig. 4a). The amino acid sequence of the “Mirror-image biology” 50-aa D-peptide is as shown in SEQ ID NO: 7.

[0157] “Mirror-image biology” 50-aa D-peptide (SEQ ID NO: 7)

[0158] AHWAGGHGHRFGMGHMGAGRGGSASANAWRHYAAYAGGMRLGTGMAN QSR

[0159] The information-storing 50-aa D-peptide was digested by the synthetic D-trypsin into five 10-aa D-peptide fragments, and the desalted digestion products were analyzed by electrospray ionization-mass spectrometry (ESI-MS) (Fig. 4b), followed by LC-MS / MS (Fig. 4c). The de novo sequencing results (Fig. 4d) were filtered (see Materials and Methods), and the sums of average local confidence (ALC) of the potential 10-aa D-peptide sequences were sorted to determine their sequences (Fig. 4e, Fig. 9), which were arranged according to the index amino acids and decoded into the original phrase: “Mirror-image biology” (Fig. 4a).

[0160] Therefore, using D-trypsin, information can be retrieved from long D-peptides, expanding the amount of information that can be stored in long D-peptides and securing the information with the D-trypsin as a key, despite their lack of ability to be amplified by mirror-image polymerase chain reaction (PCR) as the information-storing L-DNAs do—.

[0161] Discussion

[0162] In this work, the inventors realized mirror-image protein sequencing by chemically synthesizing the mirror-image version of a site-specific protease, trypsin, enabling the sequencing and distinguishing of long D-proteins and information storage in long D- peptides, echoing the first chemically synthesized D-enzyme, the mirror-image HIV-1 protease-. The mirror-image trypsin digestion and D-protein sequencing method may become a useful tool for the quality control of chemically synthesized D-peptide and D- protein drugs^2^. Being able to distinguishing different mutants of D-proteins, the mirrorimage trypsin digestion and D-protein sequencing method may also be applied to the identification of D-proteins with modified amino acids. Combined with de novo sequencing, mirror-image trypsin digestion may also become a helping hand in discovering D-peptide drugs when applied in conjunction with screening methods such as affinity selection-mass spectrometry (AS-MS) as well as in selecting small-molecule drugs with informationstoring D-peptide barcodes—. Moreover, the mirror-image versions of trypsin and other proteases can eliminate D-peptides and D-proteins after use as a containment strategy. In addition, the chemically synthesized L-trypsin, free from protease and other protein contamination from traditional protein expression systems——, may provide a contamination-free tool for sample preparation prior to LC-MS / MS analysis in proteomics.

[0163] The next key step in establishing the mirror-image central dogma of molecular biology is to realize mirror-image translation through synthesizing a mirror-image ribosome-’—’—’——. Since LC-MS / MS is widely used in the quantitative and semi- quantitative analyses of different protein components in protein and RNA-protein complexes44^ including the ribosome, the mirror-image trypsin digestion and D-protein sequencing method may become a useful tool in validating mirror-image ribosome assembly, as well as the D-protein products of mirror-image translation.

[0164] References

[0165] 1. Kent, S.B.H. Novel protein science enabled by total chemical synthesis. Protein Sci 28, 313-328 (2019).

[0166] 2. Aebersold, R. & Mann, M. Mass spectrometry-based proteomics. Nature 422, 198-207 (2003).

[0167] 3. Zhang, Y.Y., Fonslow, B.R., Shan, B., Baek, M.C. & Yates, J.R. Protein Analysis by Shotgun / Bottom-up Proteomics. Chemical Reviews 113, 2343-2394 (2013).

[0168] 4. Merrifield, R.B. Solid Phase Peptide Synthesis .1. Synthesis of a Tetrapeptide. Journal of the American Chemical Society 85, 2149-2154 (1963).

[0169] 5. Dawson, P.E., Muir, T.W., Clark-Lewis, I. & Kent, S.B. Synthesis of proteins by native chemical ligation. Science 266, 776-9 (1994).

[0170] 6. Milton, R., Milton, S. & Kent, S. Total chemical synthesis of a D-enzyme: the enantiomers of HIV-1 protease show reciprocal chiral substrate specificity. Science 256, 1445-1448 (1992).

[0171] 7. Weinstock, M.T., Jacobsen, M.T. & Kay, M.S. Synthesis and folding of a mirror-image enzyme reveals ambidextrous chaperone activity. Proc Natl Acad Sci U S A 111, 11679-84 (2014).

[0172] 8. Vinogradov, A. A., Evans, E.D. & Pentelute, B.L. Total synthesis and biochemical characterization of mirror image barnase. Chemical Science 6, 2997-3002 (2015).

[0173] 9. Wang, Z., Xu, W., Liu, L. & Zhu, T.F. A synthetic molecular system capable of mirror-image genetic replication and transcription. Nat Chem 8, 698-704 (2016).

[0174] 10. Jiang, W. et al. Mirror-image polymerase chain reaction. Cell Discov 3, 17037 (2017).

[0175] 11. Pech, A. et al. A thermostable d-polymerase for mirror-image PCR. Nucleic Acids Res 45, 3997-4005 (2017).

[0176] 12. Fan, C., Deng, Q. & Zhu, T.F. Bioorthogonal information storage in L-DNA with a high-fidelity mirror-image Pfu DNA polymerase. Nat Biotechnol 39, 1548-1555 (2021).

[0177] 13. Xu, Y. & Zhu, T.F. Mirror-image T7 transcription of chirally inverted ribosomal and functional RNAs. Science 378, 405-412 (2022).

[0178] 14. Vestling, M.M., Murphy, C.M. & Fenselau, C. Recognition of trypsin autolysis products by high-performance liquid chromatography and mass spectrometry. Anal Chem 62, 2391-4 (1990).

[0179] 15. Bunkenborg, J., Espadas, G. & Molina, H. Cutting edge proteomics: benchmarking of six commercial trypsins. J Proteome Res 12, 3631-41 (2013).

[0180] 16. Kassell, B. & Kay, J. Zymogens of proteolytic enzymes. Science 180, 1022-7 (1973).

[0181] 17. Sahin-Toth, M. Human cationic trypsinogen. Role of Asn-21 in zymogen activation and implications in hereditary pancreatitis. J Biol Chem 275, 22750-5 (2000).

[0182] 18. Kukor, Z., Toth, M. & Sahin-Toth, M. Human anionic trypsinogen: properties of autocatalytic activation and degradation and implications in pancreatic diseases. Eur J Biochem 270, 2047-58 (2003).

[0183] 19. Zhao, M., Wu, F. & Xu, P. Development of a rapid high-efficiency scalable process for acetylated Sus scrofa cationic trypsin production from Escherichia coli inclusion bodies. Protein Expr Purif 116, 120-6 (2015).

[0184] 20. Fang, G.M. et al. Protein chemical synthesis by ligation of peptide hydrazides. Angew Chem Int Ed Engl 50, 7645-9 (2011).

[0185] 21. Wan, Q. & Danishefsky, S.J. Free-radical-based, specific desulfurization of cysteine: a powerful advance in the synthesis of polypeptides and glycopolypeptides. Angew Chem Int Ed Engl 46, 9248-52 (2007).

[0186] 22. Yan, L.Z. & Dawson, P.E. Synthesis of peptides and proteins without cysteine residues by native chemical ligation combined with desulfurization. J Am Chem Soc 123, 526-33 (2001).

[0187] 23. Slechtova, T., Gilar, M., Kalikova, K. & Tesarova, E. Insight into Trypsin Miscleavage: Comparison of Kinetic Constants of Problematic Peptide Sequences. Anal Chem 87, 7636-43 (2015).

[0188] 24. Ling, J. J. et al. Mirror-Image 5S Ribonucleoprotein Complexes. Angew Chem Int Ed Engl 59, 3724-3731 (2020).

[0189] 25. Keil, B.i. Specificity of proteolysis, (Springer-Verlag, Berlin, Germany ; New York, New York, 1992).

[0190] 26. Yang, H. et al. Precision De Novo Peptide Sequencing Using Mirror Proteases of Ac-LysargiNase and Trypsin for Large-scale Proteomics. Mol Cell Proteomics 18, 773- 785 (2019).

[0191] 27. Xu, W. et al. Total chemical synthesis of a thermostable enzyme capable of polymerase chain reaction. Cell Discov 3, 17008 (2017).

[0192] 28. Wang, M. et al. Mirror-Image Gene Transcription and Reverse Transcription. Chem 5, 848-857 (2019).

[0193] 29. Ng, C.C.A. et al. Data storage using peptide sequences. Nat Commun 12, 4242 (2021).

[0194] 30. Zheng, J.S. et al. A mirror-image protein-based information barcoding and storage technology. Science Bulletin 66, 1542-1549 (2021).

[0195] 31. Rossler, S.L., Grob, N.M., Buchwald, S.L. & Pentelute, B.L. Abiotic peptides as carriers of information for the encoding of small-molecule library synthesis. Science 379, 939-945 (2023).

[0196] 32. Mandal, K. et al. Chemical synthesis and X-ray structure of a heterochiral {D- protein antagonist plus vascular endothelial growth factor} protein complex by racemic crystallography. Proc Natl Acad Sci U S A 109, 14779-84 (2012).

[0197] 33. Uppalapati, M. et al. A Potent d-Protein Antagonist of VEGF-A is Nonimmunogenic, Metabolically Stable, and Longer-Circulating in Vivo. ACS Chem Biol 11, 1058-65 (2016).

[0198] 34. Marinec, P.S. et al. A Non-immunogenic Bivalent d-Protein Potently Inhibits Retinal Vascularization and Tumor Growth. ACS Chem Biol 16, 548-556 (2021).

[0199] 35. Eckert, D.M., Malashkevich, V.N., Hong, L.H., Carr, P.A. & Kim, P.S. Inhibiting HIV-1 entry: discovery of D-peptide inhibitors that target the gp41 coiled-coil pocket. Cell 99, 103-15 (1999).

[0200] 36. Chang, H.N. et al. Blocking of the PD-1 / PD-L1 Interaction by a D-Peptide Antagonist for Cancer Immunotherapy. Angew Chem Int Ed Engl 54, 11760-4 (2015).

[0201] 37. Zuckermann, R.N., Kerr, J.M., Siani, M.A., Banville, S.C. & Santi, D.V. Identification of highest-affinity ligands by affinity selection from equimolar peptide mixtures generated by robotic synthesis. Proc Natl Acad Sci U S A 89, 4505-9 (1992).

[0202] 38. Maaty, W.S. & Weis, D.D. Label-Free, In-Solution Screening of Peptide Libraries for Binding to Protein Targets Using Hydrogen Exchange Mass Spectrometry. J Am Chem Soc 138, 1335-43 (2016).

[0203] 39. Quartararo, A. J. et al. Ultra-large chemical libraries for the discovery of high- affinity peptide binders. Nat Commun 11, 3183 (2020).

[0204] 40. Burkhart, J.M., Schumbrutzki, C., Wortelkamp, S., Sickmann, A. & Zahedi, R.P. Systematic and quantitative comparison of digest efficiency and specificity reveals the impact of trypsin quality on MS-based proteomics. J Proteomics 75, 1454-62 (2012).

[0205] 41. Peplow, M. A Conversation with Ting Zhu. ACS Cent Sci 4, 783-784 (2018).

[0206] 42. Chen, J., Chen, M. & Zhu, T.F. Translating protein enzymes without ami noacyl -tRN A synthetases. Chem 7, 786-798 (2021).

[0207] 43. Cravatt, B.F., Simon, G.M. & Yates, J.R., 3rd. The biological impact of massspectrometry-based proteomics. Nature 450, 991-1000 (2007).

[0208] 44. Wang, X. et al. Mass spectrometric characterization of the affinity-purified human 26S proteasome complex. Biochemistry 46, 3553-65 (2007).

[0209] 45. Ori, A. et al. Cell type-specific nuclear pores: a case in point for context- dependent stoichiometry of molecular machines. Mol Syst Biol 9, 648 (2013).

[0210] 46. Chen, S.S. & Williamson, J.R. Characterization of the ribosome biogenesis landscape in E. coli using quantitative mass spectrometry. J Mol Biol 425, 767-79 (2013).

[0211] 47. Coin, I. The depsipeptide method for solid-phase synthesis of difficult peptides. J Pept Sci 16, 223-30 (2010).

[0212] 48. Huang, Y.-C. et al. Facile synthesis of C-terminal peptide hydrazide and thioester of NY-ESO-1 (A39-A68) from an Fmoc-hydrazine 2-chlorotrityl chloride resin. Tetrahedron 70, 2951-2955 (2014).

[0213] 49. Huang, Y.C. et al. Synthesis of 1- and d-Ubiquitin by One-Pot Ligation and Metal-Free Desulfurization. Chemistry 22, 7623-8 (2016).

[0214] 50. Fang, G.M., Wang, J.X. & Liu, L. Convergent chemical synthesis of proteins by ligation of peptide hydrazides. Angew Chem Int Ed Engl 51, 10347-50 (2012).

[0215] 51. Wan, Q. & Danishefsky, S.J. Free-radical-based, specific desulfurization of cysteine: a powerful advance in the synthesis of polypeptides and glycopolypeptides. Angew Chem Int Ed Engl 46, 9248-52 (2007).

[0216] 52. Maity, S.K., Jbara, M., Laps, S. & Brik, A. Efficient Palladium-Assisted One- Pot Deprotection of (Acetamidomethyl)Cysteine Following Native Chemical Ligation and / or Desulfurization To Expedite Chemical Protein Synthesis. Angew Chem Int Ed Engl 55, 8108-12 (2016).

[0217] 53. Ling, J. J. et al. Mirror-Image 5S Ribonucleoprotein Complexes. Angew Chem Int Ed Engl 59, 3724-3731 (2020).

[0218] 54. Xu, W. et al. Total chemical synthesis of a thermostable enzyme capable of polymerase chain reaction. Cell Discov 3, 17008 (2017).

[0219] 55. Jiang, W. et al. Mirror-image polymerase chain reaction. Cell Discov 3, 17037 (2017).

[0220] 56. Fan, C., Deng, Q. & Zhu, T.F. Bioorthogonal information storage in L-DNA with a high-fidelity mirror-image Pfu DNA polymerase. Nat Biotechnol 39, 1548-1555 (2021).

[0221] 57. Wang, M. et al. Mirror-Image Gene Transcription and Reverse Transcription. Chem 5, 848-857 (2019).

[0222] 58. Slechtova, T., Gilar, M., Kalikova, K. & Tesarova, E. Insight into Trypsin Miscleavage: Comparison of Kinetic Constants of Problematic Peptide Sequences. Anal Chem 87, 7636-43 (2015).

[0223] The Examples of the present invention indicate various aspects of the disclosure and are given by way of illustration only. From the above discussion and these Examples, one skilled in the art can ascertain the essential characteristics of the aspects of the disclosure, and without departing from the spirit and scope thereof, can make various changes and modifications of them to adapt to various usages and conditions. Thus, various modifications in addition to those described herein will be apparent to those skilled in the art from the foregoing description. Such modifications are also intended to fall within the scope of the appended claims.

Claims

CLAIMS1. A mirror-image trypsin which digests D-peptides and / or D-proteins.

2. The mirror-image trypsin of claim 1, wherein the mirror-image trypsin is derived from a mirror-image trypsinogen; preferably, wherein the mirror-image trypsinogen is a mirror-image porcine trypsinogen.

3. The mirror-image trypsin of claim 1 or 2, wherein the mirror-image trypsin comprises an amino acid sequence having at least 70% identity, such as at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, or at least 99% identity, to SEQ ID NO: 1; preferably, wherein the mirror-image trypsin comprises, essentially consists of, or consists of the amino acid sequence as set forth in SEQ ID NO:1.

4. A mirror-image trypsinogen which is capable of being activated into a mirrorimage trypsin.

5. The mirror-image trypsinogen of claim 4, wherein the mirror-image trypsinogen is a mirror-image porcine trypsinogen.

6. The mirror-image trypsinogen of claim 4 or 5, wherein the mirror-image trypsinogen comprises an amino acid sequence having at least 70% identity, such as at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, or at least 99% identity, to SEQ ID NO: 2; preferably, wherein the mirror-image trypsinogen comprises, essentially consists of, or consists of the amino acid sequence as set forth in SEQ ID NO:2.

7. A method for chemically synthesizing a trypsin, which comprises the steps of: a) preparing the peptide segments of a trypsinogen by solid-phase peptide synthesis (SPPS); b) assembling the peptide segments into trypsinogen by native chemical ligation (NCL); and c) folding and activating the trypsinogen into active trypsin.

8. The method of claim 7, wherein in step a) the trypsinogen is divided into 4, 5, 6, 7, 8, 9, or 10 peptide segments, preferably, the trypsinogen is divided into 5 peptidesegments, and then each peptide segment is synthesized respectively.

9. The method of claim 8, wherein an ( -acyl isopeptide bond is incorporated before at least one serine or threonine residue of at least one of the peptide segments instead of a normal A-acyl peptide bond.

10. The method of claim 9, wherein the ( -acyl isopeptide bond is incorporated between valine and serine residue of the at least one of the peptide segments; preferably, wherein the ( -acyl isopeptide bond is incorporated between the valine residue at position 199 and the serine residue at position 200 numbered according to SEQ ID NO: 2.

11. The method of claim 9 or 10, wherein the ( -acyl isopeptide bond is converted into a normal A-acyl peptide bond through a pH-promoted O,N-acyl shift.

12. The method of claim 11, wherein the pH-promoted O,N-acyl shift is performed by adjusting pH to pH > 7.

13. The method of any one of claims 7-12, wherein in step c), the trypsinogen is auto-activated or activated by an enzyme into active trypsin.

14. The method of claim 13, wherein the trypsinogen is auto-activated in a buffer, preferably, wherein the buffer contains Ca2+.

15. The method of claim 13, wherein the trypsinogen is activated by enterokinase, cathepsin B, or activated trypsin.

16. The method of any one of claims 7-15, wherein the method further comprises purifying the peptide segments by reversed-phase high-performance liquid chromatography (RP-HPLC) before step b).

17. The method of any one of claims 7-16, wherein the method further comprises a step of performing metal-free radical -based desulfurization to the trypsinogen of step b) to convert unprotected cysteine to alanine.

18. The method of any one of claims 7-17, wherein the NCL is step b) is hydrazide-based NCL.

19. The method of any one of claims 7-18, wherein the trypsin is a natural-chirality trypsin or a mirror-image trypsin.

20. The method of any one of claims 7-19, wherein the trypsinogen is a naturalchirality or mirror-image version of porcine trypsinogen.

21. A trypsin which is prepared according to the method of any one of claims 7-22. The trypsin of claim 21, wherein the trypsin is a natural-chirality trypsin or a mirror-image trypsin.

23. A method for sequencing D-peptide and / or D-protein, which comprises the steps of: a) digesting the D-peptide and / or D-protein with the mirror-image trypsin of any one of claims 1-3; and b) analyzing the sequence of the D-peptide and / or D-protein with LC-MS / MS.

24. The method of claim 23, wherein the D-peptide and / or D-protein is denatured in a denature buffer prior to trypsin digestion.

25. The method of claim 23 or 24, wherein the D-peptide and / or D-protein ranges from 10 to 2,000 aa, for example, from 10 to 1,500 aa, from 10 to 1,000 aa, from 10 to 500 aa, from 10 to 200 aa, or from 10 to 100 aa, in length.

26. The method of any one of claims 23-25, wherein the method is performed so as to distinguish between different mutants of D-peptide and / or D-protein, and the method further comprises a step of c) identifying each mutant of D-peptide and / or D-protein.

27. The method of claim 26, wherein the mutants of D-peptide and / or D-protein differ from each other by at least one amino acid, at least two amino acids, at least three amino acids, at least four amino acids, at least five amino acids, at least six amino acids, at least seven amino acids, at least eight amino acids, at least nine amino acids, or at least ten amino acids; preferably, wherein the mutants of D-peptide and / or D-protein differ from each other by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more amino acids.

28. A kit for sequencing D-peptide and / or D-protein which comprises the mirrorimage trypsin of any one of claims 1-3.

29. A method for writing and reading information in a D-peptide and / or D-protein, which comprises the steps of: a) designing a text and encoding the text information into a D-peptide and / or D- protein; b) chemically synthesizing the D-peptide and / or D-protein; c) sequencing the information-storing D-peptide and / or D-protein by the method of any one of claims 23-27; and d) decoding the sequence of the D-peptide and / or D-protein into the original text andobtaining the text information in the D-peptide and / or D-protein.

30. The method of claim 29, wherein the text information is encoded and decoded according to a text encoding standard; preferably, wherein the text encoding standard includes the 128 American Standard Code for Information Interchange (ASCII) codes, UTF-8 codes, Unicode codes, or other known text encoding standards.

31. The method of claim 29 or 30, wherein the D-peptide and / or D-protein comprises arginine or lysine to provide trypsin cleavage sites.

32. The method of any one of claim 29-31, wherein the D-peptide and / or D-protein has from 20 to 1,000 aa, for example, from 20 to 900 aa, from 20 to 800 aa, from 20 to 700 aa, from 20 to 600 aa, from 20 to 500 aa, from 20 to 400 aa, from 20 to 300 aa, from 20 to 200 aa, from 20 to 100 aa in length; preferably, wherein the D-peptide and / or D-protein is a 50-aa, 60-aa, 70-aa, 80-aa, 90-aa, or 100-aa D-peptide and / or D-protein.