Chemical synthesis and use of large enantiomer proteins

By identifying and mutating ligation-inducible sequences in large proteins, the method allows for the chemical synthesis of proteins exceeding 400 amino acids, enhancing enantiomer biology systems and applications in data storage and protein crystal analysis.

JP2026062886APending Publication Date: 2026-04-10TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
TSINGHUA UNIVERSITY
Filing Date
2025-12-26
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing methods struggle to chemically synthesize large proteins exceeding approximately 400 amino acid residues due to limitations in peptide segment synthesis and ligation efficiency, limiting the development of enantiomer biology systems and their applications.

Method used

A method for total chemical synthesis of large proteins involves identifying ligation-inducible sequences in the amino acid sequence, introducing mutations to reduce hydrophobicity and ligation-inducible sites, and dividing the protein into segments that can be chemically synthesized and folded together, using specific amino acid substitutions to maintain protein function.

Benefits of technology

Enables the synthesis of large proteins with at least 1-10% activity of biologically produced proteins, facilitating applications in bioorthogonal molecular data storage, SELEX for aptamer development, and X-ray protein crystal structure analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026062886000001_ABST
    Figure 2026062886000001_ABST
Patent Text Reader

Abstract

This invention provides a method for chemically producing relatively large proteins. [Solution] A method comprising: identifying ligation-inducible sequences in the amino acid sequence of a protein; obtaining multiple ligation-inducible segments by parsing the amino acid sequence of the protein with the ligation-inducible sequences; chemically synthesizing each of the ligation-inducible segments if each of the ligation-inducible segments is chemically synthesizable; identifying a loss-of-structure section in the ligation-inducible segment if any one of the ligation-inducible segments is not chemically synthesizable; introducing a ligation-inducible sequence into the loss-of-structure section by substituting amino acids in the loss-of-structure section with ligation-inducible amino acid residues; parsing the amino acid sequence of the protein with the ligation-inducible sequences; and chemically synthesizing each of the ligation-inducible segments.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Related applications This application claims priority to U.S. Provisional Patent Application No. 63 / 061,844, filed on 6 August 2020, the contents of which are incorporated herein by reference in whole.

[0002] Description regarding sequence listings An ASCII file named 87597_ST25.txt, created on May 6, 2021, containing 180,286 bytes, which was filed concurrently with the filing of this application, is incorporated herein by reference.

[0003] In some embodiments, the present invention relates to biochemistry, and more specifically, to methods for the total chemical synthesis of large proteins and their enantiomers, and to the use thereof. [Background technology]

[0004] Proteins, composed entirely of non-natural D-amino acids and the achiral amino acid glycine, are enantiomeric images of their native L-protein counterparts. Recent advances in chemical protein synthesis have made it possible to independently and easily synthesize domain-sized enantiomeric D-proteins, allowing us to enter the "mirror world" and conduct protein research in ways previously unattainable. D-proteins can facilitate the structural determination of their native L-forms, which are difficult to crystallize (racemic X-ray crystallography), serve as bait for library screening, ultimately leading to pharmacologically superior D-peptide / D-protein therapeutics (enantiomeric phage display), and can also be used as a powerful mechanistic tool for exploring molecular events in biology, drug discovery, and immunology.

[0005] Since Pasteur painstakingly separated the left and right crystals of tartrates more than 160 years ago, scientists and the general public alike have been greatly intrigued by the idea that biomolecules possess only one dominant side. In recent years, numerous theoretical and experimental investigations have facilitated the description of models for how one enantiomer became dominant over the other, perhaps from what was prebiotic racemic world. Blackmond, DG ["The Origin of Biological Homochirality", Cold Spring Harb Perspect Biol., 2010, 2(5), a002147] emphasizes enantioenrichment mechanisms involving either chemical processes, physical processes, or a combination of both. One scientific driving force behind such endeavors is the desire to understand the origin of life, given that biomolecular homochirality is a mark of life. Other motivations stem from practical and applied scientific interests, such as the need for orthogonal biological tools that can provide naturally impermeable molecular systems for secure data storage.

[0006] At the forefront of nucleic acid synthesis, phosphoramidate chemistry has enabled the synthesis of oligonucleotides (oligonucleotides) of up to approximately 150 nt from DNA and approximately 70 nt from RNA. At the forefront of protein synthesis, the collaboration between solid-phase peptide synthesis (SPPS) and native chemical ligation (NCL) has provided a powerful method for the total chemical synthesis of various proteins (5, 14-20). Specifically, mirror-image gene replication and transcription systems based on the mirror-image version of 174aa African swine fever virus polymerase X (ASFV pol X) (5), followed by the more efficient and heat-stable 352aa Sulfolobus solfataricus P2 DNA polymerase IV (Dpo4) (17-19), have been realized, leading to the realization of mirror-image polymerase chain reactions (MI-PCR) and mirror-image gene transcription and reverse transcription (21). In detail, the mutant version D-Dpo4 enzymatically transcribes a full-length 5S rRNA at 120 nt, a feat that would otherwise be too long to be chemically synthesized (21).

[0007] Enantiomer proteins are powerful tools with broad applications in structural biology, peptide / protein drug design, and the study of biological process mechanisms. As chemical protein synthesis techniques become more robust and readily accessible to scientists from different disciplines, the immense potential of enantiomer proteins in chemical, biological, and biomedical research will be fully unleashed. Two feasible techniques, native chemical ligation and enantiomer phage display, are particularly attractive and will have a significant impact on the discovery of a new class of pharmacologically superior peptide and protein therapies for the treatment of various human diseases.

[0008] The review “Mirror image proteins” [Zhao, L. and Lu, W., Current Opinion in Chemical Biology, 2014, 22, pp. 56-61] examines recent advances in the application of mirror image proteins to structural biology, drug discovery, and immunology.

[0009] Hartrampf, N. et al. ["Synthesis of proteins by automated flow chemistry", Science, 2020, 368(6494), pp. 980-987] reports an extremely efficient chemistry compatible with automated fast-flow equipment for directly producing 164-amino acid peptide chains in 327 consecutive reactions, demonstrating that peptide chain elongation is completed in a few hours, as shown here by the chemical synthesis of nine different protein chains corresponding to enzymes, structural units, and regulatory factors. The researchers report that, after purification and folding, this synthetic material exhibits biophysical and enzymatic properties comparable to biologically expressed proteins, demonstrating that high-fidelity automated flow chemistry, i.e., automated fast-flow peptide synthesis (AFPS), is an alternative technique for producing single-domain proteins without the use of ribosomes.

[0010] However, enantiomer proteins remain relatively small. On the other hand, synthesizing large proteins exceeding approximately 400 amino acid (aa) residues is far more difficult, mainly due to limitations in peptide segment synthesis and ligation efficiency. While recently developed automated fast-flow peptide synthesis (AFPS) technology can produce peptide chains more than three times longer than those previously achievable with conventional SPPS, there is no readily apparent methodology for synthesizing large enantiomer molecules. Consequently, the development of enantiomer biology systems and their applications in information storage, among other things, are severely limited. [Overview of the Initiative]

[0011] Aspects of the present invention relate to a method for the total chemical synthesis of relatively large (longer than 400 aa) proteins using both L- and D-type dominance of amino acid residues, and its application to D-amino acid proteins prepared by the method disclosed herein. According to embodiments of the present invention, large proteins are chemically synthesized without the involvement or presence of biochemical macromolecules by searching for sections in an amino acid sequence in which amino acid residues can be substituted (mutated) without adversely affecting the function of the protein, based on multiple sequence alignment and / or structural information. According to the invention disclosed herein, the introduction of mutations into the protein sequence reduces the cost of preparing D-amino acid proteins by inserting division sites and / or ligation sites into the protein sequence, as well as reducing the hydrophobicity of the ligation-inducible polypeptide and reducing the number of Ile residues in the protein. Also provided are, without limitation, uses of D-amino acid proteins such as bioorthogonal molecular data storage, SELEX for aptamer development, and crystal growth strategies in X-ray protein crystal structure analysis.

[0012] Thus, according to certain embodiments of the present invention, a method for chemically producing a protein is provided, which is carried out by linking at least two ligation-inducible segments of the protein, each of which is chemically synthesizable and i. Identify at least one ligation-inducible sequence in the amino acid sequence of a protein, and obtain multiple ligation-inducible segments by parsing the amino acid sequence of the protein with the ligation-inducible sequence, and ii. If each of the ligation-inducible segments is chemically synthesizable, then chemically synthesize each of the ligation-inducible segments. iii. If any one of the ligation-inducible segments is not chemically synthesizable, identify at least one structurally-lose section within the ligation-inducible segment, replace at least one amino acid in the structurally-lose section with a ligation-inducible amino acid residue to introduce a ligation-inducible sequence into the structurally-lose section, parse the amino acid sequence of the protein with the ligation-inducible sequence, and then chemically synthesize each of the ligation-inducible segments. It is possible to obtain it by doing so.

[0013] In some embodiments of the present invention, in step (i), at least one of the ligation-inducible sequences is located in a loss-of-structure section of the protein.

[0014] In some embodiments of the present invention, the method provided herein includes step (iii).

[0015] In some embodiments of the present invention, the method provided herein includes, before step (i), a) Dividing the amino acid sequence of the protein into at least two domain-forming segments, b) If each of the domain-forming segments is chemically synthesizable, chemically synthesize each of the domain-forming segments, and c) Fold the domain-forming segments together to obtain the protein. It also includes.

[0016] In some embodiments of the present invention, the method provided herein includes the step (a) of dividing the amino acid sequence of a protein into at least two domain-forming segments.

[0017] According to some embodiments of the present invention, if one of the domain-forming segments is not chemically synthesizable, the method further... d) Identifying at least one ligation-inducing sequence in the domain-forming segment and parsing the amino acid sequence of the domain-forming segment with the ligation-inducing sequence to obtain a plurality of chemically synthesizable ligation-inducing segments, e) If the domain-forming segment essentially lacks a ligation-inducing sequence or any one of the ligation-inducing segments is not chemically synthesizable, identifying at least one structure-losing section in the domain-forming segment or the ligation-inducing segment, f) Substituting at least one amino acid in the structure-losing section or the ligation-inducing segment with a ligation-inducing amino acid residue to introduce a ligation-inducing sequence into the structure-losing section or the ligation-inducing segment, and parsing the amino acid sequence of the domain-forming segment with the ligation-inducing sequence to obtain a plurality of sequences of chemically synthesizable ligation-inducing segments, g) Chemically synthesizing each of the chemically synthesizable ligation-inducing segments which is carried out by.

[0018] In some embodiments of the present invention, the method provided in the present application includes step (f).

[0019] According to some embodiments of the present invention, the synthetic protein exhibits at least 1%, 5% or at least 10% of the activity of the corresponding biologically produced protein.

[0020] According to some embodiments of the present invention, the activity is selected from the group consisting of catalytic activity, specific binding activity and structural activity.

[0021] According to some embodiments of the present invention, the protein comprises at least 240 amino acid residues.

[0022] According to some embodiments of the present invention, the protein comprises at least about 400 amino acid residues.

[0023] According to some embodiments of the present invention, the method provided herein further comprises substituting at least one hydrophobic amino acid residue in at least one ligation-inducible segment with a less hydrophobic amino acid in the following order of hydrophobicity: Ile>Leu>Phe>Val>Met>Pro>Trp>His(0)>Thr>Glu(0)>Gln>Cys>Tyr>Ala>Ser>Asn>Asp(0)>Arg+>Gly>His+>Glu>Lys+>Asp-.

[0024] According to some embodiments of the present invention, the synthetic protein is produced using at least 90% D-amino acid residues other than glycan.

[0025] According to some embodiments of the present invention, the protein has a three-dimensional structure that is essentially a mirror image of the three-dimensional structure of the corresponding biologically produced protein.

[0026] According to some embodiments of the present invention, the method provided herein further comprises substituting at least one Ile residue with a D-amino acid residue selected from the group consisting of D-Ala residues, D-Val residues, D-Leu residues, D-Thr residues, D-Phe residues, D-Met residues, Gly residues, and D-Pro residues.

[0027] According to another aspect of some embodiments of the present invention, a protein prepared by the method provided herein is provided, which has a length of at least about 240 amino acid residues.

[0028] According to some embodiments of the present invention, the chemically synthesized protein provided herein comprises at least two domain-forming segments, which are non-covalently attached polypeptide chains, wherein the domain-forming segments are covalently attached polypeptide chains in at least one corresponding biologically produced protein.

[0029] According to some embodiments of the present invention, the proteins provided herein are selected from the group consisting of enzymes, transport proteins, structural / mechanism proteins, hormones, signaling proteins, antibodies, fluid balance proteins, pH balance proteins, cell channels, and cell pumps.

[0030] According to some embodiments of the present invention, the protein is an enzyme capable of catalyzing a reaction catalyzed by a corresponding biologically produced enzyme.

[0031] According to some embodiments of the present invention, the chemically synthesized enzyme is an RNA polymerase capable of synthesizing RNA from ribonucleotides using a DNA template.

[0032] According to some embodiments of the present invention, the chemically synthesized RNA polymerase is T7 RNA polymerase or a Pfu DNA polymerase variant.

[0033] According to some embodiments of the present invention, a chemically synthesized Pfu DNA polymerase variant has at least one mutation selected from the group consisting of V93Q, E102A, D141A, E143A, Y410G, A486L, and E665K.

[0034] In some embodiments, the Pfu DNA polymerase further comprises at least one mutation (SEQ ID NO: 77) selected from the group consisting of D215A, A486Y, and L490W.

[0035] In some embodiments, Pfu DNA polymerase further comprises a DNA-binding structural domain, which is the sso7d structural domain (SEQ ID NO: 78).

[0036] According to some embodiments of the present invention, the chemically synthesized enzyme is a DNA polymerase capable of synthesizing DNA from deoxyribonucleotides.

[0037] According to some embodiments of the present invention, the chemically synthesized DNA polymerase is Pfu DNA polymerase.

[0038] According to another embodiment of the present invention, a method for chemically producing a D-amino acid protein (enantiomer protein) is provided, comprising ligating at least two ligation-inducible segments of a D-amino acid protein, wherein each ligation-inducible segment comprises at least 90% of D-amino acid residues other than glycan, and is chemically synthesizable, and i. Identify at least one ligation-inducible sequence in the amino acid sequence of the corresponding L-amino acid protein, and obtain multiple ligation-inducible segments by parsing the amino acid sequence with the ligation-inducible sequence, and ii. If each of the ligation-inducible segments is chemically synthesizable, then each of the ligation-inducible segments shall be chemically synthesized using at least 90% D-amino acid residues other than glycan. iii. If any one of the ligation-inducible segments is not chemically synthesizable, identify at least one loss-of-structure section in the ligation-inducible segment, replace at least one amino acid in the loss-of-structure section with a ligation-inducible amino acid residue to introduce a ligation-inducible sequence into the loss-of-structure section, parse the amino acid sequence of the ligation-inducible segment with the ligation-inducible sequence, and chemically synthesize each of the ligation-inducible segments using at least 90% non-glycemic D-amino acid residues. It is possible to obtain it by doing so.

[0039] According to some embodiments of the present invention, a method for producing a mirror image protein comprises, in step (i), that at least one ligation-inducible sequence is located in a loss-of-structure section in the corresponding L-amino acid protein.

[0040] According to some embodiments of the present invention, a method for producing enantiomer proteins includes step (iii).

[0041] According to some embodiments of the present invention, a method for producing enantiomer proteins is, before step (i), a) Dividing the amino acid sequence of an L-amino acid protein into at least two domain-forming segments, b) If each of the domain-forming segments is chemically synthesizable, then at least 90% of each of the domain-forming segments is chemically synthesized using D-amino acid residues other than glycan, c) Fold the domain-forming segments together to obtain a D-amino acid protein. It also includes.

[0042] According to some embodiments of the present invention, in a method for producing enantiomer proteins, if one of the domain-forming segments is not chemically synthesizable, d) Identify at least one ligation-inducible sequence in the domain-forming segment, and parse the amino acid sequence of the domain-forming segment with the ligation-inducible sequence to obtain multiple chemically synthesized ligation-inducible segments. e) If the domain-forming segment inherently lacks a ligation-inducible sequence or if one of the ligation-inducible segments is not chemically synthesizable, identify at least one structure-loss section in the domain-forming segment or ligation-inducible segment. f) Replace at least one amino acid in the loss-of-structure section or ligation-inducible segment with a ligation-inducible amino acid residue to introduce a ligation-inducible sequence into the loss-of-structure section or ligation-inducible segment, and parse the amino acid sequence of the domain-forming segment with the ligation-inducible sequence. g) Chemically synthesize each of the ligation-inducible segments using at least 90% D-amino acid residues other than Gly, thereby obtaining the domain-forming segments.

[0043] According to some embodiments of the present invention, a method for producing enantiomer proteins includes step (iii).

[0044] According to some embodiments of the present invention, in a method for producing enantiomer proteins, the D-amino acid protein exhibits at least 1%, at least 5%, or at least 10% of the activity of the corresponding L-amino acid protein.

[0045] According to some embodiments of the present invention, the activity of the enantiomer protein is selected from the group consisting of catalytic activity, specific binding activity, and structural activity.

[0046] According to some embodiments of the present invention, the D-amino acid protein provided herein comprises at least 240, 300, 400, or at least 500 amino acid residues.

[0047] According to some embodiments of the present invention, a method for producing a mirror image protein further comprises substituting at least one hydrophobic D-amino acid residue with a less hydrophobic amino acid in at least one ligation-inducible segment according to the following hydrophobic order: D-Ile>D-Leu>D-Phe>D-Val>D-Met>D-Pro>D-Trp>D-His(0)>D-Thr>D-Glu(0)>D-Gln>D-Cys>D-Tyr>D-Ala>D-Ser>D-Asn>D-Asp(0)>D-Arg+>Gly>D-His+>D-Glu>D-Lys+>D-Asp-.

[0048] According to some embodiments of the present invention, the D-amino acid protein exhibits a three-dimensional structure that is essentially a mirror image of the three-dimensional structure of the corresponding L-amino acid protein.

[0049] According to some embodiments of the present invention, a method for producing a mirror image protein further comprises substituting at least one Ile residue with a D-amino acid residue selected from the group consisting of D-Ala residues, D-Val residues, D-Leu residues, D-Thr residues, Gly residues, D-Phe residues, D-Met residues, and D-Pro residues.

[0050] According to another aspect of some embodiments of the present invention, a D-amino acid protein prepared by the method provided herein is provided.

[0051] In some embodiments of the present invention, the D-amino acid protein has a three-dimensional structure that is essentially a mirror image of the three-dimensional structure of the corresponding L-amino acid protein (e.g., the corresponding biologically produced protein).

[0052] According to some embodiments of the present invention, the D-amino acid protein comprises at least two domain-forming segments, which are non-covalently attached lipeptide chains, and these domain-forming segments are covalently attached polypeptide chains in at least one corresponding L-amino acid protein.

[0053] According to some embodiments of the present invention, D-amino acid proteins are selected from the group consisting of enzymes, transport proteins, structural / mechanism proteins, hormones, signaling proteins, antibodies, fluid balance proteins, pH balance proteins, cell channels, and cell pumps.

[0054] According to some embodiments of the present invention, the D-amino acid protein is a D-amino acid enzyme that can catalyze an enantiomer reaction compared to the corresponding L-amino acid enzyme, that is, it has catalytic activity comparable to the enzymatic reaction of the corresponding biologically produced enzyme, forming an enantiomorph of the corresponding product using an enantiomorph of the corresponding substrate.

[0055] According to some embodiments of the present invention, the D-amino acid enzyme is a D-amino acid RNA polymerase that can synthesize L-RNA from L-ribonucleotides using an L-DNA template.

[0056] According to some embodiments of the present invention, the D-amino acid RNA polymerase is a D-amino acid T7 RNA polymerase or a D-amino acid Pfu DNA polymerase variant.

[0057] According to some embodiments of the present invention, the D-amino acid Pfu DNA polymerase variant has at least one mutation selected from the group consisting of V93Q, E102A, D141A, E143A, Y410G, A486L, and E665K.

[0058] According to some embodiments of the present invention, the D-amino acid protein is a T7 RNA polymerase comprising at least one cleavage site, a first cleavage site between K363 and P364, and a second cleavage site between N601 and T602.

[0059] According to some embodiments of the present invention, the D-amino acid enzyme is a D-amino acid DNA polymerase capable of synthesizing L-DNA from L-deoxyribonucleotides.

[0060] According to some embodiments of the present invention, the D-amino acid DNA polymerase is a D-amino acid Pfu DNA polymerase.

[0061] According to another aspect of some embodiments of the present invention, a T7 RNA polymerase is provided comprising at least two polypeptide chains formed by fission between K363 and P364 and / or fission between N601 and T602.

[0062] In some embodiments, the T7 RNA polymerase provided herein further comprises at least one mutation selected from the group consisting of I6V, I14L, I74V, I82V, I109V, I117L, I141V, I210M, I244L, I281V, I320V, I322L, I330V, and I367L.

[0063] According to another embodiment of the present invention, a T7 RNA polymerase is provided having an amino acid sequence characterized by at least 80% or at least 90% sequence identity compared to SEQ ID NO: 83.

[0064] According to another aspect of some embodiments of the present invention, a Pfu DNA polymerase is provided comprising at least two polypeptide chains formed by fission between K467 and M468. These two polypeptide chains are not linked to each other by a covalent bond between their main chains.

[0065] In some embodiments, the Pfu DNA polymerase further comprises at least one mutation selected from the group consisting of E102A, E276A, K317G, V367L, and I540A.

[0066] In some embodiments, the Pfu DNA polymerases provided herein are I38F, I62V, I65V, I80V, I127V, I137M, I158L, I171A, I176V, I191V, I197V, I198V, I205V, I206V, I228V, I232L, I244M, I256V, I264A, I268L, I282V, I331 It further includes at least one mutation selected from the group consisting of A, I401V, I434V, I446F, I478K, I557V, I598V, I605T, I611V, I619A, I631L, I643V, I648T, I656V, I677T, I716Y, I734V, I745V, and I772P.

[0067] In some embodiments, the Pfu DNA polymerase further comprises at least one mutation selected from the group consisting of V93Q, D141A, E143A, Y410G, A486L, and E665K.

[0068] In some embodiments, Pfu DNA polymerase exhibits RNA polymerization activity.

[0069] In some embodiments, the Pfu DNA polymerase further comprises a mutation selected from the group consisting of D215A, A486Y, and / or L490W.

[0070] In some embodiments, Pfu DNA polymerase exhibits a deficiency in 3' to 5' exonuclease activity and increased dideoxynucleoside triphosphate (ddNTP) selectivity.

[0071] In some embodiments, Pfu DNA polymerase further comprises a DNA-binding structural domain, which is the sso7d structural domain (SEQ ID NO: 78).

[0072] In some embodiments, Pfu DNA polymerase modified with the sso7d structural domain exhibits improved PCR amplification activity.

[0073] According to another aspect of some embodiments of the present invention, a Pfu DNA polymerase is provided which has an amino acid sequence characterized by at least 80% or at least 90% sequence identity compared to SEQ ID NO: 51, or an amino acid sequence characterized by at least 80% or at least 90% sequence identity compared to SEQ ID NO: 79.

[0074] According to another aspect of some embodiments of the present invention, the use of a D-amino acid protein provided herein is provided, where the D-amino acid protein is an enzyme and is used to catalyze the synthesis of a product which is an enantiomorph of a molecule synthesized by a corresponding L-amino acid enzyme, or to catalyze the reaction of a substrate which is an enantiomorph of a corresponding substrate of a corresponding L-amino acid enzyme.

[0075] According to another aspect of some embodiments of the present invention, a process for enzymatically producing an L-polydeoxyribonucleic acid molecule is provided, comprising providing a D-amino acid DNA polymerase prepared by the method provided herein and capable of synthesizing L-DNA from L-deoxyribonucleotides, and a process carried out by reacting the D-amino acid DNA polymerase with a template L-DNA molecule, an L-DNA primer and a plurality of L-deoxyribonucleotides to enzymatically produce an L-DNA molecule.

[0076] In some embodiments of this process, the D-amino acid DNA polymerase is Pfu DNA polymerase.

[0077] In some embodiments of this process, the Pfu DNA polymerase is essentially as provided herein.

[0078] According to another aspect of some embodiments of the present invention, a process for enzymatically producing L-polyribonucleic acid (L-RNA) molecules is provided, comprising providing a D-amino acid RNA polymerase prepared by the method provided herein and capable of synthesizing L-RNA from L-ribonucleotides, and a process carried out by reacting the D-amino acid RNA polymerase with a template L-DNA molecule, an L-DNA / RNA primer and a plurality of L-ribonucleotides to enzymatically produce L-RNA molecules.

[0079] In some embodiments of this process, the D-amino acid RNA polymerase is a T7 RNA polymerase or a Pfu DNA polymerase variant, and the Pfu DNA polymerase variant has at least one mutation selected from the group consisting of V93Q, E102A, D141A, E143A, Y410G, A486L, and E665K.

[0080] In some embodiments of this process, the T7 RNA polymerase is essentially as provided herein.

[0081] According to another aspect of some embodiments of the present invention, a method is provided for forming a racemic crystal of a target molecule, which is carried out by cocrystallizing the target molecule and an enantiomorph of the target molecule, thereby forming a racemic crystal of an enantiomer pair, wherein the enantiomorph of the target molecule is a D-amino acid protein or a product of such a D-amino acid protein provided by the method presented herein.

[0082] According to another aspect of some embodiments of the present invention, a molecular probe comprising a D-amino acid protein as provided herein is provided, having a label portion attached thereto, and having affinity for an analyte that is an enantiomorph of the corresponding analyte of the corresponding L-amino acid protein.

[0083] According to another aspect of some embodiments of the present invention, a method for producing an L-nucleic acid aptamer or a D-peptide bonded moiety, To provide D-amino acid proteins prepared by the methods presented herein, and D-amino acid proteins are subjected to in vitro evolution to obtain L-nucleic acid aptamers or D-peptide bond moieties. A method for performing this is provided.

[0084] According to another aspect of some embodiments of the present invention, a method for amplifying a DNA sequence or RNA sequence is provided, comprising reacting a template of the DNA or RNA sequence with a DNA or RNA polymerase prepared by the method provided herein, wherein the reaction is achieved essentially without contamination of natural enzymes and / or natural DNA / RNA.

[0085] According to another aspect of some embodiments of the present invention, a method for sequencing L-DNA or L-RNA is provided using D-amino acid DNA or D-amino acid RNA polymerase as provided herein, phosphorothioate L-dNTP or phosphorothioate L-NTP, and two primers 5'-labeled with two different dyes.

[0086] According to another aspect of some embodiments of the present invention, a method for sequencing L-DNA is provided using D-amino acid DNA polymerase as provided herein, L-dideoxynucleoside triphosphate, and two primers 5'-labeled with two different dyes.

[0087] In some embodiments, the dyes are FAM and Cy5.

[0088] According to another aspect of some embodiments of the present invention, a data storage system, A molecule comprising at least one L-nucleic acid (e.g., L-DNA, L-RNA, and any chimeric molecule of these with D-nucleic acid segments) having a sequence that encodes information data, D-amino acid RNA polymerase and / or D-amino acid DNA polymerase for synthesizing and / or sequencing L-nucleic acids, wherein the D-amino acid RNA polymerase and / or D-amino acid DNA polymerase are produced by the method provided in this application. A data storage system including this is provided.

[0089] In some embodiments of this system, the L-nucleic acid molecules are prepared chemically or by enantiomer-catalyzed reactions. In some embodiments of the L-DNA data storage system, the L-DNA segments for information storage are prepared by enantiomer-assembly PCR using D-enzymes.

[0090] In some embodiments of this system, L-nucleic acid molecules are sequenced either chemically or by a sequencing-by-synthesis method using enantiomer enzymes.

[0091] In some embodiments of this system, the D-amino acid RNA polymerase is the T7 RNA polymerase provided in this application.

[0092] In some embodiments of this system, the D-amino acid DNA polymerase is the Pfu DNA polymerase provided in this application.

[0093] According to another aspect of some embodiments of the present invention, a chiral steganography method, A D-nucleic acid molecule having a sequence encoding cover information data, A combination of at least one L-nucleic acid molecule and / or a D- / L-chimeric nucleic acid molecule having a sequence that encodes an encryption key for deciphering Stego information data, D-amino acid RNA polymerase and / or D-amino acid DNA polymerase for synthesizing and / or sequencing L-DNA molecules, wherein the D-amino acid RNA polymerase and / or D-amino acid DNA polymerase are manufactured as provided herein. A chiral steganography technique is provided, which is performed by [this method].

[0094] In some embodiments, the L-nucleic acid molecule is either chemically prepared or prepared by an enantiomer-catalyzed reaction.

[0095] In some embodiments, L-nucleic acid molecules are sequenced chemically or by sequencing-by-synthesis methods using enantiomer enzymes.

[0096] In some embodiments, the D- / L-chimeric nucleic acid molecules are either chemically prepared or prepared by natural / enantiomer-catalyzed reactions.

[0097] In some embodiments, the L-DNA / RNA portion of a D- / L-chimeric nucleic acid molecule is sequenced either chemically or by a sequencing-by-synthesis method using enantiomer enzymes.

[0098] In some embodiments, the D-amino acid RNA polymerase is the T7 RNA polymerase as provided herein.

[0099] In some embodiments, the D-amino acid DNA polymerase is Pfu DNA polymerase as provided herein.

[0100] In some embodiments, the system can be combined with DNA cryptography to provide an additional layer of security using encrypted data.

[0101] According to another aspect of some embodiments of the present invention, a method for studying L-RNA hydrolysis, At least one L-RNA molecule having a higher-order structure and a long chain length sequence, D-amino acid RNA polymerase and / or D-amino acid DNA polymerase for synthesizing L-RNA molecules, which are produced by the method provided in this application. A method for performing this is provided.

[0102] According to another aspect of some embodiments of the present invention, a method for studying RNA degradation, At least one L-RNA molecule having a higher-order structure and a long chain length sequence, A D-amino acid RNA polymerase and / or D-amino acid DNA polymerase for synthesizing L-RNA molecules, wherein the D-amino acid RNA polymerase and / or D-amino acid DNA polymerase are produced by the method provided in this application. A method for performing this is provided.

[0103] In some embodiments, this method can be used to evaluate the effectiveness of RNase inhibitor reagents.

[0104] According to another aspect of some embodiments of the present invention, a transcription AND logic is provided, which is executed by a D-amino acid RNA polymerase, the D-amino acid RNA polymerase produced by the method provided herein.

[0105] In some embodiments, the D-amino acid RNA polymerase is the T7 RNA polymerase provided in this application.

[0106] In some embodiments, the D-amino acid RNA polymerase includes at least one cleavage site, a first cleavage site between K363 and P364, and a second cleavage site between N601 and T602.

[0107] In some embodiments, the D-amino acid RNA polymerase includes at least one fission site, the aforementioned site being located in the same loop, i.e., positions 357–366 and / or 564–607.

[0108] According to another aspect of some embodiments of the present invention, a method for producing an L-RNA marker / ladder, To provide a D-amino acid RNA polymerase prepared by the method provided herein, which can synthesize L-RNA from L-ribonucleotides, and The process involves reacting D-amino acid RNA polymerase with template L-DNA molecules of different lengths, L-DNA / RNA primers, and multiple L-ribonucleotides. A method is provided for enzymatically producing L-RNA molecules of different lengths, purifying them, and then mixing them together at a specific concentration.

[0109] In some embodiments, the D-amino acid RNA polymerase is essentially the T7 RNA polymerase provided herein.

[0110] Unless otherwise defined, all technical and / or scientific terms used herein have the same meaning as those generally understood by those skilled in the art to which the present invention relates. While similar or equivalent methods and materials may be used in the implementation or testing of embodiments of the present invention, exemplary methods and / or materials are described below. In case of any conflict, this specification shall prevail, including definitions. Furthermore, materials, methods, and examples are illustrative and not necessarily intended to be limiting.

[0111] Some embodiments of the present invention are described herein by reference only to the accompanying drawings, which are provided specifically as examples. While these drawings are given in detail here, it is emphasized that the details shown are illustrative and intended to illustrate embodiments of the present invention. In this regard, interpreting this description in conjunction with the drawings will make it clear to those skilled in the art how embodiments of the present invention can be carried out. [Brief explanation of the drawing]

[0112] [Figure 1] This is a flowchart illustrating the method provided in this application, according to some embodiments of the present invention. [Figure 2A]We present the design flow chart (Figure 2A) for the synthesis pathway of a mutant Pfu-N fragment in which additional NCL sites (E102A, E276A, K317G, V367L) are introduced to form a ligation-inducible segment and 25 isoleucine residues are substituted, and the design flow chart (Figure 2B) for the synthesis pathway of a mutant Pfu-C fragment in which an additional NCL site (I540A) is introduced, along with mutations in 15 other isoleucine residues. By introducing these mutations, the protein synthesis and ligation processes in SPPS are facilitated and the synthesis cost of mirror image versions is reduced. [Figure 2B] Same as above [Figure 3A] We present the design flow charts for the synthetic pathways of the substituted 369aa (including the His6 tag added to the N-terminus) mutant T7-fission-N fragment (Figure 3A), the 238aa mutant T7-fission-M fragment (Figure 3B), and the 282aa mutant T7-fission-C fragment (Figure 3C). The design flow charts include isoleucine residue substitutions, novel NCLs, and a novel fission site between K363 and P364. These mutations were introduced to facilitate protein synthesis and ligation processes in SPPS and to reduce the cost of synthesizing mirror image versions. [Figure 3B] Same as above [Figure 3C] Same as above [Figure 4] This flowchart illustrates molecular data storage according to some embodiments of the present invention, using L-DNA as an exemplary type of XNA. [Figure 5] A flowchart illustrating DNA-based steganography according to some embodiments of the present invention presents a seemingly ordinary D-DNA storage library into which chimeric D-DNA / L-DNA key molecules are embedded to carry a secret message. [Modes for carrying out the invention]

[0113] In some embodiments, the present invention relates to biochemistry, and more specifically, to methods for the total chemical synthesis of large proteins and their enantiomers, and their uses, but not limited to these.

[0114] The principles and operation of this invention can be better understood by referring to the figures and accompanying descriptions.

[0115] Before describing in detail at least one embodiment of the present invention, it should be understood that the present invention is not necessarily limited in terms of its applications to the details shown in the following description or illustrated in the examples. Other embodiments of the present invention are possible, or it can be carried out or performed in various ways.

[0116] Alpha-amino acids, the basic building blocks of proteins, are chiral molecules that exist in two forms: L-enantiomers (left-rotational or left-handed "L") and D-enantiomers (right-handed or right-handed "D"). These two incompatible forms of amino acids, differing in their handedness or chirality, are mirror images of each other and otherwise possess identical physical and chemical properties. However, life on Earth uses only L-amino acids and the achiral amino acid glycine to construct proteins that perform diverse biological functions. D-amino acids exist in nature, particularly in peptidoglycans in cell walls and bacterial peptide antibiotics, in proteins of lower animals such as insects, snails, and amphibians, and even as neurotransmitters in the brain. However, in various organisms, it is thought that they are converted from the parent L-enantiomer through enzyme-catalyzed post-translational reactions. The intriguing question of why and how life on Earth favors these left-handed molecules has been the subject of intense debate for decades, involving chemists, physicists, biologists, and even astronomers. While the origin of α-amino acid homochirality remains a mystery, scientists have already learned much by studying the physicochemical and biological properties of unnatural or artificial D-peptides and D-proteins that contain only chiral D-amino acids.

[0117] In implementing the present invention, the inventors reasonably determined that the core step in constructing a mirror-image biology system in the laboratory is to build a central dogma in chiral inversion versions of molecular biology by leveraging the advantages of the chemical synthesis of mirror-image nucleic acids and proteins as two technical pillars (5) (5-7). The inventors reasonably determined that one way to overcome the obstacles in synthesizing long L-nucleic acid molecules is through enzymatic polymerization by mirror-image polymerases, which aligns with the intent of the present invention and leads to the realization of proof of concept. Nevertheless, previous versions of mirror-image polymerase systems were selected as models of total chemical synthesis as a negative compromise between polymerase activity and size (5). The low processability and fidelity inherent in small polymerases such as ASFV pol X and Dpo4 (10 -4 ~10 -2 Due to their error rates (of a certain magnitude), these have been considered unsuitable for the accurate assembly, amplification, and transcription of long mirror image genes (5, 17, 18, 21).

[0118] Thus, the inventors have devised a method that can enable the total chemical synthesis of virtually any protein, thereby paving the way for D-amino acid proteins.

[0119] The method for the total chemical synthesis of large proteins according to embodiments of the present invention systematically removes obstacles that have remained unresolved in this field until now, and is based on introducing specific mutations into the amino acid sequence of the target protein to alleviate the length problem without invalidating the specific activity of the protein.

[0120] Design of division proteins: The inventors reasoned that, by leveraging the advantages of splitting protein design, the problem of chemical synthesis of large proteins could be dramatically simplified to the synthesis of two or even smaller protein fragments that can be folded together in vitro to form a functionally intact enzyme. Furthermore, this splitting protein strategy allows for the parallel synthesis, purification, ligation, and desulfurization of each splitting protein fragment, thereby reducing the overall time required for large protein synthesis, as well as the cost and time involved in correcting errors in one or more certain fragments. Some enzymes, including Pfu DNA polymerase, have native or engineered cleavage versions, for example, a known cleavage site between K467 and M468 in the coiled-coil motif of its finger domain, which splits the polymerase into two fragments (a Pfu-N fragment at 467aa and a Pfu-C fragment at 308aa) without significantly altering its PCR activity and fidelity. The aforementioned cleavage site may also be selected in the vicinity of the aforementioned sequence position in the coiled-coil motif of the finger domain of Pfu DNA polymerase, for example, between positions 449 and 498.

[0121] Thus, according to some embodiments of the present invention, a method for chemically producing a protein comprises dividing the amino acid sequence of the protein into at least two domain-forming segments, each of which is short enough to be chemically synthesized from ligation of smaller polypeptide segments, but long enough to fold into a functional domain in a functional protein when the domain-forming segments fold together under folding-inducible conditions.

[0122] According to some embodiments of the present invention, if a domain-forming segment is chemically synthesized by SPPS or AFPS or has a length of about 120, 150 or 200 amino acid residues or less, it typically means that it can be chemically synthesized and is suitable for folding together with other domain-forming segments, and therefore the protein can be obtained.

[0123] When used herein, the term "chemically synthesizable" primarily refers to the length of a polypeptide that can be achieved by any non-biological synthesis process, such as solid-phase peptide synthesis (SPPS) or automated fast-flow peptide synthesis (AFPS). Generally, it is known that polypeptides with a length of approximately 10–120 amino acids can be produced by solid-phase peptide synthesis (SPPS), and polypeptides with a length of approximately 10–180 amino acids can be obtained by automated fast-flow peptide synthesis (AFPS). In some embodiments, the term "chemically synthesizable" refers to polypeptide chains with a length of approximately 120, 150, or 200 amino acids. In some embodiments, the term "chemically synthesizable" also refers to the ability to purify and optionally isolate chemically synthesized polypeptides.

[0124] If the domain-forming segment is longer than what is suitable for chemical synthesis, it is further segmented into ligation-inducible segments, and these segments are linked together to form a (relatively long) domain-forming segment.

[0125] In relation to embodiments of the present invention, the term “fragment” is used synonymously with the term “domain-forming segment” throughout this description and the entire specification. When used herein, the term “domain-forming segment” refers to a continuous polypeptide chain that folds into one or more recognizable protein domains, as is known in the Art. According to some embodiments, a domain-forming segment may fold in vitro into one or more domains that are similar to, or essentially identical to, the structures of those domains when the polypeptide is folded in vivo or under biological / physiological conditions.

[0126] In relation to embodiments of the present invention, the domain-forming segment may be a multi-domain protein or may comprise a single recognizable domain. Domain recognition or identification is within the scope of those skilled in the art and is typically performed using one or more publicly available bioinformatics tools, such as multiple sequence alignment, SCOP [scop(dot)berkeley(dot)edu / ], CATH [www(dot)cathdb(dot)info], ExPASy [www(dot)expasy(dot)org], BLAST [blast(dot)ncbi(dot)nlm(dot)nih(dot)gov], PFAM [pfam(dot)xfam(dot)org], PDB [www(dot)rcsb(dot)org], etc. (all of which are within the scope and judgment of those skilled in the art).

[0127] As discussed above, some proteins are naturally constructed from two or more polypeptide chains, which are equivalent to the multidomains or domain-forming segments discussed herein. The methods presented herein can take advantage of such natural or intentional fission into domain-forming segments.

[0128] Some proteins can be constructed from a single, continuous polypeptide chain; however, their evolutionary family members may have evolved to be constructed from two or more polypeptide chains. Information about possible divisions may arise from multiple sequence alignments of family members, and from intentionally splitting family members of the target protein for chemical production. Another source of information regarding optional division sites may come from structural information of the target protein or its family members, aided by structural alignment—it may become clear that certain sections in a protein are less conserved and therefore are not expected to disrupt protein activity even if a division site is intentionally introduced into their sequence.

[0129] Sections within a protein that may function as potential cleavage sites are referred to herein as loss-of-structure sections, regardless of whether the information leading to their identification comes from sequence data and / or structural data. Thus, “loss-of-structure sections” can be identified by using multiple sequence alignment and / or from structural information from the protein of interest and / or members of the protein family.

[0130] According to some embodiments of the present invention, if a protein is too long to be directly produced chemically by SPPS or a combination of SPPS and ligation, a division site can be introduced into the sequence of the target protein, and it is expected that these domain-forming segments will fold together to form a protein when chemically synthesized.

[0131] Chemical ligation: As the inventors discovered while implementing the present invention, even if a protein can be realized by folding together, after the execution of the fission design method, each or one of the domain-forming segments may be too long to be realized by chemical synthesis.

[0132] Native chemical ligation (NCL) is an extension of the field of chemical ligation, and is a concept for constructing large polypeptides formed by assembling two or more unprotected peptide segments. In particular, NCL is a powerful ligation method for synthesizing small and medium-sized native or modified skeletal proteins. In native chemical ligation, the thiol group of the N-terminal cysteine ​​residue of an unprotected peptide attacks the C-terminal thioester of a second unprotected peptide. This reversible thioesterification step is chemoselective and regioselective, leading to the formation of a thioester intermediate. This intermediate then undergoes an intramolecular S,N-acyl shift, resulting in the formation of a native amide (peptide) bond at the ligation site.

[0133] In relation to embodiments of the present invention, the term "ligation-inducible sequence" refers to a site in a protein sequence that exhibits an amino acid sequence that can be formed by NCL. For example, using an N-terminal cysteine ​​residue, chemical ligation can be performed under known conditions. The identification and use of ligation-inducible sequences are well within the scope of those skilled in the art, and additional information is readily available in the literature (e.g., the review article “Native Chemical Ligation and Extended Methods: Mechanisms, Catalysis, Scope, and Limitations”, by Agouridas, V. et al. [Chem Rev. 2019, 119(12), pp. 7328-7443]).

[0134] Thus, according to some embodiments of the present invention, a protein or its long-chain domain-forming segment can be synthesized by first identifying ligation-inducible sequences in the amino acid sequence of the protein, and then parsing the sequence with those ligation-inducible sequences or at least a portion thereof to obtain a plurality of sequences of protein ligation-inducible segments, each of which is short enough to be actually chemically synthesized and purified. Subsequently, when each of the chemically synthesizable ligation-inducible segments is ligated together, a protein or domain-forming segment is formed.

[0135] Generally, according to some embodiments of the present invention, the ligation-inducible sequence / segment is chemically synthesizable or has a length of about 10-120, about 10-150, or about 10-200 amino acids.

[0136] If, based on the segment length, the protein does not exhibit a ligation-inducible sequence at a desired position, the ligation-inducible sequence can be introduced by mutation in the amino acid sequence of the protein. Thus, according to some embodiments of the present invention, if any one of the ligation-inducible segments is not chemically synthesizable, i.e., is longer than about 120, 150, or 200 amino acid residues, or is of other lengths that cannot actually be synthesized and purified, the method is carried out by identifying at least one loss-of-structure section in the ligation-inducible sequence, introducing the ligation-inducible sequence into the loss-of-structure section by substituting at least one amino acid in the loss-of-structure section with a ligation-inducible amino acid residue, subsequently parsing the amino acid sequence of the protein with the ligation-inducible sequence resulting from the mutation, and further subsequently chemically synthesizing each of the ligation-inducible segments.

[0137] For example, the synthesis of the 467aa (54kDa) Pfu-N fragment alone, which is much larger than the 352aa (40kDa) Dpo4, still presents significant challenges. One challenge is that the NCL of synthetic peptides prepared by SPPS requires an N-terminal cysteine ​​residue at the ligation site, but wild-type (WT) Pfu DNA polymerase has only four cysteine ​​residues (C429 and C443 of the Pfu-N fragment (SEQ ID NO: 57), and C507 and C510 of the Pfu-C fragment (SEQ ID NO: 67)). The inventors leveraged the advantages of previously reported metal free radical-based desulfurization methods to convert unprotected cysteine ​​residues to alanine residues after NCL, making eight other ligation sites containing alanine residues (A40, A163, A223, and A408 in the Pfu-N fragment, and A501, A596, A652, and A715 in the Pfu-C fragment) available. However, some of these peptide segments were still too long to be prepared by SPPS. Therefore, based on sequence alignment, the inventors designed a mutant Pfu DNA polymerase with five point mutations (E102A, E276A, K317G, and V367L in the Pfu-N fragment, and I540A in the Pfu-C fragment) to introduce additional ligation sites or ligation-inducible sequences without significantly altering the polymerase's PCR activity (Fixed Pfu-5m, SEQ ID NO: 48).

[0138] Hydrophobic and bulk: Another challenge is the synthesis and ligation of hydrophobic peptide segments under aqueous conditions. Current methods to overcome this problem primarily focus on introducing various mutations and / or chemical modifications to the target peptide to reduce the number of highly hydrophobic and / or bulky amino acid residues. According to some embodiments of the present invention, the chemical modification is, for example, Hmb-N α This can be done using protective, removable solubilization tags, pseudoproline, and depsipeptides (O-acyl isopeptides), but their practical use is often limited by cumbersome procedures, low yields, and the need for expensive amino acid derivatives.

[0139] According to some embodiments of the present invention, certain highly hydrophobic and / or bulky residues are substituted (mutated) with less hydrophobic and / or less bulky residues to facilitate the folding of various segments of chemically synthesized, ligated, and chemically produced proteins together. The criteria for such substitutions may rely on MSA, structural information, and other mutation data.

[0140] Hydrophobicity and bulkiness are related and, in most cases, closely associated, but they are not necessarily the same property, and their properties can vary depending on pH, ionic strength, counterions, water activity, temperature, and other factors under different environmental conditions. Regarding polypeptide chains, the values ​​and rankings of the hydrophobicity and bulkiness of amino acid residues vary somewhat depending on the literature cited, but the general view that isoleucine is "one of the amino acids with exceptionally high bulkiness and hydrophobicity" holds true in all cases. Exemplary sources of information regarding hydrophobicity and bulkiness include, but are not limited to, Kyte, J. and Doolittle, RF, “A simple method for displaying the hydropathic character of a protein” [J. Mol. Biol., 1982, 157(1), pp. 105-132] and Ellington, A. and Cherry, JM, “Characteristics of amino acids” [Curr Protoc Mol Biol, 2001, A.1C.1-A.1C.12]. For example, embodiments of the present invention may reduce bulkiness according to the following non-limiting exemplary order: I>L>C>T>V>P>S>A>G as a basic criterion for amino acid mutations, and reduce hydrophobicity according to the following non-limiting exemplary order: I>V>L>F>C>M>A>G>T.

[0141] Generally, as is well known in the art, the guideline for substituting residues follows the following order of hydrophobicity: Ile>Leu>Phe>Val>Met>Pro>Trp>His(0)>Thr>Glu(0)>Gln>Cys>Tyr>Ala>Ser>Asn>Asp(0)>Arg+>Gly>His+>Glu>Lys+>Asp-.

[0142] When the methods presented herein are used for the chemical synthesis of D-amino acid proteins, according to some embodiments thereof, the methods may further include substituting at least one hydrophobic D-amino acid residue in at least one ligation-inducible segment with a less hydrophobic amino acid, in accordance with the following hydrophobic order: D-Ile>D-Leu>D-Phe>D-Val>D-Met>D-Pro>D-Trp>D-His(0)>D-Thr>D-Glu(0)>D-Gln>D-Cys>D-Tyr>D-Ala>D-Ser>D-Asn>D-Asp(0)>D-Arg+>Gly>D-His+>D-Glu>D-Lys+>D-Asp-.

[0143] For example, the Pfu-C-4 segment has low solubility in acetonitrile or 6M Gn·HCl aqueous solution, making it difficult to synthesize using standard Fmoc-SPPS. Since isoleucine is considered to be one of the most bulky and hydrophobic proteinogenic amino acids, it was thought that the physicochemical properties of the peptide segment would change if one or more isoleucines in a hydrophobic peptide were replaced with alternative amino acids that are potentially less bulky or hydrophobic (e.g., valine, alanine, leucine, threonine, glycine, phenylalanine, methionine, or proline), or if one or more other bulky or hydrophobic amino acids (e.g., valine, threonine, phenylalanine, and leucine) were replaced with other amino acids that are less bulky or hydrophobic, such as more polar amino acids.

[0144] According to some embodiments of the present invention, a systematic isoleucine substitution method based on sequence alignment and structural information was developed to mutate all seven isoleucine residues of this segment (I598V, I605T, I611V, I619A, I631L, I643V, and I648T) without significantly altering the PCR activity of the polymerase. In fact, these seven point mutations facilitated the synthesis of this peptide segment, which also allowed for solubilization in acetonitrile and 6M Gn·HCl aqueous solution for downstream purification and NCL, thus avoiding the need for other chemical modifications for its synthesis.

[0145] Cost reduction: In addition to technical challenges, the synthesis of large enantiomer (D-amino acid) proteins also faces economic obstacles due to generally low yields and high reagent costs. While enantiomer versions of proteinogenic amino acids are all commercially available and mostly priced similarly to their natural counterparts, D-isoleucine is approximately 50 to 300 times more expensive than L-isoleucine and other D-amino acids. This is mainly due to the difficulty and high loss associated with its synthesis and purification, caused by the presence of two chiral centers. When synthesizing enantiomer proteins, the cost of D-amino acids accounts for 80 to 90 percent (typically around 5 percent, depending on the abundance of isoleucine in the natural protein). Thus, according to some embodiments of the present invention, by applying a systematic isoleucine substitution method based on sequence alignment and structural information, a large number of isoleucines (41 out of 71, i.e., 58%) in Pfu DNA polymerase are mutated to other amino acids such as valine, leucine, and alanine without significantly altering the PCR activity of the polymerase (Fixation Pfu-5m-30I, SEQ ID NO: 51).

[0146] As a result of this systematic Ile reduction method, the D-amino acid cost in synthesizing this polymerase is reduced by almost half, which could be beneficial for its large-scale synthesis and application in the future.

[0147] According to some embodiments, a method for chemically producing a D-amino acid protein includes substituting at least one Ile residue with an Ala, Val, Leu, Gly, Thr, Phe, Met, or Pro residue. Thus, the resulting D-amino acid protein exhibits some or all Ile residue positions as non-Ile D-amino acid residues selected from the group consisting of D-Ala, D-Val, D-Leu, Gly, D-Thr, D-Phe, D-Met, and D-Pro residues.

[0148] Chemical total synthesis methods for large proteins: As described above and demonstrated in the Examples section below, the method provided herein has resulted in the highly fidelity total chemical synthesis of a 90 kDa D-amino acid Pfu DNA polymerase. This enabled the precise writing and reading of L-DNA sequences and the precise assembly of kilobase-sized enantiomer genes. The average size of natural enzyme proteins is approximately 300–500 aa, corresponding to a coding gene sequence of approximately 0.9–1.5 kb. Thus, the ability to synthesize large enantiomer enzyme proteins such as Pfu DNA polymerase, and therefore the ability to assemble long enantiomer genes, is a key enabling technique and an important step toward constructing enantiomers of life. From the first-generation enantiomer polymerase ASFV pol X and the second-generation Dpo4 to the current third-generation Pfu DNA polymerase, technological advancements have made the total chemical synthesis of large enantiomer proteins utilizing the best enzymatic means provided by nature a reality. These efficient next-generation enantiomers open new doors to opportunities for realizing more sophisticated enantiomer biology systems and expanding the biotechnology and medical molecular toolbox.

[0149] Thus, according to certain embodiments of the present invention, a method for the total chemical synthesis of a relatively large functional protein is provided, which is carried out by ligating at least two ligation-inducible segments of the protein. Here, each of the ligation-inducible segments is chemically synthesizable or typically about 10 to 120 amino acid residues in length for SPPS, and the ligation-inducible segments can be obtained by the following:

[0150] i. Identify at least one ligation-inducible sequence in the amino acid sequence of the protein, and parse (split) the amino acid sequence of the protein by those ligation-inducible sequences to obtain multiple sequences of ligation-inducible segments. According to some embodiments, at least one of the naturally occurring ligation-inducible sequences is found in the loss-of-structure section of the protein.

[0151] ii. If each sequence of the ligation-inducible segment can be synthesized and purified by SPPS and / or AFPS, then each of the ligation-inducible segments can be chemically synthesized and prepared for ligation.

[0152] iii. If any one of the ligation-inducible segment sequences is not chemically synthesizable, i.e., longer than approximately 120, 150, or 200 amino acid residues, or of other lengths that cannot actually be synthesized and purified, such sequences are analyzed to identify at least one loss-of-structure section within them. This analysis is described above and is well known in the art. To introduce a ligation-inducible sequence by mutation, at least one amino acid in the loss-of-structure section is replaced with a ligation-inducible amino acid residue (e.g., cysteine) to introduce the ligation-inducible sequence into the loss-of-structure section. Subsequently, the amino acid sequence of the protein is parsed by this newly introduced ligation-inducible sequence, and the resulting ligation-inducible segment smaller than 120aa is chemically synthesized.

[0153] As discussed above, utilizing existing division sites or introducing division sites into the amino acid sequence of a protein facilitates the total chemical synthesis of the protein. Thus, according to some embodiments of the present invention, the method further comprises, prior to step (i) presented above, dividing the amino acid sequence of the protein into at least two domain-forming segments, and if each of the domain-forming segments is chemically synthesizable (approximately 120, 150, or 200 amino acid residues or less), chemically synthesizing each of the domain-forming segments, and then folding those domain-forming segments together to obtain a protein.

[0154] According to some embodiments, if one of the domain-forming segments is not chemically synthesizable (for example, longer than about 120, 150, or 200 amino acid residues) or is of another length that cannot actually be synthesized and purified, it is further divided into ligation-inducible segments, as discussed above.

[0155] Preferably, the domain-forming segment is parsed by its loss-of-structure sections. This begins with identifying the loss-of-structure sections within the domain-forming segment, followed by identifying at least one ligation-inducible sequence within the loss-of-structure sections, and then parsing the amino acid sequence of the domain-forming segment with those ligation-inducible sequences. Again, if the segment or loss-of-structure section inherently lacks ligation-inducible sequences, they can be introduced by mutation, as described above. Once the domain-forming segment has been parsed and yielded chemically synthesizable (approximately 10-120 aa for SPPS and approximately 10-180 for AFPS) ligation-inducible segment sequences, these are chemically synthesized and ligated to form the domain-forming segment.

[0156] Figure 1 illustrates the method provided in this application in flowchart form, where in "Box 1", the user selects a target protein, preferably a protein family and for which structural information is available, and in "Box 2", the method requires the user to use MSA and structural data to identify loss-of-structure sections for introducing ligation-inducible aa mutations, division sites, and Ile residue substitutions. If the target protein is shorter than approximately 400 aa, in "Box 3", the method requires the user to parse the protein sequence into ligation-inducible segments by finding or mutating ligation-inducible aa into the ligation-inducible sequence and / or introducing it, thereby forming multiple ligation-inducible segment sequences, each of which can be chemically synthesized. If the target protein is longer than approximately 400 aa, in "Box 4," this method requires finding or introducing at least one division site to form domain-forming segments, each less than approximately 400 aa. In "Box 5," this method requires parsing each sequence of the domain-forming segments into ligation-inducible segments by finding and / or introducing them into the ligation-inducible sequence, thereby forming multiple ligation-inducible segment sequences, each of which can be chemically synthesized. In "Box 6," this method requires substituting hydrophobic aa in each of the domain-forming segments or in the resulting ligation-inducible segments based on sequence conservation criteria according to the MSA and / or structural information.If the target protein is a D-amino acid protein, "Box 7" requires mutating as many Ile residues as the MSA and / or structural information allows with similar aa in each domain-forming segment or the resulting ligation-inducible segment, and "Box 8" requires synthesizing all ligation-inducible segments using D-amino acids and linking the segments accordingly. If the target protein is an L-amino acid protein, "Box 9" requires synthesizing all ligation-inducible segments using L-amino acids and linking the segments accordingly, and finally, "Box 10" requires folding all domain-forming segments together to obtain the target protein.

[0157] In some embodiments of the present invention, the method requires a mutation step to make the amino acid sequence of the target protein suitable for total chemical synthesis. This requirement may arise from the excessive length of the target protein, in which case mutation is necessary to introduce a division site or a ligation-inducible sequence not present in the corresponding biologically expressed protein, or to provide a ligation-inducible segment that is sufficiently short to be realized by SPPS (or other chemical methods for producing polypeptides). This requirement may arise from the excessive hydrophobicity of the ligation-inducible segment, which makes polypeptide synthesis and ligation under aqueous conditions difficult. On the other hand, reducing its hydrophobicity makes the polypeptide more suitable for this task.

[0158] In some embodiments of the present invention, when a protein is realized as a D-amino acid protein, i.e., a mirror image of its corresponding biologically produced (or expressed) protein, i.e., a mirror image of an equivalent L-amino acid protein, it is necessary to mutate the amino acid sequence of the target protein so that the total chemical synthesis cost is reduced.

[0159] In relation to embodiments of the present invention, the terms "corresponding protein," "corresponding biologically produced protein," and "corresponding biologically expressed protein" are used synonymously and refer to proteins that are equivalent to the protein produced by the method provided in this application in terms of function and, to some extent, structure, but differ in their manufacturing process and amino acid sequence, and that, as described above, can be mutated during the process of carrying out the method provided in this application. In the case of enantiomer proteins, the term "corresponding L-amino acid protein" is similar to the term "corresponding biologically produced protein" with a structural inversion compared to an equivalent L-amino acid protein. Thus, the D-amino acid proteins produced by the method provided herein are related to equivalent proteins in the following respects: they have substantially similar sequences except for mutations that may occur to introduce a fission site and result in a domain-forming segment, and / or mutations that may occur to introduce a ligation-inducible sequence, and / or mutations that may occur to reduce the hydrophobicity of residues, and / or mutations that may occur to reduce the number of Ile residues; they have a composition in which at least 90% consists of D-amino acid residues other than Gly, rather than L-amino acid residues; they have substantially inverted (mirror image) structures; and they have similar activity except for having mirror image ligands, substrates, products, etc. These sequences, compositions, structures, and activities are also present to some extent between chemically produced proteins according to some embodiments of the present invention and their corresponding biologically produced proteins, except that these two are made up of L-amino acid residues and are therefore not mirror images of each other in terms of structure and activity.

[0160] Some methods of chemically synthesizing proteins involve the purification and isolation of the resulting protein after ligation or folding of multiple chemically synthesized chains. The purification protocol can be any known protocol for such protein purification work, and if the target protein is some that is thermally stable, the protocol can take advantage of this thermal stability by including a heating step. That is, the protocol includes a synthesis / ligation step, a subsequent folding step, and a further heating and precipitation step as part of the purification of the final result. The heating and precipitation temperature is usually set between the highest stable temperature of the target protein and the lowest precipitation temperature for many impurities (misfolded polypeptide chains and polypeptide chains with incorrect amino acid sequences). For example, for Pfu DNA polymerase, the highest stable temperature is about 95°C, and therefore the heating and precipitation temperature is set to about 85°C. For Dpo4, the highest stable temperature is about 86°C, and therefore the heating and precipitation temperature is set to about 78°C. Precipitated (thermally unstable) impurities are generally removed by ultracentrifugation and / or filtration, while correctly folded, thermally stable proteins are found in the supernatant and can be isolated therefrom. It is noted herein that multiple folding and heat-precipitation rounds are performed to increase the overall yield of correctly folded proteins, and the proteins precipitated from one or more previous folding and heat-precipitation rounds are subjected to further refolding and reheat-precipitation rounds rather than being discarded, as is common in such procedures.

[0161] In addition to the foregoing, the scope of the present invention also includes cases in which biologically produced proteins and / or protein fragments are used to induce the correct folding of synthetically produced proteins and / or protein fragments. Thus, synthetic proteins and their fragments are also brought about by folding together with biologically produced proteins or their fragments according to some embodiments of the present invention, while the final result may be a chimeric multi-fragment / multi-domain protein having a biologically produced portion and a synthetically produced portion.

[0162] Chemically synthesized proteins: According to certain embodiments of the present invention, a protein is provided that is chemically synthesized by a method disclosed herein. In some embodiments, the chemically produced protein has a length of at least about 240 amino acid residues, or at least about 250 amino acid residues, or at least about 300 amino acid residues, or at least about 350 amino acid residues, or at least about 400 amino acid residues, or at least about 450 amino acid residues, or at least about 500 amino acid residues, or at least about 550 amino acid residues, or at least about 600 amino acid residues.

[0163] Chemically synthesized proteins can be any of the target proteins and can function as enzymes, transport proteins, structural / mechanical proteins, hormones, signaling proteins, antibodies, fluid balance proteins, pH balance proteins, cell channels, or cell pumps, etc.

[0164] Chemically synthesized proteins are functional in the same way as their biologically produced and / or recombinantly produced counterparts, also referred to herein as the corresponding biologically produced proteins. Chemically synthesized proteins retain at least 5% of the activity of the corresponding biologically produced protein. In some embodiments, chemically synthesized proteins retain at least 1%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or at least 90% of the activity of the corresponding biologically produced protein.

[0165] To retain at least some percentage of the activity of the corresponding biologically produced protein means that, if the biologically produced protein exhibits catalytic activity, specific binding activity, and / or any structurally related activity, the corresponding chemically produced protein of the present invention exhibits at least 5% of that activity. In the case of D-amino acid proteins, activity is defined, determined, and measured using appropriate / corresponding enantiomer substrates, enantiomer reactants, enantiomer reagents, etc., corresponding to the enantiomer protein, compared to its corresponding L-amino acid protein, whether chemically and / or biologically obtained.

[0166] According to some embodiments of the present invention, a D-amino acid protein exhibits a three-dimensional structure that is essentially a mirror image of the three-dimensional structure of its corresponding biologically produced L-amino acid protein. In this application, producing a D-amino acid protein, also referred to as a mirror image protein (to its corresponding L-amino acid protein or naturally occurring protein), means producing it using at least 75%, 80%, 90%, or at least 95% of D-amino acid residues other than glycan when chemically producing a ligation-inducible segment.

[0167] When we say that a protein contains at least two domain-forming segments, it means that, according to embodiments of the present invention, the resulting chemically produced protein contains at least two non-covalently attached (not attached by backbone atoms) polypeptide chains, each corresponding to a domain-forming segment. In some embodiments, the corresponding domain-forming segments are polypeptide chains covalently attached to at least one corresponding family member of a biologically produced protein.

[0168] In this application, it is noted that when a synthetic L- / D- protein is used in any reaction, the reaction mixture can be isolated, and the synthetic protein can be recycled by affinity purification for use in future reactions or reused due to its rare and costly amino acid residues. For example, the synthetic protein can be manufactured with any known affinity tag, such as a His6 tag, and after use, the reaction mixture can be incubated with the corresponding affinity resin or beads to isolate the synthetic L- / D- enzyme from the reaction mixture.

[0169] Exemplary proteins prepared by this method: According to another aspect of some embodiments of the present invention, a protein produced by the method provided herein is provided, having an amino acid residue length of at least about 240, 300, 350, 400, 500 or more. This protein may be an L-amino acid protein or a D-amino acid protein, depending on the amino acids used in the chemical synthesis of the corresponding ligation-inducible segment, for example by SPPS.

[0170] Tables 1 and 2 below list gene-encoded amino acids (Table 1) and non-limiting examples of non-standard / modified amino acids that can be used in the present invention (Table 2).

[0171] [Table 1]

[0172] [Table 2-1]

[0173] [Table 2-2]

[0174] [Table 2-3]

[0175] [Table 2-4]

[0176] To demonstrate a method for the total chemical synthesis of proteins, the inventors synthesized active enzymes capable of catalyzing reactions catalyzed by their corresponding biologically produced enzymes. One such enzyme is an RNA polymerase capable of synthesizing RNA from ribonucleotides using a DNA template. In the following Examples section, the exemplary RNA polymerase is T7 RNA polymerase. In another example, the enzyme is a DNA polymerase capable of synthesizing DNA from deoxyribonucleotides. In the following Examples section, the exemplary DNA polymerase is Pfu DNA polymerase.

[0177] When the method provided herein is used to produce D-amino acid RNA polymerase, this unique enantiomer enzyme can synthesize L-RNA from L-ribonucleotides using an L-DNA template. For example, the D-amino acid RNA polymerase is D-amino acid T7 RNA polymerase.

[0178] As presented below, the D-amino acid T7 RNA polymerase is prepared to include at least one cleavage site, a first cleavage site between K363 and P364 using a WT position numbering scheme, and a second cleavage site between N601 and T602. Alternatively, the D-amino acid T7 RNA polymerase and the L-amino acid T7 RNA polymerase produced by the method provided herein include at least two polypeptide chains formed by the cleavage between K363 and P364 and / or between N601 and T602. Furthermore, the cleavage sites may be selected in the vicinity of the aforementioned sites located in the same loop, i.e., positions 357-366 and / or positions 564-607.

[0179] According to certain embodiments of the present invention, the T7 RNA polymerase produced by the method provided herein may further include at least one mutation selected from the group consisting of I6V, I14L, I74V, I82V, I109V, I117L, I141V, I210M, I244L, I281V, I320V, I322L, I330V, and I367L. These mutations facilitate cost reduction strategies by substituting costly D-Ile residues with other suitable D-amino acid residues.

[0180] According to one aspect of the present invention, a D- or L-amino acid T7 RNA polymerase produced by the method provided herein is provided, which has an amino acid sequence identical to SEQ ID NO: 83 or having at least 80-90% sequence identity with SEQ ID NO: 83.

[0181] When the method provided herein is used to produce D-amino acid DNA polymerase, this unique enantiomer enzyme can synthesize L-DNA from L-deoxyribonucleotides. For example, D-amino acid DNA polymerase is D-amino acid Pfu DNA polymerase.

[0182] Thus, according to another aspect of the present invention, a Pfu DNA polymerase is provided comprising at least two polypeptide chains formed by fission between K467 and M468 (where the position numbering is based on the amino acid position numbering of the corresponding WT enzyme). It is noted herein that other fission sites may be selected in the vicinity of this site, i.e., within the coiled-coil motif of the finger domain of the Pfu DNA polymerase, for example, between positions 449 and 498.

[0183] According to some embodiments, the synthetic Pfu DNA polymerase provided herein further comprises at least one mutation selected from the group consisting of E102A, E276A, K317G, V367L, and I540A. According to other embodiments, the Pfu DNA polymerase provided herein further comprises at least one mutation selected from the group consisting of V93Q, D141A, E143A, Y410G, A486L, and E665K.

[0184] According to one aspect of the present invention, a D- or L-amino acid Pfu DNA polymerase having or not having a DNA-binding structure domain (SEQ ID NO: 78), which is produced by the method provided herein, is provided, which is selected from the group consisting of SEQ ID NO: 48, SEQ ID NO: 49, SEQ ID NO: 50, SEQ ID NO: 51, SEQ ID NO: 74, SEQ ID NO: 75, SEQ ID NO: 76, SEQ ID NO: 77, and SEQ ID NO: 79, or which has an amino acid sequence having at least 80-90% sequence identity with SEQ ID NO: 51.

[0185] Bioorthogonal data storage: The accelerating pace of data production worldwide has created a growing need for reliable, high-density media to store vast amounts of information. Natural DNA has evolved to encode, store, and propagate information.

[0186] Storage using DNA, a select group of natural molecules that code for a vast amount of genomic instructions within tightly packed chromosomes, is emerging as a promising solution (1-3). On the other hand, mirror-image DNA shows unique suitability for the challenge of bioorthogonal information storage, and for this purpose, methodologies for registering and searching L-DNA data are essential, but most of them remain uninvestigated.

[0187] The inventors envisioned that chiral-inverted (mirror image) DNA possessing the same information capacity could retain its unique ability to avoid biological degradation and contamination, and thus function as a highly robust bioorthogonal data repository. In implementing the present invention, a high-fidelity 90 kDa D-amino acid Pfu DNA polymerase according to some embodiments of the present invention was chemically synthesized for accurate writing and reading of L-DNA sequences.

[0188] The inventors demonstrated the storage of an entire passage of digital text in enantiomer DNA, one of the embodiments of the present invention. As can be seen in the following Examples section, L-DNA barcodes carrying trace amounts of messages in unpurified environmental water samples remained stable and amplified for months and potentially beyond. Furthermore, the high-fidelity D-polymerase produced by certain embodiments of the present invention enabled the precise assembly of full-length kilobase-sized enantiomer genes, a crucial step toward realizing enantiomer translation and establishing the enantiomer central dogma. This success in synthesizing next-generation enantiomer enzyme tools, and thus assembling long enantiomer genes, has transformed the development of enantiomer biology systems and the exploration of their emerging applications.

[0189] Simply put, DNA is essentially a data storage molecule. It contains all the instructions a cell (or organism as a whole) needs to maintain itself. These instructions are found in genes, which are specific sequences of nucleotides within DNA. For the instructions contained within a gene to be executed, it must be expressed or copied into a form that the cell can use to produce the proteins necessary to sustain life. The instructions stored in DNA are read and processed by the cell in two steps: transcription and translation. These steps are distinct biochemical processes involving multiple molecules. During transcription, a portion of the cell's DNA acts as a template for creating RNA molecules. In some cases, the newly created RNA molecule itself is the final product, performing an important function within the cell. In other cases, the RNA molecule carries the message from the DNA to other parts of the cell for processing. In most cases, this information is used to produce proteins. Certain types of RNA that carry information stored in DNA to other regions of the cell are called messenger RNA or mRNA.

[0190] Figure 4 is a flowchart illustrating molecular data storage according to some embodiments of the present invention, using L-DNA as an exemplary XNA.

[0191] Thus, according to one embodiment of the present invention, a method is provided for forming a bioorthogonal data storage polymer using D-amino acid RNA polymerase or D-amino acid DNA polymerase and L-ribonucleic acid or L-deoxyribonucleic acid, respectively, wherein the polymerase is produced by the method provided herein.

[0192] According to another embodiment of the present invention, a method is provided for forming a bioorthogonal data storage polymer using the D-amino acid RNA polymerase or the D-amino acid DNA polymerase provided in the present application and L-ribonucleic acid or L-deoxyribonucleic acid, respectively.

[0193] According to another embodiment of the present invention, a method is provided for decoding a bioorthogonal data storage polymer using at least one D-amino acid protein produced by the method provided herein, wherein the bioorthogonal data storage polymer comprises L-ribonucleic acid or L-deoxyribonucleic acid residues.

[0194] Furthermore, according to another embodiment of the present invention, a bioorthogonal data storage system is provided which is substantially as described above, including: at least one L-DNA that encodes informational data in a sequence using four letters A, T, G, and C; and a D-amino acid RNA / DNA polymerase for L-DNA synthesis (writing the code into the DNA sequence) and / or L-DNA sequencing (reading the code in the DNA sequence).

[0195] It should be noted hereby noted that the scope of the present invention is intended to include the use of other types of nucleotides and polymers thereof that are not naturally occurring or atypical, and which are referred to in this application and in the art as "xeno nucleic acids" or XNAs. Thus, according to some embodiments of the present invention, the systems and methods provided herein for producing and using molecular data storage include, for example, the use of XNAs as discussed by Eremeeva, E and Herdewijn, P. in the publication "Non canonical genetic material" [Current Opinion in Biotechnology, 2019, 57, pp. 25-33] and as discussed by Chaput, JC et al. [Chem. Biol., 2012, 21;19(11), pp. 1360-71].

[0196] The precise assembly, amplification, and sequencing of L-DNA offers excellent opportunities for bioorthogonal information storage, environmental and food barcoding, medical implant monitoring, forensic testing, and secure messaging. These were not possible with previous versions of enantiomer polymerase systems, such as ASFV pol X or Dpo4, due to their inefficiency and high error rates in amplifying and sequencing small amounts of signaling L-DNA molecules (5, 17, 18, 21). If the system can accurately assemble enantiomer genes, and potentially even entire genomes in the future, it could also be suitable for creating enantiomer genome backup copies of naturally occurring organisms for genome banking and interplanetary transport.

[0197] Enantiomer ribosomes: The next step in establishing the enantiomer central dogma is to achieve enantiomer translation by constructing functional enantiomer ribosomes. While we have recently overcome the limitations of L-RNA chemosynthesis (typically less than 70 nt) by transcribing synthetic L-DNA templates to 120 nt full-length 5S rRNA, obtaining 1.5 kb 16S rRNA and 2.9 kb 23S rRNA, as well as mRNA for translation, requires a more efficient enzymatic system capable of transcribing enantiomer genes to longer L-RNAs. One possibility, as demonstrated previously, is to mutate DNA polymerase into DNA-dependent RNA polymerase. Indeed, we have successfully re-engineered fission-type Pfu DNA polymerase (with seven point mutations: V93Q, E102A, D141A, E143A, Y410G, A486L, and E665K) into an efficient DNA-dependent RNA polymerase. However, the preparation and purification of long single-stranded (ss) L-DNA templates present another challenge that must be addressed first. Alternatively, if a mirror-image version of 100kDa T7 RNA polymerase using a double-stranded (ds) L-DNA template could be synthesized, it should be possible to enzymatically transcribe any mirror-image rRNA and the mRNA required for mirror-image translation. In the process of implementing the present invention, D-amino acid T7 RNA polymerase was realized by total chemical synthesis according to some embodiments of the present invention, as presented in the Examples section that follows.

[0198] Racemic crystal structure analysis: As is well known in the field of protein crystal structure analysis, the first and perhaps greatest rate-limiting step in elucidating protein structure is obtaining X-ray diffractable crystals. In crystallization experiments of small molecules, it has been observed that racemic mixtures of two enantiomers of a given molecule tend to form high-quality diffractable crystals, and at least one of the symmetry operations observed in the unit cell is inversion. In the newly emerging field of racemic crystal structure analysis in structural biology, a shortage of enantiomer protein samples is a problem due to their rarity, especially when searching for large enantiomer proteins.

[0199] Thus, according to some embodiments of the present invention, a method for forming crystals of a target protein is provided, which is carried out by co-crystallizing the target protein with an enantiomorph of the target protein obtained as provided herein, thereby forming crystals of an enantiomer protein pair, where the enantiomorph is a D-amino acid (mirror image) protein and a corresponding L-amino acid protein of the target.

[0200] In another kind of embodiment of the present invention, the enantiomorph is produced by a mirror image protein, as provided herein. For example, an L-RNA can be transcribed using a high-fidelity mirror image RNA polymerase provided as discussed herein, thereby producing an enantiomorph of its corresponding D-RNA, which can then be used for enantiomeric / racemic cocrystallization with the D-RNA to determine the RNA structure.

[0201] For additional information regarding the analysis of racemic crystal structures, see, for example, Matthews, BW, “Racemic crystallography—Easy crystals and easy structures: What's not to like?”, Protein Science, 2009, 18(6), pp. 1135-1138, Yestes, TO and Kent, SBH, “Racemic Protein Crystallography”, Annual Review of Biophysics, 2012, 41(1), pp. 41-61, and Mandal, PK et al., “Racemic DNA Crystallography”, Angewandte Chemie International Edition, 2014, 53(52), pp. 14424-14427 (these contents are incorporated herein by reference as being fully explained herein).

[0202] Sequencing: According to certain embodiments of the present invention, when the synthetic protein is used in sequencing and denaturing sequencing PAGE for separating chemically synthesized enantiomer DNA oligos, the quality of the synthetic oligos can be substantially improved by reducing the vast majority of -1 and -2nt products. Using either a D- or L-amino acid synthetic protein in this manner improves the fidelity of the sequencing process, resulting in the majority of the ultimately assembled gene sequences being correct.

[0203] According to some embodiments of the present invention, unlabeled carrier D-(or L-)DNA is added to the sample before purification by denaturing sequencing PAGE (which has a specific required amount as its "dead volume") in order to reduce the scale required for enantiomer PCR and gel purification of the PCR-amplified L-DNA product. According to some embodiments of the present invention, high-fidelity synthetic enantiomer polymerase can be used with phosphorothioate L-dNTPs for sequencing-by-synthesis of enantiomer nucleic acids such as L-DNA and L-RNA. The use of a bidirectional sequencing strategy with two primers 5'-labeled with two different dyes (FAM and Cy5, respectively) is also used to improve the read length in a single reaction to more than 160-170 bp.

[0204] In vitro evolution: The development of sequencing-by-synthesis using, for example, the enantiomer Pfu DNA polymerase provided in this application, according to some embodiments of the present invention, represents a new step toward realizing a more efficient L-DNA sequencing technique compared to cumbersome L-DNA chemical sequencing methods.

[0205] In vitro evolution (SELEX), also known as in vitro selection or in vitro evolution, is a combinatorial chemistry technique in molecular biology for producing oligonucleotides of either single-stranded DNA or RNA that specifically bind to one or more target ligands. This process begins with the synthesis of a large oligonucleotide library consisting of fixed-length, randomly generated sequences with constant 5' and 3' ends adjacent, which act as primers. For a randomly generated region of length n, the number of possible sequences in the library is 4. n This is the case (for each position, there are n positions with four possibilities (A, T, C, and G)). Sequences in the library are exposed to a target ligand (which can be a protein or a small organic compound), and those that do not bind to the target are usually removed by affinity chromatography or target capture with paramagnetic beads. Binding sequences are eluted, amplified by PCR, and prepared for a subsequent selection round. By increasing the stringency of the elution conditions, the tightest binding sequences can be identified. SELEX has been used to develop several aptamers that bind to targets of interest for both clinical and research purposes. For this purpose, several nucleotides with chemically modified sugars and bases have been incorporated into the SELEX reaction. These modified nucleotides allow for the selection of aptamers with novel binding properties and potentially improved stability.

[0206] Further efforts to re-engineer high-fidelity enantiomer polymerases (e.g., through the synthesis of mutant or truncated versions lacking 3'-5' exonuclease activity) for enantiomer Sanger sequencing and more automated high-throughput L-DNA sequencing techniques could lead to novel applications such as multiplex L-DNA sequencing and enantiomer in vitro evolution (MI-SELEX) for the direct selection of L-aptamer drugs (17, 18).

[0207] During the term of this patent from the date of application until its expiration, it is anticipated that many relevant large synthetic D / L proteins will be developed, and the term "large synthetic D / L protein" is intended to a priori encompass all such novel technologies.

[0208] As used herein, the term "about" refers to ±10% (for example, "about 30" means 27 to 33 or 30 ± 3).

[0209] The terms "to include," "to contain," "to encompass," "to possess," and their conjugated forms all mean "to include, but not limited to."

[0210] The term "to consist of" means "to include and be limited to."

[0211] The term “essentially derived from” means that the composition, method, or structure may include additional components, steps, and / or parts, provided that such additional components, steps, and / or parts do not substantially alter the fundamental novel features of the claimed composition, method, or structure.

[0212] As used herein, in relation to certain substances, the phrases “substantially absent” and / or “essentially absent” mean a composition that is completely absent of the substance or contains less than about 5, 1, 0.5, or 0.1 percent of the substance by total weight or total volume. Alternatively, in relation to processes, methods, properties, or features, the phrases “substantially absent” and / or “essentially absent” mean a process / composition, structure, or article that completely lacks certain process / method steps or certain properties or features, or a process / method that performs less than about 5, 1, 0.5, or 0.1 percent of certain process / method steps compared to a given standard process / method, or a property or feature characterized by having less than about 5, 1, 0.5, or 0.1 percent of the properties or features compared to a given standard.

[0213] The term “exemplary” is used herein to mean “provided as an example, case, or illustration.” No embodiment described as “exemplary” should necessarily be construed as being preferable or advantageous to other embodiments, and / or should not be construed as precluding the borrowing of features from other embodiments.

[0214] The words “optionally” or “alternatively” are used herein to mean “provided in some embodiments and not provided in other embodiments.” Any detailed embodiments of the present invention may include several “optional” features, provided that such features do not conflict.

[0215] As used herein, the singular forms "a," "an," and "the" include plural referents unless otherwise specifically indicated by the context. For example, the terms "a compound" or "at least one compound" may include multiple compounds, including mixtures thereof.

[0216] Throughout this application, various embodiments of the present invention may be presented in range form. It should be understood that descriptions in range form are merely for convenience and conciseness, and should not be interpreted as inflexible limitations on the scope of the invention. Accordingly, range descriptions should be considered as specifically disclosing all possible partial ranges and the individual numbers within those ranges. For example, a range description such as 1-6 should be considered as specifically disclosing partial ranges such as 1-3, 1-4, 1-5, 2-4, 2-6, 3-6, and the individual numbers within those ranges, such as 1, 2, 3, 4, 5, and 6. This applies regardless of the width of the range.

[0217] Wherever a numerical range is indicated herein, it is meant to include any cited number (fraction or integer) within that range. The phrases “taking the range between” and “taking the range from” the first indicated number to the second indicated number are used synonymously herein and mean including the first and second indicated numbers, as well as all fractions and integers between them.

[0218] As used herein, the terms “process” and “method” mean a set of sets

[0219] As used herein, the term “treat” includes resolving, substantially inhibiting, slowing or improving the progression of a condition, substantially improving the clinical or aesthetic symptoms of a condition, or substantially preventing the appearance of the clinical or aesthetic symptoms of a condition.

[0220] When a detailed sequence listing is referenced, such reference should be understood to include sequences substantially corresponding to its complementary sequence, such as minor sequence variations resulting from sequencing errors, cloning errors, or other changes that cause base substitutions, base deletions, or base additions, provided that the frequency of such variations is less than one per 50 nucleotides, alternatively less than one per 100 nucleotides, alternatively less than one per 200 nucleotides, alternatively less than one per 500 nucleotides, alternatively less than one per 1,000 nucleotides, alternatively less than one per 5,000 nucleotides, and alternatively less than one per 10,000 nucleotides.

[0221] For clarity, it should be understood that certain features of the present invention described in relation to separate embodiments may also be provided in combination in a single embodiment. Conversely, for brevity, various features of the present invention described in relation to a single embodiment may also be provided separately, in any preferred partial combination, or as preferred in any other described embodiment of the present invention. Certain features described in relation to various embodiments should not be considered essential features of those embodiments unless the embodiment becomes unimplementable without those elements.

[0222] Experimental and / or computed support can be found in the following examples for various embodiments and aspects of the present invention as explicitly stated above and asserted in the following claims section. [Examples]

[0223] Herein, we refer to the following embodiments, which, in conjunction with the above description, illustrate some embodiments of the present invention in a non-limiting manner.

[0224] Example 1 Chemical total synthesis of Pfu DNA polymerase The chemical total synthesis of both natural (L-amino acid protein) and enantiomer Pfu DNA polymerases demonstrated the concept of some embodiments of the present invention.

[0225] The first step in implementing the method provided in this application was to identify existing sequence features that would lead to the total chemical synthesis of the enzyme, using available information on Pfu DNA polymerase, and to identify a position in the sequence that possesses sufficient structural flexibility (looseness) to allow the introduction of mutations without impairing structural stability and therefore the desired activity of the enzyme. For this purpose, multiple sequence alignment (MSA) was performed using Pfu-WT (SEQ ID NO: 47), Pfu-5m (SEQ ID NO: 48), Pfu-5m-55I (SEQ ID NO: 49), Pfu-5m-46I (SEQ ID NO: 50), Pfu-5m-30I (SEQ ID NO: 51), Pfu-5m-0I (SEQ ID NO: 52), KOD1 (SEQ ID NO: 53), Tgo (SEQ ID NO: 54), 9°N-7 (SEQ ID NO: 55), and Tok (SEQ ID NO: 56) polymerases. MSA revealed highly conserved amino acids, which were left unchanged, while other parts of the MSA showed diversity leading to mutations that introduce additional NCL sites, cleavage sites, hydrophobicity-reducing mutations, and Ile-reducing mutations. Thus, based on MSA, E102A, E276A, K317G, V367L, and I540A were selected as mutations to introduce ligation-inducible amino acids into diverse amino acid sections of the sequence (and to substitute isoleucine at position 540). Based on MSA analysis and protein structure information, the isoleucine WT residues I38, I62, I65, I80, I127, I137, I158, I171, I176, I191, I197, I198, I205, I206, I228, I232, I244, I256, I264, I268, I282, I331, I401, I434, I446, I478, I557, I598, I605, I611, I619, I631, I643, I648, I656, I677, I716, I734, I745, and I772 were replaced with other suitable residues. In addition, to convert Pfu DNA polymerase into an efficient RNA polymerase in both L-amino acid and D-amino acid forms, we introduced the V93Q, D141A, E143A, Y410G, A486L, and E665K mutations.

[0226] The amino acid sequence of Pfu DNA polymerase was divided into two domain-forming segments, referred to in this application as the Pfu-N fragment (SEQ ID NO: 57) and the Pfu-C fragment (SEQ ID NO: 67), according to some embodiments of the present invention. As shown in Figures 2A to 2B below, the Pfu-N fragment was divided into nine peptide segments (SEQ ID NOs: 58 to 66) in the range of 40 to 62 aa in length, and the Pfu-C fragment was divided into six segments (SEQ ID NOs: 68 to 73) in the range of 33 to 63 aa.

[0227] Figures 2A and 2B show the design flow chart for the synthesis pathway of the mutant Pfu-N fragment, which has additional NCL sites (E102A, E276A, K317G, V367L) to form a ligation-inducible segment and has 25 isoleucine residues substituted (Figure 2A), and the design flow chart for the synthesis pathway of the mutant Pfu-C fragment, which has an additional NCL site (I540A) and mutations in 15 other isoleucine residues (Figure 2B). On the other hand, introducing these mutations facilitates protein synthesis and ligation processes in SPPS and reduces the cost of mirror image synthesis.

[0228] Peptide segments were prepared by Fmoc-based SPPS, purified by reverse-phase high-performance liquid chromatography (RP-HPLC), assembled by hydrazide-based NCL using a convergent assembly strategy, and subsequently desulfurized using a metal-free radical-based method. For L-polymerase, 4.3 mg of L-Pfu-N fragment was obtained with a measured molecular weight (MW) of 54830.0 Da (calculated MW of 54829.9 Da, determined by analytical HPLC and ESI-MS, not shown), and 2.2 mg of L-Pfu-C fragment was obtained with a measured MW of 35563.2 Da (calculated MW of 35563.02 Da). For D-polymerase, 16.5 mg of D-Pfu-N fragment was obtained with a measured MW of 54829.5 Da, and 11.9 mg of D-Pfu-C fragment was obtained with a measured MW of 35561.9 Da. Both synthetic L-polymerase and D-polymerase were folded by continuous dialysis followed by heat precipitation at 85°C, and this heat precipitation further improved the purity of the correctly folded protein (ESI-MS, not shown). Next, the PCR activity of these polymerases was tested with a short 100 bp synthetic D- or L-DNA template (SEQ ID NO: 12), and comparable amplification efficiencies were measured between recombinant and synthetic L-polymerase and D-polymerase (analysis by 3% sieving agarose gel electrophoresis, stained with ExRed.M, DNA marker, ImageLab software (Bio-Rad Laboratories, California, USA). M is the DNA marker). The fidelity of synthetic L-polymerase was also quantified with a 1.2 kb D-DNA sequence from the pUC19 plasmid (SEQ ID NO: 80), and Sanger sequencing of the PCR product yielded 3.6 × 10⁻¹⁶. -6 An error rate of less than 1% was measured (see Table 3 below), which is consistent with that of WT Pfu DNA polymerase reported in previous studies.

[0229] [Table 3]

[0230] material: L-DNA oligos were synthesized using an H-8 oligo synthesizer (K&A Laborgeraete, Germany) with L-deoxynucleoside phosphoramidite (ChemGenes, Massachusetts, USA). Primers for recombinant protein expression were ordered from Genewiz (Beijing, China). Primers for bacterial 16S rRNA gene assembly were purified by denaturing sequencing (PAGE). Other DNA oligos were purified using oligonucleotide purification cartridges (OPCs) (Ruibiotech, Beijing, China). The PAGE DNA purification kit was purchased from Tiandz Inc. (Beijing, China). Tris bases, NP-40, Tween-20, KCl, guanidine hydrochloride (Gn·HCl), and β-mercaptoethanol (β-ME) were purchased from Amresco Inc. (Pennsylvania, USA). Imidazole and EDTA were purchased from Solarbio Life Sciences (Beijing, China). 2-chlorotrityl chloride resin (loading = 0.6 mmol / g) was purchased from Tianjin Nankai Hecheng Science & Technology Co. (Tianjin, China). Wang Chemmatrix resin was purchased from CSBio Ltd (Shanghai, China). Fmoc-D-amino acids, Fmoc-L-amino acids, and O-(6-chlorobenzotriazol-1-yl)-N,N,N',N'-tetramethyluronium hexafluorophosphate (HCTU) were purchased from GL Biochem Co. (Shanghai, China). N,N-diisopropylethylamine (DIEA), trifluoroacetic acid (TFA), N,N-dimethylformamide (DMF), thioanisole, triisopropylsilane (TIPS), 1,2-ethanedithiol (EDT), palladium chloride (PdCl2), sodium 2-mercaptoethanesulfonate (MESNa), and 2,2'-azobis[2-(2-imidazolin-2-yl)propane] dihydrochloride (VA-044) were purchased from J&K Scientific (Beijing, China). 4-mercaptophenylacetic acid (MPAA) was purchased from Alfa Aesar Chemicals Co. (Shanghai, China).Piperidine, Na2HPO4·12H2O, NaH2PO4·2H2O, sodium nitrite (NaNO2), and acetic anhydride were purchased from Sinopharm Chemical Reagent Co. (Shanghai, China). NaCl, NaOH, and hydrochloric acid were purchased from Sinopharm Chemical Reagent (Beijing, China). Dichloromethane (DCM) was purchased from Shanghai Titan Scientific Co. (Shanghai, China). Tris(2-carboxyethyl)phosphine hydrochloride (TCEP·HCl), 9-fluorenylmethyl carbacate (Fmoc-NHNH2), ethyl-2-oxime cyanoglycylate (Oxyma), N,N'-diisopropylcarbodiimide (DIC), and DL-1,4-dithiothreitol (DTT) were purchased from Adamas Reagent Co. (Shanghai, China). Reduced glutathione (GSH) was purchased from Acros Organics (New Jersey, USA). Anhydrous ether was purchased from Beijing Tongguang Fine Chemicals Company (Beijing, China). Acetonitrile (HPLC grade) was purchased from JTBaker (New Jersey, USA).

[0231] Fmoc-based solid-phase peptide synthesis (Fmoc-SPPS): All peptides were synthesized using Fmoc-based SPPS on a Liberty Blue automated microwave peptide synthesizer (CEM Corporation, North Carolina, USA) and a Prelude X automated peptide synthesizer (Protein Technologies Inc., Arizona, USA). Peptides with C-terminal carboxylic acids, such as Pfu-N-9 and Pfu-C-6, were synthesized on Wang Chemmatrix resin (CSBio Ltd, Shanghai, China) with the first C-terminal residue pre-loaded. All other peptides were synthesized on Fmoc-hydrazine 2-chlorotrityl chloride resin to prepare peptide hydrazides. For each peptide acid, the first residue was manually attached to the Wang Chemmatrix resin by double coupling: In the first coupling reaction, the amino acids were coupled at 30°C for 1 hour using 4 equivalents of amino acid, 3.8 equivalents of HCTU, and 8 equivalents of DIEA. The resin was washed with DMF and DCM, and without deprotection, the second coupling reaction was carried out overnight at 25°C with 4 equivalents of amino acid, 4 equivalents of Oxyma, and 4 equivalents of DIC. All resins were used after swelling in DMF for 5-10 minutes. Both the Fmoc groups in the resin and the assembled amino acids were removed by treatment with 20% piperidine and 0.1 mol / L of Oxyma in DMF at 85°C. Coupling of amino acids other than Fmoc-Cys(Trt)-OH and Fmoc-His(Trt)-OH was carried out at 85°C using 4 equivalents of amino acid, 4 equivalents of Oxyma, and 8 equivalents of DIC. The coupling reaction of Fmoc-Cys(Trt)-OH and Fmoc-His(Trt)-OH was carried out at 50°C for 10 minutes to avoid side reactions at high temperatures. Trifluoroacetylthiazolidine-4-carboxylic acid-OH (Tfa-Thz-OH) was coupled at room temperature using Oxyma / DIC activation. After completion of peptide chain assembly, the peptide was cleaved from the resin using H2O / thioanisole / triisopropylsilane / 1,2-ethanedithiol / trifluoroacetic acid (0.5 / 0.5 / 0.5 / 0.25 / 8.25). The cleavage reaction took 2.5 hours at 27°C with stirring.Most of the TFA in the mixture was removed by N2 blowing, and the crude peptide was precipitated by adding cold ether. After centrifugation, the supernatant was discarded, and the precipitate was washed twice with ether. The crude peptide was dissolved in CH3CN / H2O and analyzed by RP-HPLC and ESI-MS, and purified by semi-preparative HPLC.

[0232] Native Chemical Ligation (NCL): The C-terminal peptide hydrazide segment was dissolved in acidified ligation buffer (6 M Gn·HCl and 0.1 M NaH2PO4 aqueous solution, pH 3.0). This mixture was cooled in an ice salt bath (-10°C), and 10 equivalents of NaNO2 were added to the acidified ligation buffer (pH 3.0). The activation reaction system was placed in the ice salt bath with stirring for 25 minutes, after which 40 equivalents of MPAA and 1 equivalent of the N-terminal cysteine ​​peptide were added to the ligation buffer, and the pH of the solution was adjusted to 6.5 at room temperature. After the reaction overnight, the reaction system was diluted 2-fold by adding 150 mM TCEP in ligation buffer (pH adjusted to 7.0), and this reaction system was placed at room temperature with stirring for 30 minutes. Finally, the ligation product was analyzed by HPLC and ESI-MS and purified by semi-preparative HPLC. Notably, during the ligation of the Pfu-C-1 and Pfu-C-2 segments, it was found that the ligation was extremely inefficient due to the insoluble Pfu-C-2 segment. Therefore, increasing the initial concentration of Gn·HCl to 8M (final Gn·HCl concentration was approximately 7M) significantly improved the solubility and ligation efficiency of these two peptide segments.

[0233] Desulfurization: Cys-containing peptide (3 mg / ml) was dissolved in a desulfurization buffer (0.1 M phosphate-buffered aqueous solution containing 6 M Gn·HCl, 200 mM TCEP, 40 mM reduced L-glutathione, and 20 mM VA-044, pH 6.8). This mixture was left to stand overnight at 37°C with stirring, and the desulfurization product was analyzed by HPLC and ESI-MS, and purified by semi-preparative HPLC.

[0234] Acm deprotection: The acetamidomethyl (Acm) group was removed using a Pd-assisted deprotection strategy. The Acm-protected peptide was dissolved in Acm deprotection buffer (6M Gn·HCl, 0.1M phosphate, and 40mM TCEP aqueous solution, pH 7.0) to a final concentration of 1mM, and then 20 equivalents of PdCl2 were added. This reaction mixture was incubated overnight at 25°C with stirring. The reaction was quenched by adding DTT to a final concentration of 50mM. This reaction mixture was left to stand for 1 hour with stirring and purified by semi-preparative HPLC.

[0235] Folding of fission-type Pfu DNA polymerase in vitro: Lyophilized N and C fragments of Pfu DNA polymerase were dissolved in 4M and 5M Gn·HCl, respectively, containing 10 mM β-ME. After mixing the two fragments at equal concentrations (0.5 μM), in vitro protein folding was performed by dialyzing overnight at 4°C in a buffer containing 40 mM Tris-HCl (pH 7.5), 1 mM EDTA, 100 mM KCl, and 10% glycerol. The folded Pfu DNA polymerase was heated at 85°C for 15 minutes to precipitate the thermally unstable peptide, which was then removed by centrifugation at 20,000 × g for 40 minutes at 4°C. The supernatant was concentrated and dialyzed with a storage buffer consisting of 100 mM Tris-HCl (pH 8.0), 50% glycerol, 0.2 mM EDTA, 0.2% NP-40 nonionic surfactant, 0.2% Tween 20, and 2 mM DTT.

[0236] RP-HPLC and ESI-MS: All RP-HPLC analysis and purification were performed using a Shimadzu Prominence HPLC system (Shimadzu Corporation, Kyoto, Japan) equipped with an SPD-20A UV-Vis detector and an LC-20AT solvent delivery unit. For the analysis, the ligation reaction was monitored using an Ultimate XB-C4 column (5 μm, 4.6 × 250 mm) (Welch Materials, Shanghai, China) at a flow rate of 1 ml / min, and the purity of the peptide product was analyzed. Crude peptides and ligation products were separated using Ultimate XB-C4 and C18 columns (5 μm, 21.2 × 250 mm or 5 μm, 10 × 250 mm) (Welch Materials, Shanghai, China) at flow rates of 4–8 ml / min, respectively. The purified products were characterized by ESI-MS using a Shimadzu LC / MS-2020 system (Shimadzu Corporation, Kyoto, Japan).

[0237] Protein expression and purification: The Pfu DNA polymerase gene was cloned into the pET-28c plasmid, and mutants were constructed using the pEASY-Uni seamless cloning and assembly kit (TransGen Biotech, Beijing, China). Proteins fused to the N-terminal His6 tag were expressed using E. coli strain BL21(DE3) in LB medium. Induced cells were harvested and resuspended in lysis buffer (40 mM Tris-HCl, 300 mM NaCl, 10 mM imidazole, 10 mM β-ME, 10 mg / ml lysozyme, pH 8.0). Cell lysates were heated at 85°C for 15 minutes, followed by removal of heat-unstable proteins by centrifugation at 20,000 × g for 40 minutes at 4°C. The supernatant was incubated in Ni-NTA Superflow resin (Senhui Microsphere Tech, Suzhou, China) at 4°C for 1 hour. The resin was washed with a buffer containing 40 mM Tris-HCl (pH 8.0), 300 mM NaCl, 40 mM imidazole, and 10 mM β-ME, and then eluted with a buffer containing 40 mM Tris-HCl (pH 8.0), 300 mM NaCl, 250 mM imidazole, and 10 mM β-ME. The purified and concentrated Pfu DNA polymerase and mutants were dialyzed with a storage buffer containing 100 mM Tris-HCl (pH 8.0), 50% glycerol, 0.2 mM EDTA, 0.2% NP-40 nonionic surfactant, 0.2% Tween 20, and 2 mM DTT.

[0238] PCR activity and fidelity: Native and enantiomer PCR reactions were performed in a 50 μl reaction system containing 1×Pfu buffer (Solarbio Life Sciences, Beijing, China) along with 200 μM of each dNTP, 0.2 μM of each primer, template, and polymerase. To quantify the PCR activity of Pfu DNA polymerase and its mutants, the polymerase was adjusted to the same concentration as wild-type (WT) Pfu DNA polymerase by 12% SDS-PAGE. SDS-PAGE analysis confirmed that the molecular weight of a fragment of recombinant fission mutant Pfu DNA polymerase expressed and purified from E. coli was similar to that of the naturally synthesized native and enantiomer Pfu DNA polymerases of the same sequence (results not shown). The PCR program settings were 3 minutes at 94°C (initial denaturation), 30 seconds at 94°C, 30 seconds at 50–65°C (Tm-dependent), and 1–7 minutes at 72°C (amplicon length-dependent) for 10–35 cycles, followed by 10 minutes at 72°C (final extension). To quantify the amplification efficiency of synthetic Pfu DNA polymerase, a 100 bp DNA sequence was used as a template. PCR amplification using recombinant, synthetic L-, and synthetic D-Pfu DNA polymerase (fission-Pfu-5m-30I) was analyzed by 3% sieved agarose gel electrophoresis and stained with ExRed (results not shown). The PCR amplification efficiency of synthetic D-Pfu DNA polymerase was estimated to be approximately 1.5 based on the intensity of the product band. The amplification products from the first 9 cycles were analyzed using ImageJ software (Bio-Rad Laboratories, California, USA). To investigate the fidelity of synthetic Pfu DNA polymerase, the product after 45 cycles of natural PCR (1.2kb D-DNA) was purified using the V-elute gel mini purification kit (Beijing Zoman Biotech, China), cloned for Sanger sequencing using the zero-background ZT4 Simple-Blunt Fast Clone Kit (Beijing Zoman Biotech, China), and calculated according to the method described above.

[0239] Example 2 Chemical total synthesis of T7 RNA polymerase and its use As discussed above, synthesizing a mirror-image version of RNA polymerase using a double-stranded (ds) L-DNA template could enable the enzymatic transcription of any mirror-image rRNA and the mRNA required for mirror-image translation. Therefore, as another step in the proof of concept of some aspects of the present invention, both the native (L-amino acid protein) and mirror-image versions of 100 kDa T7 RNA polymerase were chemically synthesized, with a binary site design.

[0240] T7 RNA polymerase has known cleavage morphologies; for example, Segall-Shapiro et al. [Mol Syst Biol., 2014, 30(10), pp. 742] identified several cleavage sites in T7 RNA polymerase using a transposon-based method. Tiyun Han et al. [ACS Synth Biol., 2017, 6(2), pp. 357-366.] designed photoactivatable gene switches based on cleavage T7 RNA polymerase to implement photoactivatable gene expression under different conditions. However, the fission sites used in these natural enzymes are not always suitable for the chemical synthesis of T7 RNA polymerase: some fission sites of T7 RNA polymerase significantly alter its enzymatic activity, and some are located near the N-terminus or C-terminus of the protein peptide chain, resulting in the formation of one or more large protein fragments (over 400-500 aa), which would still be too large for chemical synthesis.

[0241] To obtain a practical domain-forming segment, a second previously unproposed cleavage site, namely between K363 and P364, was identified in some embodiments of the present invention using criteria of low sequence conservation and structural flexibility. The cleavage site between N601 and T602 reported by Segall-Shapiro et al., combined with the cleavage site in the solvent-exposed loop of the T7 RNA polymerase structure (between K363 and P364) discovered during the implementation of the present invention, split the polymerase into three fragments of approximately equal length (typically less than 400-500 aa) suitable for chemosynthesis, without significantly altering enzyme activity and fidelity: a 369 aa T7-cleavage-N fragment (with a His6 tag attached to the N-terminus), a 238 aa T7-cleavage-M fragment, and a 282 aa T7-cleavage-C fragment. The aforementioned cleavage sites can be selected to be in the vicinity of the aforementioned sites in the same loop, i.e., positions 357–366 and / or 564–607. Simultaneously, cleavage T7 RNA polymerase can be used as transcriptional AND logic. For example, with an engineering strategy of splitting a protein into fragments and regulating its recombination using a regulatory domain, a gene switch can be obtained in which the activity of T7 RNA polymerase is directly controlled by an external signal. A robust switchable system with excellent light-shielded off / light-in-light characteristics can be obtained using a photoactivatable VVD domain and its variants as the regulatory domain.

[0242] Furthermore, a systematic isoleucine substitution method was also performed based on multiple sequence alignment (MSA) and structural information using T7-WT (SEQ ID NO: 82), T7-37I (SEQ ID NO: 83), YenP (SEQ ID NO: 84), phiEap (SEQ ID NO: 85), and KpnP (SEQ ID NO: 86) polymerases. This method mutated several isoleucines (14 out of 51 or 27% of Ile residues) in the T7 RNA polymerase with other amino acids such as valine, leucine, and methionine (I6V, I14L, I74V, I82V, I109V, I117L, I141V, I210M, I244L, I281V, I320V, I322L, I330V, I367L). This method reduced the amino acid cost required to synthesize this D-polymerase. This will facilitate its large-scale synthesis and practical applications in the future.

[0243] Figures 3A to 3C present design flow charts of the synthetic pathways for the 369aa mutant T7-fission-N fragment (SEQ ID NO: 87) (Figure 3A), the 238aa mutant T7-fission-M fragment (SEQ ID NO: 94) (Figure 3B), and the 282aa mutant T7-fission-C fragment (SEQ ID NO: 101) (Figure 3C), including isoleucine residue substitutions, a novel NCL, and a novel fission site between K363 and P364 (these were introduced to facilitate protein synthesis and ligation processes in SPPS and reduce the cost of synthesizing mirror image versions).

[0244] Further chemical synthesis of T7 RNA polymerase was performed by introducing ligation-inducible residue substitutions. The T7-fission-N fragment was divided into seven peptide segments ranging in length from 32 to 76 aa (SEQ ID NOs. 88-94), the T7-fission-M fragment into six peptide segments ranging in length from 23 to 45 aa (SEQ ID NOs. 96-101), and the T7-fission-C fragment into five peptide segments ranging in length from 41 to 75 aa (SEQ ID NOs. 103-107). These peptide segments were prepared by Fmoc-based SPPS, purified by reverse-phase high-performance liquid chromatography (RP-HPLC), assembled by hydrazide-based NCL using a convergent assembly strategy, and subsequently desulfurized using a metal-free radical-based method. After synthesis, ligation, purification, and freeze-drying, approximately 3 mg of L-polymerase was obtained as a T7-fission-N fragment with an measured molecular weight (MW) of 41369.0 Da (calculated MW of 41372.6 Da), approximately 2.5 mg of T7-fission-M fragment with an MW of 26786.0 Da (calculated MW of 26787.4 Da), and approximately 4.8 mg of T7-fission-C fragment with an MW of 31459.0 Da (calculated MW of 31459.9 Da). For D-polymerase, approximately 9 mg of D-T7-fission-N fragment with an measured molecular weight (MW) of 41373.0 Da, approximately 8 mg of T7-fission-M fragment with an MW of 26787.0 Da, and approximately 15 mg of T7-fission-C fragment with an MW of 31459.0 Da.

[0245] Folding of synthetic polymerase in vitro: The synthetic polymerase was folded by precipitating impurities through continuous dialysis followed by ultrafiltration.

[0246] Lyophilized synthetic N, M, and C fragments of T7 RNA polymerase were dissolved in denaturation buffers containing 6 M Gn·HCl and 20 mM DTT, respectively. The N, M, and C fragments were mixed equally (0.5 nmol / ml), and protein folding was performed by dialyzing with gentle agitation at 4°C for 24 hours in a reconstitution buffer (50 mM Tris-HCl, 100 mM KCl, 10% glycerol, 1 mM EDTA, 10 mM DTT, pH 8.0). After reconstitution, the enzyme was dialyzed at 4°C for 12 hours with gentle agitation in a storage buffer containing 50% glycerol, 50 mM Tris-HCl (pH 8.0), 100 mM NaCl, 1 mM EDTA, 0.1% Triton X-100, and 10 mM DTT, followed by ultrafiltration using an Amicon Utra centrifugal filter (0.5 ml, 100,000 MWCO).

[0247] Transcriptional activity and fidelity of synthetic T7 RNA polymerase: Native and mirror-image transcription were carried out in a 10 μl reaction system containing 1× T7 reaction buffer (New England Biolabs, Beijing, China), 500 μM each of rNTPs, 10% DMSO, 5 mM DTT, template, and polymerase. To quantify the transcriptional activities of T7 RNA polymerase and its mutants, the polymerase was adjusted to the same concentration as wild-type (WT) T7 RNA polymerase by 12% SDS-PAGE (results not shown). The reaction mixture was incubated at 37 °C for various times. According to what was shown by the transcriptional activities of native and mirror-image T7 RNA polymerase, this polymerase can successfully transcribe a 160 bp DNA template (SEQ ID NO: 108) and a 1.5 kb DNA template (SEQ ID NO: 109), and it is pointed out that a wide range of lengths of L-RNA molecules can be produced from a 1.5 kb L-DNA template by synthetic mirror-image T7 RNA polymerase (results not shown). A mixture of purified and concentration-determined single-stranded L-RNA transcripts of different lengths can be used as an RNA marker (or RNA ladder) during size processing and quantification of RNA on native or denaturing gels, which is superior to commercially available D-RNA markers (D-RNA ladders) due to its native RNase resistance. The fidelity of synthetic T7 RNA polymerase was also examined by reverse transcription of DNase I-digested transcripts with Superscript IV high-fidelity reverse transcriptase, followed by PCR amplification with high-fidelity Pfu DNA polymerase and sequencing of the amplicon by Sanger sequencing, and an error rate (10 -6 orders of magnitude) consistent with the error rate of WT T7 RNA polymerase reported in previous studies was measured.

[0248] L-tRNA Ser Charged: L-tDNA Ser (SEQ ID NO: 110) was assembled by mutant mirror-image Dpo4 (D-Dpo4-5m). L-tRNA SerThe L-tRNA was transcribed using high-fidelity enantiomer T7 RNA polymerase, and the reaction system containing 1×T7 reaction buffer A (40 mM Tris-HCl, 25 mM MgCl2, 1 mM spermidine, 2 mM DTT, pH 8.0) along with 2 mM of each L-rNTP, 10% DMSO, 0.3 μM template, and 2 μM polymerase was incubated overnight at 37°C. The product was purified at single-nucleotide resolution by denaturing PAGE, and the purified product was analyzed by 10% denaturing PAGE (results not shown). 25 mM HEPES-KOH (pH 7.5), 50 mM KCl, 2 μM L-tRNA Ser and L-tRNA in 10 μM L-dFx Ser A charge was performed. The reaction system was heated at 95°C for 2 minutes and slowly cooled to room temperature for annealing. Next, 100 mM MgCl2 was added to the system, and the system was incubated at room temperature for 10 minutes, then at 4°C for 10 minutes. Finally, 5 mM D-Ser-DBE was added to the system, and the system was incubated at 4°C for 6 hours. Ethanol precipitation was performed by adding 1 / 10 volume of 3 M NaOAc and 2.5 volume of ethanol, and the system was incubated overnight at -20°C. The product was analyzed by 8% acidic PAGE (results are not shown).

[0249] L-16S rRNA purification: L-16S rDNA (SEQ ID NO: 109) was assembled using high-fidelity enantiomer Pfu DNA polymerase. L-16S rRNA was transcribed using high-fidelity enantiomer T7 RNA polymerase, and a reaction system containing 500 μM of each L-rNTP, 10% DMSO, 5 mM DTT, template, and polymerase in 1×T7 reaction buffer (New England Biolabs, Beijing, China) was incubated overnight at 37°C. The transcript was purified by β-agarase digestion from a 2% low-melting-point agarose gel (Amersco, USA). Gel sections containing RNA samples were equilibrated in 10 volumes of 1×β-agarase buffer at room temperature for 60 minutes, then thawed at 70°C for 15 minutes, and cooled to 45°C. The molten agarose solution was incubated with 2 units of β-agarase (New England Biolabs, Beijing, China) at 45°C for 60 minutes, followed by incubation at -20°C for 15 minutes and centrifugation at 4°C for 15 minutes. The supernatant was transferred to a new microcentrifuge tube and precipitated with ethanol by adding 1 / 10 volume of 3M NaOAc and 2.5 volume of ethanol, and incubated overnight at -20°C. The purified product was analyzed by 3% agarose gel (results not shown).

[0250] L-guanine sensor: The molecular discriminative ability of guanine sensors was demonstrated by tracking the specificity of D- and L-guanine sensors transcribed by synthetic L- and D-T7 RNA polymerases. An L-guanine sensor DNA template (SEQ ID NO: 111) was assembled using D-Dpo4-5m. The L-guanine sensor was transcribed using high-fidelity enantiomer T7 RNA polymerase, and a reaction system containing 1×T7 reaction buffer A (40 mM Tris-HCl, 25 mM MgCl2, 1 mM spermidine, 2 mM DTT, pH 8.0) along with 2 mM each of the L-rNTPs, 10% DMSO, 0.2 μM template, and 2 μM polymerase was incubated overnight at 37°C. The product was purified using a polyacrylamide gel in 8 M urea, and the purified product was analyzed by 10% denatured PAGE (results not shown). A 1 μM L-guanine sensor and 10 μM DFHBI were incubated at 37°C in a buffer containing 40 mM HEPES (pH 7.4), 125 mM KCl, and 1 mM MgCl2. Next, 1 mM guanine was rapidly added to this solution, and fluorescence emission was recorded for 15 minutes at 37°C under continuous illumination using the following instrument parameters: excitation wavelength 460 nm, emission wavelength 500 nm, and slit width 12 nm. 0.1 μM RNA and 10 μM DFHBI were incubated with 100 μM guanine or a competing molecule and assayed for fluorescence emission at 500 nm. The guanine sensor saturated with 100 μM guanine and showed high molecular discrimination ability for GTP and adenine at the same concentration (results not shown).

[0251] L-38-6 RNA polymerization reaction: DNA templates for L-38-6 ribozyme (SEQ ID NO: 112) and L-class I ligase DNA templates (SEQ ID NO: 113) were assembled using D-Dpo4-5m. RNA was transcribed using high-fidelity enantiomer T7 RNA polymerase, and a reaction system containing 1×T7 reaction buffer A (40 mM Tris-HCl, 25 mM MgCl2, 1 mM spermidine, 2 mM DTT, pH 8.0) along with 2 mM of each L-rNTP, 10% DMSO, 0.3 μM of the template, and 2 μM of polymerase was incubated overnight at 37°C. The product was purified by polyacrylamide gel in 8 M urea (results not shown). For the RNA polymerization reaction, 100 nM L-38-6 ribozyme (SEQ ID NO: 114), 80 nM L-5'-FAM labeled primer (SEQ ID NO: 115), and 100 nM L-class I ligase template (SEQ ID NO: 116) were used. To anneal the RNA, the mixture was first heated at 80°C for 30 seconds, then slowly cooled to 17°C, and then added to a reaction mixture containing 4 mM L-rNTP, 200 mM MgCl2, 25 mM Tris-HCl, pH 8.3, and 0.05% Tween-20, which was incubated at 17°C for varying times. The product was concentrated using an ssDNA / RNA Clean & Concentrator kit (ZYMO RESEARCH, California, USA), then mixed with denaturation buffer (98% formamide, 0.25 mM EDTA), followed by heating to 65°C for 10 minutes, and then rapidly placed on ice. This sample was separated using a 10% polyacrylamide gel in 8M urea and scanned using a Typhoon Trio+ system operating in Cy2 mode.

[0252] Reaction rates of RNA degradation in natural and enantiomer 16S rRNA: To evaluate the integrity of RNA under controlled conditions, three prepared transcripts containing native 16S rRNA, native 16S rRNA with RNase inhibitor, and mirror-image 16S rRNA were detected and analyzed by the Bioanalyzer method. Native and mirror-image 16S rRNA were transcribed by native and mirror-image T7 RNA polymerases, respectively, and purified from a 2% low-melting-point agarose gel by β-agarase I digestion. The purified RNA was placed at 37 °C for 5 minutes, 30 minutes, 1 hour, 2 hours, 4 hours, 8 hours, 18 hours, 24 hours, 48 hours, 72 hours, 7 days, 15 days, 30 days, 60 days, and 100 days, and the quality of the RNA was determined based on the electrophoretogram of microchip gel electrophoresis. For native 16S rRNA, when placed at 37 °C for 30 minutes, only minimal signs of degradation were observed, and degradation became more evident at 1 hour, with a substantial increase in the baseline. After 6 hours at 37 °C, the peak completely disappeared due to the progression of degradation. In the sample of native 16S rRNA with RNase inhibitor, when placed at 37 °C for 4 hours, only minimal signs of degradation were observed, and RNA degradation became more evident at 8 hours, with a substantial increase in the baseline. After 48 hours at 37 °C, the peak completely disappeared due to the progression of degradation. In the sample of mirror-image 16S rRNA, no signs of degradation were detected even when placed at 37 °C for 15 days. This indicates that RNA has greater stability under conditions where RNase is completely removed. Measuring the hydrolysis reaction rate of RNA under different conditions using the L-RNA system can provide a control for evaluating the effectiveness of RNase inhibitor reagents.

[0253] Example 3 Mirror-image DNA information storage After obtaining highly faithful mirror-image Pfu DNA polymerase, a proof-of-concept of mirror-image DNA information storage according to some embodiments of the present invention was conducted by exploring its application in mirror-image DNA information storage through accurate writing and reading of L-DNA sequences.

[0254] Encode the following passage from a 1860 publication by Louis Pasteur, from whom the concepts of mirror-image molecules and mirror-image biological systems were first proposed, into DNA sequences (see Table 4), and archive them into 11 L-DNA segments, 220 bp in length, assembled from four short synthetic L-DNA oligos, 70 - 90 nt each (Table 5). Pasteur: “And consequently, if the mysterious influence to which the asymmetry of natural products is due should change its sense or direction, the constitutive elements of all living beings would assume the opposite asymmetry. Perhaps a new world would present itself to our view. Who could foresee the organisation of living things if cellulose, right as it is, became left; if the albumen of the blood, now left, became right? These are mysteries which furnish much work for the future, and demand henceforth the most serious consideration from science.”

[0255]

Table 4

[0256] An L-DNA storage library (L-library) containing all 11 segments and a 220 bp double-stranded L-DNA segment for information storage, assembled using enantiomer Pfu DNA polymerase via enantiomer assembly PCR from four short 70-90 nt synthetic L-DNA oligos, was analyzed by 2.5% agarose gel electrophoresis and stained with ExRed.M DNA marker (results not shown), and is shown in Table 5. Table 5 shows the sequences used for L-DNA information storage; lowercase letters are M13-F and M13-R sequences for amplification, and underlined letters are unique sequences for sequencing individual segments.

[0257] [Table 5-1]

[0258] [Table 5-2]

[0259] L-DNA reading can be achieved by sequencing-bi-synthesis using a phosphorothioate method (with cleavage using L-deoxynucleoside α-thiotriphosphate (L-dNTPαS) and 2-iodoethanol) with enantiomer Pfu DNA polymerase, or by a linkage termination method using L-dideoxynucleoside triphosphate (L-ddNTP) with mutant enantiomer Pfu DNA polymerase. Bidirectional sequencing was also applied, and using 5'-labeled primers with two different dyes (FAM and Cy5, respectively), the maximum read length in a single reaction by denatured polyacrylamide gel electrophoresis (PAGE, PCR amplification) improved to approximately 180 bp. The 203bp sequence of the information-carrying L-DNA in the storage medium was amplified using D-Dpo4-5m segment-specific sequencing primers from an L-DNA storage library treated with DNase I, analyzed by 2.5% agarose gel electrophoresis, stained with ExRed.M, a DNA marker (results not shown), and encoded digital data was obtained by sequencing the L-DNA storage segment S1 (SEQ ID NO: 1) using enantiomer DNA polymerase via the phosphorothioate method. Specifically, the L-DNA S1 segment was specifically amplified using D-Dpo4-5m with 5'-FAM-labeled (forward) and 5'-Cy5-labeled (reverse) sequencing primers in four separate PCR reactions. During these reactions, one L-dNTP was replaced with the corresponding L-dNTPαS, and each was cleaved with 2-iodoethanol. Analysis was performed by 10% denaturation PAGE and scanned using a Typhoon Trio+ system operating in Cy2 and Cy5 modes. Sequencing chromatograms of the information-storing L-DNA segment S1 with L-dNTPαS and D-Dpo4-5m with 5'-labeled forward and reverse sequencing primers were processed using ImageJ software (results not shown). Although enantiomer Pfu DNA polymerase can amplify and sequence the L-DNA storage segment, D-Dpo4 was used in the actual experiment due to its ease of synthesis.

[0260] Chiral Steganography: Steganography is a well-known technique and science for concealing messages so that no one other than the recipient can see or even know of their existence. This is in contrast to cryptography, which conceals only the content of information, not the existence of the information itself. The L-DNA information storage system provided in this application can also be applied to secure communications through the design of a chiral steganography experiment in which a D-DNA storage library encoding a passage from Louis Pasteur's 1860 work acts as the "cover text," and an L-DNA key assists in deciphering the "stegotext" (secret message). To further disguise the secret message, a chimeric D-DNA / L-DNA key molecule (SEQ ID NO: 46) is designed to transmit either a false message "error" or a secret message "mirror" depending on the chirality of the reading. The D-DNA storage library was sequenced by Sanger sequencing to obtain the "cover text." Using natural PCR, only the D-DNA portion of the chimeric key embedded in the storage library can be amplified and sequenced, revealing the false message. On the other hand, using mirror-image PCR, the L-DNA portion of the chimeric key can be amplified and sequenced, revealing the secret message. Steganography and cryptography are two excellent techniques for protecting data secrecy. Steganography is a technique for concealing the existence of a secret message, while cryptography refers to the conventional method of converting a secret message into an unreadable format. The chiral steganography developed here, when combined with DNA cryptography, could potentially provide an additional layer of security using encrypted data.

[0261] Figure 5 presents a flowchart illustrating DNA-based steganography for carrying secret messages by embedding chimeric D-DNA / L-DNA key molecules into a seemingly ordinary D-DNA storage library, according to some embodiments of the present invention.

[0262] To demonstrate the ability of the L-DNA information storage medium to avoid biological degradation and contamination from the natural environment, freshwater samples were collected from local ponds, and a trace amount of 100 bp L-DNA barcode (SEQ ID NO: 12) (50 μg / L or 770 pM) (Table 5), encoding information about the sample collection location ("Lianchi, Beijing"), was added to the collected water samples. Notably, the L-DNA barcode on the message carrier remained stable and amplified for up to 7 months (arbitrarily selected time) and potentially beyond. In comparison, a D-DNA barcode of the same sequence and concentration became unamplified after only 1 day. Specifically, agarose gel electrophoresis was performed after amplification of the D-DNA barcode with L-Dpo4-5m after 24 hours, and after amplification of the L-DNA barcode with D-Dpo4-5m after 1 year. In a 40 ml pond water sample, PCR amplification of the D-DNA barcode was performed using L-Dpo4-5m after 24 hours, and MI-PCR amplification of the L-DNA barcode was performed using D-Dpo4-5m after 1 year in a 40 ml pond water sample. The results were analyzed by 3% sieved agarose gel electrophoresis and stained with ExRed.M, a DNA marker (results are not shown).

[0263] Furthermore, the L-DNA barcoding of microbial DNA extracted from water samples is bioorthogonal in that it can be specifically amplified by mirror-image PCR using D-polymerase and L-DNA primers, and it did not affect the results of D-DNA metagenomic microbial sequencing.

[0264] Motivated by the accurate writing and reading of L-DNA sequences, we assembled the full-length 1.5-kb mirror-image bacterial 16S rRNA gene using high-fidelity mirror-image Pfu DNA polymerase. This attempt began with testing gene assembly using synthetic L-polymerase for D-DNA using the following two-step assembly procedure: First, DNA blocks of 450 - 600 bp were assembled from short synthetic oligos of approximately 90 nt (Table 6), followed by a second step of assembling the DNA blocks into the full-length 16S rRNA gene (SEQ ID NO: 81).

[0265]

Table 6-1

[0266]

Table 6-2

[0267] Initial attempts showed that only about 40% of the assembled sequences were correct in Sanger sequencing of full-length D-DNA products (Table 3), with most errors being nucleotide deletions, which appeared to originate from -1nt and -2nt products from oligo synthesis. Therefore, modifying the oligo purification method using denaturing PAGE at single-nucleotide resolution substantially improved the quality of the synthesized oligos by removing the majority of -1nt and -2nt products. Subsequently, most deletion errors were eliminated, and approximately 90% of the final assembled sequences were correct (the remaining sequences contained only one randomly occurring mutation). Accordingly, the assembly of a full-length 1.5kb mirror image 16S rRNA gene was performed using the same oligo purification method and mirror image assembly PCR. This gene will serve as a template for enzymatic transcription into mirror image 16S rRNA, which will be crucial for constructing functional mirror image ribosomes in the future. Specifically, enantiomer 16S rRNA genes were assembled using enantiomer Pfu DNA polymerase, followed by agarose gel electrophoresis. The full-length 1.5kb enantiomer bacterial 16S rRNA gene was obtained by enantiomer assembly PCR using enantiomer Pfu DNA polymerase, analyzed by 1.5% agarose gel electrophoresis, and stained with ExRed.M, a DNA marker. (Results are not shown).

[0268] RNA polymerization using DNA templates: RNA polymerization was carried out using 1×Thermopol buffer (New England Biolabs, Massachusetts, USA), 3 mM MgSO4, 0.625 mM each of the NTPs, 0.5 μM 5'-FAM labeled DNA primer (21 nt), and 1 μM ssDNA template (41 nt), along with polymerase. Before adding the polymerase, the reaction system was heated at 94°C for 30 seconds for annealing and then slowly cooled to 4°C. The primer extension reaction was carried out at 65°C for 10 minutes. The reaction was stopped by adding a loading buffer containing 98% formamide, 0.25 mM EDTA, and 0.0125% SDS, and the product was analyzed by denaturing PAGE in 8 M urea with 20% denaturation. Specifically, RNA polymerization activity assays using DNA templates of various mutant Pfu DNA polymerases were followed by PAGE analysis. Here, a 41nt single-stranded DNA template, a 21nt 5'-FAM-labeled DNA primer, and NTPs were used, and DNA template-specific primer extension was performed with various Pfu DNA polymerase mutants by incubation at 65°C for 10 minutes. The results were then analyzed by PAGE in 8M urea at 20% (results are not shown).

[0269] L-DNA writing and reading: A passage from a 1860 publication by Louis Pasteur containing 550 characters (see the text above) was converted into a 1650-nucleotide DNA sequence (Table 4), and encoded into 11 220-bp L-DNA segments assembled from four short synthetic L-DNA oligos, each 70-90 nt long (Table 5). The assembly PCR program settings were 35 cycles of 3 minutes at 94°C (initial denaturation), 30 seconds at 94°C, 30 seconds at 55°C, and 1 minute at 72°C (depending on amplicon length), followed by 10 minutes at 72°C (final extension). For the phosphorothioate method, 5'-FAM-labeled (forward) and 5'-Cy5-labeled (reverse) primers were used to amplify the L-DNA segment in four separate PCR reactions using D-Dpo4-5m (a mutant Dpo4 to facilitate its chemosynthesis). In each reaction, one L-dNTP was replaced with the corresponding L-dNTPαS. The PCR program was set to 86°C for 3 minutes (initial denaturation), 86°C for 30 seconds, 54°C (Tm-dependent) for 1 minute, and 65°C for 1-2.5 minutes (dependent on amplicon length) for 45 cycles, followed by 65°C for 5 minutes (final extension). The PCR product (mixed with unlabeled carrier dsDNA of the same length at a 1:20 w / w ratio) was purified by 8% PAGE and dissolved in water to a concentration of approximately 200 ng / μl. For each sequencing reaction, 2.5 μl of double-labeled L-DNA was mixed with 2.5 μl of denaturing buffer (98% formamide, 0.25 mM EDTA) containing 2% (v / v) 2-iodoethanol, followed by heating at 95°C for 3 minutes, and then rapidly placed on ice. For the chain termination technique, L-DNA segments were amplified in four separate PCR reactions using enantiomer Pfu DNA polymerase mutants (D215A, L490W) (SEQ ID NO: 77) with 5'-FAM-labeled (forward) and / or 5'-Cy5-labeled (reverse) primers, with one L-dNTP being replaced with the corresponding L-ddNTP in a fixed ratio within each reaction. The PCR program settings were as follows: 3 minutes at 94°C (initial denaturation), 30 seconds at 94°C, 30 seconds at 54°C (depending on Tm), and 30-60 seconds at 72°C (depending on amplicon length), for 20 cycles, followed by 5 minutes at 72°C (final extension).The dual-labeled PCR products were each mixed with an equal volume of denaturing buffer (98% formamide, 0.25 mM EDTA), then heated at 95°C for 3 minutes, and then rapidly placed on ice. Sequencing gels of D-DNA segment S1 obtained by linkage arrest using expressed Pfu DNA polymerase mutants (D215A, L490W) with ddNTP and 5'-Cy5-labeled (reverse) sequencing primers, and D-DNA segment S1 amplification products using Pfu DNA polymerase mutants (D215A, L490W) with ddNTP and 5'-Cy5-labeled reverse sequencing primers were analyzed by 10% denaturing PAGE and scanned with a Typhoon Trio+ system operating in Cy5 mode. A represents a sample in which dATP is partially replaced by ddATP, C represents a sample in which dCTP is partially replaced by ddCTP, G represents a sample in which dGTP is partially replaced by ddGTP, and T represents a sample in which dTTP is partially replaced by dTTP (results are not shown). Sequencing samples were loaded onto 0.4 mm × 340 mm × 300 mm slabs and separated using a 10% polyacrylamide gel in 8 M urea. The gels were pre-electrophoresed at 50 W (constant power) for 2 hours until heated to 30-40°C. After loading, the gels were electrophoresed at 50 W (constant power) for 1.5 hours, interrupted for fluorescence scanning, and then the gel electrophoresis was continued, scanning every hour until the total electrophoresis time reached a maximum of 5 hours. The polyacrylamide gels were scanned using Typhoon Trio operating in Cy2 and Cy5 modes, respectively. + The images were scanned by the system. Gel quantification and chromatographic analysis were performed using ImageJ software.

[0270] Chiral Steganography: Chimeric D-DNA / L-DNA oligos were synthesized using D- and L-deoxynucleoside phosphoramidites using the method described above. Oligos D-F1, D-R1, D / L-F2, and D / L-R2 (Table 7) were heated to 95°C for 3 minutes for annealing, slowly cooled to 4°C, and the annealed double-stranded DNA was ligated using T3 DNA ligase (New England Biolabs, Massachusetts, USA) at 25°C for 1.5 hours. A D-DNA storage library, which would serve as "cover text," was prepared using TransStart FastPfu Fly polymerase (TransGen Biotech, Beijing, China) in a similar manner to that used for the L-DNA storage library. Chimeric double-stranded D-DNA / L-DNA keys purified by agarose gel were added to the D-DNA storage library as each D-DNA segment in a 1:1 concentration ratio. Eleven information-storing D-DNA segments and the D-DNA portion of the chimeric key were amplified from the storage library using segment-specific primers and cloned for Sanger sequencing using the Zero Background ZT4 Simple-Blunt Fast Cloning Kit (Beijing Zoman Biotech, China) (Supplementary Table S6). The L-DNA portion of the chimeric key was amplified from D-Dpo4-5m from the storage library using L-M13F and L-M13R primers and sequenced using the phosphorothioate method.

[0271] Table 7 shows the sequences used for chiral steganography, where lowercase letters represent D-DNA sequences, uppercase letters represent L-DNA sequences, and underlined letters are unique sequences for amplification and sequencing of individual segments.

[0272] [Table 7]

[0273] L-DNA barcoding: Unpurified environmental water samples were collected from the Lotus Pond (40°0'27"N, 116°19'34"E) at Tsinghua University on December 8, 2019. Synthetic D- and L-DNA oligos were heated to 95°C for 5 minutes for annealing, slowly cooled to 4°C, and the annealed dsDNA was added to the water sample to a concentration of 50 μg / L. To amplify the DNA barcode (SEQ ID NO: 12), 2 ml of the water sample was filtered through a 0.22 μm filter (Pall Corporation, Wisconsin, USA), resuspended in DEPC-treated water using an Amicon Utra centrifugal filter unit (0.5 ml, 10,000 MWCO), and then amplified using D- / L-Pfu DNA polymerase. The PCR program settings were 3 minutes at 94°C (initial denaturation), 30 seconds at 94°C, 30 seconds at 55°C, and 1 minute at 72°C for 25 cycles, and 10 minutes at 72°C (final extension). For metagenomic microbial DNA extraction, the water sample was filtered through a 0.2 μm Supor 200 PES membrane disc filter (Pall, New York, USA), and microbial DNA was extracted using the DNeasy PowerSoil kit (Qiagen, Maryland, USA).

[0274] 16S rRNA gene assembly: Synthetic oligonucleotides approximately 90 nt long, at concentrations of 0.005–0.02 μM (inner) or 0.2 μM (outer), were assembled into full-length genes in two steps. In the first step, the assembly PCR program was set to 35 cycles of 94°C for 3 minutes (initial denaturation), 94°C for 30 seconds, 60°C for 30 seconds, and 72°C for 3 minutes, followed by 10 minutes at 72°C (final extension). In the second step, pre-assembled DNA blocks approximately 450–550 bp long were purified using a 1.5% agarose gel and then subjected to assembly PCR. The assembly PCR program was set to 35 cycles of 94°C for 3 minutes (initial denaturation), 94°C for 30 seconds, 60°C for 30 seconds, and 72°C for 7 minutes, followed by 10 minutes at 72°C (final extension). The assembled product was further amplified using a PCR program with the following settings: 3 minutes at 94°C (initial denaturation), 30 seconds at 94°C, 30 seconds at 60°C, and 7 minutes at 72°C for 35 cycles, followed by 10 minutes at 72°C (final extension). The final D-DNA product (SEQ ID NO: 81) from the natural assembly PCR was purified using the V-elute gel mini purification kit (Beijing Zoman Biotech, Beijing, China) and cloned for Sanger sequencing using the Zero Background ZT4 Simple-Blunt Fast Cloning Kit (Beijing Zoman Biotech, Beijing, China).

[0275] Although the present invention is described in conjunction with its specific embodiments, it will be obvious to those skilled in the art that many alternative, improved, and modified forms will be apparent. Accordingly, it is intended that all such alternative, improved, and modified forms that fall within the spirit and broad scope of the appended claims are encompassed.

[0276] All publications, patents, and patent applications referenced herein are incorporated herein by reference in the same manner as each individual publication, patent, or patent application is specifically and individually indicated herein. In addition, any citation or specification of any reference in this application should not be construed as an endorsement that such reference is available as prior art of the present invention. Section headings, insofar as they are used, should be construed as not necessarily limiting. In addition, any one or more priority documents of this application are incorporated herein by reference in their entirety.

[0277] In addition, any one or more priority documents of this application are incorporated herein by reference in whole.

[0278] References 1. L. Ceze, J. Nivala, K. Strauss, Molecular digital data storage using DNA. Nat Rev Genet 20, 456-466 (2019). 2. N. Goldman et al., Towards practical, high-capacity, low-maintenance information storage in synthesized DNA. Nature 494, 77-80 (2013). 3. GM Church, Y. Gao, S. Kosuri, Next-generation digital information storage in DNA. Science 337, 1628 (2012). 4. L. Pasteur, Researches on the Molecular Asymmetry of Natural Organic Products. Soc. Chim. Paris, (1860). 5. Z. Wang, W. Xu, L. Liu, T. F. Zhu, A synthetic molecular system capable of mirror-image genetic replication and transcription. Nature Chemistry 8, 698-704 (2016). 6. M. Peplow, A Conversation with Ting Zhu. ACS Cent Sci 4, 783-784 (2018). 7. M. Peplow, Mirror-image enzyme copies looking-glass DNA. Nature 533, 303-304 (2016). 8. S. L. Beaucage, M. H. Caruthers, Deoxynucleoside Phosphoramidites - a New Class of Key Intermediates for Deoxypolynucleotide Synthesis. Tetrahedron Lett 22, 1859-1862 (1981). 9. Y. Liu et al., Synthesis and applications of RNAs with position-selective labelling and mosaic composition. Nature 522, 368-372 (2015). 10. R. B. Merrifield, Solid Phase Peptide Synthesis .1. Synthesis of a Tetrapeptide. Journal of the American Chemical Society 85, 2149-& (1963). 11. L. Z. Yan, P. E. Dawson, Synthesis of peptides and proteins without cysteine residues by native chemical ligation combined with desulfurization. J Am Chem Soc 123, 526-533 (2001). 12. P. Dawson, T. Muir, I. Clark-Lewis, S. Kent, Synthesis of proteins by native chemical ligation. Science 266, 776-779 (1994). 13. G.-M. Fang et al., Protein Chemical Synthesis by Ligation of Peptide Hydrazides. Angewandte Chemie International Edition 50, 7645-7649 (2011). 14. R. Milton, S. Milton, S. Kent, Total chemical synthesis of a D-enzyme: the enantiomers of HIV-1 protease show reciprocal chiral substrate specificity. Science 256, 1445-1448 (1992). 15. A. A. Vinogradov, E. D. Evans, B. L. Pentelute, Total synthesis and biochemical characterization of mirror image barnase. Chemical Science 6, 2997-3002 (2015). 16. M. T. Weinstock, M. T. Jacobsen, M. S. Kay, Synthesis and folding of a mirror-image enzyme reveals ambidextrous chaperone activity. Proceedings of the National Academy of Sciences of the United States of America 111, 11679-11684 (2014). 17. W. Xu et al., Total chemical synthesis of a thermostable enzyme capable of polymerase chain reaction. Cell discovery 3, 17008 (2017). 18. W. Jiang et al., Mirror-image polymerase chain reaction. Cell discovery 3, 17037 (2017). 19. A. Pech et al., A thermostable d-polymerase for mirror-image PCR. Nucleic Acids Res 45, 3997-4005 (2017). 20. L. E. Zawadzke, J. M. Berg, A Racemic Protein. Journal of the American Chemical Society 114, 4002-4003 (1992). 21. M. Wang et al., Mirror-image gene transcription and reverse transcription. Chem 5, 848-857 (2019). 22. B. J. Lamarche, S. Kumar, M. D. Tsai, ASFV DNA polymerse X is extremely error-prone under diverse assay conditions and within multiple DNA sequence contexts. Biochemistry 45, 14826-14833 (2006). 23. H. Ling, F. Boudsocq, R. Woodgate, W. Yang, Crystal structure of a Y-family DNA polymerase in action: a mechanism for error-prone and lesion-bypass replication. Cell 107, 91-102 (2001). 24. F. Boudsocq, S. Iwai, F. Hanaoka, R. Woodgate, Sulfolobus solfataricus P2 DNA polymerase IV (Dpo4): an archaeal DinB-like DNA polymerase with lesion-bypass properties akin to eukaryotic pol eta. Nucleic Acids Research 29, 4607-4616 (2001). 25. J. Cline, J. C. Braman, H. H. Hogrefe, PCR fidelity of pfu DNA polymerase and other thermostable DNA polymerases. Nucleic Acids Res 24, 3546-3551 (1996). 26. C. J. Hansen, L. Wu, J. D. Fox, B. Arezi, H. H. Hogrefe, Engineered split in Pfu DNA polymerase fingers domain improves incorporation of nucleotide gamma-phosphate derivative. Nucleic Acids Res 39, 1801-1810 (2011). 27. Q. Wan, S. J. Danishefsky, Free-radical-based, specific desulfurization of cysteine: a powerful advance in the synthesis of polypeptides and glycopolypeptides. Angew Chem Int Ed Engl 46, 9248-9252 (2007). 28. J. T. Hyde C, Owen D, Quibell M, Sheppard RC., Some ‘difficult sequences’ made easy. International journal of peptide and Protein Research 43, 431-440 (1994). 29. T. Johnson, M. Quibell, R. C. Sheppard, N,O-bisFmoc derivatives of N-(2-hydroxy-4-methoxybenzyl)-amino acids: Useful intermediates in peptide synthesis. Journal of Peptide Science 1, 11-25 (1995). 30. J. S. Zheng et al., Robust Chemical Synthesis of Membrane Proteins through a General Method of Removable Backbone Modification. J Am Chem Soc 138, 3553-3561 (2016). 31. M. T. Jacobsen et al., A Helping Hand to Overcome Solubility Challenges in Chemical Protein Synthesis. J Am Chem Soc 138, 11775-11782 (2016). 32. F. W. Torsten Wohr, Adel Nefzi, Barbara Rohwedder, Tatsunori Sato, Xicheng Sun, Manfred Mutter, Pseudo-Prolines as a Solubilizing, Structure-Disrupting Protection Technique in Peptide Synthesis. J Am Chem Soc 118, 9218-9227 (1996). 33. M. K. Pascal Dumy, Declan E. Ryan, Barbara Rohwedder, Torsten Wohr, Manfred Mutter, Pseudo-Prolines as a Molecular Hinge: Reversible Induction of cis Amide Bonds into Peptide Backbones. J. Am. Chem. Soc. 119, 918-925 (1997). 34. Y. Sohma et al., ‘O-Acyl isopeptide method’ for the efficient synthesis of difficult sequence-containing peptides: use of ‘O-acyl isodipeptide unit’. Tetrahedron Letters 47, 3013-3017 (2006). 35. I. Coin, The depsipeptide method for solid-phase synthesis of difficult peptides. Journal of peptide science: an official publication of the European Peptide Society 16, 223-230 (2010). 36. G. M. Fang, J. X. Wang, L. Liu, Convergent chemical synthesis of proteins by ligation of peptide hydrazides. Angew Chem Int Ed Engl 51, 10347-10350 (2012). 37. J. S. Zheng, S. Tang, Y. K. Qi, Z. P. Wang, L. Liu, Chemical synthesis of proteins using peptide hydrazides as thioester surrogates. Nat Protoc 8, 2483-2495 (2013). 38. N. K. L., G. Gerald, E. Fritz, V. Hans-Peter, Direct sequencing of polymerase chain reaction amplified DNA fragments through the incorporation of deoxynucleoside α-thiotriphosphates. Nucleic Acids Research, 21 (1988). 39. G. Gish, F. Eckstein, DNA and RNA sequence determination based on phosphorothioate chemistry. Science 240, 1520-1522 (1988). 40. C. Y. Chen, DNA polymerases drive DNA sequencing-by-synthesis technologies: both past and present. Front Microbiol 5, 305 (2014). 41. A. S. Xiong et al., A simple, rapid, high-fidelity and cost-effective PCR-based two-step DNA synthesis method for long gene sequences. Nucleic Acids Res 32, e98 (2004). 42. A. Tiessen, P. Perez-Rodriguez, L. J. Delaye-Arredondo, Mathematical modeling and comparison of protein size distribution in different plant, animal, fungal and microbial species reveals a negative correlation between protein size and protein number, thus providing insight into the evolution of proteomes. BMC Res Notes 5, 85 (2012). 43. C. Cozens, V. B. Pinheiro, A. Vaisman, R. Woodgate, P. Holliger, A short adaptive path from DNA to RNA polymerases. Proc Natl Acad Sci U S A 109, 8067-8072 (2012). 44. X. Liu, TF Zhu, Sequencing mirror-Image DNA chemically. Cell Chemical Biology 25, 1151-1156 e1153 (2018). 45. D. Wade et al., All-D amino acid-containing channel-forming antibiotic peptides. Proc Natl Acad Sci USA 87, 4761-4765 (1990). [Sequence Listing Free Text]

[0279] Sequence ID 1: L-DNA nucleic acid sequence Sequence ID 2: L-DNA nucleic acid sequence Sequence ID 3: L-DNA nucleic acid sequence Sequence ID 4: L-DNA nucleic acid sequence Sequence ID 5: L-DNA nucleic acid sequence Sequence ID 6: L-DNA nucleic acid sequence Sequence ID 7: L-DNA nucleic acid sequence Sequence ID 8: L-DNA nucleic acid sequence Sequence ID 9: L-DNA nucleic acid sequence Sequence ID 10: L-DNA nucleic acid sequence Sequence ID 11: L-DNA nucleic acid sequence Sequence ID 12: DNA barcode nucleic acid sequence Sequence ID 13: Short synthetic oligonucleotide sequence Sequence ID 14: Short synthetic oligonucleotide sequence Sequence ID 15: Short synthetic oligonucleotide sequence Sequence ID 16: Short synthetic oligonucleotide sequence Sequence ID 17: Short synthetic oligonucleotide sequence Sequence ID 18: Short synthetic oligonucleotide sequence Sequence ID 19: Short synthetic oligonucleotide sequence Sequence ID 20: Short synthetic oligonucleotide sequence Sequence ID 21: Short synthetic oligonucleotide sequence Sequence ID 22: Short synthetic oligonucleotide sequence Sequence ID 23: Short synthetic oligonucleotide sequence Sequence ID 24: Short synthetic oligonucleotide sequence Sequence ID 25: Short synthetic oligonucleotide sequence Sequence ID 26: Short synthetic oligonucleotide sequence Sequence ID 27: Short synthetic oligonucleotide sequence Sequence ID 28: Short synthetic oligonucleotide sequence Sequence ID 29: Short synthetic oligonucleotide sequence Sequence ID 30: Short synthetic oligonucleotide sequence Sequence ID 31: Short synthetic oligonucleotide sequence Sequence ID 32: Short synthetic oligonucleotide sequence Sequence ID 33: Short synthetic oligonucleotide sequence Sequence ID 34: Short synthetic oligonucleotide sequence Sequence ID 35: Short synthetic oligonucleotide sequence Sequence ID 36: Single-stranded DNA oligonucleotide Sequence ID 37: Single-stranded DNA oligonucleotide Sequence ID 38: Short synthetic oligonucleotide sequence Sequence ID 39: Short synthetic oligonucleotide sequence Sequence ID 40: Short synthetic D- / L-chimeric oligonucleotide sequence Sequence ID 41: Short synthetic D- / L-chimeric oligonucleotide sequence Sequence ID 42: Single-strand DNA oligonucleotide Sequence ID 43: Single-stranded DNA oligonucleotide Sequence ID No. 44: Single-stranded L-DNA oligonucleotide Sequence ID No. 45: Single-stranded L-DNA oligonucleotide Sequence ID 46: D- / L-chimeric DNA nucleic acid sequence Sequence ID 47: Pfu DNA polymerase Sequence ID 48: Mutant of Pfu DNA polymerase SEQ ID NO: 49: Pfu-5m-55I amino acid sequence Sequence ID 50: Pfu-5m-46I amino acid sequence Sequence ID 51: Mutant of Pfu DNA polymerase Sequence ID 52: Mutant form of Pfu DNA polymerase Sequence ID 53: KOD1 polymerase Sequence ID 54: Tgo polymerase Sequence ID 55: Amino acid sequence of N-7 polymerase at degree 9 Sequence ID 56: Tok polymerase Sequence ID 57: N fragment of Pfu DNA polymerase Sequence ID 58: N fragment of Pfu DNA polymerase Sequence ID 59: N fragment of Pfu DNA polymerase Sequence ID 60: N fragment of Pfu DNA polymerase, with a thiazolidine-4-carboxylic acid (Tfa-Thz) bond at position 1. Sequence ID 61: N fragment of Pfu DNA polymerase, with a thiazolidine-4-carboxylic acid (Tfa-Thz) bond at position 1. Sequence ID 62: N fragment of Pfu DNA polymerase Sequence ID 63: N fragment of Pfu DNA polymerase, with a thiazolidine-4-carboxylic acid (Tfa-Thz) bond at position 1. Sequence ID 64: N fragment of Pfu DNA polymerase, with a thiazolidine-4-carboxylic acid (Tfa-Thz) bond at position 1. Sequence ID 65: N fragment of Pfu DNA polymerase, with a thiazolidine-4-carboxylic acid (Tfa-Thz) bond at position 1. Sequence ID 66: N fragment of Pfu DNA polymerase Sequence ID 67: C fragment of Pfu DNA polymerase Sequence ID 68: C fragment of Pfu DNA polymerase Sequence ID 69: C fragment of Pfu DNA polymerase Sequence ID 70: C fragment of Pfu DNA polymerase Sequence ID 71: C fragment of Pfu DNA polymerase, with a thiazolidine-4-carboxylic acid (Tfa-Thz) bond at position 1 (N-terminal). Sequence ID 72: C fragment of Pfu DNA polymerase, with a thiazolidine-4-carboxylic acid (Tfa-Thz) N-terminal bond at position 1. Sequence ID 73: C fragment of Pfu DNA polymerase Sequence ID 74: Mutant of Pfu DNA polymerase Sequence ID 75: Mutant of Pfu DNA polymerase Sequence ID 76: Mutant of Pfu DNA polymerase Sequence ID 77: Mutant of Pfu DNA polymerase Sequence ID 78: Amino acid sequence of the sso7d structural domain Sequence ID 79: Amino acid sequence of Pfu DNA polymerase Sequence ID 80: Nucleic acid sequence of the pUC19 plasmid Sequence ID 81: DNA template encoding the bacterial 16S rRNA gene Sequence ID 82: T7-WT amino acid sequence Sequence ID 83: T7-37I (I6V, I14L, I74L, I82V, I109V, I117L, I141V, I219M, I244L, I281V, I320V, I322L, I330V, I367L) amino acid sequence Sequence ID 84: YenP amino acid sequence Sequence ID 85: phiEap amino acid sequence Sequence ID 86: KpnP amino acid sequence Sequence ID 87: Amino acid sequence of the T7-fission-N fragment Sequence ID 88: T7-N-1 amino acid sequence Sequence ID 89: T7-N-2 amino acid sequence Sequence ID 90: T7-N-3 amino acid sequence Sequence ID 91: T7-N-4 amino acid sequence Sequence ID 92: T7-N-5 amino acid sequence Sequence ID 93: T7-N-6 amino acid sequence, with a thiazolidine-4-carboxylic acid (Tfa-Thz) bond at position 1 (N-terminal trifluoroacetate). Sequence ID 94: T7-N-7 amino acid sequence Sequence ID 95: Amino acid sequence of the T7-fission-M fragment Sequence ID 96: T7-M-1 amino acid sequence Sequence ID 97: T7-M-2 amino acid sequence Sequence ID 98: T7-M-3 amino acid sequence Sequence ID 99: T7-M-4 amino acid sequence, with a thiazolidine-4-carboxylic acid (Tfa-Thz) bond at position 1 (N-terminal trifluoroacetate). Sequence ID 100: T7-M-5 amino acid sequence, position 1 is N-terminal thiazolidine-4-carboxylic acid (Tfa-Thz) bond. Sequence ID 101: T7-M-6 amino acid sequence Sequence ID 102: Amino acid sequence of the T7-split-C fragment Sequence ID 103: T7-C-1 amino acid sequence Sequence ID 104: T7-C-2 amino acid sequence Sequence ID 105: T7-C-3 amino acid sequence Sequence ID 106: T7-C-4 amino acid sequence, with a thiazolidine-4-carboxylic acid (Tfa-Thz) bond at position 1 (N-terminal trifluoroacetate). Sequence ID 107: Amino acid sequence of T7-C-5 Sequence ID 108: Nucleic acid sequence of a DNA template Sequence ID 109: Nucleic acid sequence of the Tt 16S DNA template Sequence ID 110: DNA template for tRNA(Ser) Sequence ID 111: DNA template for L-guanine sensor Sequence ID 112: DNA template for L-38-6 ribozyme Sequence ID 113: DNA template for L-class I ligase Sequence ID 114: L-38-6 Ribozyme Sequence ID 115: L-5'-FAM-labeled primer, FAM-labeled at position 1, FAM-bound at position 1. Sequence ID 116: Template for L-class I ligase

Claims

1. A method for chemically producing a protein, comprising linking at least two ligation-inducible segments of the protein, wherein each of the ligation-inducible segments is chemically synthesizable, and i. Identifying at least one ligation-inducible sequence in the amino acid sequence of the protein, and obtaining multiple ligation-inducible segments by parsing the amino acid sequence of the protein with the ligation-inducible sequence, and ii. If each of the ligation-inducible segments is chemically synthesizable, chemically synthesize each of the ligation-inducible segments. iii. If any one of the ligation-inducible segments is not chemically synthesizable, identify at least one loss-of-structure section in the ligation-inducible segment, replace at least one amino acid in the loss-of-structure section with a ligation-inducible amino acid residue to introduce a ligation-inducible sequence into the loss-of-structure section, parse the amino acid sequence of the protein with the ligation-inducible sequence, and chemically synthesize each of the ligation-inducible segments. It can be obtained by, In step (i), at least one of the ligation-inducible sequences is located in a loss-of-structure section of the protein, Before step (i), a) Dividing the amino acid sequence of the protein into at least two domain-forming segments, b) If each of the domain-forming segments is chemically synthesizable, chemically synthesize each of the domain-forming segments, and c) Fold the domain-forming segments together to obtain the protein. Methods that further include the above.

2. If one of the domain-forming segments is not chemically synthesizable, d) Identify at least one ligation-inducible sequence in the domain-forming segment, and parse the amino acid sequence of the domain-forming segment with the ligation-inducible sequence to obtain a plurality of chemically synthesizable ligation-inducible segments. e) If the domain-forming segment lacks a ligation-inducible sequence, or if any one of the ligation-inducible segments is not chemically synthesizable, identify at least one structure-loss section in the domain-forming segment or the ligation-inducible segment. f) Substituting at least one amino acid in the structure loss section or the ligation-inducible segment with a ligation-inducible amino acid residue to introduce a ligation-inducible sequence into the structure loss section or the ligation-inducible segment, and parsing the amino acid sequence of the domain-forming segment with the ligation-inducible sequence to obtain a plurality of sequences of chemically synthesizable ligation-inducible segments. g) Chemically synthesize each of the chemically synthesizable ligation-inducible segments. The method according to claim 1.

3. The method according to claim 1 or 2, wherein the protein comprises at least 240 amino acid residues.

4. The method according to any one of claims 1 to 3, wherein the protein is produced using at least 90% D-amino acid residues other than glycerides, and the protein has a three-dimensional structure that is a mirror image of the three-dimensional structure of a corresponding biologically produced protein.

5. A protein prepared by the method according to any one of claims 1 to 4, wherein the protein has a length of at least 240 amino acid residues, and the protein is a D- or L-amino acid RNA polymerase capable of synthesizing RNA from ribonucleotides using a DNA template, or a D- or L-amino acid DNA polymerase capable of synthesizing DNA from deoxyribonucleotides.

6. The protein according to claim 5, wherein the D- or L-amino acid RNA polymerase is T7 RNA polymerase or a Pfu DNA polymerase variant.

7. The protein according to claim 5, wherein the D- or L-amino acid DNA polymerase is Pfu DNA polymerase.

8. A process for enzymatically producing L-polydeoxyribonucleic acid molecules, To provide a D-amino acid DNA polymerase that can be prepared by the method described in claim 4 and can synthesize L-DNA from L-deoxyribonucleotides, and The L-DNA molecule is enzymatically produced by reacting the D-amino acid DNA polymerase with a template L-DNA molecule, an L-DNA primer, and a plurality of L-deoxyribonucleotides. A process that includes this.

9. The process according to claim 8, wherein the D-amino acid DNA polymerase is Pfu DNA polymerase.

10. The process according to claim 9, wherein the Pfu DNA polymerase comprises at least one mutation selected from the group consisting of E102A, E276A, K317G, V367L, and I540A, or the Pfu DNA polymerase comprises at least two polypeptide chains formed by fission between K467 and M468, wherein the position numbering is based on the amino acid position numbering of the corresponding WT enzyme.

11. A process for enzymatically producing L-polyribonucleic acid (L-RNA) molecules, To provide a D-amino acid RNA polymerase that can be prepared by the method described in claim 4 and can synthesize L-RNA from L-ribonucleotides, and The L-RNA molecule is enzymatically produced by reacting the D-amino acid RNA polymerase with a template L-DNA molecule, an L-DNA / RNA primer, and a plurality of L-ribonucleotides. A process that includes this.