Methods for generating miniaturized enzymes and enzymes produced thereby

By transforming vectors with mutated active-site residues into hosts, miniaturized enzymes with catalytic activity are produced, addressing the lack of in vivo urzyme generation and enhancing our understanding of PK's role in health and disease.

WO2026156147A1PCT designated stage Publication Date: 2026-07-23THE UNIV OF NORTH CAROLINA AT CHAPEL HILL
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
THE UNIV OF NORTH CAROLINA AT CHAPEL HILL
Filing Date
2026-01-15
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

Current methods have not successfully produced miniaturized enzymes, known as urzymes, that retain catalytic activity in vivo, particularly for enzymes like pyruvate kinase (PK) which are crucial for glycolytic pathways and have implications in health and disease.

Method used

A method involving vectors encoding enzymes with mutations at multiple catalytic active-site residues is transformed into an exogenous host for expression, resulting in the production of recombinant deletions that form miniaturized enzymes with substantial catalytic activity.

Benefits of technology

The method generates miniaturized enzymes, or urzymes, that maintain functional activity, filling the gap in understanding regulatory and evolutionary mechanisms across various organisms and providing insights into PK's role in health and disease.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2026011400_23072026_PF_FP_ABST
    Figure US2026011400_23072026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to methods for generating miniaturized enzymes ("urzymes") and miniaturized enzymes produced using the method. The methods involve delivering vectors encoding an enzyme containing mutations of multiple catalytic active-site residues into an exogenous host for recombination induction and expression, producing nucleic acids encoding miniaturized enzymes through in vivo recombinant deletions.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No. 5470.977.WOMETHODS FOR GENERATING MINIATURIZED ENZYMES AND ENZYMES PRODUCED THEREBYSTATEMENT REGARDING ELECTRONIC FILING OF A SEQUENCE LISTING

[0001] A Sequence Listing in XML format, entitled 5470-977WO_ST26.xml, 78,443 bytes in size, generated on January 14, 2026, and filed herewith, is hereby incorporated by reference into the specification for its disclosures.STATEMENT OF PRIORITY

[0002] This application claims the benefit of U.S. Provisional Application Serial No.63 / 745,471, filed January 15, 2025, the entire contents of which are incorporated by reference herein.FIELD OF THE INVENTION

[0003] The present invention relates to methods for generating miniaturized enzymes (“urzymes”) and miniaturized enzymes produced using the method. The methods involve delivering vectors encoding an enzyme containing mutations of multiple catalytic active-site residues into an exogenous host for recombination induction and expression, producing nucleic acids encoding miniaturized enzymes through in vivo recombinant deletions.BACKGROUND OF THE INVENTION

[0004] Aminoacyl-tRNA synthetases (AARS) are essential and play a critical role in protein synthesis, pairing tRNAs with their cognate amino acids for decoding mRNAs according to the genetic code. AARSs are believed to have originated very early in evolution and were already present in their mature form with the last universal common ancestor (LUCA). There are twenty -three known AARSs, which are divided into two major classes based on the architecture of their active site, Class I and Class II. The classes of synthetases approach the tRNA from different sides and it is possible to simultaneously model the docking of pairs of enzymes of each class to a single tRNA without major steric hinderances. This complementary recognition of the major and minor grooves of the tRNA acceptor appears to be the consequence of an evolutionary model in which both ancestors of each class arose from a single gene, often known as the Rodin-Ohno hypothesis, such that the gene of the ancestral AARS could be read bidirectionally, and each of the opposite strands would code for the ancestor of class I and class II, respectively. This evolutionary model emerged when it became clear that short segments ofAttorney Docket No. 5470.977.WOthe genes for actives sites of Class I and II AARS exhibit exceptionally high base pairing when aligned in frame, opposite one another.

[0005] Beyond their central role in translation, tRNA synthetases are also emerging as key players in an increasing number of other cellular processes, with far-reaching consequences in health and disease. The biochemical versatility of the synthetases has also proven pivotal in efforts to expand the genetic code, further emphasizing the wide-ranging roles of the AARS family in synthetic and natural biology.

[0006] Current study of the evolution of AARSs involves comparing enzyme anatomy by superimposing tridimensional models aiming to unveil the basic functional and invariant core of the enzyme and removing long, structurally variable inserted segments that intersect the conserved region at residues whose backbone residues can be connected by a single peptide bond. This procedure yields extremely reduced versions of both AARS classes, which have been termed urzymes. These urzymes are created in vitro and contain little more than the active site and can be as small as 15% of the length of the corresponding full-length enzymes. To date, urzymes created in vivo that retain catalytic activity have yet to be identified.

[0007] Pyruvate kinase (PK) is central to the most rudimentary glycolytic pathway that all cells use (glycolysis). In the last step of glycolysis, PK catalyzes the conversion of phosphoenolpyruvate and ADP to pyruvate and ATP. It also plays an important regulatory role in controlling whether that pathway is turned on or off, depending on the nutritional state of the cell. PK is regulated by a variety of mechanisms including allosteric modulators, heterotropic effectors, and other effectors. Given its central role in glycolysis, PK has far-reaching consequences in health and disease, including in cancer and blood disorders such as pyruvate kinase deficiency.

[0008] While PK is believed to be highly conserved, the phylogenetic tree structure of PK does not coincide with the three domains of life, Bacteria, Archaea, and Eukarya, and regulation of PK varies in different organisms and different tissues. Current study of PK involves structural and functional studies of PK from select organisms with known structures of PK in its ligand-free state and in complex with various ligands. More work is needed to understand the regulatory and evolutionary mechanisms for PK across a variety of organisms. Urzymes can help fill this gap. To date, urzymes of PK that retain catalytic activity have yet to be identified.SUMMARY OF THE INVENTION

[0009] The present invention is based on the discovery that when vectors encoding an enzyme are amplified to introduce mutations at multiple catalytic active-site residues and transformedAttorney Docket No. 5470.977.WOinto an exogenous host for expression, the transformation produces, in addition to the expected full-length double mutant vector, a series of recombinant deletions with chimeric features in vivo. In particular, the inventors have shown that the vectors encode enzymes similar in size to a miniaturized enzyme (urzyme) and those enzymes retain substantial catalytic activity.

[0010] Thus, one aspect of the present invention relates to a method of generating a polynucleotide encoding a miniaturized enzyme comprising: i) preparing a vector comprising a polynucleotide encoding an enzyme; wherein the enzyme comprises an active site composed of amino acid residues that are divided into two or more separate sections on the primary amino acid sequence of the enzyme; ii) hybridizing to the vector at least one pair of oligonucleotides that specifically hybridize to nucleotides encoding the amino acid residues in separate sections of the active site, wherein at least one of the oligonucleotides is designed to introduce an amino acid mutation into the enzyme; iii) performing a nucleic acid amplification reaction on the vector to produce an amplification product; iv) introducing the amplification product into a prokaryotic, archaean, or eukaryotic cell suitable for recombination; and v) isolating vectors that have undergone recombination to encode a miniaturized enzyme.

[0011] In an additional aspect, the present invention relates to a composition comprising a vector comprising a polynucleotide encoding an enzyme; wherein the enzyme comprises an active site composed of amino acid residues that are divided into two or more separate sections on the primary amino acid sequence of the enzyme; and at least one pair of oligonucleotides hybridized to the vector, wherein the oligonucleotides specifically hybridize to nucleotides encoding amino acid residues in separate sections of the active site, and wherein at least one of the oligonucleotides is designed to introduce an amino acid mutation into the enzyme.

[0012] In additional aspects, the invention relates to amplification products produced from performing an amplification reaction on the compositions of the invention, prokaryotic, archaean, or eukaryotic cells comprising the amplification products of the invention, and vectors encoding a miniaturized enzyme, produced by the methods of the invention (e.g., plasmids derived from transformation with genes containing multiple mutant sites that encode a miniaturized enzyme).

[0013] These and other aspects of the invention are set forth in more detail in the description of the invention below.Attorney Docket No. 5470.977.WOBRIEF DESCRIPTION OF THE DRAWINGS

[0014] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments of the invention and, together with the description, serve to explain principles of the invention.

[0015] FIG. 1 shows a schematic for formation and selection of LeuRS urzyme-like deletions, which is an example of using multi- mutant multi-domain protein genes for ancestral gene discovery. A shows full-length LeuRS with an allosteric network of interactions between Dom A (CPI), Dom B (ABD) and WT active-site catalytic histidine and lysine residues to enforce specificity. B shows a cytotoxic DNA and protein created by a double active-site mutant which corrupts the allosteric effects of the two domains (domains indicated by asterisk). C shows deletion of the domains in vivo whose functions have been corrupted to produce variants with substantially less cytotoxicity. The full-length LeuRS with double mutations resembles the evolutionary precursor of the full-length protein created in vitro (solid arrow).

[0016] FIG. 2 shows a schematic of miniature enzyme (“urzyme”) deconstruction using TrpRS urzyme as an example. The sequence of modules referenced in the figure are shown in the bar schematic on top. The Class I AARS active site signatures are TIGN ((SEQ ID NO:1) in TrpRS; a variant of consensus HIGH (SEQ ID NO:2)) and (consensus) KMSKS (SEQ ID NO:3) The activated tryptophanyl-5’ AMP is shown as sticks. Dashed lines indicate genetic manipulations carried out in the deconstruction. Triangles and crossed circles indicate chain breaks. The two chain breaks between residues 46 and 121 are reconnected by a single peptide bond.

[0017] FIG. 3 shows the bidirectional mutagenesis of a full-length P. horikoshii LeuRS with a pair of antiparallel AVGA (SEQ ID NO:4) and AMSAS (SEQ ID NO:5) mutant primers (SEQ ID NOs:10 and 11). Active-site catalytic signatures are HVGH (SEQ ID NO:6) and KMSKS (SEQ ID NO:3). They are oversized for emphasis and not to scale.

[0018] FIGS. 4A-4D show discrete LeuAC urzyme-like and half-protozyme coding regions created in vivo in E. coli. FIG. 4A shows a linear schematic of the 3D structure of the mosaic structure of P. horikoshii LeuRS and the LeuAC urzyme. “CPI” denotes connecting peptide 1 as described in Burbaum, et al. 1990 and Burbaum and Schimmel 1991. “Editing” denotes the editing domain embedded in CPI. “ABD” denotes the anticodon-binding domain. “A,” “Ci,” and “C2” are discontinuous segments of LeuRS that form LeuAC. They are each joined by single peptide bonds following excision of the intervening sequences to make a single polypeptide. FIG. 4B shows a 3D schematic of LeuRS, with regions corresponding to those in FIG. 4A indicated with numbers. The LeuAC urzyme, a model for the early ancestor of LeuRS,Attorney Docket No. 5470.977.WOis indicated by the arrow pointing to the portions indicated with numbers “1” and “5”. FIG. 4C shows results of the transformation of E. coli with the mutant plasmid (e.g., the megaprimer PCR-released plasmid or a synthetic plasmid comprising the mutated sites), resulting in a reproducible distribution of plasmids in which the expected full-length double mutant P. horikoshii LeuRS represented about 60% of the sequenced plasmids. The remaining 40% of the sequenced plasmids were about evenly split between three recombinant subsets of short, fragmentary orphaned coding sequences (OCSs) formed in vivo from the starting full-length double mutant gene. The 2ndCrossover connection (2ndXvr) of the Rossmann fold is shown as a filled box here for clarity. The different N-terminal segments in the truncated fragments are not to scale. FIG. 4D shows a linear schematic of the 3D structure of the mosaic structure of P. horikoshii LeuRS and the LeuAC urzyme created in vitro, and the two new constructs made in vivo.

[0019] FIGS. 5A-5C show shortened OCSs generated in vivo from full-length LeuRS double active-site mutants. FIG. 5A shows modular deletions and shuffling in multiple sequence alignment. Panels (i) and (ii) show the LeuAC sequence and corresponding segments from full-length P. horikoshii leucyl-tRNA synthetase in rows 1 and 2 (at the top). Boxes “X” and “Z” indicate the two mutant signatures. Region with “#” highlights uniquely retained sequences of the LeuAC -like urzyme. These lack the second half of the protozyme (“$” region). Regions with highlight urzyme sequences uniquely retained in the Goldilocks variant. It lacks the second crossover connection of the Rossman dinucleotide-binding fold (2nd Xvr). FIG. 5B shows a tree constructed from the alignment in FIG. 5A using the default neighbor joining algorithm in Clustal X. It highlights three principal clades of related sequences. The“$,” and symbols correspond to those in FIG. 5A. For reference, the Goldilocks deletion variant was excerpted from one of the second crossover deletions (denoted A2nd Xvr). The LeuAC-like urzyme contains the second crossover lacking in the Goldilocks urzyme and has an extended segment preceding it (underlined region under sequence numbers). The Goldilocks urzyme cannot be a subset of the larger LeuAC-like urzyme. FIG.5C shows genomic properties of the urzyme-like (LeuAC-like) clade (right), Goldilocks clade (middle), and half protozyme clade (left). The filled panels show that incubating transformant colonies from the plates incubated at lower temperatures led to the appearance of more Goldilocks deletions. Y-axes show the fraction of total sample given for each bar within each section. Thus, the total of all bars in all three filled panels, and totals for all six reading frames in each open panel equal 1.0. The open panels show the distribution of reading frames relative to that of the input gene as follows. OCS 0 is the same reading frame as the intact LeuRS plasmid. OCSs 1,2 are advancedAttorney Docket No. 5470.977.WOby 1 and 2 frames. OCSs -1, -2, -3 are corresponding frames in the 3 ’-5’ direction. The different temperature-dependence and mutually exclusive sequences in Goldilocks and Urzyme-like coding regions imply that they arise via different intracellular DNA rearrangements.

[0020] FIGS. 6A-6F show the results of single turnover analysis of the recombinant deletion urzyme variants. The Goldilocks and LeuAC-like urzyme variants were subcloned with both mutant and wild-type active site sequences as MBP fusions into pMAL-c2x in BL21 Star (DE3) (Invitrogen) as described in (Tang et al. 2024). FIG. 6A shows a time course for ATP consumption by all four variants (Goldilocks WT, Goldilocks Dbl (double mutant AVGA (SEQ ID NO:4) / AMSAS (SEQ ID NO:5)), LeuAC-like WT, and LeuAC-like Dbl). FIG. 6B shows a time course for ADP production by all four variants. FIG. 6C shows a time course for AMP production by all four variants. FIG. 6D shows a regression model for the dependence of the first-order rate constant, kchem, on variant and mutant properties. FIG. 6E shows a table of regression coefficients for the regression model in FIG. 6D, along with their estimated error and P values. FIG. 6F shows a histogram of the ratio of AMP to ADP produced by the four urzyme variants. Goldilocks WT urzyme has the highest ratio, and hence uses ATP most efficiently. This ratio is consistent with lower burst size of this variant in ADP production and the higher steady-state rate of AMP production compared to the other variants.

[0021] FIGS. 7A-7D show Michaelis-Menten plots of replicated aminoacylation assays of the leucine TTC minihelix by all four variants: Goldilocks WT (HVGH (SEQ ID NO:6) and KMSKS (SEQ ID NO:3) active site signatures) (FIG. 7A), Goldilocks Dbl (double mutant AVGA (SEQ ID NO:4) / AMSAS (SEQ ID NO:5)) (FIG. 7B), LeuAC-like WT (FIG. 7C), and LeuAC-like Dbl (FIG. 7D).

[0022] FIG. 8 shows a schematic of a full-length pyruvate kinase (PK).

[0023] FIG. 9A shows a sequence alignment of the wild type (WT) PK gene (top, “PK”) and a truncated plasmid (bottom, “Urz”) observed when mutations are made to the sites indicated by boxes. FIG. 9B shows a schematic of miniature enzyme (“urzyme”) deconstruction using PK urzyme as an example. The excerpts shown in the bottom schematic correspond with the full-length PK shown in the top schematic. The N-terminal segment is indicated with “N ” The central segment is indicated with “M,” and the C-terminal segment is indicated with “C ” Spheres show important ligands. At the upper is glucose phosphate; in the lower right is Mg++ oxalate, a competitive inhibitor. ID and CTD are insertion and C-terminal domains.

[0024] FIGS. 10A-10G show a schematic of megaprimer-PCR combined with PCR mutagenesis. FIG. 10A shows a first round of PCR (left) creating a megaprimer with selected mutations encoding amino acid sequences AVGA (SEQ ID NO:4) and AMSAS (SEQ IDAttorney Docket No. 5470.977.WONO:5) from a target wild-type plasmid DNA, and an agarose gel purification of the PCR megaprimer product (right). FIG. 10B shows a second round of PCR, where the purified megaprimer is then annealed onto a target plasmid DNA. FIG. IOC shows PCR extension of the megaprimer during the second round of PCR. The desired PCR extension is initiated and proceeded using the sense-single strand and antisense-single strand of the megaprimer as indicated with dashed circular lines. FIG. 10D shows restriction enzyme Dpnl digestion of the second round PCR product and destruction of original intact wild-type plasmid DNA as indicated with dashed lines. FIG. 10E shows single colonies of transformants on agar plates obtained via transformation with / J / wI-digested PCR mixture into competent E. coli cells. FIG.10F shows culturing of single colonies for plasmid miniprep. FIG. 10G shows sequencing and sequence analysis and to identify mutant plasmids.

[0025] FIG. 11 shows a schematic of fragment-overlapping PCR ligation for megaprimer acquisition. Both the fragment amplicons and the megaprimer amplicon were gel-purified.

[0026] FIG. 12 shows a schematic of PCR mutagenesis and fragment-overlapping PCR ligation for megaprimer preparation for obtaining pyruvate kinase urzymes.

[0027] FIG. 13 shows a gel image of the PCR amplicons generated by three PCR polymerases (from left to right, Fisher DreamTaq, Fisher Phusion Polymerase, NEB Phusion polymerase) in step 1 of FIG. 12. Multiple PCR polymerases were tested for amplification from the construct because of the high GC content of the PK sequence. The white squares indicated the failures of PCRs. Only the bands outside of designated squares were excised and subjected to purification.

[0028] FIG. 14 shows a gel image of the PCR amplicons generated in step 1 of FIG. 12 (the individual fragments). All the labeled and squared bands were excised and subjected to purification using QiaQuick Gel Extraction Kit.

[0029] FIG. 15 shows a gel image of the PCR amplicon (megaprimer) generated in step 2 of FIG. 12, using three PCR polymerases (from left to right, Fisher DreamTaq, Fisher Phusion Polymerase, NEB Phusion polymerase). All squared and bright bands from ThermoFisher Phusion Polymerase were excised and subjected to purification using QiaQuick Gel Extraction Kit. The NEB brand Phusion polymerase failed because of expiration.

[0030] FIG. 16 shows sequence annotations of a PK urzyme generated by the method shown in FIG. 12. The top panel shows the raw nucleotide sequence, and the bottom panel shows its translation, with mutations as indicated by boxes.. The sites of the mutations are indicated by the labels “Mutation site 1,” Catalytic domain 1, and Catalytic domain 2 (a urzyme-like sequence).Attorney Docket No. 5470.977.WO

[0031] FIG. 17 shows sequence alignments indicating that the amino acid sequence in FIG.16 (bottom) (SEQ ID NO:7) is of pyruvate kinase origin. Mutations of catalytic signatures are indicated by boxes.

[0032] FIG. 18 shows a schematic of bidirectional genes for ancestral AARS. The second crossover connection of the Class 1 Rossmann fold is not required for aminoacylation of minihelix RNA substrates. That discovery enables antiparallel alignments of three Class 1 / Class 2 antiparallel alignments with exceptionally high codon middle-base pairing. The high pairing frequencies, however, contradict the conventional sub classification, as indicated by crossing-over of the horizontal lines of descent for subclasses a and b. Placement along the horizontal time axis is uncertain. However, the subclass c synthetases for aromatic amino acids appear to have come later. The subclass a and b bidirectional genes could have enforced a code with an alphabet of four non-overlapping sets of letters. Underlined AARS on the right indicate the contemporary sequences that retain the highest codon middle-base pairing frequencies.

[0033] FIG. 19 shows a schematic of how the N- and C-terminal parts of the LeuRS active site can be connected in three distinct ways (dashed arrows) without changing the catalytic activity in RNA minihelix aminoacylation by either WT or double mutant (Dbl) forms, as shown in the data plot. The three active tRNA synthetases have, respectively, 81 (Goldilocks), 102 (LeuAC-like), and 129 (LeuAC) residues. All structures are based on the crystal structure PDB 1WZ2. The construct on the right was previously designed via computer simulation to mimic the ancestral LeuRS. The two on the left were created in vivo by deletions to a full-length double mutant LeuRS plasmid. Wild type histidine (filled stars) and lysine (open stars) residues were added back by PCR mutagenesis for purposes of comparison.

[0034] FIGS. 20A-20F show data plots from Michaelis-Menten assays indicating that Michaelis-Menten parameters of AARS urzyme minihelix acylation activities depend on the structural differences between constructs. Each plot is ordered in descending favorability from left to right. FIG. 20A shows a bar plot of AG1(kcat / KM), the free energy for the second order rate constant for aminoacylation by reaction of free T'PC-minihelix with free substrate at saturating concentration of isoleucine, the amino acid substrate. FIG. 20B shows a bar plot of AGt(kcat), the free energy for the first-order rate constant. FIG. 20C shows a bar plot of AGi (KM), the free energy for the apparent dissociation constant for the minihelix substrate. FIG.20D shows a regression model that relates steady state parameters in FIG.20A to the structural differences between the constructs. FIG.20E shows a regression model that relates steady state parameters in FIG. 20B to the structural differences between the constructs. FIG.20F shows a regression model that relates steady state parameters in FIG. 20C to the structural differencesAttorney Docket No. 5470.977.WObetween the constructs. Independent variables denote the following: WT refers to the presence of wild type active-site histidine and lysine residues; 2ndXvr denotes the presence of the second crossover; Protoz denotes the presence of the intact protozyme at the N-terminus.

[0035] FIG. 21 shows structural computations indicating that corresponding segments of the active site that bind RNA substrates in Class I LeuRS (PDB ID 1WZ2) and Class II GlyRS (PDB ID 7YSE) have very similar secondary structures. Approximately orthogonal views of the complexes between RNA binding motifs in the two Classes with their respective activated amino acyl-5’AMP and RNA ligands. Dashed circles highlight the aminoacyl groups. The vertices of the square in the background indicate the 5’-phoshporyl group. They are located remarkably similarly with respect to the P-strand and loop. A chlorine atom that is part of the adenine ring in 7YSE has been omitted.

[0036] FIGS. 22A-22B shows schematics indicating that the Goldilocks urzyme suggests a paradigm for designing bidirectional urzyme genes for Class I and II ancestors on opposing DNA strands. FIG.22A shows a schematic of sense / anti sense encoding of Class I and II AARS urzyme genes, consistent with the architecture of the two AARS Classes. Vertical lines between the two genes suggest bases that would be paired in a bidirectional gene but are only partially paired in sequences derived from contemporary genes. FIG. 22B shows a data plot with percentages of codon middle-base pairing in antiparallel alignments, calculated as indicated in FIG. 22A. The percentages are summarized for three groups derived from the three consensus synthetase subclasses (Carter, et al. 2025). Values greater than -35% are highly unlikely under the null hypothesis (Chandrasekaran, et al. 2013).

[0037] FIG. 23 shows a schematic denoting key steps in the emergence of protein AARS. In vivo fragmentation of the LeuRS gene described here provides overlapping reinforcement for the idea that amino acid activating and RNA aminoacylation activities can be attributed to three polypeptide fragments each about 25 amino acids long encircled by short dash lined, solid lined, and long dash lined backgrounds, as described in the text. The two reactions catalyzed by all AARS are indicated with their respective rate accelerations, as shown in FIG. 20.DETAILED DESCRIPTION OF THE INVENTION

[0038] The present invention now will be described hereinafter with reference to the accompanying drawings and examples, in which embodiments of the invention are shown. This invention may, however, be embodied in many different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the inventionAttorney Docket No. 5470.977.WOto those skilled in the art. For example, features illustrated with respect to one embodiment can be incorporated into other embodiments, and features illustrated with respect to a particular embodiment can be deleted from that embodiment. In addition, numerous variations and additions to the embodiments suggested herein will be apparent to those skilled in the art in light of the instant disclosure, which do not depart from the instant invention.

[0039] Unless the context indicates otherwise, it is specifically intended that the various features of the invention described herein can be used in any combination. Moreover, the present invention also contemplates that in some embodiments of the invention, any feature or combination of features set forth herein can be excluded or omitted. To illustrate, if the specification states that a complex comprises components A, B and C, it is specifically intended that any of A, B or C, or a combination thereof, can be omitted and disclaimed singularly or in any combination.

[0040] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the specification and relevant art and should not be interpreted in an idealized or overly formal sense unless expressly so defined herein. Well-known functions or constructions may not be described in detail for brevity and / or clarity.

[0041] Nucleotide sequences are presented herein by single strand only, in the 5' to 3' direction, from left to right, unless specifically indicated otherwise. Nucleotides and amino acids are represented herein in the manner recommended by the IUPAC-IUB Biochemical Nomenclature Commission, or (for amino acids) by either the one-letter code, or the three-letter code, both in accordance with 37 C.F.R. §1.822 and established usage.

[0042] Except as otherwise indicated, standard methods known to those skilled in the art may be used for cloning genes, amplifying and detecting nucleic acids, and the like. Such techniques are known to those skilled in the art. See, e.g., Sambrook et al., Molecular Cloning: A Laboratory Manual 4th Ed. (Cold Spring Harbor, NY, 2012); Ausubel et al. Current Protocols in Molecular Biology (Green Publishing Associates, Inc. and John Wiley & Sons, Inc., New York).

[0043] As used herein, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, steps, operations, elements, and / or components, but do not precludeAttorney Docket No. 5470.977.WOthe presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0044] The transitional phrase “consisting essentially of’ means that the scope of a claim is to be interpreted to encompass the specified materials or steps recited in the claim and those that do not materially affect the basic and novel characteristic(s) of the claimed invention.

[0045] The term “consists essentially of’ (and grammatical variants), as applied to a polynucleotide or polypeptide sequence of this invention, means a polynucleotide or polypeptide that consists of both the recited sequence (e.g., SEQ ID NO) and a total of ten or less (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10) additional nucleotides or amino acids on the 5' and / or 3' or N-terminal and / or C-terminal ends of the recited sequence such that the function of the polynucleotide or polypeptide is not materially altered. The total of ten or less additional nucleotides or amino acids includes the total number of additional nucleotides or amino acids on both ends added together. The term “materially altered,” as applied to polynucleotides of the invention, refers to an increase or decrease in ability to express the encoded polypeptide of at least about 50% or more as compared to the expression level of a polynucleotide consisting of the recited sequence. The term “materially altered,” as applied to polypeptides of the invention, refers to an increase or decrease in enzymatic activity of at least about 50% or more as compared to the activity of a polypeptide consisting of the recited sequence.

[0046] It will be understood that, although the terms "first," "second," etc. may be used herein to describe various elements, and these elements should not be limited by these terms. These terms are only used to distinguish one element from another. Thus, a "first" element discussed below could also be termed a "second" element without departing from the teachings of the present invention. The sequence of operations (or steps) is not limited to the order presented in the claims or figures unless specifically indicated otherwise.

[0047] Furthermore, the term "about," as used herein when referring to a measurable value such as an amount of a compound or agent of this invention, dose, time, temperature, and the like, is meant to encompass variations of ± 10%, ± 5%, ± 1%, ± 0.5%, or even ± 0.1% of the specified amount.

[0048] “Amino acid sequence,” as used herein, refers to an oligopeptide, peptide, polypeptide, or protein sequence, and fragment thereof, and to naturally occurring or synthetic molecules. Where “amino acid sequence” is recited herein to refer to an amino acid sequence of a naturally occurring protein molecule, amino acid sequence, and like terms, are not meant to limit theAttorney Docket No. 5470.977.WOamino acid sequence to the complete, native amino acid sequence associated with the recited protein molecule.

[0049] “Nucleic acid,” “nucleic acid molecule,” “nucleic acid construct,” “nucleotide sequence”, “nucleic acid enzyme” and “polynucleotide”, as used herein, refers to a compound comprising a nucleobase and an acidic moiety, e.g., a nucleoside, a nucleotide, or a polymer of nucleotides. Typically, polymeric nucleic acids, e.g., nucleic acid molecules comprising three or more nucleotides are linear molecules, in which adjacent nucleotides are linked to each other via a phosphodiester linkage. In some embodiments, “nucleic acid” refers to individual nucleic acid residues (e.g., nucleotides and / or nucleosides). The term “nucleic acid” refers to an oligonucleotide chain comprising three or more individual nucleotide residues. As used herein, the terms “oligonucleotide” and “polynucleotide” can be used interchangeably to refer to a polymer of nucleotides (e.g., a string of at least three nucleotides). The term “nucleic acid” encompasses RNA as well as single and / or double-stranded DNA. Nucleic acids may be naturally occurring, for example, in the context of a genome, a transcript, an mRNA, tRNA, rRNA, siRNA, snRNA, a plasmid, cosmid, chromosome, chromatid, or other naturally occurring nucleic acid molecule. On the other hand, a nucleic acid molecule may be a non-naturally occurring molecule, e.g., a recombinant DNA or RNA, an artificial chromosome, an engineered genome, or fragment thereof, or a synthetic DNA, RNA, DNA / RNA hybrid, or including non-naturally occurring nucleotides or nucleosides. Furthermore, the terms “nucleic acid,” “DNA,” “RNA,” and / or similar terms include nucleic acid analogs, i.e., analogs having other than a phosphodiester backbone. Nucleic acids can be purified from natural sources, produced using recombinant expression systems and optionally purified, chemically synthesized, etc. Where appropriate, e.g., in the case of chemically synthesized molecules, nucleic acids can comprise nucleoside analogs such as analogs having chemically modified bases or sugars, and backbone modifications' A nucleic acid sequence is presented in the 5' to 3' direction unless otherwise indicated. The term nucleic acid enzyme is composed of single strands of either DNA (deoxyribozymes or DNAzymes) or RNA (ribozymes) that are organized into domains required for enzymatic activity (catalytic core domains) and for substrate recognition (substrate binding domains). DNAzymes and ribozymes can catalyse a broad range of chemical reactions including cleavage, ligation, phosphorylation, and deglycoslyation of RNA or DNA.

[0050] “ Gene,” as used herein, refers to a nucleic acid molecule capable of being used to produce mRNA, antisense RNA, RNAi (miRNA, siRNA, shRNA), anti-microRNA antisense oligodeoxyribonucleotide (AMO), and the like. Genes may or may not be capable of beingAttorney Docket No. 5470.977.WOused to produce a functional protein or gene product. Genes can include both coding and noncoding regions (e.g., introns, regulatory elements, promoters, enhancers, termination sequences and / or 5’ and 3’ untranslated regions). A gene may be “isolated” by which is meant a nucleic acid that is substantially or essentially free from components normally found in association with the nucleic acid in its natural state. Such components include other cellular material, culture medium from recombinant production, and / or various chemicals used in chemically synthesizing the nucleic acid.

[0051] An “isolated polynucleotide” is a nucleotide sequence (e.g., DNA or RNA) that is not immediately contiguous with nucleotide sequences with which it is immediately contiguous (one on the 5' end and one on the 3' end) in the naturally occurring genome of the organism from which it is derived. Thus, in one embodiment, an isolated nucleic acid includes some or all of the 5' non-coding (e.g., promoter) sequences that are immediately contiguous to a coding sequence. The term therefore includes, for example, a recombinant DNA that is incorporated into a vector, into an autonomously replicating plasmid or virus, or into the genomic DNA of a prokaryote or eukaryote, or which exists as a separate molecule (e.g., a cDNA or a genomic DNA fragment produced by PCR or restriction endonuclease treatment), independent of other sequences. It also includes a recombinant DNA that is part of a hybrid nucleic acid encoding an additional polypeptide or peptide sequence. An isolated polynucleotide that includes a gene is not a fragment of a chromosome that includes such gene, but rather includes the coding region and regulatory regions associated with the gene, but no additional genes naturally found on the chromosome.

[0052] The term “isolated” can refer to a nucleic acid, nucleotide sequence or polypeptide that is substantially free of cellular material, viral material, and / or culture medium (when produced by recombinant DNA techniques), or chemical precursors or other chemicals (when chemically synthesized). Moreover, an “isolated fragment” is a fragment of a nucleic acid, nucleotide sequence or polypeptide that is not naturally occurring as a fragment and would not be found in the natural state. “Isolated” does not mean that the preparation is technically pure (homogeneous), but it is sufficiently pure to provide the polypeptide or nucleic acid in a form in which it can be used for the intended purpose.

[0053] An “isolated cell” refers to a cell that is separated from other components with which it is normally associated in its natural state. For example, an isolated cell can be a cell in culture medium and / or a cell in a pharmaceutically acceptable carrier of this invention. Thus, an isolated cell can be delivered to and / or introduced into a subject. In some embodiments, anAttorney Docket No. 5470.977.WOisolated cell can be a cell that is removed from a subject and manipulated as described herein ex vivo and then returned to the subject.

[0054] A “fragment” or “portion” of a nucleotide sequence or polypeptide sequence will be understood to mean a nucleotide sequence or polypeptide of reduced length relative (e.g., reduced by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides or amino acids) to a reference nucleotide sequence or amino acid and comprising, consisting essentially of and / or consisting of a nucleotide sequence of contiguous nucleotides / amino acids identical (100% identical) or substantially identical (e.g., about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% identical) to the reference nucleotide sequence or amino acid sequence. Such a nucleic acid or amino acid “fragment” or “portion” according to the invention may be, where appropriate, included in a larger polynucleotide or amino acid of which it is a constituent.

[0055] A “heterologous” or a “recombinant” nucleic acid, as used herein, refers to a nucleotide sequence not naturally associated with a host cell into which it is introduced, including non-naturally occurring multiple copies of a naturally occurring nucleotide sequence. Alternatively, a heterologous nucleotide sequence can be one that does not naturally occur with another nucleotide sequence to which it is associated.

[0056] An “enzyme,” as used herein, refers to a biological catalyst, protein, or RNA, that speeds up the rate of a specific chemical reaction in the cell.

[0057] “Active site” or “catalytic site,” as used herein, refers to a small area, a cavity or hole on the surface of the enzyme, that catalyzes a transformation of a molecule that binds to the active site. The active site consists of 10-15 amino acid residues brought together by folding from different parts of the primary structure. One part of the active site is responsible for stereospecific binding of the substrate in proximity to the (second) transforming part of the active site.

[0058] “Transfer RNAs” (tRNAs) and “aminoacyl-tRNA” (aa-tRNA), are RNA molecules that translate mRNA into proteins. The term tRNA includes a cloverleaf structure that comprises a 3' acceptor site, 5' terminal phosphate, D arm, T arm, and anticodon arm. When an amino acid is bound to tRNA, the tRNA is considered an aminoacyl-tRNA. The type of amino acid on a tRNA is dependent on the mRNA codon destined to be matched by codon: anti codon pairing on the ribosome. The anticodon arm of the tRNA is the site of the anticodon, which is complementary to an mRNA codon and dictates which amino acid to carry. The tRNA may also be a minihelical RNA, also referred to as a “minihelix,” “RNA minihelix,” “tRNA minihelix,” or a “half-tRNA.” The minihelical RNA may comprise the acceptor stemAttorney Docket No. 5470.977.WOand / or the T C stem loop of the full-length tRNA. In some embodiments, the minihelix may be the preferred RNA substrate for AARS urzymes.

[0059] “Aminoacyl-tRNA synthetase” (AARS), as used herein, refers to a polypeptide or fragment thereof that catalyzes the aminoacylation of transfer RNAs. This activity is referred to as “charging a tRNA molecule with an amino acid”.

[0060] A “deletion,” as used herein, refers to a change in the amino acid or nucleotide sequence and results in the absence of one or more amino acid residues or nucleotides.

[0061] A “primer,” as used herein, refers to those for either amplification and / or detection, which are oligonucleotides (including naturally occurring oligonucleotides such as DNA and synthetic and / or modified oligonucleotides) of any suitable length, but are typically from 5, 6, or 8 nucleotides in length up to 40, 50, 60, 100, 150, or 200 nucleotides in length, or more. Such primers may be immobilized on or coupled to a solid support such as a bead, chip, pin, or microtiter plate well, and / or coupled to or labelled with a detectable group such as a fluorescent compound, a chemiluminescent compound, a radioactive element, or an enzyme.

[0062] A “vector,” as used herein, refers to a nucleic acid molecule capable of transporting another nucleic acid to which it has been linked (including naturally occurring vectors and synthetic and / or modified vectors). One type of vector is a “plasmid”, which refers to a circular double stranded DNA loop into which additional DNA segments may be ligated. Another type of vector is a viral vector, wherein additional DNA segments may be ligated into the viral genome. Vectors can also be supercoiled, nicked double stranded circular DNA, singlestranded circular DNA, relaxed, circular vector, or linearized double stranded vector. Certain vectors are capable of autonomous replication in a host cell into which they are introduced (e.g., bacterial vectors having a bacterial origin of replication and episomal mammalian vectors). Other vectors (e.g., non-episomal mammalian vectors) can be integrated into the genome of a host cell upon introduction into the host cell, and thereby are replicated along with the host genome. Moreover, certain vectors are capable of directing the expression of genes to which they are operatively linked. Such vectors are referred to herein as “expression vectors.” Vectors can also be, for example, a construct, transposon, cosmid, artificial chromosome, or capsid.

[0063] Vectors may be introduced into the desired cells by methods known in the art, e.g., transfection, electroporation, microinjection, transduction, cell fusion, DEAE dextran, calcium phosphate precipitation, lipofection (lysosome fusion), use of a gene gun, or a nucleic acid vector transporter (see, e.g., Wu et al., J. Biol. Chem. 267:963 (1992); Wu et al., J. Biol. Chem.263: 14621 (1988); and Hartmut et al., Canadian Patent Application No. 2,012,311, filed Mar.15, 1990).Attorney Docket No. 5470.977.WO

[0064] “Exogenous,” as used herein, in the context of nucleic acids, refers to, e.g., expression constructs, cDNAs, and nucleic acid vectors, that have artificially been introduced into a cell in which it does not naturally occur.

[0065] “Mutation,” “nucleic acid variant,” or “variant,” as used herein, refers to a variation from a known reference sequence and includes mutations such as, for example, single nucleotide variants (SNVs), copy number variants or variations (CNVs) / aberrations, insertions, or deletions (indels), truncation, gene fusions, transversions, translocations, frame shifts, duplications, repeat expansions, and / or epigenetic variants. A mutation, as used herein, refers to a deletion, and insertion, or a substitution or conversion of one or more amino acid or nucleotide to another. A mutation, as used herein, can refer to a mutation introduced into a nucleic acid sequence that, when translated, results in a mutant amino acid sequence.

[0066] A “miniature enzyme,” or “urzyme,” as used herein, refers to catalysts derived from invariant cores of protein families. Urzymes from both AARS classes are known to possess sophisticated catalytic mechanisms: pre-steady state bursts, significant transit! on- state stabilization of both amino acid activation, and tRNA acylation. Urzymology uses three-dimensional structural superposition to identify invariant cores. The miniature enzyme comprises one or more segments of the parent full-length enzyme, the amino acid sequence of each segment being at least 90% identical to the sequence of the segment in the full-length enzyme, e.g., at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical.

[0067] “Polymerase chain reaction” (PCR), as used herein, refers to the technique of treating a nucleic acid sample with one oligonucleotide primer for each strand of the specific sequence to be detected under hybridizing conditions so that an extension product of each primer is synthesized which is complementary to each nucleic acid strand, with the primers sufficiently complementary to each strand of the specific sequence to hybridize therewith so that the extension product synthesized from each primer, when it is separated from its complement, can serve as a template for synthesis of the extension product of the other primer, and then treating the sample under denaturing conditions to separate the primer extension products from their templates if the sequence or sequences to be detected are present. These steps are cyclically repeated until the desired degree of amplification is obtained. Although embodiments according to the present invention are described with respect to PCR reactions, it should be understood that other nucleic acid amplification methods can be used.

[0068] “Catalytic activity” as used herein, refers to the chemical free energy of transition state stabilization. It is proportional to the negative logarithm of the ratio, kcat / KM, of the ratio of the two steady state kinetic parameters determined from a Michaelis-Menton experiment.Attorney Docket No. 5470.977.WO

[0069] Nucleic acid molecules useful in the methods of the invention include any nucleic acid molecule that encodes an enzyme, miniaturized enzyme, amino acid residue or polypeptide of the invention or a fragment thereof. Such nucleic acid molecules need not be 100% identical with an endogenous nucleic acid sequence, but will typically exhibit substantial identity. In particular, sequences modified by protein design programs for the purpose of enhancing various properties including solubility, activity, specificity, and stability are to be considered. Nucleic acid molecules useful in the methods of the invention include any nucleic acid molecule that encodes a polypeptide of the invention or a fragment thereof.

[0070] The present invention is based on the discovery that when vectors encoding an enzyme are amplified to introduce mutations at widely separated active-site residues and transformed into an exogenous host for expression, the transformation produces, in addition to the expected full-length double mutant vector, a series of recombinant deletions in vivo. In particular, the inventors have shown that the vectors encode enzymes similar in size to a miniaturized enzyme (urzyme) and although reduced in size, the enzymes retain substantial catalytic activity. As used herein, the term “retain substantial catalytic activity” means that the miniaturized enzyme retains at least 20% of the catalytic activity of the wild-type enzyme, e.g., at least 30%, 40%, 50%, 60%, 70%, 80%, or more of the wild-type activity.

[0071] One advantage of the present invention is the deletion of long sequences in vivo which extends the range of constructs possible with analytical structure-based deconstruction alone. Another advantage is the retention of functionality and reducing the size of large and complex proteins.

[0072] Thus, one aspect of the present invention relates to a method of generating a polynucleotide encoding a miniaturized enzyme comprising: i) preparing a vector comprising a polynucleotide encoding an enzyme; wherein the enzyme comprises an active site composed of amino acid residues that are divided into two or more separate sections on the primary amino acid sequence of the enzyme; ii) hybridizing to the vector at least one pair of oligonucleotides that specifically hybridize to nucleotides encoding the amino acid residues in separate sections of the active site, wherein at least one of the oligonucleotides is designed to introduce an amino acid mutation into the enzyme; iii) performing a nucleic acid amplification reaction on the vector to produce an amplification product; iv) introducing the amplification product into a prokaryotic, archaean, or eukaryotic cell suitable for recombination; and v) isolating vectors that have undergone recombination to encode a miniaturized enzyme. In this method, the amplification reaction produces an amplicon from the vector that is then introduced into a cellAttorney Docket No. 5470.977.WOfor recombination. Vectors that have undergone recombination, which encode a miniaturized enzyme, are then isolated.

[0073] Another aspect of the present invention relates to a method of generating a polynucleotide encoding a miniaturized enzyme comprising: i) preparing a vector comprising a polynucleotide encoding an enzyme; wherein the enzyme comprises an active site composed of amino acid residues that are divided into two or more separate sections on the primary amino acid sequence of the enzyme; ii) hybridizing to the vector at least one pair of oligonucleotides that specifically hybridize to nucleotides encoding the amino acid residues in separate sections of the active site, wherein at least one of the oligonucleotides is designed to introduce an amino acid mutation into the enzyme; iii) performing a first nucleic acid amplification reaction on the vector to produce a first amplification product; iv) combining the first amplification product with the vector of step i); v) performing a second nucleic acid amplification reaction on the vector to produce a second amplification product; vi) introducing the second amplification product into a prokaryotic, archaean, or eukaryotic cell suitable for recombination; and vii) isolating vectors that have undergone recombination to encode a miniaturized enzyme. In this method, the oligonucleotides are hybridized to the vector, and the first amplification step produces a megaprimer. Then, the second amplification step uses the megaprimer to amplify the vector. The resulting amplicon is a vector that is then introduced into a cell for recombination. Vectors that have undergone recombination, which encode a miniaturized enzyme, are then isolated.

[0074] Another aspect of the present invention relates to a method of generating a polynucleotide encoding a miniaturized enzyme comprising: i) preparing a vector comprising a polynucleotide encoding an enzyme; wherein the enzyme comprises an active site composed of amino acid residues that are divided into two or more separate sections on the primary amino acid sequence of the enzyme; ii) hybridizing to the vector a first set of oligonucleotides, wherein the first set of oligonucleotides comprises at least two pairs of oligonucleotides that specifically hybridize to nucleotides encoding the amino acid residues in separate sections of the active site, wherein at least one of the at least two pairs of oligonucleotides is designed to introduce an amino acid mutation into the enzyme; iii) performing a first nucleic acid amplification reaction on the vector to produce a first set of amplification products, wherein the first set of amplification products comprises at least two amplification products, wherein the at least two amplification products of the first set of amplification products overlap at a 5’ end or 3’ end; iv) combining the first set of amplification products with a second set of oligonucleotides (e.g., a subset of the oligonucleotides from the first set of oligonucleotides, or differentAttorney Docket No. 5470.977.WOoligonucleotides); v) performing a second nucleic acid amplification reaction on the combination of step iv) to produce a second set of amplification products, wherein the second set of amplification products comprises at least one amplification product (e.g., an amplification product that differs from the amplification products in step iii); vi) combining the second set of amplification products with the vector of step i); vii) performing a third nucleic acid amplification reaction on the combination of step vi); viii) introducing the amplification product of step vii) into a prokaryotic, archaean, or eukaryotic cell suitable for recombination; and ix) isolating vectors that have undergone recombination to encode a miniaturized enzyme. In this method, the oligonucleotides are hybridized to the vector, and the first amplification step produces multiple fragments. Then, the second amplification step uses the fragments to produce a megaprimer. The megaprimer is used in the third amplification step to amplify the vector. The resulting amplicon is then introduced to a cell for recombination. Vectors that have undergone recombination, which encode a miniaturized enzyme, are then isolated.

[0075] In some embodiments, the oligonucleotides bind to one or more catalytic sites. In some embodiments, the catalytic sites can be highly conserved catalytic sites. In some embodiments, e.g., in leucyl-tRNA synthetase (LeuRS), the oligonucleotides encode amino acid residues AVGA (SEQ ID NO:4) and AMSAS (SEQ ID NO:5) and / or a mutation within amino acid residues AVGA (SEQ ID NO:4) and AMSAS (SEQ ID NO:5). In some embodiments, e.g., in pyruvate kinase, the oligonucleotides encode amino acid residues 38H, 74R, 217A, 218K, 220E, and 221K and / or mutations 38L, 74A, 217Q, 218T, 220A, and 221A. In some embodiments, the first set of oligonucleotides will comprise oligonucleotides that are the same as oligonucleotides in the second set of oligonucleotides. In some embodiments, the first set of oligonucleotides will comprise oligonucleotides that are different from the oligonucleotides in the second set of oligonucleotides.

[0076] In another aspect, the present invention relates to a composition comprising a vector comprising a polynucleotide encoding an enzyme; wherein the enzyme comprises an active site composed of amino acid residues that are divided into two or more separate sections on the primary amino acid sequence of the enzyme; and at least one pair of oligonucleotides hybridized to the vector, wherein the oligonucleotides specifically hybridize to nucleotides encoding amino acid residues in separate sections of the active site, and wherein at least one of the oligonucleotides is designed to introduce an amino acid mutation into the enzyme.

[0077] The vector may be any type of vector that is suitable for introducing a polynucleotide into a cell. In some embodiments, the vector is a non-viral vector, e.g., a plasmid or cosmid. In some embodiments, the vector is a viral vector, e.g., retrovirus, lentivirus, adeno-associatedAttorney Docket No. 5470.977.WOvirus, poxvirus, alphavirus, baculovirus, vaccinia virus, herpes virus, Epstein-Barr virus, and / or adenovirus vectors.

[0078] In some embodiments, the method comprises hybridizing to the vector at least two pairs of oligonucleotides (e.g., at least 3, 4, 5, or 6 pairs of oligonucleotides) that specifically hybridize to nucleotides encoding the amino acid residues in separate sections of the active site of the enzyme. In some embodiments, one or more of the oligonucleotides is a primer, e.g., for an amplification reaction. In some embodiments, each oligonucleotide may be about 10-50 nucleotides in length, e.g., about 15-30 nucleotides in length. Each oligonucleotide may be at least 80% complementary to the nucleic acid sequence encoding the enzyme, e.g., at least 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% complementary. The level of complementarity should be sufficiently high for the oligonucleotide to specifically hybridize to the intended target site.

[0079] In some embodiments, at least two pairs of oligonucleotides that specifically hybridize with nucleotides encoding the amino acid residues in separate sections of the active site are hybridized to the vector. In some embodiments, the oligonucleotides that specifically hybridize to nucleotides encoding the amino acid residues in separate sections of the active site create mutations encompassing the catalytic site.

[0080] One or more of the oligonucleotides is designed to introduce an amino acid mutation into the enzyme. The mutation may be one or more substitution, insertion, deletion, or any combination thereof. The mutation to be introduced may be selected for any purpose, e.g., to introduce or delete an active site signature or to make synonymous codon changes. In some embodiments, PCR mutagenesis may be used, which relates to site-directed mutagenesis. Custom design oligonucleotide primers, an in vitro procedure, may be used to confer a desired mutation. Oligonucleotide primers can be designed in either an overlapping or a back-to-back orientation. Another example relates to an oligonucleotide-directed mutagenesis technique that makes use of synthetic oligonucleotides that share homology with a target sequence(s) with the exception of the nucleotide(s) to be modified. When PCR is used for site-directed mutagenesis, the oligonucleotide primers are designed to include the desired change, which could be a base substitution, addition, deletion, or any combination thereof. The mutation is incorporated into the amplicon during the PCR protocol, replacing the original sequence. Site-directed mutagenesis by primer extension involves incorporating mutagenic primers in independent, nested PCRs before combining them in the final product. The reaction requires flanking primers to be positioned on either side of the region to be deleted. Inverse PCR mutagenesis enables amplification of a region of unknown sequence using primers oriented in the reverse direction.Attorney Docket No. 5470.977.WOAn adaptation of this method can be used to introduce mutations in previously cloned sequences. Using primers incorporating the desired change, for example, an entire circular plasmid may be amplified to delete, substitute, or insert the desired sequence.

[0081] In some embodiments, PCR mutagenesis may be used to create at least two amplification products. In some embodiments, the at least two amplification products will be joined via a second amplification reaction (e.g., overlap extension polymerase chain reaction (OE-PCR)). As used herein, OE-PCR is used interchangeably with fragment-overlapping PCR ligation and overlapping PCR ligation.

[0082] In some embodiments, the enzyme comprises an active site composed of amino acid residues that are divided into two or more separate sections (e.g., 3, 4, or more separate sections) on the primary amino acid sequence of the enzyme. In some embodiments, the separate sections of the active site may be at least about 10 amino acids apart, e.g., at least 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200 or more amino acids apart. In many enzymes the catalytic site or active site is constituted within two or more domains. It is mutation of the two or more domains that leads the cell, e.g., E. coli, to generate miniature genes with the same or comparable functionality.

[0083] The enzyme may be any suitable enzyme for which a miniaturized version is desired and that comprises an active site composed of amino acid residues that are divided into two or more separate sections on the primary amino acid sequence of the enzyme. In some embodiments, the enzyme is an AARS, e.g., a Class I tRNA synthetase. In other embodiments, the enzyme is pyruvate kinase.

[0084] The nucleic acid amplification may be any suitable amplification technique known in the art. Examples include, without limitation, polymerase chain reaction (PCR), rolling circle amplification (RCA), duplex-specific nuclease (DSN)-based amplification, loop-mediated isothermal amplification (LAMP), exponential amplification reaction (EXPAR), stranddisplacement amplification (SDA) and some enzyme-free amplifications. In some embodiments, the nucleic acid amplification reaction is a PCR.

[0085] In one embodiment, the product from the amplification reaction is introduced into a prokaryotic, archaean, or eukaryotic cell suitable for recombination. For example, a mechanism used for introducing an amplification product into a cell is transformation. Transformation is a process in which the cells take up foreign DNA from the environment. One example is bacterial transformation, which commonly uses a plasmid to carry a gene of interest and introduce it into a bacterial cell. During transformation, the bacterial cell takes up the plasmid and will then contain the genes in that plasmid and can express those genes. Other methods used forAttorney Docket No. 5470.977.WOintroducing an amplification product include microinjection, biolistic or gene gun, alternate cooling and heating, and use of calcium ions.

[0086] Cells suitable for recombination are prokaryotic cells, which are single-celled microorganisms and include bacteria and archaea, and eukaryotic cells, which have a nucleus enclosed within the nuclear membrane and form large and complex organisms and include protozoa, fungi, plants, and animals. In some embodiments, the cell is a bacterium such as E. coli. In some embodiments, the cell is a yeast cell.

[0087] In another aspect of the invention, the recombination results in deletion of nucleotides. In some embodiments, the deletion of nucleotides comprises deletion of a single nucleotide. In some embodiments, the deletion of nucleotides comprises deletion of a short nucleotide sequence (e.g., a nucleotide sequence comprising 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, or 20 to 50, 60, 70, 80, 90, or 100 nucleotides). In some embodiments, the deletion of nucleotides comprises a deletion of a long nucleotide sequence (e.g., a nucleotide sequence comprising 101, 200, 300, 400, or 500, to 600, 700, 800, 900, 1000, 1250, 1500, 1750, 2000, 3000, 4000, 5000, 7500, or 10000 nucleotides). In some embodiments, the deletion of nucleotides comprises a series of deletions (e.g., a deletion of more than one segment of nucleotides). The series of deletions may be a nested series of deletions or a non-nested series of deletions. When the amplification product (e.g., recombinant DNA or RNA) is introduced into a recipient host cell, it gets multiplied and is expressed. Expression enables determination of the nucleotide sequences, determination of series of deletions, and obtaining large amounts of proteins for structural and functional characteristics.

[0088] In an additional embodiment, the enzyme is an aminoacyl-tRNA synthetase (AARS). For example, the AARS can be one of the twenty -three AARSs identified; one for each of the 20 proteinogenic amino acids (except for lysine, for which there are two) plus pyrrolysyl-tRNA synthetase (PylRS) and phosphoseryl-tRNA synthetase (SepRS), enzymes with a more restricted distribution that are only found in some bacterial and archaeal genomes. In another embodiment, the AARS is leucyl-tRNA synthetase (LeuRS). In an additional embodiment, the enzyme is pyruvate kinase.

[0089] In another embodiment, the at least two oligonucleotides encode amino acid residues AVGA (SEQ ID NO:4) and AMSAS (SEQ ID NO:5). For example, antiparallel AVGA (SEQ ID NO:4) and AMSAS (SEQ ID NO: 5) mutant primers are used simultaneously in PCR mutagenesis, which is illustrated in FIG. 3.

[0090] Another aspect of the invention relates to a composition comprising the vector of the invention, e.g., a vector comprising a polynucleotide encoding an enzyme; wherein the enzymeAttorney Docket No. 5470.977.WOcomprises an active site composed of amino acid residues that are divided into two or more separate sections on the primary amino acid sequence of the enzyme; and at least one pair of oligonucleotides hybridized to the vector, wherein the oligonucleotides specifically hybridize to nucleotides encoding amino acid residues in separate sections of the active site, and wherein at least one of the oligonucleotides is designed to introduce an amino acid mutation into the enzyme. In some embodiments, the oligonucleotides are primers. In some embodiments, all of the oligonucleotides are designed to introduce an amino acid mutation into the enzyme. In some embodiments, the enzyme is an AARS (e.g., LeuRS) or pyruvate kinase. In some embodiments, the at least two oligonucleotides encode amino acid residues AVGA (SEQ ID NO:4) and AMSAS (SEQ ID NO:5)

[0091] A further aspect of the invention relates to an amplification product produced from performing an amplification reaction on the composition of the invention. In some embodiments, the amplification reaction is a PCR.

[0092] An additional aspect of the invention relates to a prokaryotic, archaean, or eukaryotic cell comprising the amplification product of the invention. In some embodiments, the cell is a bacterium, e.g., E. coli. In some embodiments, the cell is a yeast or mammalian cell.

[0093] A further aspect of the invention relates to a vector encoding a miniaturized enzyme, produced by the method of the invention. In another embodiment, the vectors that have undergone recombination to encode a miniaturized enzyme are isolated. The vectors may be isolated using any technique known in the art, e.g., using standard commercial purification kits. Miniaturized enzymes are catalysts derived from invariant cores of protein families. These miniaturized enzymes can be used, without limitation, to approximate the evolutionary history of a particular gene, gene therapy, useful in treating inborn errors of metabolism, protein engineering, gene delivery, and aid research into the evolution of the proteome.

[0094] One use for isolating many such recombinant plasmids is to accumulate statistics on the frequency with which each class of recombinant deletion occurs. Without being bound by theory, it is believed that the extent of the recombinant deletions is, itself a novel kind of evidence about how the smallest deletions came to assimilate new, advantageous coding information. That is to say, it is believed the pattern of deletions may reflect the evolutionary history of modular additions leading to the eventual evolutionary maturation of the enzyme family into its contemporary form. If that turns out to be true, it introduces a new window on the ancient molecular history of the gene family. Thus, it is ultimately likely to be transformative.Attorney Docket No. 5470.977.WO

[0095] The present invention is more particularly described in the following examples that are intended as illustrative only since numerous modifications and variations therein will be apparent to those skilled in the art.EXAMPLES

[0096] Deconstruction based on 3D superposition computer modeling: First, the 3D coordinates of X-ray structures for the active sites of all 10 Class I AARS were superimposed, which highlighted a small (130-residue) set of secondary structures that were nearly the same in all 10 families, in contrast to the remaining variable segments representing 40-85% of the total length. The sharp contrast between the conserved and variable components narrowed the focus to segments that could be aligned according to the Rodin-Ohno hypothesis. For Class II AARS, the conserved segments were continuous and contained Motifs 1 and 2 but were missing Motif 3 1. Class I AARS cores in contrast were interrupted at one or more places with long and variable insertions. The insertion points invariably occurred were the entering and leaving chains could be bridged by a single peptide bond. As such, a shortened gene for the tryptophanyl-tRNA synthetase (TrpRS) was created. That construct is illustrated in FIG. 2.

[0097] Conserved active-site amino acids contribute to catalysis only in the full-length AARS: Relevant to the invention is the observation that highly conserved active-site residues that contribute substantially to the increased catalytic activity of full-length enzymes do not play so significant a role in AARS urzymes. Class I AARS have two active-site signatures — HVGH (SEQ ID NO:6) and KMSKS (SEQ ID NO:3) — that together contribute ~5 kcal / mole of transition-state stabilization free energy in the full length LeuRS. Notwithstanding, the AVGA (SEQ ID NO:4) and AMSAS (SEQ ID NO:5) mutations in the LeuRS urzyme, LeuAC have little effect or in fact enhance catalysis 60-fold in the case of AVGA (SEQ ID NO:4).Biosynthesis of both histidine and lysine require an unusually large number of different enzymes, and neither is produced in Miller syntheses. Thus, it is believed to have become incorporated into the genetic code long after a primordial code was established.

[0098] These AARS urzymes contain little more than the active site and can be as small as 15% of the length of the corresponding full-length enzymes. For example, a resultant urzyme can be approximately 130 amino acids, whereas the original full-length enzyme is 967 amino acids long, such that it is about 13% of the full length, its “Goldilocks Zone” is about 82 amino acids long, about 8.5% of the full length. They retain >60% of the catalytic proficiency for both of the reactions by which the full-length enzymes translate the genetic code. A smaller construct called the protozyme and containing only the ATP binding determinant retains 40% of theAttorney Docket No. 5470.977.WOcatalytic proficiency for ATP-dependent amino acid activation. The analytical process used entails constructing genes after: (i) superimposing 3D coordinates for related enzymes and identifying highly conserved secondary structures and (ii) removing long, structurally variable inserted segments that intersect the conserved region at residues whose backbone residues can be connected by a single peptide bond. Those inserted segments greatly increase the distance between different structural components of the active site.

[0099] Assays of AARS function: Relevant to the invention are the conventional assays for AARS functionality. Assays include amino acid activation (pyrophosphate exchange, single turnover active-site titration); aminoacylation of RNA substrates, both tRNA and T'PCminihelix; and Michaelis-Menten assays for both reactions. The latter afford the opportunity to measure the steady-state kinetic parameters for both types of substrates. These assays must be carried out with higher concentrations of enzyme and for longer times, because of the reduced catalytic proficiency of AARS urzymes.Example 1: An in vivo production of the minigene DNA sequences encoding highly active miniaturized enzymes1. Materials

[0100] The plasmid vector was pETlla (Novagen, Sacramento, CA, United States). DNA oligos were ordered from IDT (Integrated DNA Technologies, Coralville, IA, United States). The detailed sequences of oligo DNAs are provided in the Table 1 below. Phusion™ Plus PCR Master Mix (Cat# F631 S) was purchased from Thermo Fisher Scientific (Waltham, MA, United States). E coli competent cells were partially from Agilent (XLIO-Gold ultracompetent cells, Cat# 200315, Santa Clara, CA, United States), and partially using DH5a, home-made, prepared according to the method described by Sharma in 2017. Restriction enzyme Dpnl (Cat# 500402, United States) was purchased from Agilent. Purifications and handling of DNA fragment and plasmid used the QIAquick Gel Extraction Kit (Cat# 28704, Qiagen, Hilden, Germany) and QIAprep Spin Miniprep Kit (Cat# 27104, Qiagen, Hilden, Germany) according to instruction, unless stated otherwise.Table 1Attorney Docket No. 5470.977.WOAttorney Docket No. 5470.977.WO2. MethodsAttorney Docket No. 5470.977.WO2-1. Method for preparation of double mutant plasmid

[0101] DNA oligo sequences were designed according to the manual of QuikChange II-E Site-Directed Mutagenesis Kit (Santa Clara, CA, United States) as published online at agilent.com / cs / library / usermanuals / public / 200555.pdf. In line with the oligo DNA design, and for creating the double mutant plasmid DNA, a megaprimer PCR method described by Picard in 1994 was implemented, and a plasmid was created that contained a double mutant full length LeuRS gene. A clone of full-length LeuRS from Pyrococcus horikoshii in pET-1 la provided the starting material for the procedure. A doubly mutant LeuRS was prepared using antiparallel AVGA (SEQ ID NO:4) and AMSAS (SEQ ID NO:5) mutant primers simultaneously in PCR mutagenesis, as illustrated in FIG. 3.

[0102] Specifically, a megaprimer was first prepared in a first round of PCR. Phusion™ Plus PCR Master Mix (Cat# F631S) was employed according to instruction from the manufacturer (ThermoFisher Scientific, Waltham, MA, United States). The primer pair for the first round PCR was LeuRS AVGAal F (SEQ ID NO: 10) and LeuRS=AMSASbl_R (SEQ ID NO: 11) as listed in Table 1 above. In the PCR reaction mix, the plasmid template concentration was 1-4 ng / pL. The primers were used at a concentration of 0.66 pM. The PCR was started at 98°C, 30 seconds for initial template denaturation and was following by 44 cycles (98°C, 12 seconds; 61 °C, 30 seconds; 72°C, 3 minutes 20 seconds). After the final extension at 72°C, 5 minutes, the PCR reactions were kept at 4°C overnight. After the PCR run, following by horizontal electrophoresis in a IxTAE, 1.2% Agarose-gel, and visualization by UV transmitter, the megaprimer fragment from the first round of PCR was collected from gel cutting out the desired band using a clean razor blade. After gel extraction purification procedures using QIAquick Gel Extraction Kit (Cat# 28704, Hilden, Germany), to yield the megaprimer for the downstream operations.

[0103] In the second phase, the entire plasmid DNA was amplified using the megaprimer to generate the resultant plasmid harboring the mutant. A second round of PCR was performed using Phusion™ Plus PCR Master Mix (Cat#F631S) and the parameters were setup according to the instructions from the manufacturer (ThermoFisher Scientific, Waltham, MA). The PCR procedures were the same as the first PCR run. Additionally, instead of oligo primers, the megaprimer was added at specific desired concentration. For example, in a 25-40 pL PCR reaction volume, the amount of megaprimer fragment was within a range of 0.16 pM-1.25 pM, and the template plasmid DNA was 0.5-1.5 ng.Attorney Docket No. 5470.977.WO

[0104] After Dpnl digestion of the PCR mixture at 37°C from hours to overnight according to instructions provided by Agilent (Santa Clara, CA, United States), the DpnI-digested PCR mixture is ready for E. coli. transformation. The transformation was done using 0.3 pL DpnI-digested PCR mixture per transformation.2-2. Method for E. coli. transformation and the sample submission for sequencing

[0105] Subsequently, E. coli wa transformed with the PCR mix, which produced a reproducible spectrum of different combinations of the various modules, as illustrated in FIGS.4A-4D. Specifically, E coli transformation was performed according to conventional procedure (Molecular Cloning: A Laboratory Manual. 1982. Book by E. F. Fritsch, Joseph Sambrook, and Tom Maniatis), and spread onto LB (Lysogeny broth) agar plates with antibiotics (Ampicillin 50 pg / mL, or Carbenicillin 100 pg / mL). After incubation overnight at 37°C, individual colonies on the LB agar plate were readied for miniprep by following the standard protocol described by QIAprep Spin Miniprep Kit (Cat# 27104, Qiagen, Hilden, Germany).

[0106] The resultant plasmid DNA samples were subjected to sequencing by following standard sequencing sample preparation and submission described by Eton Bioscience (San Diego, CA) online (genewiz.com / en / Public / Resources / Sample-Submission-Guidelines / Sanger-Sequencing-Sample-Submission-Guidelines / Sample-Preparation#sanger-sequence) with minor modifications. For example, each sequencing sample was in 15 uL, containing 100-200 ng plasmid DNA with 5-10 picomole sequencing primer, i.e., seqnl_AVGA (SEQ ID NO: 12) or AMSAS_seqn2rv (SEQ ID NO: 13) as detailed sequences in Table 1 above.2-3. Method for plasmid sequencing and sequencing analysis

[0107] A sequence analysis was conducted for about 50 plasmids isolated by comparable procedures over the course of several months, which is illustrated in FIGS.5A-5C. Specifically, all the DNA sequences were translated into six reading frames in amino acid sequence using Expasy translate online at web.expasy.org / translate / . For screening the tentative urzyme candidate, amino acid sequences from six reading frames from a single sequenced plasmid DNA sample were aligned with the sequence of urzyme template using conventional sequencing alignment software such as MOFFT online at ebi.ac.uk / jdispatcher / msa / mafft?stype=protein (Katch K, et al., 2019) or MUSCLE online at ebi.ac.uk / jdispatcher / msa / muscle?stype=protein (Edgar RC, 2004). The list of criteria for selection of tentative candidate was according to the sequence length, similarity and identify in addition to the highly conserved catalytic motifs such as motif I HVGH (SEQ ID NO:6) and motif II KMSKS (SEQ ID NO:3) or their-like.Attorney Docket No. 5470.977.WO

[0108] The resulting new candidates were designated urzyme, urzyme-like and goldilocks, according to their sequence content and were collected and subjected for further downstream identification at the enzymology level after protein biochemistry procedures.3. Results

[0109] When plasmids containing double mutations of the widely separated active site residues are transformed back into E. coli for expression, the double mutant plasmids are unstable. Transformation reproducibly produces, in addition to the expected full-length double mutant plasmids, a nested series of recombinant deletions. One of these is similar to the urzyme created in vitro, by the procedure outline above (“LeuAC-like,” FIG. 5B). Another matches the first half of the protozyme, mentioned above, in length (“Half-protozyme,” FIG. 5B). A third is intermediate in size between the protozyme and the urzyme (“Goldilocks,” FIG. 5B).

[0110] We observed the retained coding sequences in all six reading frames of the sequenced plasmids, and so refer to them as orphaned coding sequences (OCSs). Yet, when aligned to the full length LeuRS gene reading frame, they formed a coherent alignment, shown in FIG. 5A.Genomic features, together with the temperature at which they were incubated and the respective reading frames are shown schematically in FIG. 5C. A neighbor joining tree based on the alignment in FIG. 5A reveals only three clades as shown in FIG. 5B. We observed each clade in FIG. 5B in each of the six reading frames. Thus, whatever process generated the deletions had no obvious local sequence specificity. Second, and in contrast to the lack of positional specificity, all OCSs deleted the same major segments from the full-length LeuRS gene. Notably, there are no remnants of the two large domains we deleted to generate LeuAC. The first missing domain is the 444-residue Connecting Peptide 1 (CPI; residues 83-527 of SEQ ID NO: 71) that provides the editing function. The second domain missing in all deletions is the C-terminal 304-residue anticodon-binding domain (ABD; residues 663-967 of SEQ ID NO: 71). These domains are precisely the same domains that we deleted to construct the LeuAC urzyme in vitro (FIG. 4D)[OHl] The coding sequences retained in the shortened OCSs comprise three discrete clades. Each clade differs in a distinct way from the analytical LeuAC urzyme (Carter, et al. 2014; Hobson, et al. 2022) we studied earlier (Tang, et al. 2023; Tang, et al. 2024). The shortest subset retains only a 28-residue segment (residues 41-68 of SEQ ID NO: 71) containing the AVGA (SEQ ID NO:4) sequence and which forms the ATP binding site. About 11% (8) of the 34 sequences encode truncated protozymes, retaining only these sequences. We refer to this set asAttorney Docket No. 5470.977.WOthe half-protozyme (from “Proto”, relating to a precursor). These sequences are indicated by the “$” region in FIG. 5A. These residues are present in nearly all deletions.

[0112] A longer subset contains, in addition, to the intact, 50-residue protozyme, residues 650-674 of SEQ ID NO: 71 that contain the AMSAS motif (SEQ ID NO:5). This 25-residue segment is denoted with in FIG.5A. Although it contains the entire length of the protozyme, it lacks all of the specificity helix (residues 537-552 of SEQ ID NO: 71), a second leucinespecific connecting peptide (residues 552-604 of SEQ ID NO: 71), and the second crossover connection of the Rossmann fold (residues 605-635 of SEQ ID NO: 71). It thus is the shorter of the two deletions that contain both mutant catalytic signature motifs. About -10% (7) of the sequences in FIG. 5A have this intermediate length. It is -81 residues long. We expressed this OCS as the “Goldilocks” urzyme We call the 81 -residue OCS the “Goldilocks” urzyme because it is “just right” for antiparallel alignment with the shortest Class II urzymes (Patra, et al. 2024). Its high catalytic proficiency of aminoacylation implies that catalyzing acyl transfer to RNA substrates requires only the Protozyme and a short, 25-residue C-terminal segment containing the KMSKS (SEQ ID NO: 3) / AMSAS (SEQ ID NO:5) signature.

[0113] The longest OCS in FIG.5A (300 - 450nt; about 19% (14) of the sequences), includes residues 588-636 of SEQ ID NO: 71 which code for the second crossover of the Rossmann fold (which are denoted by “#” regions in FIG. 5A). That segment contributes a third motif, GKDL (SEQ ID NO: 14) (GxDQ in other Class I AARS) to the active site. Without wishing to be bound by theory, we presume that the aspartate side chain coordinates the active-site Mg++ion transiently during catalysis by full-length LeuRS (Weinreb and Carter 2008; Weinreb, et al.2009; Williams, et al. 2015). Curiously, these sequences also have only a half-protozyme. They also have an extra segment that is part of the leucine-specific CP2, but not part of LeuAC (underlined region under the sequence numbers). We refer to this OCS as the “LeuAC-like urzyme”.

[0114] Curiously, the lengths of the half-protozyme OCSs match exactly that portion of the protozyme in the LeuAC-like OCSs. The protozyme segment is twice as long in the intermediate-sized OCS which lacks the entire second crossover segment of the Rossmann fold. Each clade in the ensemble is remarkably discrete and reproducible. Occasional point mutations that differentiate the different truncated plasmids from one another are not highlighted.

[0115] Controls were performed to test whether or not the reverse process of restoring the wild type enzyme sequence by PCR mutagenesis of the double mutant plasmid. The controls were able to establish that recombinant deletions occurred in vivo, as none of the controls produced recombinant deletions. The experiments conducted, together with other controls inAttorney Docket No. 5470.977.WOwhich the output from PCR was selected for full-length plasmids before transformation of E. coli established to a high degree of certainty that the recombinant deletions occurred after transformation, in vivo, and not during the PCR. Without wishing to be bound by theory, we note here that they are likely coupled to DNA replication because mRNA synthesis from the pET-1 la host plasmid is minimal without induction. Notably, those incubated exclusively at 37° C had no Goldilocks deletions and fewer tandemly repeated and multi -frame encoding than did those at 4° C. The enrichment of deletions at low temperature (FIG. 5C) is consistent with the known accessibility of D-loops to recombination proteins and replication restart at low temperature (Michel and Sandler 2017).

[0116] These findings suggest that E. coli uses recombinant deletions of an explicit set of segments to reduce stresses experienced by the presence of a foreign variant of a gene required for proper cellular function, especially when the presence of an antibiotic resistance gene assures that the plasmid cannot be deleted in its entirety. Moreover, the nested recombinant deletions correspond almost exactly the same as the nested set that was created by application of structural biology and analytical deconstruction. Whenever the mutant gene had alanine mutations to both HVGH (SEQ ID NO:6) (resulting in AVGA (SEQ ID NO:4)) and KMSKS (SEQ ID NO:3) (resulting in AMSAS (SEQ ID NO:5)) signatures (Tang, et al. 2023), E. coli deleted both the long, connecting peptide 1 insertion and the C-terminal anticodon-binding domain. These are the very segments we initially deleted in designing AARS urzymes and protozymes (Pham, Li, Kim, Erdogan, et al. 2007; Pham, et al. 2010). The highly limited spectrum of deletions suggests, in turn, that the nested set is in some sense pre-ordained, and that such preordination may reflect three-dimensional contacts that bring regions of the coding nucleic acid close enough together to favor specific recombinational deletion events.

[0117] AARS urzymes retained three catalytic properties seen in full length AARS. (i) They accelerate amino acid activation, (ii) They show significant pre-steady state bursts in singleturnover kinetic assays (Fersht, et al. 1975; Francklyn, et al. 2008). (iii) They accelerate aminoacylation of tRNA and minihelix substrates (Tang, et al. 2024). The Goldilocks and LeuAC-like urzymes (FIG.5B), retain all three catalytic activities at levels comparable to those observed for the urzymes generated intentionally on the basis of structural biology (Carter 2014). Experimental kinetics data confirm all three criteria (FIGS. 6A-6F and 7A-7D).Example 2: An in vivo production of the minigene DNA sequences encoding highly active miniaturized enzymes from pyruvate kinase (PK)1. MaterialsAttorney Docket No. 5470.977.WO

[0118] The plasmid vector was pET28a(+) (Novagen, Sacramento, CA, United States). DNA oligos were ordered from IDT (Integrated DNA Technologies, Coralville, IA, United States). The detailed sequences of the oligo DNAs are provided in Table 2 below. Phusion™ Plus PCR Master Mix (Cat# F631S) and DreamTaq PCR Master Mixes (2X) (Cat# KI 081) were purchased from Thermo Fisher Scientific (Waltham, MA, United States). Phusion™ High-Fidelity PCR Master Mix with HF Buffer (Cat#M0531S, Ipswich, MA 01938-2723). E coli competent cells used were DH5a (for cloning) and BL21CODONPLUS DE3RIPL (for expression), home-made, and prepared according to the methods reported by Sharma, et al.2017 and described by W. Zwerschke's PhD thesis, DKFZ, Heidelberg, Germany, 1997 with care about cold conditions for all supplies including pre-cold-tips and pipettes. The expression E coli strain BL21(DE3)pLysS was purchased from Promega (Madison, WI 53719, USA). Restriction enzyme Dpnl was from Thermo Fisher (Cat# ER1701, United States). Purifications and handling of DNA fragment and plasmid used QIAquick Gel Extraction Kit (Cat# 28704, Qiagen, Hilden, Germany) and QIAprep Spin Miniprep Kit (Cat# 27104, Qiagen, Hilden, Germany) according to instruction, unless stated otherwise. DNA sequencing including Sanger sequencing was by Eton BioSciences (San Diego, CA92121). Initial construct plasmid DNA was synthesized by Twist BioScience (South San Francisco CA 94080).Table 2>Attorney Docket No. 5470.977.WOAttorney Docket No. 5470.977.WOAttorney Docket No. 5470.977.WO2. Methods2-1. Method for preparation of triple mutant plasmid

[0119] DNA oligo sequences were designed according to the manual of QuikChange II-E Site-Directed Mutagenesis Kit (Santa Clara, CA, United States) as published online at agilent.com / cs / library / usermanuals / public / 200555.pdf. In line with the oligo DNA design, and for creating double mutant plasmid DNA, a megaprimer PCR method described by Picard in 1994 was implemented, and a plasmid was created that contained a double mutant full length PK gene. A clone of full-length PK from M. marinarum in pET28a(+) provided the starting material for the procedure. A triple mutant PK was prepared using antiparallel mutant primers simultaneously in PCR mutagenesis.Fragment-overlapping PCR Ligation for Megaprimer Acquisition

[0120] For megaprimer acquisition when multiple catalytic sites are embedded in a given gene, as in PK, a fragment-overlapping PCR ligation is performed as in FIG. 11. This is a key technique for triggering “reverse evolution” via megaprimer PCR over a sequence having multiple catalytic sites or catalytic signatures.

[0121] First, DNA fragments were prepared using PCR with Phusion Plus PCR mix (Cat#F631S) according to manufacturer’s instruction (ThermoFisher Scientific). The primerAttorney Docket No. 5470.977.WOpairs utilized are in Table 3 below. Primer pairs were designed such that each fragmentamplicon was greater than 200 bp, or such that the distance between two catalytic signatures was greater than 200 bp. The primers were also designed such that the resulting megaprimer following fragment-overlapping PCR ligation would be less than lOkb. Because of the high GC content of the PK sequence, high-fidelity polymerases were needed (FIG. 14).Table 3.

[0122] After PCR and electrophoresis using 2%Agarose IxTAE gel, the individual desired bands were excised from gel and subjected to purification using QIAquick Gel Extraction Kit (Cat# 28704, Hilden, Germany).

[0123] After all DNA fragment preparations were purified, PCR reaction for megaprimer construction was carried out. At 3-10ng / reaction in a PCR tube, all the three purified PCRAttorney Docket No. 5470.977.WOfragments were mixed with the following primers as shown in Table 4. The running PCR was followed using the instruction from Thermo Fisher Scientific as above.Table 4.

[0124] After PCR, the product was subjected to electrophoresis using 2% Agarose IxTAE, and the resultant megaprimer band was eluted from gel extraction using the Qiagen kit as above.

[0125] The megaprimer was then used in an amplification of entire plasmid DNA to generate the resultant plasmid harboring the three catalytic site mutations. The megaprimer round of PCR was performed using Phusion™ Plus PCR Master Mix (Cat#F631S) and with parameters according to the instructions from the manufacturer (ThermoFisher Scientific, Waltham, MA), except that, instead of oligo primers, the megaprimer was added at a desired concentration. Specifically, in a 25-40 L PCR reaction volume, the megaprimer was added in an amount between 0.16 M-1.25 M, and the template plasmid DNA was added in an amount between 0.5-1.5ng.

[0126] After Dpnl digestion of the PCR mixture at 37°C from hours to overnight according to instructions provided by Agilent (Santa Clara, CA, United States), the DpnI-digested PCR mixture was ready for E coli transformation.2-2. Method for E. coli. transformation and the sample submission for sequencing

[0127] E coli transformation was proceeded according to conventional procedure (Molecular Cloning: A Laboratory Manual. 1982. E. F. Fritsch, Joseph Sambrook, and Tom Maniatis), and spread onto LB (Lysogeny broth) agar plates with antibiotics (Ampicillin 50pg / mL, or Carbenicillin IOOpg / mL). After incubation overnight to 72hrs at 37°C, individual colonies onAttorney Docket No. 5470.977.WOthe LB agar plate are ready for miniprep by following the standard protocol described by QIAprep Spin Miniprep Kit (Cat# 27104, Qiagen, Hilden, Germany).

[0128] The resultant plasmid DNA samples were subjected to sequencing by following standard sequencing sample preparation and submission described by Eton Bioscience (San Diego, CA) online at genewiz.com / en / Public / Resources / Sample-Submission-Guidelines / Sanger-Sequencing-Sample-Submission-Guidelines / Sample-Preparation#sanger-sequence with minor modifications. For example, each sequencing sample was in 15uL, containing 100-200ng plasmid DNA with 5-10 picomole sequencing primer, i.e., pET28aT7fd (SEQ ID NO:23), pET28aRBS-lfd (SEQ ID NO:24), or PyK3’end_rv (SEQ ID NO:25), as detailed in Table 2 above.2-3. Method for plasmid sequencing and sequencing analysis

[0129] Sequence analysis was performed using conventional methods as follows. First, all the DNA sequences were translated into six reading frames in amino acid sequence using Expasy translate online at expasy.org / translate / . For screening the tentative urzyme candidate, amino acid sequences from six reading frames from a single sequenced plasmid DNA sample were aligned with the sequence of urzyme template using conventional sequencing alignment software such as MOFFT online at ebi.ac.uk / jdispatcher / msa / mafft?stype=protein (Katch K, et al., 2019) or MUSCLE online at ebi.ac.uk / jdispatcher / msa / muscle?stype=protein (Edgar RC, 2004). The list of criteria for selection of tentative candidates included the sequence length, similarity, and identity, in addition to the highly conserved catalytic motifs such as including at amino acid positions 38, 74, 217-218, and 220-221 or the like.

[0130] The resulting new candidates for PK urzymes were designated urzyme and = protozyme as shown in Table 5 below, according to their sequence content and were collected and subjected for further downstream identification at the enzymology level after protein biochemistry procedures.Table 5. A summary of sequence annotated ancestral sequences of pyruvate kinase (PK).>Attorney Docket No. 5470.977.WO>> >Results

[0131] The full-length PK is shown in FIG. 8. As shown in FIGS. 9A and 9B, the truncated plasmid bears a strong similarity to what we observe with LeuRS (FIGS.3-5C). FIG.9B shows that the OCS includes extensive parts of the active site. The gap between residues 76-166 of SEQ ID NO: 68 includes the entire insertion domain , as well as part of the catalytic domain. The CTD is also missing. In all these respects, the OCS resembles what we observed in LeuAC. These parallels support the applicability of the present invention to multiple types of enzymes, as PK is a distantly related enzyme to LeuRS, is not part of the translation system, and has distinct sequences and domain structures as compared to LeuRS. Furthermore, because the AL marinarum PK gene has 64% GC content, which typically is associated with problems with PCR and sequencing (FIG. 13), the examples herein demonstrate that the methods of the invention can work across a broad spectrum of enzymes, even where the genes have high GC content.Conclusion

[0132] Many important protein families — enzymes, motor proteins, regulatory proteins — share a complex multidomain configuration with LeuRS. It has long been apparent that using intramolecular communication to induce functional cooperativity is an emergent property of such proteins (Monod, et al. 1963; Monod, et al. 1965; Changeux and Edelstein 2005; Sadovsky and Yifrach 2007; Zandany, et al. 2008; Ben-Abu, et al. 2009 ; Carter 2017). It appears to be an important source of the fitness that supported the (rapid) evolutionary gain of new genetic sequences.Attorney Docket No. 5470.977.WO

[0133] Although the details of coupling differ in many ways from protein family to protein family, the specific fragmentation we observed for LeuRS and PK suggest that disrupting internal coupling networks can perturb the fitness of cloned proteins, favoring loss of interacting domains when the active site is perturbed. Other protein families may use coupling mechanisms that can be made cytotoxic by active-site mutations. Our experiments show that one way to eliminate cytotoxic effects of cloned variant full-length proteins is to reduce those variants to their simplest functional forms.

[0134] Our models for ancestral AARS show that the most highly conserved modules retain substantial catalytic activity. They also reduce putative cytotoxicity. Experiments similar to those we describe here may produce similar sets of much shorter, but still functional variants of many other proteins. The possibility of a wider generality of our protocol could provide a rapid access to urzyme-like variants. Such variants could have evolutionary significance. Protein engineering (Kuhlman and Baker 2004; Romeroa, et al. 2018; Dauparas, et al. 2022; Dauparas, et al. 2023) can further enhance function (Patra, et al. 2025). Enhanced functionality could thus provide a generalized procedure to reduce the large size of genes for delivery by viral capsid systems (Hirsch, et al. 2010; Lewis 2014).

[0135] Thus, the LeuRS and PK OCSs signify a new approach to the discovery of ancestral genes, for AARS and also for other important protein families.Example 3: A study of the evolution of AARSs using urzymes created in vivoMethods

[0136] tRNALeuand iA>d-miniheHx,'-!lpreparation. A plasmid encoding the P. horikoshii tRNALeu(UAG anticodon) was synthesized by Integrated DNA Technologies and used as template for PCR amplification of the tRNA and upstream T7 promoter and downstream Hepatitis Delta Virus (HDV) ribozyme. The PCR product was used directly as template for T7 transcription. Following a 4-hour transcription at 37°C the RNA was cycled five times (90°C for 1 min, 60°C for 2 min, 25°C for 2 min) to increase the cleavage by HDV. The tRNA was purified by urea PAGE and crush and soak extraction. The tRNA 2’ -3’ cyclic phosphate was removed by treatment with T4 PNK (New England Biolabs) following the manufacturer’s protocol. The tRNA was then phenol chloroform isoamyl alcohol extracted, filter concentrated, aliquoted, and stored at -20°C.

[0137] The original sequence information for composing and designing Pyrococcus horikoshii TT'C-minihelix-Leu was based on the mature P. horikoshii tRNA-Leu, the tRNAAttorney Docket No. 5470.977.WOsequence was collected from the complement strand of the genomic sequence of P. horikoshii (OT3, GB# BA000001.2, between nucleotides 1448081 and 1448168). The designated minihelix sequence was obtained by combining the acceptor stem sequence including DCCA with the T C -stem-loop according to Schimmel (Schimmel and Alexander 1998). A plasmid harboring the minihelix, an upstream T7 promoter, and downstream Hepatitis Delta Virus (HDV) ribozyme was acquired from previous co-worker Jessica Elder. To reduce challenges posed by the high GC content in the stem portion of minihelix during the preparation, we implemented the Phi29 DNA polymerase-mediated isothermal amplification approach according to the manufacture’s protocol (New England Biolabs, Ipswich, MA). After the purification of the product from isothermal amplification, T7 transcription was carried out. Following a 4-6-hour transcription at 37 °C, the reaction mixture was subjected to urea PAGE fractionation, crush, and soak extraction. After 2’ -3’ cyclic phosphate removal by T4 PNK, the minihelix RNA was then phenol chloroform isoamyl alcohol extracted, filter-concentrated, quantitated, and aliquoted, and stored at -80 °C.

[0138] Expression and purification of deletion variants. We expressed the Goldilocks and LeuAC-like urzymes as MBP fusions from pMAL-c2x in BL21Star (DE3) (Invitrogen) as described in Tang, et al. 2024). The lysis buffer contained 20 mM Tris, pH 7.4, 1 mM EDTA, 5 mM p-mercaptoethanol ( ME), 17.5% Glycerol, 0.1% NP40, 33 mM (NH^SC , 1.25% Glycine, 300 mM Guanidine Hydrochloride) plus complete protease inhibitor (Roche). After elution from Amylose FF resin (Cytiva) with 10 mM maltose in lysis buffer (200 mM HEPES, pH 7.4, 450 mM NaCl, 100 mM KC1, 10 mM P-ME). Fractions containing protein were concentrated and mixed to 50% glycerol and stored at -20°C. All protein concentrations were determined using the Pierce™ Detergent-Compatible Bradford Assay Kit (Thermo Scientific). We determined purity by running samples on PROTEAN® TGX (Bio-RAD) gels and active fractions for all variants were measured as described in the next section.

[0139] Aminoacylation assays. We determined the active fraction of minihelix RNA by following extended acylation assays using the AVGA (SEQ ID NO:4) LeuAC mutant until they reached a plateau. That plateau value was used to compute minihelix concentrations in assays with all variants.

[0140] Aminoacylations withs minihelix were performed as described (Tang, et al. 2024) in 50 mM HEPES, pH 7.5, 10 mM MgCh, 20 mM KC1, 5 mM DTT with indicated amounts of ATP and isoleucine. The high affinity of The LeuAC-like and Goldilocks urzymes for minihelix meant that we had to mix increased amounts of [a32P] A76-labeled tRNA for assays. The RNAAttorney Docket No. 5470.977.WOsubstrates were heated in 30 mM HEPES, pH 7.5, 30 mM KC1 to 90°C for 2 minutes and then cooled linearly (drop l°C / 30 seconds) until it reached 80°C when MgCh was added to a final concentration of 10 mM. The tRNA continued to cool linearly until it reached 20°C. Michaelis-Menten experiments were performed by repeating these assays at the indicated RNA concentrations. Concentration-dependence fitted to Ksp= kcat / KM and kcat according to the modified formula (Eqn. 2) introduced by Johnson (Johnson 2019).

[0141] Data processing and Statistical analysis. Phosphor imaging screens of TLC plates were densitometered using ImageJ. Data were transferred to JMP16PRO™ Pro 16 via Microsoft Excel (version 16.49), after intermediate calculations. We fitted both active-site titration curves and Michaelis-Menten assays using the JMP16 PRO™ nonlinear fitting module.

[0142] Factorial design matrices were processed using the Fit Model multiple regression analysis module of JMP16PRO™ Pro, using an appropriate form of equation (3) (Box, et al.1978)AG{Yobs = Po + SPi*Pi + EPij*Pi*Pj + a (3)where Yobs is a dependent variable, usually an experimental observation, Po is a constant derived from the average value of Yobs, Pi and Pij are coefficients to be fitted, Pi,j are independent predictor variables from the design matrix, and a is a residual to be minimized. All rates and apparent affinities were converted to free energies of activation or binding, AG = - RTln(k), before regression analysis. Free energies are additive, whereas rates and apparent affinities are multiplicative. For example, the activation free energy for the first-order decay rate in singleturnover experiments is AGJ(kchem).

[0143] Multiple regression analyses of factorial designs exploit the replication inherent in the full collection of experiments to estimate experimental variances. In most cases we estimated errors on the basis of t-test P-values. In some cases, we added error bars to histograms to show the variance of individual datapoints. Multiple regression analyses reported here also entail triplet experimental replicates, which enhance the associated analysis of variance.

[0144] Amino acid sequence alignment and computation of codon-middle-base pairing. Sequences for 12-15 bacterial AARS for Class 1 LeuRS, IleRS, ValRS, MetRS; Class lb GlnRS, GluRS, GlxRS, and Class 1c TrpRS, and TyrRS were taken from the aars. online database. Segments corresponding to the protozyme (an ~50-residue segment at the N-terminus of Class 1 AARS containing the HIGH (SEQ ID NO:2) catalytic signature) and an ~25-residueAttorney Docket No. 5470.977.WOsegment containing the KMSKS (SEQ ID NO:3) signature were identified by inspection of 3D structures. Bidirectional genes were assembled from this curated dataset as described in the text.Results

[0145] The LeuAC-like and Goldilocks OCSs urzymes are 30-fold more active tRNA synthetases than is LeuAC. FIG. 4D shows that the two in vivo urzyme configurations omit longer and different segments from those present in LeuAC. The 3D consequences of the longer deletions are shown schematically in the structural cartoons in FIG. 19, together with their free energies of activation for kcat / KM. The active-site signatures, HVGH (SEQ ID NO:6) and KMSKS (SEQ ID NO:3), do not appear to contribute to the catalytic rate accelerations. The bar graph in FIG. 19 shows that the relative rate enhancements of both WT and mutant in vivo constructs are more proficient aminoacyl-tRNA synthetases than either WT or mutant LeuAC.

[0146] In turn, the significantly different linkages between the two parts of the active site (i.e. dashed arrows in FIG. 19 introduce quite a new element into our thinking about ancestral Class I AARS. LeuAC, designed in vitro, has 129 amino acid residues. The two in vivo LeuRS urzymes are both smaller. The LeuAC-like urzyme has residues 41-68 of SEQ ID NO: 71 + 601-674 of SEQ ID NO: 71 = 102 residues. The Goldilocks urzyme has 41-95 of SEQ ID NO: 71 + 650-674 of SEQ ID NO: 71 = 81 residues. Yet, the averages of the four bars representing the WT and double mutant in vivo constructs are ~ 2 kcal / mole more negative than the corresponding average for LeuAC (Tang, et al. 2024). Thus, they both catalyze minihelix acylation about 30 times better than the original LeuAC. This comparative catalytic advantage is achieved despite being significantly shorter than the original LeuAC.

[0147] Regression modeling identifies and attributes selectable functions. The in vivo fragmentation allows us to analyze the enzymology of LeuAC at unprecedented resolution. The second crossover connection (missing in the Goldilocks urzyme) and second P-strand of the protozyme (missing in the LeuAC-like urzyme), together with the WT vs double mutant active site residues form an incomplete factorial design (Carter and Carter 1979). The logarithmic transformation from rates to free energies of activation allow testing of linear models (Weinreb, et al. 2012; Li and Carter 2013; Tang, et al. 2023; Tang, et al. 2024). Regression models for the Michaelis-Menten constants in Table 6, below, show that the second half of the protozyme and the 2ndcrossover of the Rossmann fold have significant effects on the AG^kcat, the in situ first-order process that transfers the acyl group to the cognate minihelix. Such effects offer an experimental basis on which to assess why particular modules might be selected.Attorney Docket No. 5470.977.WOTable 6. Design matrix for regression analysis of active-site titration parameters. Independent parameters in the last three columns are: WT, the presence of full LeuAC; 2ndXvr, the presence of the second crossover of the Rossmann dinucleotide binding fold; Protoz, the presence of the full length of the protozyme.Attorney Docket No. 5470.977.WO

[0148] Both Goldilocks and LeuAC-like urzymes aminoacylate minihelices significantly better than WT LeuAC does (FIG. 19). The enhanced values of kcat / KM result both from increased kcat and higher affinity for minihelix RNA. However, active-site mutations have opposite effects on the two in vivo constructs. The double mutant (Dbl) LeuAC-like urzyme is the best catalyst of aminoacylation. It is nearly matched by the WT Goldilocks urzyme. The WT LeuAC-like urzyme is the least active variant. The decreased activity of WT variants that contain the second crossover connection of the Rossmann fold (i.e. LeuAC and LeuAC-like) has a P value <10'4. We noted a similar effect on the burst size (Supplementary Fig. S8 in (Tang and Carter 2025)).

[0149] The wild-type active-site variants bind more tightly to substrate minihelix in both Goldilocks and LeuAC-like urzymes, (FIG. 20C), but not in LeuAC. All four in vivo variants are better catalysts in all respects than the double mutant LeuAC. All of the effects influencing AG^kcat have highly significant P-values (FIGS. 20B and 20E). The WT active side residues change the effects of all three parameters. The second crossover connection reverses the effect of active-site mutations on the LeuAC-like urzyme. This is apparent from the placement of the various symbols on the graph in both FIG. 20D and FIG. 20E.

[0150] The filled-in shading in all three plots shows that the kinetic properties of WT LeuAC are all significantly worse than all of the constructs produced in vivo. Estimation of higher-order effects is limited by the poor reproducibility of KM. Notwithstanding, significant evidence remains for a negative interaction between the second crossover connection and the second crossover connection of the Rossmann fold (FIG.20D and FIG.20E). That interaction reduces both AG^kcat and AGd<cat / K\i by about 2 kcal / mole. The Goldilocks and LeuAC-like variants therefore enhance studies of the evolution of protein modularity at unrivalled resolution.

[0151] Class I and II AARS urzymes use similar secondary structures for aminoacylation. The Goldilocks urzyme is the smallest functional Class I AARS gene ever produced. The LeuAC, LeuAC-like and Goldilocks urzymes all share the AMSAS (SEQ ID NO:5) signature. The two new variant urzymes differ in both the protozyme and second crossover. The C-terminal fragment shared by LeuAC and both LeuAC-like and Goldilocks variants is key to binding RNA substrates. It is strong evidence that acquisition of that signature converted Class I protozymes into aminoacylating enzymes. Curiously, the secondary structure of the segment is a P-strand followed by a loop. The Motif 2 loop in Class 2 AARS also forms much of the RNA binding site. A striking resemblance of the RNA-binding secondary structures of Class I and II AARS is shown in FIG. 21. There are thus extraordinary similarities between the stereochemistry in both sets of polypeptide»RNA interactions leading to aminoacylation. WeAttorney Docket No. 5470.977.WOhave suggested elsewhere how the contrasting substrate-binding interactions of Class I and II AARS can arise from the inversion symmetry of base-pairing in bidirectional ancestral genes (Carter and Wills 2019a; Carter and Wills 2019b, 2021; Carter 2024; Carter, et al. 2025).

[0152] The modest size of the Class I AMSAS (SEQ ID NO:5)-containing module is reminiscent of the N-terminal fragment of the protozyme, which also is ~25 residues long and has a similar P->tum secondary structure and is key to the binding of ATP. In fact, the corresponding fragment in Class II AARS is actually the same fragment, residues 54-78 of SEQ ID NO: 74 in FIG. 21. As we noted previously (Tang, et al. 2023), neither the histidine side chains in the HVGH (SEQ ID NO:6) nor the lysine residues in the KMSKS (SEQ ID NO:3) signatures are necessary for amino acid activation or aminoacylation. The full functionality of the Lys-to-Ala mutations implies that the function of that segment arises solely from polypeptide secondary structure.

[0153] The Goldilocks urzyme suggests how one gene can encode both Class I and II AARS. The Rodin-Ohno hypothesis that bidirectional ancestral AARS genes coded for Class I and II urzymes on opposite strands (Rodin and Ohno 1995) led to a surprising number of retrodictions about the Class distinction (Carter, et al. 2025). The bidirectional protozyme gene was readily constructed (Martinez-Rodriguez, et al. 2015) because the corresponding segments from Class I and II AARS have the same number of amino acids. Others have verified its activation activity (Onodera, et al. 2021). We have failed, however, to design a bidirectional gene encoding Class I and II AARS urzymes on opposite strands because there are several loci where the combination of the two signatures has different numbers of amino acids and cannot be fitted into a consistent antiparallel alignment.

[0154] Because it lacks the second crossover connection of the Rossmann fold, The Goldilocks urzyme is the same length (81 residues) as the Class I GlyCA urzyme (Patra, et al.2024). Further, the 81-residue Goldilocks excerpt is a catalytically active AARS. Moreover, the Class-defining signature sequences of Class I and II AARS can both be readily aligned opposite one another in antiparallel alignments by the homology suggested by the GlyCA structure (Carter, et al. 2025). Thus, the Goldilocks urzyme may show how to explore designed bidirectional Class I / II AARS urzyme genes.

[0155] We compare in FIGS.22A-22B antiparallel alignments of Goldilocks-length urzymes drawn from 10-12 species of each synthetase from those in aars. online (Douglas, et al. 2024). The three Class I AARS subclassses match uniquely with Class II subclass urzyme genes. Superior matches entail both the total length and significantly elevated codon middle-base pairing. From the table in FIG.22B, Class IA IleRS pairs best with IIB LysRS; Class IIA ProRSAttorney Docket No. 5470.977.WOpairs best with Class IB GlnRS. Class IC TrpRS pairs best with Class IIC HisRS. We previously outlined extensive tests of the pairing frequencies under the null hypothesis of no bidirectional coding ancestry for such alignments (Chandrasekaran, et al. 2013). Standard errors for all mean values in FIG. 22B are derived from -140 bidirectional alignments. They are about 1 percent of the values themselves. The middle-base pairing frequencies are thus all highly non-random and statistically different one from another. Moreover, the different lengths of the three sets of related pairings make it hard to evaluate cross pairing in the matrix in FIG. 22B.

[0156] Pairings in FIG. 22B conflict with the conventional sub-classification by pairing subclass A in Class I with subclass B in Class II and vice versa. Each such pairing brings along two additional AARS from the same respective subclasses. Thus, 16 of the 20 bacterial AARS sequence alignments show analogous pairings between Class I and II AARS genes (Carter, et al. 2025).Conclusions

[0157] The longer-term impact for the study of the origin of genetics is hard to underestimate. These results, together with previous work, provide new and broadly-based support for a model for the emergence of the first Class I protein AARS. We envision assembly from three peptide fragments, each about 25 residues long. There is significant evidence for exceptionally high codon middle-base pairing between antiparallel alignments of each fragment. Bidirectional ancestral genes could thus have provided each fragment, establishing a basis for the joint origin of aminoacylating urzymes of both AARS classes (FIG. 23).

[0158] One fragment (sand background) provides an ATP binding site. The new result that both the half-protozyme and LeuAC-like ORFs retain this fragment unexpectedly supports its key role. To date we have only the structural evidence in the molecular cartoons that this fragment actually binds ATP. It does, however, contain the highly conserved ATP binding surfaces found in crystal structures from both Classes.

[0159] We proposed the half-protozyme as a progenitor of the Class I AARS protozyme via gene duplication and fusion to form an inverted repeat (Chandrasekaran, et al. 2013). Protozymes with both contemporary and bidirectionally coded sequences accelerate amino acid activation by 106-fold. The third fragment, indicated by the long dashed-line box and denoted (75 aa), endows the protozyme with the additional ability to recognize RNA substrates and catalyze acyl-transfer. It is described here for the first time (FIG. 21) showing the similarity between polypeptide :RNA interactions in the two AARS Classes. It appears to be both necessary and sufficient to enable the protozyme to catalyze aminoacylation (FIG. 23).Attorney Docket No. 5470.977.WOReferencesBen-Abu Y, Zhou Y, Zilberberg N, Yifrach O. 2009 Inverse coupling in leak and voltage-activated K+ channel gates underlies distinct roles in electrical signaling. Nature Structural & Molecular Biology 16 71-79.Box GEP, Hunter WG, Hunter JS. 1978. Statistics for Experimenters. New York: Wiley Interscience. Burbaum J, Schimmel P. (Minireview co-authors). 1991. Structural Relationships and the Classification of Aminoacyl-tRNA Synthetases. The Journal of Biological Chemistry 266:16965-16968.Burbaum JJ, Starzyk RM, Schimmel P. (Review Article co-authors). 1990. Understanding Structural Relationships in Proteins of Unsolved Three-Dimensional Structure. PROTEINS: Structure, Function, and Genetics 7:99-111.Carter CW, Jr, Tang GQ, Patra SK, Betts L, Dieckhaus H, Kuhlman B, Douglas J, Wills PR, Bouckaert R, Popovic M, Ditzler M. 2025. Structural Enzymology, Phylogenetics, Differentiation, and Symbolic Reflexivity at the Dawn of Biology. Genome Biology and Evolution 17: evaf095.Carter CW, Jr., Li L, Weinreb V, Collier M, Gonzales-Rivera K, Jimenez-Rodriguez M, Erdogan O, Chandrasekharan SN. 2014. The Rodin-Ohno Hypothesis That Two Enzyme Superfamilies Descended from One Ancestral Gene: An Unlikely Scenario for the Origins of Translation That Will Not Be Dismissed. Biology Direct 9:11.Carter CW, Jr. 2014. Urzymology: Experimental Access to a Key Transition in the Appearance of Enzymes. J. Biol. Chem. 289:30213-30220.Carter CW, Jr. 2017. High-Dimensional Mutant and Modular Thermodynamic Cycles, Molecular Switching, and Free Energy Transduction. Annual Review of Biophysics 46:433-453.Carter CW, Jr., Carter CW. 1979. Protein Crystallization Using Incomplete Factorial Experiments. Journal of Biological Chemistry 254:12219-12223.Carter CW, Jr, Wills PR. 2019a. Class I and II aminoacyl -tRNA synthetase tRNA groove discrimination created the first synthetase*tRN A cognate pairs and was therefore essential to the origin of genetic coding. IUBMB Life 71: 1088-1098.Carter CW, Jr., Wills PR. 2019b. Experimental Solutions to Problems Defining the Origin of Codon-Directed Protein Synthesis. Biosystems 183:103979.Carter CW, Jr., Wills PR. 2021. The Roots of Genetic Coding in Aminoacyl-tRNA Synthetase Duality Annual Review of Biochemistry 90:349-373.Carter CW, Jr. 2024. Base Pairing Promoted the Self-Organization of Genetic Coding, Catalysis, and Free-Energy Transduction. MDPI Life 14:199.Chandrasekaran SN, Yardimci G, Erdogan O, Roach JM, Carter CW, Jr. 2013. Statistical Evaluation of the Rodin-Ohno Hypothesis: Sense / Antisense Coding of Ancestral Class I and II Aminoacyl-tRNA Synthetases. Molecular Biology and Evolution 30: 1588-1604.Changeux J-P, Edelstein SJ. 2005. Allosteric Mechanisms of Signal Transduction. Science 308: 1424-1428.Dauparas J, Anishchenko I, Bennett N, Bai H, Ragotte RJ, Milles LF, Wicky BIM, Courbet A, de Haas RJ, Bethel N, et al. 2022. Robust deep learning-based protein sequence design using ProteinMPNN. Science 378, :49-56 (2022)Dauparas J, Lee GR, Pecoraro R, An L, Anishchenko I, Glasscock C, Baker D. 2023. Atomic context-conditioned protein sequence design using LigandMPNN. BioRxiv.Douglas J, Cui H, Perona JJ, Vargas-Rodriguez O, Tyynismaa H, Carreno CA, Ling J, Ribas-de-Pouplana L, Yang X-L, Ibba M, et al. 2024. AARS Online: a collaborative database on the structure, function, and evolution of the aminoacyl -tRNA synthetases. Life 4: 1091-1105.Attorney Docket No. 5470.977.WODouglas J, Bouckaert R, Carter CWJ, Wills PR. 2024. Enzymic recognition of amino acids drove the evolution of primordial genetic codes. Nucleic Acids Research 52:558-571.Edgar RC. (2004) MUSCLE: multiple sequence alignment with high accuracy and high throughput. Nucleic Acids Res. 32(5): 1792-7. doi: 10.1093 / nar / gkh340. Print 2004.Fersht AR, Ashford, J.S., Bruton, C.J., Jakes, R., Koch, G.L.E. and Hartley, B.S. 1975. Active Site Titration and Amino Acyladenylate Binding Stoichiometry of Amino Acyl Synthetases. Biochemistry 14:1-4.Francklyn CS, First EA, Perona JJ, Hou Y-M. 2008. Methods for kinetic and thermodynamic analysis of aminoacyl -tRNA synthetases. Methods 44:100-118.Fritsch EF, Sambrook J, Maniatis T. 1982. Molecular Cloning: A Laboratory Manual. New York.: Cold Spring Harbor Laboratory.Hirsch ML, Agbandje-McKenna M, Samulski RJ. 2010. Little Vector, Big Gene Transduction:Fragmented Genome Reassembly of Adeno-associated Virus. Molecular Therapy 18 6-8.Hobson JJ, Li Z, Carter CW, Jr. 2022. A leucyl -tRNA synthetase urzyme: authenticity of tRNA Synthetase urzyme catalytic activities and production of a non-canonical product. International Journal of Molecular Sciences 23:4229.Johnson KA. 2019. New standards for collecting and fitting steady state kinetic data. Beilstein J. Org. Chem. 15:16-29.Katoh K, Rozewicki J, Yamada KD. 2019. MAFFT online service: multiple sequence alignment, interactive sequence choice and visualization. Brief. Bioinform. 20:1160-1166.Kuhlman B, Baker D. 2004. Exploring folding free energy landscapes using computational protein design. Current Opinion in Structural Biology 14:89-95.Lewis R. 2014. Gene Therapy's Second Act. Scientific American 310:52-57.Li L, Carter CW, Jr. 2013. Full Implementation of the Genetic Code by Tryptophanyl-tRNA Synthetase Requires Intermodular Coupling. J. Biol. Chem. 288:34736-34745.Martinez-Rodriguez L, Jimenez-Rodriguez M, Gonzalez-Rivera K, Williams T, Li L, Weinreb V, Chandrasekaran SN, Collier M, Ambroggio X, Kuhlman B, et al. 2015. Functional Class I and II Amino Acid Activating Enzymes Can Be Coded by Opposite Strands of the Same Gene. J. Biol.Chem. 290:19710-19725.Michel B, Sandler SJ. 2017. Replication Restart in Bacteria. Journal of Bacteriology 199:e00102-00117Monod J, Changeux J-P, Jacob F. 1963. Allosteric Proteins and Cellular Control Systems. Journal of Molecular Biology 6: 306-329.Monod J, Wyman J, Changeux J-P. 1965. On the nature of allosteric transitions: A plausible model. Journal of Molecular Biology 12:88-118.Onodera K, SuganumaN, Takano H, Sugita Y, Shoji T, Minobe A, Yamaki N, Otsuka R, Mutsuro-Aoki H, Umehara T, Tamura K. 2021. Amino acid activation analysis of primitive aminoacyl -tRNA synthetases encoded by both strands of a single gene using the malachite green assay. Biosystems 208:104481.Patra SK, Betts L, Tang GQ, Douglas J, Wills PR, Bouckaert R, Carter CW, Jr. . 2024. A genomic database furnishes minimal functional glycyl-tRNA synthetases homologous to other, designed class II urzymes. Nucleic Acids Research 52: 13305-13324.Patra SK, Randolph N, Dieckhaus H, Kuhlman B, Douglas J, Carter CW, Jr. 2025. Aminoacyl -tRNA synthetase urzymes optimized by deep learning behave as a quasispecies. Structural Dynamics 12:024701.Attorney Docket No. 5470.977.WOPham Y, Li L, Kim A, Erdogan 0, Weinreb V, Butterfoss G, Kuhlman B, Carter CW, Jr. 2007. A Minimal TrpRS Catalytic Domain Supports Sense / Antisense Ancestry of Class I and II Aminoacyl-tRNA Synthetases. Mol. Cell 25:851-862.Pham Y, Kuhlman B, Butterfoss GL, Hu H, Weinreb V, Carter CW, Jr. 2010. Tryptophanyl-tRNA synthetase Urzyme: a model to recapitulate molecular evolution and investigate intramolecular complementation. J. Biol. Chem. 285:38590-38601.Picard V, Ersdal-Badju E, Lu A, Bock SC. 1994. A rapid and efficient one-tube PCR-based mutagenesis technique using Pfu DNA polymerase. Nucleic Acids Research 22:2587-2591.Rodin SN, Ohno S. 1995. Two Types of Aminoacyl -tRNA Synthetases Could be Originally Encoded by Complementary Strands of the Same Nucleic Acid. Origins of Life and Evolution of the Biosphere 25:565-589.Romeroa MLR, Yang F, Lind Y-R, Toth-Petroczy A, Berezovsky IN, Goncearenco A, Yang W, Wellner A, Kumar-Deshmukh F, Sharon M, et al. 2018. Simple yet functional phosphate-loop proteins. PNAS 115:E11943— El 1950.Sadovsky E, Yifrach O. 2007. Principles underlying energetic coupling along an allosteric communication trajectory of a voltage-activated K channel. Proceedings of the National Academy of Sciences, USA 104:19813-19818.Sharma, N., Anleu Gil, M. X. and Wengier, D. (2017). A Quick and Easy Method for Making Competent Escherichia coli Cells for Transformation Using Rubidium Chloride. Bio-101: e2590. DOI: 10.21769 / BioProtoc.2590.Schimmel P, Alexander R. 1998. Diverse RNA substrates for aminoacylation: Clues to origins? Proc. Nat. Acad. Sci. USA 95: 10351-10353.Tang GQ, Elder JJH, Douglas J, Carter CW, Jr, . 2023. Domain Acquisition by Class I Aminoacyl -tRNA Synthetase Urzymes Coordinated the Catalytic Functions of HVGH and KMSKS Motifs.Nucleic Acids Research 51:8070-8084.Tang GQ, Hu H, Douglas J, Carter CW, Jr. 2024. Primordial aminoacyl-tRNA synthetases preferred tRNA minihelix substrates over full-length tRNA. Nucleic Acids Research 52:7096-7111.Tang GQ, Carter CW, Jr. 2025. Escherichia coli deletes in vivo the same domains from a doublemutant leucyl -tRNA synthetase gene that were deleted in vitro to make the LeuAC urzyme. BioRxiv doi: https: / / doi.org / 10.1101 / 2025.04.05.647346.Weinreb V, Li L, Kaguni LS, Campbell CL, Carter CW, Jr. 2009. Mg2+-Assisted Catalysis by B. stearothermophilus TrpRS is Promoted by Allosteric Effects. Structure 17:952-964.Weinreb V, Carter CW, Jr. 2008. Mg2+-free B. stearothermophilus Tryptophanyl-tRNA Synthetase Activates Tryptophan With a Major Fraction of the Overall Rate Enhancement. Journal of the American Chemical Society 130:1488-1494.Weinreb V, Li L, Carter CW, Jr. 2012. A Master Switch Couples Mg2+-Assisted Catalysis to Domain Motion in B. stearothermophilus Tryptophanyl-tRNA Synthetase. Structure 20: 128-138.Williams T, Yin WY, Carter Jr. CW. 2015. Selective Inhibition of Bacterial Tryptophanyl-tRNA Synthetases by Indolmycin is Mechanism Based. Journal Biologial Chemistry 291:255-265.W. Zwerschke's PhD thesis, DKFZ, Heidelberg, Germany, 1997: Competent cells: Rubidium Chloride. https: / / cbdm.hms.harvard.edu / assets / Protocols / DNA%20prep / Competent%20cells.pdfZandany N, Ovadia M, Orr I, Yifrach O. 2008 Direct analysis of cooperativity in multisubunit allosteric proteins. Proc Nat. Acad. Sci, USA 105 11697-11702.Attorney Docket No. 5470.977.WO

[0160] The foregoing examples are illustrative of the present invention and are not to be construed as limiting thereof. Although the invention has been described in detail with reference to preferred embodiments, variations and modifications exist within the scope and spirit of the invention as described and defined in the following claims.

Claims

1. Attorney Docket No. 5470.977.WOTHAT WHICH IS CLAIMED IS:

1. A method of generating a polynucleotide encoding a miniaturized enzyme comprising:i) preparing a vector comprising a polynucleotide encoding an enzyme; wherein the enzyme comprises an active site composed of amino acid residues that are divided into two or more separate sections on the primary amino acid sequence of the enzyme;ii) hybridizing to the vector at least one pair of oligonucleotides that specifically hybridize to nucleotides encoding the amino acid residues in separate sections of the active site, wherein at least one of the oligonucleotides is designed to introduce an amino acid mutation into the enzyme;iii) performing a nucleic acid amplification reaction on the vector to produce an amplification product;iv) introducing the amplification product into a prokaryotic, archaean, or eukaryotic cell suitable for recombination; andv) isolating vectors that have undergone recombination to encode a miniaturized enzyme.

2. A method of generating a polynucleotide encoding a miniaturized enzyme comprising:i) preparing a vector comprising a polynucleotide encoding an enzyme; wherein the enzyme comprises an active site composed of amino acid residues that are divided into two or more separate sections on the primary amino acid sequence of the enzyme;ii) hybridizing to the vector at least one pair of oligonucleotides that specifically hybridize to nucleotides encoding the amino acid residues in separate sections of the active site, wherein at least one of the oligonucleotides is designed to introduce an amino acid mutation into the enzyme;iii) performing a first nucleic acid amplification reaction on the vector to produce a first amplification product;iv) combining the first amplification product with the vector of step i);v) performing a second nucleic acid amplification reaction on the vector to produce a second amplification product;vi) introducing the second amplification product into a prokaryotic, archaean, or eukaryotic cell suitable for recombination; andAttorney Docket No. 5470.977.WOvii) isolating vectors that have undergone recombination to encode a miniaturized enzyme.

3. The method of claim 1 or 2, comprising hybridizing to the vector at least two pairs of oligonucleotides that specifically hybridize to nucleotides encoding the amino acid residues in separate sections of the active site.

4. The method of any one of claims 1-3, comprising hybridizing to the vector at least three pairs of oligonucleotides that specifically hybridize to nucleotides encoding the amino acid residues in separate sections of the active site.

5. A method of generating a polynucleotide encoding a miniaturized enzyme comprising:i) preparing a vector comprising a polynucleotide encoding an enzyme; wherein the enzyme comprises an active site composed of amino acid residues that are divided into two or more separate sections on the primary amino acid sequence of the enzyme;ii) hybridizing to the vector a first set of oligonucleotides, wherein the first set of oligonucleotides comprises at least two pairs of oligonucleotides that specifically hybridize to nucleotides encoding the amino acid residues in separate sections of the active site, wherein at least one of the at least two pairs of oligonucleotides is designed to introduce an amino acid mutation into the enzyme;iii) performing a first nucleic acid amplification reaction on the vector to produce a first set of amplification products, wherein the first set of amplification products comprises at least two amplification products, wherein the at least two amplification products of the first set of amplification products overlap at a 5’ end or 3’ end;iv) combining the first set of amplification products with a second set of oligonucleotides (e.g., a subset of the oligonucleotides from the first set of oligonucleotides, or different oligonucleotides);v) performing a second nucleic acid amplification reaction on the combination of step iv) to produce a second set of amplification products, wherein the second set of amplification products comprises at least one amplification product (e.g., an amplification product that differs from the amplification products in step iii));vi) combining the second set of amplification products with the vector of step i); vii) performing a third nucleic acid amplification reaction on the combination of step vi);Attorney Docket No. 5470.977.WOviii) introducing the amplification product of step vii) into a prokaryotic, archaean, or eukaryotic cell suitable for recombination; andix) isolating vectors that have undergone recombination to encode a miniaturized enzyme.

6. The method of any one of claims 1-5, wherein the oligonucleotides are primers.

7. The method of any one of claims 1-6, wherein the oligonucleotides are about 20-200 nucleotides long.

8. The method of any one of claims 1-7, wherein the amplification reaction is a polymerase chain reaction (PCR).

9. The method of claim 5, wherein the second nucleic acid amplification reaction is an overlap extension polymerase chain reaction (OE-PCR) (e.g., overlapping PCR ligation).

10. The method of any one of claims 1-9, wherein all the oligonucleotides are designed to introduce an amino acid mutation into the enzyme.

11. The method of any one of claims 1-10, wherein the recombination comprises deletion of nucleotides.

12. The method of claim 11, wherein the deletion of nucleotides comprises deletion of a single nucleotide.

13. The method of claim 11, wherein the deletion of nucleotides comprises deletion of a short nucleotide sequence (e.g., a nucleotide sequence comprising 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, or 20 to 50, 60, 70, 80, 90, or 100 nucleotides).

14. The method of claim 11, wherein the deletion of nucleotide comprises deletion of a long nucleotide sequence (e.g., a nucleotide sequence comprising 101, 200, 300, 400, or 500, to 600, 700, 800, 900, 1000, 1250, 1500, 1750, 2000, 3000, 4000, 5000, 7500, or 10000 nucleotides).Attorney Docket No. 5470.977.WO15. The method of any one of claims 1-14, wherein the enzyme is an aminoacyl-tRNA synthetase (AARS).

16. The method of claim 15, wherein the AARS is leucyl-tRNA synthetase (LeuRS).

17. The method of any one of claims 1-14, wherein the at least two oligonucleotides encode amino acid residues AVGA (SEQ ID NO:4) and AMSAS (SEQ ID NO:5), optionally wherein the at least two oligonucleotides encode a mutation within amino acid residues AVGA (SEQ ID NO:4) and AMSAS (SEQ ID NO:5)18. The method of any one of claims 1-14, wherein the enzyme is pyruvate kinase.

19. The method of claim 18, wherein the at least two oligonucleotides encode amino acid residues 38H, 74R, 217A, 218K, 220E, and 22 IK, optionally wherein the at least two oligonucleotides encode mutations within amino acid residues 38H, 74R, 217A, 218K, 220E, and 221K (e.g., 38L, 74A, 217Q, 218T, 220A, and / or 221A).

20. The method of any one of claims 1-19, wherein the vector is a plasmid.

21. The method of any one of claims 1-20, wherein the prokaryotic, archaean, or eukaryotic cell is a bacterium.

22. The method of claim 21, wherein the bacterium is E. coli.

23. The method of any one of claims 1-20, wherein the prokaryotic, archaean, or eukaryotic cell is a yeast.

24. The method of any one of claims 1-20, wherein the prokaryotic, archaean, or eukaryotic cell is a mammalian cell.

25. A composition comprising a vector comprising a polynucleotide encoding an enzyme; wherein the enzyme comprises an active site composed of amino acid residues that are divided into two or more separate sections on the primary amino acid sequence of the enzyme; and at least one pair of oligonucleotides hybridized to the vector, wherein the oligonucleotidesAttorney Docket No. 5470.977.WOspecifically hybridize to nucleotides encoding amino acid residues in separate sections of the active site, and wherein at least one of the oligonucleotides is designed to introduce an amino acid mutation into the enzyme.

26. The composition of claim 25, wherein at least two pairs of oligonucleotides that specifically hybridize to nucleotides encoding the amino acid residues in separate sections of the active site are hybridized to the vector.

27. The composition of claim 25 or 26, wherein at least three pairs of oligonucleotides that specifically hybridize to nucleotides encoding the amino acid residues in separate sections of the active site are hybridized to the vector.

28. The composition of any one of claims 25-27, wherein the oligonucleotides are primers.

29. The composition of any one of claims 25-28, wherein all the oligonucleotides are designed to introduce an amino acid mutation into the enzyme.

30. The composition of any one of claims 25-29, wherein the enzyme is an aminoacyl-tRNA synthetase (AARS).

31. The composition of claim 30, wherein the AARS is leucyl-tRNA synthetase (LeuRS).

32. The composition of claim 29, wherein the at least two oligonucleotides encode amino acid residues AVGA (SEQ ID NO:4) and AMSAS (SEQ ID NO:5), optionally wherein the at least two oligonucleotides encode a mutation within amino acid residues AVGA (SEQ ID NO:4) and AMSAS (SEQ ID NO:5).

33. The composition of any one of claims 25-29, wherein the enzyme is pyruvate kinase.

34. The composition of claim 33, wherein the at least two oligonucleotides encode amino acid residues 38H, 74R, 217A, 218K, 220E, and 22 IK, optionally wherein the at least two oligonucleotides encode mutations within amino acid residues 38H, 74R, 217A, 218K, 220E, and 221K (e.g., 38L, 74A, 217Q, 218T, 220A, and / or 221A).Attorney Docket No. 5470.977.WO35. The composition of any one of claims 25-34, wherein the vector is a plasmid.

36. An amplification product produced from performing an amplification reaction on the composition of any one of claims 25-35.

37. The amplification product of claim 36, wherein the amplification reaction is a polymerase chain reaction (PCR).

38. A prokaryotic, archaean, or eukaryotic cell comprising the amplification product of claim 36.

39. The prokaryotic, archaean, or eukaryotic cell of claim 38, wherein the prokaryotic or eukaryotic cell is a bacterium.

40. The prokaryotic, archaean, or eukaryotic cell of claim 39, wherein the bacterium is E. coli.

41. The prokaryotic, archaean, or eukaryotic cell of claim 38, wherein the prokaryotic or eukaryotic cell is a yeast.

42. The prokaryotic, archaean, or eukaryotic cell of claim 38, wherein the prokaryotic or eukaryotic cell is a mammalian cell.

43. A vector encoding a miniaturized enzyme, produced by the method of any one of claims 1-24.