Genomic Insertion in Cells
The two-RNA delivery system with nrRT and modified uridines addresses integration challenges in existing gene therapies, achieving efficient and safe site-specific insertion of heterologous polynucleotides into eukaryotic genomes, including post-mitotic cells.
Patent Information
- Application Number
- JP2025502992
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-20
- Filing Date
- 2023-07-19
- Publication Date
- 2025-07-25
AI Technical Summary
Current methods for inserting DNA into host cell genomes face issues such as immune responses, mutagenicity, and non-specific integration, especially in post-mitotic cells like neurons, and existing gene therapy techniques have limitations.
A two-RNA delivery system using a non-LTR retrotransposon reverse transcriptase protein (nrRT) and a template RNA with modified uridines to facilitate site-specific integration of heterologous polynucleotides into eukaryotic genomes, avoiding DNA delivery-related issues and enabling targeted insertion in various cell types.
Enhances insertion efficiency and fidelity, reduces cytotoxicity, and allows for targeted gene therapy in post-mitotic cells, providing a safer and more effective method for introducing therapeutic proteins or regulatory RNAs into target cells.
Smart Images

Figure 2025523992000001_ABST
Abstract
Description
Technical Field
[0001] Cross - reference to related applications This application claims the benefit of priority of U.S. Provisional Patent Application No. 63 / 390,863, filed Jul. 20, 2022, which is incorporated by reference in its entirety for all purposes.
Background Art
[0002] Background of the Invention Insertion of DNA - introduced genes into the genomic DNA of organisms is associated with several undesirable side effects. For example, introduction of DNA into the cytoplasm of cells can induce an immune response that can be harmful to the cells or organisms. Further, current methods for integrating DNA into a target site in the host cell genome via homologous recombination require introducing potentially mutagenic double - strand breaks into the genomic DNA. Additionally, DNA integration in post - mitotic cells such as neurons can occur at non - specific locations due to the fact that homologous recombination occurs more efficiently in dividing cells.
[0003] The present disclosure provides compositions and methods for improving gene editing at target sites in the host cell genome. The methods can be used for gene therapy applications and can provide advantages over current DNA - based and viral vector - based gene therapy methods.
Summary of the Invention
Means for Solving the Problems
[0004] Brief Summary of the Invention The present disclosure provides compositions and methods for improving gene therapy techniques for introducing heterologous polynucleotides into target cells.
[0005] In one aspect, the present disclosure provides a method of inserting a heterologous polynucleotide at a target site in a eukaryotic genome, the method comprising transfecting a eukaryotic cell with (a) an RNA encoding a non-LTR retrotransposon reverse transcriptase protein (nrRT) comprising a reverse transcriptase domain and an endonuclease domain; and (b) a template RNA. In some embodiments, the template RNA comprises a promoter, a payload sequence, a polyA sequence, and an nrRT binding sequence. In some embodiments, the template RNA comprises one or more modified uridine (U) nucleosides selected from the group consisting of N1-methyl-pseudouridine (N1mΨU), pseudouridine (ΨU), 5-methyluridine (5meU), 5-methoxyuridine (5moU), and mixtures thereof. In some embodiments, the template RNA comprises unmodified uridine and a mixture comprising one or more modified uridines selected from the group consisting of N1mΨU, ΨU, 5meU, and 5moU. In some embodiments, the template RNA comprises a modified uridine that is not cleavable by a ribozyme.
[0006] In some embodiments, the nrRT is expressed in the cell and catalyzes the insertion of a double-stranded heterologous polynucleotide comprising the payload sequence at the target site in the eukaryotic genome.
[0007] In some embodiments, the template RNA comprising modified U increases the efficiency of insertion of the payload sequence into the eukaryotic genome as compared to a template RNA comprising unmodified U.
[0008] In some embodiments, the template RNA further comprises a 5' ribozyme sequence selected from an active ribozyme, a partially active ribozyme, a ribozyme having reduced catalytic activity, or a catalytically inactive ribozyme. In some embodiments, the 5' ribozyme is selected from an HDV ribozyme, a TriCasA ribozyme, or a natural cognate ribozyme, semi-cognate ribozyme, or variant thereof. In some embodiments, the 5' ribozyme sequence is a sequence selected from any one of SEQ ID NO: 3 or SEQ ID NOs: 13-22 (without the pp7 binding sequence), or a sequence having a sequence identity greater than or equal to 60% (e.g., greater than or equal to 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity) with a sequence selected from any one of SEQ ID NO: 3 or SEQ ID NOs: 13-22 (without the pp7 binding sequence).
[0009] In some embodiments, the template RNA does not contain a functional 5' ribozyme sequence or does not contain a 5' ribozyme sequence.
[0010] In some embodiments, the cytotoxicity decreases when the template RNA contains a modified U.
[0011] In some embodiments, the template RNA further comprises a 5' sequence that protects the 5' end from degradation.
[0012] In some embodiments, the template RNA further comprises a 5' sequence that promotes site-specific insertion of the heterologous polynucleotide into the target site in the eukaryotic genome.
[0013] In some embodiments, the nrRT binding sequence includes a 3’UTR sequence. In some embodiments, the 3’UTR sequence is isolated from an organism selected from the group consisting of G. aculeatus, D. melanogaster, L. polyphemus, P. pungitis, N. vitripennis, G. fortis, O. latipes, Z. albicollis, T. guttata, T. castaneum, T. guttatus, D. simulans, B. mori, and A. vaga. In some embodiments, the 3’UTR includes a sequence having a sequence identity greater than or equal to 60% (e.g., greater than or equal to 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%) with any one of SEQ ID NOs: 26-39.
[0014] In some embodiments, the template RNA further includes a 3’ sequence that promotes site-specific insertion of the heterologous polynucleotide into the eukaryotic genome and / or enhances the efficiency and fidelity of target-primed reverse transcription.
[0015] In some embodiments, the template RNA further includes one or more of: i) an RNA polymerase terminator; ii) a sequence useful for purification; iii) a sequence encoding a protein useful for enrichment; iv) a Kozak sequence located 5’ to the payload sequence; and / or v) a polyA sequence located 3’ to the nrRT binding sequence.
[0016] In some embodiments, the template RNA further includes: (a) a 5’ sequence homologous to a DNA sequence located 5’ to the target insertion site in the eukaryotic genome; or (b) a 3’ sequence homologous to a DNA sequence located 3’ to the target insertion site in the eukaryotic genome; or both (a) and (b).
[0017] In some embodiments, the template RNA lacks a 5' phosphate.
[0018] In some embodiments, the payload sequence encodes a therapeutic protein that replaces or complements a defective gene or protein. In some embodiments, the therapeutic protein is selected from the group consisting of Factor VIII, Factor IX, and phenylalanine hydroxylase (PAH).
[0019] In some embodiments, the payload sequence encodes an inhibitor of another protein. In some embodiments, the inhibitor is a single-chain antibody.
[0020] In some embodiments, the payload sequence encodes a regulatory RNA.
[0021] In some embodiments, the payload sequence encodes a protein selected from the genes in Table 7.
[0022] In some embodiments, adjusting i) the molar ratio of the nrRT mRNA to the template RNA and / or ii) the amount of total RNA delivered to the target cells increases the insertion efficiency.
[0023] In some embodiments, the RNA encoding the nrRT comprises one or more modified uridine (U) nucleosides selected from the group consisting of N1-methyl-pseudouridine (N1mΨU), pseudouridine (ΨU), 5-methyluridine (5meU), 5-methoxyuridine (5moU), and mixtures thereof. In some embodiments, the RNA encoding the nrRT comprises unmodified uridine, as well as a mixture of modified Us selected from the group consisting of N1mΨU, ΨU, 5meU, and 5moU.
[0024] In some embodiments, the eukaryotic cell is transfected in vitro. In some embodiments, the eukaryotic cell is transfected in vivo. In some embodiments, the eukaryotic cell is a mammalian cell. In some embodiments, the eukaryotic cell is a human cell. In some embodiments, the human cell is removed from a human subject and transfected (e.g., ex vivo) with the RNAs of (a) and (b) above to insert the heterologous polynucleotide into the human cell genome and administered to the human subject.
[0025] In some embodiments, the cell is transfected with an LNP formulation, a lipofection reagent, or by electroporation.
[0026] In another aspect, the disclosure provides a composition comprising (a) an RNA encoding a non-LTR retrotransposon reverse transcriptase protein (nrRT) comprising a reverse transcriptase domain and an endonuclease domain; and (b) a template RNA. In some embodiments, the template RNA comprises a promoter, a payload sequence, a polyA sequence, and an nrRT binding sequence. In some embodiments, the template RNA comprises one or more modified uridine (U) nucleosides selected from the group consisting of N1-methyl-pseudouridine (N1mΨU), pseudouridine (ΨU), 5-methyluridine (5meU), 5-methoxyuridine (5moU), and mixtures thereof. In some embodiments, the template RNA comprises unmodified uridine and a mixture comprising one or more modified uridines selected from the group consisting of N1mΨU, ΨU, 5meU, and 5moU. In some embodiments, the template RNA comprising modified uridine is uncleavable by ribozymes.
[0027] In some embodiments, the template RNA further comprises a 5' ribozyme sequence selected from an active ribozyme, a partially active ribozyme, a ribozyme having reduced catalytic activity, or a catalytically inactive ribozyme. In some embodiments, the 5' ribozyme is selected from an HDV ribozyme, a TriCasA ribozyme, or a natural cognate ribozyme, semi-cognate ribozyme, or variants thereof. In some embodiments, the 5' ribozyme sequence is a sequence selected from any one of SEQ ID NO: 3 or SEQ ID NOs: 13 to 22 (without the pp7 binding sequence), or a sequence having a sequence identity greater than or equal to 60% (e.g., greater than or equal to 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity) with a sequence selected from any one of SEQ ID NO: 3 or SEQ ID NOs: 13 to 22 (without the pp7 binding sequence).
[0028] In some embodiments, the template RNA does not contain a functional 5' ribozyme sequence or does not contain a 5' ribozyme sequence.
[0029] In some embodiments, the template RNA further comprises a 5' sequence that protects the 5' end from degradation.
[0030] In some embodiments, the template RNA further comprises a 5' sequence that promotes site-specific insertion of the heterologous polynucleotide into a target site in the eukaryotic genome.
[0031] In some embodiments, the nrRT binding sequence includes a 3’UTR sequence. In some embodiments, the 3’UTR sequence is isolated from an organism selected from the group consisting of G. aculeatus, D. melanogaster, L. polyphemus, P. pungitis, N. vitripennis, G. fortis, O. latipes, Z. albicollis, T. guttata, T. castaneum, T. guttatus, D. simulans, B. mori, and A. vaga. In some embodiments, the 3’UTR includes a sequence having a sequence identity greater than or equal to 60% (e.g., greater than or equal to 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity) with any one of SEQ ID NOs: 26-39.
[0032] In some embodiments, the template RNA further includes a 3’ sequence that promotes site-specific insertion of the heterologous polynucleotide into the eukaryotic genome and / or enhances the efficiency and fidelity of target-primed reverse transcription.
[0033] In some embodiments, the template RNA further includes one or more of: i) an RNA polymerase terminator; ii) a sequence useful for purification; iii) a sequence encoding a protein useful for enrichment; iv) a Kozak sequence located 5’ to the payload sequence; and / or v) a polyA sequence located 3’ to the nrRT binding sequence.
[0034] In some embodiments, the template RNA further includes: a) a 5’ sequence that is homologous to a DNA sequence located 5’ to the target insertion site in the eukaryotic genome; or (b) a 3’ sequence that is homologous to a DNA sequence located 3’ to the target insertion site in the eukaryotic genome; or both (a) and (b).
[0035] In some embodiments, the template RNA lacks a 5' phosphate.
[0036] In some embodiments, the payload sequence encodes a therapeutic protein that replaces or complements a defective gene or protein. In some embodiments, the therapeutic protein is selected from the group consisting of Factor VIII, Factor IX, and phenylalanine hydroxylase (PAH).
[0037] In some embodiments, the payload sequence encodes an inhibitor of another protein. In some embodiments, the inhibitor is a single-chain antibody.
[0038] In some embodiments, the payload sequence encodes a regulatory RNA.
[0039] In some embodiments, the payload sequence encodes a protein selected from the genes in Table 7.
[0040] In some embodiments, the RNA encoding the nrRT comprises one or more modified uridine (U) nucleosides selected from the group consisting of N1-methyl-pseudouridine (N1mΨU), pseudouridine (ΨU), 5-methyluridine (5meU), 5-methoxyuridine (5moU), and mixtures thereof. In some embodiments, the RNA encoding the nrRT comprises unmodified uridine, as well as a mixture of modified Us selected from the group consisting of N1mΨU, ΨU, 5meU, and 5moU.
[0041] In another aspect, the present disclosure provides a pharmaceutical composition. The pharmaceutical composition may include the compositions described herein. In some embodiments, the pharmaceutical composition is formulated in a lipid nanoparticle formulation selected from liposomes or lipid nanoparticles (LNPs). In some embodiments, the pharmaceutical composition further comprises a pharmaceutically acceptable excipient or salt.
[0042] In another aspect, the present disclosure provides a method of treating a disease or condition in a subject in need thereof. In some embodiments, the method includes administering to the subject an effective amount of the pharmaceutical composition of the present disclosure.
[0043] In some embodiments, the disease or condition is selected from the group consisting of sickle cell anemia, severe combined immunodeficiency (ADA - SCID / X - SCID), cystic fibrosis, hemophilia, Duchenne muscular dystrophy, Huntington's disease, Parkinson's disease, hypercholesterolemia, α1 - antitrypsin deficiency, chronic granulomatous disease, Fanconi anemia, and Gaucher's disease. In some embodiments, the disease or condition is selected from Table 7. Brief Description of the Drawings
Brief Description of the Drawings
[0044]
Figure 1-1
Figure 1-2
[0045]
Figure 2
[0046]
Figure 3
[0047]
Figure 4
[0048]
Figure 5
[0049]
Figure 6
[0050]
Figure 7
Mode for Carrying Out the Invention
[0051] Definitions Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, only exemplary methods and materials are described. For the purposes of the present invention, the following terms are defined below.
[0052] The terms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise.
[0053] As used herein, the term "cognate" refers to an nrRT protein and a template RNA, where the nrRT protein preferentially binds to a specific template RNA. The nrRT protein and its cognate template RNA may occur naturally (referred to as natural proteins and templates), or one or both of the nrRT protein and the template RNA may be modified to preferentially bind to another nrRT protein and / or template RNA.
[0054] As used herein, the term "natural" refers to a nucleic acid or protein found in nature or in its native configuration when present in another organism or cell.
[0055] As used herein, the term "ribozyme" refers to an RNA molecule having enzymatic activity. The term includes self-cleaving ribozymes that catalyze sequence-specific intramolecular cleavage of RNA, including cleavage in cis (on the same strand).
[0056] As used herein, the term "natural ribozyme" refers to ribozymes found in nature, such as wild-type ribozymes, and includes different ribozymes found in different organisms.
[0057] As used herein, the term "cognate ribozyme" refers to a ribozyme sequence that preferentially associates with a natural or naturally occurring nrRT protein.
[0058] As used herein, the term "semi-cognate" ribozyme refers to a ribozyme derived from a closely related species that associates with an nrRT protein.
[0059] As used herein, the term "HDV RZ fold" refers to an RNA sequence that includes the fold of the hepatitis delta virus (HDV) ribozyme and retains ribozyme function.
[0060] The term "non-LTR retrotransposon reverse transcriptase protein" or "nrRT protein" refers to a reverse transcriptase protein that can copy template RNA into cDNA at a target site in the host cell genome, where cDNA synthesis is primed by a nick induced by the nrRT protein at the target site and which leads to stable double-stranded transgene insertion. The term also includes modified variants of the nrRT protein that have increased efficiency or modified nicking activity or modified binding properties (affinity) to template RNA.
[0061] The term "template RNA" refers to single-stranded RNA that binds to the nrRT protein and serves as a template for first-strand cDNA synthesis at a target site in the host cell genome.
[0062] The term "payload" refers to a compound, protein, inhibitor, or nucleic acid that is inserted into the genome of a host cell using the compositions and methods of the present disclosure.
[0063] The terms "encode", "encodes", or "encoding" refer to the transcription and / or translation of an RNA sequence to produce a product. The product can be a polypeptide, protein, or functional RNA.
[0064] The term "operably linked" refers to a sequence that is linked in a functional relationship to another sequence. For example, a promoter or enhancer is operably linked to a payload sequence when it regulates the transcription of the payload sequence. The term includes nucleic acid sequences that are covalently linked in a plasmid or vector, regardless of the number of nucleotides between the sequences. For example, a promoter is operably linked to a polyA sequence even if the payload sequence is present between the promoter and the polyA sequence.
[0065] The term "junction" refers to the position in the host cell genome where it is connected to the double-stranded cDNA into which the genomic DNA has been inserted.
[0066] The term "lipid nanoparticle" or "LNP" refers to a delivery vehicle comprising one or more lipids (e.g., cationic lipids, non-cationic lipids, PEG-modified lipids).
[0067] The term "liposome" generally refers to a vesicle composed of one or more lipids (e.g., amphiphilic lipids) arranged in a spherical bilayer or bilayers.
[0068] As used herein, "percentage of sequence identity" is determined by comparing two optimally aligned sequences over a comparison window, wherein a portion of the sequences in the comparison window may include additions or deletions (i.e., gaps) as compared to the reference sequence (which does not include additions or deletions) for the optimal alignment of the two sequences. The percentage is calculated by determining the number of positions at which the identical nucleic acid base or amino acid residue occurs in both sequences to obtain the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window, and multiplying the result by 100 to obtain the percentage of sequence identity.
[0069] As used herein, the terms "identical" or "identity" refer to two or more nucleic acid or polypeptide sequences that are the same, or portions thereof, in the context of two or more sequences. Sequences are considered "substantially identical" to each other if they have a specified percentage of the same nucleotide or amino acid residues (e.g., at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identity) over a specified region when the sequences are compared and aligned for maximum correspondence over a comparison window, or using one of the following sequence comparison algorithms, or by manual alignment and visual inspection. These definitions also refer to the complement of a test sequence.
[0070] For sequence comparison, typically one sequence acts as a reference sequence, to which a test sequence is compared. When using a sequence comparison algorithm, the test and reference sequences are input into a computer, subsequence coordinates are designated (if necessary), and sequence algorithm program parameters are designated. Default program parameters are generally used, or alternative parameters can be specified. The sequence comparison algorithm then calculates the percent sequence identity or similarity of the test sequence to the reference sequence based on the program parameters.
[0071] As used herein, a "comparison window" includes reference to any one segment of a contiguous number of positions selected from the group consisting of from 20 to 600, usually from about 50 to about 200, more usually from about 100 to about 150, in which the two sequences can be compared to a reference sequence of the same number of contiguous positions after the two sequences have been optimally aligned. Methods of sequence alignment for comparison are well known in the art. Optimal alignment of sequences for comparison can be conducted, e.g., by the local homology algorithm of Smith and Waterman (Adv. Appl. Math. 2:482, 1970), by the homology alignment algorithm of Needleman and Wunsch (J. Mol. Biol. 48:443, 1970), by the search for similarity method of Pearson and Lipman (Proc. Natl. Acad. Sci. USA 85:2444, 1988), by computerized implementations of these algorithms (e.g., GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package (Genetics Computer Group, 575 Science Dr., Madison, Wis.)), or by manual alignment and visual inspection (see, e.g., Ausubel et al., Current Protocols in Molecular Biology (1995 Supplement)).
[0072] Algorithms suitable for determining percent sequence identity and percent sequence similarity are the BLAST and BLAST 2.0 algorithms, which are described in Altschul et al. (Nuc. Acids Res. 25:3389-402, 1977) and Altschul et al. (J. Mol. Biol. 215:403-10, 1990), respectively. Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information (http: / / www.ncbi.nlm.nih.gov / ). This algorithm involves first identifying high-scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence that either match or satisfy some positive-valued threshold score T when aligned with words of the same length in the database sequence. T is referred to as the neighborhood word score threshold (Altschul et al., supra). These initial neighborhood word hits act as seeds for initiating a search to find longer HSPs that contain them. The word hits are extended in both directions along each sequence as far as the cumulative alignment score can be increased. The cumulative score is calculated using, for nucleotide sequences, the parameters M (reward score for pairs of matching residues; always >0) and N (penalty score for mismatching residues; always <0). For amino acid sequences, a scoring matrix is used to calculate the cumulative score. Extension of the word hits in each direction is stopped when: the cumulative alignment score drops by an amount X from its maximum achieved value; the cumulative score goes to zero or below due to the accumulation of one or more negative-scoring residue alignments; or the end of either sequence is reached. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment. The BLASTN program (for nucleotide sequences) uses, by default, a word length (W) of 11, an expectation value (E) of 10, M = 5, N = -4, and a comparison of both strands.Regarding amino acid sequences, the BLASTP program, by default, uses a word length of 3, an expectation value (E) of 10, and an alignment (B) of 50, an expectation value (E) of 10, M = 5, N = -4 with the BLOSUM62 scoring matrix (see Henikoff and Henikoff, Proc. Natl. Acad. Sci. USA 89:10915, 1989).
[0073] The BLAST algorithm also performs a statistical analysis of the similarity between two sequences (see, for example, Karlin and Altschul, Proc. Natl. Acad. Sci. USA 90:5873-87, 1993). One measure of similarity provided by the BLAST algorithm is the minimum total probability (P(N)), which provides an indication of the probability that a match between two nucleotide or amino acid sequences occurs by chance. For example, a nucleic acid is considered similar to a reference sequence if the minimum total probability in a comparison of the test nucleic acid to the reference nucleic acid is less than about 0.2, typically less than about 0.01, and more typically less than about 0.001.
[0074] The term "heterologous" refers to any polynucleotide or polypeptide sequence that does not occur naturally in a host cell or organism or that is inserted at a position that does not occur naturally in the host cell or organism.
[0075] The term "vector" refers to DNA that contains foreign or heterologous DNA, typically double-stranded DNA. The above term includes plasmids and viral vectors. A vector may contain a polynucleotide sequence that promotes autonomous replication of the vector in a host cell. The above vector can be used to replicate foreign or heterologous DNA in a suitable host cell. Further, the above vector may also contain elements that enable the inserted DNA to be transcribed into one or more mRNA molecules. An expression vector further contains sequence elements operably linked to the inserted DNA that increase the half-life of the expressed mRNA and / or enable the mRNA to be translated into protein molecules.
[0076] Detailed Description of the Invention The present disclosure provides compositions and methods for improving gene therapy techniques for introducing heterologous polynucleotides into target cells. The present disclosure provides a method (site-specific integration) for inserting a heterologous polynucleotide into a target site in the genome of a target cell. The heterologous polynucleotide may include a transgene encoding a therapeutic protein or a non-protein regulatory element. In some embodiments, the cell is a eukaryotic cell (e.g., a mammalian cell).
[0077] The present disclosure provides many advantages over current gene therapy techniques, including: 1) the technique is an RNA-based therapy that uses gene synthesis of an RNA template into the target cell genome, thereby avoiding problems associated with DNA delivery to cells, such as unwanted genetic changes that can impair cell function and promote carcinogenesis; 2) the heterologous polynucleotide can be inserted into so-called "safe harbor" sites that do not cause harmful or undesirable changes to the target cell genome or cell physiology; 3) there are no known limitations on the size or length of the heterologous polynucleotide inserted into the target cell genome; and 4) there is no requirement for cell division, and as a result, post-mitotic cells (e.g., neurons) can be targeted. Some of the advantages of the present disclosure, referred to as THERAPEUTIC ADDITION by CONTROLLED SYNTHESIS INSERTION (TASCI™), are shown in Table 1. Table 1. [Table 1]
[0078] The compositions and methods of the present disclosure utilize a two-RNA delivery system for introducing a heterologous polynucleotide into a target cell; 1) a first RNA (e.g., mRNA) encoding a non-LTR retrotransposon reverse transcriptase protein (nrRT); and 2) a second RNA (also referred to as template RNA) comprising a protein coding sequence (or open reading frame "ORF") and a sequence that binds to the nrRT. The system may further comprise a delivery system for introducing the two RNAs into the cytoplasm of the target cell. In some embodiments, the delivery system comprises lipid nanoparticles (LNPs).
[0079] After delivering the two RNAs into the cytoplasm of the target cells, the mRNA is translated by the endogenous protein synthesis components of the cells to produce the nrRT protein. The nrRT protein then binds to the template RNA to form a ribonucleoprotein (RNP) complex and enters the nucleus of the target cells. Without being bound by theory, it is currently believed that after delivery into the nucleus, the endonuclease (EN) domain of the nrRT protein cleaves the bottom strand of the target genomic DNA, thereby providing a 3'-hydroxyl end that serves as a primer for the reverse transcription of the template RNA by the reverse transcriptase (RT) domain of the nrRT protein. After the first-strand synthesis to generate cDNA, the EN domain or a host endonuclease cleaves the opposite side (e.g., the top strand) of the genomic DNA. The nick in the top strand generates another 3'-hydroxyl end that serves as a primer for the second-strand cDNA synthesis. Whether the second-strand DNA synthesis is performed by the nrRT or by cellular polymerases is currently unknown. The nick is then repaired, resulting in the integration of the double-stranded cDNA into the target site in the genomic DNA. The proposed mechanism is shown in Figure 1.
[0080] Although the two RNAs do not necessarily include the nrRT protein and its naturally occurring cognate template RNA or its modified variant, it is understood by those skilled in the art that both the nrRT protein and the template RNA can be separately engineered to bind to different nrRTs and / or template RNAs.
[0081] Eukaryotic non-LTR retrotransposon reverse transcriptase protein (nrRT) In some embodiments, the present disclosure provides an RNA (e.g., mRNA) encoding an nrRT protein. In some embodiments, the nrRT protein comprises one or more of a DNA binding domain, an RNA binding domain, a reverse transcriptase domain, and an endonuclease domain, or a combination thereof. The endonuclease domain of the nrRT protein of the present disclosure generates a single-stranded nick in genomic DNA at the target site, generating a free 3' end of the genomic DNA that serves as a primer for reverse transcribing the template RNA into cDNA. After first-strand cDNA synthesis, the nrRT protein introduces a nick in the second strand, creating another 3' end of the genomic DNA that serves as a primer for second-strand cDNA synthesis at the target site. This results in a double-stranded DNA molecule that is inserted into the target site in the host cell genomic DNA.
[0082] It is understood that the present disclosure encompasses any eukaryotic nrRT protein capable of binding and reverse transcribing template RNA at a target site in the host cell genome. In some embodiments, the nrRT protein is an nrRT protein isolated from Zonotrichia albicollis, Taeniopygia guttata, Tinamus guttatus, Geospiza fortis, Pungitis pungitis, Oryzias latipes, Daniorerio, Oryzias melastigma, Petromyzon marinus, Salmo trutta, Salmo salar, Gasterosteus aculeatus, Drosophila mercatorum, Drosophila melanogaster, Nasonia vitripennis, Tribolium castaneum, Drosophila simulans, Apis cerana, Bombyx mori, Lepidurus couesii, Triops cancriformis, Limulus polyphemus, Hydra magnipapillata, Adineta vaga, or Ciona intestinalis, or a modified functional variant thereof.In some embodiments, the mRNA encodes an amino acid sequence that is substantially identical to an nrRT protein isolated from Zonotrichia albicollis, Taeniopygia guttata, Tinamus guttatus, Geospiza fortis, Pungitis pungitis, Oryzias latipes, Daniorerio, Oryzias melastigma, Petromyzon marinus, Salmo trutta, Salmo salar, Gasterosteus aculeatus, Drosophila mercatorum, Drosophila melanogaster, Nasonia vitripennis, Tribolium castaneum, Drosophila simulans, Apis cerana, Bombyx mori, Lepidurus couesii, Triops cancriformis, Limulus polyphemus, Hydra magnipapillata, Adineta vaga, or Ciona intestinalis.In some embodiments, the mRNA encodes an amino acid sequence having at least 60% sequence identity (e.g., identity greater than or equal to 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95%) with an nrRT protein isolated from Zonotrichia albicollis, Taeniopygia guttata, Tinamus guttatus, Geospiza fortis, Pungitis pungitis, Oryzias latipes, Danio rerio, Oryzias melastigma, Petromyzon marinus, Salmo trutta, Salmo salar, Gasterosteus aculeatus, Drosophila mercatorum, Drosophila melanogaster, Nasonia vitripennis, Tribolium castaneum, Drosophila simulans, Apis cerana, Bombyx mori, Lepidurus couesii, Triops cancriformis, Limulus polyphemus, Hydra magnipapillata, Adineta vaga, or Ciona intestinalis. In some embodiments, the nrRT protein includes nrRT proteins isolated from other animals.
[0083] In some embodiments, the RNA encoding the nrRT includes one or more of a 5' cap, 5' UTR, open reading frame (ORF) encoding the nrRT, 3' URT, or polyA sequence at the 3' end.
[0084] In some embodiments, the RNA encoding the nrRT includes one or more modified uridine (U) nucleosides as described herein.
[0085] An illustration of an exemplary mRNA encoding the nrRT of the present disclosure is shown in FIG. 2.
[0086] Template RNA In some aspects, the template RNA of the present disclosure includes (i) a promoter, (ii) a payload sequence, (iii) a polyA sequence, and (iv) an nrRT binding sequence. In some embodiments, the elements of the template RNA are operably linked to each other. It is understood that the relative positions of the individual elements in the template RNA can vary in the 5' to 3' direction. For example, in some embodiments, the template RNA includes elements (i), (ii), (iii), and (iv) in the 5' to 3' direction. In some embodiments, the template RNA includes elements (iv), (i), (ii), and (iii) in the 5' to 3' direction.
[0087] It is further understood that the individual elements in the template RNA can vary in their 5' to 3' orientation relative to other elements. For example, in some embodiments, the promoter (i), the payload sequence (ii), and / or the poly sequence (iii) are in the reverse 5' to 3' orientation relative to element (iv). Further, in some embodiments, the direction of transcription of the payload sequence in the template can be reversed. As a result, in one orientation, the promoter (i) is closest to the 5' end of the template RNA, or in a second orientation, the promoter (i) is closest to the 3' end of the template RNA.
[0088] An illustration of an exemplary template RNA of the present disclosure is shown in FIG. 2.
[0089] In some embodiments, the promoter is an RNA polymerase (Pol) II promoter. In some embodiments, the promoter is selected from the EFS promoter, the ABPnat mini promoter, the CRNM-TTR enhancer promoter, the AAV-rDNA TTR promoter, or the CBh promoter.
[0090] In some embodiments, the payload array encodes a reporter protein such as GFP or luciferase. In some embodiments, the payload array encodes a therapeutic protein that replaces or complements a defective gene or protein. In some embodiments, the therapeutic protein is used to treat a disease or condition in a subject or patient. In some embodiments, the therapeutic protein is selected from the group consisting of Factor VIII, Factor IX, and phenylalanine hydroxylase (PAH).
[0091] In some embodiments, the payload array encodes a protein in the "Gene Name" column in Table 7 below. In some embodiments, the therapeutic protein is used to treat a disease or condition shown in Table 7 below.
[0092] In some embodiments, the payload array encodes an inhibitor of another protein. In some embodiments, the inhibitor is a single-chain antibody.
[0093] In some embodiments, the payload array encodes a regulatory RNA. In some embodiments, the regulatory RNA is selected from a ligand-binding riboswitch (e.g., a ligand-activated riboswitch or an allosteric ribozyme (aptazyme)), a small RNA (sRNA), a small interfering RNA (siRNA), or a short hairpin RNA (shRNA).
[0094] In some embodiments, the polyA sequence is selected from a short SV40 poly, SNRP1 polyA, a synthetic polyA, BHG polyA, or BGH polyA min. In some embodiments, the template RNA includes a WPRE3 3' enhancer.
[0095] In some embodiments, the nrRT binding sequence comprises a sequence isolated from the 3' region of a natural non-LTR retroelement or an organism comprising a non-LTR retroelement. In some embodiments, the nrRT binding sequence comprises a 3'UTR sequence. In some embodiments, the 3'UTR sequence is isolated from an organism comprising a non-LTR retroelement. In some embodiments, the 3'UTR sequence is isolated from an organism selected from the group consisting of G. aculeatus, D. melanogaster, L. polyphemus, P. pungitis, N. vitripennis, G. fortis, O. latipes, Z. albicollis, T. guttata, T. castaneum, T. guttatus, D. simulans, B. mori, and A. vaga. In some embodiments, the nrRT binding sequence comprises a sequence having a sequence identity greater than or equal to 60% (e.g., greater than or equal to 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95% or 99%) with a sequence isolated from G. aculeatus, D. melanogaster, L. polyphemus, P. pungitis, N. vitripennis, G. fortis, O. latipes, Z. albicollis, T. guttata, T. castaneum, T. guttatus, D. simulans, B. mori, or A. vaga. In some embodiments, the 3'UTR comprises a sequence having a sequence identity greater than or equal to 60% (e.g., greater than or equal to 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100%) with a sequence selected from any one of SEQ ID NOs: 26-39.
[0096] In some embodiments, the nrRT binding sequence comprises a modified (non-natural) sequence. For example, the nrRT binding sequence can be modified to increase or decrease binding to the nrRT protein of the present disclosure.
[0097] Modified uridine In some embodiments, the mRNA encoding the nrRT and / or the template RNA comprise one or more modified uridine (U) nucleosides. RNA containing unmodified uridine can activate the innate immune response and is less stable in cells. Modified uridines can provide the following advantages: i) they reduce the innate immune response in the host organism when the cells are transfected with the nrRT mRNA and template RNA of the present disclosure, ii) increase RNA stability, and iii) increase the amount of protein produced when the RNA is transcribed.
[0098] In some embodiments, the mRNA encoding the nrRT protein comprises one or more modified uridine (U) nucleosides selected from the group consisting of N1-methyl-pseudouridine (N1mΨU), pseudouridine (ΨU), 5-methyluridine (5meU), 5-methoxyuridine (5moU), and mixtures thereof. In some embodiments, the ORF encoding the nrRT comprises a modified uridine (U) selected from one of the following: N1-methyl-pseudouridine (N1mΨU), pseudouridine (ΨU), 5-methyluridine (5meU), or 5-methoxyuridine (5moU). In some embodiments, the ORF encoding the nrRT comprises N1-methyl-pseudouridine (N1mΨU). In some embodiments, the ORF encoding the nrRT comprises unmodified uridine, and a mixture or combination of modified uridines selected from the group consisting of N1mΨU, ΨU, 5meU, and 5moU. The structure of the modified uridine is shown in FIG. 3.
[0099] In some embodiments, the template RNA comprises one or more modified uridine nucleosides. Unexpectedly, the inventors have determined that template RNAs containing one or more modified uridines result in successful incorporation and expression of the payload sequence at the target site in the genome.
[0100] In some embodiments, the template RNA comprises one or more modified uridines selected from the group consisting of N1-methyl-pseudouridine (N1mΨU), pseudouridine (ΨU), 5-methyluridine (5meU), 5-methoxyuridine (5moU), and mixtures thereof. In some embodiments, the template RNA comprises a single type of modified uridine selected from among: N1-methyl-pseudouridine (N1mΨU), pseudouridine (ΨU), 5-methyluridine (5meU), or 5-methoxyuridine (5moU). In some embodiments, the template RNA comprises N1-methyl-pseudouridine (N1mΨU). In some embodiments, the template RNA comprises unmodified uridine, as well as a mixture or combination of modified uridines selected from the group consisting of N1mΨU, ΨU, 5meU, and 5moU.
[0101] In some embodiments, the template RNA containing modified uridine is uncleavable by ribozymes. In some embodiments, the template RNA containing the modified uridine N1-methyl-pseudouridine (N1mΨU) or pseudouridine (ΨU) is uncleavable by ribozymes. In some embodiments, the template RNA containing modified uridine increases the efficiency of insertion into the eukaryotic genome compared to template RNA containing unmodified uridine.
[0102] In some embodiments, cytotoxicity is reduced when the template RNA contains modified uridine.
[0103] The modified uridine is distributed throughout the template RNA sequence, and in some embodiments, it will be understood by those skilled in the art that all of the uridines may contain the same modified uridine (e.g., all of the above uridines are N1-methyl-pseudouridine (N1mΨU), or all of the modified uridines are pseudouridine (ΨU)).
[0104] 5’ ribozyme As is known in the art, natural RNA templates (which bind to their cognate nrRT proteins) contain an active ribozyme at the 5’ end. The self-cleavage function of the ribozyme was previously thought to be extremely important for genomic insertion. Thus, in some embodiments, the template RNA further comprises, or optionally comprises, an active or functional 5’ ribozyme sequence. In some embodiments, the 5’ ribozyme is selected from an HDV ribozyme (e.g., HDV_ac2, HDV_gu1, HDV_gu5b, HDV_gu6, HDV_gu5b_NP2), TriCasA ribozyme, L8 ribozyme (e.g., L8_gu6), SL28 ribozyme, or a natural cognate or semi-cognate ribozyme, or a modified variant thereof. In some embodiments, the ribozyme sequence comprises a sequence selected from any one of SEQ ID NO: 3 or 13-22 (without the pp7 binding sequence), or a sequence having a sequence identity greater than or equal to 60% (e.g., greater than or equal to 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity) with a sequence selected from any one of SEQ ID NO: 3 or 13-22 (without the pp7 binding sequence).
[0105] However, in contrast to the teachings in the art, the inventors have unexpectedly determined that template RNAs engineered to have 5' ribozymes with reduced activity, catalytically inactive ribozymes, and ribozymes that are not cleaved can be successfully used to insert heterologous polynucleotides at target sites in the genomic DNA of target cells. Accordingly, in some embodiments, the 5' ribozyme sequence is selected from the sequences of partially active ribozymes, ribozymes with reduced catalytic activity, or catalytically inactive ribozymes. In some embodiments, the template RNA does not contain a functional 5' ribozyme sequence. In some embodiments, the template RNA does not contain a 5' ribozyme sequence.
[0106] Additional components In some embodiments, the template RNA includes, further includes, or optionally includes 5' elements and 3' elements that regulate transcription, translation, and / or insertion of the payload sequence at a target located in the host cell genome. Non-limiting examples of these elements are described below.
[0107] In some embodiments, the template RNA includes a Kozak consensus translation initiation site upstream or 5' of the payload sequence. In some embodiments, the Kozak sequence includes the sequence (5'-GCCACC-3' SEQ ID NO: 7).
[0108] In some embodiments, the template RNA includes an RNA polymerase (RNAP) terminator sequence located 5' of a promoter sequence. The RNAP terminator sequence functions to stop RNA polymerase read-through from a gene at the target insertion site. In some embodiments, the RNAP terminator sequence includes the sequence 5'-AGGTCGACCAGATGTCCGAGGTCGACCAGTTGTCCG-3' (SEQ ID NO: 4).
[0109] In some embodiments, the template RNA comprises a 5’ sequence or 5’ modification that protects the 5’ end from degradation. In some embodiments, the 5’ modification comprises a 5’ cap structure.
[0110] In some embodiments, the template RNA comprises a 5’ sequence that facilitates site-specific insertion of the heterologous polynucleotide into a target site in the eukaryotic genome.
[0111] In some embodiments, the template RNA comprises a 3’ sequence that facilitates site-specific insertion of the heterologous polynucleotide into the eukaryotic genome. In some embodiments, the template RNA comprises a 3’ sequence that enhances the efficiency and fidelity of target-primed reverse transcription.
[0112] In some embodiments, the template RNA comprises a sequence useful for purification of the template RNA. In some embodiments, the sequence useful for purification of the template RNA comprises a hairpin structure that binds to the PP7 coat protein or a truncated version thereof. See, e.g., Hogg, J.R. & Collins, K. RNA-based affinity purification reveals 7SK RNPs with distinct composition and regulation. RNA 13, 868-880 (2007).
[0113] In some embodiments, the template RNA comprises a sequence that binds to a DNA-binding protein, which enables enrichment of the inserted double-stranded sequence in the target DNA by purifying a genomic DNA fragment comprising a sequence that binds to the DNA-binding protein. In some embodiments, the payload sequence is adjacent to a sequence that binds to a DNA-binding protein, such that one sequence is located on the 5' side of the payload sequence (e.g., upstream of the promoter sequence) and the other sequence is located on the 3' side of the payload sequence (e.g., downstream of the polyA sequence). In some embodiments, the template RNA comprises a lacO operator sequence that binds to the LacI protein. In some embodiments, the template RNA comprises a first lacO operator sequence located on the 5' side of the payload sequence and a second lacO operator sequence located on the 3' side of the payload sequence.
[0114] In some embodiments, the template RNA comprises a polyA sequence located on the 3' side of the nrRT binding sequence.
[0115] In some embodiments, the template RNA comprises: (a) a 5' sequence that is homologous to a DNA sequence located 5' to the target insertion site in the eukaryotic genome; or (b) a 3' sequence that is homologous to a DNA sequence located 3' to the target insertion site in the eukaryotic genome; or both (a) and (b). In some embodiments, the 5' homologous sequence comprises about 1 to 36 nucleotides of a sequence that base pairs with a complementary sequence at the target site. In some embodiments, the 3' homologous sequence comprises about 1 to 30 nucleotides of a sequence that base pairs with a complementary sequence at the target site.
[0116] In some embodiments, the template RNA does not contain a 5' phosphate.
[0117] Method for inserting a polynucleotide into a target site in a genome The present disclosure also provides a method for inserting a heterologous polynucleotide into a eukaryotic genome at a target site. In some embodiments, the method comprises transfecting a eukaryotic cell with (a) RNA encoding a non-LTR retrotransposon reverse transcriptase protein (nrRT) comprising a reverse transcriptase domain and an endonuclease domain; and (b) a template RNA. In some embodiments, the template RNA comprises a promoter, a payload sequence, a polyA sequence, and an nrRT binding sequence.
[0118] In some embodiments, the template RNA comprises one or more modified uridine (U) nucleosides selected from the group consisting of N1-methyl-pseudouridine (N1mΨU), pseudouridine (ΨU), 5-methyluridine (5meU), 5-methoxyuridine (5moU), and mixtures thereof. In some embodiments, the template RNA comprises unmodified uridine, and a mixture of one or more modified uridines selected from the group consisting of N1mΨU, ΨU, 5meU, and 5moU.
[0119] In some embodiments, the template RNA comprising modified uridine is uncleavable by ribozymes.
[0120] In some embodiments, the nrRT is expressed in the cell and catalyzes the insertion of a double-stranded heterologous polynucleotide comprising the payload sequence at a target site in the eukaryotic genome.
[0121] The method provides the advantage that the insertion efficiency of the payload sequence into the eukaryotic genome is increased when the template RNA comprising modified uridine is compared to a template RNA comprising unmodified uridine.
[0122] In some embodiments, the template RNA further comprises a 5' ribozyme sequence selected from active ribozymes. In some embodiments, the ribozyme is selected from HDV ribozymes, TriCasA ribozymes, natural cognate ribozymes, semi-cognate ribozymes, or variants thereof.
[0123] The method also provides the unexpected advantage that the template RNA does not require a functional ribozyme for insertion and expression of the payload sequence. Thus, in some embodiments, the template RNA comprises a 5' ribozyme sequence selected from partially active ribozymes, ribozymes having reduced catalytic activity, or catalytically inactive ribozymes. In some embodiments, the template RNA does not comprise a functional 5' ribozyme sequence.
[0124] The method also provides the unexpected advantage that the template RNA does not require a 5' ribozyme sequence for insertion and expression of the payload sequence. Thus, in some embodiments, the template RNA does not comprise a 5' ribozyme sequence.
[0125] Template RNAs containing modified uridines can also reduce cytotoxicity compared to template RNAs containing unmodified uridines. Thus, in some embodiments, cytotoxicity is reduced when the template RNA contains a modified uridine selected from the group consisting of N1-methyl-pseudouridine (N1mΨU), pseudouridine (ΨU), 5-methyluridine (5meU), 5-methoxyuridine (5moU), and mixtures thereof.
[0126] In some embodiments of the above method, when the molar ratio of the nrRT mRNA to the template RNA delivered to the target cells is increased, the insertion efficiency of the payload sequence at the target site in the genome increases compared to the equimolar (1:1) ratio. In some embodiments, when the total amount of RNA delivered to the target cells is increased, the insertion efficiency of the payload sequence at the target site in the genome increases. In some embodiments, when both the molar ratio of the nrRT mRNA to the template RNA and the total amount of RNA delivered to the target cells are increased, the insertion efficiency of the payload sequence at the target site in the genome increases. Representative non-limiting examples showing the results of the molar ratio of nrRT to template RNA and total RNA for payload expression are described in the Examples.
[0127] In some embodiments of the above method, the payload sequence encodes a therapeutic protein that replaces or complements a defective gene or protein. In some embodiments, the therapeutic protein is used to treat a disease or condition in a subject or patient. In some embodiments, the therapeutic protein is selected from the group consisting of Factor VIII, Factor IX, and phenylalanine hydroxylase (PAH).
[0128] In some embodiments of the above method, the payload sequence encodes an inhibitor of another protein. In some embodiments, the inhibitor is a single-chain antibody.
[0129] In some embodiments of the above method, the payload sequence encodes a regulatory RNA. In some embodiments, the regulatory RNA is selected from a ligand-binding riboswitch (e.g., a ligand-activated riboswitch or an allosteric ribozyme (aptazyme)), a small RNA (sRNA), a small interfering RNA (siRNA), or a short hairpin RNA (shRNA).
[0130] In some embodiments, the method includes the step of transfecting a eukaryotic cell. In some embodiments, the eukaryotic cell is transfected in vitro. In some embodiments, the eukaryotic cell is transfected in vivo. In some embodiments, the eukaryotic cell is a mammalian cell. In some embodiments, the eukaryotic cell is a human cell.
[0131] In some embodiments, the cell is transfected with an LNP formulation, a lipofection reagent, or by electroporation. In some embodiments, the cell is neither transduced nor transfected with a viral vector.
[0132] In some embodiments of the method, the template RNA comprises, further comprises, or optionally comprises 5' elements and 3' elements that regulate the transcription, translation, and / or insertion of the payload sequence at a target located in the host cell genome. Non-limiting examples of these elements are described below.
[0133] In some embodiments, the template RNA comprises a Kozak consensus translation initiation site upstream or 5' of the payload sequence.
[0134] In some embodiments, the template RNA comprises an RNA polymerase (RNAP) terminator sequence located 5' of a promoter sequence. The RNAP terminator sequence functions to stop RNA polymerase read-through from a gene at the target insertion site.
[0135] In some embodiments, the template RNA comprises a 5' sequence or 5' modification that protects the 5' end from degradation. In some embodiments, the 5' modification comprises a 5' cap structure.
[0136] In some embodiments, the template RNA comprises a 5' sequence that promotes site-specific insertion of the heterologous polynucleotide into a target site in the eukaryotic genome.
[0137] In some embodiments, the template RNA comprises a 3' sequence that promotes site-specific insertion of the heterologous polynucleotide into the eukaryotic genome. In some embodiments, the template RNA comprises a 3' sequence that enhances the efficiency and fidelity of target-primed reverse transcription.
[0138] In some embodiments, the template RNA comprises a sequence useful for purification of the template RNA. In some embodiments, the sequence useful for purification of the template RNA comprises a hairpin structure that binds to the PP7 coat protein or a truncated version thereof. See, for example, Hogg, J.R. & Collins, K. RNA-based affinity purification reveals 7SK RNPs with distinct composition and regulation. RNA 13, 868-880 (2007).
[0139] In some embodiments, the template RNA includes a sequence that binds to a DNA-binding protein, which enables enrichment of the inserted double-stranded sequence in the target DNA by purifying a genomic DNA fragment that includes a sequence that binds to the DNA-binding protein. In some embodiments, the payload sequence is adjacent to a sequence that binds to a DNA-binding protein, such that one sequence is located on the 5' side of the payload sequence (e.g., upstream of a promoter sequence), and the other sequence is located on the 3' side of the payload sequence (e.g., downstream of the polyA sequence). In some embodiments, the template RNA includes a lacO operator sequence that binds to the LacI protein. In some embodiments, the template RNA includes a first lacO operator sequence located on the 5' side of the payload sequence and a second lacO operator sequence located on the 3' side of the payload sequence.
[0140] In some embodiments, the template RNA further includes a polyA sequence located on the 3' side of the nrRT binding sequence. In some embodiments, the template RNA does not include a 5' phosphate.
[0141] In some embodiments, the template RNA includes: (a) a 5' sequence that is homologous to a DNA sequence located 5' to the target insertion site in the eukaryotic genome; or (b) a 3' sequence that is homologous to a DNA sequence located 3' to the target insertion site in the eukaryotic genome; or both (a) and (b). In some embodiments, the 5' homologous sequence includes a homologous sequence of about 1 to 36 nucleotides that base pairs with a complementary sequence at the target site. In some embodiments, the 3' homologous sequence includes a homologous sequence of about 1 to 30 nucleotides that base pairs with a complementary sequence at the target site.
[0142] In some embodiments, the target insertion site is located within ribosomal RNA genes or ribosomal DNA (rDNA). In some embodiments, the target insertion site is located within genomic DNA encoding ribosomal RNA (rRNA). In some embodiments, the target insertion site is located within the 5S, 8S, 18S, or 28S rDNA sequence.
[0143] In some embodiments of the above method, the nrRT binding sequence comprises a sequence isolated from the 3' region of a natural non-LTR retroelement or an organism containing a non-LTR retroelement. In some embodiments, the nrRT binding sequence comprises a 3'UTR sequence. In some embodiments, the 3'UTR sequence is isolated from an organism containing a non-LTR retroelement. In some embodiments, the 3'UTR sequence is isolated from an organism selected from the group consisting of G. aculeatus, D. melanogaster, L. polyphemus, P. pungitis, N. vitripennis, G. fortis, O. latipes, Z. albicollis, T. guttata, T. castaneum, T. guttatus, D. simulans, B. mori, and A. vaga. In some embodiments, the nrRT binding sequence comprises a sequence having a sequence identity greater than or equal to 60% (e.g., greater than or equal to 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%) with a sequence isolated from G. aculeatus, D. melanogaster, L. polyphemus, P. pungitis, N. vitripennis, G. fortis, O. latipes, Z. albicollis, T. guttata, T. castaneum, T. guttatus, D. simulans, B. mori, or A. vaga. In some embodiments, the 3'UTR comprises a sequence having a sequence identity greater than or equal to 60% (e.g., greater than or equal to 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%) with a sequence selected from any one of SEQ ID NOs: 26 to 39.
[0144] In some embodiments, the nrRT binding sequence comprises a modified (non-natural) sequence. For example, the nrRT binding sequence can be modified to increase or decrease binding to the nrRT protein of the present disclosure.
[0145] In some embodiments of the above method, the RNA encoding the nrRT comprises one or more modified uridine (U) nucleosides selected from the group consisting of N1-methyl-pseudouridine (N1mΨU), pseudouridine (ΨU), 5-methyluridine (5meU), 5-methoxyuridine (5moU), and mixtures thereof, or unmodified U, and a mixture of modified U selected from the group consisting of N1mΨU, ΨU, 5meU, and 5moU.
[0146] Safe harbor insertion site In some embodiments, the heterologous polynucleotide is inserted into a so-called "safe harbor" site in the host cell genome, which does not change normal cell physiology or metabolism. Examples of safe harbor sites include regions of the genome with high-copy-number repeated genes such that disruption of one gene does not significantly change normal cell physiology or metabolism. Examples of high-copy-number regions include rDNA genes encoding rRNA. Thus, in some embodiments, the target insertion site is located within a ribosomal RNA gene or ribosomal DNA (rDNA). In some embodiments, the heterologous polynucleotide is inserted into genomic DNA encoding ribosomal RNA (rRNA). In some embodiments, the heterologous polynucleotide is inserted into a 5S, 8S, 18S, or 28S rDNA sequence.
[0147] Delivery method The compositions of the present disclosure can be introduced into target cells using methods compatible with RNA delivery. In some embodiments, the mRNA encoding the nrRT protein and the template RNA are introduced into the target cells using lipid nan formulations (e.g., liposomes or lipid nanoparticles (LNP)), lipofection reagents, or by electroporation. In some embodiments, the target cells are not transduced with a virus. Viral transduction is associated with various undesirable effects on cells, including mutations in the host cell chromosome, random integration, and the presence of double-strand breaks that can cause cytotoxicity.
[0148] Pharmaceutical composition Also provided are pharmaceutical compositions comprising the mRNA encoding the nrRT protein and the template RNA described herein. In some embodiments, the pharmaceutical composition comprises a lipid nan formulation (e.g., liposome or lipid nanoparticle (LNP)). In some embodiments, the pharmaceutical composition comprises a pharmaceutically acceptable excipient or salt. Examples of pharmaceutically acceptable excipients are described in the United States Pharmacopeia (USP), European Pharmacopeia (EP), British Pharmacopeia, and International Pharmacopeia.
[0149] Treatment method Also provided is a method of treating a subject or patient with the RNA compositions described herein. The method can be used to treat diseases associated with a defective or mutated gene in a subject (e.g., diseases caused by single gene defects such as sickle cell anemia, severe combined immunodeficiency (ADA-SCID / X-SCID), cystic fibrosis, hemophilia, Duchenne muscular dystrophy, Huntington's disease, Parkinson's disease, hypercholesterolemia, α1-antitrypsin deficiency, chronic granulomatous disease, Fanconi anemia, and Gaucher disease, but not limited thereto). In some embodiments, the method can be used to treat spinal muscular atrophy and hereditary retinal dystrophy.
[0150] In some embodiments, the method can be used to treat polygenic disorders including, but not limited to, heart disease, cancer, diabetes, schizophrenia, Parkinson's disease, and Alzheimer's disease.
[0151] In some embodiments, the method can be used to treat infectious diseases such as HIV.
[0152] For example, in a patient with hemophilia A, the payload can encode wild-type factor VIII protein. In a patient with hemophilia B, the payload can encode wild-type factor IX protein. In some embodiments, the payload can encode wild-type p53 gene in a subject having a defective p53 gene to help prevent tumor growth.
[0153] Representative examples of diseases or conditions that can be treated by the methods of the present disclosure are shown in Table 7.
[0154] In some aspects, the method is an in vivo method. In some embodiments, the method is an ex vivo method.
[0155] In some embodiments, the method includes administering an effective dose of the pharmaceutical composition of the present disclosure to a patient in need of treatment. The pharmaceutical composition can be administered via any suitable method that results in targeted integration of the payload sequence into one or more cells of the subject. In some embodiments, the pharmaceutical composition is administered intravenously, intramuscularly, subcutaneously, intravitreally, intravascularly, into the CNS or other neural tissue, or intranasally.
[0156] The effective dosage can range from 0.1 to 100 mg of active ingredient / kg body weight (including the endpoint and any sub-range thereof) of the subject or patient. The effective dosage can also range from 1 microgram to 200 micrograms of active ingredient (including the endpoint and any sub-range thereof) per dose for adults. The effective dosage can be readily determined by a skilled medical professional.
[0157] In some embodiments, the cells are removed from the subject or patient before being transfected ex vivo with the mRNA encoding the nrRT protein of the present disclosure and the template RNA. In some embodiments, the subject or patient is human, the cells are removed from a human, and transfected with the mRNA encoding the nrRT protein of the present disclosure and the template RNA. After transfection ex vivo, the accurate insertion of the heterologous polynucleotide containing the payload sequence can be determined, for example, by amplifying the sequence at the 5' and / or 3' insertion junctions and / or by amplifying the payload sequence. The accurately targeted insertion can also be determined by sequencing the genomic target site. The expression of the payload sequence can also be determined, for example, by detecting the expression of the product (e.g., protein or regulatory RNA) encoded by the payload sequence. After the accurate integration and / or expression of the payload sequence is determined, the accurately targeted cells are administered to the subject (autologous treatment).
Example
[0158] Example The following examples are provided to illustrate, but not to limit, the claimed invention.
[0159] Example 1. This example provides a representative method for generating template RNA containing modified uridine.
[0160] The plasmid DNA used for in vitro transcription (IVT) to generate template RNA is digested to completion using the restriction enzymes BbsI-HF and PvuI. The linearized plasmid is purified using phenol chloroform isoamyl alcohol (PCI) extraction and quantified using a Nanodrop. In vitro transcription is performed using the HiScribe T7 High Yield RNA Synthesis Kit (NEB, cat# E2040) in the presence of the recommended amount of T7 RNA polymerase mix, the corresponding reaction buffer, 50 ng / ul linearized plasmid DNA, and 10 mM each of ATP, GTP, CTP, and their corresponding modified UTP. The above reaction mixture is incubated at 37 °C for 2 hours and subsequently DNase-treated to remove the DNA template. For each 20 ul IVT reaction, 2 ul of DNase I (NEB, cat# M0303S), 10 ul of 10× DNase buffer, and 68 ul nuclease-free water are added to a final volume of 100 ul. Incubate at 37 °C for 3 hours.
[0161] Purify the resulting RNA transcripts using Oligo(d)T25 magnetic beads (NEB Cat#S1419S). Use 5 mg of beads for each 20 μl IVT reaction. Equilibrate the beads by washing them three times with 250 μl of 1× wash buffer (20 mM Tris-HCL, pH 7.5, 500 mM LiCl, and 1 mM EDTA). Mix the DNase-treated RNA with 2× binding buffer (0.1% Triton® X-100 in 2× wash buffer) in a 1:1 (v:v) ratio and pipette to mix with the corresponding amount of equilibrated beads. Incubate at 37 °C for 5 minutes and then at room temperature for 15 minutes on a rotator. Wash the beads three times with 250 μl of 1× wash buffer and then once with 250 μl of 1× low-salt wash buffer (20 mM Tris-HCL, pH 7.5, 200 mM LiCl, and 1 mM EDTA). Elute the RNA by adding 100 μl of nuclease-free H2O to the beads and incubating at 37 °C for 5 minutes.
[0162] Use CIAP (Promega Cat# M2825) to remove 5’ triphosphate from the RNA transcripts. Mix the eluted dT-purified RNA (100 μl) with 0.5 μl of CIAO, 30 μl of 5× CIAP buffer, and 19.5 μl of nuclease-free water. Incubate the mixture (150 μl) at 37 °C for 30 minutes and stop by adding 6 μl of 10% SDS and 1.5 μl of 0.5 M EDTA. Purify the treated RNA by PCI extraction and precipitation and then resuspend it in nuclease-free H2O in a volume equal to the original input volume of the dT-purified RNA. Quantify the resuspended RNA using a Nanodrop and check its integrity by Tapestation.
[0163] Example 2. This example provides a representative method for transfecting cells with mRNA encoding the nrRT protein and template RNA encoding the GFP reporter gene.
[0164] Prior to transfection, hTERT RPE-1 cells are harvested from 30% - 50% confluent plates using trypsin-EDTA (0.25%) and phenol red (Gibco, 25200056), and seeded into 6-well plates at a density of 500,000 cells / well. Each transfection is performed in duplicate. Dilute 10 μL of Messenger Max (Invitrogen Lipofectamine MessengerMAX, LMRNA003) in 250 μL of Opti-MEM and incubate for 10 minutes at room temperature. Dilute a total of 5 μg of nrRT mRNA and template RNA in a 1:3 molar ratio in 250 μL of Opti-MEM. Then, mix the RNA diluted in Opti-MEM with the diluted and incubated Messenger Max and incubate for 5 minutes at room temperature. Then, add the resulting mixture to two wells seeded with 500,000 cells (250 μL each). Place the transfected cells in an incubator at 37°C with 5% CO2. Image the cells on day 1 and day 2 after transfection to evaluate cell health and transfection efficiency via image analysis. On day 2, wash the cells with 1 mL of 1× PBS and 500 μL of trypsin-EDTA (0.25%) and incubate for 3 minutes in an incubator at 37°C with 5% CO2.
[0165] Example 3 This example provides a representative method for analyzing the ribozyme cleavage efficiency of template RNA containing uridine modifications.
[0166] A minimized version of template RNA containing only the 5’ module array (HDV_gu6_GFP) is generated according to the protocol described in Example 1 by replacing 100% of uridine with various modified uridines (see Table 4). After completion of the DNase treatment step, 200 μl of oligo binding buffer is added to the 100 μl of the RNA sample after DNase treatment, and then this is mixed with 800 μl of ethanol (95 - 100%). Approximately 750 μL of the above mixture is transferred to a Zymo-Spin IC Column (Zymo, Cat#D4060) placed in a Collection Tube and centrifuged. The flow-through is discarded. The remaining sample is transferred to the Zymo-Spin IC Column and centrifuged at 10,000 - 16,000 × g. The flow-through is discarded. 750 μl of DNA wash buffer is added to the column and centrifuged for 1 minute to ensure complete removal of the wash buffer. The column is carefully transferred to a nuclease-free tube. 15 μl of water is added directly to the column matrix and centrifuged. Quantify with Nanodrop. 25 ng of purified RNA / lane is electrophoresed on a 10% Criterion TBE-Urea Polyacrylamide Gel (Bio-Rad, Cat#3450089) at 120 V until bromophenol blue reaches the bottom of the gel. The gel is stained by adding a 1:10,000 dilution of SYBR Gold (ThermoFisher, Cat#S11494) in water and shaking for 10 minutes at room temperature in the dark. Wash the gel with water before taking an image.
[0167] The cleavage efficiency is quantified using the densitometry analysis feature of ImageJ with background subtraction. The results (Figure 4 and Table 2) show that the use of different uridine substitutions led to different efficiencies of ribozyme cleavage. The use of unmodified uridine results in nearly complete cleavage. The use of 5mU or 5moU leads to >80% ribozyme cleavage efficiency. In contrast, the use of 5mC, ΨU, or N1mΨU resulted in very low or undetectable cleavage products. Table 2. Ribozyme cleavage efficiency associated with different uridine modifications [Table 2]
[0168] The corresponding full-length versions of HDV_gu6_GFP template RNAs containing either 5meU or N1mpU modifications are co-transfected into hTERT RPE-1 cells together with the above nrRT (TaGu RT mRNA) using the protocol described in Example 2. GFP image analysis is performed on the second day after transfection. The above template RNA with 5mU modification resulted in low payload incorporation as reflected by a small number of GFP-positive cells and high cytotoxicity (Figure 4). In contrast, the template RNA with N1mpU modification resulted in a considerably larger number of GFP-positive cells and low toxicity (Figure 5).
[0169] A second template RNA, HDV_ac2_GFP (with different uridine modifications), is generated using the protocol described in Example 1 and transfected into hTERT RPE-1 cells as described in Example 2 to evaluate the efficiency of payload expression and the impact of uridine modifications on cellular health. The results (Figure 6) show that both unmodified U and 5mU lead to a small number of GFP-positive cells and high cytotoxicity. The use of 5moU modification resulted in a very small number of GFP-positive cells and similarly low toxicity. Consistent with the HDV_gu6_GFP template RNA, the use of N1mΨU or ΨU resulted in a considerably high percentage of GFP-positive cells without causing significant cytotoxicity.
[0170] Example 4 This example describes junction analysis to evaluate the integration efficiency of the above payload sequence at the target site in the genome and a comparison of the integration efficiency of template RNAs with uridine modifications.
[0171] gDNA extraction and qPCR The transfected cells are washed with PBS, pelleted by centrifugation, and snap-frozen. The cells are lysed with cell lysis buffer (0.1 M EDTA, 0.5% SDS, 10 mM Tris-HCl pH 7.5, 0.2 mg / mL RNaseA) at 56 °C for 10 minutes, followed by lysis at 37 °C for 1 - 3 hours. An equal volume of phenol:chloroform:isoamyl alcohol (25:24:1) is added to the cell lysate, vortexed at maximum speed for 10 seconds, and centrifuged at 21,000 × g for 5 minutes at room temperature. The aqueous layer containing genomic DNA is removed, mixed with an equal volume of 100% isopropanol + 300 mM sodium chloride, and centrifuged at 21,000 × g for 10 minutes to precipitate the genomic DNA. The genomic DNA pellet is washed with 70% ethanol and centrifuged at 21,000 × g for 5 minutes. The genomic DNA pellet is air-dried for 5 - 10 minutes before resuspending in nuclease-free water. Total genomic DNA is quantified using the 1× DNA HS Quantification Assay Kit (Invitrogen Cat #Q33231) according to the manufacturer's instructions. Quantitative PCR is performed using the NEB Luna Universal One-Step qPCR Kit (NEB Cat #M3003). 5 nanograms of gDNA is used as a template for each reaction, and each sample is run technically in duplicate or triplicate. The relevant forward and reverse primers are used at a concentration of 0.5 uM each per reaction. The cycling conditions are as follows: 1 cycle of 95 °C for 5 minutes, 40 cycles of (95 °C for 15 seconds, 60 °C for 30 seconds), followed by a melting curve analysis step of heating from 65 °C to 95 °C. Quantitative analysis is performed as described in the following section "qPCR Data Analysis".
[0172] Cell Direct qPCR The transfected cells are washed with PBS and frozen at -80 °C in tissue culture plates. The cells are lysed directly in cell lysis buffer (5 mM EDTA, 0.5% SDS, 10 mM Tris-HCl pH 7.5, 40 μg / mL proteinase K) at 37 °C for 10 minutes. The cell lysate is diluted 1:1 with nuclease-free water and then heated at 37 °C for 5 minutes, followed by heating at 95 °C for 5 minutes. The cell lysate is further diluted 1:10 in nuclease-free water. Quantitative PCR is performed using the NEB Luna Universal One-Step qPCR Kit (NEB Cat #M3003). 5 microliters of the diluted cell lysate is used as a template for each reaction, and each sample is run technically in duplicate or triplicate. The relevant forward and reverse primers (see Table 3) are used at a concentration of 0.5 μM each per reaction. The cycling conditions are as follows: 1 cycle of 95 °C for 5 minutes, 40 cycles of (95 °C for 15 seconds, 60 °C for 30 seconds), followed by a melting curve analysis step of heating from 65 °C to 95 °C. Quantitative analysis is performed as described in the following section "qPCR Data Analysis".
[0173] qPCR Data Analysis Quantification is performed by setting a uniform fluorescence signal across all primer sets and samples and determining at which cycle number the fluorescence signal exceeds the threshold for each well (referred to as the Cq value). Quantification of the 3' junction from each sample is normalized to the quantification value of Tbp1 (a single-copy gene) from the same sample by subtracting the average Cq value of Tbp1 from the average Cq value of the relevant junction to obtain the ΔCq value. The data is presented as "fold over Tbp1" by converting ΔCq: 2^(-ΔCq). Table 3. Primers used for junction analysis
Table 3
[0174] Three different template RNAs containing five different uracil nucleotides (U, 5meU, 5moU, N1mpU, and pU) are transfected into hTERT RPE-1 cells together with TaGu RT mRNA according to the protocol described in Example 3. The cells are harvested and genomic DNA is extracted from each sample according to the protocol described above. The qPCR-based junction analysis is performed according to the process described above. The results are summarized in Table 4. For each of the template RNAs, the use of N1mΨU or ΨU resulted in at least 2- to 5-fold higher 3’ insertion efficiency compared to those with unmodified U. The other two modifications, 5meU or 5moU, resulted in similar 3’ insertion efficiencies. Table 4. Comparison of 3’ integration efficiencies of template RNAs with different U modifications
Table 4
[0175] Example 5 This example describes that the functional 5’ ribozyme in the template RNA is not required for the integration of the payload sequence at the target site in the genome.
[0176] Template RNAs containing various 5’ module sequences (see Table 5) and encoding the GFP reporter gene as the payload are generated using the in vitro transcription (IVT) protocol described in Example 1. Uridine is replaced with N1mΨU in the IVT RNA. The resulting RNA is co-transfected into hTERT RPE-1 cells together with TaGu RT mRNA as described in Example 2. The GFP image analysis of the transfected cells is summarized in Table 5. The results show that with the N1mΨU modification, the integration of the GFP gene at the target site in the genome does not require the 5’ module of the template RNA at all, because it contains an active ribozyme, a complete ribozyme structure, or any ribozyme sequence. Table 5. Summary of the effect of the 5’ module of template RNA on gene insertion efficiency [Table 5]
[0177] Example 6 This example describes that the molar ratio of the above nrRT mRNA to the above template RNA and / or the amount of total RNA delivered to the target cells affects the insertion efficiency.
[0178] Before transfection, hTERT RPE-1 cells were harvested from 30% - 50% confluent plates using trypsin-EDTA (0.25%) and phenol red (Gibco, 25200056), and placed in an incubator at 37°C with 5% CO2 until a dilution series was performed (within 30 minutes). The total volume of Messenger Max (Invitrogen Lipofectamine MessengerMAX, LMRNA003) was diluted into 140 uL (number of wells × 30 uL × number of plates) of Opti-MEM and incubated for 10 minutes. The total volume of Messenger Max required was based on a volume-to-weight ratio of 2 uL Messenger Max to 1 ug RNA. TaGu-RT mRNA and HDV_gu5b-Luciferase-n1mpU RNA were mixed at a specific molar ratio (see Table 6), and then diluted in 140 uL of Opti-MEM. The diluted RNA in Opti-MEM was mixed with the diluted Messenger Max and incubated at room temperature for 5 minutes. Serial dilutions were performed in a 96-well plate across the columns of the plate starting from the highest dose (1.25 ug) to the lowest dose (0.01 ug) for each molar ratio. Then, 20,000 cells were added per well. Luciferase assays were performed on Day 1 and Day 2 using the Bright-Glo Luciferase Assay System (Promega, Cat#E2620). For luminescence quantification, Agilent's Cytation5 with Gen5 software was used with the following settings: endpoint / dynamic read type with a luminescence fiber, integration time of 1 second, gain of 135, and a read height of 4.50 mm. Mixing on the platform was achieved by clicking "Shake", and for the shaking mode, "Linear" was selected with a duration of "0:04". The intensity of the luminescence signal reflects the expression level of the luciferase protein, which is the result of the integration of the luciferase gene encoded by the template RNA in the genomic site described above. The results are shown in Table 6.The highest luminescence signal was observed when the molar ratio of nrRT to template RNA was 1:6 and the total RNA dosage was 0.08 μg / well. Table 6. The effect of the molar ratio of nrRT to template RNA and total dosage affects payload expression. [Table 6] Table 7. Representative diseases and conditions that can be treated by the methods of the present disclosure [Table 7-1] [Table 7-2]
[0179] It is understood that the examples and embodiments described herein are for illustrative purposes only, and various modifications or changes suggested thereby to those skilled in the art should be included within the spirit and scope of this application and the appended claims. All publications, sequence accession numbers, patents, and patent applications cited herein are hereby incorporated by reference in their entirety for all purposes. Abbreviated sequence listing: Template RNA sequence: HDVRZ-28_gu5b_GFP_GeFo full-length template RNA (SEQ ID NO: 1): [Chemical formula] pp7 sequence (SEQ ID NO: 2): [Chemical formula] HDV_gu5b ribozyme having XbaI at the 3' (SEQ ID NO: 3): [Chemical formula] Polymerase terminator (SEQ ID NO: 4): [Chemical formula] LacI binding site (SEQ ID NO: 5):
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Claims
1. A method for inserting a heterologous polynucleotide at a target site in a eukaryotic genome, the method comprising: transfecting a eukaryotic cell with a) RNA encoding a non-LTR retrotransposon reverse transcriptase protein (nrRT) comprising a reverse transcriptase domain and an endonuclease domain; and b) a template RNA wherein the template RNA comprises a promoter, a payload sequence, a polyA sequence, and an nrRT binding sequence, wherein the template RNA comprises one or more modified uridine (U) nucleosides selected from the group consisting of N1-methyl-pseudouridine (N1mΨU), pseudouridine (ΨU), 5-methyluridine (5meU), 5-methoxyuridine (5moU), and mixtures thereof, or the template RNA comprises a mixture comprising unmodified uridine and one or more modified uridines selected from the group consisting of N1mΨU, ΨU, 5meU, and 5moU; wherein the template RNA comprising modified uridine is not cleavable by ribozymes, wherein the nrRT is expressed in the cell and catalyzes the insertion of a double-stranded heterologous polynucleotide comprising the payload sequence at the target site in the eukaryotic genome, method.
2. The method of claim 1, wherein the template RNA comprising modified U increases the insertion efficiency of the payload sequence into the eukaryotic genome as compared to a template RNA comprising unmodified U.
3. The method of claim 2, wherein the template RNA further comprises a 5' ribozyme sequence selected from an active ribozyme, a partially active ribozyme, a ribozyme having reduced catalytic activity, or a catalytically inactive ribozyme.
4. The method of claim 3, wherein the 5' ribozyme is selected from an HDV ribozyme, a TriCasA ribozyme, or a natural cognate ribozyme, semi-cognate ribozyme, or variant thereof.
5. The method according to claim 3, wherein the 5' ribozyme sequence is a sequence selected from any one of SEQ ID NO: 3 or SEQ ID NOs: 13 to 22 (without the pp7 binding sequence), or a sequence having a sequence identity greater than or equal to 60% (e.g., greater than or equal to 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100%) with a sequence selected from any one of SEQ ID NO: 3 or SEQ ID NOs: 13 to 22 (without the pp7 binding sequence). **Claim 6** The method according to claim 1, wherein the template RNA does not contain a functional 5' ribozyme sequence or does not contain a 5' ribozyme sequence. **Claim 7** The method according to claim 1, wherein the cytotoxicity is reduced when the template RNA contains a modified U. **Claim 8** The method according to claim 1, wherein the template RNA further comprises a 5' sequence that protects the 5' end from degradation. **Claim 9** The method according to claim 1, wherein the template RNA further comprises a 5' sequence that promotes site-specific insertion of the heterologous polynucleotide into the target site in the eukaryotic genome. **Claim 10** The method according to claim 1, wherein the nrRT binding sequence comprises a 3'UTR sequence. **Claim 11** The method according to claim 10, wherein the 3'UTR sequence is isolated from an organism selected from the group consisting of G. aculeatus, D. melanogaster, L. polyphemus, P. pun gitis, N. vitripennis, G. fortis, O. latipes, Z. albicol lis, T. guttata, T. castaneum, T. gutatta, D. simulans, B. mori, and A. vaga. **Claim 12** The method according to claim 11, wherein the 3'UTR comprises a sequence having a sequence identity greater than or equal to 60% (e.g., greater than or equal to 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100%) with a sequence selected from any one of SEQ ID NOs: 26 to 39. **Claim 13** The method according to claim 1, wherein the template RNA further comprises a 3' sequence that promotes site-specific insertion of the heterologous polynucleotide into the eukaryotic genome and / or enhances the efficiency and fidelity of target-primed reverse transcription.
14. The method according to claim 1, wherein the template RNA further comprises one or more of: i) an RNA polymerase terminator; ii) a sequence useful for purification; iii) a sequence encoding a protein useful for enrichment; iv) a Kozak sequence located 5' to the payload sequence; and / or v) a polyA sequence located 3' to the nrRT binding sequence.
15. The template RNA a) a 5' sequence homologous to a DNA sequence located 5' to the target insertion site in the eukaryotic genome; or b) a 3' sequence homologous to a DNA sequence located 3' to the target insertion site in the eukaryotic genome; or c) both (a) and (b), The method according to claim 1, further comprising.
16. The method according to claim 1, wherein the payload sequence encodes: i) a therapeutic protein that replaces or complements a defective gene or protein; or ii) an inhibitor of another protein.
17. The method according to claim 16, wherein the therapeutic protein is selected from the group consisting of factor VIII, factor IX, and phenylalanine hydroxylase (PAH).
18. The method according to claim 16, wherein the inhibitor is a single-chain antibody.
19. The method according to claim 1, wherein the payload sequence encodes a regulatory RNA.
20. The method according to claim 1, wherein the payload sequence encodes a protein selected from the genes in Table 7.
21. The method according to claim 1, wherein modulating: i) the molar ratio of the nrRT mRNA to the template RNA; and / or ii) the amount of total RNA delivered to the target cells increases the insertion efficiency.
22. The method according to claim 1, wherein the template RNA lacks a 5' phosphate.
23. The RNA encoding the nrRT comprises one or more modified uridine (U) nucleosides selected from the group consisting of N1-methyl-pseudouridine (N1mΨU), pseudouridine (ΨU), 5-methyluridine (5meU), 5-methoxyuridine (5moU), and mixtures thereof, or unmodified U, and a mixture of modified U selected from the group consisting of N1mΨU, ΨU, 5meU, and 5moU, the method of claim 1. **Claim 24** The method of claim 1, wherein the eukaryotic cell is transfected in vitro. **Claim 25** The method of claim 1, wherein the eukaryotic cell is transfected in vivo. **Claim 26** The method of claim 1, wherein the eukaryotic cell is a mammalian cell. **Claim 27** The method of claim 1, wherein the eukaryotic cell is a human cell. **Claim 28** The method of claim 27, wherein the human cell is removed from a human subject, transfected with the RNAs of (a) and (b) to insert the heterologous polynucleotide into the human cell genome, and administered to the human subject. **Claim 29** The method according to any one of claims 24 to 28, wherein the cell is transfected with an LNP formulation, a lipofection reagent, or by electroporation. **Claim 30** a composition comprising: (a) an RNA encoding a non-LTR retrotransposon reverse transcriptase protein (nrRT) comprising a reverse transcriptase domain and an endonuclease domain; and (b) a template RNA, wherein the template RNA comprises a promoter, a payload sequence, a poly A sequence, and an nrRT binding sequence, wherein the template RNA comprises one or more modified uridine (U) nucleosides selected from the group consisting of N1-methyl-pseudouridine (N1mΨU), pseudouridine (ΨU), 5-methyluridine (5meU), 5-methoxyuridine (5moU), and mixtures thereof, or the template RNA comprises unmodified uridine, and a mixture comprising one or more modified uridines selected from the group consisting of N1mΨU, ΨU, 5meU, and 5moU; The template RNA containing the modified uridine herein is not cleavable by ribozymes, Composition. **Claim 31** The composition according to claim 30, wherein the template RNA further comprises a 5' ribozyme sequence selected from an active ribozyme, a partially active ribozyme, a ribozyme having reduced catalytic activity, or a catalytically inactive ribozyme. **Claim 32** The composition according to claim 31, wherein the 5' ribozyme is selected from an HDV ribozyme, a TriCasA ribozyme, or a natural cognate ribozyme, a semi-cognate ribozyme, or variants thereof. **Claim 33** The composition according to claim 31, wherein the 5' ribozyme sequence is a sequence selected from any one of SEQ ID NO: 3 or SEQ ID NOs: 13-22 (without the pp7 binding sequence), or a sequence having a sequence identity greater than or equal to 60% (e.g., greater than or equal to 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity) with a sequence selected from any one of SEQ ID NO: 3 or SEQ ID NOs: 13-22 (without the pp7 binding sequence). **Claim 34** The composition according to claim 30, wherein the template RNA does not contain a functional 5' ribozyme sequence or does not contain a 5' ribozyme sequence. **Claim 35** The composition according to claim 30, wherein the template RNA further comprises a 5' sequence that protects the 5' end from degradation. **Claim 36** The composition according to claim 30, wherein the template RNA further comprises a 5' sequence that promotes site-specific insertion of the heterologous polynucleotide into a target site in the eukaryotic genome. **Claim 37** The composition according to claim 30, wherein the nrRT binding sequence comprises a 3'UTR sequence. **Claim 38** The composition according to claim 38, wherein the 3'UTR sequence is isolated from an organism selected from the group consisting of G. aculeatus, D. melanogaster, L. polyphemus, P. pungitis, N. vitripennis, G. fortis, O. latipes, Z. albicolli, T. guttata, T. castaneum, T. gutatta, D. simulans, B. mori, and A. vaga. **Claim 39** The 3'UTR of the composition according to claim 38 comprises a sequence selected from any one of SEQ ID NOs: 26 to 39 and having a sequence identity greater than or equal to 60% (e.g., greater than or equal to 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity).
40. The composition according to claim 30, wherein the template RNA further comprises a 3' sequence that promotes site-specific insertion of the heterologous polynucleotide into the eukaryotic genome and / or enhances the efficiency and fidelity of target-primed reverse transcription.
41. The composition according to claim 30, wherein the template RNA further comprises one or more of: i) an RNA polymerase terminator; ii) a sequence useful for purification; iii) a sequence encoding a protein useful for enrichment; iv) a Kozak sequence located 5' to the payload sequence; and / or v) a polyA sequence located 3' to the nrRT binding sequence.
42. The template RNA is a) a 5' sequence homologous to a DNA sequence located 5' to the target insertion site in the eukaryotic genome; or b) a 3' sequence homologous to a DNA sequence located 3' to the target insertion site in the eukaryotic genome; or c) both a) and b), and the composition according to claim 30 further comprises the same.
43. The composition according to claim 30, wherein the payload sequence encodes: i) a therapeutic protein that replaces or complements a defective gene or protein; or ii) an inhibitor of another protein.
44. The composition according to claim 43, wherein the therapeutic protein is selected from the group consisting of Factor VIII, Factor IX, and phenylalanine hydroxylase (PAH).
45. The composition according to claim 43, wherein the inhibitor is a single-chain antibody.
46. The composition according to claim 30, wherein the payload sequence encodes a regulatory RNA.
47. The composition according to claim 30, wherein the payload sequence encodes a protein selected from the genes in Table 7.
48. The composition according to claim 30, wherein the template RNA lacks a 5' phosphate.
49. The RNA encoding the nrRT comprises one or more modified uridine (U) nucleosides selected from the group consisting of N1-methyl-pseudouridine (N1mΨU), pseudouridine (ΨU), 5-methyluridine (5meU), 5-methoxyuridine (5moU), and mixtures thereof, or the composition according to claim 30, comprising unmodified U, as well as a mixture of modified U selected from the group consisting of N1mΨU, ΨU, 5meU, and 5moU.
50. A pharmaceutical composition comprising the composition according to any one of claims 30 to 49.
51. The pharmaceutical composition according to claim 50, wherein the composition is formulated in a lipid nanoparticle formulation selected from liposomes or lipid nanoparticles (LNP).
52. The pharmaceutical composition according to claim 50 or 51, further comprising a pharmaceutically acceptable excipient or salt.
53. A method of treating a disease or condition in a subject in need of treatment, the method comprising administering to the subject an effective amount of the pharmaceutical composition according to any one of claims 50 to 52.
54. The disease or condition is selected from the group consisting of sickle cell anemia, severe combined immunodeficiency (ADA-SCID / X-SCID), cystic fibrosis, hemophilia, Duchenne muscular dystrophy, Huntington's disease, Parkinson's disease, hypercholesterolemia, α1-antitrypsin deficiency, chronic granulomatous disease, Fanconi anemia, and Gaucher's disease, according to the method of claim 53.
55. The disease or condition is selected from Table 7, according to the method of claim 53.