Methods and compositions for genomic integration

JP2024518100A5Pending Publication Date: 2025-05-20MYELOID THERAPEUTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023570260
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-11-02
Filing Date
2022-05-11
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

Current methods for delivering large nucleic acid cargoes into cells for therapeutic purposes, such as in cell and gene therapy, face challenges related to safety, efficacy, and stability, with viral delivery mechanisms posing immunogenic risks and affecting cell health.

Method used

A pharmaceutical composition comprising polynucleic acids, including mobile genetic elements like LINE polypeptides, facilitates the integration of exogenous therapeutic polypeptides into the human genome using insertion sequences that are substantially non-immunogenic, enabling stable integration and expression, with methods utilizing endonucleases and reverse transcription for precise genome targeting.

Benefits of technology

The solution provides safe and efficient integration and long-term expression of therapeutic polypeptides in human cells, including immune cells, without inducing an immune response, addressing the limitations of viral delivery and ensuring stable genetic manipulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Methods and compositions are disclosed for modulating a target genome and the stable integration of a transgene of interest into the genome of a cell.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] cross reference

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 187,117, filed May 11, 2021, U.S. Provisional Application No. 63 / 254,791, filed October 12, 2021, and U.S. Provisional Application No. 63 / 274,907, filed November 2, 2021, each of which is incorporated by reference in its entirety into this specification. [Background technology]

[0002]

[0002] Cell therapy is a rapidly developing field addressing difficult-to-treat diseases, such as cancer, persistent infections, and certain diseases that resist other forms of treatment. Cell therapy often utilizes cells that are modified ex vivo and administered to an organism to correct defects within the body. An effective and reliable system for manipulating a cell's genome is crucial to ensure that the modified cells function optimally with long-term efficacy when administered to an organism. Similarly, a reliable mechanism for genetic manipulation forms the basis for successful gene therapy. However, serious deficiencies exist in methods for safely and effectively delivering nucleic acid cargoes (e.g., large cargoes) for therapeutic purposes. Viral delivery mechanisms are frequently used to deliver large nucleic acid cargoes to cells, but they are associated with safety concerns and cannot be used to express the cargo in some cell types. In addition, subjecting cells to repeated genetic manipulation may affect the health of the cells, induce cell cycle alterations, and render the cells unsuitable for therapeutic use. Advances in this area are constantly being sought for the effective delivery and stabilization of exogenous transgene material for therapeutic purposes. Summary of the Invention [Means for solving the problem]

[0003]

[0003] Provided herein is a pharmaceutical composition comprising a therapeutically effective amount of one or more polynucleic acids, or at least one vector encoding one or more polynucleic acids, wherein the one or more polynucleic acids comprise: a mobile genetic element comprising a sequence encoding a polypeptide, and an insert sequence, the insert sequence comprising a sequence that is the reverse complement of a sequence encoding an exogenous therapeutic polypeptide, wherein the polypeptide encoded by the sequence of the mobile genetic element promotes integration of the insert sequence into the genome of a cell; and the pharmaceutical composition is substantially non-immunogenic in human subjects.

[0004]

[0004] In some embodiments, the polypeptide encoded by the sequence of the mobile genetic element comprises one or more long interspersed nuclear element (LINE) polypeptides, wherein the one or more LINE polypeptides comprise: human ORF1p or a functional fragment thereof, and human ORF2p or a functional fragment thereof.

[0005]

[0005] In some embodiments, the insertion sequence stably integrates and / or retrotransposes into the genome of the human cell.

[0006] In some embodiments, the human cell is an immune cell selected from the group consisting of a T cell, a B cell, a myeloid cell, a monocyte, a macrophage, and a dendritic cell.

[0007] In some embodiments, the insertion sequence is integrated into the genome by (i) cleavage of the DNA strand at the target site by an endonuclease encoded by one or more polynucleic acids, (ii) via target-primed reverse transcription (TPRT), or (iii) via reverse splicing of the insertion sequence into the DNA target site in the genome. In some embodiments, the insertion sequence is integrated into the genome at a polyT site using the specificity of the endonuclease domain of human ORF2p. In some embodiments, the polyT site comprises the sequence TTTTTA. In some embodiments, the one or more polynucleic acids comprise homology arms complementary to the target site in the genome. In some embodiments, the insertion sequence: (a) is integrated into the genome at a locus that is not a ribosomal RNA locus; (b) is integrated into a gene or a regulatory region of a gene in the genome, thereby disrupting the gene or downregulating gene expression; (c) is integrated into a gene or a regulatory region of a gene in the genome, thereby upregulating gene expression; or (d) is integrated into the genome and replaces a gene in the genome. In some embodiments, the pharmaceutical composition further comprises (i) one or more siRNAs and / or (ii) an RNA guide sequence or a polynucleic acid encoding the RNA guide sequence, wherein the RNA guide sequence targets a DNA target site in the genome and the insertion sequence is integrated into the genome at the DNA target site in the genome. In some embodiments, the one or more polynucleic acids have a total length of 3 kb to 20 kb. In some embodiments, the one or more polynucleic acids comprise one or more polyribonucleic acids, one or more RNAs, or one or more mRNAs. In some embodiments, the exogenous therapeutic polypeptide is selected from the group consisting of a ligand, an antibody, a receptor, an enzyme, a transport protein, a structural protein, a hormone, a contractile protein, a storage protein, and a transcription factor. In some embodiments, the exogenous therapeutic polypeptide is a receptor selected from the group consisting of a chimeric antigen receptor (CAR) and a T cell receptor (TCR).In some embodiments, the one or more polynucleic acids comprise a first expression cassette comprising a promoter sequence, a 5'UTR sequence, a 3'UTR sequence and a polyA sequence: the promoter sequence is upstream of the 5'UTR sequence, the 5'UTR sequence is upstream of the sequence of the mobile genetic element encoding the polypeptide, and the 3'UTR sequence is downstream of the insertion sequence; the 3'UTR is upstream of the polyA sequence; and the 5'UTR sequence, 3'UTR sequence or polyA sequence comprises a binding site for human ORF2p or a functional fragment thereof. In some embodiments, the insert sequence comprises a second expression cassette comprising a sequence that is the reverse complement of the promoter sequence, a sequence that is the reverse complement of the 5'UTR sequence, a sequence that is the reverse complement of the 3'UTR sequence, and a sequence that is the reverse complement of the polyA sequence: (i) the sequence that is the reverse complement of the promoter sequence is downstream of the sequence that is the reverse complement of the 5'UTR sequence, (ii) the sequence that is the reverse complement of the 5'UTR sequence is downstream of the sequence that is the reverse complement of the sequence encoding the exogenous therapeutic polypeptide, (iii) the sequence that is the reverse complement of the 3'UTR sequence is upstream of the sequence that is the reverse complement of the sequence encoding the exogenous therapeutic polypeptide, and (iv) the sequence that is the reverse complement of the polyA sequence is upstream of the sequence that is the reverse complement of the 3'UTR sequence and downstream of the sequence of a mobile genetic gene encoding the polypeptide. In some embodiments, the promoter sequence of the first expression cassette is different from the promoter sequence of the second expression cassette. In some embodiments, the one or more LINE polypeptides comprise a first LINE polypeptide comprising human ORF1p or a functional fragment thereof and a second LINE polypeptide comprising human ORF2p or a functional fragment thereof, wherein the first LINE polypeptide and the second LINE polypeptide are translated from different open reading frames (ORFs). In some embodiments, the one or more polynucleic acids comprise a first polynucleic acid molecule encoding human ORF1p or a functional fragment thereof and a second polynucleic acid molecule encoding human ORF2p or a functional fragment thereof.In some embodiments, the one or more polynucleic acids comprise a 5'UTR sequence and a 3'UTR sequence, wherein the 5'UTR comprises a sequence having at least 80% sequence identity to the 5'UTR from LINE-1 or ACUCCUCCCCAUCCUCUCCCUCUGUCCCUCUGUCCCUCUGACCCUGCACUGUCCCAGCACC; and / or the 3'UTR comprises a sequence having at least 80% sequence identity to the 3'UTR from LINE-1 or CAGGACACAGCCUUGGAUCAGGACAGAGACUUGGGGGCCAUCCUGCCCCUCCAACCCGACAUGUGUACCUCAGCUUUUUCCCUCACUUGCAUCAAUAAAGCUUCUGUGUUUGGAACAG. In some embodiments, the sequence encoding the exogenous therapeutic polypeptide does not contain introns. In some embodiments, the polypeptide encoded by the sequence of the mobile genetic element comprises a C-terminal nuclear localization signal (NLS), an N-terminal NLS, or both. In some embodiments, the sequence encoding the foreign polypeptide is not in frame with the sequence encoding ORF1p or a functional fragment thereof, and / or is not in frame with the sequence encoding ORF2p or a functional fragment thereof. In some embodiments, one or more polynucleic acids comprise a sequence encoding a nuclease domain, a nuclease domain not derived from ORF2p, a megaTAL nuclease domain, a TALEN domain, a Cas9 domain, a Cas6 domain, a Cas7 domain, a Cas8 domain, a zinc finger binding domain derived from an R2 retroelement, or a DNA binding domain that binds to a repeat sequence. In some embodiments, one or more polynucleic acids comprise a sequence encoding a nuclease domain, wherein the nuclease domain does not have nuclease activity or comprises a mutation that reduces the activity of the nuclease domain compared to a nuclease domain without the mutation. In some embodiments, ORF2p or a functional fragment thereof lacks endonuclease activity or comprises a mutation selected from the group consisting of S228P and Y1180A, and / or ORF1p or a functional fragment comprises a K3R mutation.In some embodiments, the inserted sequence comprises a sequence that is the reverse complement of a sequence encoding two or more foreign therapeutic polypeptides. In some embodiments, the one or more polynucleic acids comprise one or more polyribonucleic acids, and the foreign therapeutic polypeptides are receptors selected from the group consisting of chimeric antigen receptors (CARs) and T cell receptors (TCRs), and the pharmaceutical composition is formulated for systemic administration to a human subject. In some embodiments, the one or more polynucleic acids are (i) formulated into nanoparticles selected from the group consisting of lipid nanoparticles and polymeric nanoparticles; and / or (ii) comprise one or more polynucleic acids selected from the group consisting of glycosylated RNA, circular RNA, and self-replicating RNA.

[0008]

[0008] Provided herein are (i) a method for treating a disease or condition in a human subject in need thereof, comprising administering to the human subject a pharmaceutical composition described herein; or (ii) a method for ex vivo modifying a population of human cells, comprising contacting a population of human cells ex vivo with the composition, thereby forming an ex vivo modified population of human cells, wherein the composition comprises one or more polynucleic acids or at least one vector encoding one or more polynucleic acids, wherein the one or more polynucleic acids comprise: a mobile genetic element comprising a sequence encoding a polypeptide; and an insertion sequence that is the reverse complement of a sequence encoding an exogenous therapeutic polypeptide, wherein the ex vivo modified population of human cells is substantially non-immunogenic to the human subject. In some embodiments, the one or more polynucleic acids further comprise (i) a sequence encoding an integrase or a fragment thereof for site-specific integration of the insertion sequence into the genome, and (ii) an integrase genomic landing site sequence operable by the integrase, wherein the genomic landing sequence is more than four contiguous nucleotides in length. In some embodiments, ORF2 and the integrase are on separate polynucleotides. In some embodiments, ORF2 and the integrase are on a single polynucleotide. In some embodiments, the integrase is not integrated into the genome of the cell. In some embodiments, the integrase is a mutated or truncated recombination protein. In some embodiments, the integrase genomic landing sequence operable by the integrase is more than 20 nucleotides in length or more than 30 nucleotides in length. In some embodiments, the insertion sequence comprises an attachment site operable by the integrase. In some embodiments, the integrase genomic landing site is inserted into the genome using a guide RNA and a Cas system. In some embodiments, the guide RNA, the Cas system, and the genomic landing sequence are on a separate polynucleotide from the polynucleotide comprising the sequence encoding the LINE-1 ORF and the insertion sequence.In some embodiments, the one or more ORF polynucleotide sequences comprise mutations.A method for site-specific integration of a heterologous genomic insert sequence into the genome of a mammalian cell, the method comprising: (i) introducing into the cell (a) a polynucleotide comprising a sequence encoding one or more human retrotransposon elements associated with the heterologous insert sequence, and (b) a polynucleotide comprising a sequence encoding a guide RNA, an RNA-guided integrase or a fragment thereof, and a landing sequence operable by the integrase; (ii) verifying integration of the heterologous insert sequence into the site of the genome.

[0009]

[0009] Provided herein are methods for site-specific integration of a heterologous genomic insert using a LINE retrotransposon system, wherein the LINE retrotransposon system is modified to incorporate a fragment of an integrase protein capable of recognizing a genomic landing sequence greater than 10 contiguous nucleotides in length, and the LINE retrotransposon system integrates the heterologous genomic insert into the genomic landing sequence recognized by the fragment of the integrase protein. In some embodiments, the method further includes integrating a genomic landing sequence greater than four contiguous nucleotides in length into the genome. In some embodiments, the integrating step of the genomic landing sequence into the genome is performed by an RNA-guided CRISPR-Cas system. In some embodiments, the RNA-guided CRISPR-Cas system has an editing function capable of integrating a sequence greater than four contiguous nucleotides in length into a specific genomic site. In some embodiments, the RNA-guided CRISPR-Cas system integrates an ORF-mRNA binding sequence at a specific location within the genome that shares sequence homology with the sequence of the guide RNA. In some embodiments, the insert is about 10 kilobases or greater than 10 kilobases. In some embodiments, the polynucleotide is mRNA.

[0010]

[0010] Provided herein are methods for stably integrating an insert sequence into the genomic DNA of a target cell, the methods comprising: contacting the target cell with a composition, the composition comprising a polynucleic acid, the polynucleic acid comprising: an insert sequence, the insert sequence comprising a sequence that is the reverse complement of a sequence encoding a foreign polypeptide, and a mobile genetic element comprising a sequence encoding the polypeptide, wherein the polypeptide encoded by the sequence of the mobile genetic element promotes integration of the insert sequence into the genomic DNA; stably integrating the insert sequence into the genomic DNA of the target cell; and expressing the foreign polypeptide in the target cell, the target cell being a human hepatocyte. In some embodiments, the human hepatocytes are primary cells. In some embodiments, the human hepatocytes are derived from a cultured hepatocyte cell line. In some embodiments, the incorporating step comprises electroporation under conditions optimal for human hepatocytes. In some embodiments, the method further comprises culturing the human hepatocytes in vitro for about 2 hours, about 3 hours, about 4 hours, about 5 hours, about 6 hours, about 8 hours, about 10 hours, or about 24 hours after the incorporating step. In some embodiments, the method further comprises introducing human hepatocytes expressing the foreign polypeptide into a human subject in need thereof, in some embodiments, at least 2% of the human hepatocytes express the foreign polypeptide 10 days after incorporation.

[0011]

[0011] Provided herein are methods for stably integrating an insertion sequence into the genomic DNA of a target cell, the methods comprising: contacting the target cell with a composition, the composition comprising a polynucleic acid, the polynucleic acid comprising: an insertion sequence, the insertion sequence comprising a sequence that is the reverse complement of a sequence encoding a foreign polypeptide, and a mobile genetic element comprising a sequence encoding the polypeptide, wherein the polypeptide encoded by the sequence of the mobile genetic element promotes integration of the insertion sequence into the genomic DNA; stably integrating the insertion sequence into the genomic DNA of the target cell; and expressing the foreign polypeptide in the target cell, the target cell being a human cardiomyocyte. In some embodiments, the human cardiomyocytes are primary cells. In some embodiments, the human cardiomyocytes are derived from a cultured cardiomyocyte cell line. In some embodiments, the incorporating step comprises electroporation under conditions optimal for human cardiomyocytes. In some embodiments, the method further comprises culturing the cardiomyocytes in vitro for about 2 hours, about 3 hours, about 4 hours, about 5 hours, about 6 hours, about 8 hours, about 10 hours, or up to 24 hours after the incorporating step. In some embodiments, the method further comprises introducing human cardiomyocytes expressing the exogenous polypeptide into a human subject in need thereof, in some embodiments, at least 2% of the human cardiomyocytes express the exogenous polypeptide 10 days after incorporation.

[0012]

[0012] Provided herein are methods for stably integrating an insertion sequence into the genomic DNA of a target cell, the methods comprising: contacting the target cell with a composition, the composition comprising a polynucleic acid, the polynucleic acid comprising: an insertion sequence, the insertion sequence comprising a sequence that is the reverse complement of a sequence encoding a foreign polypeptide, and a mobile genetic element comprising a sequence encoding the polypeptide, wherein the polypeptide encoded by the sequence of the mobile genetic element promotes integration of the insertion sequence into the genomic DNA; stably integrating the insertion sequence into the genomic DNA of the target cell; and expressing the foreign polypeptide in the target cell, wherein the target cell is a human retinal pigment epithelial cell. In some embodiments, the human retinal pigment epithelial cell is a primary cell. In some embodiments, the human retinal pigment epithelial cell is derived from a cultured retinal pigment epithelial cell line. In some embodiments, the incorporating step comprises electroporation under conditions optimal for human retinal pigment epithelial cells. In some embodiments, the method further comprises culturing the human retinal pigment epithelial cells in vitro for about 2 hours, about 3 hours, about 4 hours, about 5 hours, about 6 hours, about 8 hours, about 10 hours, or up to 24 hours after the incorporation. In some embodiments, the method further comprises introducing the human retinal pigment epithelial cells expressing the exogenous polypeptide into a human subject in need thereof. In some embodiments, 10 days after the incorporation, at least 2% of the human retinal pigment epithelial cells express the exogenous polypeptide.

[0013]

[0013] Provided herein are methods for stably integrating an insert sequence into the genomic DNA of a target cell, the methods comprising: contacting the target cell with a composition, the composition comprising a polynucleic acid, the polynucleic acid comprising: an insert sequence, the insert sequence comprising a sequence that is the reverse complement of a sequence encoding a foreign polypeptide, and a mobile genetic element comprising a sequence encoding the polypeptide, wherein the polypeptide encoded by the sequence of the mobile genetic element promotes integration of the insert sequence into the genomic DNA; stably integrating the insert sequence into the genomic DNA of the target cell; and expressing the foreign polypeptide in the target cell, the target cell being a human neuronal cell. In some embodiments, the human neuronal cell is a primary cell. In some embodiments, the human neuronal cell is derived from a cultured neuronal cell line. In some embodiments, the incorporating step comprises electroporation under conditions optimal for human neuronal cells. In some embodiments, the method further comprises culturing the neuronal cell in vitro for about 2 hours, about 3 hours, about 4 hours, about 5 hours, about 6 hours, about 8 hours, about 10 hours, or up to about 24 hours after the incorporating step. In some embodiments, the method further comprises introducing human neural cells expressing the exogenous polypeptide into a human. In some embodiments, at least 2% of the human neural cells express the exogenous polypeptide 10 days after introduction. In some embodiments, the insert sequence is a human insert sequence. In some embodiments, the exogenous polypeptide is an exogenous therapeutic polypeptide. In some embodiments, the exogenous polypeptide is an exogenous human polypeptide. In some embodiments, the polypeptide encoded by the sequence of the mobile genetic element promotes integration of the insert sequence into genomic DNA via target-primed reverse transcription (TPRT). In some embodiments, the polynucleic acid is an mRNA or mRNA molecule. In some embodiments, the mobile genetic element comprises a human LINE-1 retrotransposon element. In some embodiments, the ORF2p is selected from a non-human species. In some embodiments, the ORF2p selected from a non-human species is further modified to enhance retrotransposition and / or translation efficiency.In some embodiments, the cell is an immune cell, a hepatocyte, a cardiomyocyte, a retinal pigment epithelial cell, or a neuron. In some embodiments, ORF2p comprises a nuclear localization sequence (NLS). In some embodiments, ORF2p comprises at least two NLSs, which are the same or different. In some embodiments, the NLS is N-terminal to the sequence encoding ORF1p, ORF2p, or both. In some embodiments, the NLS is C-terminal to the sequence encoding ORF1p, ORF2p, or both. In some embodiments, the NLS is derived from SV40. In some embodiments, the NLS is derived from nucleoplasmin. In some embodiments, a first NLS of the at least two NLSs is derived from SV40, and a second NLS of the at least two NLSs is derived from nucleoplasmin. In some embodiments, the first and second NLSs of the at least two NLSs are derived from SV40. In some embodiments, the first and second NLSs of the at least two NLSs are derived from nucleoplasmin. In some embodiments, each of the at least two NLSs is the same.

[0014] INCORPORATION BY REFERENCE

[0014] All publications, patents, and patent applications mentioned in this specification are incorporated herein by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent that the publications and patents or patent applications incorporated by reference conflict with the disclosure contained herein, it is intended that the present specification supersede and / or take precedence over any such conflicting material.

[0015] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also referred to herein as "Figures"). [Brief explanation of the drawings]

[0016] [Figure 1A]

[0016] Figure 1A illustrates the general mechanism of action of retrotransposons. (I) is a schematic diagram depicting the overall life cycle of an autonomous retrotransposon. (II) LINE-1 retrotransposons contain a LINE-1 element, which encodes two proteins, ORF1p and ORF2p, expressed as mRNA. The bicistronic mRNA is translated into two proteins, and ORF2p binds to the 3' end of its own mRNA via a poly(A) tail when translated by a ribosomal read-through event (III). ORF2p is cleaved at the consensus sequence TAAAA, where the poly(A) at the 3' end of the mRNA hybridizes and stimulates the reverse transcriptase activity of the ORF2 protein. This protein reverse transcribes the mRNA into DNA, resulting in the insertion of the LINE-1 sequence into a new location in the genome (IV). [Figure 1B]

[0017] FIG. 1B is a schematic diagram of an mRNA construct containing a gene payload (left) that can be designed for integration into the genome (right). [Figure 1C]

[0018] Figure 1C illustrates various exemplary designs for integrating mRNA encoding a transgene into a cell's genome, where the boxed GFP is an exemplary transgene. [Figure 1D]

[0019] Figure ID illustrates various exemplary designs for integrating mRNA encoding a transgene into the genome of a cell, where the boxed GFP is an exemplary transgene. [Figure 1E]

[0020] Figure 1E is a schematic diagram of the LINE-1 retrotransposition cycle, showing the mechanism of LINE transposon action and the introduction of transgene cargo into retrotransposon sites. LINE-1 retrotransposons are genomic sequences encoding two proteins, ORF1 and ORF2. These elements are transcribed and translated into proteins that form an RNA-protein complex with LINE-1 mRNA, ORF1 trimer, and ORF2, a reverse transcriptase endonuclease. This complex then transfers back to the nuclease, which cleaves DNA at a 5'-TTTTN-3' motif, creating an RNA-DNA hybrid between the poly(A) tail of the mRNA and the cleaved, excised DNA, priming the ORF2 protein for reverse transcription of the LINE-1 RNA. Reverse transcription of LINE-1 into cDNA results in a new LINE-1 integration event. [Figure 2A]

[0021] Figure 2A illustrates three exemplary designs for expressing an exemplary transgene, GFP, using constructs to stably integrate a sequence encoding GFP. Expected GFP expression levels at 72 hours are shown on the right. [Figure 2B]

[0022] Figure 2B illustrates three exemplary designs for expressing an exemplary transgene, GFP, using constructs to stably integrate sequences encoding RFP, RFP and GFP, or ORF2p and GFP. Expected GFP and RFP expression levels at 72 hours are shown on the right. [Figure 3A]

[0023] Figure 3A is an exemplary illustration of conventional circRNA structure and formation. [Figure 3B]

[0024] FIG. 3B shows two diagrams of an exemplary RL-GAAA tectoRNA motif design. [Figure 3C]

[0025] Figure 3C illustrates an exemplary structure of the Chip-Flow fragment RNA as a platform for testing potential tectoRNAs. [Figure 4A]

[0026] FIG. 4A is an exemplary schematic diagram showing ORF2p binding to the ORF2 polyA region. [Figure 4B]

[0027] FIG. 4B is an exemplary schematic diagram showing how a fusion of ORF2p and the MS2 RNA-binding domain binds to an MS2-binding RNA sequence in the 3′ UTR of the mRNA encoding ORF2, improving specificity. [Figure 4C]

[0028] Figure 4C illustrates an exemplary design of a retrotransposon system for stably integrating a nucleic acid into a cell's genome at a specific site. The top panel shows a design using an ORFp2-MegaTAL DNA binding domain fusion in which the DNA binding and endonuclease activities of ORF2p are mutated to be inactive. The middle panel shows a chimeric ORF2p in which the endonuclease domain is replaced with a highly specific and highly accurate nuclease domain of another protein. The bottom panel shows a fusion of ORF2p with the DNA binding domain of a heterologous protein, such that the fusion protein binds to the ORF2 binding site and additional DNA sequences near the ORF2 site. [Figure 5-1]

[0029] FIG. 5 illustrates exemplary constructs (I) through (X) for integrating mRNA encoding a transgene into the genome of a cell. [Figure 5-2]

[0029] Figure 5 illustrates exemplary constructs (I)-(X) for integrating mRNA encoding a transgene into the genome of a cell. [Figure 5-3]

[0029] Figure 5 illustrates exemplary constructs (I)-(X) for integrating mRNA encoding a transgene into the genome of a cell. [Figure 6A]

[0030] FIG. 6A illustrates an exemplary construct having a sequence encoding ORF1p for integrating mRNA encoding a transgene into the genome of a cell. [Figure 6B]

[0031] FIG. 6B illustrates an exemplary construct for integrating mRNA encoding a transgene into the genome of a cell, without the sequence encoding ORF1p. [Figure 7A]

[0032] FIG. 7A illustrates an exemplary method for improving mRNA half-life by inhibiting degradation by 5′-3′ exonucleases such as XRN1 or 3′-5′ exosomal degradation, introducing a G4 structure or a structure corresponding to a pseudoknot in the 5′ UTR, and / or xrRNA and / or non-A nucleotide residues of a triplex motif in the 3′ UTR. [Figure 7B]

[0033] FIG. 7B is an exemplary schematic diagram of bone marrow cells expressing a transgene encoding a chimeric receptor that binds to cancer cells and induces anti-cancer activity. [Figure 7C]

[0034] Figure 7C is a graph showing expected results regarding increased and prolonged expression of the chimeric receptor by introducing bulk or purified RNA encoding the chimeric receptor that binds to the cancer cells described in Figure 7B. [Figure 8A]

[0035] Figure 8A shows an exemplary plasmid design and the expected cargo nucleic acid sequence-containing LINE-1 mRNA transcript. The plasmid has a LINE-1 sequence (including ORF1 and ORF2 protein coding sequences) and a cargo sequence that is a nucleic acid sequence encoding GFP, where the coding sequence for GFP is separated by an intron. GFP is not expressed until the sequence is integrated into the genome and the intron is spliced ​​out. [Figure 8B]

[0036] Figure 8B is a graph showing exemplary results demonstrating successful integration of the mRNA transcripts encoded by the plasmids shown in Figure 8A and expression of GFP compared to mock-transfected cells (showing the fold increase in mean fluorescence intensity of GFP-positive cells). Mock-transfected cells were transfected with a vector lacking the GFP cargo sequence. [Figure 8C]

[0037] FIG. 8C is a graph showing exemplary flow cytometry results from the results shown in FIG. 8B. [Figure 9A]

[0038] Figure 9A shows an exemplary plasmid design and the expected cargo nucleic acid sequence-containing LINE-1 mRNA transcript. The plasmid contains a LINE-1 sequence (including ORF1 and ORF2 protein coding sequences) and a cargo sequence that is a nucleic acid sequence encoding a recombinant chimeric fusion receptor protein (ATAK receptor) having an extracellular region capable of binding to CD5 and an intracellular region containing an FCR intracellular domain and a PI3 kinase recruitment domain. The coding sequence for the ATAK receptor is separated by an intron. [Figure 9B]

[0039] Figure 9B is a graph showing exemplary results demonstrating successful integration of the mRNA transcript encoded by the plasmid shown in Figure 9A and expression of ATAK compared to mock-transfected cells (the fold increase in mean fluorescence intensity of ATAK-positive cells is shown). Mock-transfected cells were transfected with a vector lacking the ATAK cargo sequence. Expression of the ATAK receptor protein was detected by binding with labeled CD5 antibody. [Figure 9C]

[0040] FIG. 9C is a graph showing exemplary flow cytometry results from the results shown in FIG. 9B. [Figure 10A]

[0041] Figure 10A shows an exemplary plasmid design and the expected cargo nucleic acid sequence-containing LINE-1 mRNA transcript. The plasmid contains the LINE-1 sequence (including ORF1 and ORF2 protein coding sequences) and the cargo sequence, which is a nucleic acid sequence encoding a recombinant chimeric fusion receptor protein (ATAK receptor), followed by a T2A self-cleaving sequence, followed by a disrupted GFP sequence (all in the reverse orientation relative to the LINE-1 sequence). The coding sequence for GFP is separated by an intron. The expected mRNA after reverse transcription and integration of the cargo is shown. [Figure 10B]

[0042] Figure 10B is a graph showing exemplary results demonstrating successful integration of the mRNA transcript encoded by the plasmid shown in Figure 10A and expression of ATAK-T2A-GFP compared to mock-transfected cells (fold change in GFP and ATAK double-positive cells is shown). Mock-transfected cells were transfected with a vector lacking the ATAK cargo sequence. ATAK receptor protein expression was detected by binding with labeled CD5 antibody. [Figure 10C]

[0043] FIG. 10C shows representative flow cytometry data from two separate experimental runs for the expression of both GFP and a CD5-binding agent (ATAK) using the experimental setup shown in FIG. 10A. [Figure 10D]

[0044] FIG. 10D is a graph showing representative flow cytometry data from two separate experimental runs for the expression of both GFP and a CD5-binding agent (ATAK) using the experimental setup shown in FIG. 10A. [Figure 11A]

[0045] Figure 11A shows an exemplary mRNA construct for retrotransposition-based gene delivery. The ORF1 and ORF2 sequences are present in two different mRNA molecules. The ORF2p (ORF2)-encoding mRNA contains the GFP-encoding sequence and is inverted. [Figure 11B]

[0046] Figure 11B is a graph showing exemplary data showing GFP expression when both ORF1-mRNA and ORF2-FLAG-GFPai mRNA were electroporated, normalized to that when only ORF2-FLAG-GFPai mRNA was electroporated (fold increase in mean fluorescence intensity of GFP-positive cells is shown). [Figure 12A]

[0047] Figure 12A is a graph showing exemplary data showing the expression of GFP upon electroporation with various amounts of ORF1-mRNA and ORF2-FLAG-GFPai mRNA (fold increase in mean fluorescence intensity of GFP-positive cells is shown). The fold increase is relative to 1x ORF2-GFPao and 1x ORF1 mRNA. [Figure 12B]

[0048] FIG. 12B is an exemplary fluorescence microscopy image of GFP+ cells after electroporation with the mRNA shown in FIG. 11A. [Figure 13A]

[0049] Figure 13A shows an exemplary mRNA construct for gene delivery in which ORF1 and ORF2 sequences are present in two different mRNA molecules (top panel), and a LINE-1 mRNA transcript containing the ORF1 and ORF2 protein-coding sequences in a single mRNA molecule (bottom panel). The mRNA contains bicistronic ORF1 and ORF2 sequences and a CMV-GFP sequence oriented 3' to 5' in the 3'UTR. After retrotransposition of the delivered ORF2-cmv-GFP antisense (LINE-1 mRNA), cells are expected to express GFP. [Figure 13B]

[0050] FIG. 13B is a graph showing exemplary data demonstrating the expression of GFP upon electroporation of the construct shown in FIG. 13A (fold increase in mean fluorescence intensity of GFP-positive cells is shown). [Figure 14A]

[0051] Figure 14A shows an exemplary experimental design for testing whether multiple electroporations improve retrotransposition efficiency. HEK293T cells were electroporated every 48 h using the Maxcyte system and assessed for GFP-positive cells using flow cytometry after 24-72 h of culture. [Figure 14B]

[0052] Figure 14B is a graph showing exemplary data showing GFP expression at the indicated times when electroporated 1 to 5 times according to Figure 14A (fold increase in mean fluorescence intensity of GFP-positive cells is shown). [Figure 15A]

[0053] Figure 15A shows exemplary constructs for enhancing retrotransposition via mRNA delivery. In one construct, a nuclear localization signal (NLS) sequence is fused to the C-terminus of the ORF2 sequence (ORF2-NLS fusion). In another construct, the minke whale ORF2 sequence was used instead of human ORF2. In another construct, a minimal Alu element sequence (AJL-H33 delta) was inserted into the 3' UTR of the LINE-1 sequence. In another construct, an MS2 hairpin was inserted into the 3' UTR of the LINE-1 sequence, and an MS2 hairpin-binding protein (MCP) sequence was fused to the ORF2 sequence. [Figure 15B]

[0054] FIG. 15B is a graph showing exemplary data demonstrating GFP expression (fold increase in mean fluorescence intensity of GFP-positive cells shown) using the constructs shown in FIG. 15A. [Figure 16A]

[0055] Figure 16A shows exemplary plasmid constructs for gene delivery in which ORF1 and ORF2 sequences are present on two different plasmid molecules (top panel), and a plasmid encoding a LINE-1 mRNA transcript containing the ORF1 and ORF2 protein-coding sequences of a single mRNA molecule with various permutations of the inter-ORF sequence between ORF1 and ORF2 (bottom panel). [Figure 16B]

[0056] FIG. 16B is a graph showing exemplary data demonstrating GFP expression (fold increase in mean fluorescence intensity of GFP-positive cells shown) using the constructs shown in FIG. 16A. [Figure 17A]

[0057] Figure 17A shows an exemplary plasmid construct encoding a LINE-1 mRNA transcript comprising the ORF1 and ORF2 protein coding sequences and a GFP sequence on a single mRNA molecule (top panel), and an exemplary LINE-1 mRNA transcript comprising the ORF1 and ORF2 protein coding sequences and a GFP sequence on a single mRNA molecule. [Figure 17B]

[0058] Figure 17B is a graph showing exemplary data demonstrating the expression of GFP in Jurkat cells using the constructs shown in Figure 17A (fold increase in mean fluorescence intensity of GFP-positive cells is shown). Plasmid constructs were transfected and mRNA constructs were electroporated. [Figure 18A]

[0059] Figure 18A shows an exemplary plasmid design and the expected cargo nucleic acid sequence-containing LINE-1 mRNA transcript. The plasmid has a LINE-1 sequence (including ORF1 and ORF2 protein coding sequences) and a cargo sequence that is a nucleic acid sequence encoding a recombinant chimeric fusion receptor protein (ATAK receptor), followed by a T2A self-cleaving sequence, followed by a disrupted GFP sequence (all in the reverse orientation relative to the LINE-1 sequence). The coding sequence for GFP is separated by an intron. The expected mRNA after reverse transcription and integration of the cargo is shown. [Figure 18B]

[0060] Figure 18B is a graph showing exemplary results demonstrating successful integration of the mRNA transcript encoded by the plasmid shown in Figure 10A and expression of ATAK-T2A-GFP in a myeloid cell line (THP-1) compared to mock-transfected cells (fold change in GFP and ATAK double-positive cells is shown). Data represent expression 6 days after transfection normalized to mock-plasmid-transfected cells, where the mock plasmid does not carry the GFP coding sequence. [Figure 19]

[0061] 19 illustrates an exemplary experimental setup for cell synchronization. A heterogeneous cell population is sorted based on the cell cycle stage before delivery of exogenous nucleic acid. Cell cycle synchronization is expected to result in higher expression and stabilization of the delivered exogenous nucleic acid. If the cells after cell sorting are not homogeneous, the cells can be further incubated with a suitable agent that stops the cell cycle at a certain stage. [Figure 20]

[0062] FIG. 20 illustrates an exemplary method for improving retrotransposon efficiency by inducing DNA double-strand breaks, with or without inhibiting DNA repair pathways, such as by inducing the DNA ligase inhibitor SCR7 or inhibiting host surveillance proteins using miRNAs against HUSH complex TASOR proteins. [Figure 21]

[0063] FIG. 21 is a diagram illustrating an exemplary construct for integrating mRNA encoding a transgene into the genome of a cell. [Figure 22]

[0064] FIG. 22 is a diagram illustrating an exemplary construct for integrating mRNA encoding a transgene into the genome of a cell. [Figure 23]

[0065] FIG. 23 is a diagram illustrating an exemplary construct for integrating mRNA encoding a transgene into the genome of a cell. [Figure 24]

[0066] FIG. 24 is a diagram illustrating an exemplary construct for integrating mRNA encoding a transgene into the genome of a cell. [Figure 25]

[0067] FIG. 25 is a diagram illustrating an exemplary construct for integrating mRNA encoding a transgene into the genome of a cell. [Figure 26]

[0068] FIG. 26 is a diagram illustrating an exemplary construct for integrating mRNA encoding a transgene into the genome of a cell. [Figure 27]

[0069] FIG. 27 is a diagram illustrating an exemplary construct for integrating mRNA encoding a transgene into the genome of a cell. [Figure 28]

[0070] FIG. 28 is a diagram illustrating an exemplary construct for integrating mRNA encoding a transgene into the genome of a cell. [Figure 29]

[0071] Figure 29 shows an exemplary retrotransposon construct (left) containing a 2.4 kb cargo with a general mechanism of action of retrotransposons, and representative data (right) for the expression of a fluorescent GFP marker encoded by the cargo derived from a nucleic acid sequence integrated into the genome of HEK293 cells. The placement of an antisense GFP gene interrupted by an intron in the sense orientation and a promoter sequence in the 3'UTR of LINE-1 results in reconstitution of the GFP cargo and retrotransposition. GFP expression in 293T cells transfected with the constructs shown on the left was measured by flow cytometry (right) and quantified (bottom left). Data were collected 35 days after doxycycline induction of the ORF. [Figure 30]

[0072] Figure 30 shows an exemplary retrotransposon construct (left) containing a 3.0 kb cargo containing a membrane protein (CD5-binding chimeric antigen receptor, CD5-CAR) and representative flow cytometry data (right) for the expression of the CD5-binding agent from the nucleic acid sequence integrated into the genome of HEK293 cells. The % of CD5-binding agent positive (+) cells is indicated within the figure. [Figure 31]

[0073] Figure 31 shows an exemplary retrotransposon construct containing a 3.7 kb cargo comprising membrane proteins (CD5-binding chimeric antigen receptor, CD5-CAR and GFP separated by a self-cleavable T2A element) (top), and representative flow cytometry data demonstrating expression of the CD5-binding agent and GFP (bottom). [Figure 32]

[0074] Figure 32 shows an exemplary retrotransposon construct containing a 3.9 kb cargo comprising membrane proteins (HER2-binding chimeric antigen receptor and GFP separated by a self-cleavable T2A element) (top), and representative flow cytometry data demonstrating expression of the HER2-binding agent and GFP (bottom). [Figure 33A]

[0075] FIG. 33A shows exemplary data for delivery of retrotransposon elements delivered as mRNA. [Figure 33B]

[0076] Figure 33B is a schematic diagram (top panel) showing the trans-mRNA and cis-mRNA designs for delivery of LINE1 mRNA containing GFP cargo. Representative results from electroporation of 293T cells with trans-mRNA containing separate ORF1 and ORF2 mRNAs. 293T cells were electroporated with 100 μg / mL of mRNA containing either ORF2 alone, ORF1 + ORF2 mRNA, or a GFP-encoding mRNA with the same 5' and 3' UTRs as ORF1 mRNA (left panel of data plots). Retrotransposition events result in GFP-positive cells. Cells were assayed for GFP fluorescence by flow cytometry 4 and 10 days after electroporation. Mock-electroporated cells serve as a negative control population for gating. The bar graph on the right shows results from a representative experiment showing titration of the concentrations of trans-mRNA and cis-ORF1 and ORF2-containing mRNA during electroporation. Trans-mRNA is represented by a solid bar, and cis-mRNA is represented by a striped bar. 20X is 2000ug / mL in the electroporation reaction. [Figure 33C]

[0077] Figure 33C shows titration of ORF1 and ORF2-GFPai trans mRNA. Increasing the concentration to 200ug / mL, separately and together during electroporation, increases retrotransposition of the GFP gene cargo. [Figure 33D]

[0078] Figure 33D shows exemplary flow cytometry data plots for the various constructs indicated above, with the top panel showing day 4 and the bottom panel showing day 13. The right panel illustrates light and fluorescence microscopy images of GFP-expressing cells in culture. The number of integrated cargo copies per construct at day 13 is shown in the bottom right. qPCR assays for genomic DNA integration from various transfected LINE-1 plasmids, LINE-1 mRNA (retro-mRNA), and ORF1- and ORF2-GFP mRNA electroporated cells are shown. Two qPCR primer-probe sets were used: one against the housekeeping gene RPS30 and the other against the GFP gene. Plasmid-transfected cells use a plasmid that does not contain SV40 maintenance sequences. Integration per cell was calculated from the copy number per sample via interpolation of plasmid and genomic DNA standard curves and normalized to two copies of RPS30 per 293T cell. Error bars indicate the standard deviation of three technical replicates. [Figure 34]

[0079] FIG. 34 shows exemplary retrotransposon constructs (left) and expression data (right) in the indicated cell lines. [Figure 35-1]

[0080] Figure 35 shows flow cytometry data showing expression of LINE-1 GFP constructs in K562, 293T, and THP1 cells (top panel); and the number of LINE-2-GFP mRNA integrations per cell in K562 and THP-1 cell lines (bottom panel). [Figure 35-2] This is a continuation of Figure 35. [Figure 36]

[0081] Figure 36 shows flow cytometry data showing expression of LINE 1 GFP constructs in primary T cells (left). Integration per cell is shown in the graph on the right. Data were collected 6 days after electroporation. [Figure 37A]

[0082] FIG. 37A is a schematic representation of the activation, culture period, electroporation and GFP expression assay of isolated primary T cells. [Figure 37B]

[0083] Figure 37B shows flow cytometry data showing expression of LINE 1 GFP mRNA constructs in primary T cells at the concentrations shown and before and after freeze-thawing as indicated in the figure. Integration per cell is shown in a bar diagram. GFP expression using retro-mRNA electroporation with GFP cargo. GFP expression was assayed 4 days after electroporation and 15 days after electroporation in culture. Primary T cells were cryopreserved and thawed at this time. qPCR integration assay for GFP integration. Genomic DNA from 20X samples was isolated and assayed for copies of GFP. [Figure 38]

[0084] FIG. 38 shows a summary of retrotransposon integration and expression results across cell types. [Figure 39]

[0085] FIG. 39 is a diagram illustrating various applications of the techniques described herein, including but not limited to, the use of CART cells, NK cells, neurons, and other cells for cell therapy, as well as in vivo applications, including but not limited to, gene therapy, gene editing, transcriptional regulation, and genome modification. [Figure 40]

[0086] Figure 40 shows exemplary flow cytometry data demonstrating sorting and enrichment of GFP+ 293T cells electroporated with 2000 ng / μL LINE1-GFP mRNA. The first panel shows flow cytometry data for mock-electroporated cells in the absence of LINE1-GFP mRNA. The second panel shows flow cytometry data collected 5 days after electroporation for unsorted cells electroporated with LINE1-GFP mRNA. GFP+ cells from the second panel were sorted, and the flow cytometry data is shown in the third panel. GFP+ cells from the third panel were cultured for 9 days after sorting and re-sorted using a 103 or 104 GFP fluorescence intensity gate. The fourth panel shows flow cytometry data for cells collected 4 days after re-sorting and re-sorted using a 103 GFP gate. The fifth panel shows flow cytometry data for cells collected 4 days after re-sorting and re-sorted using the GFP+ 103GFP gate. [Figure 41A]

[0087] FIG. 41A is a graph showing standard curves for GFP (NB2 plasmid) and housekeeping gene (FAU) to assess genomic integration of GFP-encoding nucleic acid per cell using quantitative PCR. [Figure 41B]

[0088] FIG. 41B shows exemplary graphical results showing the interpolation of the standard curve of FIG. 41A for quantification of genomic integration. [Figure 41C]

[0089] Figure 41C shows the number of GFP genes integrated into the genome of 293T cells following LINE1-GFP mRNA electroporation and double sorting as shown in Figure 40. qPCR shows the average number of GFP integrations per cell when gated at 10 3 GFP+ cells and 10 4 GFP+ cells. [Figure 42]

[0090] FIG. 42 shows exemplary flow cytometry data showing GFP+ 293T cells after electroporation with titrated amounts of LINE1-GFP mRNA, shown in ng / μL, in the electroporation solution and cultured for 3 days post-electroporation. [Figure 43]

[0091] Figure 43 shows exemplary flow cytometry data showing GFP+ 293T cells after electroporation with titrated amounts of LINE1-GFP mRNA, shown in ng / μL, in the electroporation solution and cultured for 5 days post-electroporation. [Figure 44]

[0092] Figure 44 shows exemplary flow cytometry data showing GFP+ 293T cells after electroporation with titrated amounts of LINE1-GFP mRNA, shown in ng / μL, in the electroporation solution and cultured for 7 days post-electroporation. [Figure 45]

[0093] Figure 45 shows a graph of the number of GFP integrations per genome by qPCR (top) and a graph of integration kinetics (bottom) from data from Figures 42-44 for 293T cells electroporated with titrated amounts of LINE1-GFP mRNA shown in ng / μL in the electroporation solution, after 3, 5, or 7 days of culture after electroporation from Figures 42-44. [Figure 46]

[0094] Figure 46 shows exemplary flow cytometry data (right) showing GFP+ K562 cells electroporated with titrated amounts of LINE1-GFP mRNA in ng / μL in the electroporation solution and cultured for 6 days after electroporation, and a graph of the number of GFP integrations per genome by qPCR (left). [Figure 47]

[0095] Figure 47 shows exemplary flow cytometry data (top) showing GFP+ human primary monocytes electroporated with the indicated titrated amounts of LINE1-GFP mRNA and cultured for 3 days post-electroporation, and a graph of the number of GFP integrations per genome by qPCR (bottom). [Figure 48]

[0096] Figure 48 shows exemplary flow cytometry data (bottom) showing GFP+ 293T cells electroporated with 2000 ng / μL LINE1-GFP mRNA and 100 ng / μL, 200 ng / μL, or 300 ng / μL siRNA targeting BRCA1 (siBRCA1) and cultured for 4 days after electroporation, as well as a graph of the number of GFP integrations per genome by qPCR (top). [Figure 49]

[0097] Figure 49 shows exemplary flow cytometry data (bottom) showing GFP+ 293T cells electroporated with 2000 ng / μL LINE1-GFP mRNA and 100 ng / μL siRNA targeting RNASEL (siRNASEL), ADAR1 (siADAR1), or ADAR2 (siADAR2) and cultured for 6 days after electroporation, as well as a graph of the number of GFP integrations per genome by qPCR (top). [Figure 50]

[0098] Figure 50 shows exemplary flow cytometry data (bottom) showing GFP+ 293T cells electroporated with 2000 ng / μL LINE1-GFP mRNA and 100 ng / μL siRNA targeting APOBEC3C (siAPOBEC3C) or FAM208A (siFAM208A) and cultured for 6 days after electroporation, as well as a graph of the number of GFP integrations per genome by qPCR (top). [Figure 51]

[0099] Figure 51 shows exemplary flow cytometry data (bottom) showing GFP+ 293T cells electroporated with 1000 ng / μL or 1500 ng / μL LINE1-GFP mRNA and an siRNA cocktail containing 25 ng / μL, 50 ng / μL, or 75 ng / μL of each siRNA targeting RNASEL (siRNASEL), ADAR1 (siADAR1), ADAR2 (siADAR2), and BRCA1 (siBRCA1) and cultured for 6 days after electroporation, as well as a graph of the number of GFP integrations per genome by qPCR (top). [Figure 52]

[0100] Figure 52 shows exemplary flow cytometry data (bottom) showing GFP+ K562 cells electroporated with an siRNA cocktail containing 1000 ng / μL LINE1-GFP mRNA and 25 ng / μL, 50 ng / μL, or 75 ng / μL each of siRNAs targeting RNASEL (siRNASEL), ADAR1 (siADAR1), ADAR2 (siADAR2), and BRCA1 (siBRCA1) and cultured for 5 days after electroporation, as well as a graph of the number of GFP incorporations per cell by qPCR (top). [Figure 53]

[0101] Figure 53 is a schematic diagram showing exemplary locations of the exogenous nuclear localization sequence (NLS) and exemplary ORF1p and ORF2p mutations of an exemplary LINE1-GFP mRNA construct. [Figure 54A]

[0102] Figure 54A is a schematic diagram showing an exemplary LINE1-GFP construct in which an NLS was inserted at the N-terminus of the sequence encoding ORF1. [Figure 54B]

[0103] Figure 54B is a bar graph showing GFP incorporation per cell 4 days after electroporation of the indicated constructs into 293T cells. [Figure 54C]

[0104] FIG. 54C shows an exemplary flow cytometry image showing GFP+ 293T cells 4 days after electroporation of the indicated constructs. [Figure 55A]

[0105] Figure 55A is a schematic diagram showing an exemplary LINE1-GFP construct in which an NLS was inserted at the C-terminus of the ORF1-encoding sequence. [Figure 55B]

[0106] Figure 55B is a bar graph showing GFP incorporation per cell 4 days after electroporation of the indicated constructs into 293T cells. [Figure 55C]

[0107] FIG. 55C shows an exemplary flow cytometry image showing GFP+ 293T cells 4 days after electroporation of the indicated constructs. [Figure 56A]

[0108] Figure 56A is a schematic diagram showing an exemplary LINE1-GFP construct in which an NLS was inserted at the N-terminus of the sequence encoding ORF2. [Figure 56B]

[0109] Figure 56B is a bar graph showing GFP incorporation per cell 4 days after electroporation of the indicated constructs into 293T cells. [Figure 56C]

[0110] FIG. 56C shows an exemplary flow cytometry showing GFP+ 293T cells 4 days after electroporation of the indicated constructs. [Figure 57A]

[0111] Figure 57A is a schematic diagram showing an exemplary LINE1-GFP construct in which an NLS and linker were inserted at the N-terminus of the sequence encoding ORF2. [Figure 57B]

[0112] Figure 57B is a bar graph showing GFP incorporation per cell 5 days after electroporation of the indicated constructs into 293T cells. [Figure 57C]

[0113] FIG. 57C shows an exemplary flow cytometry showing GFP+ 293T cells 5 days after electroporation of the indicated constructs. [Figure 58A]

[0114] Figure 58A is a schematic diagram showing an exemplary LINE1-GFP construct in which an NLS was inserted at the C-terminus of the sequence encoding ORF2. [Figure 58B]

[0115] Figure 58B is a bar graph showing GFP incorporation per cell 5 days after electroporation of the indicated constructs into 293T cells. [Figure 58C]

[0116] FIG. 58C shows an exemplary flow cytometry showing GFP+ 293T cells 5 days after electroporation of the indicated constructs. DETAILED DESCRIPTION OF THE INVENTION

[0017]

[0117] The present invention arises, in part, from the exciting discovery that polynucleotides can be designed and developed to effect the transfer and integration of genetic cargo (e.g., large genetic cargo) into the genome of a cell. In some embodiments, the polynucleotide comprises (i) genetic material for stable expression, and (ii) a self-integrating genomic integration mechanism that allows for stable integration of the genetic material into a cell by safe and effective non-viral means. Furthermore, it is believed that the genetic material can be integrated into a locus other than a ribosomal locus, the genetic material can be integrated site-specifically, and / or the integrated genetic material will be expressed without triggering the cell's natural silencing mechanisms.

[0018]

[0118] Clustered regularly interspaced short palindromic repeats (CRISPR) have revolutionized the field of molecular biology and have been developed into a powerful gene editing system. It utilizes homology-directed repair (HDR) and can be targeted to genomic sites. CRISPR / Cas9 is a naturally occurring RNA-guided endonuclease. While the CRISPR / Cas9 system has demonstrated great promise for site-specific gene editing and other applications, several factors affecting its efficacy must be addressed, particularly when used for in vivo human gene therapy. These factors include target DNA site selection, sgRNA design, off-target cleavage, the incidence / efficiency of HDR versus NHEJ, Cas9 activity, and delivery method. Delivery remains a major barrier to the use of CRISPR for in vivo applications. Zinc finger nucleases (ZFNs) are fusion proteins of a Cys2-His2 zinc finger protein (ZFP) and a nonspecific DNA restriction enzyme derived from the FokI endonuclease. Challenges associated with ZFPs include the design and engineering of ZFPs for high-affinity binding of desired sequences, which is important. Furthermore, site selection is limited because not all sequences are available for ZFP binding. Another significant challenge is off-target cleavage. Transcription activator-like effector nucleases (TALENs) are fusion proteins composed of a TALE and a FokI nuclease. While off-target cleavage remains a concern, TALENs have been shown to be more specific and less cytotoxic than ZFNs in a side-by-side comparison study. However, TALENs are substantially larger, with a cDNA encoding only a TALEN being 3 kb. This makes delivery of a pair of TALENs more challenging than a pair of ZFNs due to limited delivery vehicle cargo size. Furthermore, packaging and delivery of some TALENs into viral vectors can be problematic due to the high level of repetition in the TALEN sequence.A mutant Cas9 system, a fusion protein of inactive dCas9 and a FokI nuclease dimer, improves specificity and reduces off-target cleavage, and the number of potential target sites is smaller due to PAM and other sgRNA design constraints.

[0019]

[0119] The present invention addresses the above-described problems by providing new, effective and efficient compositions containing transposon-based vectors for providing therapy, including gene therapy, to animals and humans. The present invention also provides methods for using these compositions for providing therapy to animals and humans. These transposon-based vectors can be used in the preparation of pharmaceuticals useful for producing desired effects in recipients after administration. Gene therapy includes the introduction of genes, such as foreign genes, into animals using transposon-based vectors. These genes can perform various functions in the recipient, such as encoding the production of nucleic acids, e.g., RNA, or encoding the production of proteins and peptides. The present invention can facilitate the efficient integration of polynucleotide sequences containing a gene of interest, a promoter, an insertion sequence, a polyA, and any regulatory sequences. The present invention is based on the discovery that human LINE-1 elements are retrotransposable in human cells and cells of other animal species and can be manipulated in various ways to achieve efficient delivery and integration of genetic cargo into the genome of cells. Such LINE-1 elements have diverse applications in human and animal genetics, including, but not limited to, applications in the diagnosis and treatment of genetic disorders and in cancer. The LINE-1 elements of the present invention are also useful for treating various phenotypic effects of various diseases. For example, LINE-1 elements can be used to transfer DNA encoding anti-tumor gene products into cancer cells. Other uses of the LINE-1 elements of the present invention will become apparent to those skilled in the art upon reading this specification.

[0020]

[0120] Generally, human LINE-1 elements contain a 5' UTR with an internal promoter, two non-overlapping reading frames (ORF1 and ORF2), a 200-bp 3' UTR, and a 3' polyA tail. LINE-1 retrotransposons can also contain an endonuclease domain at the N-terminus of LINE-1 ORF2. The finding that LINE-1 encodes an endonuclease demonstrates that this element is capable of autonomous retrotransposition. LINE-1 is a modular protein containing non-overlapping functional domains that mediate LINE-1 reverse transcription and integration. In some embodiments, the sequence specificity of the LINE-1 endonuclease itself can be altered, or the LINE-1 endonuclease can be replaced with another site-specific endonuclease.

[0021]

[0121] LINE-1 retrotransposons can be engineered using recombinant DNA techniques to contain and / or be contiguous with other nucleic acid elements that make the retrotransposon suitable for insertion of heterologous or homologous nucleic acid sequences of substantial length (up to 1 kb, or more than 1 kb, e.g., more than 5, 6, 7, 8, 9, or 10 kb) into the genome of a cell. LINE-1 retrotransposons can also be engineered using the same types of techniques to site-specifically insert the nucleic acid sequence of the heterologous or homologous nucleic acid into the genome of a cell (the site into which such DNA is inserted is known). Alternatively, LINE-1 retrotransposons can be engineered to insert the DNA at a random site. Retrotransposons can also be engineered to achieve insertion of a desired DNA sequence into a region of DNA that is not normally transcribed, such that the DNA sequence is expressed in a manner that does not interfere with normal expression of genes in the cell. In some embodiments, integration or retrotransposition is in the trans orientation. In some embodiments, integration or retrotransposition is in the cis orientation.

[0022]

[0122] Because LINE-1 is native to human cells, the constructs, when placed in human cells, should not be rejected as non-self by the immune system. Additionally, the mechanism of LINE-1 retrointegration ensures that only one copy of the gene is integrated at any specific chromosomal location; therefore, copy number control is built into the system. In contrast, conventional plasmid-based gene transfer procedures offer little or no control over copy number and often result in complex arrays of DNA molecules integrated in tandem at the same genomic location.

[0023]

[0123] All terms are intended to be understood in the same manner as would be understood by one of ordinary skill in the art. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.

[0024]

[0124] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.

[0025]

[0125] As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.

[0026]

[0126] In this application, the use of "or" means "and / or" unless stated otherwise. The terms "and / or" and "any combinations thereof," as well as their grammatical equivalents, can be used interchangeably when used herein. These terms can convey that any combination is specifically contemplated. For illustrative purposes only, the following phrases "A, B, and / or C" or "A, B, C, or any combinations thereof" can mean "A only, B only, C only, A and B, B and C, A and C, and A, B and C." The term "or" can be used conjunctively or disjunctively unless the context clearly dictates a reference to disjunctive use.

[0027]

[0127] The term "about" or "approximately" can mean within an acceptable error range of a particular value as determined by one of ordinary skill in the art, and will depend, in part, on how the value is measured or determined, i.e., the limitations of the measurement system. For example, "about" can mean within 1 or more than 1 standard deviation, as is customary in the art. Alternatively, "about" can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Alternatively, particularly with respect to biological systems or processes, the term can mean within 5-fold, more preferably within 2-fold, of a value. When particular values ​​are described in this application and claims, unless otherwise stated, the term "about" should be assumed to mean within an acceptable error range of the particular value.

[0028]

[0128] As used in the specification and claims, the words "comprising" (and any form of comprising, e.g., "comprise" and "comprises"), "having" (and any form of having, e.g., "have" and "has"), "including" (and any form of including, e.g., "includes" and "include"), or "containing" (and any form of containing, e.g., "contains" and "contain") are non-exclusive or open-ended and do not exclude additional, unrecited elements or method steps. It is contemplated that any embodiment discussed herein can be implemented with respect to any method or composition of the disclosure, and vice versa. Furthermore, the compositions of the disclosure can be used to achieve the methods of the disclosure.

[0029]

[0129] References herein to "some embodiments," "an embodiment," "one embodiment," or "other embodiments" mean that the particular feature, structure, or characteristic described in connection with the embodiment is included in at least some embodiments of the present disclosure, but not necessarily in all embodiments. To facilitate understanding of this disclosure, several terms and phrases are defined below.

[0030]

[0130] Various features of the present disclosure may be described in the context of a single embodiment, but may also be described separately or in any suitable combination. Conversely, the present disclosure may, for clarity of explanation, be described herein in the context of separate embodiments, but may also be implemented in a single embodiment.

[0031]

[0131] The present disclosure encompasses, but is not limited to, methods and compositions related to the expression of foreign nucleic acids in cells. In some embodiments, the foreign nucleic acid is configured for stable integration into the genome of a cell, such as a bone marrow cell. In some embodiments, stable integration of the foreign nucleic acid can be at a specific target within the genome. In some embodiments, the foreign nucleic acid comprises one or more coding sequences. In some embodiments, the foreign nucleic acid can comprise one or more coding sequences, including a nucleic acid sequence encoding an immune receptor. In some embodiments, the present disclosure provides methods and compositions for the stable integration of a nucleic acid encoding a transmembrane receptor associated with an immune response function (e.g., a phagocytic receptor or a synthetic chimeric antigen receptor) into human macrophages or dendritic cells or suitable bone marrow or myeloid progenitor cells. Foreign nucleic acid can refer to a nucleic acid that is not native to a cell and is added exogenously, regardless of whether it includes sequences that may already be endogenously present in the cell. Foreign nucleic acid can be a DNA or RNA molecule. Foreign nucleic acid can include sequences encoding a transgene. Foreign nucleic acid can encode a recombinant receptor or a recombinant protein, such as a chimeric antigen receptor (CAR). Foreign nucleic acid can also be referred to as "gene cargo" in the context of delivery of the foreign nucleic acid into a cell. The genetic cargo can be DNA or RNA. Genetic material can generally be delivered into cells ex vivo by several different known techniques, using either chemical (CaCl-mediated transfection), physical (electroporation), or biological (e.g., viral infection or transduction) means.

[0032]

[0132] Compositions and methods are provided herein for the stable, non-viral transfer and integration of genetic material into cells. In one embodiment, the genetic material is a self-integrating polynucleotide. The genetic material can be stably integrated into the genome of the cell. The cell can be a human cell. The method is designed for safe and reliable integration of the genetic material into the genome of the cell.

[0033]

[0133] Provided herein is a pharmaceutical composition comprising a therapeutically effective amount of one or more polynucleic acids, or at least one vector encoding the one or more polynucleic acids, wherein the one or more polynucleic acids comprise: (a) a mobile genetic element comprising a sequence encoding a polypeptide, and (b) an insert sequence, the insert sequence comprising a sequence that is the reverse complement of a sequence encoding an exogenous therapeutic polypeptide, wherein the polypeptide encoded by the sequence of the mobile genetic element promotes integration of the insert sequence into the genome of a cell; and the pharmaceutical composition is substantially non-immunogenic in a human subject.

[0034]

[0134] In some embodiments, the polypeptide encoded by the sequence of the mobile genetic element comprises one or more long interspersed nuclear element (LINE) polypeptides, wherein the one or more LINE polypeptides comprise: (i) human ORF1p or a functional fragment thereof, and (ii) human ORF2p or a functional fragment thereof.

[0035]

[0135] In some embodiments, the insertion sequence stably integrates and / or retrotransposes into the genome of the human cell.

[0036]

[0136] In some embodiments, the human cell is an immune cell selected from the group consisting of a T cell, a B cell, a myeloid cell, a monocyte, a macrophage, and a dendritic cell.

[0037]

[0137] In some embodiments, the insert sequence is integrated into the genome by (i) DNA strand cleavage at the target site by an endonuclease encoded by one or more polynucleic acids, (ii) via target-primed reverse transcription (TPRT), or (iii) via reverse splicing of the insert sequence into the DNA target site in the genome.

[0038]

[0138] In some embodiments, the insertion sequence is integrated into the genome at a polyT site using the specificity of the endonuclease domain of human ORF2p.

[0039]

[0139] In some embodiments, the poly-T site comprises the sequence TTTTTA.

[0040]

[0140] In some embodiments, the one or more polynucleic acids comprise homology arms that are complementary to a target site in the genome.

[0041]

[0141] In some embodiments, the insertion sequence: (a) integrates into the genome at a locus that is not a ribosomal locus; (b) integrates into a gene or a regulatory region of a gene in the genome, thereby disrupting the gene or downregulating expression of the gene; (c) integrates into a gene or a regulatory region of a gene in the genome, thereby upregulating expression of the gene; or (d) integrates into the genome and replaces a gene in the genome.

[0042]

[0142] In some embodiments, the pharmaceutical composition further comprises (i) one or more siRNAs and / or (ii) an RNA guide sequence or a polynucleic acid encoding the RNA guide sequence, wherein the RNA guide sequence targets a DNA target site in the genome and the insertion sequence is integrated into the genome at the DNA target site in the genome.

[0043]

[0143] In some embodiments, one or more genes are knocked down in the methods provided herein. In some embodiments, one or more siRNAs are used in the compositions or methods described herein. For example, one or more genes can be knocked down to enhance integration, such as by modulating pathways that can inhibit LINE-1. In some embodiments, the one or more genes that are knocked down include ADAR1, ADAR2 (ADAR1B), APOBEC3C, BRCA1, let-7 miRNA, RNaseL, TASHOR (HUSH complex), and / or RAD51. For example, knocking down RNaseL can be used to enhance integration by inhibiting or preventing the degradation of mRNA, such as mRNA transcribed from LINE-1. For example, knocking down ADAR1, ADAR2 (ADAR1B), and / or BRCA1 can be used to enhance integration by inhibiting or preventing ADAR1, ADAR2 (ADAR1B), and / or BRCA1 from cis-binding ORF2p to the polyA tail for L1 RNP assembly. For example, knockdown of let-7 miRNA can be used to enhance integration by inhibiting or preventing let-7 miRNA from inhibiting translation, such as translation of ORF2p. let-7 miRNA. For example, knockdown of RAD51 and / or BRCA1 can be used to enhance integration by inhibiting or preventing repair of broken DNA by RAD51 and / or BRCA1.

[0044]

[0144] In some embodiments, the one or more polynucleic acids have a total length of between 3 kb and 20 kb.

[0045]

[0145] In some embodiments, the one or more polynucleic acids comprise one or more polyribonucleic acids, one or more RNAs, or one or more mRNAs.

[0046]

[0146] In some embodiments, the exogenous therapeutic polypeptide is selected from the group consisting of a ligand, an antibody, a receptor, an enzyme, a transport protein, a structural protein, a hormone, a contractile protein, a storage protein, and a transcription factor.

[0047]

[0147] In some embodiments, the exogenous therapeutic polypeptide is a receptor selected from the group consisting of a chimeric antigen receptor (CAR) and a T cell receptor (TCR).

[0048]

[0148] In some embodiments, the one or more polynucleic acids comprise a first expression cassette comprising a promoter sequence, a 5'UTR sequence, a 3'UTR sequence, and a polyA sequence: (i) the promoter sequence is upstream of the 5'UTR sequence, (ii) the 5'UTR sequence is upstream of the sequence of the mobile genetic element encoding the polypeptide, (iii) the 3'UTR sequence is downstream of the insertion sequence; (iv) the 3'UTR is upstream of the polyA sequence; and the 5'UTR sequence, 3'UTR sequence, or polyA sequence comprises a binding site for human ORF2p or a functional fragment thereof.

[0049]

[0149] In some embodiments, the insert sequence comprises a second expression cassette comprising a sequence that is the reverse complement of a promoter sequence, a sequence that is the reverse complement of a 5'UTR sequence, a sequence that is the reverse complement of a 3'UTR sequence, and a sequence that is the reverse complement of a polyA sequence: (i) the sequence that is the reverse complement of the promoter sequence is downstream of the sequence that is the reverse complement of the 5'UTR sequence; (ii) the sequence that is the reverse complement of the 5'UTR sequence is downstream of the sequence that is the reverse complement of a sequence encoding an exogenous therapeutic polypeptide; (iii) the sequence that is the reverse complement of the 3'UTR sequence is upstream of the sequence that is the reverse complement of a sequence encoding an exogenous therapeutic polypeptide; and (iv) the sequence that is the reverse complement of the polyA sequence is upstream of the sequence that is the reverse complement of the 3'UTR sequence and downstream of the sequence of a mobile gene encoding a polypeptide.

[0050]

[0150] In some embodiments, the promoter sequence of the first expression cassette is different from the promoter sequence of the second expression cassette.

[0051]

[0151] In some embodiments, the one or more LINE polypeptides comprise a first LINE polypeptide comprising human ORF1p or a functional fragment thereof and a second LINE polypeptide comprising human ORF2p or a functional fragment thereof, wherein the first LINE polypeptide and the second LINE polypeptide are translated from different open reading frames (ORFs).

[0052]

[0152] In some embodiments, the one or more polynucleic acids comprise a first polynucleic acid molecule encoding human ORF1p or a functional fragment thereof and a second polynucleic acid molecule encoding human ORF2p or a functional fragment thereof.

[0053]

[0153] In some embodiments, the one or more polynucleic acids comprise a 5'UTR sequence and a 3'UTR sequence, wherein (a) the 5'UTR comprises a sequence having at least 80% sequence identity to the 5'UTR from LINE-1 or ACUCCUCCCCAUCCUCUCCCUCUGUCCCUCUGUCCCUCUGACCCUGCACUGUCCCAGCACC; and / or (b) the 3'UTR comprises a sequence having at least 80% sequence identity to the 3'UTR from LINE-1 or CAGGACACAGCCUUGGAUCAGGACAGAGACUUGGGGGCCAUCCUGCCCCUCCAACCCGACAUGUGUACCUCAGCUUUUUCCCUCACUUGCAUCAAUAAAGCUUCUGUGUUUGGAACAG.

[0054]

[0154] In some embodiments, the sequence encoding the exogenous therapeutic polypeptide is intron-free.

[0055]

[0155] In some embodiments, the polypeptide encoded by the sequence of the mobile genetic element comprises a C-terminal nuclear localization signal (NLS), an N-terminal NLS, or both.

[0056]

[0156] In some embodiments, the sequence encoding the foreign polypeptide is not in frame with the sequence encoding ORF1p or a functional fragment thereof and / or is not in frame with the sequence encoding ORF2p or a functional fragment thereof.

[0057]

[0157] In some embodiments, the one or more polynucleic acids comprise a sequence encoding a nuclease domain, a nuclease domain not derived from ORF2p, a megaTAL nuclease domain, a TALEN domain, a Cas9 domain, a Cas6 domain, a Cas7 domain, a Cas8 domain, a zinc finger binding domain derived from an R2 retroelement, or a DNA binding domain that binds to a repeat sequence.

[0058]

[0158] In some embodiments, the one or more polynucleic acids comprise a sequence encoding a nuclease domain, wherein the nuclease domain comprises a mutation that reduces the activity of the nuclease domain compared to a nuclease domain that does not have nuclease activity or does not have the mutation.

[0059]

[0159] In some embodiments, ORF2p or a functional fragment thereof lacks nuclease activity or comprises a mutation selected from the group consisting of S228P and Y1180A, and / or ORF1p or a functional fragment thereof comprises a K3R mutation.

[0060]

[0160] In some embodiments, the insert sequence comprises a sequence that is the reverse complement of a sequence encoding two or more exogenous therapeutic polypeptides.

[0061]

[0161] In some embodiments, the one or more polynucleic acids comprise one or more polyribonucleic acids, the exogenous therapeutic polypeptide is a receptor selected from the group consisting of a chimeric antigen receptor (CAR) and a T-cell receptor (TCR), and the pharmaceutical composition is formulated for systemic administration to a human subject.

[0062]

[0162] In some embodiments, the one or more polynucleic acids (i) are formulated in nanoparticles selected from the group consisting of lipid nanoparticles and polymeric nanoparticles; and / or (ii) comprise one or more polynucleic acids selected from the group consisting of glycosylated RNA, circular RNA, and self-replicating RNA.

[0063]

[0163] Also provided herein is a method of treating a disease or condition in a human subject in need thereof, the method comprising administering to the human subject a pharmaceutical composition described herein.

[0064]

[0164] Also provided herein is a method for ex vivo modifying a population of human cells, comprising contacting a composition with a population of human cells ex vivo, thereby forming an ex vivo modified population of human cells, wherein the composition comprises one or more polynucleic acids or at least one vector encoding one or more polynucleic acids, wherein the one or more polynucleic acids comprise: (a) a mobile genetic element comprising a sequence encoding a polypeptide; and (b) an insert sequence that is the reverse complement of a sequence encoding an exogenous therapeutic polypeptide, wherein the ex vivo modified population of human cells is substantially non-immunogenic to a human subject.

[0065]

[0165] In one aspect, compositions and methods are provided herein that allow for the integration of genetic material into the genome of a cell, where the genetic material that can be integrated is not specifically limited in size. In some aspects, the methods described herein provide a one-step, single polynucleotide-mediated delivery and integration of genetic "cargo" into the genome of a cell. The genetic material can include coding sequences, such as sequences encoding transgenes, peptides, recombinant proteins, or antibodies, or fragments thereof, where the methods and compositions ensure stable expression of the transcript encoded by the coding sequence. The genetic material can also include non-coding sequences, such as regulatory RNA sequences, e.g., regulatory small interfering RNAs (siRNAs), microRNAs (miRNAs), long non-coding RNAs (lncRNAs), or one or more transcriptional regulators, such as promoters and / or enhancers, and can also include, but are not limited to, structurally related biomolecules, such as ribosomal RNAs (rRNAs), transfer RNAs (tRNAs), or fragments thereof, or combinations thereof.

[0066]

[0166] In another aspect, provided herein are methods and compositions for site-specific integration of genetic material into the genome of a cell, without specific size limitations, via non-viral delivery, ensuring both safety and efficacy of transfer.The provided methods and compositions can be particularly useful for developing therapeutic agents, such as therapeutic agents comprising a polynucleotide, which comprises genetic material and a mechanism that allows the polynucleotide or mRNA encoding the polynucleotide to be transferred into the cell and stably integrated into the genome of the cell.In some embodiments, the therapeutic agent can be a cell that comprises a polynucleotide that is stably integrated into the genome of a cell using the methods and compositions described herein.

[0067]

[0167] In one aspect, the present disclosure provides compositions and methods for stable gene transfer into cells. In some embodiments, the compositions and methods are for stable gene transfer into immune cells. In some cases, the immune cells are bone marrow cells. In some cases, the methods described herein relate to the development of bone marrow cells for immunotherapy.

[0068]

[0168] Provided herein is a method for treating a disease in a subject in need thereof, comprising administering to the subject a pharmaceutical composition: the pharmaceutical composition comprises a polycistronic mRNA sequence encoding a gene or a fragment thereof operably linked to a sequence encoding an L1 retrotransposon; and the gene or fragment thereof is at least 10.1 kb in length.

[0069]

[0169] Provided herein is a method for integrating a nucleic acid sequence into the genome of a cell, comprising contacting the cell with a composition comprising a polycistronic mRNA sequence encoding a gene or a fragment thereof, operably linked to a sequence encoding an L1 retrotransposon; the gene or a fragment thereof is at least 10.1 kb in length. In some embodiments, the gene or a fragment thereof (e.g., payload) is at least about 10.2 kb, 10.3 kb, 10.4 kb, 10.5 kb, 10.6 kb, 10.7 kb, 10.8 kb, 10.9 kb, 11 kb, 12 kb, 13 kb, 14 kb, 15 kb, 16 kb, 17 kb, 18 kb, 19 kb, 20 kb or longer.

[0070]

[0170] Provided herein is a method for integrating a nucleic acid sequence into the genome of a cell, comprising contacting the cell with a composition comprising a polycistronic mRNA sequence encoding a gene or fragment thereof operably linked to a sequence encoding an L1 retrotransposon; wherein the gene or fragment thereof is selected from the group consisting of ABCA4, MY07A, CEP290, CDH23, EYS, USH2a, GPR98, ALMS1, GDE, OTOF, and F8.

[0071]

[0171] Provided herein is a method for expressing a protein encoded by a recombinant nucleic acid in a cell, the method comprising: integrating the nucleic acid sequence into the genome of the cell by contacting the cell with a composition comprising a polycistronic mRNA sequence encoding a gene or fragment thereof operably linked to a sequence encoding an L1 retrotransposon; and expressing the protein encoded by the gene or fragment thereof, wherein expression of the protein is detectable more than 30 days after (a).

[0072]

[0172] In one embodiment of the methods described herein, the disease is a genetic disease.

[0073]

[0173] Provided herein are methods for treating Stargardt disease, LCA10, USH1D, DFNB12, retinitis pigmentosa (RP) USH2A, USH2C, Alström syndrome, glycogen storage disease type III, nonsyndromic hearing loss, hemophilia A, or Leber congenital amaurosis in a subject, the methods comprising: (i) introducing into the subject mRNA encoding a suitable gene or a fragment thereof operably linked to a human L1 transposon, or (ii) introducing into the subject a population of cells comprising mRNA encoding a suitable gene or a fragment thereof operably linked to a human L1 transposon.

[0074]

[0174] In one embodiment of the methods described herein, the method comprises treating Stargardt's disease in a subject in need thereof, wherein the mRNA encodes the ABCA4 gene or a fragment thereof.

[0075]

[0175] In one embodiment of the methods described herein, the method includes treating Usher syndrome type 1b (Usher 1b) disease in a subject in need of treatment, wherein the mRNA encodes the MY07A gene or a fragment thereof.

[0076]

[0176] In one embodiment of the methods described herein, the method comprises treating Leber congenital amaurosis (LCA) type 10 disease in a subject in need thereof, wherein the mRNA encodes the CEP290 gene or a fragment thereof.

[0077]

[0177] In one embodiment of the methods described herein, the method includes treating Usher Syndrome type 1D (USH1D) non-syndromic hearing loss or hearing loss USH1D, DFN12 disease in a subject in need of treatment, wherein the mRNA encodes the CDH23 gene or a fragment thereof.

[0078]

[0178] In one embodiment of the methods described herein, the method comprises treating retinitis pigmentosa (RP) in a subject in need thereof, wherein the mRNA encodes the EYS gene or a fragment thereof.

[0079]

[0179] In one embodiment of the methods described herein, the method comprises treating Usher Syndrome type 2A (USH2A), wherein the mRNA encodes the USH2a gene or a fragment thereof.

[0080]

[0180] In one embodiment of the methods described herein, the method comprises treating Usher Syndrome type 2C (USH2C), wherein the mRNA encodes the GPR98 gene or a fragment thereof.

[0081]

[0181] In one embodiment of the methods described herein, the method comprises treating Alström syndrome, wherein the mRNA encodes the ALMS1 gene or a fragment thereof.

[0082]

[0182] In one embodiment of the methods described herein, the method comprises treating glycogen storage disease type III, wherein the mRNA encodes a GDE gene or a fragment thereof.

[0083]

[0183] In one embodiment of the methods described herein, the method comprises treating non-syndromic hearing loss or hearing impairment, wherein the mRNA encodes the OTOF gene or a fragment thereof.

[0084]

[0184] In one embodiment of the methods described herein, the method comprises treating hemophilia A, and the mRNA encodes a factor VIII (F8) gene or a fragment thereof.

[0085]

[0185] Provided herein is a method for targeted replacement of a cellular genomic nucleic acid sequence, comprising: (A) introducing into the cell a polynucleotide sequence encoding a first protein complex comprising a targeted excision mechanism for excising a nucleic acid sequence comprising one or more mutations from the genome of the cell; and (B) a recombinant mRNA encoding a second protein complex, the recombinant mRNA comprising: (i) a nucleic acid sequence comprising the nucleic acid sequence excised in (A) that does not contain the one or more mutations, and (ii) a sequence encoding an L1 retrotransposon ORF2 protein under the influence of an independent promoter.

[0086]

[0186] In one embodiment of the methods described herein, the nucleic acid sequence comprising one or more mutations comprises a pathogenic variant of a gene of the cell.

[0087]

[0187] In one embodiment of the methods described herein, the nucleic acid sequence in (B), including the nucleic acid sequence that does not contain the one or more mutations, is operably linked to the ORF2 sequence.

[0088]

[0188] In one embodiment of the methods described herein, the method further comprises introducing a sequence comprising multiple thymidine residues at the excision site.

[0089]

[0189] In some embodiments, the step of introducing the sequence comprises introducing at least four thymidine residues.

[0090]

[0190] In one embodiment of the methods described herein, the targeted excision mechanism comprises a sequence-guided site-specific excision endonuclease.

[0091]

[0191] In one embodiment of the methods described herein, the targeted excision mechanism comprises a CRISPR-Cas system.

[0092]

[0192] In some embodiments, the targeted excision mechanism is a modified recombinant LINE 1 (L1) endonuclease.

[0093]

[0193] In some embodiments, introducing a sequence comprising multiple thymidine residues comprises base extension by prime editing at the excision site.

[0094]

[0194] In some embodiments, the mRNA sequence encoding the L1 retrotransposon ORF2 protein further comprises a sequence encoding the L1 retrotransposon ORF1 protein.

[0095]

[0195] In some embodiments, the mRNA includes a sequence for an inducible promoter.

[0096]

[0196] In one embodiment of the methods described herein, the excision sequence is longer than 1000 bases.

[0097]

[0197] In one embodiment of the methods described herein, the excision sequence is longer than 6 kb.

[0098]

[0198] In one embodiment of the methods described herein, the excision sequence is about 10 kb.

[0099]

[0199] In some embodiments, the cell is a lymphocyte. In some embodiments, the cell is a bone marrow cell. In some embodiments, the cell is an epithelial cell. In some embodiments, the cell is a cancer cell.

[0100]

[0200] In some embodiments, the nucleic acid sequence encodes an ATP-binding cassette (ABC) transporter gene, (ABCA4) gene, or a fragment thereof.

[0101]

[0201] In some embodiments, the nucleic acid sequence encodes a MY07A, CEP290, CDH23, EYS, USH2a, GPR98, ALMS1, GDE, OTOF, or F8 gene or a fragment thereof.

[0102]

[0202] In some embodiments, introducing comprises introducing into a cell ex vivo. In some embodiments, introducing comprises electroporation. In some embodiments, introducing comprises introducing into a cell in vivo. In some embodiments, expression of the nucleic acid sequence comprising a sequence that does not contain one or more mutations is detectable for at least 35 days after introduction into the cell. In some embodiments, introducing into a subject comprises direct systemic administration of mRNA.

[0103]

[0203] In some embodiments, introducing into a subject comprises local administration of the mRNA.

[0104]

[0204] In some embodiments, the mRNA sequence includes a portion that targets a cell.

[0105]

[0205] In some embodiments, the cell-targeting moiety is an aptamer.

[0106]

[0206] In some embodiments, introducing into the subject comprises introducing mRNA into the retina of the subject.

[0107]

[0207] Provided herein is a method for integrating a nucleic acid sequence into the genome of a cell, the method comprising introducing into the cell a recombinant mRNA or a vector encoding the mRNA, wherein the mRNA: (a) is an insert sequence, the insert sequence comprising (i) a foreign sequence or (ii) a sequence that is the reverse complement of the foreign sequence; (b) a 5'UTR sequence and a 3'UTR sequence downstream of the 5'UTR sequence; the 5'UTR sequence or the 3'UTR sequence comprises a binding site for a human ORF protein, and the insert sequence is integrated into the genome of the cell, wherein the insert sequence is a gene selected from the group consisting of ABCA4, MY07A, CEP290, CDH23, EYS, USH2a, GPR98, ALMS1, GDE, OTOF, and F8.

[0108]

[0208] In some embodiments, the 5'UTR sequence or the 3'UTR sequence comprises a binding site for human ORF2p.

[0109]

[0209] Provided herein is a method for integrating a nucleic acid sequence into the genome of an immune cell, comprising the step of introducing a recombinant mRNA or a vector encoding the mRNA, wherein the mRNA: (a) is an insert sequence, the insert sequence comprising (i) a foreign sequence or (ii) a sequence that is the reverse complement of the foreign sequence; (b) a 5'UTR sequence and a 3'UTR sequence downstream of the 5'UTR sequence; the 5'UTR sequence or the 3'UTR sequence comprises an endonuclease binding site and / or a reverse transcriptase binding site, wherein the insert sequence is integrated into the genome of the immune cell, and the insert sequence is a gene selected from the group consisting of ABCA4, MY07A, CEP290, CDH23, EYS, USH2a, GPR98, ALMS1, GDE, OTOF, and F8.

[0110]

[0210] Provided herein is a method for integrating a nucleic acid sequence into the genome of a cell, the method comprising introducing a recombinant mRNA or a vector encoding the mRNA, wherein the mRNA comprises: (a) an insertion sequence, the insertion sequence comprising (i) a foreign sequence or (ii) a sequence that is the reverse complement of the foreign sequence; (b) a 5'UTR sequence, a sequence of a human retrotransposon downstream of the 5'UTR sequence, and a 3'UTR sequence downstream of the sequence of the human retrotransposon; wherein the 5'UTR sequence or the 3'UTR sequence comprises an endonuclease binding site and / or a reverse transcriptase binding site, and the sequence of the human retrotransposon encodes two proteins that are translated from a single RNA containing two ORFs, wherein the insertion sequence is integrated into the genome of the cell, and the insertion sequence is a gene selected from the group consisting of ABCA4, MY07A, CEP290, CDH23, EYS, USH2a, GPR98, ALMS1, GDE, OTOF, and F8.

[0111]

[0211] In some embodiments, the 5'UTR sequence or the 3'UTR sequence comprises an ORF2p binding site. In some embodiments, the ORF2p binding site is a polyA sequence in the 3'UTR sequence.

[0112]

[0212] In some embodiments, the mRNA comprises a sequence of a human retrotransposon. In some embodiments, the sequence of the human retrotransposon is downstream of the 5'UTR sequence.

[0113]

[0213] In some embodiments, the sequence of human retrotransposon is located upstream of 3'UTR sequence.In some embodiments, the sequence of human retrotransposon encodes two proteins that are translated from a single RNA that contains two ORFs.In some embodiments, the two ORFs are non-overlapping ORFs.

[0114]

[0214] In some embodiments, the sequence of the human retrotransposon comprises a sequence of a non-LTR retrotransposon. In some embodiments, the sequence of the human retrotransposon comprises a LINE-1 retrotransposon. In some embodiments, the LINE-1 retrotransposon is a human LINE-1 retrotransposon. In some embodiments, the sequence of the human retrotransposon comprises a sequence encoding an endonuclease and / or a reverse transcriptase.

[0115]

[0215] In some embodiments, the endonuclease and / or reverse transcriptase is ORF2p.

[0116]

[0216] In some embodiments, the reverse transcriptase is a group II intron reverse transcriptase domain.

[0117]

[0217] In some embodiments, the endonuclease and / or reverse transcriptase is a minke whale endonuclease and / or reverse transcriptase.

[0118]

[0218] In some embodiments, the sequence of the human retrotransposon comprises a sequence encoding ORF2p. In some embodiments, the insertion sequence is integrated into the genome at a poly-T site using the specificity of the endonuclease domain of ORF2p. In some embodiments, the poly-T site comprises the sequence TTTTTA. In some embodiments, the retrotransposon comprises ORF1p and / or ORF2p fused to a nuclear retention sequence. In some embodiments, the nuclear retention sequence is an Alu sequence. In some embodiments, ORF1p and / or ORF2p are fused to an MS2 coat protein. In some embodiments, the 5'UTR sequence or the 3'UTR sequence comprises at least one, two, three, or more MS2 hairpin sequences.

[0119]

[0219] Provided herein is a composition comprising a recombinant mRNA or a vector encoding the mRNA, wherein the mRNA comprises a human LINE-1 transposon sequence comprising: (i) a human LINE-1 transposon 5'UTR sequence, (ii) a sequence encoding ORF1p downstream of the human LINE-1 transposon 5'UTR sequence, (iii) an inter-ORF linker sequence downstream of the sequence encoding ORF1p, (iv) a sequence encoding ORF2p downstream of the inter-ORF linker sequence, and (v) a 3'UTR sequence derived from a human LINE-1 transposon downstream of the sequence encoding ORF2p, wherein the 3'UTR sequence comprises an insertion sequence, wherein the insertion sequence is the reverse complement of a sequence encoding a foreign polypeptide or the reverse complement of a sequence encoding a foreign regulatory element, and wherein the insertion sequence is a gene selected from the group consisting of ABCA4, MY07A, CEP290, CDH23, EYS, USH2a, GPR98, ALMS1, GDE, OTOF, and F8.

[0120]

[0220] Provided herein is a composition comprising a nucleic acid comprising a nucleotide sequence encoding (a) a long interspersed nuclear element (LINE) polypeptide, the LINE polypeptide comprising human ORF1p and human ORF2p, and (b) an insert sequence that is the reverse complement of a sequence encoding a foreign polypeptide or the reverse complement of a sequence encoding a foreign regulatory element, wherein the composition is substantially non-immunogenic, and the insert sequence is a gene selected from the group consisting of ABCA4, MY07A, CEP290, CDH23, EYS, USH2a, GPR98, ALMS1, GDE, OTOF, and F8.

[0121]

[0221] Immunotherapy using phagocytes involves the creation and use of engineered bone marrow cells, e.g., macrophages or other phagocytes, that attack and kill diseased or infected cells, such as cancer cells. Engineered bone marrow cells, e.g., macrophages and other phagocytes, are prepared by incorporating into bone marrow cells, via recombinant nucleic acid technology, engineered proteins, e.g., chimeric antigen receptors, that contain targeting antigen-binding extracellular domains designed to bind to specific antigens on the surface of targets, e.g., target cells, e.g., cancer cells. Binding of the engineered chimeric receptor to the target antigen, e.g., a cancer antigen (or similarly, a disease target), initiates phagocytosis of the target. This triggers a two-component process: 1. engulfment and lysis of the target by phagocytes destroys the target and eliminates it as the first line of immune defense; and 2. Target-derived antigens are digested in the phagolysosomes of the bone marrow cells and presented on the surface of the bone marrow cells, which then triggers T cell activation, further activation of the immune response, and the generation of immunological memory. The chimeric receptor is modified to enhance phagocytosis and immune activation of bone marrow cells into which it is incorporated and expressed. The chimeric antigen receptor of the present disclosure is variously referred to herein as a chimeric fusion protein, CFP, phagocytic receptor (PR) fusion protein (PFP), or chimeric antigen receptor for phagocytosis (CAR-P), with each term covering the concept of a recombinant chimeric and / or fusion receptor protein. In some embodiments, a gene encoding a non-receptor protein is also typically co-expressed in bone marrow cells to enhance chimeric antigen receptor function. In summary, contemplated herein are various modified receptor and non-receptor recombinant proteins designed to enhance the phagocytosis and / or immune response of bone marrow cells against disease targets, as well as methods and compositions for creating and incorporating recombinant nucleic acids encoding the modified receptor or non-receptor recombinant proteins, which result in modified bone marrow cells suitable for immunotherapy.

[0122]

[0222] In one aspect, the present disclosure provides compositions and methods for stable gene transfer into cells, and the cells can be any somatic cells.In some embodiments, compositions and methods are designed for cell-specific or tissue-specific delivery.In some cases, the methods described herein relate to providing functional protein or its fragment to correct missing or defective (mutated) protein in vivo, for example, for protein replacement therapy.

[0123]

[0223] The incorporation of recombinant nucleic acid into cells can be achieved by one or more gene transfer techniques available in the latest technology.However, the therapeutic integration of foreign gene (for example, nucleic acid) elements into genome still faces several challenges.Some of them are achieving stable integration in a safe and reliable manner, and efficient and long-term expression.Most successful gene transfer systems for the genome integration of cargo nucleic acid sequences rely on viral delivery mechanisms, which have some inherent problems regarding safety and effectiveness.The delivery and integration of long nucleic acid sequences cannot be achieved by current gene editing systems.

[0124]

[0224] To date, little attention has been paid to the creation and use of modified bone marrow cells for stable, long-term gene transfer and expression of transgenes. For example, gene transfer into ex vivo differentiated mammalian cells for cell therapy can be achieved via viral gene transfer mechanisms. However, the use of viral gene transfer vectors has several strategic drawbacks, including the undesirable possibility of transgene silencing over time, preferential integration into transcriptionally active sites of the genome with associated undesired activation of other genes (e.g., oncogenes), and genotoxicity. In addition to safety issues, the increased costs and inefficient efforts involved in producing, storing, and handling integrated viruses often hinder the large-scale use of viral vector-mediated gene-modified cells in therapeutic applications. These persistent concerns about safety and the cost and scale of vector production associated with viral vectors necessitate alternative methods for effective therapy.

[0125]

[0225] Integration of a transgene into the genome of cells used for immunotherapy can be advantageous in that the integration is stable and fewer cells are required for delivery during therapy. On the other hand, integration of a transgene into non-dividing cells can be challenging in that it affects the health and function of the cells and their ultimate lifespan in vivo, thus affecting their overall usefulness as a therapy. In some embodiments, the methods described herein for generating bone marrow cells for immunotherapy can be the cumulative product of several steps and compositions, including, but not limited to, selecting bone marrow cells for modification; methods and compositions for incorporating recombinant nucleic acids into bone marrow cells; methods and compositions for enhancing expression of recombinant nucleic acids; methods and compositions for selecting and modifying vectors; and methods for preparing recombinant nucleic acids suitable for in vivo administration for uptake and incorporation of recombinant nucleic acids by bone marrow cells in vivo, thereby generating bone marrow cells for therapy. In some aspects, one or more embodiments of the various inventions described herein are transferable with one another, and it is expected that one of skill in the art will use the embodiments alternatively, in combination, or interchangeably without undue experimentation. All such variations of the disclosed elements are contemplated and fully encompassed herein.

[0126]

[0226] In one aspect, transposons, or transposable elements (TEs), are contemplated herein as a means of integrating heterologous, synthetic, or recombinant nucleic acids encoding a transgene of interest into bone marrow cells. Transposons, or transposable elements, are genetic elements capable of transferring segments of genetic material into a genome through the use of enzymes known as transposases. Mammalian genomes contain numerous transposable element (TE)-derived sequences, with up to 70% of our genomes being TE-derived sequences (de Koning et al., 2011; Richardson et al., 2015). These elements can be utilized to introduce genetic material into a cell's genome. TE elements are capable of moving, often described as "jumping," genetic material within the genome. TEs generally exist in eukaryotic genomes in a reversibly inactive, epigenetically silenced form. This disclosure describes methods and compositions for the efficient and stable integration of transgenes into macrophages and other phagocytic cells. The methods are based on the use of transposases and transposable element mRNAs encoding transposases. In some embodiments, long interspersed repeat 1 (L1) RNA is used for stable integration and / or retrotransposition of transgenes into cells (e.g., macrophages or phagocytes).

[0127]

[0227] Contemplated herein are methods for the stable integration of foreign nucleic acid sequences into the genome of a cell mediated by retrotransposons. The methods can utilize the random genome integration mechanism of retrotransposons without causing adverse effects to the cell. The methods described herein can be used for the robust and versatile integration of foreign nucleic acid sequences into cells, resulting in the foreign nucleic acid being integrated into a safe locus within the genome and being expressed without being silenced by the cell's inherent defense mechanisms. The methods described herein can be used to integrate foreign nucleic acids of about 1 kb, about 2 kb, about 3 kb, about 4 kb, about 5 kb, about 6 kb, about 7 kb, about 8 kb, about 9 kb, about 10 kb, or larger in size. In some embodiments, the foreign nucleic acid is not integrated into a ribosomal locus. In some embodiments, the foreign nucleic acid is not integrated into the ROSA26 locus or another safe harbor locus. In some embodiments, the methods and compositions described herein can integrate foreign nucleic acid sequences into any location within the genome of a cell. Furthermore, retrotransposition systems are contemplated herein that are developed to integrate foreign nucleic acid sequences into specific, predetermined sites within the genome of a cell without causing adverse effects. The disclosed methods and compositions incorporate several mechanisms for modifying retrotransposons for highly specific integration of foreign nucleic acids into cells with high accuracy. The retrotransposon selected for this purpose can be a human retrotransposon.

[0128]

[0228] The methods and compositions described herein represent a significant breakthrough in the molecular systems and mechanisms for manipulating the genome of a cell. For the first time, it is shown here how to utilize the human retrotransposon system for the non-viral delivery and stable integration of large fragments of foreign nucleic acid sequences (at least 100 nucleobases, at least 1 kb, at least 2 kb, at least 3 kb, etc.) into non-conserved regions of the genome that are not rDNA, ribosomal loci, or designated safe harbor loci such as the ROSA26 locus.

[0129]

[0229] In some embodiments, a retrotransposition system is used to stably integrate and express a non-endogenous nucleic acid into a genome, wherein the non-endogenous nucleic acid comprises a retrotransposition element within the nucleic acid sequence. In some embodiments, the cell's endogenous retrotransposition system (e.g., proteins and enzymes) is used to stably express the non-endogenous nucleic acid in the cell. In some embodiments, the cell's endogenous retrotransposition system (e.g., proteins and enzymes, e.g., the LINE-1 retrotransposition system) is used, but one or more components of the retrotransposition system may be further expressed to stably express the non-endogenous nucleic acid in the cell.

[0130]

[0230] In some embodiments, provided herein are synthetic nucleic acids that encode a transgene and encode one or more components for genomic integration and / or retrotransposition.

[0131]

[0231] In one aspect, the present disclosure provides a method for integrating a nucleic acid sequence into the genome of a cell, the method comprising: introducing recombinant mRNA or a vector encoding mRNA into a cell, wherein the mRNA comprises an insertion sequence comprising a foreign sequence or the reverse complement of the sequence of the foreign sequence; a 5'UTR sequence and a 3'UTR sequence downstream of the 5'UTR sequence, wherein the 5'UTR sequence or the 3'UTR sequence comprises the binding site for human ORF protein, and the insertion sequence is integrated into the genome of a cell.In some embodiments, the 5'UTR sequence or the 3'UTR sequence comprises the binding site for human ORF2p.

[0132]

[0232]

[0010] In one aspect, provided herein is a method for integrating a nucleic acid sequence into the genome of an immune cell, comprising introducing a recombinant mRNA or a vector encoding the mRNA, wherein the mRNA comprises an insertion sequence, the insertion sequence comprising (i) the foreign sequence or (ii) a sequence that is the reverse complement of the foreign sequence; a 5'UTR sequence and a 3'UTR sequence downstream of the 5'UTR sequence, wherein the 5'UTR sequence or the 3'UTR sequence comprises an endonuclease binding site and / or a reverse transcriptase binding site, and wherein the transgene sequence is integrated into the genome of the immune cell.

[0133]

[0233]

[0010] In one aspect, provided herein is a method for integrating a nucleic acid sequence into the genome of a cell, comprising the step of introducing a recombinant mRNA or a vector encoding the mRNA, wherein the mRNA comprises an insertion sequence, the insertion sequence comprising (i) a foreign sequence or (ii) a sequence that is the reverse complement of the foreign sequence; a 5'UTR sequence, a sequence of a human retrotransposon downstream of the 5'UTR sequence, and a 3'UTR sequence downstream of the human retrotransposon sequence, wherein the 5'UTR sequence or the 3'UTR sequence comprises an endonuclease binding site and / or a reverse transcriptase binding site, and the human retrotransposon sequence encodes two proteins that are translated from a single RNA containing two ORFs, and wherein the insertion sequence is integrated into the genome of the cell.

[0134]

[0234] In some embodiments, the 5'UTR sequence or the 3'UTR sequence comprises an ORF2p binding site. In some embodiments, the ORF2p binding site is a polyA sequence in the 3'UTR sequence.

[0135]

[0235] In some embodiments, the mRNA comprises a human retrotransposon sequence. In some embodiments, the human retrotransposon sequence is downstream of the 5' UTR sequence. In some embodiments, the human retrotransposon sequence is upstream of the 3' UTR sequence. In some embodiments, a polynucleotide sequence (e.g., an insert) desired to be transferred and integrated into the genome of a cell is inserted into a recombinant nucleic acid construct at a site 3' to the sequence encoding ORF1. In some embodiments, a polynucleotide sequence desired to be transferred and integrated into the genome of a cell is inserted into a recombinant nucleic acid construct at a site 3' to the sequence encoding ORF2. In some embodiments, a sequence desired to be transferred and integrated into the genome of a cell is inserted into the 3' UTR of ORF1 or ORF2, or both. In some embodiments, a polynucleotide sequence desired to be transferred and integrated into the genome of a cell is inserted into a recombinant nucleic acid construct upstream of the polyA tail of ORF2.

[0136]

[0236] In some embodiments, the sequence of human retrotransposon encodes two proteins that are translated from a single RNA containing two ORFs.In some embodiments, the two ORFs are non-overlapping ORFs.In some embodiments, the two ORFs are ORF1 and ORF2.In some embodiments, ORF1 encodes ORF1p, and ORF2 encodes ORF2p.

[0137]

[0237] In some embodiments, the sequence of the human retrotransposon comprises a sequence of a non-LTR retrotransposon. In some embodiments, the sequence of the human retrotransposon comprises a LINE-1 retrotransposon. In some embodiments, the LINE-1 retrotransposon is a human LINE-1 retrotransposon. In some embodiments, the sequence of the human retrotransposon comprises a sequence encoding an endonuclease and / or reverse transcriptase. In some embodiments, the endonuclease and / or reverse transcriptase is ORF2p. In some embodiments, the reverse transcriptase is a group II intron reverse transcriptase domain. In some embodiments, the endonuclease and / or reverse transcriptase is a minke whale endonuclease and / or reverse transcriptase. In some embodiments, the sequence of the human retrotransposon comprises a sequence encoding ORF2p. In some embodiments, the insertion sequence is integrated into the genome at a poly-T site using the specificity of the endonuclease domain of ORF2p. In some embodiments, the poly-T site comprises the sequence TTTTTA.

[0138]

[0238] In some embodiments, provided herein is a polynucleotide construct comprising an mRNA, wherein the mRNA comprises a sequence encoding a human retrotransposon, and (i) the sequence of the human retrotransposon comprises a sequence encoding ORF1p, (ii) the mRNA does not comprise a sequence encoding ORF1p, or (iii) the mRNA comprises a replacement sequence for the sequence encoding ORF1p with a 5'UTR sequence from a complementary gene. In some embodiments, the mRNA comprises a first mRNA molecule encoding ORF1p and a second mRNA molecule encoding an endonuclease and / or reverse transcriptase. In some embodiments, the mRNA is an mRNA molecule comprising a first sequence encoding ORF1p and a second sequence encoding an endonuclease and / or reverse transcriptase. In some embodiments, the first sequence encoding ORF1p and the second sequence encoding the endonuclease and / or reverse transcriptase are separated by a linker sequence.

[0139]

[0239] In some embodiments, the linker sequence comprises an internal ribosome entry sequence (IRES). In some embodiments, the IRES is an IRES from CVB3 or EV71. In some embodiments, the linker sequence encodes a self-cleaving peptide sequence. In some embodiments, the linker sequence encodes a T2A, E2A, or P2A sequence.

[0140]

[0240] In some embodiments, the sequence of the human retrotransposon comprises a sequence encoding ORF1p fused to an additional protein sequence and / or a sequence encoding ORF2p fused to an additional protein sequence. In some embodiments, ORF1p and / or ORF2p are fused to a nuclear retention sequence. In some embodiments, the nuclear retention sequence is an Alu sequence. In some embodiments, ORF1p and / or ORF2p are fused to an MS2 coat protein. In some embodiments, the 5'UTR or 3'UTR sequence comprises at least one, two, three, or more MS2 hairpin sequences. In some embodiments, the 5'UTR or 3'UTR sequence comprises a sequence that promotes or enhances interaction of an mRNA polyA tail with an endonuclease and / or reverse transcriptase. In some embodiments, the 5'UTR or 3'UTR sequence comprises a sequence that promotes or enhances interaction of a polyA-binding protein (e.g., PABP) with an endonuclease and / or reverse transcriptase. In some embodiments, the 5' or 3' UTR sequence comprises a sequence that increases the specificity of an endonuclease and / or reverse transcriptase for the mRNA relative to other mRNAs expressed by the cell, hi some embodiments, the 5' or 3' UTR sequence comprises an Alu element sequence.

[0141]

[0241] In some embodiments, the first sequence encoding ORF1p and the second sequence encoding the endonuclease and / or reverse transcriptase have the same promoter. In some embodiments, the insert sequence has a promoter that is different from the promoter of the first sequence encoding ORF1p. In some embodiments, the insert sequence has a promoter that is different from the promoter of the second sequence encoding the endonuclease and / or reverse transcriptase. In some embodiments, the first sequence encoding ORF1p and / or the second sequence encoding the endonuclease and / or reverse transcriptase have a promoter or transcription start site selected from the group consisting of an inducible promoter, a CMV promoter or transcription start site, a T7 promoter or transcription start site, an EF1a promoter or transcription start site, and combinations thereof. In some embodiments, the insert sequence has a promoter or transcription start site selected from the group consisting of an inducible promoter, a CMV promoter or transcription start site, a T7 promoter or transcription start site, an EF1a promoter or transcription start site, and combinations thereof.

[0142]

[0242] In some embodiments, the first sequence encoding ORF1p and the second sequence encoding the endonuclease and / or reverse transcriptase are codon-optimized for expression in human cells.

[0143]

[0243] In some embodiments, the mRNA comprises a WPRE element. In some embodiments, the mRNA comprises a selectable marker. In some embodiments, the mRNA comprises a sequence encoding an affinity tag. In some embodiments, the affinity tag is linked to a sequence encoding an endonuclease and / or a reverse transcriptase.

[0144]

[0244] In some embodiments, the 3'UTR comprises a polyA sequence, or the polyA sequence is added to the mRNA in vitro. In some embodiments, the polyA sequence is downstream of the sequence encoding the endonuclease and / or reverse transcriptase. In some embodiments, the insertion sequence is upstream of the polyA sequence.

[0145]

[0245] In some embodiments, the 3'UTR sequence comprises an insertion sequence. In some embodiments, the insertion sequence comprises a sequence that is the reverse complement of the sequence encoding the foreign polypeptide. In some embodiments, the insertion sequence comprises a polyadenylation site. In some embodiments, the insertion sequence comprises an SV40 polyadenylation site. In some embodiments, the insertion sequence comprises a polyadenylation site upstream of the sequence that is the reverse complement of the sequence encoding the foreign polypeptide. In some embodiments, the insertion sequence is integrated into the genome at a locus that is not a ribosomal locus. In some embodiments, the insertion sequence is integrated into the genome at a locus that is not an rDNA locus. In some embodiments, the insertion sequence is integrated into a gene or a regulatory region of a gene, thereby disrupting the gene or down-regulating gene expression. In some embodiments, the insertion sequence is integrated into a gene or a regulatory region of a gene, thereby up-regulating gene expression. In some embodiments, the insertion sequence is integrated into the genome and replaces the gene. In some embodiments, the insertion sequence is stably integrated into the genome. In some embodiments, the insertion sequence is retrotransposed into the genome. In some embodiments, the insertion sequence is integrated into the genome by the cleavage of the DNA strand at the target site by the endonuclease encoded by the mRNA.In some embodiments, the insertion sequence is integrated into the genome through target-primed reverse transcription (TPRT).In some embodiments, the insertion sequence is integrated into the genome through reverse splicing of the mRNA to the DNA target site of the genome.

[0146]

[0246] In some embodiments, the cell is an immune cell. In some embodiments, the immune cell is a T cell or a B cell. In some embodiments, the immune cell is a bone marrow cell. In some embodiments, the immune cell is selected from the group consisting of monocytes, macrophages, dendritic cells, dendritic precursor cells, and macrophage precursor cells.

[0147]

[0247] In some embodiments, the mRNA is a self-integrating mRNA. In some embodiments, the method comprises introducing the mRNA into a cell. In some embodiments, the method comprises introducing a vector encoding the mRNA into the cell. In some embodiments, the method comprises introducing the mRNA or a vector encoding the mRNA into the cell ex vivo. In some embodiments, the method further comprises administering the cell to a human subject. In some embodiments, the method comprises administering the mRNA or a vector encoding the mRNA to the human subject. In some embodiments, an immune response is not elicited in the human subject. In some embodiments, the mRNA or vector is substantially non-immunogenic.

[0148]

[0248] In some embodiments, the vector is a plasmid or a viral vector. In some embodiments, the vector comprises a non-LTR retrotransposon. In some embodiments, the vector comprises a human L1 element. In some embodiments, the vector comprises the L1 retrotransposon ORF1 gene. In some embodiments, the vector comprises the L1 retrotransposon ORF2 gene. In some embodiments, the vector comprises an L1 retrotransposon. In some embodiments, provided herein is an mRNA comprising a sequence encoding a human LINE 1 retrotransposable element and a payload comprising a nucleic acid sequence capable of retrotransposing and integrating into the genome of a cell comprising the mRNA. In some embodiments, provided herein is an mRNA that can be delivered to a living cell, e.g., a human cell, comprising a payload comprising a sequence encoding a human LINE 1 retrotransposable element and a nucleic acid sequence capable of retrotransposing and integrating into the genome of the cell. In some embodiments, the sequence encoding the human LINE 1 retrotransposable element comprises an L1 retrotransposon ORF1 sequence or a fragment thereof. In some embodiments, the sequence encoding the human LINE 1 retrotransposable element comprises an L1 retrotransposon ORF2 sequence or a fragment thereof. In some embodiments, the sequence encoding the human LINE 1 retrotransposable element comprises an L1 retrotransposon ORF1 sequence or a fragment thereof, and an L1 retrotransposon ORF2 sequence or a fragment thereof, and a nucleic acid "payload" sequence, which is a heterologous sequence that is integrated into the genome of a cell by retrotransposition (see, e.g., Figure 1B).

[0149]

[0249] In some embodiments, the mRNA is at least about 1, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 2.7, 2.8, 2.9, or 3 kilobases, hi some embodiments, the mRNA is at most about 2.5, 2.6, 2.7, 2.8, 2.9, 3, 3.1, 3.2, 3.3, 3.4, 3.5, 3.6, 3.7, 3.8, 3.9, 4, 4.1, 4.2, 4.3, 4.4, 4.5, 4.6, 4.7, 4.8, 4.9, or 5 kilobases. In some embodiments, the mRNA is at least about 5.1, 5.2, 5.3, 5.4, 5.5, 5.6, 5.7, 5.8, 5.9, or 6 kilobases. In some embodiments, the mRNA is at least about 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, or 7 kilobases. In some embodiments, the mRNA is at least about 7.1, 7.2, 7.3, 7.4, 7.5, 7.6, 7.7, 7.8, 7.9, or 8 kilobases. In some embodiments, the mRNA is at least about 8.1, 8.2, 8.3, 8.4, 8.5, 8.6, 8.7, 8.8, 8.9, or 9 kilobases. In some embodiments, the mRNA is at least about 9.1, 9.2, 9.3, 9.4, 9.5, 9.6, 9.7, 9.8, 9.9, or 10 kilobases.

[0150]

[0250] In some embodiments, the mRNA comprises a sequence that inhibits or prevents mRNA degradation. In some embodiments, the sequence that inhibits or prevents mRNA degradation inhibits or prevents mRNA degradation by exonucleases or RNAses. In some embodiments, the sequence that inhibits or prevents mRNA degradation is a G4 structure, a pseudoknot, or a triplex sequence. In some embodiments, the sequence that inhibits or prevents mRNA degradation is an exoribonuclease-resistant RNA structure derived from flavivirus RNA or an ENE element derived from KSV. In some embodiments, the sequence that inhibits or prevents mRNA degradation inhibits or prevents mRNA degradation by deadenylases. In some embodiments, the sequence that inhibits or prevents mRNA degradation comprises non-adenosine nucleotides within or at the end of the polyA tail of the mRNA. In some embodiments, the sequence that inhibits or prevents mRNA degradation improves mRNA stability. In some embodiments, the foreign sequence comprises a sequence encoding a foreign polypeptide. In some embodiments, the sequence encoding the foreign polypeptide is not in frame with a sequence encoding an endonuclease and / or a reverse transcriptase. In some embodiments, the sequence encoding the foreign polypeptide is not in frame with a sequence encoding an endonuclease and / or a reverse transcriptase. In some embodiments, the foreign sequence does not contain an intron. In some embodiments, the foreign sequence comprises a sequence encoding a foreign polypeptide selected from the group consisting of an enzyme, a receptor, a transport protein, a structural protein, a hormone, an antibody, a contractile protein, and a storage protein. In some embodiments, the foreign sequence comprises a sequence encoding a foreign polypeptide selected from the group consisting of a chimeric antigen receptor (CAR), a ligand, an antibody, a receptor, and an enzyme. In some embodiments, the foreign sequence comprises a regulatory sequence. In some embodiments, the regulatory sequence comprises a cis-acting regulatory sequence. In some embodiments, the regulatory sequence comprises a cis-acting regulatory sequence selected from the group consisting of an enhancer, a silencer, a promoter, or a response element. In some embodiments, the regulatory sequence comprises a trans-acting regulatory sequence. In some embodiments, the regulatory sequence comprises a trans-acting regulatory sequence encoding a transcription factor.

[0151]

[0251] In some embodiments, integration of the insertion sequence does not adversely affect the health of the cell. In some embodiments, the endonuclease, reverse transcriptase, or both are capable of site-specific integration of the insertion sequence.

[0152]

[0252] In some embodiments, the retrotransposon system used herein is further modified for precise site-specific integration.In some embodiments, the retrotransposon system used herein is combined with CRISPR-Cas system to improve specificity.In some embodiments, ORF polypeptide binding sequence, for example, TTTTTA, can be the modification site specific to the genome sequence of cell.

[0153]

[0253] In some embodiments, the mRNA comprises a sequence encoding an additional nuclease domain or a nuclease domain not derived from ORF2. In some embodiments, the mRNA comprises a sequence encoding a DNA binding domain that binds to a repeat sequence, such as a megaTAL nuclease domain, a TALEN domain, a Cas9 domain, a zinc finger binding domain derived from an R2 retroelement, or Rep78 derived from AAV. In some embodiments, the endonuclease comprises a mutation that reduces the activity of the endonuclease compared to an endonuclease that does not have the mutation. In some embodiments, the endonuclease is an ORF2p endonuclease, and the mutation is S228P. In some embodiments, the mRNA comprises a sequence encoding a domain that improves the accuracy and / or processivity of the reverse transcriptase. In some embodiments, the reverse transcriptase is a reverse transcriptase derived from a retroelement other than ORF2, or a reverse transcriptase that has higher accuracy and / or processivity compared to the reverse transcriptase of ORF2p. In some embodiments, the reverse transcriptase is a group II intron reverse transcriptase. In some embodiments, the group II intron reverse transcriptase is a group IIA intron reverse transcriptase, a group IIB intron reverse transcriptase, or a group IIC intron reverse transcriptase, hi some embodiments, the group II intron reverse transcriptase is TGIRT-II or TGIRT-III.

[0154]

[0254] In some embodiments, the mRNA comprises a sequence comprising an Alu element and / or a ribosome-binding aptamer. In some embodiments, the mRNA comprises a sequence encoding a polypeptide comprising a DNA-binding domain. In some embodiments, the 3'UTR sequence is derived from a viral 3'UTR or a beta-globin 3'UTR.

[0155]

[0255] In one aspect, provided herein is a composition comprising a recombinant mRNA or a vector encoding the mRNA, wherein the mRNA comprises a human LINE-1 transposon sequence comprising a human LINE-1 transposon 5'UTR sequence, a sequence encoding ORF1p downstream of the human LINE-1 transposon 5'UTR sequence, an inter-ORF linker sequence downstream of the sequence encoding ORF1p, a sequence encoding ORF2p downstream of the inter-ORF linker sequence, and a 3'UTR sequence derived from the human LINE-1 transposon downstream of the sequence encoding ORF2p, wherein the 3'UTR sequence comprises an insertion sequence, and the insertion sequence is the reverse complement of a sequence encoding a foreign polypeptide or the reverse complement of a sequence encoding a foreign regulatory element.

[0156]

[0256] In some embodiments, when the insertion sequence is introduced into a cell, it is integrated into the genome of the cell.In some embodiments, the insertion sequence is integrated into a gene associated with a condition or disease, thereby disrupting the gene or down-regulating the expression of the gene.In some embodiments, the insertion sequence is integrated into a gene, thereby up-regulating the expression of the gene.In some embodiments, the recombinant mRNA or the vector encoding the mRNA is isolated or purified.

[0157]

[0257] In one aspect, provided herein are compositions comprising a nucleic acid comprising a nucleotide sequence encoding (a) a long interspersed nuclear element (LINE) polypeptide, the LINE polypeptide comprising human ORF1p and human ORF2p, and (b) an insertion sequence, the insertion sequence being the reverse complement of a sequence encoding a foreign polypeptide or a sequence encoding a foreign regulatory element, the composition being substantially non-immunogenic. In some embodiments, integration of the insertion sequence does not adversely affect the health of a cell.

[0158]

[0258] In some embodiments, the composition comprises human ORF1p and human ORF2p proteins. In some embodiments, the composition comprises a ribonucleoprotein (RNP) comprising human ORF1p and human ORF2p complexed with a nucleic acid. In some embodiments, the nucleic acid is mRNA.

[0159]

[0259] In one aspect, provided herein is a composition comprising a cell comprising the composition described herein. In some embodiments, the cell is an immune cell. In some embodiments, the immune cell is a T cell or a B cell. In some embodiments, the immune cell is a bone marrow cell. In some embodiments, the immune cell is selected from the group consisting of monocytes, macrophages, dendritic cells, dendritic precursor cells, and macrophage precursor cells. In some embodiments, the inserted sequence is the reverse complement of the sequence encoding the foreign polypeptide, and the foreign polypeptide is a chimeric antigen receptor (CAR).

[0160]

[0260] In one aspect, provided herein is a pharmaceutical composition comprising a composition described herein and a pharmaceutically acceptable excipient. In some embodiments, the pharmaceutical composition is for use in gene therapy. In some embodiments, the pharmaceutical composition is for use in the manufacture of a medicament for treating a disease or condition. In some embodiments, the pharmaceutical composition is for use in treating a disease or condition. In one aspect, provided herein is a method of treating a disease in a subject, the method comprising administering to a subject having the disease or condition a pharmaceutical composition described herein. In some embodiments, the method increases the amount or activity of a protein or functional RNA in the subject. In some embodiments, the subject has an insufficient amount or activity of the protein or functional RNA. In some embodiments, the insufficient amount or activity of the protein or functional RNA is associated with or causes the disease or condition.

[0161]

[0261] In some embodiments, the method further comprises administering an agent that inhibits the human silencing hub (HUSH) complex, an agent that inhibits FAM208A, or an agent that inhibits TRIM28. In some embodiments, the agent that inhibits the human silencing hub (HUSH) complex is an agent that inhibits periphilin, TASOR, and / or MPP8. In some embodiments, the agent that inhibits the human silencing hub (HUSH) complex inhibits HUSH complex assembly. In some embodiments, the agent inhibits the Fanconi anemia complex. In some embodiments, the agent inhibits FANCD2-FANC1 heterodimer monoubiquitination. In some embodiments, the agent inhibits FANCD2-FANC1 heterodimer formation. In some embodiments, the agent inhibits the Fanconi anemia (FA) core complex. The FA core complex is a component of the Fanconi anemia DNA damage repair pathway, for example, in chemotherapy-induced DNA interstrand crosslinking. The FA core complex contains two major dimers: the FANCB subunit and the 100 kDa FA-associated protein (FAAP100) subunit, flanked by two copies of the FANCL RING finger subunit. These two heterotrimers act as a scaffold to assemble the remaining five subunits, resulting in an extended, asymmetric structure. Destabilization of the scaffold can disrupt the entire complex, resulting in a non-functional FA pathway. Examples of drugs that can inhibit the FA core complex include bortezomib and the curcumin analogs EF24 and 4H-TTD.

[0162]

[0262] Therefore, it is an object of the present invention to provide novel transposon-based vectors useful for providing gene therapy to animals. It is an object of the present invention to provide novel transposon-based vectors for use in preparing medicaments useful for providing gene therapy to animals or humans. Another object of the present invention is to provide novel transposon-based vectors that encode for the production of a desired protein or peptide in a cell. Yet another object of the present invention is to provide novel transposon-based vectors that encode for the production of a desired nucleic acid in a cell. It is a further object of the present invention to provide a method for cell- and tissue-specific integration of a transposon-based DNA or RNA construct, comprising targeting a selected gene to a particular cell or tissue of an animal. It is a further object of the present invention to provide a method for cell- and tissue-specific expression of a transposon-based DNA or RNA construct, comprising designing a DNA or RNA construct with a cell-specific promoter that enhances stable integration of the selected gene by a transposase, and expressing the selected gene in the cell. It is an object of the present invention to provide multi-generational gene therapy through germline administration of transposon-based vectors. It is another object of the present invention to provide gene therapy in animals through non-germline administration of transposon-based vectors. Another object of the present invention is to provide gene therapy in animals by administration of transposon-based vectors, whereby the animal produces a desired protein, peptide, or nucleic acid. Yet another object of the present invention is to provide gene therapy in animals by administration of transposon-based vectors, whereby the animal produces a desired protein or peptide that is recognized by a receptor on a target cell.It is yet another object of the present invention to provide gene therapy in animals by administering a transposon-based vector, wherein the animal produces a desired fusion protein or fusion peptide, a portion of which is recognized by a receptor on the target cell to deliver other protein or peptide components of the fusion protein or fusion peptide to the cell and induce a biological response.It is yet another object of the present invention to provide a method for gene therapy in animals by administering a transposon-based vector containing a tissue-specific promoter and a gene of interest, which facilitates tissue-specific integration and expression of the gene of interest to produce the desired protein, peptide, or nucleic acid.It is another object of the present invention to provide a method for gene therapy in animals by administering a transposon-based vector containing a cell-specific promoter and a gene of interest, which facilitates cell-specific integration and expression of the gene of interest to produce the desired protein, peptide, or nucleic acid. It is yet another object of the present invention to provide a method for gene therapy in animals by administration of a transposon-based vector containing a cell-specific promoter and a gene of interest that facilitates cell-specific integration and expression of the gene of interest to produce a desired protein, peptide, or nucleic acid, wherein the desired protein, peptide, or nucleic acid has a desired biological effect in the animal.

[0163]

[0263] In one aspect, provided herein are methods and compositions for the delivery and stable integration of one or more nucleic acids, including nucleic acid sequences encoding one or more proteins, into cells, such as bone marrow cells, where the stable integration can be achieved via a non-viral mechanism. In some embodiments, the delivery of the nucleic acid composition to the bone marrow cells is achieved via a non-viral mechanism. In some embodiments, the delivery of the nucleic acid can further avoid plasmid-mediated delivery. As used herein, "plasmid" refers to a non-viral expression vector, e.g., a nucleic acid molecule encoding a gene and / or regulatory elements necessary for gene expression. As used herein, "viral vector" refers to a viral-derived nucleic acid capable of transporting another nucleic acid into a cell. When present in the appropriate environment, a viral vector can direct the expression of one or more proteins encoded by one or more genes carried by the vector. Examples of viral vectors include, but are not limited to, retroviral, adenoviral, lentiviral, and adeno-associated viral vectors.

[0164]

[0264] In some embodiments, provided herein are methods for delivering a composition into a cell, such as a bone marrow cell, wherein the composition comprises one or more nucleic acid sequences encoding one or more proteins, and the one or more nucleic acid sequences are RNA. In some embodiments, the RNA is mRNA. In some embodiments, one or more mRNAs comprising one or more nucleic acid sequences are delivered. In some embodiments, the one or more mRNAs may comprise at least one modified nucleotide. The term "nucleotide," as used herein, refers to a base-sugar-phosphate combination. A nucleotide may include synthetic nucleotides. A nucleotide may include synthetic nucleotide analogs. A nucleotide may be a monomeric unit of a nucleic acid sequence (e.g., deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term nucleotide may include ribonucleoside triphosphates adenosine triphosphate (ATP), uridine triphosphate (UTP), cytosine triphosphate (CTP), guanosine triphosphate (GTP), and deoxyribonucleoside triphosphates, such as dATP, dCTP, dITP, dUTP, dGTP, or derivatives thereof. Such derivatives can include, for example, [aS]dATP, 7-deaza-dGTP, and 7-deaza-dATP, as well as nucleotide derivatives that confer nuclease resistance to nucleic acid molecules containing them. The term nucleotide, as used herein, can refer to dideoxyribonucleoside triphosphates (ddNTPs) and their derivatives. Illustrative examples of dideoxyribonucleoside triphosphates can include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP. Nucleotides can be unlabeled or detectably labeled using well-known techniques. Labeling can also be performed using quantum dots. Detectable labels can include, for example, radioisotopes, fluorescent labels, chemiluminescent labels, bioluminescent labels, and enzyme labels.Fluorescent labels for nucleotides include, but are not limited to, fluorescein, 5-carboxyfluorescein (FAM), 2'7'-dimethoxy-4'5-dichloro-6-carboxyfluorescein (JOE), rhodamine, 6-carboxyrhodamine (R6G), N,N,NcN'-tetramethyl-6-carboxyrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), 4-(4'dimethylaminophenylazo)benzoic acid (DABCYL), Cascade Blue, Oregon Green, Texas Red, cyanine, and 5-(2'-aminoethyl)aminonaphthalene-1-sulfonic acid (EDANS).Specific examples of fluorescently labeled nucleotides include [R6G]dUTP, [TAMRA]dUTP, [R110]dCTP, [R6G]dCTP, [TAMRA]dCTP, [JOE]ddATP, [R6G]ddATP, [FAM]ddCTP, [R110]ddCTP, [TAN1RA]ddGTP, [ROX]ddTTP, [dR6G]ddATP, [dR110]ddCTP, [dTAMRA]ddGTP, and [dROX]ddTTP available from Perkin Elmer, Foster City, Calif.; FluoroLink deoxynucleotides, FluoroLink Cy3-dCTP, FluoroLink Cy5-dCTP, FluoroLink Fluor X-dCTP, FluoroLink Cy3-dUTP, and FluoroLink Cy5-dUTP available from Amersham, Arlington Heights, Ill.; and Boehringer Fluorescein-15-dATP, fluorescein-12-dUTP, tetramethyl-rhodamine-6-dUTP, TR770-9-dATP, fluorescein-12-ddUTP, fluorescein-12-UTP, and fluorescein-15-2'-dATP available from Mannheim, Indianapolis, Ind.; and Molecular Examples of chromosome-labeled nucleotides include BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, BODIPY-TMR-14-UTP, BODIPY-TMR-14-dUTP, BODIPY-TR-14-UTP, BODIPY-TR-14-dUTP, Cascade Blue-7-UTP, Cascade Blue-7-dUTP, Fluorescein-12-UTP, Fluorescein-12-dUTP, Oregon Green 488-5-dUTP, Rhodamine Green-5-UTP, Rhodamine Green-5-dUTP, Tetramethylrhodamine-6-UTP, Tetramethylrhodamine-6-dUTP, Texas Red-5-UTP, Texas Red-5-dUTP, and Texas Red-12-dUTP, all available from Probes, Eugene, Oreg. Nucleotides can also be labeled or tagged by chemical modification.Chemically modified single nucleotide can be biotin-dNTP.Some non-limiting examples of biotinylated dNTP can include biotin-dATP (for example, bio-N6-ddATP, biotin-14-dATP), biotin-dCTP (for example, biotin-11-cICTP, biotin-14-dCTP) and biotin-dUTP (for example, biotin-11-dUTP, biotin-1.6-dUTP, biotin-20-dUTP).

[0165]

[0265] The terms "polynucleotide," "oligonucleotide," and "nucleic acid" are used interchangeably to refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or their analogs, in either single-, double-, or multi-stranded form. A polynucleotide may be exogenous or endogenous to a cell. A polynucleotide may be present in a cell-free environment. A polynucleotide may be a gene or a fragment thereof. A polynucleotide may be DNA. A polynucleotide may be RNA. A polynucleotide may have any three-dimensional structure and may perform any function, known or unknown. A polynucleotide may contain one or more analogs (e.g., altered backbones, sugars, or nucleobases). If present, modifications to the nucleotide structure may be imparted before or after assembly of the polymer. Some non-limiting examples of modified nucleotides or analogs include pseudouridine, 5-bromouracil, 5-methylcytosine, peptide nucleic acids, xenonucleic acids, morpholinos, locked nucleic acids, glycol nucleic acids, threose nucleic acids, dideoxynucleotides, cordycepin, 7-deaza-GTP, florophores (e.g., sugar-linked rhodamine or fluorescein), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudourdine, dihydrouridine, queusine, and wyosine.Non-limiting examples of polynucleotide include the coding or non-coding region of gene or gene fragment, the locus defined by linkage analysis, exon, intron, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozyme, eDNA, recombinant polynucleotide, branched polynucleotide, plasmid, vector, isolated DNA of any sequence, isolated RNA of any sequence, cell-free polynucleotide including cell-free DNA (cfDNA) and cell-free RNA (cfRNA), nucleic acid probe and primer.Nucleotide sequence can be separated by non-nucleotide components.

[0166]

[0266] In some embodiments, the nucleic acid composition may comprise one or more mRNAs, including at least one mRNA encoding a transmembrane receptor associated with an immune response function (e.g., a phagocytic receptor or a synthetic chimeric antigen receptor), that enters human macrophages or dendritic cells or suitable myeloid cells or myeloid progenitor cells. In some embodiments, the nucleic acid composition comprises one or more mRNAs and one or more lipids for delivery of the nucleic acid to cells of hematopoietic origin, such as myeloid cells or myeloid progenitor cells. In some embodiments, the one or more lipids may form a liposome complex.

[0167]

[0267] As used herein, the compositions described herein can be used for intracellular delivery. Cells may originate from any organism having one or more cells. Some non-limiting examples include prokaryotic cells, eukaryotic cells, bacterial cells, archaeal cells, unicellular eukaryotic cells, protozoan cells, plant-derived cells (e.g., cells from plant crops, fruits, vegetables, grains, soybeans, corn, maize, wheat, seeds, tomatoes, rice, cassava, sugarcane, pumpkin, hay, potatoes, cotton, hemp, tobacco, angiosperms, conifers, gymnosperms, ferns, club mosses, hornworts, mosses), algae cells (e.g., Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, Sargassum patens, etc.). Examples of cells that may be used include: seaweed (e.g., kelp), fungal cells (e.g., yeast cells, cells from mushrooms), animal cells, cells from invertebrates (e.g., fruit flies, cnidarians, echinoderms, nematodes, etc.), cells from vertebrates (e.g., fish, amphibians, reptiles, birds, mammals), cells from mammals (e.g., pigs, cows, goats, sheep, rodents, rats, mice, non-human primates, humans, etc.). In some cases, the cells may not originate from a natural organism (e.g., the cells may be synthetically produced and may be referred to as artificial cells). In some embodiments, the cells referred to herein are mammalian cells. In some embodiments, the cells are human cells. The methods and compositions described herein relate to incorporating genetic material into cells, more specifically human cells, where the human cells may be any human cells. As used herein, human cells can be cells of any origin, such as somatic cells, neurons, fibroblasts, muscle cells, epithelial cells, cardiac cells, or hematopoietic cells. The methods and compositions described herein can also be applicable and useful for incorporating foreign nucleic acids into difficult-to-transfect human cells. The methods are simple and universally applicable once a suitable foreign nucleic acid construct has been designed and developed.The methods and compositions described herein are applicable to the exogenous nucleic acid incorporation into cells ex vivo. In some embodiments, the compositions may be applicable to systemic administration to an organism, where the nucleic acid material in the composition can be taken up by cells in vivo and then incorporated into cells in vivo.

[0168]

[0268] In some embodiments, the methods and compositions described herein may be directed to incorporating exogenous nucleic acids into human hematopoietic cells, such as human cells of hematopoietic origin, such as human bone marrow cells or bone marrow cell precursors. However, the methods and compositions described herein may be used in any biological cell, or may be suitable for use in any biological cell with minimal modification. Thus, a cell may refer to any cell that is the basic structural, functional, and / or biological unit of a living organism.

[0169]

[0269] In one aspect, provided herein are methods and compositions for utilizing transposable elements for stable integration of one or more nucleic acids into the genome of a cell, wherein the cell is a type of hematopoietic cell, such as a bone marrow cell. In some embodiments, the one or more nucleic acids include at least one nucleic acid sequence encoding a transmembrane receptor protein that plays a role in the immune response. In some embodiments, the methods and compositions are directed to using retrotransposable elements to integrate one or more nucleic acid sequences into bone marrow cells. The nucleic acid composition can include one or more nucleic sequences, such as genes, where the gene is a transgene. The term "gene," as used herein, refers to a nucleic acid (e.g., DNA, such as genomic DNA and cDNA) and the corresponding nucleotide sequence of a nucleic acid that is involved in encoding an RNA transcript. When used herein in reference to genomic DNA, the term includes intervening non-coding regions and regulatory regions, and may include the 5' and 3' ends. In some uses, the term encompasses the transcribed sequence, including the 5' and 3' untranslated regions (5'UTR and 3'UTR), exons, and introns. In some genes, the transcribed region may contain an "open reading frame" that encodes a polypeptide. In some uses of the term, a "gene" includes only the coding sequence (e.g., an "open reading frame" or "coding region") necessary to encode a polypeptide. In some cases, a gene does not encode a polypeptide, such as a ribosomal RNA gene (rRNA) and a transfer RNA (tRNA) gene. In some cases, the term "gene" not only includes the transcribed sequence but also includes non-transcribed regions, including upstream and downstream regulatory regions, enhancers, and promoters. A gene may refer to an "endogenous gene" or a natural gene present in its natural location in the genome of an organism. A gene may also refer to a "foreign gene" or a non-native gene. A non-native gene can refer to a gene not normally found in a host organism and that is introduced into the host organism by gene transfer. A non-native gene may also refer to a gene that is not present in its natural location in the genome of an organism.A non-native gene may also refer to a naturally occurring nucleic acid or polypeptide sequence (eg, a non-natural sequence) that contains mutations, insertions, and / or deletions.

[0170]

[0270] The term "transgene" refers to any nucleic acid molecule introduced into a cell, which may sometimes be referred to herein as a recipient cell. The resulting cell after receiving a transgene can be classified as a transgenic cell. A transgene may include a gene that is partially or completely heterologous (i.e., non-autologous) to the transgenic organism or cell, or may be a gene homologous to an endogenous gene of the organism or cell. In some cases, a transgene includes any polynucleotide, such as a gene encoding a polypeptide or protein, a polynucleotide that is transcribed into an inhibitory polynucleotide, or a polynucleotide that is not transcribed (e.g., lacking an expression control element, such as a promoter, to drive transcription). The transcript and encoded polypeptide may be collectively referred to as a "gene product." If the polynucleotide is derived from genomic DNA, expression may include splicing of mRNA in eukaryotic cells. "Upregulated" in relation to expression refers to an increase in the expression level of a polynucleotide (e.g., RNA such as mRNA) and / or polypeptide sequence compared to the expression level in the wild-type state, and "downregulated" refers to a decrease in the expression level of a polynucleotide (e.g., RNA such as mRNA) and / or polypeptide sequence compared to the expression level in the wild-type state. Expression of a transfected gene can be transient or stable in a cell. During "transient expression," the transfected gene is not transferred to daughter cells during cell division. Expression of the gene is restricted to the transfected cell and is lost over time. In contrast, stable expression of a transfected gene can be achieved when the gene is co-transfected with another gene that confers a selective advantage to the transfected cell. Such a selective advantage can be resistance to a certain toxin presented to the cell. When the transfected gene needs to be expressed, the present application contemplates the use of codon-optimized sequences. An example of a codon-optimized sequence may be a sequence optimized for expression in a eukaryote, such as a human (i.e., optimized for expression in a human), or optimized for another eukaryote, animal, or mammal.Codon optimization for non-human host species or for specific organs is known. In some embodiments, the coding sequence encoding a protein can be codon-optimized for expression in a specific cell, for example, a eukaryotic cell. The eukaryotic cell can be a specific organism, for example, a plant or a mammal, for example, but not limited to, a human, or a non-human eukaryotic organism or animal or mammal discussed herein, for example, a mouse, a rat, a rabbit, a dog, a livestock animal, or a non-human mammal or primate, or can be derived from these specific organisms. Codon optimization refers to the process of modifying a nucleic acid sequence while maintaining the natural amino acid sequence by replacing at least one codon (for example, about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50, or more than about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more codons) of the natural sequence with a codon that is more frequently or most frequently used in the gene of the host cell, for enhanced expression in the target host cell. Various species exhibit specific biases toward certain codons for specific amino acids. Codon bias (differences in codon usage among organisms) often correlates with the efficiency of messenger RNA (mRNA) translation, which is thought to depend, among other things, on the characteristics of the codon being translated and the availability of specific transfer RNA (tRNA) molecules. The cellular dominance of selected tRNAs may generally reflect the codons most frequently used in peptide synthesis. Thus, genes can be tailored for optimal gene expression in a given organism based on codon optimization. Codon usage tables are readily available, for example, in the "Codon Usage Database" available at www.kazusa.orjp / codon / , and these tables may be modified in some respects. Computer algorithms for codon-optimizing specific sequences for expression in specific host cells are also available, for example, Gene Forge (Aptagen, Jacobus, PA).

[0171]

[0271] As used herein, a "multicistronic transcript" refers to an mRNA molecule containing two or more protein coding regions or cistrons. An mRNA containing two coding regions is designated a "bicistronic transcript." A "5'-proximal" coding region or cistron is one in which the translation initiation codon (usually AUG) is closest to the 5' end of a multicistronic mRNA molecule. A "5'-distal" coding region or cistron is one in which the translation initiation codon (usually AUG) is not the closest initiation codon to the 5' end of the mRNA.

[0172]

[0272] The term "transfection" or "transfected" refers to the introduction of nucleic acid into a cell by non-viral or viral-based methods. The nucleic acid molecule may be a gene sequence encoding an entire protein or a functional portion thereof. See, e.g., Sambrook et al., 1989, Molecular Cloning: A Laboratory Manual, 18.1-18.88.

[0173]

[0273] The term "promoter," as used herein, refers to a polynucleotide sequence capable of driving transcription of a coding sequence in a cell. Thus, promoters used in the polynucleotide constructs of the present disclosure include cis-acting transcriptional control elements and regulatory sequences involved in regulating or modulating the timing and / or rate of gene transcription. For example, a promoter can be a cis-acting transcriptional control element, including an enhancer, promoter, transcription terminator, origin of replication, chromosomal integration sequence, 5' and 3' untranslated region, or intron sequence, which are involved in transcriptional regulation. These cis-acting sequences typically interact with proteins or other biomolecules that execute (turn on / off, regulate, modulate, etc.) gene transcription. A "constitutive promoter" is a promoter that can initiate transcription in almost all tissue types, while a "tissue-specific promoter" initiates transcription only in one or a few specific tissue types. An "inducible promoter" is a promoter that initiates transcription only under specific environmental, developmental, or drug or chemical conditions. An exemplary inducible promoter can be a doxycycline- or tetracycline-inducible promoter. Tetracycline-regulated promoters can be either tetracycline-inducible or tetracycline-repressible, referred to as the tet-on system and tet-off system. The tet regulatory system relies on two components: a tetracycline-controlled regulator (also called a transactivator) (tTA or rtTA) and a tTA / rtTA-dependent promoter that controls downstream cDNA expression in a tetracycline-dependent manner. tTA is a fusion protein containing the repressor of the Escherichia coli Tn10 tetracycline resistance operon and the carboxyl-terminal portion of herpes simplex virus protein 16 (VP16). The tTA-dependent promoter consists of a minimal RNA polymerase II promoter fused to the tet operator (tetO) sequence (a series of seven homologous operator sequences). This fusion converts the tet repressor into a potent transcriptional activator in eukaryotic cells.In the absence of tetracycline or its derivatives (e.g., doxycycline), tTA binds to the tetO sequence, allowing transcriptional activation of tTA-dependent promoters. However, in the presence of doxycycline, tTA cannot interact with its target, and transcription does not occur. Tet systems using tTA are called tet-OFF because tetracycline or doxycycline allows downregulation of transcription. In contrast, in the tet-ON system, a mutant form of tTA called rtTA has been isolated using random mutagenesis. In contrast to tTA, rtTA is not functional in the absence of doxycycline and requires the presence of a ligand for transactivation. The term "exon" refers to a nucleic acid sequence found in genomic DNA that is bioinformatically predicted and / or experimentally confirmed to contribute to a contiguous sequence in a mature mRNA transcript. The term "intron" refers to a sequence present in genomic DNA that is bioinformatically predicted and / or experimentally confirmed not to encode part or all of an expressed protein, that is endogenously transcribed into an RNA (e.g., pre-mRNA) molecule, but that is spliced ​​out of the endogenous RNA (e.g., pre-mRNA) before the RNA is translated into a protein.

[0174]

[0274] The term "splice acceptor site" refers to sequences present in genomic DNA that are bioinformatically predicted and / or experimentally confirmed to be acceptor sites during splicing of pre-mRNA, and can include identified and unidentified, naturally occurring and artificially obtained or obtainable splice acceptor sites.

[0175]

[0275] An "internal ribosome entry site" or "IRES" refers to a nucleotide sequence that allows 5'-end / cap-independent translation initiation, thereby increasing the possibility of expressing two proteins from a single messenger RNA (mRNA) molecule. IRESs are typically located in the 5'UTR of positive-strand RNA viruses with uncapped genomes. Another means of expressing two proteins from a single mRNA molecule is by inserting a 2A peptide(-like) sequence between the coding sequences of those proteins. 2A peptide(-like) sequences mediate self-processing of the primary translation product by a process variously referred to as "ribosome skipping," "stop-go" translation, and "stop-carry-on" translation. 2A peptide(-like) sequences are present in various groups of positive-strand and double-stranded RNA viruses, including Picornaviridae, Flaviviridae, Tetraviridae, Dicistroviridae, Reoviridae, and Totiviridae.

[0176]

[0276] The term "2A peptide" refers to a class of viral oligopeptides, 18–22 amino acids (AA) long, that mediate the "cleavage" of polypeptides during translation in eukaryotic cells. The "2A" designation refers to a specific region of the viral genome, and various viral 2As are generally named after the viruses from which they are derived. The first 2A discovered was F2A (foot-and-mouth disease virus), followed by E2A (equine rhinitis A virus), P2A (porcine teschovirus-1 2A), and T2A (thosea asigna virus 2A). The mechanism of 2A-mediated "self-cleavage" is thought to involve ribosomal skipping, which results in the formation of a glycyl-prolyl peptide bond at the C-terminus of the 2A sequence. 2A peptide(-like) sequences mediate the self-processing of primary translation products by processes variously referred to as "ribosomal skipping," "stop-go" translation, and "stop-carry-on" translation. 2A peptide(-like) sequences are present in various groups of positive-stranded and double-stranded RNA viruses, including Picornaviridae, Flaviviridae, Tetraviridae, Dicistroviridae, Reoviridae, and Totiviridae.

[0177]

[0277] As used herein, the term "operably linked" refers to a functional relationship between two or more segments, e.g., nucleic acid segments or polypeptide segments. Typically, the term refers to the functional relationship of a transcriptional regulatory sequence to a transcribed sequence.

[0178]

[0278] The term "termination sequence" refers to a nucleic acid sequence recognized by a host cell polymerase and resulting in the termination of transcription. A termination sequence is a DNA sequence at the 3' end of a natural or synthetic gene that effects the termination of mRNA transcription or both mRNA transcription and ribosomal translation of an upstream open reading frame. Prokaryotic termination sequences generally contain a GC-rich region with dyad symmetry followed by an AT-rich sequence. A commonly used termination sequence is the T7 termination sequence. A variety of termination sequences are known in the art and can be used in the nucleic acid constructs of the present invention, including the TINT3, TL13, TL2, TR1, TR2, and T6S termination signals from bacteriophage lambda, as well as termination signals from bacterial genes such as the trp gene of Escherichia coli.

[0179]

[0279] The term "polyadenylation sequence" (also referred to as "poly A site" or "poly A sequence") refers to a DNA sequence that directs both the termination and polyadenylation of nascent RNA transcripts. Efficient polyadenylation of recombinant transcripts is desirable because transcripts lacking a poly A tail are typically unstable and rapidly degraded. The poly A signal utilized in an expression vector may be "heterologous" or "endogenous." An endogenous poly A signal is one that is naturally found at the 3' end of the coding region of a given gene in the genome. A heterologous poly A signal is one that is isolated from one gene and placed 3' to the coding sequence of another gene, e.g., a protein. A commonly used heterologous poly A signal is the SV40 poly A signal. The SV40 poly A signal is contained in a 237-bp BamHI / BclI restriction fragment and directs both termination and polyadenylation; many vectors contain the SV40 poly A signal. Another commonly used heterologous poly(A) signal is derived from the bovine growth hormone (BGH) gene, and the BGH poly(A) signal is also available in some commercially available vectors. The poly(A) signal derived from the herpes simplex virus thymidine kinase (HSV tk) gene is also used as the poly(A) signal in some commercial expression vectors. Polyadenylation signals facilitate the transport of RNA from the cell nucleus to the cytosol, increasing the cellular half-life of such RNA. Polyadenylation signals are present at the 3' end of mRNA.

[0180]

[0280] The terms "complement," "complements," "complementary," and "complementarity," as used herein, refer to a sequence that is complementary to and hybridizable with a given sequence. In some cases, a sequence that hybridizes with a given nucleic acid is referred to as the "complement" or "reverse complement" of a given molecule if the sequence of bases across a given region can complementarily bind with the sequence of the binding partner, for example, so that AT, AU, GC, and GU base pairs are formed. Generally, a first sequence that can hybridize to a second sequence can specifically or selectively hybridize to the second sequence, such that hybridization to the second sequence or set of second sequences is more favorable (e.g., thermodynamically more stable under a given set of conditions, such as stringent conditions commonly used in the art) than hybridization to non-target sequences during a hybridization reaction. Typically, hybridizable sequences share some degree of sequence complementarity over all or part of their respective lengths, for example, between 25% and 100%, including at least 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, and 100% sequence complementarity.Sequence identity, for example, for purposes of assessing percent complementarity, can be measured by any suitable alignment algorithm, including, but not limited to, the Needleman-Wunsch algorithm (see, for example, the EMBOSS Needle aligner available at www.ebi.ac.uk / Tools / psa / embossneedle / nucleotide.html), the BLAST algorithm (see, for example, the BLAST alignment tool available at blast.ncbi.nlm.nih.gov / Blast.cgi, optionally with default settings), or the Smith-Waterman algorithm (see, for example, the EMBOSS Water aligner available at www.ebi.ac.ukaools / psa / emboss_water / nucleotide.html, optionally with default settings). Optimal alignment can be assessed using any suitable parameters of the selected algorithm, including the default parameters.

[0181]

[0281] Complementarity can be perfect or substantial / sufficient. Perfect complementarity between two nucleic acids can mean that the two nucleic acids can form a duplex in which every base is bound to a complementary base by Watson-Crick pairing. Substantial or sufficient complementarity can mean that the sequence in one strand is not completely and / or perfectly complementary to the sequence in the opposing strand, but that sufficient binding occurs between the bases of the two strands to form a stable hybrid complex under a set of hybridization conditions (e.g., salt concentration and temperature). Such conditions can be determined by calculating the melting temperature (T) of the hybridized strands using the sequences and standard mathematical calculations. m ) or by using conventional methods. m can be predicted by empirical determination of

[0182]

[0282] As used herein, a "transposon," also known as a "jumping gene," is a segment within a chromosome that can change position within a genome. Two distinct classes of transposons exist: Class 1 or retrotransposons, which move via an RNA intermediate and a "copy-and-paste" mechanism, and Class II or DNA transposons, which move via an excision-integration or "cut-and-paste" mechanism (Ivics Nat Methods 2009). Bacterial, lower eukaryotic (e.g., yeast), and invertebrate transposons are largely species-specific and are thought to be unable to efficiently transfer DNA in vertebrate cells. "Sleeping Beauty" (Ivics Cell 1997) was the first active transposon artificially reconstructed by sequence shuffling of an inactive TE derived from fish. This allowed successful DNA integration by transposition into vertebrate cells, including human cells. Sleeping Beauty is a Class II DNA transposon belonging to the Tcl / Mariner family of transposons (Ni Genomics Proteomics 2008). On the other hand, additional functional transposons have been identified or reconstructed from different species, including Drosophila, frog, and even human genomes, and all of them have been shown to be capable of DNA transfer into vertebrate and even human host cell genomes.Each of these transposons has advantages and disadvantages in terms of transfer efficiency, expression stability, gene payload capacity, etc.Exemplary class II transposases that have been produced include Sleeping Beauty, PiggyBac, Frog Prince, Himarl, Passport, Minos, hAT, Toll, Tol2, Acids, PIF, Harbinger, Harbinger3-DR, and Hsmarl.

[0183]

[0283] As used herein, "heterologous" includes molecules such as DNA and RNA that are not naturally found in the cell into which they are inserted. For example, when mouse or bacterial DNA is inserted into the genome of a human cell, such DNA is referred to herein as heterologous DNA. In contrast, the term "homologous" as used herein refers to molecules such as DNA and RNA that are naturally found in the cell into which they are inserted. For example, the insertion of mouse DNA into the genome of a mouse cell constitutes the insertion of homologous DNA into the cell. In the latter case, it is not necessary for the homologous DNA to be inserted into the site of the cell genome where it is naturally found; rather, the homologous DNA may be inserted into a site other than the site where it is naturally found, thereby causing genetic alterations (mutations) at the inserted site.

[0184]

[0284] A "transposase" is an enzyme that can form a functional complex with a transposon end-containing composition (e.g., a transposon, a transposon end) and catalyze the insertion or transposition of the transposon end-containing composition into double-stranded DNA that is incubated for an in vitro transposon reaction. The term "transposon end" refers to double-stranded DNA that contains the nucleotide sequences ("transposon end sequences") necessary to form a complex with a transposase or integrase enzyme that is functional in an in vitro transposition reaction.

[0185]

[0285] The transposon ends form a complex, or synaptic complex, or transposon complex, or transposon composition, with a transposase or integrase that recognizes and binds to the transposon ends, and these complexes can insert or transfer the transposon ends into target DNA incubated together in an in vitro transposition reaction. The transposon ends exhibit two complementary sequences: an imported transposon end sequence or imported strand and a non-imported transposon end sequence or non-imported strand. For example, a transposon end complexed with a hyperactive Tn5 transposase active in an in vitro transposition reaction includes an imported strand exhibiting the following imported transposon end sequence: 5'AGATGTGTATAAGAGACAG3', and a non-imported strand exhibiting the following "non-imported transposon end sequence": 5'CTGTCTCTTATACACATCT3'. The 3' end of the imported strand binds to or is imported into target DNA in an in vitro transposition reaction. The non-imported strand exhibiting a transposon end sequence complementary to the imported transposon end sequence does not bind to or is not imported into target DNA in an in vitro transposition reaction.

[0186]

[0286] In some embodiments, the import strand and the non-import strand are covalently linked. For example, in some embodiments, the import and non-import strand sequences are provided in a single oligonucleotide, for example, in a hairpin configuration. Thus, the free end of the non-import strand does not directly bind to the target DNA through the transposition reaction, but the non-import strand is indirectly linked to the DNA fragment by being linked to the import strand through the loop of the hairpin structure. As used herein, "cleavage domain" refers to a nucleic acid sequence that is sensitive to cleavage by an agent, for example, an enzyme.

[0187]

[0287] "Restriction site domain" refers to a tag domain that exhibits a sequence intended to facilitate cleavage using a restriction endonuclease. For example, in some embodiments, a restriction site domain is used to generate two tagged linear ssDNA fragments. In some embodiments, a restriction site domain is used to generate a compatible double-stranded 5' end in the tag domain so that the compatible double-stranded 5' end can be ligated to another DNA molecule using a template-dependent DNA ligase. In some embodiments, the restriction site domain in the tag exhibits the sequence of a restriction site that is rarely, if ever, present in the target DNA (e.g., a restriction site for a rare-cutting restriction endonuclease such as NotI or AscI).

[0188]

[0288] As used herein, the term "recombinant nucleic acid molecule" refers to a recombinant DNA molecule or a recombinant RNA molecule. A recombinant nucleic acid molecule is any nucleic acid molecule containing a combination of nucleic acid molecules from different original sources that are not naturally associated together. Recombinant RNA molecules include RNA molecules transcribed from recombinant DNA molecules. Recombinant nucleic acids can be synthesized in the laboratory. Recombinant nucleic acids can be prepared using recombinant DNA techniques by using enzymatic modification of DNA, such as enzyme restriction digestion, ligation, and DNA cloning. Recombinant DNA can be transcribed in vitro to produce messenger RNA (mRNA), which can be isolated, purified, and used to transfect cells. Recombinant nucleic acids can encode proteins or polypeptides. Under appropriate conditions, recombinant nucleic acids can be incorporated into and expressed in living cells. As used herein, "expression" of a nucleic acid typically refers to the transcription and / or translation of the nucleic acid. The product of nucleic acid expression is typically a protein but can also be mRNA. Detection of mRNA encoded by a recombinant nucleic acid in a cell that has incorporated the recombinant nucleic acid is considered clear evidence that the nucleic acid is "expressed" in the cell. The process of inserting or incorporating a nucleic acid into a cell can be via transformation, transfection, or transduction. Transformation is the process of uptake of non-self nucleic acids by bacterial cells. This process has been modified for plasmid DNA amplification, protein production, and other applications. Transformation involves introducing recombinant plasmid DNA into competent bacterial cells that take up extracellular DNA from the environment. Some bacterial species are naturally competent under certain environmental conditions, but competence can be artificially induced in laboratory settings. Transfection is the forced introduction of small molecules, such as DNA, RNA, or antibodies, into eukaryotic cells. Simply complicating the situation, "transfection" also refers to the introduction of bacteriophage into bacterial cells."Infection" refers to the natural infection of a human or animal with a wild-type virus, whereas "transduction" is most often used to describe the introduction of recombinant viral vector particles into target cells.

[0189]

[0289] A "stem-loop" sequence refers to a nucleic acid sequence (e.g., an RNA sequence) that has sufficient self-complementarity to hybridize to form a stem and a region of non-complementarity that bulges out into a loop. The stem may contain mismatches or bulges.

[0190]

[0290] The term "vector" refers to a nucleic acid molecule capable of transporting or mediating the expression of heterologous nucleic acid. As used herein, a "vector sequence" refers to a nucleic acid sequence comprising at least one origin of replication and at least one selectable marker gene. A vector capable of directing the expression of an operatively linked gene and / or nucleic acid sequence is referred to herein as an "expression vector."

[0191]

[0291] A plasmid is a species of the genus encompassed by the term "vector." Generally, useful expression vectors often exist in the form of "plasmids," which refer to circular double-stranded DNA molecules that are not bound to a chromosome in vector form and typically contain components for stable or transient expression of encoded DNA. Other expression vectors that can be used in the methods disclosed herein include, but are not limited to, plasmids, episomes, bacterial artificial chromosomes, yeast artificial chromosomes, bacteriophages, or viral vectors, which can be integrated into the host genome or replicate autonomously in cells. Vectors can be DNA or RNA vectors. Other forms of expression vectors known by those skilled in the art that perform equivalent functions, such as self-replicating extrachromosomal vectors or vectors that can integrate into the host genome, can also be used. Exemplary vectors are vectors capable of autonomous replication and / or expression of linked nucleic acids. A safe harbor locus is a region within a genome into which additional foreign or heterologous nucleic acid sequences can be inserted and into which the host genome can accommodate the inserted genetic material. Exemplary safe harbor sites include, but are not limited to, AAVS1 site, GGTA1 site, CMAH site, B4GALNT2 site, B2M site, ROSA26 site, COLA1 site, and TIGRE site. For example, the heterologous nucleic acid described in the present disclosure can be integrated into one or more sites in the genome of a cell, wherein one or more positions are selected from the group consisting of AAVS1 site, GGTA1 site, CMAH site, B4GALNT2 site, B2M site, ROSA26 site, COLA1 site, and TIGRE site. In some embodiments, the nucleic acid cargo containing the transgene can be delivered to the R2D locus.

[0192]

[0292] In some embodiments, the nucleic acid cargo comprising a transgene can be delivered to the genome in an intergenic or intragenic region. In some embodiments, the nucleic acid cargo comprising a transgene is integrated into the genome within 0.1 kb, 0.25 kb, 0.5 kb, 0.75 kb, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 7.5 kb, 10 kb, 15 kb, 20 kb, 25 kb, 50 kb, 75 kb, or 100 kb 5' or 3' of the endogenous active gene. In some embodiments, the nucleic acid cargo comprising the transgene is integrated into the genome within 0.1 kb, 0.25 kb, 0.5 kb, 0.75 kb, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 7.5 kb, 10 kb, 15 kb, 20 kb, 25 kb, 50, 75 kb, or 100 kb 5' or 3' of the endogenous promoter or enhancer. In some embodiments, the nucleic acid cargo comprising the transgene is 50 to 50,000 base pairs, e.g., between 50 and 40,000 bp, between 500 and 30,000 bp, between 500 and 20,000 bp, between 100 and 15,000 bp, between 500 and 10,000 bp, between 50 and 10,000 bp, or between 50 and 5,000 bp. In some embodiments, the nucleic acid cargo comprising the transgene is less than 1,000, 1,300, 1500, 2,000, 3,000, 4,000, 5,000, or 7,500 nucleotides in length.

[0193] L1 and non-L1 retrotransposon systems

[0293] Retrotransposons can contain transposable elements, which are active participants in reorganizing the native genome. Broadly speaking, retrotransposons can refer to DNA sequences that have the ability to be transcribed into RNA, translated into protein, and reverse transcribe themselves into DNA. Approximately 45% of the human genome is composed of sequences resulting from transposition events. Retrotransposition can result in target site deletions or add non-retrotransposon DNA to the genome through a process called 5' and 3' transduction. Recombination between non-homologous retrotransposons can result in deletion, duplication, or rearrangement of gene sequences. Ongoing retrotransposition can generate novel splice sites, polyadenylation signals, and promoters, thus constructing new transcriptional modules.

[0194]

[0294] In general, retrotransposons can be grouped into two classes: retrovirus-like long-term repeat (LTR) retrotransposons and non-long-term repeat (LTR) elements such as human L1 elements, Neurospora TAD elements (Kinsey, 1990, Genetics 126:317-326), I elements from Drosophila (Bucheton et al., 1984, Cell 38:153-163), and R2Bm from the silkworm (Bombyx mori) (Luan et al., 1993, Cell 72:595-605). These two types of retrotransposons are structurally distinct and use fundamentally different mechanisms for retrotransposition. Illustrative, non-limiting examples of LINE-encoded polypeptides are found in GenBank Accession Nos. AAC51261, AAC51262, AAC51263, AAC51264, AAC51265, AAC51266, AAC51267, AAC51268, AAC51269, AAC51270, AAC51271, AAC51272, AAC51273, AAC51274, AAC51275, AAC51276, AAC51277, AAC51278, and AAC51279.

[0195]

[0295] The decision to focus on LINE-1 and develop the system described in this disclosure was due to several reasons, exemplified at least in part as follows: (a) LINE-1 (or L1) elements are autonomous because they solely encode all of the machinery to complete this reverse transcription and integration process; (b) L1 elements are sufficiently abundant in the human genome that they can be considered naturalized elements of the genome; and (c) L1 retrotransposons retrotranspose their own mRNA with a high degree of specificity relative to other mRNAs floating around in the cell.

[0196]

[0296] L1 expresses a 6-kb bicistronic RNA encoding a 40-kDa open reading frame-1 RNA-binding protein (ORF1p), which has essential but uncertain functions, and a 150-kDa ORF2 protein with endonuclease and reverse transcriptase (RT) activities. L1 retrotransposition is a complex process involving transcription of L1, transport of the RNA to the cytoplasm, translation of the bicistronic RNA, formation of a ribonucleoprotein (RNP) particle, its retranslocation to the nucleus, and target-primed reverse transcription at the integration site. Several transcription factors that interact with L1 have been identified. The transcribed L1 RNA forms RNPs in cis with the proteins translated from the transcript. L1 is integrated into genomic DNA by target-site-primed reverse transcription (TPRT) via ORF2p cleavage at 5'-TTTT-3', where the poly(A) sequence of L1 RNA anneals and stimulates reverse transcriptase (RT) activity to generate L1 cDNA.

[0197]

[0297] Other mobile elements in the genome can "hijack" the L1 ORF for retrotransposition. For example, Alu elements are non-autonomous retrotransposons, mobile DNA elements belonging to the class of short interspersed nuclear elements (SINEs) that acquire trans-factors for integration. Alu and SINE-1 elements can also associate in trans with the L1 ribonucleoprotein and be retrotransposed by ORF1p and ORF2p. Similar to L1 RNA, Alu elements often terminate with a long stretch of A, often referred to as an A tail, and also have a smaller A-rich region (designated AA) that bisects the branched dimeric structure. Alu elements likely contain internal components of RNA polymerase III promoters (e.g., commonly referred to as A-box and B-box promoters), but do not encode RNA polymerase III termination factors. Alu elements can terminate transcription using stretches of T nucleotides located at various distances downstream from the Alu element. A typical Alu transcript encompasses the entire Alu, including the A-tail, and has a 3' region unique to each locus. Alu RNA folds into a distinct structure for each monomeric unit. The RNA has been shown to bind to the 7SL RNA SRP9 and 14 heterodimers and poly(A)-binding protein (PABP). The Alu poly(A) tail primes a T-rich (TTTT) region of the genome, attracting ORF2p to bind to the primer region and cleaving at the T-rich region via ORF2p's endonuclease activity. The T-rich region primes reverse transcription by ORF2p binding to the 3' A-tail region of the Alu element. This creates a cDNA copy of the Alu element's body. A nick is generated in the second strand by an unknown mechanism, priming second-strand synthesis. The new Alu element is then flanked by short direct repeats, a duplication of the DNA sequence between the first and second nicks. Alu elements are highly prevalent within RNA molecules due to their preference for gene-rich regions.Full-length Alu (approximately 300 bp) is derived from the signal recognition particle RNA 7SL and consists of two similar monomers with an A-rich linker between them, the A and B boxes present in the 5' monomer, and a polyA tail lacking the previously mentioned polyadenylation signal, resulting in an extended tail (up to 100 bp long). Alu can be transcribed by RNA polymerase III using internal promoters within the A and B boxes, but does not contain an ORF and therefore does not encode a protein product.

[0198]

[0298] Other non-L1 transposons include SVA and HERV-K. The full-length SVA (SINE-VNTR-Alu) element (approximately 2-3 kb) is a synthetic unit containing a CCCTCT repeat sequence, two Alu-like sequences, a VNTR, a SINE-R region containing the env (envelope) gene, the 3' LTR of HERV-K10, and a polyadenylation signal followed by a poly(A) tail. It is unknown whether the SVA element possesses an internal promoter, but SVA is most likely transcribed by RNA polymerase II.

[0199]

[0299] The full-length HERV-K element (approximately 9-10 kb) consists of an ancient relic of endogenous retroviral sequences and contains two flanking long terminal repeat (LTR) regions flanking three retroviral ORFs: (1) gag, which encodes the structural proteins of the retroviral capsid; (2) pol-pro, which encodes the enzymes protease, RT, and integrase; and (3) env, which encodes the proteins that enable horizontal transmission. The HERV-K long terminal repeat (LTR) contains an internal bidirectional promoter that is thought to be under the transcriptional control of RNA polymerase II.

[0200]

[0300] L1 retrotransposition and RNA binding may occur at or near the polyA tail. The 3'UTR plays a role in the recognition of stringent LINE RNA by ORF1 protein (ORF1p). Stringent LINEs can contain stem-loop structures located at the end of the 3'UTR. Branched molecules consisting of the junction of the transposon 3' end cDNA with the target DNA and the specific positioning of L1 RNA within ORF2 protein (ORF2p) were detected during the early stages of L1 retrotransposition in vitro. Secondary or tertiary RNA structures shared by L1 and Alu, possibly together with the polyA tail, are likely responsible for recognition and binding by ORF2. In some embodiments, the stem-loop structure located downstream of the polyA sequence correlates with cleavage strength.

[0201]

[0301] Mechanisms to restrict or resolve L1 integration have also evolved to maintain the genetic integrity and stability of the genome. Nonhomologous end-joining repair proteins, such as XRCC1, Ku70, and DNA-PK, are involved in the resolution of integrated L1s during insertion. In addition, cells have evolved several proteins to counteract uncontrolled retrotransposition, including members of the APOBEC3 family of cytosine deaminases, the adenosine deaminase ADAR1, chromatin remodeling factors, and the piRNA pathway for post-transcriptional gene silencing that functions in the male germline.

[0202] I. Compositions and methods comprising nucleic acid constructs involved in stable expression of encoded proteins

[0302] Provided herein is a recombinant nucleic acid encoding one or more proteins for expression in cells such as bone marrow cells. In one embodiment, the recombinant nucleic acid is designed for stable expression of one or more proteins or polypeptides encoded by the recombinant nucleic acid. In some embodiments, stable expression is achieved by integrating the recombinant nucleic acid into the genome of the cell.

[0203]

[0303] Those skilled in the art will readily appreciate that the compositions and methods described herein can be utilized to design products in which a recombinant nucleic acid is not translated as a protein or polypeptide component, but may contain one or more sequences that may encode an oligonucleotide that may be a regulatory nucleic acid, e.g., an inhibitory oligonucleotide product, e.g., an activator oligonucleotide.

[0204]

[0304] In one aspect, provided herein is a composition comprising a synthetic nucleic acid comprising a nucleic acid sequence encoding a gene of interest and one or more retrotransposable elements for stably integrating non-endogenous nucleic acid into cells.In some embodiments, the cell is a hematopoietic cell.In some embodiments, the cell is a bone marrow cell.In some embodiments, the cell is a progenitor cell.In some embodiments, the cell is undifferentiated.In some embodiments, the cell has the potential for further differentiation.In some embodiments, the cell is not a stem cell.

[0205] A.LINE / Alu retrotransposon constructs

[0305] In some embodiments, the present disclosure may utilize a retrotransposition system for stably integrating and expressing a non-endogenous nucleic acid into a genome, wherein the non-endogenous nucleic acid comprises a retrotransposition element within the nucleic acid sequence. In some embodiments, the present disclosure may utilize a cell's endogenous retrotransposition system (e.g., proteins and enzymes) for stably expressing the non-endogenous nucleic acid in the cell. In some embodiments, the present disclosure may utilize a cell's endogenous retrotransposition system (e.g., proteins and enzymes, e.g., the LINE1 retrotransposition system), but may further express one or more components of the retrotransposition system for stably expressing the non-endogenous nucleic acid in the cell.

[0206]

[0306] In some embodiments, provided herein is a synthetic nucleic acid that encodes a transgene and encodes one or more components for retrotransposition. The synthetic nucleic acids described herein are interchangeably referred to as nucleic acid constructs, transgenes, or foreign nucleic acids.

[0207]

[0307]

[0010] In one aspect, provided herein is a method for integrating a nucleic acid sequence into the genome of a cell, comprising the step of introducing a recombinant mRNA or a vector encoding the mRNA into a cell, wherein the mRNA comprises an insert sequence comprising the foreign sequence, or a sequence that is the reverse complement of the foreign sequence; a 5'UTR sequence and a 3'UTR sequence downstream of the 5'UTR sequence, wherein the 5'UTR sequence or the 3'UTR sequence comprises a binding site for a human ORF protein; and wherein the insert sequence is integrated into the genome of the cell.

[0208]

[0308] In some embodiments, the 5'UTR sequence or the 3'UTR sequence comprises a binding site for human ORF2p.

[0209]

[0309]

[0010] In one aspect, provided herein is a method for integrating a nucleic acid sequence into the genome of an immune cell, comprising introducing a recombinant mRNA or a vector encoding the mRNA, wherein the mRNA comprises an insertion sequence, the insertion sequence comprising (i) the foreign sequence or (ii) a sequence that is the reverse complement of the foreign sequence; a 5'UTR sequence and a 3'UTR sequence downstream of the 5'UTR sequence, wherein the 5'UTR sequence or the 3'UTR sequence comprises an endonuclease binding site and / or a reverse transcriptase binding site, and wherein the transgene sequence is integrated into the genome of the immune cell.

[0210]

[0310]

[0010] In one aspect, provided herein is a method for integrating a nucleic acid sequence into the genome of a cell, comprising the step of introducing a recombinant mRNA or a vector encoding the mRNA, wherein the mRNA comprises an insertion sequence, the insertion sequence comprising (i) a foreign sequence or (ii) a sequence that is the reverse complement of the foreign sequence; a 5'UTR sequence, a sequence of a human retrotransposon downstream of the 5'UTR sequence, and a 3'UTR sequence downstream of the human retrotransposon sequence, wherein the 5'UTR sequence or the 3'UTR sequence comprises an endonuclease binding site and / or a reverse transcriptase binding site, and the human retrotransposon sequence encodes two proteins that are translated from a single RNA containing two ORFs, and wherein the insertion sequence is integrated into the genome of the cell.

[0211]

[0311] In some embodiments, the 5'UTR sequence or the 3'UTR sequence comprises an ORF2p binding site. In some embodiments, the ORF2p binding site is a polyA sequence in the 3'UTR sequence.

[0212]

[0312] In some embodiments, the mRNA comprises a human retrotransposon sequence. In some embodiments, the human retrotransposon sequence is downstream of the 5'UTR sequence. In some embodiments, the human retrotransposon sequence is upstream of the 3'UTR sequence.

[0213]

[0313] In some embodiments, the sequence of human retrotransposon encodes two proteins that are translated from a single RNA containing two ORFs.In some embodiments, the two ORFs are non-overlapping ORFs.In some embodiments, the two ORFs are ORF1 and ORF2.In some embodiments, ORF1 encodes ORF1p, and ORF2 encodes ORF2p.

[0214]

[0314] In some embodiments, the sequence of the human retrotransposon comprises a sequence of a non-LTR retrotransposon. In some embodiments, the sequence of the human retrotransposon comprises a LINE-1 retrotransposon. In some embodiments, the LINE-1 retrotransposon is a human LINE-1 retrotransposon. In some embodiments, the sequence of the human retrotransposon comprises a sequence encoding an endonuclease and / or reverse transcriptase. In some embodiments, the endonuclease and / or reverse transcriptase is ORF2p. In some embodiments, the reverse transcriptase is a group II intron reverse transcriptase domain. In some embodiments, the endonuclease and / or reverse transcriptase is a minke whale endonuclease and / or reverse transcriptase. In some embodiments, the sequence of the human retrotransposon comprises a sequence encoding ORF2p. In some embodiments, the insertion sequence is integrated into the genome at a poly-T site using the specificity of the endonuclease domain of ORF2p. In some embodiments, the poly-T site comprises the sequence TTTTTA.

[0215]

[0315] In some embodiments, (i) the sequence of the human retrotransposon includes a sequence encoding ORF1p, (ii) the mRNA does not include a sequence encoding ORF1p, or (iii) the mRNA includes a replacement sequence for the sequence encoding ORF1p with a 5'UTR sequence from a complementary gene. In some embodiments, the mRNA includes a first mRNA molecule encoding ORF1p and a second mRNA molecule encoding an endonuclease and / or reverse transcriptase. In some embodiments, the mRNA is an mRNA molecule including a first sequence encoding ORF1p and a second sequence encoding an endonuclease and / or reverse transcriptase. In some embodiments, the first sequence encoding ORF1p and the second sequence encoding the endonuclease and / or reverse transcriptase are separated by a linker sequence.

[0216]

[0316] In some embodiments, the linker sequence comprises an internal ribosome entry sequence (IRES). In some embodiments, the IRES is an IRES from CVB3 or EV71. In some embodiments, the linker sequence encodes a self-cleaving peptide sequence. In some embodiments, the linker sequence encodes a T2A, E2A, or P2A sequence.

[0217]

[0317] In some embodiments, the sequence of the human retrotransposon comprises a sequence encoding ORF1p fused to an additional protein sequence and / or a sequence encoding ORF2p fused to an additional protein sequence. In some embodiments, ORF1p and / or ORF2p are fused to a nuclear retention sequence. In some embodiments, the nuclear retention sequence is an Alu sequence. In some embodiments, ORF1p and / or ORF2p are fused to an MS2 coat protein. In some embodiments, the 5'UTR or 3'UTR sequence comprises at least one, two, three, or more MS2 hairpin sequences. In some embodiments, the 5'UTR or 3'UTR sequence comprises a sequence that promotes or enhances the interaction of an mRNA polyA tail with an endonuclease and / or reverse transcriptase. In some embodiments, the 5'UTR or 3'UTR sequence comprises a sequence that promotes or enhances the interaction of a polyA-binding protein (PABP) with an endonuclease and / or reverse transcriptase. In some embodiments, the 5' or 3' UTR sequence comprises a sequence that increases the specificity of an endonuclease and / or reverse transcriptase for the mRNA relative to other mRNAs expressed by the cell, hi some embodiments, the 5' or 3' UTR sequence comprises an Alu element sequence.

[0218]

[0318] In some embodiments, the first sequence encoding ORF1p and the second sequence encoding the endonuclease and / or reverse transcriptase have the same promoter. In some embodiments, the insert sequence has a promoter that is different from the promoter of the first sequence encoding ORF1p. In some embodiments, the insert sequence has a promoter that is different from the promoter of the second sequence encoding the endonuclease and / or reverse transcriptase. In some embodiments, the first sequence encoding ORF1p and / or the second sequence encoding the endonuclease and / or reverse transcriptase have a promoter or transcription start site selected from the group consisting of an inducible promoter, a CMV promoter or transcription start site, a T7 promoter or transcription start site, an EF1a promoter or transcription start site, and combinations thereof. In some embodiments, the insert sequence has a promoter or transcription start site selected from the group consisting of an inducible promoter, a CMV promoter or transcription start site, a T7 promoter or transcription start site, an EF1a promoter or transcription start site, and combinations thereof.

[0219]

[0319] In some embodiments, the first sequence encoding ORF1p and the second sequence encoding the endonuclease and / or reverse transcriptase are codon-optimized for expression in human cells.

[0220]

[0320] In some embodiments, the mRNA comprises a WPRE element. In some embodiments, the mRNA comprises a selectable marker. In some embodiments, the mRNA comprises a sequence encoding an affinity tag. In some embodiments, the affinity tag is linked to a sequence encoding an endonuclease and / or a reverse transcriptase.

[0221]

[0321] In some embodiments, the 3'UTR comprises a polyA sequence, or the polyA sequence is added to the mRNA in vitro. In some embodiments, the polyA sequence is downstream of the sequence encoding the endonuclease and / or reverse transcriptase. In some embodiments, the insertion sequence is upstream of the polyA sequence.

[0222]

[0322] In some embodiments, the 3'UTR sequence comprises an insertion sequence. In some embodiments, the insertion sequence comprises a sequence that is the reverse complement of the sequence encoding the foreign polypeptide. In some embodiments, the insertion sequence comprises a polyadenylation site. In some embodiments, the insertion sequence comprises an SV40 polyadenylation site. In some embodiments, the insertion sequence comprises a polyadenylation site upstream of the sequence that is the reverse complement of the sequence encoding the foreign polypeptide. In some embodiments, the insertion sequence is integrated into the genome at a locus that is not a ribosomal locus. In some embodiments, the insertion sequence is integrated into a gene or a gene regulatory region, thereby disrupting the gene or down-regulating gene expression. In some embodiments, the insertion sequence is integrated into a gene or a gene regulatory region, thereby up-regulating gene expression. In some embodiments, the insertion sequence is integrated into the genome and replaces the gene. In some embodiments, the insertion sequence is stably integrated into the genome. In some embodiments, the insertion sequence is retrotransposed into the genome. In some embodiments, the insertion sequence is integrated into the genome by cleavage of the DNA strand at the target site by an endonuclease encoded by the mRNA. In some embodiments, the insert is integrated into the genome via target-primed reverse transcription (TPRT). In some embodiments, the insert is integrated into the genome via reverse splicing of mRNA into a DNA target site in the genome.

[0223]

[0323] In some embodiments, the cell is an immune cell. In some embodiments, the immune cell is a T cell or a B cell. In some embodiments, the immune cell is a bone marrow cell. In some embodiments, the immune cell is selected from the group consisting of monocytes, macrophages, dendritic cells, dendritic precursor cells, and macrophage precursor cells.

[0224]

[0324] In some embodiments, the mRNA is a self-integrating mRNA. In some embodiments, the method comprises introducing the mRNA into a cell. In some embodiments, the method comprises introducing a vector encoding the mRNA into the cell. In some embodiments, the method comprises introducing the mRNA or a vector encoding the mRNA into the cell ex vivo. In some embodiments, the method further comprises administering the cell to a human subject. In some embodiments, the method comprises administering the mRNA or a vector encoding the mRNA to the human subject. In some embodiments, an immune response is not elicited in the human subject. In some embodiments, the mRNA or vector is substantially non-immunogenic.

[0225]

[0325] In some embodiments, the vector is a plasmid or a viral vector. In some embodiments, the vector comprises a non-LTR retrotransposon. In some embodiments, the vector comprises a human L1 element. In some embodiments, the vector comprises an L1 retrotransposon ORF1 gene. In some embodiments, the vector comprises an L1 retrotransposon ORF2 gene. In some embodiments, the vector comprises an L1 retrotransposon.

[0226]

[0326] In some embodiments, the mRNA is at least about 1, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 2.7, 2.8, 2.9, or 3 kilobases, hi some embodiments, the mRNA is at most about 2.5, 2.6, 2.7, 2.8, 2.9, 3, 3.1, 3.2, 3.3, 3.4, 3.5, 3.6, 3.7, 3.8, 3.9, 4, 4.1, 4.2, 4.3, 4.4, 4.5, 4.6, 4.7, 4.8, 4.9, or 5 kilobases.

[0227]

[0327] In some embodiments, the mRNA comprises a payload that is at least about 1, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 2.7, 2.8, 2.9, or 3 kilobases. In some embodiments, the mRNA is at most about 2.5, 2.6, 2.7, 2.8, 2.9, 3, 3.1, 3.2, 3.3, 3.4, 3.5, 3.6, 3.7, 3.8, 3.9, 4, 4.1, 4.2, 4.3, 4.4, 4.5, 4.6, 4.7, 4.8, 4.9, or 5 kilobases. In some embodiments, the mRNA is at least about 5.1, 5.2, 5.3, 5.4, 5.5, 5.6, 5.7, 5.8, 5.9, or 6 kilobases. In some embodiments, the mRNA is at least about 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, or 7 kilobases. In some embodiments, the mRNA is at least about 7.1, 7.2, 7.3, 7.4, 7.5, 7.6, 7.7, 7.8, 7.9, or 8 kilobases. In some embodiments, the mRNA is at least about 8.1, 8.2, 8.3, 8.4, 8.5, 8.6, 8.7, 8.8, 8.9, or 9 kilobases. In some embodiments, the mRNA is at least about 9.1, 9.2, 9.3, 9.4, 9.5, 9.6, 9.7, 9.8, 9.9, or 10 kilobases. In some embodiments, the mRNA is at least about 10.1, 10.2, 10.3, 10.4, 10.5, 10.6, 10.7, 10.8, 10.9, or 11 kilobases. In some embodiments, the mRNA is at least about 11.1, 11.2, 11.3, 11.4, 11.5, 11.6, 11.7, 11.8, 11.9, or 12 kilobases. In some embodiments, the mRNA comprises a payload of about 6.8 kB, e.g., a sequence encoding an ABCA4 gene product. In some embodiments, the mRNA comprises a payload of about 6.7 kB, e.g., a sequence encoding an MY07A gene product. In some embodiments, the mRNA comprises a payload of about 7.5 kB, e.g., a sequence encoding the CEP290 gene product. In some embodiments, the mRNA comprises a payload of about 10.1 kB, e.g., a sequence encoding the CDH23 gene product.In some embodiments, the mRNA comprises a payload of about 9.4 kB, e.g., a sequence encoding the EYS gene product. In some embodiments, the mRNA comprises a payload of about 15.6 kB, e.g., a sequence encoding the USH2a gene product. In some embodiments, the mRNA comprises a payload of about 12.5 kB, e.g., a sequence encoding the ALMS1 gene product. In some embodiments, the mRNA comprises a payload of about 4.6 kB, e.g., a sequence encoding the GDE gene product. In some embodiments, the mRNA comprises a payload of about 6 kB, e.g., a sequence encoding the OTOF gene product. In some embodiments, the mRNA comprises a payload of about 7.1 kB, e.g., a sequence encoding the F8 gene product.

[0228]

[0328] One advantage of using retrotransposition to integrate nucleic acid into genome is that it can be designed as described herein to deliver longer nucleic acid cargo than any other existing method.For example, lentivirus and adeno-associated virus (AAV) gene delivery methods are not expected to deliver nucleic acid cargo larger than 4 kB.In addition, lentivirus delivery carries the risk of insertional mutagenesis and other toxicities.AAV-mediated delivery carries unknown liver and CNS toxicities.On the other hand, the retrotransposition-mediated method (retro-T) using mRNA described herein is quick, safe, and less complicated than these viral methods.

[0229]

[0329] In some embodiments, the mRNA comprises a sequence that inhibits or prevents mRNA degradation. In some embodiments, the sequence that inhibits or prevents mRNA degradation inhibits or prevents mRNA degradation by exonucleases or RNAses. In some embodiments, the sequence that inhibits or prevents mRNA degradation is a G4 structure, a pseudoknot, or a triplex sequence. In some embodiments, the sequence that inhibits or prevents mRNA degradation is an exoribonuclease-resistant RNA structure derived from flavivirus RNA or an ENE element derived from KSV. In some embodiments, the sequence that inhibits or prevents mRNA degradation inhibits or prevents mRNA degradation by deadenylases. In some embodiments, the sequence that inhibits or prevents mRNA degradation comprises non-adenosine nucleotides within or at the end of the polyA tail of the mRNA. In some embodiments, the sequence that inhibits or prevents mRNA degradation improves mRNA stability. In some embodiments, the foreign sequence comprises a sequence encoding a foreign polypeptide. In some embodiments, the sequence encoding the foreign polypeptide is not in frame with a sequence encoding an endonuclease and / or a reverse transcriptase. In some embodiments, the sequence encoding the foreign polypeptide is not in frame with a sequence encoding an endonuclease and / or a reverse transcriptase. In some embodiments, the foreign sequence does not contain an intron. In some embodiments, the foreign sequence comprises a sequence encoding a foreign polypeptide selected from the group consisting of an enzyme, a receptor, a transport protein, a structural protein, a hormone, an antibody, a contractile protein, and a storage protein. In some embodiments, the foreign sequence comprises a sequence encoding a foreign polypeptide selected from the group consisting of a chimeric antigen receptor (CAR), a ligand, an antibody, a receptor, and an enzyme. In some embodiments, the foreign sequence comprises a regulatory sequence. In some embodiments, the regulatory sequence comprises a cis-acting regulatory sequence. In some embodiments, the regulatory sequence comprises a cis-acting regulatory sequence selected from the group consisting of an enhancer, a silencer, a promoter, or a response element. In some embodiments, the regulatory sequence comprises a trans-acting regulatory sequence. In some embodiments, the regulatory sequence comprises a trans-acting regulatory sequence encoding a transcription factor.

[0230]

[0330] In some embodiments, integration of the insertion sequence does not adversely affect the health of the cell. In some embodiments, the endonuclease, reverse transcriptase, or both are capable of site-specific integration of the insertion sequence.

[0231]

[0331] In some embodiments, the mRNA comprises a sequence encoding an additional nuclease domain or a nuclease domain not derived from ORF2. In some embodiments, the mRNA comprises a sequence encoding a DNA binding domain that binds to a repeat sequence, such as a megaTAL nuclease domain, a TALEN domain, a Cas9 domain, a zinc finger binding domain derived from an R2 retroelement, or Rep78 derived from AAV. In some embodiments, the endonuclease comprises a mutation that reduces the activity of the endonuclease compared to an endonuclease that does not have the mutation. In some embodiments, the endonuclease is an ORF2p endonuclease, and the mutation is S228P. In some embodiments, the mRNA comprises a sequence encoding a domain that improves the accuracy and / or processivity of the reverse transcriptase. In some embodiments, the reverse transcriptase is a reverse transcriptase derived from a retroelement other than ORF2, or a reverse transcriptase that has higher accuracy and / or processivity compared to the reverse transcriptase of ORF2p. In some embodiments, the reverse transcriptase is a group II intron reverse transcriptase. In some embodiments, the group II intron reverse transcriptase is a group IIA intron reverse transcriptase, a group IIB intron reverse transcriptase, or a group IIC intron reverse transcriptase, hi some embodiments, the group II intron reverse transcriptase is TGIRT-II or TGIRT-III.

[0232]

[0332] In some embodiments, the mRNA comprises a sequence comprising an Alu element and / or a ribosome-binding aptamer. In some embodiments, the mRNA comprises a sequence encoding a polypeptide comprising a DNA-binding domain. In some embodiments, the 3'UTR sequence is derived from a viral 3'UTR or a beta-globin 3'UTR.

[0233]

[0333] In one aspect, provided herein is a composition comprising a recombinant mRNA or a vector encoding the mRNA, wherein the mRNA comprises a human LINE-1 transposon sequence comprising a human LINE-1 transposon 5'UTR sequence, a sequence encoding ORF1p downstream of the human LINE-1 transposon 5'UTR sequence, an inter-ORF linker sequence downstream of the sequence encoding ORF1p, a sequence encoding ORF2p downstream of the inter-ORF linker sequence, and a 3'UTR sequence derived from the human LINE-1 transposon downstream of the sequence encoding ORF2p; and wherein the 3'UTR sequence comprises an insertion sequence, wherein the insertion sequence is the reverse complement of a sequence encoding a foreign polypeptide or the reverse complement of a sequence encoding a foreign regulatory element.

[0234]

[0334] In some embodiments, when the insertion sequence is introduced into a cell, it is integrated into the genome of the cell.In some embodiments, the insertion sequence is integrated into a gene associated with a condition or disease, thereby disrupting the gene or down-regulating the expression of the gene.In some embodiments, the insertion sequence is integrated into a gene, thereby up-regulating the expression of the gene.In some embodiments, the recombinant mRNA or the vector encoding the mRNA is isolated or purified.

[0235]

[0335] In one aspect, provided herein is a composition comprising a nucleic acid comprising a nucleotide sequence encoding (a) a long interspersed nuclear element (LINE) polypeptide, the LINE polypeptide comprising human ORF1p and human ORF2p, and (b) an insert sequence, the insert sequence being the reverse complement of a sequence encoding a foreign polypeptide or the reverse complement of a sequence encoding a foreign regulatory element, wherein the composition is substantially non-immunogenic.

[0236]

[0336] In some embodiments, the composition comprises human ORF1p and human ORF2p proteins. In some embodiments, the composition comprises a ribonucleoprotein (RNP) comprising human ORF1p and human ORF2p complexed with a nucleic acid. In some embodiments, the nucleic acid is mRNA.

[0237]

[0337] In one aspect, provided herein is a composition comprising a cell comprising the composition described herein. In some embodiments, the cell is an immune cell. In some embodiments, the immune cell is a T cell or a B cell. In some embodiments, the immune cell is a bone marrow cell. In some embodiments, the immune cell is selected from the group consisting of monocytes, macrophages, dendritic cells, dendritic precursor cells, and macrophage precursor cells. In some embodiments, the inserted sequence is the reverse complement of the sequence encoding the foreign polypeptide, and the foreign polypeptide is a chimeric antigen receptor (CAR).

[0238]

[0338] In one aspect, provided herein is a pharmaceutical composition comprising a composition described herein and a pharmaceutically acceptable excipient. In some embodiments, the pharmaceutical composition is for use in gene therapy. In some embodiments, the pharmaceutical composition is for use in the manufacture of a medicament for treating a disease or condition. In some embodiments, the pharmaceutical composition is for use in treating a disease or condition. In one aspect, provided herein is a method of treating a disease in a subject, the method comprising administering to a subject having the disease or condition a pharmaceutical composition described herein. In some embodiments, the method increases the amount or activity of a protein or functional RNA in the subject. In some embodiments, the subject has an insufficient amount or activity of the protein or functional RNA. In some embodiments, the insufficient amount or activity of the protein or functional RNA is associated with or causes the disease or condition.

[0239]

[0339] In some embodiments, the method further comprises administering an agent that inhibits the human silencing hub (HUSH) complex, an agent that inhibits FAM208A, or an agent that inhibits TRIM28. In some embodiments, the agent that inhibits the human silencing hub (HUSH) complex is an agent that inhibits periphilin, TASOR, and / or MPP8. In some embodiments, the agent that inhibits the human silencing hub (HUSH) complex inhibits the assembly of the HUSH complex.

[0240]

[0340] In some embodiments, the agent inhibits the Fanconi anemia complex. In some embodiments, the agent inhibits FANCD2-FANC1 heterodimer monoubiquitination. In some embodiments, the agent inhibits FANCD2-FANC1 heterodimer formation. In some embodiments, the agent inhibits the Fanconi anemia (FA) core complex. The FA core complex is a component of the Fanconi anemia DNA damage repair pathway, for example, in chemotherapy-induced DNA interstrand crosslinking. The FA core complex contains two major dimers: a FANCB subunit and a 100 kDa FA-associated protein (FAAP100) subunit, flanked by two copies of a RING finger subunit called FANCL. These two heterotrimers act as a scaffold to assemble the remaining five subunits, resulting in an extended, asymmetric structure. Destabilization of the scaffold can disrupt the entire complex, resulting in a non-functional FA pathway. Examples of agents that can inhibit the FA core complex include bortezomib and the curcumin analogs EF24 and 4H-TTD.

[0241]

[0341] In some embodiments, the inserted sequence can be placed under the control of tissue-specific elements, such that the entire inserted DNA is functional only in cells in which the tissue-specific elements are active.

[0242]

[0342] In one aspect, provided herein are methods and compositions for stable gene transfer into cells by introducing into the cell a heterologous nucleic acid or gene of interest (e.g., a transgene, a regulatory sequence, e.g., a sequence for an interfering nucleic acid, siRNA, miRNA) flanked by sequences that cause the heterologous nucleic acid sequence to retrotranspose into the genome of the cell. In some embodiments, the heterologous nucleic acid is descriptively referred to herein as an insert, and an insert is a nucleic acid sequence that can be reverse transcribed and inserted into the genome of a cell by the intended design of the construct described herein. In some embodiments, the heterologous nucleic acid is descriptively referred to herein as a cargo or cargo sequence. The cargo can comprise the sequence of the heterologous nucleic acid to be inserted into the genome. In some embodiments, the cell can be a mammalian cell. The mammalian cell can be of epithelial, mesothelial, or endothelial origin. In some embodiments, the cell can be a stem cell. In some embodiments, the cell can be a progenitor cell. In some embodiments, the cell can be a terminally differentiated cell. In some embodiments, the cells may be muscle cells, cardiac cells, epithelial cells, hematopoietic cells, mucosal cells, epidermal cells, squamous cells, chondrocytes, bone cells, or any cells of mammalian origin. In some embodiments, the cells are hematopoietic lineage cells. In some embodiments, the cells are myeloid lineage cells, or phagocytic cells, such as monocytes, macrophages, dendritic cells, or myeloid progenitor cells. In some embodiments, the nucleic acid encoding the transgene is mRNA.

[0243]

[0343] In some embodiments, the retrotransposable element may be derived from a non-LTR retrotransposon.

[0244]

[0344] Provided herein is a method for integrating a nucleic acid sequence into the genome of a cell, the method comprising introducing a recombinant mRNA or a vector encoding the mRNA into a cell, wherein the mRNA comprises an insertion sequence, and the insertion sequence is integrated into the genome of the cell. In some embodiments, the insertion sequence comprises (i) a foreign sequence, or (ii) a sequence that is the reverse complement of the foreign sequence; a 5'UTR sequence and a 3'UTR sequence downstream of the 5'UTR sequence, wherein the 5'UTR sequence or the 3'UTR sequence comprises a binding site for a human ORF protein. In some embodiments, the ORF protein is a human LINE1 ORF2 protein. In some embodiments, the ORF protein is a non-human ORF protein. In some embodiments, the ORF protein is a chimeric protein, a recombinant protein, or a modified protein.

[0245]

[0345] Provided herein is a method for integrating a nucleic acid sequence into the genome of an immune cell, comprising the step of introducing a recombinant mRNA or a vector encoding the mRNA, wherein the mRNA comprises: (a) an insert sequence, the insert sequence comprising (i) a foreign sequence or (ii) a sequence that is the reverse complement of the foreign sequence; (b) a 5'UTR sequence and a 3'UTR sequence downstream of the 5'UTR sequence, wherein the 5'UTR sequence or the 3'UTR sequence comprises an endonuclease binding site and a reverse transcriptase binding site; and wherein the transgene sequence is integrated into the genome of the immune cell.

[0246]

[0346] In some embodiments, structural elements mediating RNA integration or transposition may be encoded in a synthetic construct, which is expected to deliver a heterologous gene of interest to a cell. In some embodiments, the synthetic construct may include a nucleic acid encoding the heterologous gene of interest and structural elements that cause the heterologous gene of interest to integrate or retrotranspose into the genome. In some embodiments, the structural elements that cause integration or retrotransposition may include a 5'L1 RNA region and a 3'L1 region, the latter of which includes a polyA 3' region for priming. In some embodiments, the 5'L1 RNA region may include one or more stem-loop regions. In some embodiments, the L1 3' region may include one or more stem-loop regions. In some embodiments, the 5' and 3'L1 regions are constructed adjacent to a nucleic acid sequence encoding a heterologous gene of interest (transgene). In some embodiments, the structural element may include a region derived from L1 or Alu RNA that includes a hairpin loop structure containing A box and B box elements, which are ribosome binding sites. In some embodiments, the synthetic nucleic acid may include an L1-Ta promoter.

[0247]

[0347] Two types of LINE RNA recognition by ORF2p can exist: strict and lax. In the strict type, the RT recognizes its own 3'UTR tail, while in the lax type, the RT does not require any specific recognition other than the polyA tail. The classification into strict and lax types arose from the observation that some LINE / SINE pairs share the same 3' end. Regarding the strict type, experimental studies have shown that the 3'UTR stem-loop promotes retrotransposition. The 5'UTR of LINE retrotransposition sequences has been shown to contain three conserved stem-loop regions.

[0248]

[0348] In some embodiments, the transgene or transcript of interest may be flanked at the 5' and 3' ends by transposable elements derived from L1 or Alu sequences. In some embodiments, the 5' region of the retrotransposon comprises an Alu sequence. In some embodiments, the 3' region of the retrotransposon comprises an Alu sequence. In some embodiments, the 5' region of the retrotransposon comprises an L1 sequence. In some embodiments, the 3' region of the retrotransposon comprises an L1 sequence. In some embodiments, the transgene or transcript of interest is flanked by SVA transposon sequences.

[0249]

[0349] In some embodiments, the transcript of interest may include an L1 or Alu sequence encoding a binding region for ORF2p and a 3' polyA priming region. In some embodiments, the heterologous nucleic acid encoding the transgene of interest may be flanked by an L1 or Alu sequence encoding a binding region for ORF1p and a 3' polyA priming region. The 3' region may include one or more stem-loop structures. In some embodiments, the transcript of interest is structured for cis integration or retrotransposition. In some embodiments, the transcript of interest is structured for trans integration or retrotransposition.

[0250]

[0350] In some embodiments, the retrotransposon is a human retrotransposon.The sequence of the human retrotransposon can include a sequence encoding an endonuclease and / or a reverse transcriptase.The sequence of the human retrotransposon can encode two proteins that are translated from a single RNA that contains two non-overlapping ORFs.In some embodiments, the two ORFs are ORF1 and ORF2.

[0251]

[0351] Thus, provided herein is a method for stably integrating a heterologous nucleic acid encoding a transgene into the genome of a cell, such as a bone marrow cell, comprising the step of introducing into a cell a nucleic acid encoding the transgene; one or more 5' nucleic acid sequences adjacent to the region encoding the transgene, comprising the 5' region of a retrotransposon; and one or more 3' nucleic acid sequences adjacent to the region encoding the transgene, comprising the 3' region of a retrotransposon, wherein the 3' region of the retrotransposon comprises a genomic DNA priming sequence and a LINE transposase binding sequence, and has respective endonuclease and reverse transcriptase (RT) activities.

[0252]

[0352] Provided herein is a method for integrating a nucleic acid sequence into the genome of a cell, comprising the step of introducing a recombinant mRNA or a vector encoding the mRNA, wherein the mRNA comprises an insertion sequence, the insertion sequence comprising (i) a foreign sequence or (ii) a sequence that is the reverse complement of the foreign sequence; (b) a 5'UTR sequence, a sequence of a human retrotransposon downstream of the 5'UTR sequence, and a 3'UTR sequence downstream of the human retrotransposon sequence, wherein the 5'UTR sequence or the 3'UTR sequence comprises an endonuclease binding site and a reverse transcriptase binding site, and the human retrotransposon sequence encodes two proteins that are translated from a single RNA containing two ORFs, and wherein the insertion sequence is integrated into the genome of the cell.

[0253]

[0353] In some embodiments, the method includes using a single nucleic acid molecule to deliver and integrate the insert into the genome of the cell. The single nucleic acid molecule can be a plasmid vector. The single nucleic acid can be a DNA or RNA molecule. The single nucleic acid can be mRNA.

[0254]

[0354] In some embodiments, the method includes introducing one or more polynucleotides comprising a human retrotransposon and a heterologous nucleic acid sequence into a cell. In some embodiments, the one or more polynucleotides include (i) a first nucleic acid molecule encoding ORF1p, and (ii) a second nucleic acid molecule encoding ORF2p and a cargo-encoding sequence. In some embodiments, the first nucleic acid and the second nucleic acid are mRNA. In some embodiments, the first nucleic acid and the second nucleic acid are DNAs, for example, encoded by separate plasmid vectors.

[0255]

[0355] Provided herein is a self-integrating polynucleotide comprising a sequence to be inserted into the genome of a cell, wherein the insert is stably integrated into the genome by the naked self-integrating polynucleotide.In some embodiments, the polynucleotide is RNA.In some embodiments, the polynucleotide is mRNA.In some embodiments, the polynucleotide is mRNA with modification.In some embodiments, the modification ensures protection from RNAse in the intracellular environment.In some embodiments, the modification includes a substitution modified nucleotide, such as 5-methylcytidine, pseudouridine, or 2-thiouridine.

[0256]

[0356] In some embodiments, a single polynucleotide is used for delivery and genome integration of an insert (or cargo) nucleic acid. In some embodiments, the single polynucleotide is bicistronic. In some embodiments, the single polynucleotide is tricistronic. In some embodiments, the single polynucleotide is multicistronic. In some embodiments, two or more polynucleotide molecules are used for delivery and genome integration of an insert (or cargo) nucleic acid.

[0257]

[0357] In some embodiments, a retrotransposable element can be generated that comprises: (i) a heterologous nucleic acid (insert) encoding a transgene or non-coding sequence to be inserted into the genome of a cell; (ii) a nuclear sequence encoding one or more retrotransposon ORF coding sequences; and (iii) one or more UTR regions of the ORF coding sequence such that the heterologous nucleic acid encoding the inserted transgene or non-coding sequence is contained within the UTR sequence, wherein the 3' region of the retrotransposon ORF coding sequence comprises a genomic DNA priming sequence.

[0258]

[0358] In some embodiments, retrotransposable genetic elements can be introduced into cells to stably integrate transgenes into genomic DNA. In some embodiments, retrotransposable genetic elements include (a) retrotransposon protein coding sequences and 3'UTRs, and (b) sequences that contain heterologous nucleic acids to be inserted (e.g., integrated) into the genome of cells. Retrotransposon protein coding sequences and 3'UTRs are a complete and sufficient unit for delivering heterologous nucleic acid sequences into the genome of cells, and can include sequences in 3'UTRs that bind to and prime genomic DNA at the region cut by retrotransposable elements, such as endonucleases, reverse transcriptases, and initiate reverse transcription and integration of heterologous nucleic acids.

[0259]

[0359] In some embodiments, the coding sequence of the insert is in a forward orientation relative to the coding sequence of one or more ORFs. In some embodiments, the coding sequence of the insert is in a reverse orientation relative to the coding sequence of one or more ORFs. The coding sequence of the insert and the coding sequence of one or more ORFs may comprise separate regulatory elements, including 5' UTRs, 3' UTRs, promoters, enhancers, etc. In some embodiments, the 3' UTR or 5' UTR of the insert may comprise the coding sequence of one or more ORFs; similarly, the coding sequence of the insert may be located within the 3' UTR of the coding sequence of one or more ORFs.

[0260]

[0360] In some embodiments, a retrotransposable element can be generated that includes: (a) an insertion sequence, which includes: (i) a foreign sequence, a sequence that is the reverse complement of the foreign sequence; a 5'UTR sequence and a 3'UTR sequence downstream of the 5'UTR sequence, wherein the 5'UTR sequence or the 3'UTR sequence includes a binding site for a human ORF protein.

[0261]

[0361] In some embodiments, the retrotransposon may comprise a SINE or LINE element, hi some embodiments, the retrotransposon comprises a SINE or LINE stem-loop structure, such as an Alu element.

[0262]

[0362] In some embodiments, the retrotransposon is a LINE-1 (L1) retrotransposon. In some embodiments, the retrotransposon is a human LINE-1. Human LINE-1 sequences are abundant in the human genome. There are approximately 13,224 human L1s in total, of which 480 are active, accounting for approximately 3.6%. Therefore, human L1 proteins are well tolerated and non-immunogenic in humans. Furthermore, strict regulation of random transposition in humans ensures that random transposase activity is not elicited by the introduction of the L1 system described herein. In addition, the retrotransposable constructs designed herein may include targeted and specific integration of insertion sequences. In some embodiments, the retrotransposable genetic element may include a design intended to overcome silencing mechanisms that are active and widespread in human cells while taking care not to initiate random integration that would result in genomic instability.

[0263]

[0363]

[0264]

[0364]

[0265]

[0365]

[0266]

[0366] In some embodiments, the construct comprises a nucleic acid sequence encoding a nuclear localization sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to PAAKRVKLD. In some embodiments, the nuclear localization sequence is fused to the ORF2p sequence. In some embodiments, the construct comprises a nucleic acid sequence encoding a Flag tag having the sequence DYKDDDDK. In some embodiments, the Flag tag is fused to the ORF2p sequence. In some embodiments, the Flag tag is fused to the nuclear localization sequence.

[0267]

[0367] In some embodiments, the construct is ASNFTQFVLVDNGGTGDVTVAPSNFANGIAEWISSNSRSQAYKVTCSVRQSSAQNRKYTIKVEVPKGAWRSYLNMELTIPIFATNSDCELIVKAMQGLLKDGNPIPSAIAANSGIYAMASNFTQFVLVDNGGTGDVTVAPSNFANGIAEWISSNSRSQAYKVTCSVRQSSAQNRKYTIKVEVPKGAWRSYLNMELTIPIFATNSDCELIVKAMQGLLKDGN In some embodiments, the nucleic acid sequence encoding an MS2 coat protein has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to PIPSAIAANSGIY. In some embodiments, the MS2 coat protein sequence is fused to an ORF2p sequence.

[0268]

[0368] In some embodiments, the transgene may include flanking sequences that include an Alu ORF2p recognition sequence.

[0269]

[0369] In some embodiments, additional elements may be introduced into the mRNA. In some embodiments, the additional elements may be an IRES element or a T2A element. In some embodiments, the mRNA transcript comprises one, two, three, or more stop codons at the 3' end.

[0270]

[0370] In some embodiments, one, two, three, or more stop codons are designed to be in tandem. In some embodiments, one, two, three, or more stop codons are designed to be present in all three reading frames. In some embodiments, one, two, three, or more stop codons may be designed to be present in multiple reading frames and in tandem.

[0271]

[0371] In some embodiments, one or more target-specific nucleotides can be added at the priming end of the L1 or Alu RNA priming region.

[0272]

[0372] In some embodiments, the 5' or 3' UTR sequence, in addition to being capable of binding to the ORF protein, may also be capable of binding to one or more endogenous proteins that regulate gene retrotransposition and / or stable integration. In some embodiments, the flanking sequence is capable of binding to a PABP protein.

[0273]

[0373] In some embodiments, the 5' region adjacent to the transcript may contain a strong promoter, hi some embodiments, the promoter is a CMV promoter.

[0274]

[0374] In some embodiments, an additional nucleic acid encoding L1 ORF2p is introduced into the cell. In some embodiments, the sequence encoding L1 ORF1 is removed, and only L1-ORF2 is included. In some embodiments, the nucleic acid encoding the transgene with flanking elements is mRNA. In some embodiments, endogenous L1-ORF1p function can be suppressed or inhibited.

[0275]

[0375] In some embodiments, the nucleic acid encoding the transgene having a retrotransposition flanking element comprises one or more nucleic acid modifications. In some embodiments, the nucleic acid encoding the transgene having a retrotransposition flanking element comprises one or more nucleic acid modifications in the transgene. In some embodiments, the modifications comprise codon optimization of the transgene sequence. In some embodiments, the codon optimization is for more efficient recognition by the human translation machinery, resulting in more efficient expression in human cells. In some embodiments, the one or more nucleic acid modifications are made in the 5' or 3' flanking sequence, including one or more stem-loop regions. The nucleic acid encoding the transgene having a retrotransposition flanking element comprises one, two, three, four, five, six, seven, eight, nine, ten, or more nucleic acid modifications.

[0276]

[0376] In some embodiments, the retrotransferred transgene is stably expressed throughout the life of the cell. In some embodiments, the cell is a bone marrow cell. In some embodiments, the bone marrow cell is a monocyte precursor cell. In some embodiments, the bone marrow cell is an immature monocyte. In some embodiments, the monocyte is an undifferentiated monocyte. In some embodiments, the bone marrow cell is a CD14+ cell. In some embodiments, the bone marrow cell does not express the CD16 marker. In some embodiments, the bone marrow cell can remain functionally active under suitable conditions for a desired period of more than 3 days, more than 4 days, more than 5 days, more than 6 days, more than 7 days, more than 8 days, more than 9 days, more than 10 days, more than 11 days, more than 12 days, more than 13 days, more than 14 days, or longer. Suitable conditions may represent in vitro conditions or in vivo conditions, or a combination of both.

[0277]

[0377] In some embodiments, the retrotransferred transgene can be stably expressed in cells for about 2 days, about 3 days, about 4 days, about 5 days, about 6 days, about 7 days, about 8 days, about 9 days, or about 10 days. In some embodiments, the retrotransferred transgene is stably expressed in cells for more than 10 days. In some embodiments, the retrotransferred transgene is stably expressed in cells for more than 2 weeks. In some embodiments, the retrotransferred transgene is stably expressed in cells for about 1 month.

[0278]

[0378] In some embodiments, the retrotransferred transgene may be modified for stable expression, hi some embodiments, the retrotransferred transgene may be modified for resistance to in vivo silencing.

[0279]

[0379] In some embodiments, the expression of the retrotransposed transgene can be controlled by a strong promoter. In some embodiments, the expression of the retrotransposed transgene can be controlled by a moderately strong promoter. In some embodiments, the expression of the retrotransposed transgene can be controlled by a strong promoter that can be regulated in an in vivo environment. In some embodiments, the promoter is a CMV promoter. In some embodiments, the promoter is an L1-Ta promoter.

[0280]

[0380] In some embodiments, ORF1p may be overexpressed. In some embodiments, ORF2 may be overexpressed. In some embodiments, ORF1p or ORF2p, or both, are overexpressed. In some embodiments, when ORF1 is overexpressed, ORF1p is at least 1.1 times, 1.5 times, 2 times, 3 times, 4 times, 5 times, 6 times, 7 times, 8 times, 9 times, 10 times, 12 times, 14 times, 16 times, 18 times, 20 times, 30 times, 40 times, 50 times, 60 times, 70 times, 80 times, 90 times, or at least 100 times greater than that of cells that do not overexpress ORF1.

[0281]

[0381] In some embodiments, in the case of overexpression of the ORF2 sequence, ORF2p is at least 1.1-fold, 1.5-fold, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 12-fold, 14-fold, 16-fold, 18-fold, 20-fold, 30-fold, 40-fold, 50-fold, 60-fold, 70-fold, 80-fold, 90-fold, or at least 100-fold greater than that of cells and ORF2p that are not overexpressed.

[0282] Retrometastatic fidelity and target specificity

[0382] LINE-1 elements can bind to the poly(A) tails of their own mRNAs to initiate retrotransposition. LINE-1 elements preferentially retrotranspose their own mRNAs over random mRNAs (Dewannieux et al., 2013; 3,000-fold higher LINE-1 retrotransposition compared to random mRNAs). Additionally, LINE-1 elements can also integrate nonspecific poly(A) sequences into the genome.

[0283]

[0383] In one aspect, the present disclosure provides a retrotransposition composition and a method for using the same with improved retrotransposition specificity.For example, a retrotransposition composition with high specificity can be used for highly specific and efficient reverse transcription and subsequent integration into the genome of target cells, such as bone marrow cells.In some embodiments, the retrotransposition composition provided herein comprises a retrotransposition cassette that contains one or more additional components that improve integration or retrotransposition specificity.For example, the retrotransposon cassette can encode one or more additional elements that enable high-affinity RNA-protein interaction, which competes with and eliminates non-specific binding between polyA sequence and ORF2.

[0284]

[0384] Thus, several means for enhancing integration or retrotransposition efficiency are disclosed herein.

[0285]

[0385] One exemplary means for enhancing integration or retrotransposition efficiency is external manipulation of cells. The endonuclease function of the retrotransposition machinery delivered to cells is likely to be inhibited by cellular transposition silencing mechanisms, such as DNA repair pathways. For example, small molecules can be used to modulate or inhibit DNA repair pathways in cells before introducing nucleic acids. For example, cell cycle-synchronized cell populations have been shown to increase gene transfer into cells, so cell sorting and / or synchronization can be used before introducing nucleic acids, such as by electroporation. Cell sorting can be used to synchronize or homogenize cell types and increase uniform transfer and expression of foreign nucleic acids. Uniformity can be achieved by selecting stem cells from non-stem cells. Another exemplary means for enhancing integration or retrotransposition efficiency is enhancing biochemical activity. For example, this can be achieved by improving reverse transcriptase processing ability or DNA cleavage (endonuclease) activity. Another exemplary means for enhancing integration or retrotransposition efficiency is disrupting endogenous silencing mechanisms. For example, this can be achieved by replacing the entire LINE-1 sequence with a LINE-1 from a different organism. Another exemplary means for enhancing integration or retrotransposition efficiency is to enhance translation and ribosome binding. For example, this can be achieved by increasing LINE-1 protein expression, increasing LINE protein binding to LINE-1 mRNA, or increasing LINE-1 complex binding to ribosomes. Another exemplary means for enhancing integration or retrotransposition efficiency is to increase nuclear import or retention. For example, this can be achieved by fusing the LINE-1 sequence with a nuclear retention signal sequence. Another exemplary means for enhancing integration or retrotransposition efficiency is to enhance sequence-specific insertion. For example, this can be achieved by fusing a targeting domain to ORF2 to increase sequence-specific retrotransposition.

[0286]

[0386] In one embodiment, the method includes enhancing a retrotransposon by modifying the UTR sequence of the LINE-1 ORF to improve the specificity and robustness of cargo expression. In some embodiments, the 5' UTR upstream of the ORF1 or ORF2 coding sequence may be further modified to include a sequence complementary to a target region in the genome, which facilitates homologous recombination at a specific site where ORF nuclease can act and retrotransposition can occur. In some embodiments, the sequence capable of binding to the target sequence by homology is 2 to 15 nucleotides in length. In some embodiments, the sequence homologous to the genomic target contained in the 5' UTR of ORF1 mRNA can be about 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides in length. In some embodiments, the sequence homologous to the genomic target is about 12 or 15 nucleotides in length. In some embodiments, the sequence having homology to the genomic target is at least about 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 1120, or 125 nucleotides in length. In some embodiments, the sequence having homology to the genomic target comprises about 2-5, about 2-6, about 2-8, or about 2-10, or about 2-12 contiguous nucleotides that share complementarity with the respective target region in the genome. In some embodiments, the sequence having homology to the genomic target is at least about or about 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 1120 or 125 contiguous nucleotides that share complementarity with the respective target region in the genome.

[0287]

[0387] In some embodiments, ORF2 is associated with or fused to an additional protein domain comprising RNA-binding activity. In some embodiments, the retrotransposon cassette comprises a cognate RNA sequence with affinity for the additional protein domain associated with or fused to ORF2. In some embodiments, ORF2 is associated with or fused to an MS2-MCP coat protein. In some embodiments, the retrotransposon cassette further comprises an MS2 hairpin RNA sequence in the 3' or 5' UTR sequence that interacts with the MS2-MCP coat protein. In some embodiments, ORF2 is associated with or fused to a PP7 coat protein. In some embodiments, the retrotransposon cassette further comprises a PP7 hairpin RNA sequence in the 3' or 5' UTR sequence that interacts with the MS2-MCP coat protein. In some embodiments, the one or more additional elements improve retrotransposition specificity by at least 1.5-fold, at least 2-fold, at least 3-fold, at least 4-fold, at least 5-fold, at least 10-fold, at least 20-fold, at least 30-fold, at least 50-fold, at least 100-fold, at least 200-fold, at least 300-fold, at least 500-fold, at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 3000-fold, at least 5000-fold, or more, compared to a retrotransposon cassette without the one or more additional elements.

[0288]

[0388] The DNA endonuclease domain consists of a series of purines 3' to the target site and a series of pyrimidines (Py) following the target site. n ↓(Pu) n An exemplary sequence is (adenosine) n ↓(thymidine) n It could be.

[0289]

[0389] In one aspect, a method for using retrotransposition with high target specificity is provided herein. In some embodiments, a CRISPR-Cas guide RNA system is combined with the LINE-retrotransposon system used herein to improve the accuracy of site-specific retrotransposition. In some embodiments, the system incorporates a prime editing guide RNA (pegRNA) to incorporate one or more ORF binding sequences into a specific genomic locus. In some embodiments, the pegRNA incorporates a sequence that binds to a human ORF, such as TTTTTA, in a site-specific manner. In some embodiments, the CRISPR-Cas system comprises a Cas9 enzyme. In some embodiments, the CRISPR-Cas comprises a Cfp1 enzyme. In some embodiments, the Cas9 is a dCas9 combined with a nickase system.

[0290]

[0390] Therefore, provided herein is a method and composition for the stable integration of transgene into the genome of bone marrow cells, such as monocytes or macrophages, wherein the method comprises using a non-LTR retrotransposon system to integrate transgene, and the retrotransposition with target specificity, high accuracy and precision is carried out at a specific genomic locus.Therefore, in some embodiments, the method comprises administering to cells a composition comprising a system that has at least one transgene flanked by one or more retrotransposable elements and one or more nucleic acids that code for one or more proteins for improving transposition specificity, and / or further comprises modifying one or more genes associated with retrotransposition.

[0291]

[0391] A nucleic acid containing a transgene located in the 3'UTR region of a retrotransposable element is often referred to as a retrotransposition cassette. Thus, in some embodiments, a retrotransposition cassette comprises a nucleic acid encoding a transgene and flanking an Alu transposable element. A retrotransposable element comprises a sequence for binding to a retrotransposon, such as an L1-transposon, for example, an L1-ORF protein, i.e., ORF1p and ORF2p. ORF proteins are known to bind to their own mRNA sequence for retrotransposition. Thus, a retrotransposition cassette comprises a nucleic acid encoding a transgene; an adjacent L1-ORF2p binding sequence, and / or an L1-ORF1p binding sequence, and further comprises a sequence encoding the L1-ORF1p and L1-ORF2p coding sequences outside the transgene sequence. In some embodiments, a spacer region, also referred to as the ORF1-ORF2 region, is located between L1-ORF1 and L1-ORF2. In some embodiments, the L1-ORF1 and L1-ORF2 coding sequences are in opposite orientation relative to the transgene coding region. The retrotransposition cassette can include a polyA region downstream of the L1-ORF2 coding sequence, with the transgene sequence being located downstream of the polyA sequence. L1-ORF2 includes nucleic acid sequences encoding an endonuclease (EN) and a reverse transcriptase (RT), followed by a polyA sequence. In some embodiments, the L1-ORF2 sequence in the retrotransposition cassettes described herein is a complete (intact) sequence, i.e., encodes the full-length native (WT) L1-ORF2 sequence. In some embodiments, the L1-ORF2 sequence in the retrotransposition cassettes described herein includes a partial or modified sequence.

[0292]

[0392] The system described herein can include promoters for expressing L1-ORF1p and L1-ORF2p. In some embodiments, transgene expression is driven by separate promoters. In some embodiments, the transgene and ORF are in tandem orientation. In some embodiments, the transgene and ORF are in opposite orientation.

[0293]

[0393] In some embodiments, the method includes incorporating one or more elements in addition to the retrotransposon cassette. In some embodiments, the one or more additional elements include a nucleic acid sequence encoding one or more domains of a heterologous protein. The heterologous protein can be a sequence-specific nucleic acid-binding protein, such as a sequence-specific DNA-binding protein domain (DBD). In some embodiments, the heterologous protein is a nuclease or a fragment thereof. In some embodiments, the additional element includes a nucleic acid sequence encoding one or more nuclease domains or fragments thereof from a heterologous protein. In some embodiments, the heterologous nuclease domain has reduced nuclease activity. In some embodiments, the heterologous nuclease domain is inactive. In some embodiments, the ORF2 nuclease is inactive, while the one or more nuclease domains from the heterologous protein are configured to confer specificity to retrotransposition. In some embodiments, the one or more nuclease domains or fragments thereof from the heterologous protein target the retrotransposition and integration of the polynucleotide of interest to a specific desired polynucleotide within the genome to be integrated. In some embodiments, one or more nuclease domains from a heterologous protein comprise a megaTAL nuclease domain, a TALEN, or a zinc finger nuclease domain, for example, a megaTAL, TALE, or zinc finger domain fused or associated with a nuclease domain, for example, a FokI nuclease domain. In some embodiments, one or more nuclease domains from a heterologous protein comprise a CRISPR-Cas protein domain with a specific guide nucleic acid, for example, a guide RNA (gRNA) for a specific target locus. In some embodiments, the CRISPR-Cas protein is a Cas9, Cas12a, Cas12b, Cas13, CasX, or CasY protein domain. In some embodiments, one or more nuclease domains from a heterologous protein have target specificity.

[0294]

[0394] In some embodiments, an additional nuclease domain may be incorporated into the ORF2 domain. In some embodiments, the additional nuclease domain may be fused to the ORF2p domain. In some embodiments, the additional nuclease domain may be fused to ORF2p, where the ORF2p comprises a mutation in the ORF2p endonuclease domain. In some embodiments, the mutation inactivates the ORF2p endonuclease domain. In some embodiments, the mutation is a point mutation. In some embodiments, the mutation is a deletion. In some embodiments, the mutation is an insertion. In some embodiments, the mutation suppresses ORF2 endonuclease (nickase) activity. In some embodiments, the mutation inactivates DNA target recognition of the ORF2p endonuclease. In some embodiments, the mutation spans a region associated with DNA recognition by the ORF2p nuclease. In some embodiments, the mutation reduces DNA target recognition of the ORF2p endonuclease. In some embodiments, the ORF2p endonuclease domain mutation is in the N-terminal region of the protein. In some embodiments, the ORF2p endonuclease domain mutation occurs in a conserved region of the protein. In some embodiments, the ORF2p endonuclease domain mutation occurs in a conserved N-terminal region of the protein. In some embodiments, the mutation comprises the N14 amino acid in the L1 endonuclease domain. In some embodiments, the mutation comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more consecutive amino acids including the N14 amino acid in the L1 endonuclease domain. In some embodiments, the mutation comprises the E43 amino acid in the L1 endonuclease domain. In some embodiments, the mutation comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more consecutive amino acids including the E43 amino acid in the L1 endonuclease domain. In some embodiments, the mutation comprises two or more amino acids in the L1 endonuclease domain, including N14 or E43, or a combination thereof. In some embodiments, the mutation comprises D145 of the L1 endonuclease domain. In some embodiments, the mutation may be D145A.In some embodiments, the mutation may include D205 in the L1 endonuclease domain. In some embodiments, the mutation may be D205G. In some embodiments, the mutation may include H230 in the L1 endonuclease domain. In some embodiments, the mutation may include S228 in the L1 endonuclease domain. In some embodiments, the mutation may be S228P.

[0295]

[0395] In some embodiments, the mutation reduces DNA target recognition of ORF2p endonuclease by at least 50%. In some embodiments, the mutation reduces DNA target recognition of ORF2p endonuclease by at least 60%. In some embodiments, the mutation reduces DNA target recognition of ORF2p endonuclease by at least 70%. In some embodiments, the mutation reduces DNA target recognition of ORF2p endonuclease by 80%. In some embodiments, the mutation reduces DNA target recognition of ORF2p endonuclease by 90%. In some embodiments, the mutation reduces DNA target recognition of ORF2p endonuclease by 95%. In some embodiments, the mutation reduces DNA target recognition of ORF2p by 100%.

[0296]

[0396] In some embodiments, the mutation is a deletion. In some embodiments, the deletion is complete, i.e., 100% of the L1 endonuclease domain is deleted. In some embodiments, the deletion is partial. In some embodiments, about 98%, about 95%, about 94%, about 93%, about 92%, about 91%, about 90%, about 85%, about 80%, about 75%, about 70%, about 65%, about 60%, or about 50% of the ORF2 endonuclease domain is deleted.

[0297]

[0397] In some embodiments, an additional nuclease domain is inserted into the ORF2 protein sequence.In some embodiments, the ORF2 endonuclease domain is deleted and replaced with an endonuclease domain derived from a heterologous protein.In some embodiments, the ORF2 endonuclease is partially deleted and replaced with an endonuclease domain derived from a heterologous protein.The endonuclease domain derived from a heterologous protein can be a megaTAL nuclease domain.The endonuclease domain derived from a heterologous protein can be a TALEN.The endonuclease domain derived from a heterologous protein can be Cas9 with a gRNA specific to a certain locus.

[0298]

[0398] In some embodiments, the endonuclease (i) is an endonuclease that has a specific target in the genome and (ii) generates 5'-P and 3'-OH ends at the cleavage site.

[0299]

[0399] In some embodiments, the additional endonuclease domain from a heterologous protein is an endonuclease domain from a related retrotransposon.

[0300]

[0400] In some embodiments, the endonuclease domain from a heterologous protein may comprise a bacterial endonuclease modified to target a specific site. In some embodiments, the endonuclease domain from a heterologous protein may comprise a homing endonuclease domain or a fragment thereof. In some embodiments, the endonuclease is a homing endonuclease. In some embodiments, the homing endonuclease is a modified LAGLIDADG homing endonuclease (LHE) or a fragment thereof. In some embodiments, the additional endonuclease may be a restriction endonuclease, Cre, Cas TAL, or a fragment thereof. In some embodiments, the endonuclease may comprise a group II intron-encoded protein (ribozyme) or a fragment thereof.

[0301]

[0401] The modified or modified L1-ORF2p discussed in the preceding paragraph, which confers specific DNA targeting ability for an additional / heterologous endonuclease, is expected to be highly advantageous in driving the targeted stable integration of a transgene into the genome. When expressed in cells, the modified L1-ORF2p can achieve significantly reduced off-target effects compared to the use of natural, i.e., unmodified, L1-ORF2p. In some embodiments, the modified L1-ORF2p does not produce off-target effects.

[0302]

[0402] In some embodiments, the altered or modified L1-ORF2p is a normal (Py) n ↓(Pu) n In some embodiments, the modified L1-ORF2p targets a recognition site other than the (Py) n ↓(Pu) n In some embodiments, the modified L1-ORF2p targets a recognition site, e.g., a hybrid target site, including a TTTT / AA site. n ↓(Pu) n In some embodiments, the modified L1-ORF2p targets a recognition site having at least one nucleotide in addition to the site, for example, TTTT / AAG, or TTTT / AAC, or TTTT / AAT, TTTT / AAA, GTTTT / AA, CTTTT / AA, ATTTT / AA, or TTTTT / AA. ... n ↓(Pu) n In some embodiments, the modified L1-ORF2p targets a conventional L1-ORF2p(Py) site, as well as another recognition site. n ↓(Pu) n In some embodiments, the modified L1-ORF2p targets a recognition site other than the target site. In some embodiments, the modified L1-ORF2p targets a recognition site that is 4, 5, 6, 7, 8, 9, 10, or longer than ...

[0303]

[0403] The modified L1-ORF2p can be modified to retain the ability to bind to its own mRNA after translation and reverse transcribe with high efficiency. In some embodiments, the modified L1-ORF2p has enhanced reverse transcription efficiency compared to native (WT) L1-ORF2p.

[0304]

[0404] In some embodiments, the system comprising a retrotransposable element further comprises a genetic modification that reduces nonspecific retrotransposition. In some embodiments, the genetic modification may comprise a sequence encoding L1-ORF2p. In some embodiments, the modification may comprise a mutation of one or more amino acids essential for binding to a protein that assists ORF2p in binding to target genomic DNA. The protein that assists ORF2p in binding to target genomic DNA may be part of the chromatin-ORF interactome. In some embodiments, the modification may comprise one or more amino acids essential for binding to a protein that assists ORF2p DNA endonuclease activity. In some embodiments, the modification may comprise one or more amino acids essential for binding to a protein that assists ORF2p RT activity. In some embodiments, the modification may comprise a protein binding site of ORF2p such that the association of the protein with ORF2p is altered, where the binding of the protein with ORF2p is required for binding to chromatin. In some embodiments, the modification may comprise a protein binding site of ORF2p such that the association of the protein with ORF2p is more stringent and / or specific than in the absence of the modification. In some embodiments, the binding of ORF2p to target DNA has improved specificity as a result of altered association of ORF2p with proteins due to modifications in protein binding sites in the ORF2p coding sequence, hi some embodiments, the modifications may reduce binding of ORF2 to one or more proteins that are part of the ORF2p chromatin interactome.

[0305]

[0405] In some embodiments, the genetic modification may be in the PIP domain of ORF2p.

[0306]

[0406] In some embodiments, the genetic modification may be present in one or more genes encoding proteins that bind to ORF2p and assist ORF2p in its recognition, binding, endonuclease, or RT activity. In some embodiments, the genetic modification may be present in one or more genes encoding PCNA, PARP1, PABP, MCM, TOP1, RPA, PURA, PURB, RUVBL2, NAP1, ZCCHC3, UPF1, or MOV10 proteins, at the ORF2p interaction site of each protein, or at a site that affects the interaction of the protein with ORF2p or the interaction of ORF2p with target DNA. In some embodiments, the modification may be present in the ORF2p binding domain of PCNA, at an ORF2p interaction site, or at a site that affects the interaction of the protein with ORF2p or the interaction of ORF2p with target DNA. In some embodiments, the modification may be present in the ORF2p binding domain of TOP1. In some embodiments, the modification may be present in the ORF2p binding domain of RPA. In some embodiments, the modification may be present in the ORF2p-binding domain of PARP1 at an ORF2p interaction site or at a site that affects the interaction of a protein with ORF2p or the interaction of ORF2p with target DNA. In some embodiments, the modification may be present in the ORF2p-binding domain of PABP (e.g., PABPC1) at an ORF2p interaction site or at a site that affects the interaction of a protein with ORF2p or the interaction of ORF2p with target DNA. In some embodiments, the genetic modification may be present in an MCM gene. In some embodiments, the genetic modification may be present in the gene encoding the MCM3 protein at an ORF2p interaction site or at a site that affects the interaction of a protein with ORF2p or the interaction of ORF2p with target DNA. In some embodiments, the genetic modification may be present in the gene encoding the MCM5 protein at an ORF2p interaction site or at a site that affects the interaction of a protein with ORF2p or the interaction of ORF2p with target DNA.In some embodiments, the genetic modification may be present at an ORF2p interaction site in the gene encoding the MCM6 protein, or at a site that affects the interaction of the protein with ORF2p or the interaction of ORF2p with target DNA. In some embodiments, the genetic modification may be present at an ORF2p interaction site in the gene encoding the MEPCE protein, or at a site that affects the interaction of the protein with ORF2p or the interaction of ORF2p with target DNA. In some embodiments, the genetic modification may be present at an ORF2p interaction site in the gene encoding the RUVBL1 or RUVBL2 protein, or at a site that affects the interaction of the protein with ORF2p or the interaction of ORF2p with target DNA. In some embodiments, the genetic modification may be present at an ORF2p interaction site in the gene encoding the TROVE protein, or at a site that affects the interaction of the protein with ORF2p or the interaction of ORF2p with target DNA.

[0307]

[0407] In some embodiments, the retrotransposition systems disclosed herein comprise one or more elements that improve the fidelity of reverse transcription.

[0308]

[0408] In some embodiments, the L1-ORF2 RT domain is modified, in some embodiments, to improve fidelity, improve processivity, improve DNA-RNA substrate affinity, or inactivate RNase H activity.

[0309]

[0409] In some embodiments, the modification comprises introducing one or more mutations into the L1-ORF2 RT domain to improve RT fidelity. In some embodiments, the mutation comprises a point mutation. In some embodiments, the mutation comprises an alteration, such as a substitution of one, two, three, four, five, six, or more amino acids in the L1-ORF2p RT domain. In some embodiments, the mutation comprises a deletion of one or more amino acids, e.g., one, two, three, four, five, six, seven, eight, nine, ten, or more amino acids, in the L1-ORF2p RT domain. In some embodiments, the mutation may comprise an indel mutation. In some embodiments, the mutation may comprise a frameshift mutation.

[0310]

[0410] In some embodiments, the modification may include the inclusion of an additional RT domain or fragment thereof from a second protein. In some embodiments, the second protein is a viral reverse transcriptase. In some embodiments, the second protein is a non-viral reverse transcriptase. In some embodiments, the second protein is a retrotransposable element. In some embodiments, the second protein is a non-LTR retrotransposable element. In some embodiments, the second protein is a group II intron protein. In some embodiments, the group II intron is TGIRTII. In some embodiments, the second protein is a Cas nickase, wherein the retrotransposable system further comprises introducing a guide RNA. In some embodiments, the second protein is a Cas9 endonuclease, wherein the retrotransposable system further comprises introducing a guide RNA. In some embodiments, the second protein or fragment thereof is fused to the N-terminus of the L1-ORF2 RT domain or modified L1-ORF2 RT domain. In some embodiments, the second protein or fragment thereof is fused to the C-terminus of the L1-ORF2 RT domain or modified L1-ORF2 RT domain.

[0311]

[0411] In some embodiments, an additional RT domain or fragment thereof from a second protein is incorporated into the retrotransposition system in addition to the full-length WT L1-ORF2p RT domain. In some embodiments, the additional RT domain or fragment thereof from a second protein is incorporated in the presence of a modified (altered) L1-ORF2p RT domain or fragment thereof, where the modification (or alteration) may include mutations to enhance L1-ORF2p RT processing ability, stability, and / or fidelity of the modified L1-ORF2p RT compared to native or WT ORF2p.

[0312]

[0412] In some embodiments, the reverse transcriptase domain can be replaced with other more processive and highly accurate RT domains from other retroelements or group II introns, such as TGIRTII.

[0313]

[0413] In some embodiments, the modification may include fusion with an additional RT domain or fragment thereof from a second protein. In some embodiments, the second protein may include a retroelement. The additional RT domain or fragment thereof from the second protein is configured to improve the accuracy of reverse transcription of the fused L1-ORF2p RT domain. In some embodiments, a nucleic acid encoding the additional RT domain or fragment thereof is fused to a native or WT L1-ORF2 coding sequence. In some embodiments, a nucleic acid encoding an additional RT domain or fragment thereof from a second protein is fused to a modified L1-ORF2 coding sequence. In some embodiments, the modification includes introducing one or more mutations into the L1-ORF2 RT domain or fragment thereof that improve the accuracy of the fused RT. In some embodiments, the mutation in the L1-ORF2 RT domain or fragment thereof comprises a point mutation. In some embodiments, the mutation includes an alteration, such as a substitution of one, two, three, four, five, six, or more amino acids, in the L1-ORF2p RT domain. In some embodiments, the mutation comprises a deletion of one or more amino acids in the L1-ORF2p RT domain, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more amino acids. In some embodiments, the mutation may comprise an indel mutation. In some embodiments, the mutation may comprise a frameshift mutation.

[0314]

[0414] In some embodiments, the modified L1-ORF2p RT domain has improved processivity relative to the WT L1-ORF2p RT domain.

[0315]

[0415] In some embodiments, the modified L1-ORF2p RT domain has at least 10% greater processing ability and / or accuracy than the WT L1-ORF2p RT domain. In some embodiments, the modified L1-ORF2p RT domain has at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 110%, 120%, 150%, 200%, 300%, 400%, 500%, 1000%, or more greater processing ability and / or accuracy than the WT L1-ORF2p RT domain. In some embodiments, the modified RT is capable of processing a nucleic acid span of greater than 6 kb. In some embodiments, the modified RT is capable of processing a nucleic acid span of greater than 7 kb. In some embodiments, the modified RT is capable of processing a nucleic acid span of greater than 8 kb. In some embodiments, the modified RT is capable of processing a nucleic acid span of greater than 9 kb. In some embodiments, the modified RT is capable of processing nucleic acid stretches of greater than 10 kb.

[0316] B. Group II Introns and Ribozymes

[0416] Group II enzymes are mobile ribozymes that self-splice RNA precursors, generating excised intron lariat RNAs. The intron encodes a reverse transcriptase, which can stabilize the RNA for forward and reverse splicing and then converts the incorporated intron RNA into DNA.

[0317]

[0417] Group II RNAs are characterized by a conserved secondary structure spanning 400–800 nucleotides. The secondary structure is formed by six domains, DI–VI, organized into a wheel-like structure with the domains radiating from a central point. The domains interact to form a conserved tertiary structure, which brings together distinct sequences to form the active site. The active site binds splice site and branch point nucleotides and activates splicing catalysis upon the association of Mg2+ cations. The DV domain, located within the active site, contains the conserved catalytic AGC and AY bulges, both of which bind the Mg2+ ions required for catalysis. DI is the largest domain, with its upper and lower halves separated by kappa and zeta motifs. The lower half contains an ε' motif associated with the active site. The upper half contains sequence elements that bind the 5' and 3' exons in the active site. DIV encodes an intron-encoded protein (IEP), and subdomain IVa near the 5' end contains a high-affinity binding site for the IEP. Group II introns have conserved 5' and 3' terminal sequences, GUGYG and AY, respectively.

[0318]

[0418] Group II RNA introns can be used to retrotransfer sequences of interest into DNA via target-primed reverse transcription. This process of transfer by group II RNA introns is often referred to as retrohoming. Group II introns recognize DNA target sites through base pairing between the intron RNA and the DNA target sequence, and can be modified to retarget specific sequences contained within the intron to desired DNA sites.

[0319]

[0419] In some embodiments, the methods and compositions for retrotransposition described herein may include a group II intron sequence, a modified group II intron sequence, or a fragment thereof. Exemplary group II IEPs (maturases) include, but are not limited to, bacterial, fungal, and yeast IEPs that are functional in human cells. In particular, the nuclease leaves a 3'-OH residue at the DNA cleavage site that can be utilized by another RT for priming and reverse transcription. An exemplary group II maturase may be TGIRT (thermostable group II intron maturase).

[0320]

[0420] In one or more embodiments of some aspects described herein, the nucleic acid construct comprises RNA. In one or more embodiments of some aspects of the present disclosure, the nucleic acid construct is RNA. In one or more embodiments of some aspects of the present disclosure, the nucleic acid construct is mRNA. In one aspect, the mRNA comprises a sequence of a heterologous gene or portion thereof, where the heterologous gene or portion thereof encodes a polypeptide or protein. In some embodiments, the mRNA comprises a sequence encoding a fusion protein. In some embodiments, the mRNA comprises a sequence encoding a recombinant protein. In some embodiments, the mRNA comprises a sequence encoding a synthetic protein. In some embodiments, the nucleic acid comprises one or more sequences, where the one or more sequences encode one or more heterologous proteins, one or more recombinant proteins, or one or more synthetic proteins, or a combination thereof. In some embodiments, the nucleic acid comprises one or more sequences, where the one or more sequences encode one or more heterologous proteins, including synthetic or recombinant proteins. In some embodiments, the synthetic or recombinant protein is a recombinant fusion protein.

[0321] C. Retrotransposon Systems Containing Site-Specific Editing and / or Integrases

[0421] In one aspect, provided herein is a method for using retrotransposition with the aid of guide RNA and Cas protein, which has higher target specificity after modification than the site-specific pegRNA-mediated integration of LINE binding sequences into genomes. In some embodiments, the CRISPR-Cas guide RNA system is combined with the LINE-retrotransposon system used herein to improve the accuracy of site-specific retrotransposition; for example, the system incorporates a prime editing guide RNA (pegRNA) to integrate one or more ORF binding sequences into specific genomic loci. In some embodiments, the pegRNA integrates a sequence that binds to a human ORF, such as TTTTTA, in a site-specific manner. In some embodiments, the CRISPR-Cas system includes a Cas9 enzyme. In some embodiments, the CRISPR-Cas includes a Cfp1 enzyme. In some embodiments, the Cas9 is a dCas9 combined with a nickase system.

[0322]

[0422] In some embodiments, the retrotransposon system described herein comprises (i) a LINE1 retrotransposon element and (ii) an integrase system or a portion thereof. Some integrase systems are capable of site-specific integration of double-stranded DNA. To avoid double-stranded DNA delivery and / or integration into the genome, a recombinant hybrid system is provided herein in which the integrase or a fragment thereof is incorporated into a recombinant ORF protein or delivered separately as a separate nucleic acid (e.g., mRNA) encoding the integrase or a fragment thereof that recognizes a specific genomic site; this is coupled with LINE1 reverse transcription and insertion of a cargo sequence into the genome of a cell or organism at a precise location guided by the specificity of the integrase. This can be achieved in a few alternative ways. In some embodiments, the cargo sequence contains an attachment site that is recognized and utilized by the integrase to guide the cargo to a landing site in the genome and is also recognized by the same integrase. The integrase is capable of single-strand cleavage. The integrase DNA recognition site, i.e., genomic landing sequence, can be 10 nucleotides in length, e.g., 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, or more nucleotides in length, thereby conferring greater specificity than any other system. The integrase can be truncated or otherwise mutated so that the ORF is reverse transcribed and incorporated into a cargo sequence at an integrase-specific genomic site. Conversely, the ORF protein may be mutated with an RNA recognition site to enable the integrase to recognize a genomic integration sequence (also referred to as a "genomic landing sequence or site") that is preferentially recognized by the integrase.In an alternative embodiment, the integrase may be encoded by a separate polynucleotide and driven by the CRISPR Cas system and guide RNA at the nickable site, and an integrase landing sequence further comprising an ORF binding site comprising 4 nucleotides may be introduced, after which the integrase pulls in a cargo sequence comprising an attachment sequence to the landing sequence, followed by LINE1 activity that results in genome integration at the site specified by the integrase system. Any catalytic activity of the integrase that results in double-stranded DNA integration into the genome is mutated, truncated, or otherwise silenced.

[0323]

[0423] In one or more embodiments of some aspects of the present disclosure, the nucleic acid construct is developed for expression in eukaryotic cells. In some embodiments, the nucleic acid construct is developed for expression in human cells. In some embodiments, the nucleic acid construct is developed for expression in hematopoietic cells. In some embodiments, the nucleic acid construct is developed for expression in bone marrow cells. In some embodiments, the bone marrow cells are human cells.

[0324] II. Modifications in Nucleic Acid Constructs for Methods of Enhanced Expression of Encoded Proteins

[0424] In some embodiments of the present disclosure, recombinant nucleic acids are modified to enhance the expression of proteins encoded by the nucleic acid sequence. Enhanced expression of the encoded proteins can be a function of nucleic acid stability, translation efficiency, and the stability of the translated protein. Several modifications are contemplated herein for incorporation into the design of nucleic acid constructs that can confer nucleic acid stability, for example, the stability of messenger RNA encoding a foreign or heterologous protein, which can be a synthetic recombinant protein or a fragment thereof.

[0325]

[0425] In some embodiments, the nucleic acid is an mRNA comprising one or more sequences, wherein the one or more sequences encode one or more heterologous proteins, including synthetic or recombinant fusion proteins.

[0326]

[0426] In some embodiments, one or more modifications are made in an mRNA that includes a sequence encoding a recombinant or fusion protein to increase mRNA half-life.

[0327]

[0427] Structural elements for preventing exonucleolytic 5'- and 3'-degradation: 5' cap and 3' UTR modifications

[0428] A proper 5' cap structure is important in the synthesis of functional messenger RNA. In some embodiments, the 5' cap comprises guanosine triphosphate organized as GpppG at the 5' end of the nucleic acid. In some embodiments, the mRNA comprises a 5' 7-methylguanosine cap, m7-GpppG. The 5' 7-methylguanosine cap improves mRNA translation efficiency and prevents mRNA 5'-3' exonuclease degradation. In some embodiments, the mRNA comprises an "anti-reverse" cap analog (ARCA, m 7,3’-OThe guanosine cap contains a guanosine cap (GpppG). However, translation efficiency can be significantly improved by using the ARCA method. In some embodiments, the guanosine cap is a Cap 0 structure. In some embodiments, the guanosine cap is a Cap 1 structure. In addition to its essential role in cap-dependent initiation of protein synthesis, the mRNA cap also functions as a protecting group against 5' to 3' exonuclease cleavage and as a unique identifier for recruiting protein factors for pre-mRNA splicing, polyadenylation, and nuclear export. The mRNA cap acts as an anchor for recruiting initiation factors that initiate 5' to 3' loop formation of the mRNA during protein synthesis and translation. Three enzymatic activities are required to generate the Cap 0 structure: RNA triphosphatase (TPase), RNA guanylyltransferase (GTase), and guanine-N7 methyltransferase (guanine-N7MTase). Each of these enzymatic activities performs an essential step in the conversion of the nascent RNA's 5' triphosphate to the Cap 0 structure. RNA TPase removes the gamma phosphate from the 5' triphosphate to generate 5' diphosphate RNA. GTase transfers the GMP group from GTP to the 5' diphosphate via a lysine-GMP covalent intermediate. Guanine-N7MTase then adds a methyl group to the N7 amine of the guanine cap to form the Cap 0 structure. For the Cap 1 structure, m7G-specific 2'O methyltransferase (2'OMTase) methylates the +1 ribonucleotide at the 2'O position of the ribose to generate the Cap 1 structure. Nuclear RNA capping enzymes interact with the polymerase subunit of the RNA polymerase II complex at the phosphorylated Ser5 of the C-terminal heptad repeat. RNA guanine-N7 methyltransferase also interacts with the RNA polymerase II phosphorylated heptad repeat. In some embodiments, the cap is a G4 structure cap.

[0328]

[0429] In some embodiments, mRNA is synthesized by in vitro transcription (IVT). In some embodiments, mRNA synthesis and capping can be carried out in one step. Capping can be carried out in the same reaction mixture as IVT. In some embodiments, mRNA synthesis and capping can be carried out in separate steps. The mRNA thus formed by IVT is purified and then capped.

[0329]

[0430] In some embodiments, nucleic acid constructs, e.g., mRNA constructs, containing one or more sequences encoding a protein or polypeptide of interest can be designed to contain elements that protect, prevent, inhibit, or reduce mRNA degradation by endogenous 5'-3' exoribonucleases, such as Xrn1. Xrn1 is a cellular enzyme in the normal RNA degradation pathway that degrades 5'-monophosphorylated RNA. However, some viral RNA structural elements have been found to be particularly resistant to such RNAses, such as Xrn1-resistant structures in flavivirus sfRNAs called "xrRNAs." For example, the mosquito-borne flavivirus (MBFV) genome contains distinct RNA structures in its 3' untranslated region (UTR) that block Xrn1 progression. These RNA elements are sufficient to block Xrn1 without the use of accessory proteins. The xrRNA stops the enzyme at a defined location, thereby protecting viral RNAs located downstream of the xrRNA from degradation. For example, xrRNA from Zika virus or Murray Valley encephalitis virus contains three-way junctions and multiple pseudoknot interactions that result in an unusual, complex fold that requires a set of nucleotides conserved across the MBFV structure. The xrRNA stalls the enzyme at a defined location, thereby protecting the viral RNA downstream of the xrRNA from degradation. The 5' end of the RNA traverses the ring-like structure of the fold and is thought to remain protected from Xrn1-like exonucleases.

[0330]

[0431] In some embodiments, a nucleic acid construct containing one or more sequences encoding a protein of interest may include one or more xrRNA structures incorporated therein. In some embodiments, the xrRNA is a nucleotide section containing a conserved region of the 3'UTR of one or more viral xrRNA sequences. In some embodiments, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more xrRNA elements are incorporated into the nucleic acid construct. In some embodiments, two or more xrRNA elements are incorporated in tandem into the nucleic acid construct. In some embodiments, the xrRNA comprises one or more regions containing a conserved sequence, or a fragment thereof, or a modification thereof. In some embodiments, the xrRNA is located in the 3'UTR of a retrotransposon element. In some embodiments, the xrRNA is located upstream of a sequence encoding one or more proteins or polypeptides. In some embodiments, the xrRNA is located upstream of the 3'UTR of a retrotransposon element, such as an ORF2 sequence, and a sequence encoding one or more proteins or polypeptides.

[0331]

[0432] In some embodiments, the xrRNA structure comprises an MBFV xrRNA sequence, or a sequence at least 90% identical thereto. In some embodiments, the xrRNA structure comprises a tick-borne flavivirus (TBFV) xrRNA sequence, or a sequence at least 90% identical thereto. In some embodiments, the xrRNA structure comprises a tick-borne flavivirus (TBFV) xrRNA sequence, or a sequence at least 90% identical thereto. In some embodiments, the xrRNA structure comprises a tick-borne flavivirus (TBFV) xrRNA sequence, or a sequence at least 90% identical thereto. In some embodiments, the xrRNA structure comprises an xrRNA sequence from a member of the unknown vector arthropod flavivirus (NKVFV) or a sequence at least 90% identical thereto. In some embodiments, the xrRNA structure comprises an xrRNA sequence from a member of the insect-specific flavivirus (ISFV) or a sequence at least 90% identical thereto. In some embodiments, the xrRNA structure comprises a Zika virus xrRNA sequence, or a sequence at least 90% identical thereto. It is hereby contemplated that any known xrRNA structural element or non-obvious conceivable variation thereof may be used for the purposes described herein.

[0332]

[0433] Some messenger RNAs from various organisms exhibit one or more pseudoknot structures that confer resistance to 5'-3' exonucleases. A pseudoknot is an RNA structure that, at a minimum, consists of two helical segments connected by a single-stranded region or loop, although several distinct folding topologies of pseudoknots exist.

[0333] Poly A tail modification

[0434] The polyA structure in the 3' UTR of an mRNA is an important regulator of mRNA half-life. Deadenylation of the 3' end of the polyA tail is the first step in intracellular mRNA degradation. In some embodiments, the length of the polyA tail of an mRNA construct is carefully considered and designed to maximize expression of the protein encoded by the mRNA coding region and mRNA stability. In some embodiments, a nucleic acid construct comprises one or more polyA sequences. In some embodiments, the polyA sequence in the 3' UTR of one or more protein or polypeptide encoding sequences comprises 20 to 200 adenosine nucleobases. In some embodiments, the polyA sequence comprises 30 to 200 adenosine nucleobases. In some embodiments, the polyA sequence comprises 50 to 200 adenosine nucleobases. In some embodiments, the polyA sequence comprises 80 to 200 adenosine nucleobases. In some embodiments, an mRNA segment comprising a sequence encoding one or more proteins or polypeptides comprises a 3'UTR having a poly-A tail comprising about 180 adenosine nucleobases, or about 140 adenosine nucleobases, or about 120 adenosine nucleobases. In some embodiments, the poly-A tail comprises about 122 adenosine nucleobases. In some embodiments, the poly-A sequence comprises 50 adenosine nucleobases. In some embodiments, the poly-A sequence comprises 30 adenosine nucleobases. In some embodiments, the adenosine nucleobases in the poly-A tail are arranged in tandem with or without intervening non-adenosine bases. In some embodiments, one or more non-adenosine nucleobases are incorporated into the poly-A tail to confer additional resistance to certain exonucleases.

[0334]

[0435] In some embodiments, the adenosine stretch in the polyA tail of the construct comprises one or more non-adenosine (A) nucleobases. In some embodiments, the non-A nucleobases are present at positions -3, -2, -1, and / or +1 of the polyA 3'-terminal region. In some embodiments, the non-A base comprises guanosine (G), cytosine (C), or uracil (U). In some embodiments, the non-A base is G. In some embodiments, two or more non-A bases in tandem, e.g., GG. In some embodiments, modification of the 3'-end of the polyA tail with one or more non-A bases is directed to disrupting A base stacking in the polyA tail. PolyA base stacking promotes deadenylation by various deadenylases, and therefore, a 3'-end of a polyA tail terminating in -AAAG, -AAAGA, or -AAAGGA is effective in conferring stability against deadenylation. In some organisms, GC sequences intervening in polyA sequences have been shown to effectively attenuate 3'-5' exonuclease-mediated degradation. Modifications contemplated herein include intervening non-A residues, or non-A residue duplexes intervening in the 3'-terminal polyA stretch.

[0335]

[0436] In some embodiments, a triplex structure is introduced into the 3'UTR that effectively stops or slows down exonuclease activity associated with the 3' end.

[0336]

[0437] In some embodiments, mRNAs with the modifications described above have an extended half-life and demonstrate stable expression for longer periods than unmodified mRNAs. In some embodiments, the mRNA is stably expressed for more than 2, 3, 4, 5, 6, 7, 8, 9, or 10 days or longer, and the mRNA or its protein product is detectable in vivo. In some embodiments, the mRNA is detectable in vivo for up to 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 days. In some embodiments, the protein product of the mRNA is detectable in vivo for up to 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 25, or 30 days.

[0337] CircRNAs and tectoRNAs

[0438] Circular RNAs are useful for designing and producing stable forms of RNA to be used as messenger RNAs to guide synthetic protein chains, such as long, multiply repeated protein chains. Few methods exist for generating circular RNA (circRNA). Methods include protein-mediated ligation of RNA ends using an RNA ligase, which uses a bisected, self-splicing intron to splice itself together, leaving a ligated product, when located at both ends of a transcribed mRNA (Figure 3A). Another technique relies on the ability of T4 DNA ligase to act as an RNA ligase when the RNA ends to be ligated are held together by an oligonucleotide. Both of these techniques suffer from inefficiency and require large amounts of enzyme. A third technique utilizes the circularization or cyclization activity of group I introns, which is believed to ensure that the majority of the intron sequence that carries out the reaction remains as part of the circle. Group I introns share a set of complex secondary and tertiary structures containing a series of conserved RNA stem-loops that form the catalytic core. Many of these introns have self-splicing activity in vitro and can splice to form two ligated exons as RNA without accessory protein factors. The products generated by Group I autocatalysis are (1) an upstream exon ligated at the 5' splice site with the 3' splice site of the downstream exon, and (2) a linear intron that can undergo further reversible autocatalysis to form a circular intron. The presence of such highly structured, large nucleic acid sequences severely limits the types of RNA sequences that can be circularized by this technique. In addition, the catalytic activity of the intron remains and can interfere with the structure and function of the circular RNA.

[0338]

[0439] It is useful to increase the reaction rate, and therefore the overall efficiency, by bringing the ends of the RNA closer together. Previous studies have achieved this by including complementary RNA sequences on the 3' and 5' ends of the mRNA, so that upon hybridization of these sequences, the ends of the mRNA are brought closer together, allowing the mRNA to undergo ligation or self-splicing reactions at an overall faster rate than without the complementary sequences. These sequences are called the homology arms of the self-splicing circularization reaction (Figure 3A). A major problem with such hybridization strategies is that if sequences complementary to either of the homology arms are present within the coding region, hybridization may actually inhibit the splicing reaction, requiring the arms to be optimized for each new coding region. An alternative to this strategy, described herein, is the use of RNA sequences that fold into a three-dimensional structure to form stable, sequence-independent binding interactions.

[0339]

[0440] Non-Watson-Crick RNA tertiary interactions can be exploited to construct "tectoRNA" molecular units, defined as RNA molecules capable of self-assembly. The use of such types of tertiary interactions is highly dependent on the cation concentration (e.g., Mg). 2+ ) and / or suitable temperature manipulation, as well as the use of modularly designed "selected" RNA molecules, allows the assembly process to be controlled and modulated. For the self-assembly of one-dimensional arrays, a basic modular unit was designed that contains a four-way junction with an interaction module for each helical arm. In some embodiments, the interaction module is a GAAA loop or a specific GAAA loop receptor. Each tectoRNA can interact with two other tectoRNAs, two with each partner molecule, through the formation of four loop-receptor interactions.

[0340]

[0441] In some embodiments, the tectoRNA structure is appropriately selected and incorporated into RNA containing exons and introns to form circRNA. In some embodiments, the incorporation is performed by well-known molecular biology techniques such as ligation. In some embodiments, the tectoRNA forms a stable structure at high temperatures. The tectoRNA structure does not compete with internal RNA sequences, thereby resulting in highly efficient circularization and splicing.

[0341]

[0442] The circRNA can comprise a coding sequence described in any of the preceding sections. For example, the circRNA can comprise a sequence encoding a fusion protein comprising a linked or receptor molecule. The receptor can be a phagocytic receptor fusion protein.

[0342]

[0443] In some embodiments, the intron is a self-splicing intron.

[0343]

[0444] In some embodiments, the terminal region having a tertiary structure, also referred to as a scaffold region for the circRNA, is about 30 to about 100 nucleotides in length. In some embodiments, the tertiary structure motif is about 45, 50, 55, 60, 65, 70, or 75 nucleotides in length. In some embodiments, the tertiary motif is formed at high temperatures. In some embodiments, the tertiary motif is stable.

[0344]

[0445] In some embodiments, nucleic acid constructs having one or more modifications described herein and comprising one or more sequences encoding one or more proteins or polypeptides are stable when administered in vivo. In some embodiments, the nucleic acid is mRNA. In some embodiments, mRNA comprising one or more sequences encoding one or more proteins or polypeptides is stable in vivo for more than 2 days, more than 3 days, more than 4 days, more than 5 days, more than 6 days, more than 7 days, more than 8 days, more than 9 days, more than 10 days, more than 11 days, more than 12 days, more than 13 days, more than 14 days, more than 15 days, more than 16 days, more than 17 days, more than 18 days, more than 19 days, or more than 20 days. In some embodiments, proteins encoded by sequences in mRNA can be detected in vivo for more than 3 days, 4 days, 5 days, 6 days, 7 days, 8 days, 9 days, 10 days, 11 days, 12 days, 13 days, 14 days, 15 days, 16 days, 17 days, 18 days, 19 days, or more than 20 days. In some embodiments, the protein encoded by the sequence in the mRNA can be detected in vivo for about 7 days after the mRNA is administered. In some embodiments, the protein encoded by the sequence in the mRNA can be detected in vivo for about 14 days after the mRNA is administered. In some embodiments, the protein encoded by the sequence in the mRNA can be detected in vivo for about 21 days after the mRNA is administered. In some embodiments, the protein encoded by the sequence in the mRNA can be detected in vivo for about 30 days after the mRNA is administered. In some embodiments, the protein encoded by the sequence in the mRNA can be detected in vivo for more than about 30 days after the mRNA is administered.

[0345]

[0446] In some aspects, enhancing nucleic acid uptake or incorporation into cells is contemplated to enhance the expression of retrotransposition. One method involves obtaining a homogeneous cell population and initiating nucleic acid incorporation, for example, via transfection in the case of a plasmid vector construct, or via electroporation or any other means that can be suitably used to deliver nucleic acid molecules to cells. In some embodiments, cell cycle synchronization may be explored. Cell cycle synchronization can be achieved by sorting cells for a specific common phenotype. In some embodiments, a cell population may be treated with a reagent that can arrest cell cycle progression of all cells at a specific stage. Exemplary reagents can be found in commercial databases, for example, at www.tocris.com / cell-biology / cell-cycle-inhibitors or www.scbt.com / browse / chemicals-Other-Chemicals-cell-cycle-arresting-compounds. For example, itraconazole or nocodazole, to name a few, inhibit the cell cycle at the G1 phase, or agents that arrest the cell cycle at the G0 / G1 phase, such as 5-[(4-ethylphenyl)methylene]-2-thioxo-4-thiazolidinone (compound 10058-F4) (Tocris Bioscience), or G2M cell cycle inhibitors such as AZD5438 (chemical name: 4-[2-methyl-1-(1-methylethyl)-1H-imidazol-5-yl]-N-[4-(methylsulfonyl)phenyl]-2-pyrimidinamine), which block the cell cycle at the G2M, G1, or S phase. Cyclosporine, hydroxyurea, and thymidine are well-known agents that can cause cell cycle arrest. Some agents may irreversibly alter the cellular state or be toxic to cells. Serum deprivation of cells for approximately 2-16 hours prior to electroporation or transfection, depending on the cell type, can also be an easy and reversible strategy for cell synchronization.

[0346]

[0447] In some embodiments, retrotransposition efficiency can be improved by promoting the generation of DNA double-strand breaks in cells transfected or electroporated with the retrotransposition constructs described herein and / or modulating DNA repair mechanisms.The application of these techniques may be limited depending on the end use of cells that may be genetically engineered ex vivo for stable integration of nucleic acid sequences by this method.In some cases, the use of such techniques can be considered when robust expression of the protein or transcript encoded by the nucleic acid to be integrated is expected as a result for a determined period of time.A method for introducing double-strand breaks into cells includes subjecting cells to controlled ionizing radiation of about 0.1 Gy or less for a short period of time.

[0347]

[0448] In some embodiments, the efficiency of LINE-1-mediated retrotransposition can be improved by treating cells with small molecule inhibitors of DNA repair proteins to increase the opportunity for reverse transcriptase to act. Exemplary small molecule inhibitors of DNA repair proteins include benzamide (CAS55-21-0), olaparib (Lynparza) (CAS763113-22-0), rucaparib (Clovis-AG014699, PF-01367338 Pfizer), niraparib (MK-827 Tesaro) CAS1038915-60-4); veliparib (ABT-888 Abbvie) (CAS 912444-00-9); camptothecin (CPT) (CAS 7689-03-4); irinotecan (CAS 100286-90-6); topotecan (Hycamtin® GlaxoSmithKline) (CAS 123948-87-8); NSC19630 (CAS 72835-26-8); NSC617145 (CAS 203115-63-3); ML 216 (CAS 1430213-30-1); 6-hydroxy DL-dopa (CAS 21373-30-8); D-103; D-G23; DIDS (CAS 67483-13-0); B02 (CAS 1290541-46-6); RI-1 (CAS 415713-60-9); RI-2 (CAS 1417162-36-7); streptonigrin (SN) (CAS 3930-19-6).

[0348] III. Nucleic Acid Cargo: A. Transgene

[0449] In one embodiment, the transgene or non-coding sequence, which is a heterologous nucleic acid sequence inserted into the genome of a cell, is delivered as mRNA. The mRNA may contain more than about 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10,000 bases. In some embodiments, the mRNA may be more than 10,000 bases long. In some embodiments, the mRNA may be about 11,000 bases long. In some embodiments, the mRNA may be about 12,000 bases long. In some embodiments, the mRNA comprises a transgene sequence encoding a fusion protein. In some embodiments, the nucleic acid is delivered as a plasmid.

[0349]

[0450] In some embodiments, the nucleic acid is delivered to the cell by transfection. In some embodiments, the nucleic acid is delivered to the cell by electroporation. In some embodiments, the transfection or electroporation is repeated two or more times to enhance the incorporation of the nucleic acid into the cell.

[0350]

[0451] Contemplated herein is the stable integration of recombinant nucleic acids encoding phagocytic or ligated receptor (PR) fusion proteins (CFPs) via retrotransposon mediation. In some embodiments, the CFP comprises a PR subunit that comprises an intracellular domain that comprises a transmembrane domain and an intracellular signaling domain, and an extracellular domain that comprises an antigen-binding domain specific to an antigen of a target cell, wherein the transmembrane domain and the extracellular domain are operably linked.

[0351]

[0452] In some embodiments, the nucleic acid comprises a sequence encoding a chimeric fusion protein (CFP), wherein the CFP comprises an extracellular domain comprising a CD5-binding domain and a transmembrane domain operably linked to the extracellular domain. In some embodiments, the CD5-binding domain is a CD5-binding protein, e.g., an antigen-binding fragment of an antibody, a Fab fragment, an scFv domain, or an sdAb domain. In some embodiments, the CD5-binding domain comprises an scFv comprising (i) a variable heavy chain (VH) sequence having at least 90% sequence identity with EIQLVQSGGGLVKPGGSVRISCAASGYTFTNYGMNWVRQAPGKGLEWMGWINTHTGEPTYADSFKGRFTFSLDDSKNTAYLQINSLRAEDTAVYFCTRRGYDWYFDVWGQGTTVTV, and (ii) a variable light chain (VL) sequence having at least 90% sequence identity with DIQMTQSPSSLSASVGDRVTITCRASQDINSYLSWFQQKPGKAPKTLIYRANRLESGVPSRFSGSGSGTDYTLTISSLQYEDFGIYYCQQYDESPWTFGGGTKLEIK. In some embodiments, the CFP further comprises an intracellular domain, wherein the intracellular domain comprises one or more intracellular signaling domains, and wherein a wild-type protein comprising the intracellular domain does not comprise an extracellular domain. In some embodiments, the one or more intracellular signaling domains comprise a phagocytic signaling domain. In some embodiments, the phagocytosis signaling domain comprises an intracellular ...

Claims

1. 1. A composition for use in expressing an exogenous therapeutic polypeptide encoded by a genomically integrated sequence at a target site in cells of a population of human cells, comprising: (a) an RNA sequence encoding a Cas nickase; (b) a first guide RNA, or a sequence encoding a first guide RNA, that specifically targets a sequence upstream of a DNA target site in a cell of the population of human cells; (c) a second guide RNA, or a sequence encoding a second guide RNA, that specifically targets a sequence downstream of the DNA target site in a cell of the population of human cells; (d) an RNA sequence that encodes an exogenous therapeutic polypeptide or an RNA sequence that is the reverse complement of a sequence that encodes an exogenous therapeutic polypeptide; and (e) an RNA sequence encoding a human ORF2p polypeptide that lacks endonuclease activity; The composition comprising one or more polynucleic acid molecules comprising:

2. The composition of claim 1, wherein the human ORF2p polypeptide comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:59, or the RNA sequence encoding the human ORF2p polypeptide comprises a sequence having at least 80% sequence identity to SEQ ID NO:

60.

3. The composition of claim 1 , wherein the one or more polynucleic acid molecules further comprise an RNA sequence encoding a human ORF1p polypeptide.

4. The composition of claim 3 , wherein the human ORF1p polypeptide comprises a sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:

57.

5. The composition of claim 1 , wherein the human ORF2p polypeptide comprises a mutation that inactivates endonuclease activity.

6. The composition of claim 5, wherein the mutation comprises a mutation at D145, D205, H230, or S228 relative to the amino acid sequence of SEQ ID NO:

59.

7. The composition of claim 1 , wherein the one or more polynucleic acid molecules are one or more RNA molecules.

8. 2. The composition of claim 1, wherein the Cas nickase is a Cas9 nickase.

9. The composition of claim 1 , wherein the cells of the population of human cells comprise immune cells.

10. 10. The composition of claim 9, wherein the immune cell is a T cell, a B cell, a bone marrow cell, a monocyte, a macrophage, or a dendritic cell.

11. One or more polynucleic acid molecules (a) a 5' homology arm that contains a sequence complementary to the upstream sequence of the target site; and (b) a 3' homology arm that contains a sequence complementary to the sequence downstream of the target site. The composition of claim 1 , comprising:

12. The composition of claim 1 , wherein the exogenous therapeutic polypeptide is selected from the group consisting of a ligand, an antibody, a receptor, an enzyme, a transport protein, a structural protein, a hormone, a contractile protein, a storage protein, and a transcription factor.

13. 13. The composition of claim 12, wherein the receptor is selected from the group consisting of a T cell receptor (TCR) or a chimeric antigen receptor (CAR).

14. The composition of claim 1 , wherein the human ORF2p polypeptide is fused to a nuclear localization signal (NLS).

15. The composition of any one of claims 1 to 14, wherein the one or more polynucleic acid molecules are formulated in nanoparticles.

16. The composition of claim 15 , wherein the nanoparticles are lipid nanoparticles or polymeric nanoparticles.

17. A pharmaceutical composition comprising the composition according to any one of claims 1 to 16 and a pharma- ceutically acceptable excipient.

18. 20. The pharmaceutical composition of claim 17, formulated for systemic administration to a human subject.

19. 17. Use of a composition according to any one of claims 1 to 16 in the manufacture of a medicament for the treatment of a disease or condition in a human in need thereof.