RNA tails for target primed reverse transcription

A nucleic acid construct with a four-nucleotide complementarity region and adenine-rich region enhances the efficiency and stability of target primed reverse transcription by improving insertion into a genome.

WO2025264587A1PCT designated stage Publication Date: 2025-12-26ADDITION THERAPEUTICS INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/033843
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-17
Filing Date
2025-06-16
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Existing target primed reverse transcription (TPRT) methods face inefficiencies and challenges in inserting exogenous nucleic acid sequences into a subject genome using non-long terminal repeat (non-LTR) retrotransposons.

Method used

A nucleic acid construct with a 5' region, a template sequence, and a 3' region comprising a tail region, where the tail region includes a four-nucleotide complementarity region complementary to an rRNA gene and an adenine-rich region with at least 19 nucleotides, where at least 60% of the nucleobases are adenines, is used to enhance insertion efficiency.

Benefits of technology

The nucleic acid construct exhibits increased stability and exonuclease resistance, leading to more efficient insertion into a genome compared to conventional constructs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025033843_26122025_PF_FP_ABST
    Figure US2025033843_26122025_PF_FP_ABST
Patent Text Reader

Abstract

Provided herein are nucleic acid constructs and systems for target primed reverse transcription gene insertion.
Need to check novelty before this filing date? Find Prior Art

Description

RNA TAILS FOR TARGET PRIMED REVERSE TRANSCRIPTIONCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 660,760 filed on June 17, 2024, which is incorporated herein by reference in its entirety.INCORPORATION BY REFERENCE OF SEQUENCE LISTING

[0002] The present application is being filed along with a Sequence Listing in electronic format. The Sequence Listing is provided as a file entitled 67098-710_601_SL.xml, created June 12, 2025, which is 205,525 bytes in size. The information in the electronic format of the Sequence Listing is incorporated by reference in its entirety.BACKGROUND

[0003] Target primed reverse transcription (TPRT) inserts a nucleic acid sequence into a subject genome using non-long terminal repeat (non-LTR) retrotransposons. The insertion of an exogenous nucleic acid sequence into a subject via TPRT faces various difficulties and drawbacks. Additional methods and systems are needed to improve the efficiency and other characteristics of TPRT methods and systems.SUMMARY

[0004] In one aspect disclosed herein is a nucleic acid construct comprising (a) a 5’ region, (b) a template sequence, and (c) a 3’ region comprising a tail region, wherein: the tail region comprises (i) a four-nucleotide complementarity region that is complementary to a region of an rRNA gene, and (ii) an adenine-rich region having at least 19 nucleotides; the adenine-rich region is located 3’ of the four-nucleotide complementarity region in the nucleic acid construct; at least 60% of nucleobases in the adenine-rich region are adenines; and at least 10% of nucleobases in the adenine-rich region are pyrimidine nucleobases.

[0005] In some embodiments, the adenine-rich region comprises a sequence having at least 80%, at least 90%, or at least 95% sequence identity to any one of SEQ ID NOs: 174, 173, 143, 171, 130, 135, 165, 170, 139, 137, 177, 129, 178, or 167. In some embodiments, the adenine-rich region comprises a sequence having at least 80%, at least 90%, at least 95%, or 100% sequence identity to any one of SEQ ID NOs: 202, 208, 194, 193, 192, 204, 209, 195, 191, 189, 193, 205, or 198. In some embodiments, the adenine-rich region comprises any one of SEQ ID NOs: 174, 173, 143, 171, 130, 135, 165, 170, 139, 137, 177, 129, 178, or 167.

[0006] In some embodiments, at least 20% of nucleobases in the adenine-rich region are pyrimidine nucleobases. In some embodiments, the adenine-rich region comprises a sequence having at least 80%, at least 90%, or at least 95% sequence identity to any one of SEQ ID NOs: 174, 173, 143, 171, 135, 170, 139, 137, 177, or 178. In some embodiments, the adenine-rich region comprises a sequence having at least 80%, at least 90%, at least 95%, or 100% sequence identity to any one of SEQ ID NOs: 202, 208, 194, 192, 209, 195, 191, 189, or 205. In some embodiments, the adenine-rich region comprises any one of SEQ ID NOs: 174, 173, 143, 171, 135, 170, 139, 137, 177, or 178.

[0007] In some embodiments, at least 30% of nucleobases in the adenine-rich region are pyrimidine nucleobases. In some embodiments, the adenine-rich region comprises a sequence having at least 80%, at least 90%, or at least 95% sequence identity to any one of SEQ ID NOs: 174, 173, 143, 139, 177, or 178. In some embodiments, the adenine-rich region comprises a sequence having at least 80%, at least 90%, at least 95%, or 100% sequence identity to any one of SEQ ID NOs: 202, 208, 195, 189, or 205. In some embodiments, the adenine-rich region comprises any one of SEQ ID NOs: 174, 173, 143, 139, 177, or 178.

[0008] In some embodiments, the pyrimidine nucleobases comprise uracil. In some embodiments, all of the pyrimidine nucleobases of the adenine-rich region are uracils. In some embodiments, the pyrimidine nucleobases comprise cytosine. In some embodiments, all of the pyrimidine nucleobases of the adenine-rich region are cytosines.

[0009] In some embodiments, the adenine-rich region comprises, at its 3’ end, a tail end sequence, wherein the tail end sequence is selected from the group consisting of: AAA, AAX, and AXX; wherein X is C, G, or U. In some embodiments, there is no additional sequence that is 3’ to the adenine-rich region.

[0010] In some embodiments, the nucleic acid construct further comprises a tail end sequence that is located 3’ of the adenine rich region, wherein the tail end sequence is selected from the group consisting of: AAA, AAX, and AXX; wherein X is C, G, or U. In some embodiments, there is no additional sequence that is 3’ to the tail end sequence.

[0011] In some embodiments, there is no intervening sequence between the four-nucleotide complementarity region and the adenine-rich region.

[0012] In some embodiments, the nucleic acid construct comprises a sequence having at least 80%, at least 90%, or at least 95% sequence identity to any one of SEQ ID NOs: 37, 58, 48, 35, 73, 68, 43, 51, 32, 81, 16, 69, 44, 42, or 56. In some embodiments, nucleic acid construct comprises any one of SEQ ID NOs: 37, 58, 48, 35, 73, 68, 43, 51, 32, 81, 16, 69, 44, 42, or 56.

[0013] In some embodiments, the nucleic acid construct comprises RNA. In some embodiments, the nucleic acid construct comprises a modified uridine. In some embodiments, the modifieduridine is a pseudouridine or a N1 -methylpseudouridine. In some embodiments, the modified uridine is a pseudouridine. In some embodiments, the modified uridine is a Nl- methylpseudouridine. In some embodiments, the nucleic acid construct is a single-stranded nucleic acid construct.

[0014] In some embodiments, the adenine-rich region comprises at least 22 nucleotides. In some embodiments, the adenine-rich region comprises at least 29 nucleotides.

[0015] In some embodiments, the complementarity region is complementary to a region in a 28S rRNA gene. In some embodiments, the complementarity region has a sequence of UAGC or TAGC.

[0016] In some embodiments, the nucleic acid construct has increased stability as compared to an otherwise identical nucleic acid construct having a tail region consisting of SEQ ID NO: 1. In some embodiments, the nucleic acid construct has increased exonuclease resistance as compared to an otherwise identical nucleic acid construct having a tail region consisting of SEQ ID NO: 1. In some embodiments, the nucleic acid construct is more efficiently inserted into a genome by target primed reverse transcription as compared to an otherwise identical nucleic acid construct having a tail region consisting of SEQ ID NO: 1.

[0017] In some embodiments, the template sequence encodes a gene. In some embodiments, the gene is a therapeutic gene or a diagnostic gene.

[0018] In some embodiments, the 5’ region comprises a 5’-module, a promoter, and a 5’ UTR. In some embodiments, the 5’ module comprises a 5 ’-flanking self-cleaving ribozyme motif. In some embodiments, the 5’ module comprises a 5’ UTR derived from a same R2 retroelement as the 5’- flanking self-cleaving ribozyme motif.

[0019] In some embodiments, the 3’ region further comprises a 3’ UTR and a 3’ module.

[0020] In another aspect disclosed herein is a system for modifying a target nucleic acid in a cell of a subject, the system comprising: (a) a nucleic acid construct as described above; and (b) an R2 retroelement polypeptide or a second nucleic acid encoding the R2 retroelement polypeptide; wherein the R2 retroelement polypeptide inserts the template sequence or a reverse complement thereof into the target nucleic acid.

[0021] In another aspect disclosed herein is a cell comprising a nucleic acid construct as described above or a system as described above. In some embodiments, the cell is a liver cell. In some embodiments, the cell is a cancer cell. In some embodiments, the cell is a retinal cell. In some embodiments, the cell comprises an rRNA gene. In some embodiments, the nucleic acid construct is not endogenous to the cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell.

[0022] In another aspect disclosed herein is a method of modifying a target nucleic acid in a cell, the method comprising providing the cell with: (a) a nucleic acid construct as described above; and (b) an R2 retroelement polypeptide or a second nucleic acid encoding the R2 retroelement polypeptide; wherein the R2 retroelement polypeptide inserts the template sequence of the nucleic acid construct, or a reverse complement thereof, into the target nucleic acid. In some embodiments, the target nucleic acid comprises an rRNA gene. In some embodiments, the cell is derived from a human subject.

[0023] In another aspect disclosed herein is a nucleic acid construct comprising (a) a 5’ region, (b) a template sequence, and (c) a 3’ region comprising a tail region, wherein the tail region comprises a four-nucleotide complementarity region and an adenine-rich region of at least 22 nucleotides, wherein at least 70% of the nucleobases in the adenine-rich region are adenine.

[0024] In some embodiments, the tail region further comprises a stability element. In some embodiments, the stability element is 5’ to the adenine-rich region. In some embodiments, the stability element is 3’ to the adenine-rich region. In some embodiments, the four-nucleotide complementarity region is 5’ to the adenine-rich region and / or the stability element.

[0025] In some embodiments, the complementarity region is complementary to a region in a 28S rRNA gene. In some embodiments, the complementarity region has a sequence of UAGC or TAGC.

[0026] In some embodiments, the adenine-rich region is between 22 and 41 nucleotides in length. In some embodiments, the adenine-rich region is 22, 32, or 41 nucleotides in length. In some embodiments, the adenine-rich region comprises one or more non-adenine nucleobases. In some embodiments, the adenine-rich region has no more than 10%, 15%, or 20% or 25% non- adenine nucleobases. In some embodiments, the one or more non-adenine nucleobases comprise any one of cytosine, guanine, uridine, or a modified uridine. In some embodiments, the modified uridine is pseudouridine or N1 -methylpseudouridine. T In some embodiments, the adenine-rich region comprises any one of SEQ ID NOs: 118-178.

[0027] In some embodiments, the terminal three nucleotides of the adenine-rich region each comprise adenine. In some embodiments, the terminal nucleotide of the adenine-rich region does not comprise adenine. In some embodiments, the two most terminal nucleotides of the adenine- rich region do not comprise adenine.

[0028] In some embodiments, the stability element is selected from K3, K4, eK5, MALATl mm, MALATl hs, MALAT1 Jis-core, PAN KSHV, dENE os, HSL, U7, K9, K16, KEMCV, or KBDV1. In some embodiments, the stability element comprises any one of SEQ ID NOs: 90-117.

[0029] In some embodiments, the tail region comprises any one of SEQ ID NOs: 2-89.

[0030] In some embodiments, the nucleic acid construct has increased stability as compared to an otherwise identical nucleic acid construct having a tail region consisting of SEQ ID NO: 1. In some embodiments, the nucleic acid construct has increased exonuclease resistance as compared to an otherwise identical nucleic acid construct having a tail region consisting of SEQ ID NO: 1. In some embodiments, the nucleic acid construct is more efficiently inserted into a genome by target primed reverse transcription as compared to an otherwise identical nucleic acid construct having a tail region consisting of SEQ ID NO: 1.

[0031] In some embodiments, the template sequence encodes a gene. In some embodiments, the gene is a therapeutic gene or a diagnostic gene.

[0032] In some embodiments, the 5’ region comprises a 5’-module, a promoter, and a 5’ UTR. In some embodiments, the 5’ module comprises a 5 ’-flanking self-cleaving ribozyme motif. In some embodiments, the 5’ module comprises a 5’ UTR derived from a same R2 retroelement as the 5’- flanking self-cleaving ribozyme motif.

[0033] In some embodiments, the 3’ region further comprises a 3’ UTR and a 3’ module.

[0034] In another aspect disclosed herein is a system for modifying a target nucleic acid in a cell of a subject, the system comprising: (a) a nucleic acid construct as described above; and (b) an R2 retroelement polypeptide or a second nucleic acid encoding the R2 retroelement polypeptide; wherein the R2 retroelement polypeptide inserts the template sequence or a reverse complement thereof into the target nucleic acid.

[0035] In another aspect disclosed herein is a method of modifying a target nucleic acid in a cell of a subject, the method comprising providing the cell with: (a) a nucleic acid construct as described above; and (b) an R2 retroelement polypeptide or a second nucleic acid encoding the R2 retroelement polypeptide; wherein the R2 retroelement polypeptide inserts the template sequence of the nucleic acid construct, or a reverse complement thereof, into the target nucleic acid.

[0036] In another aspect disclosed herein is a nucleic acid construct comprising (a) a 5’ region, (b) a template sequence, and (c) a 3’ region comprising a tail region, wherein the tail region comprises (i) a four-nucleotide complementarity region that is complementary to a region of an rRNA gene, and (ii) an adenine-rich region having at least 19 nucleotides, wherein the adenine- rich region is located 3’ of the four-nucleotide complementarity region in the nucleic acid construct, wherein at least 60% of the nucleobases in the adenine-rich region are adenines, and wherein at least one nucleobase in the adenine-rich region is a non-adenine nucleobase.

[0037] In another aspect disclosed herein is a nucleic acid construct comprising (a) a 5’ region, (b) a template sequence, and (c) a 3’ region comprising a tail region, wherein the tail region comprises (i) a four-nucleotide complementarity region that is complementary to a region of anrRNA gene, and (ii) an adenine-rich region having at least 22 nucleotides, wherein the adenine- rich region is located 3’ of the four-nucleotide complementarity region in the nucleic acid construct, wherein at least 50% of the nucleobases in the adenine-rich region are adenines, and wherein at least one nucleobase in the adenine-rich region is a non-adenine nucleobase.

[0038] In some embodiments, the non-adenine nucleobase is a pyrimidine nucleobase. In some embodiments, the non-adenine nucleobase is a uracil. In some embodiments, the non-adenine nucleobase is a cytosine.

[0039] In some embodiments, the nucleic acid construct comprises a sequence having at least 80%, at least 90%, or at least 95% sequence identity to any one of SEQ ID NOs: 118-178. In some embodiments, the adenine-rich region comprises a sequence having at least 80%, at least 90%, at least 95%, or 100% sequence identity to any one of SEQ ID NOs: 185-214. In some embodiments, the nucleic acid construct comprises any one of SEQ ID NOs: 118-178.

[0040] In some embodiments, the nucleic acid construct comprises a sequence having at least 80%, at least 90%, or at least 95% sequence identity to any one of SEQ ID NOs: 174, 173, 143, 171, 118, 120-122, 125, 126, 129, 130, 135, 137, 139, 165, 167, 170, 177, and 178. In some embodiments, the adenine-rich region comprises a sequence having at least 80%, at least 90%, at least 95%, or 100% sequence identity to any one of SEQ ID NOs: 202, 108, 194, 185, 207, 193, 192, 191, 195, 204, 198, 209, 189, or 205. In some embodiments, the nucleic acid construct comprises any one of SEQ ID NOs: 174, 173, 143, 171, 118, 120-122, 125, 126, 129, 130, 135, 137, 139, 165, 167, 170, 177, and 178.

[0041] In some embodiments, the nucleic acid construct comprises a sequence having at least 80%, at least 90%, or at least 95% sequence identity to any one of SEQ ID NOs: 2-89. In some embodiments, the nucleic acid construct comprises any one of SEQ ID NOs: 2-89.

[0042] In some embodiments, at least 10% of the nucleobases in the adenine-rich region are non- adenine nucleobases. In some embodiments, the at least 10% of the nucleobases in the adenine- rich region are pyrimidine nucleobases. In some embodiments, the pyrimidine nucleobases comprise uracil. In some embodiments, all of the pyrimidine nucleobases of the adenine-rich region are uracils. In some embodiments, the pyrimidine nucleobases comprise cytosine. In some embodiments, all of the pyrimidine nucleobases of the adenine-rich region are cytosines.

[0043] In some embodiments, the nucleic acid construct comprises a sequence having at least 80%, at least 90%, or at least 95% sequence identity to any one of SEQ ID NOs: 128-178. In some embodiments, the adenine-rich region comprises a sequence having at least 80%, at least 90%, at least 95%, or 100% sequence identity to any one of SEQ ID NOs: 186, 193, 214, 203, 192, 191, 195, 201, 208, 200, 206, 199, 187, 190, 213, 196, 212, 210, 188, 198, 204, 211, 209,194, 202, 189, or 205. In some embodiments, the nucleic acid construct comprises any one of SEQ ID NOs: 128-178.

[0044] In some embodiments, at least 20% of the nucleobases in the adenine-rich region are nonadenine nucleobases. In some embodiments, the at least 20% of the nucleobases in the adenine- rich region are pyrimidine nucleobases. In some embodiments, the pyrimidine nucleobases comprise uracil. In some embodiments, all of the pyrimidine nucleobases of the adenine-rich region are uracils. In some embodiments, the pyrimidine nucleobases comprise cytosine. In some embodiments, all of the pyrimidine nucleobases of the adenine-rich region are cytosines.

[0045] In some embodiments, the nucleic acid construct comprises a sequence having at least 80%, at least 90%, or at least 95% sequence identity to any one of SEQ ID NOs: 133-143, 151- 162, or 168-178. In some embodiments, the adenine-rich region comprises a sequence having at least 80%, at least 90%, at least 95%, or 100% sequence identity to any one of SEQ ID NOs: 203, 192, 191, 195, 201, 208, 187, 190, 213, 196, 212, 210, 211, 209, 194, 202, 189, or 205. In some embodiments, the nucleic acid construct comprises any one of SEQ ID NOs: 133-143, 151- 162, or 168-178.

[0046] In some embodiments, at least 30% of the nucleobases in the adenine-rich region are nonadenine nucleobases. In some embodiments, the at least 30% of the nucleobases in the adenine- rich region are pyrimidine nucleobases. In some embodiments, the pyrimidine nucleobases comprise uracil. In some embodiments, all of the pyrimidine nucleobases of the adenine-rich region are uracils. In some embodiments, the pyrimidine nucleobases comprise cytosine. In some embodiments, all of the pyrimidine nucleobases of the adenine-rich region are cytosines.

[0047] In some embodiments, the nucleic acid construct comprises a sequence having at least 80%, at least 90%, or at least 95% sequence identity to any one of SEQ ID NOs: 139-143, 157- 162, or 173-178. In some embodiments, the adenine-rich region comprises a sequence having at least 80%, at least 90%, at least 95%, or 100% sequence identity to any one of SEQ ID NOs:195, 201, 208, 196, 212, 210, 202, 189, or 205. In some embodiments, the nucleic acid construct comprises any one of SEQ ID NOs: 139-143, 157-162, or 173-178.

[0048] In some embodiments, the adenine-rich region comprises, at its 3’ end, a tail end sequence, wherein the tail end sequence is selected from the group consisting of: AAA, AAX, and AXX, wherein X is C, G, or U. In some embodiments, there is no additional sequence that is 3’ to the adenine-rich region.

[0049] In some embodiments, the nucleic acid construct further comprises a tail end sequence that is located 3’ of the adenine rich region, wherein the tail end sequence is selected from the group consisting of: AAA, AAX, and AXX, wherein X is C, G, or U. In some embodiments, there is no additional sequence that is 3’ to the tail end sequence.

[0050] In some embodiments, there is no intervening sequence between the four-nucleotide complementarity region and the adenine-rich region.

[0051] In some embodiments, the nucleic acid construct comprises RNA. In some embodiments, the nucleic acid construct comprises a modified uridine. In some embodiments, the modified uridine is a pseudouridine or a N1 -methylpseudouridine. In some embodiments, the modified uridine is a pseudouridine. In some embodiments, the modified uridine is a Nl- methylpseudouridine.

[0052] In some embodiments, the tail region further comprises a stability element.

[0053] In some embodiments, the complementarity region is complementary to a region in a 28S rRNA gene. In some embodiments, the complementarity region has a sequence of UAGC or TAGC.

[0054] In some embodiments, the nucleic acid construct has increased stability as compared to an otherwise identical nucleic acid construct having a tail region consisting of SEQ ID NO: 1. In some embodiments, the nucleic acid construct has increased exonuclease resistance as compared to an otherwise identical nucleic acid construct having a tail region consisting of SEQ ID NO: 1. In some embodiments, the nucleic acid construct is more efficiently inserted into a genome by target primed reverse transcription as compared to an otherwise identical nucleic acid construct having a tail region consisting of SEQ ID NO: 1.

[0055] In some embodiments, the template sequence encodes a gene. In some embodiments, the gene is a therapeutic gene or a diagnostic gene.

[0056] In some embodiments, the 5’ region comprises a 5’-module, a promoter, and a 5’ UTR. In some embodiments, the 5’ module comprises a 5 ’-flanking self-cleaving ribozyme motif. In some embodiments, the 5’ module comprises a 5’ UTR derived from a same R2 retroelement as the 5’- flanking self-cleaving ribozyme motif.

[0057] In some embodiments, the 3’ region further comprises a 3’ UTR and a 3’ module.

[0058] In another aspect disclosed herein is a system for modifying a target nucleic acid in a cell of a subject, the system comprising: (a) a nucleic acid construct as described above; and (b) an R2 retroelement polypeptide or a second nucleic acid encoding the R2 retroelement polypeptide; wherein the R2 retroelement polypeptide inserts the template sequence or a reverse complement thereof into the target nucleic acid.

[0059] In another aspect disclosed herein is a cell comprising a nucleic acid construct as described above or a system as described above. In some embodiments, the cell is a liver cell. In some embodiments, the cell is a cancer cell. In some embodiments, the cell is a retinal cell. In some embodiments, the cell comprises an rRNA gene. In some embodiments, the nucleic acidconstruct is not endogenous to the cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell.

[0060] In another aspect disclosed herein is a method of modifying a target nucleic acid in a cell, the method comprising providing the cell with: (a) a nucleic acid construct as described above; and (b) an R2 retroelement polypeptide or a second nucleic acid encoding the R2 retroelement polypeptide; wherein the R2 retroelement polypeptide inserts the template sequence of the nucleic acid construct, or a reverse complement thereof, into the target nucleic acid.

[0061] In some embodiments, the target nucleic acid comprises an rRNA gene.

[0062] In some embodiments, the cell is derived from a human subject.INCORPORATION BY REFERENCE

[0063] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Various features of the disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings of which:

[0065] FIG. 1 illustrates a schematic of an example 2-RNA system for TPRT -mediated insertion and tail architectures of a template RNA.

[0066] FIG. 2 illustrates a schematic of an example experimental setup and reporting format to test template RNA tail architectures for TPRT-mediated insertion.

[0067] FIG. 3 illustrates example data of the performance of template RNA tail architectures for TPRT-mediated insertion with different RNA tails across cell lines.

[0068] FIG. 4 illustrates example data of the performance of template RNA tail architectures for TPRT-mediated insertion with different RNA tails separated by cell lines (RPE1, Huh7, and BNL.CL.2).

[0069] FIG. 5 illustrates example data of the performance of template RNA tail architectures for TPRT-mediated insertion with different RNA tails separated by uridine modification (N1 -methyl pseudouridine, pseudouridine, and uridine).

[0070] FIG. 6 illustrates example data of the performance of template RNA tail architectures for TPRT-mediated insertion with N1 -methyl pseudouridine modified RNA tails across cell lines.

[0071] FIG. 7 illustrates example data of the performance of template RNA tail architectures for TPRT-mediated insertion with N1 -methyl pseudouridine modified RNA tails separated by cell lines (RPE1, Huh7, and BNL.CL.2).

[0072] FIG. 8 illustrates example data of the performance of template RNA tail architectures for TPRT-mediated insertion with pseudouridine modified RNA tails across cell lines.

[0073] FIG. 9 illustrates example data of the performance of template RNA tail architectures for TPRT-mediated insertion with pseudouridine modified RNA tails separated by cell lines (RPE1, Huh7, and BNL.CL.2).

[0074] FIG. 10 illustrates example data of the performance of template RNA tail architectures for TPRT-mediated insertion with unmodified (uridine) RNA tails across cell lines.

[0075] FIG. 11 illustrates example data of the performance of template RNA tail architectures for TPRT-mediated insertion with unmodified (uridine) RNA tails separated by cell lines (RPE1, Huh7, and BNL.CL.2).

[0076] FIG. 12 illustrates example data of the performance of template RNA tail architectures for TPRT-mediated insertion with RNA tails separated by intervening non-adenine (non-A) nucleotide (U, G, C, and none) and uridine modification (N1 -methyl pseudouridine, pseudouridine, and uridine) across cell lines.

[0077] FIG. 13 illustrates example data of the performance of template RNA tail architectures for TPRT-mediated insertion with Nl-methyl pseudouridine or pseudouridine modified RNA tails across cell lines.

[0078] FIG. 14 illustrates example data of the performance of template RNA tail architectures for TPRT-mediated insertion with Nl-methyl pseudouridine or pseudouridine modified RNA tails separated by cell lines (RPE1, Huh7, and BNL.CL.2).

[0079] FIG. 15 illustrates example data of the performance of template RNA tail architectures for TPRT-mediated insertion RNA tails separated by intervening non-A nucleotide (U, G, C, and none) and intervening stability elements (eK5, K3, K4, none, and other K-element) across cell lines.

[0080] FIG. 16 illustrates example data of the performance of template RNA tail architectures for TPRT-mediated insertion without intervening stability elements between the 3’ module and adenine-rich region across cell lines.

[0081] FIG. 17 illustrates example data of the performance of template RNA tail architectures for TPRT-mediated insertion without intervening stability elements between the 3’ module and adenine-rich region transcribed using Nl-methyl pseudouridine or pseudouridine across cell lines.

[0082] FIG. 18 illustrates example data of the performance of template RNA tail architectures for TPRT -mediated insertion without intervening stability elements between the 3’ module and adenine-rich region separated by cell lines (RPE1, Huh7, and BNL.CL.2).

[0083] FIG. 19 illustrates example data of the performance of template RNA tail architectures for TPRT -mediated insertion without intervening stability elements between the 3’ module and adenine-rich region transcribed using N1 -methyl pseudouridine or pseudouridine separated by cell lines (RPE1, Huh7, and BNL.CL.2).

[0084] FIG. 20 illustrates example data of the performance of template RNA tail architectures for TPRT -mediated insertion without intervening stability elements between the 3’ module and adenine-rich region transcribed using N1 -methyl pseudouridine across cell lines.

[0085] FIG. 21 illustrates example data of the performance of template RNA tail architectures for TPRT -mediated insertion without intervening stability elements between the 3’ module and adenine-rich region transcribed using N1 -methyl pseudouridine separated by cell lines (RPE1, Huh7, and BNL.CL.2).

[0086] FIG. 22 illustrates example data of the performance of template RNA tail architectures for TPRT -mediated insertion without intervening stability elements between the 3’ module and adenine-rich region transcribed using pseudouridine across cell lines.

[0087] FIG. 23 illustrates example data of the performance of template RNA tail architectures for TPRT -mediated insertion without intervening stability elements between the 3’ module and adenine-rich region transcribed using pseudouridine separated by cell lines (RPE1, Huh7, and BNL.CL.2).

[0088] FIG. 24 illustrates example data of the performance of template RNA tail architectures for TPRT -mediated insertion without intervening stability elements between the 3’ module and adenine-rich region transcribed using uridine across cell lines.

[0089] FIG. 25 illustrates example data of the performance of template RNA tail architectures for TPRT -mediated insertion without intervening stability elements between the 3’ module and adenine-rich region transcribed using uridine separated by cell lines (RPE1, Huh7, and BNL.CL.2).

[0090] FIG. 26 illustrates example data of the performance of template RNA tail architectures for TPRT -mediated insertion without intervening stability elements between the 3’ module separated by intervening non-A nucleotide (U, G, C, and none) and uridine modification (Nl- methyl pseudouridine, pseudouridine, and uridine) across cell lines.

[0091] FIG. 27 illustrates example data of the performance of template RNA tail architectures for TPRT -mediated insertion without intervening stability elements between the 3’ module separated by intervening non-A nucleotide (U, G, C, and none), uridine modification (N1 -methylpseudouridine, pseudouridine, and uridine), and fraction of intervening non-A nucleotides (0.1, 0.2, 0.3, and 0) across cell lines.DETAILED DESCRIPTION

[0092] While various embodiments of the disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the disclosure. It should be understood that various alternatives to the embodiments of the disclosure described herein may be employed.I. Compositions

[0093] Provided herein are nucleic acid constructs for introducing a transgene or other nucleic acid to a subject.

[0094] The nucleic acid construct comprises the following elements from 5’ to 3’: a 5’ region, a template sequence, and a 3’ region comprising a tail region. In some embodiments, the tail region comprises a four-nucleotide complementarity region and an adenine-rich region of at least 22 nucleotides. In some embodiments, at least about 50% of the nucleobases in the adenine-rich region are adenine. In some embodiments, at least about 55% of the nucleobases in the adenine- rich region are adenine. In some embodiments, at least about 60% of the nucleobases in the adenine-rich region are adenine. In some embodiments, at least about 65% of the nucleobases in the adenine-rich region are adenine. In some embodiments, at least about 70% of the nucleobases in the adenine-rich region are adenine. In some embodiments, at least about 80% of the nucleobases in the adenine-rich region are adenine. In some embodiments, at least about 90% of the nucleobases in the adenine-rich region are adenine. In some embodiments, the nucleic acid construct comprises RNA.

[0095] In some embodiments, the tail region further comprises a stability element. In some embodiments, the stability element is 5’ to the adenine-rich region. In some embodiments, the stability element is 3’ to the adenine-region. In some embodiments, the stability element is 3’ to the four-nucleotide complementarity region. In some embodiments, the stability element is 3’ to the four-nucleotide complementarity region and 5’ to the adenine-rich region. In some embodiments, there is no intervening sequence between the four-nucleotide complementarity region and the adenine-rich region.

[0096] In some embodiments, the tail region does not comprise a stability element.

[0097] In some embodiments, the four-nucleotide complementarity region is 5’ to the adenine- rich region and / or the stability element. In some embodiments, the adenine-rich region is located 3’ of the four-nucleotide complementarity region in the nucleic acid construct. The four- nucleotide complementarity region is complementary to a target site in the genome. In someembodiments, the complementarity region is complementary to a region of a rRNA gene. In some embodiments, the rRNA gene is a 28S rRNA gene. In some embodiments, the complementarity region has a sequence of UAGC or TAGC.

[0098] In some embodiments, the tail region comprises the following elements from 5’ to 3’ : a four-nucleotide complementarity region and an adenine-rich region. In some embodiments, the tail region comprises the following elements from 5’ to 3’ : a four-nucleotide complementarity region, a stability element, and an adenine-rich region. In some embodiments, the tail region comprises the following elements from 5’ to 3’ : a four-nucleotide complementarity region, an adenine-rich region, and a stability element.

[0099] In some embodiments, the adenine-rich region is between 19 and 41 nucleotides in length. In some embodiments, the adenine-rich region is 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, or 41 nucleotides in length. In some embodiments, the adenine-rich region is 22, 32, or 41 nucleotides in length. In some embodiments, the adenine-rich region is 19, 29, or 38 nucleotides in length. In some embodiments, the adenine-rich region is 19 nucleotides in length. In some embodiments, the adenine-rich region is 22 nucleotides in length. In some embodiments, the adenine-rich region is 29 nucleotides in length. In some embodiments, the adenine-rich region is 32 nucleotides in length. In some embodiments, the adenine-rich region is 38 nucleotides in length. In some embodiments, the adenine-rich region is 41 nucleotides in length.

[0100] In some embodiments, at least one nucleobase in the adenine-rich region is a non-adenine nucleobase. In some embodiments, the adenine-rich region has between 0% and about 50% non- adenine nucleobases. In some embodiments, the adenine-rich region has 0% to about 5% non- adenine nucleobases. In some embodiments, the adenine-rich region has about 5% to about 10% non-adenine nucleobases. In some embodiments, the adenine-rich region has about 10% to about 15% non-adenine nucleobases. In some embodiments, the adenine-rich region has about 15% to about 20% non-adenine nucleobases. In some embodiments, the adenine-rich region has about 20% to about 25% non-adenine nucleobases. In some embodiments, the adenine-rich region has about 25% to about 30% non-adenine nucleobases. In some embodiments, the adenine-rich region has about 30% to about 35% non-adenine nucleobases. In some embodiments, the adenine-rich region has about 35% to about 40% non-adenine nucleobases. In some embodiments, the adenine-rich region has about 40% to about 45% non-adenine nucleobases. In some embodiments, the adenine-rich region has about 45% to about 50% non-adenine nucleobases. In some embodiments, the adenine-rich region has 0% non-adenine nucleobases. In some embodiments, the adenine-rich region has no more than about 10% non-adenine nucleobases. In some embodiments, the adenine-rich region has no more than about 20% non-adenine nucleobases. In some embodiments, the adenine-rich region has no more than about 30% non-adenine nucleobases. In some embodiments, the adenine-rich region has no more than about 40% non-adenine nucleobases. In some embodiments, the adenine-rich region has no more than about 50% non-adenine nucleobases. In some embodiments, the adenine-rich region has about 10% non-adenine nucleobases. In some embodiments, the adenine-rich region has about 20% non-adenine nucleobases. In some embodiments, the adenine-rich region has about 30% non- adenine nucleobases. In some embodiments, the adenine-rich region has about 35% non-adenine nucleobases. In some embodiments, the adenine-rich region has about 40% non-adenine nucleobases. In some embodiments, the adenine-rich region has about 45% non-adenine nucleobases. In some embodiments, the adenine-rich region has about 50% non-adenine nucleobases.

[0101] In some embodiments, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40% or at least 45% of nucleobases in the adenine-rich region are pyrimidine nucleobases. In some embodiments, the pyrimidine nucleobases comprise uracil. In some embodiments, all of the pyrimidine nucleobases of the adenine-rich region are uracils. In some embodiments, the pyrimidine nucleobases comprise cytosine. In some embodiments, all of the pyrimidine nucleobases of the adenine-rich region are cytosines.

[0102] In some embodiments, the adenine-rich region has at least about 10% non-adenine nucleobases. In some embodiments, the tail region comprises a sequence having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 13, 15, 20, 23, 28, 29, 31, 42, 43, 45, 46, 60, 69, 72, 73, 74, 77, 78, 83, 86, 89, 14, 17, 18, 19, 22, 27, 35, 39, 41, 51, 53, 59, 63, 67, 68, 75, 76, 81, 82, 85, 16, 21, 24, 25, 32, 33, 34, 36, 37, 38, 40, 44, 48, 52, 54, 55, 56, 58, 61, and 79. In some embodiments, the adenine-rich region comprises a sequence having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 128-178. In some embodiments, the adenine-rich region comprises a sequence having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or 100% sequence identity to any one of SEQ ID NOs: 186, 193, 214, 203, 192, 191, 195, 201, 208, 200, 206, 199, 187, 190, 213, 196, 212, 210, 188, 198, 204, 211, 209, 194, 202, 189, or 205. In some embodiments, the adenine-rich region has at least about 20% non-adenine nucleobases. In some embodiments, the tail region comprises a sequence having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 14, 17, 18, 19, 22, 27, 35, 39, 41, 51, 53, 59, 63, 67, 68, 75, 76, 81, 82, 85, 16, 21, 24, 25, 32, 33, 34, 36, 37, 38, 40, 44, 48, 52, 54, 55, 56, 58, 61, and 79. In some embodiments, the adenine-rich region comprises a sequence having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100%sequence identity to any one of SEQ ID NOs: 133-143, 151-162, or 168-178. In some embodiments, the adenine-rich region comprises a sequence having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or 100% sequence identity to any one of SEQ ID NOs: 203, 192, 191, 195, 201, 208, 187, 190, 213, 196, 212, 210, 211, 209, 194, 202, 189, or 205. In some embodiments, the adenine-rich region has at least about 30% non-adenine nucleobases. In some embodiments, the tail region comprises a sequence having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 16, 21, 24, 25, 32, 33, 34, 36, 37, 38, 40, 44, 48, 52, 54, 55, 56, 58, 61, and 79. In some embodiments, the adenine-rich region comprises a sequence having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 139-143, 157-162, or 173-178. In some embodiments, the adenine-rich region comprises a sequence having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or 100% sequence identity to any one of SEQ ID NOs: 195, 201, 208, 196, 212, 210, 202, 189, or 205. In some embodiments, the adenine-rich region has at least about 40% non-adenine nucleobases.

[0103] In some embodiments, the non-adenine nucleobases comprise any one of cytosine, guanine, uridine, or a modified uridine. In some embodiments, the non-adenine nucleobases are pyrimidines (e.g., cytosines, uridines, or a modified uridines). In some embodiments, the non- adenine nucleobases comprise modified uridines. In some embodiments, the modified uridine is pseudouridine or N1 -methylpseudouridine. In some embodiments, the non-adenine nucleobases comprise pseudouridine. In some embodiments, the non-adenine nucleobases comprise Nl- methylpseudouridine. In some embodiments, the non-adenine nucleobases are the same. In some embodiments, the non-adenine nucleobases are different.

[0104] In some embodiments, the adenine-rich region comprises, at its 3’ end, a tail end sequence. In some embodiments, the tail end sequence is selected from the group consisting of: AAA, AAX, and AXX, wherein X is a nucleotide comprising a non-adenine nucleobase. In some embodiments, the tail end sequence is AAA. In some embodiments, the tail end sequence is AAX, wherein X is a nucleotide comprising a non-adenine nucleobase. In some embodiments, the tail end sequence is AXX, wherein X is a nucleotide comprising a non-adenine nucleobase. In some embodiments, X is C, G, or U. In some embodiments, there is no additional sequence that is 3’ to the adenine-rich region.

[0105] In some embodiments, the nucleic acid construct further comprises a tail end sequence that is located 3’ of the adenine rich region. In some embodiments, the tail end sequence is selected from the group consisting of: AAA, AAX, and AXX, wherein X is a nucleotide comprising a non-adenine nucleobase. In some embodiments, the tail end sequence is AAA. Insome embodiments, the tail end sequence is AAX, wherein X is a nucleotide comprising a nonadenine nucleobase. In some embodiments, the tail end sequence is AXX, wherein X is a nucleotide comprising a non-adenine nucleobase. In some embodiments, X is C, G, or U. In some embodiments, there is no additional sequence that is 3’ to the tail end sequence.

[0106] For example, an adenine-rich region with a tail end can comprise any one of SEQ ID NOs: 118-178. For example, an adenine-rich region without a tail end can comprise any one of SEQ ID NOs: 185-214.

[0107] In some embodiments, the three, 3 ’-most terminal nucleotides of the adenine-rich region each comprise adenine, e.g., the sequence of the three terminal nucleotides of the adenine-rich region is AAA, where A is a nucleotide comprising adenine. In some embodiments, 3 ’-most terminal nucleotide of the adenine-rich region does not comprise adenine. In some embodiments, the sequence of the three 3 ’-most terminal nucleotides of the adenine-rich region is AAX, where A is a nucleotide comprising adenine and X is a nucleotide comprising a non-adenine nucleobase. In some embodiments, X is C, G, or U. In some embodiments, the two 3 ’-most terminal nucleotides of the adenine region do not comprise adenine. In some embodiments, the sequence of the three 3 ’-most terminal nucleotides of the adenine-rich region is AXX, where A is a nucleotide comprising adenine and X is a nucleotide comprising a non-adenine nucleobase.

[0108] In some embodiments, the adenine-rich region comprises any one of SEQ ID NOs: 118- 178. In some embodiments, the nucleic acid construct comprises a sequence having at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity to any one of SEQ ID NOs: 118-178. In some embodiments, the adenine-rich region comprises a sequence having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or 100% sequence identity to any one of SEQ ID NOs: 185-214.

[0109] In some embodiments, the stability element is selected from K3, K4, eK5, MALATl mm, MALATl hs, MALAT1 Jis-core, PAN KSHV, dENE os, HSL, U7, K9, K16, KEMCV, or KBDV1. In some embodiments, the stability element comprises a sequence having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 90-117.

[0110] In some embodiments, the tail region of the nucleic acid construct comprises a sequence having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 2-89.[OHl] In some embodiments, the 5’ region further comprises a 5’ module, a promoter, and a 5’UTR. In some embodiments, the 5’ module may comprise a sequence derived from a native retroelement 5’ region, an rRNA sequence, a ribozyme sequence, a folding motif sequence, and / or an RNA polymerase terminator sequence.

[0112] In some embodiments, the nucleic acid construct may further comprise one or more sequences in the 3’ modules for reverse transcriptase (RT)-mediated TPRT. In some embodiments, the 3’ region may comprise an untranslated region (UTR sequence) that is cognate to the paired retroelement or modified from a native cognate. As used herein, the term “cognate” when referring to sequences means that the sequences are from the same species.

[0113] In some embodiments, the nucleic acid construct may further comprise one or more 5’ modules for RT -mediated TPRT. The 5’ module may be cognate to the paired retroelement or modified from a native cognate.

[0114] In some embodiments, the template sequence of the nucleic acid construct may have a length of at least about 200 base pairs. The template sequence may have a length of at least about 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, or 2000 base pairs. The template may have a length of longer than about 2000 base pairs.

[0115] In some embodiments, the template sequence of the nucleic acid construct comprises a transgene. In some embodiments, the transgene is a therapeutically active gene or a diagnostic gene. The transgene may comprise PKU, OTC, MMUT, ATP7B, F8, F9, or SERPINA1. The template sequence may comprise one or more non -native transgenes that rescue loss of function in a human disease or confer beneficial function.

[0116] In some embodiments, the nucleic acid construct may comprise native or non-native additions or modifications. Such additions or modifications may be additional nucleic acid or nucleic acid-like material, chemically synthetic components, natural or synthetic peptides or lipids, scaffold attachment and release capability, and others.

[0117] In some embodiments, any one of the nucleic acid constructs herein has increased stability as compared to an otherwise identical nucleic acid construct having a tail region consisting of SEQ ID NO: 1. In some embodiments, increased stability is assessed by determining a level of gene expression from the template sequence following insertion into a target nucleic acid of a cell. In some embodiments, any one of the nucleic acid constructs herein has increased exonuclease resistance as compared to an otherwise identical nucleic acid construct having a tail region consisting of SEQ ID NO: 1. In some embodiments, increased exonuclease resistance is assessed by determining a level of gene expression from the template sequence following insertion into a target nucleic acid of a cell. In some embodiments, the target nucleic acid comprises a rRNA gene.

[0118] In some embodiments, any one of the nucleic acid constructs herein is more efficiently inserted into a genome by target primed reverse transcription as compared to an otherwise identical nucleic acid construct having a tail region consisting of SEQ ID NO: 1. In someembodiments, insertion efficiency is assessed by determining a level of gene expression from the template sequence following insertion into a target nucleic acid of a cell.

[0119] In some embodiments, the adenine-rich region comprises a sequence having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 118-178. In some embodiments, the adenine-rich region comprises a sequence having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 174, 173, 143, 171, 130, 135, 165, 170, 139, 137, 177, 129, 178, or 167. In some embodiments, the adenine-rich region comprises a sequence having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or 100% sequence identity to any one of SEQ ID NOs: 202, 208, 194, 193, 192, 204, 209, 195, 191, 189, 193, 205, or 198. In some embodiments, the adenine-rich region comprises a sequence having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 174, 173, 143, 171, 135, 170, 139, 137, 177, or 178. In some embodiments, the adenine-rich region comprises a sequence having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or 100% sequence identity to any one of SEQ ID NOs: 202, 208, 194, 192, 209, 195, 191, 189, or 205. In some embodiments, the adenine-rich region comprises a sequence having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 174, 173, 143, 139, 177, or 178. In some embodiments, the adenine-rich region comprises a sequence having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or 100% sequence identity to any one of SEQ ID NOs: 202, 208, 195, 189, or 205. In some embodiments, the adenine-rich region comprises a sequence having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 37, 58, 48, 35, 73, 68, 43, 51, 32, 81, 16, 69, 44, or 42. In some embodiments, the adenine-rich region comprises a sequence having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 174, 173, 143, 171, 118, 120-122, 125, 126, 129, 130, 135, 137, 139, 165, 167, 170, 177, and 178. In some embodiments, the adenine-rich region comprises a sequence having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or 100% sequence identity to any one of SEQ ID NOs: 202, 108, 194, 185, 207, 193, 192, 191, 195, 204, 198, 209, 189, or 205. In some embodiments, the nucleic acid construct comprises a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 13, 15, 16, 18, 19, 20, 21, 22, 23, 24, 28, 32, 33, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 47, 48, 50, 51, 53, 56, 57, 58, 60, 62, 67, 68, 69, 70, 71, 73, 75, 77, 78, 81, 82, 85, 86, 87, and 89. In some embodiments, the nucleic acid construct comprises a tail region having any one of SEQ ID NOs: 13, 18, 19, 20, 23, 24, 32, 36, 38, 39, 40, 41, 47, 48,56, 57, 67, 68, 69, 70, 73, 77, 81, and 86. In some embodiments, the nucleic acid construct comprises a tail region having any one of SEQ ID NOs: 15, 16, 21, 22, 28, 33, 35, 37, 42, 43, 44, 50, 51, 53, 58, 60, 62, 71, 75, 78, 82, 85, 87, and 89. In some embodiments, the nucleic acid construct comprises a tail region having any one of SEQ ID NOs: 15, 16, 21, 22, 28, 33, 35, 37, 42, 43, 44, 50, 51, 53, 58, 60, 62, 71, 75, 78, 82, 85, 87, and 89, and comprises Nl- methylpseudouridine or pseudouridine nucleobases. In some embodiments, the nucleic acid construct comprises a tail region having any one of SEQ ID NOs: 15, 16, 21, 22, 28, 33, 35, 37, 42, 43, 44, 50, 51, 53, 58, 60, 62, 71, 75, 78, 82, 85, 87, and 89, and comprises pseudouridine nucleobases. In some embodiments, the nucleic acid construct comprises a tail region having any one of SEQ ID NOs: 15, 16, 21, 22, 28, 33, 35, 37, 42, 43, 44, 50, 51, 53, 58, 60, 62, 71, 75, 78, 82, 85, 87, and 89, and comprises N1 -methylpseudouridine nucleobases.

[0120] In some embodiments, the nucleic acid construct comprises a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 2, 3, 4, 5, 6, 7, 8, 14, 29, 35, 37, 43, 46, 47, 48, 50, 51, 52, 54, 57, 58, 64, 68, 69, 73, and 81. In some embodiments, the nucleic acid construct comprises an adenine-rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 118, 152, 148, 171, 174, 165, 144, 121, 143, 122, 170, 160, 159, 125, 173, 126, 135, 129, 130, and 137. In some embodiments, the nucleic acid construct comprises a tail region having any one of SEQ ID NOs: 35, 37, 43, 47, 48, 50, 51, 57, 58, 68, 69, 73, and 81. In some embodiments, the nucleic acid construct comprises an adenine-rich region having any one of SEQ ID NOs : 171, 174, 165, 121, 143, 122, 170, 125,173, 135, 129, 130, and 137. In some embodiments, the nucleic acid construct comprises a tail region having any one of SEQ ID NOs: 47, 48, 57, 68, 69, 73, and 81. In some embodiments, the nucleic acid construct comprises an adenine-rich region having any one of SEQ ID NOs : 121, 143, 125, 135, 129, 130, and 137. In some embodiments, the nucleic acid construct comprises a tail region having any one of SEQ ID NOs: 35, 37, 43, 50, 51, and 58. In some embodiments, the nucleic acid construct comprises an adenine-rich region having any one of SEQ ID NOs : 171,174, 165, 122, 170, and 173. In some embodiments, the nucleic acid construct comprises a tail region having any one of SEQ ID NOs: 35, 37, 43, 50, 51, and 58, and comprises Nl- methylpseudouridine or pseudouridine nucleobases. In some embodiments, the nucleic acid construct comprises a tail region having any one of SEQ ID NOs: 35, 37, 43, 50, 51, and 58, and comprises pseudouridine nucleobases. In some embodiments, the nucleic acid construct comprises a tail region having any one of SEQ ID NOs: 35, 37, 43, 50, 51, and 58, and comprises N1 -methylpseudouridine nucleobases. In some embodiments, the nucleic acid construct comprises an adenine-rich region having any one of SEQ ID NOs : 171, 174, 165, 122, 170, and173, and comprises N1 -methylpseudouridine or pseudouridine nucleobases. In some embodiments, the nucleic acid construct comprises an adenine-rich region having any one of SEQ ID NOs : 171, 174, 165, 122, 170, and 173, and comprises pseudouridine nucleobases. In some embodiments, the nucleic acid construct comprises an adenine-rich region having any one of SEQ ID NOs : 171, 174, 165, 122, 170, and 173, and comprises N1 -methylpseudouridine nucleobases.

[0121] In some embodiments, the nucleic acid construct comprises a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 3. In some embodiments, the nucleic acid construct comprises an adenine-rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 3. In some embodiments, the nucleic acid construct comprises a tail region having any one of SEQ ID NOs: 37, 58, 48, 73, 68, 43, and 35. In some embodiments, the nucleic acid construct comprises an adenine-rich region having any one of SEQ ID NOs: 174, 173, 143, 130, 135, 165, and 171. In some embodiments, the nucleic acid construct comprises a tail region having any one of SEQ ID NOs: 48, 73, and 68. In some embodiments, the nucleic acid construct comprises an adenine-rich region having any one of SEQ ID NOs: 143, 130, and 135. In some embodiments, the nucleic acid construct comprises a tail region having any one of SEQ ID NOs: 37, 58, 43, and 35. In some embodiments, the nucleic acid construct comprises an adenine-rich region having any one of SEQ ID NOs: 174, 173, 165, and 171.

[0122] In some embodiments, the composition comprises a retinal cell or an epithelial cell (e.g., RPE1 cell) and the nucleic acid construct comprising a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 4. In some embodiments, the composition comprises a retinal cell or an epithelial cell (e.g., RPE1 cell) and the nucleic acid construct comprising an adenine- rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 4. In some embodiments, the composition comprises a cancer cell and the nucleic acid construct comprising a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 4. In some embodiments, the composition comprises a cancer cell and the nucleic acid construct comprising an adenine-rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 4. In some embodiments, the composition comprises a retinal cell, an epithelial cell, or a cancer cell and the nucleic acid construct comprising a tail region having any one of SEQ ID NOs: 37, 58, 48, 73,35, 81, 16, 69, 51, 68, 44, 56, 43, and 42. In some embodiments, the composition comprises a retinal cell, an epithelial cell, or a cancer cell and the nucleic acid construct comprising an adenine-rich region having any one of SEQ ID NOs: 174, 173, 143, 130, 171, 137, 177, 129, 170, 135, 178, 165, and 167. In some embodiments, the composition comprises a retinal cell, an epithelial cell, or a cancer cell and the nucleic acid construct comprising a tail region having any one of SEQ ID NOs: 48, 73, 81, 69, 68, and 56. In some embodiments, the composition comprises a retinal cell, an epithelial cell, or a cancer cell and the nucleic acid construct comprising an adenine-rich region having any one of SEQ ID NOs: 143, 130, 137, 129, and 135. In some embodiments, the composition comprises a retinal cell, an epithelial cell, or a cancer cell and the nucleic acid construct comprising a tail region having any one of SEQ ID NOs: 37, 58, 35, 16, 51, 44, 43, and 42. In some embodiments, the composition comprises a retinal cell, an epithelial cell, or a cancer cell and the nucleic acid construct comprising an adenine-rich region having any one of SEQ ID NOs: 174, 173, 171, 177, 170, 178, 165, and 167.

[0123] In some embodiments, the composition comprises a liver cell (e.g., Huh7 cell) and the nucleic acid construct comprising a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 5. In some embodiments, the composition comprises a liver cell (e.g., Huh7 cell) and the nucleic acid construct comprising an adenine-rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 5. In some embodiments, the composition comprises a cancer cell (e.g., Huh7 cell) and the nucleic acid construct comprising a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 5. In some embodiments, the composition comprises a cancer cell (e.g., Huh7 cell) and the nucleic acid construct comprising an adenine-rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 5. In some embodiments, the composition comprises a liver cell or cancer cell and the nucleic acid construct comprising a tail region having any one of SEQ ID NOs: 37, 48, 58, 35, 68, 42, and 73. In some embodiments, the composition comprises a liver cell or cancer cell and the nucleic acid construct comprising an adenine-rich region having any one of SEQ ID NOs: 174, 143, 173, 171, 135, 167, and 130. In some embodiments, the composition comprises a liver cell or cancer cell and the nucleic acid construct comprising a tail region having any one of SEQ ID NOs: 48, 68, and 73. In some embodiments, the composition comprises a liver cell or cancer cell and the nucleic acid construct comprising an adenine-rich region having any one of SEQ ID NOs: 143, 135, and 130. In some embodiments, the composition comprises a liver cell or cancer cell and the nucleic acid constructcomprising a tail region having any one of SEQ ID NOs: 37, 58, 35, and 42. In some embodiments, the composition comprises a liver cell or cancer cell and the nucleic acid construct comprising an adenine-rich region having any one of SEQ ID NOs: 174, 173, 171, and 167.

[0124] In some embodiments, the composition comprises a liver cell, an embryonic cell, or an epithelial cell (e.g., BNL.CL.2 cell) and the nucleic acid construct comprising a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 6. In some embodiments, the composition comprises a liver cell, an embryonic cell, or an epithelial cell (e.g., BNL.CL.2 cell) and the nucleic acid construct comprising an adenine-rich having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 6. In some embodiments, the composition comprises a cancer cell and the nucleic acid construct comprising a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 6. In some embodiments, the composition comprises a cancer cell and the nucleic acid construct comprising an adenine-rich having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 6. In some embodiments, the composition comprises a liver cell, an embryonic cell, an epithelial cell, or a cancer cell and the nucleic acid construct comprising a tail region having any one of SEQ ID NOs: 37, 58, 43, 51, and 32. In some embodiments, the composition comprises a liver cell, an embryonic cell, an epithelial cell, or a cancer cell and the nucleic acid construct comprising an adenine-rich region having any one of SEQ ID NOs: 174, 173, 165, 170, and 139. In some embodiments, the composition comprises a liver cell, an embryonic cell, an epithelial cell, or a cancer cell and the nucleic acid construct comprising a tail region having any one of SEQ ID NOs: 32. In some embodiments, the composition comprises a liver cell, an embryonic cell, an epithelial cell, or a cancer cell and the nucleic acid construct comprising an adenine-rich region having any one of SEQ ID NOs: 139. In some embodiments, the composition comprises a liver cell, an embryonic cell, an epithelial cell, or a cancer cell and the nucleic acid construct comprising a tail region having any one of SEQ ID NOs: 37, 58, 43, and 51. In some embodiments, the composition comprises a liver cell, an embryonic cell, an epithelial cell, or a cancer cell and the nucleic acid construct comprising an adenine-rich region having any one of SEQ ID NOs: 174, 173, 165, and 170.

[0125] In some embodiments, the nucleic acid construct comprises a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 19. In some embodiments, the nucleic acid construct comprises an adenine-rich region having at least 60%, at least 70%, at least 80%, atleast 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 19. In some embodiments, the nucleic acid construct comprises a tail region having any one of SEQ ID NOs: 37, 58, 48, 73, 68, 43, and 35. In some embodiments, the nucleic acid construct comprises an adenine-rich region having any one of SEQ ID NOs: 174, 173, 143, 130, 135, 165, and 171. In some embodiments, the nucleic acid construct comprises a tail region having any one of SEQ ID NOs: 48, 73, and 68. In some embodiments, the nucleic acid construct comprises an adenine-rich region having any one of SEQ ID NOs: 143, 130, and 135. In some embodiments, the nucleic acid construct comprises a tail region having any one of SEQ ID NOs: 37, 58, 43, and 35. In some embodiments, the nucleic acid construct comprises an adenine-rich region having any one of SEQ ID NOs: 174, 173, 165, and 171.

[0126] In some embodiments, the composition comprises a retinal cell or an epithelial cell (e.g., RPE1 cell) and the nucleic acid construct comprising a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 20. In some embodiments, the composition comprises a retinal cell or an epithelial cell (e.g., RPE1 cell) and the nucleic acid construct comprising an adenine- rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 20. In some embodiments, the composition comprises a cancer cell and the nucleic acid construct comprising a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 20. In some embodiments, the composition comprises a cancer cell and the nucleic acid construct comprising an adenine-rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 20. In some embodiments, the composition comprises a retinal cell, an epithelial cell, or a cancer cell and the nucleic acid construct comprising a tail region having any one of SEQ ID NOs: 37, 58, 48, 73, 35, 81, 69, 51, 68, and 43. In some embodiments, the composition comprises a retinal cell, an epithelial cell, or a cancer cell and the nucleic acid construct comprising an adenine-rich region having any one of SEQ ID NOs: 174, 173, 143, 130, 137, 129, 170, 135, 165, and 171. In some embodiments, the composition comprises a retinal cell, an epithelial cell, or a cancer cell and the nucleic acid construct comprising a tail region having any one of SEQ ID NOs: 48, 73, 81, 69, and 68. In some embodiments, the composition comprises a retinal cell, an epithelial cell, or a cancer cell and the nucleic acid construct comprising an adenine-rich region having any one of SEQ ID NOs: 143, 130, 137, 129, and 135. In some embodiments, the composition comprises a retinal cell, an epithelial cell, or a cancer cell and the nucleic acid construct comprising a tail region having any one of SEQ ID NOs: 37, 58, 35, 51, and 43. In some embodiments, thecomposition comprises a retinal cell, an epithelial cell, or a cancer cell and the nucleic acid construct comprising an adenine-rich region having any one of SEQ ID NOs: 174, 173, 170, 165, and 171.

[0127] In some embodiments, the composition comprises a liver cell (e.g., Huh7 cell) and the nucleic acid construct comprising a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 21. In some embodiments, the composition comprises a liver cell (e.g., Huh7 cell) and the nucleic acid construct comprising an adenine-rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 21. In some embodiments, the composition comprises a cancer cell (e.g., Huh7 cell) and the nucleic acid construct comprising a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 21. In some embodiments, the composition comprises a cancer cell (e.g., Huh7 cell) and the nucleic acid construct comprising an adenine-rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 21. In some embodiments, the composition comprises a liver cell or cancer cell and the nucleic acid construct comprising a tail region having any one of SEQ ID NOs: 37, 48, 58, 35, 68, and 73. In some embodiments, the composition comprises a liver cell or cancer cell and the nucleic acid construct comprising an adenine-rich region having any one of SEQ ID NOs: 174, 143, 173, 171, 135, and 130. In some embodiments, the composition comprises a liver cell or cancer cell and the nucleic acid construct comprising a tail region having any one of SEQ ID NOs: 48, 68, and 73. In some embodiments, the composition comprises a liver cell or cancer cell and the nucleic acid construct comprising an adenine-rich region having any one of SEQ ID NOs: 143, 135, and 130. In some embodiments, the composition comprises a liver cell or cancer cell and the nucleic acid construct comprising a tail region having any one of SEQ ID NOs: 37, 58, and 35. In some embodiments, the composition comprises a liver cell or cancer cell and the nucleic acid construct comprising an adenine-rich region having any one of SEQ ID NOs: 174, 173, and 171.

[0128] In some embodiments, the composition comprises a liver cell, an embryonic cell, or an epithelial cell (e.g., BNL.CL.2 cell) and the nucleic acid construct comprising a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 22. In some embodiments, the composition comprises a liver cell, an embryonic cell, or an epithelial cell (e.g., BNL.CL.2 cell) and the nucleic acid construct comprising an adenine-rich having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of thesequences listed in Table 22. In some embodiments, the composition comprises a cancer cell and the nucleic acid construct comprising a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 22. In some embodiments, the composition comprises a cancer cell and the nucleic acid construct comprising an adenine-rich having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 22. In some embodiments, the composition comprises a liver cell, an embryonic cell, an epithelial cell, or a cancer cell and the nucleic acid construct comprising a tail region having any one of SEQ ID NOs: 37, 58, 43, and 51. In some embodiments, the composition comprises a liver cell, an embryonic cell, an epithelial cell, or a cancer cell and the nucleic acid construct comprising an adenine-rich region having any one of SEQ ID NOs: 174, 173, 165, and 170.

[0129] In some embodiments, the nucleic acid construct comprises a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 7. In some embodiments, the nucleic acid construct comprises an adenine-rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 7. In some embodiments, the nucleic acid construct comprises a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 8. In some embodiments, the nucleic acid construct comprises an adenine-rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 8. In some embodiments, the nucleic acid construct comprises a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 9. In some embodiments, the nucleic acid construct comprises an adenine-rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 9. In some embodiments, the nucleic acid construct comprises a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 10. In some embodiments, the nucleic acid construct comprises an adenine-rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 10. In some embodiments, the nucleic acid construct comprises a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 11. In some embodiments,the nucleic acid construct comprises an adenine-rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 11. In some embodiments, the nucleic acid construct comprises a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 12. In some embodiments, the nucleic acid construct comprises an adenine-rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 12. In some embodiments, the nucleic acid construct comprises a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 13. In some embodiments, the nucleic acid construct comprises an adenine-rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 13. In some embodiments, the nucleic acid construct comprises a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 14. In some embodiments, the nucleic acid construct comprises an adenine-rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 14. In some embodiments, the nucleic acid construct comprises a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 15. In some embodiments, the nucleic acid construct comprises an adenine-rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 15. In some embodiments, the nucleic acid construct comprises a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 16. In some embodiments, the nucleic acid construct comprises an adenine-rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 16. In some embodiments, the nucleic acid construct comprises a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 17. In some embodiments, the nucleic acid construct comprises an adenine-rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 17. In some embodiments, the nucleic acid construct comprises a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 18. In some embodiments,the nucleic acid construct comprises an adenine-rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 18. In some embodiments, the nucleic acid construct comprises a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 23. In some embodiments, the nucleic acid construct comprises an adenine-rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 23. In some embodiments, the nucleic acid construct comprises a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 24. In some embodiments, the nucleic acid construct comprises an adenine-rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 24. In some embodiments, the nucleic acid construct comprises a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 25. In some embodiments, the nucleic acid construct comprises an adenine-rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 25. In some embodiments, the nucleic acid construct comprises a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 26. In some embodiments, the nucleic acid construct comprises an adenine-rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 26. In some embodiments, the nucleic acid construct comprises a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 27. In some embodiments, the nucleic acid construct comprises an adenine-rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 27. In some embodiments, the nucleic acid construct comprises a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 28. In some embodiments, the nucleic acid construct comprises an adenine-rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 28. In some embodiments, the nucleic acid construct comprises a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 29. In some embodiments,the nucleic acid construct comprises an adenine-rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 29. In some embodiments, the nucleic acid construct comprises a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 30. In some embodiments, the nucleic acid construct comprises an adenine-rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 30. In some embodiments, the nucleic acid construct comprises a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 31. In some embodiments, the nucleic acid construct comprises an adenine-rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 31. In some embodiments, the nucleic acid construct comprises a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 32. In some embodiments, the nucleic acid construct comprises an adenine-rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 32. In some embodiments, the nucleic acid construct comprises a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 33. In some embodiments, the nucleic acid construct comprises an adenine-rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 33. In some embodiments, the nucleic acid construct comprises a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 34. In some embodiments, the nucleic acid construct comprises an adenine-rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 34. In some embodiments, the nucleic acid construct comprises a tail region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 35. In some embodiments, the nucleic acid construct comprises an adenine-rich region having at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to any one of the sequences listed in Table 35.

[0130] Cells

[0131] In some aspects, provided herein is a cell comprising a nucleic acid construct described elsewhere herein or a system described elsewhere herein. In some embodiments, the cell is an epithelial cell. In some embodiments, the cell is a retinal cell. In some embodiments, the cell is an epithelial cell. In some embodiments, the cell is an embryonic cell. In some embodiments, the cell is a liver cell. In some embodiments, the cell is a hepatocyte. In some embodiments, the cell is a RPE1 cell. In some embodiments, the cell is a Huh7 cell. In some embodiments, the cell is a BNL.CL.2 cell. In some embodiments, the cell is a cancer cell. In some embodiments, the cell is a derived from a human subject. In some embodiments, the nucleic acid construct is not endogenous to the cell.II. Systems

[0132] Provided herein are systems for modifying a target nucleic acid in a cell of the subject. The systems are gene-insertion systems, which comprise any one of the nucleic acid constructs described herein and a partnered retroelement polypeptide or second nucleic acid encoding the partnered retroelement polypeptide. The gene-insertion system components cause site-specific addition of the template sequence of the nucleic acid construct to a eukaryotic genome.

[0133] In some embodiments, the partnered retroelement polypeptide, or second nucleic acid encoding the partnered retroelement polypeptide, can comprise the following elements from 5’ to 3’ : a 5’ retroelement polypeptide module, an RT module, and a 3’ retroelement polypeptide module. In some embodiments, the 5’ retroelement polypeptide module may comprise a 5’ untranslated region (5’-UTR), a Kozak sequence or an internal ribosome entry site, a non-native translation start codon, and / or a 5’ cap. The RT module encodes a reverse transcriptase. In some embodiments, the 3’ retroelement polypeptide module may comprises a reverse transcriptase translation stop codon, a 3’ untranslated region (3’ UTR), and / or a poly-A tail.

[0134] In some embodiments, the partnered retroelement polypeptide or second nucleic acid encoding the partnered retroelement polypeptide constructs of the disclosure may comprise components (interchangeably referred to as modules) which may be derived from portions of at least one non-long terminal repeat retroelement (non-LTR) or are not known in nature.

[0135] In some embodiments, the partnered retroelement polypeptide or second nucleic acid encoding the partnered retroelement polypeptide comprises or encodes at least one retroelement protein. In some embodiments, the retroelement proteins may be R2 retroelement proteins, or an R2 / R8 / R9 domain architecture of non-LTR RT proteins, or a naturally occurring protein or protein complex. In some embodiments, the retroelement protein is a reverse transcriptase.

[0136] In some embodiments, the retroelement protein may be a non-LTR retroelement protein containing a TPRT-competent reverse transcriptase (RT) or strand-nicking endonuclease activitythat is active when assayed for RT primer extension or in vitro TPRT, which may be sitespecific.

[0137] Suitable retroelements from which gene-insertion system components may be derived from non-LTR retroelements, for example of the RLE-type or APE-type or Penelope-type. An RLE-type non-LTR retrotransposon may be from any one of many clades, including but not limited to R2, R4, CRE, Genie, HERO, NeSL. An APE-type non-LTR retrotransposon may be from any one of many clades, including but not limited to I, Rl, LI , Txl, CR1, Rexl, Jockey, L2, Tad, RTE, RTEX, ingi, Vingi, TR AS, SARI’, or any combination thereof. In some embodiments, gene-insertion system components may be derived from retroelements that insert into rDNA, i.e., the so-called R elements, such as retroelements of the Rl or R2 clades. In some embodiments, the R2 clade retroelement may have canonical R2 retroelement insertion site specificity or may be derived from an R8 or R9 retroelement in the larger R2 clade that have changed target sequence relative to the canonical R2 retroelements or may be derived from R2NS retroelements that appear to have lost target site specificity.

[0138] Gene-insertion system components may be derived from portions or domains of retroelements found in any species, including those of distant evolutionary relation to the subject. For example, suitable retroelements from which gene-insertion system components may be derived may comprise those found in birds (e.g., Zonotrichia albicollis. Taeniopygia giilala, Tinamus giiUalus. and Geospiza fords), fish (e.g., Pungitis pungilis, Oryzias talipes, Danio rerio, Oryzias melasligma, Petromyzon marinus, Salmo triila, Salmo salar, or Gasterosteus aciilealus), insects (e.g., Drosophila mercalorum. Drosophila melanogasler, Nasonia vitripennis, Tribolium caslaneum. Drosophila simulans. Apis cerana, and Bombyx mori), crustaceans (e.g., Lepidurus couesii, and Triops cancriformis), other invertebrates (e.g., Limulus polyphemus, Hydra magnipapillata, or Adineta vaga), chordates (e.g., Ciona intestinalis) including mammals, and any combination thereof.

[0139] In some embodiments, the retroelement may be partnered to the template sequence such that the efficiency of the insertion of the transgene is improved.III. Methods of Use

[0140] Provided herein are methods for introducing a transgene or other nucleic acid to a subject. The methods may comprise introducing an effective amount of at least one gene-insertion system which comprises a transgene or other nucleic acid to the subject.

[0141] In some embodiments, the cells of the subject may be actively proliferating or expanding. In some embodiments, the cells of the subject may be progressing through the cell cycle. In some embodiments, the cells of the subject may comprise transcriptionally active or replicationally active rDNA.

[0142] In some embodiments, the cells of the subject are eukaryotic cells. For example, the cells may be liver cells, hepatocytes, hepatic stellate cells, Kupffer cells, or liver sinusoidal endothelial cells. In some embodiments, the cells may be epithelial cells. In some embodiments, the cells may be retinal cells. In some embodiments, the cells may be liver cells. In some embodiments, the cells may be hepatocytes. In some embodiments, the cells may be cancerous cells. In some embodiments, the cells may be non-cancerous cells.

[0143] In some embodiments, the method comprises introducing an effective amount of at least one gene-insertion system described herein to the subject. In some embodiments, the template sequence is inserted at a one or more target insertion sites. In some embodiments, the geneinsertion system comprises delivery of a transgene to the subject via delivery of the geneinsertion system described herein.

[0144] Introduction of the gene-insertion system to cells may involve standard methods, such as lipid-enabled transfection, electroporation, or other methods. The gene-insertion system may be encapsulated in one or more delivery vehicles. Delivery vehicles may facilitate in vivo or in vitro transfection of subject cells by protecting gene-insertion system components from degradation in the extracellular environment, facilitating uptake by subject cells, enhancing endosomal escape, or any combination thereof. Delivery vehicles may be lipid-based (e.g., lipid nanoparticles (LNPs), liposomes, and micelles) or non-lipid-based (e.g., virus like particles (VLPs) and polymeric delivery particles).

[0145] For example, the gene-insertion system may be introduced to cells via a lipid nanoparticle (LNP). LNPs possess an exterior lipid layer including a hydrophilic exterior surface that is exposed to the non-LNP environment, a non-aqueous or an aqueous interior space (i.e., micellelike and vesicle-like LNPs respectively), and at least one hydrophobic inter-membrane space. LNP membranes may be non-lamellar or lamellar, with 1, 2, 3, 4, 5 or more than 5 layers. LNPs may be solid or semi-solid. In some embodiments, at least one cargo or a payload (such as the gene-insertion system) is comprised in the interior space, the inter membrane space, on the exterior surface, or any combination thereof of the LNP.

[0146] In some embodiments, the LNPs may comprise an ionizable (cationic) lipid, a phospholipid, cholesterol, and a polymer-conjugated lipid. Cholesterol promotes membrane fusion and aids in LNP stability. In some embodiments, the cholesterol is a sterol. Phospholipids aid in endosomal escape and provide structure to the LNP bilayer. The phospholipid is a noncationic lipid. Polymer-conjugated lipids reduce LNP aggregation and “protect” the LNP from non-specific endocytosis by immune cells. Polymer-conjugated lipids comprise PEG-lipids. Ionizable (cationic) lipids enhance endosomal escape and complex with a negatively chargedcargo (such as polynucleotides of the gene-insertion system). Cationic lipids comprise ionizable cationic lipids.

[0147] In some embodiments, the LNPs have a selective preference for delivery to liver cells. Following injection into circulation, the LNPs may shed the PEG component. The LNPs then bind to apolipoprotein E (ApoE), and then be taken into liver cells, including hepatocytes, via the LDL receptor. Thus, the LNPs facilitate hepato-specific delivery of cargo contained in the LNPs.

[0148] In some embodiments, the method may leverage a modified R2 retroelement protein to support RNA-mediated transgene insertion, like Target Primed Reverse Transcription (TPRT)- initiated transgene insertion. TPRT may be used to modify a target nucleic acid in a mammalian cell rDNA, using a directly introduced RNA template. The systems and methods may involve other species’ genomes as targets for TPRT -mediated transgene insertion, or for non-genomic targets.

[0149] In some embodiments, the method may insert one or more transgenes in human cell 28 S rDNA. The transgenes are functionally expressed. The human rDNA is a safe harbor site for insertion of a successful transgene protein expression cassette.

[0150] The method may also employ various techniques to assess successful gene introduction. For example, In Vivo Imaging System (IVIS) imaging facilitates real-time, non-invasive monitoring of gene expression. IVIS also facilitates evaluation of biological expression patterns of the introduced gene in live animals. Further, non-secreted genes may be evaluated by immunohistochemistry. Similarly, secreted genes may be evaluated by microplate assays. These microplate assays may test blood samples.IV. DEFINITIONS

[0151] Derived from: As used herein, the term “derived from” refers to a nucleic acid or protein sequence that is isolated from or obtained from a specific source, such as a non-long terminal repeat (non-LTR) retrotransposon. The term includes native sequences isolated from or obtained from a specific source. The term also includes man-made variants of sequences from the original source that have the same or similar functional properties, e.g., the variant can comprise a nucleic or amino acid sequence that has been modified from the original source to have improved functional properties compared to the original source molecule.

[0152] If an RNA sequence is recited using deoxyribonucleotides, any thymidines (“T”s) can be replaced with uridines (“U”s) or uridine analogs (e.g., N1 -methylpseudouridine) to convert the DNA sequence to an RNA sequence.

[0153] Encapsulate: As used herein, the term “encapsulate” means to enclose, surround, or encase.

[0154] Encode: As used herein, the term “encode” refers broadly to any process whereby the information in a polymeric macromolecule is used to direct the production of a second molecule that is different from the first. The second molecule may have a chemical structure that is different from the chemical nature of the first molecule.

[0155] Flanking: As used herein, the term “flanking” refers to the positioning of one element either 5' (5' flanking) or 3' (3' flanking) to another element. Elements that are said to be flanking may be directly connected to each other or may have other elements interspaced between them.

[0156] Gene Insertion Construct: As used herein, the term “Gene Insertion Construct”, or GIC, refers to an RNA construct which comprises the RNA template for an RT protein.

[0157] Gene-Insertion System: As used herein, the term “Gene-Insertion System” or “GIS,” is a system of components (modules) which may be used to insert a genetic sequence (transgene) into a location of a subject genome via reverse transcription, including TPRT.

[0158] Liposome: As used herein, “liposome” generally refers to a vesicle composed of lipids (e.g., amphiphilic lipids) arranged in one or more spherical bilayers or bilayers.

[0159] Paired retroelement: As used herein, the term “paired retroelemenf ’ refers to the combination of a reverse transcriptase (RT) or retroelement with at least one of the modules comprising the insertion payload module. A module may be homologous to its paired RT, meaning the RT and all elements in the module are derived from the same retroelement gene. A module may be heterologous to its paired RT, meaning at least one element of the module is not derived from the same retroelement gene as the RT.

[0160] Target Cell: As used herein, the phrase “targeted cells” refers to any one or more cells of interest. The cells may be found in vitro, in vivo, in situ or in the tissue or organ of an organism. The organism may be an animal, preferably a mammal, more preferably a human and most preferably a patient.

[0161] Target Primed Reverse Transcription: As used herein, the term “target primed reverse transcription” refers to any process where a reverse transcriptase uses a genome-embedded nicked DNA 3’ end at the target site as the primer to initiate cDNA synthesis.

[0162] Unless otherwise specified, the term “about” when used before a numerical designation, e.g., time, amount, size, should be understood to include variations of ±10% around the stated value. In case of illumination level, the term “about” should be understood to include variation of ±0.5 log.

[0163] The “percent sequence identity” between a reference amino acid sequence and a query amino sequence (i.e., the amino sequence being analyzed to determine whether it is within a particular percent sequence identity with the reference amino acid sequence) is determined by optimally aligning the sequences using the Needleman-Wunsch alignment algorithm with a gapexistence penalty of 11 and a gap extension penalty of 1 and comparing the sequences. The number of exact matches, divided by the total number of positions in the alignment (which corresponds with the number of amino acids in the reference sequence plus any gaps in the reference sequence when aligned with the query sequence) is determined and expressed as a percentage. This is the percent sequence identity between the query amino acid sequence and the reference amino acid sequence (i.e., percent sequence identity = (# of exact matches / (total # of positions in alignment)* 100). An alignment using the Needleman-Wunsch alignment algorithm (with a gap existence penalty of 11 and a gap extension penalty of 1) can be generated using the “Global Align” BLAST program available at https: / / blast.ncbi.nlm.nih.gov / Blast.cgi. The “percent sequence identity” between a reference nucleic acid sequence and a query nucleic acid sequence (i.e., the nucleic acid sequence being analyzed to determine whether it is within a particular percent sequence identity with the reference nucleic acid sequence) is determined by optimally aligning the sequences using the Needleman-Wunsch alignment algorithm (with match / mismatch scores of 2,-3, a gap existence penalty of 5, and a gap extension penalty of 2) and comparing the aligned nucleic acids. The number of exact match-es divided by the total number of nucleotides in the alignment (which corresponds with the number of nucleotides in the reference sequence plus any gaps in the reference sequence when aligned with the query sequence) is determined and expressed as a percentage. This is the percent sequence identity between the query nucleic acid sequence and the reference nucleic acid sequence (i.e., percent sequence identity = (# of exact matches) / (total # of nucleotides in the alignment)* 100). An alignment using the Needleman-Wunsch alignment algorithm (with match / mismatch scores of 2,- 3, a gap existence penalty of 5, and a gap ex-tension penalty of 2) can be generated using the “Global Align” BLAST program available at https: / / blast.ncbi.nlm.nih.gov / Blast.cgi.EXAMPLES

[0164] The following examples are provided to further illustrate some embodiments of the present disclosure, but are not intended to limit the scope of the disclosure; it will be understood by their exemplary nature that other procedures, methodologies, or techniques known to those skilled in the art may alternatively be used.Example 1: Screening variant RNA tails for improved gene insertion efficiency

[0165] Template RNA molecules used in TPRT gene insertion systems currently utilize a tail configuration comprising a 4 nucleotide complementarity region to the genomic priming source, with a 22 nucleotide homopolymer of adenosine nucleotides. A screen was performed to identifytail region sequences that improve stability of the template mRNA used in a TPRT gene insertion system.

[0166] The test construct comprised of an RNA with the following orientation: 5’ module - promoter - 5’UTR - reporter gene - 3’UTR+pAS - 3’ module - TAIL, where the 5’ module was HDV-gRZ (SEQ ID NO: 179), the promoter was pCMV (SEQ ID NO: 180), the 5’UTR was Neo3-Kozak (SEQ ID NO: 181), the reporter gene was secNL_intron002_pos03 (SEQ ID NO: 182), the 3’UTR+pAS was miniPAS (SEQ ID NO: 183), and the 3 ’module was GeFo98 (SEQ ID NO: 184); the tails tested are shown in Table 2.

[0167] Table 1 shows the variables to the tail tested in the assay, which included tail length; percent, identity and distribution of non-adenosine nucleotides; use of various stability elements; and uridine modifications.Table 1: Tail region variables

[0168] Combinations of the above variables were chosen using a qualitative discrete choice model, with candidate tails shown in Table 2. Three versions of each RNA was tested: one using unmodified uridine, one using pseudouridine, and one using Nlmethyl-pseudouri dine.

[0169] Table 2: Candidate Tails

[0170] Plasmids were generated encoding the reporter RNA with variant tails downstream of theT7 promoter and upstream of a Bbsi cut site. Plasmids were then digested, purified, and used as a template for in vitro transcription to synthesize RNA. The in vitro transcribed RNA was purified using an oligo dT resin.

[0171] To assess stability of reporter RNAs with variant tails, the RNAs are complexed with Lipofectamine Messenger Max to form lipoplexes that are then transfected into cells in 384-well plates using either forward or reverse transfection. At designated readout times (usually 24, 48, 72 hours etc.) the number of GFP positive cells are counted using the Opera Phenix Plus High- Content Screening System (Revvity). The supematant / cell media is completely removed, and fresh media is applied to the cells. The extracted supernatant is used to assay for secreted Nano Luciferase using the Nano-GLO Luciferase Assay System (Promega). Background signal is subtracted, and data is reported as RLU (Relative Luminescent Units) per transfected cell.Example 2: Methods

[0172] Cloning

[0173] Constructs were cloned into a plasmid vector and linearized downstream of the tail sequence using the type IIS restriction enzyme Bbsi to ensure generation of a uniform 3' end for in vitro transcription (IVT).

[0174] In Vitro Transcription (IVT)

[0175] IVT was performed using the HiCap T7 RNA polymerase transcription system (manufacturer: [Codexis]) in accordance with the manufacturer’s instructions. Reactions werecarried out using one of the following three uridine analog compositions: (i) unmodified uridine (U), (ii) pseudouridine ( ), or (iii) N1 -methylpseudouridine (mlvP).

[0176] Each transcript was generated with only one of the three uridine or uridine analog per reaction (FIG. 2).

[0177] Transfection

[0178] Following transcription, template RNA transcripts were combined with a constant mlvP- modified R2 RNA species at a 6: 1 molar ratio (tempi ate :R2). RNA mixtures were then transfected into cultured mammalian cells using Lipofectamine MessengerMAX transfection reagent (Thermo Fisher Scientific), following the manufacturer’s protocol. Transfections were performed in RPE1, Huh7, and BNL.CL.2 cell lines (FIG. 2).

[0179] Luciferase Readout

[0180] Cell culture supernatants were harvested at 24 hours and 48 hours post-transfection. Secreted luciferase activity was quantified using the Nano-Gio Luciferase Assay System (Promega), per manufacturer’s protocol. For each construct, luminescence values were normalized to those obtained from a positive control transcript containing the R4A22 tail with N1 -methylpseudouridine modification (R4A22-mlvP), yielding a "fold-over-control" value for each time point (FIG. 2).

[0181] The geometric mean of the 24-hour and 48-hour fold-over-control values was calculated for each tail construct in each cell line. This mean value was used as the final performance metric for comparing tail sequence efficiency in the TPRT gene insertion system.Example 3: Performance of tail sequences in the TPRT gene insertion system

[0182] Assessment of the performance of different tail sequences on the template RNA was evaluated using the methods described elsewhere herein.

[0183] Each RNA tail sequence and chemical modification context was evaluated for efficiency in the TPRT gene insertion system in three mammalian cell lines: RPE1, Huh7, and BNL.CL.2. Performance was expressed as the geometric mean of fold-over-positive control (R4A22_ml-psi, SEQ ID NO: 1 transcribed using N1 -methyl pseudouridine) across the three cell lines.

[0184] Each data point represents a unique tail element. The shape of each point denotes the identity of any non-adenosine (non-A) nucleotide incorporated within the tail (circle = U, square = G, diamond = C, triangle = none). The fill color of each point indicates the fraction of non-A substitution (0.0, 0.1, 0.2, or 0.3). The intervening stability element between the 3' module and the adenine-rich region (if applicable) is denoted in the name on the X-axis (FIGs. 3-11, 13-14, and 16-25)

[0185] Rotated kernel density plots was used to denote the distributions of performances of tails containing intervening U, G, or C nucleotides (or no intervening non-A nucleotides) when transcribed using N1 -methyl pseudouridine, pseudouridine, or uridine (FIG. 12), of tails containing intervening U, G, or C nucleotides (or no intervening non-A nucleotides) and various intervening stability elements (eK5, K3, K4, none, and other K-element) between the 3’ module and the adenine-rich region (FIG. 15), of tails without intervening stability elements containing intervening U, G, or C nucleotides (or no intervening non-A nucleotides) when transcribed using N1 -methyl pseudouridine, pseudouridine, or uridine (FIG. 26), and of tails without intervening stability elements containing intervening U, G, or C nucleotides (or no intervening non-A nucleotides) and containing different fraction of intervening non-A nucleotides (0.1, 0.2, 0.3, and 0) when transcribed using N1 -methyl pseudouridine, pseudouridine, or uridine (FIG. 27).

[0186] Performance of template RNA tails are depicted in FIG. 3 and the top performing tails are detailed in Table 3.Table 3: Top performing tails aggregate across cell lines

[0187] Performance of template RNA tails separated by cell lines are depicted in FIG. 4, the top performing tails in each cell line are detailed in Table 4-6.Table 4: Top performing tails in RPE1Table 5: Top performing tails in Huh7Table 6: Top performing tails in BNL.CL.2

[0188] Performance of template RNA tails separated by uridine modifications are depicted in FIG. 5. FIG. 6 depict the performance of template RNA tails with N1 -methyl pseudouridine ranked. FIG. 8 depict the performance of template RNA tails with pseudouridine ranked. FIG.10 depict the performance of template RNA tails with uridine ranked. The top performing tails separated by uridine modifications (N1 -methyl pseudouridine and pseudouridine) are detailed inTable 7-8Table 7: Top performing tails with Nl-methyl pseudouridine aggregated across cell linesTable 8: Top performing tails with pseudouridine aggregated across cell lines

[0189] Performance of template RNA tails with N1 -methyl pseudouridine separated by cell line are depicted in FIG. 7. The top performing tails with Nl-methyl pseudouridine separated by cell line are detailed in Table 9-11.Table 9: Top performing tails with Nl-methyl pseudouridine in RPE1Table 10: Top performing tails with Nl-methyl pseudouridine in Huh7Table 11: Top performing tails with Nl-methyl pseudouridine in BNL.CL.2

[0190] Performance of template RNA tails with pseudouridine separated by cell line are depicted in FIG. 9. The top performing tails with pseudouridine separated by cell line are detailed inTable 12-14Table 12: Top performing tails with pseudouridine in RPE1Table 13: Top performing tails with pseudouridine in Huh7Table 14: Top performing tails with pseudouridine in BNL.CL.2

[0191] Performance of template RNA tails with uridine separated by cell line are depicted in FIG. 11

[0192] Performance of template RNA tails with pseudouridine or N1 -methyl pseudouridine are depicted in FIG. 13. Performance of template RNA tails with pseudouridine or N1 -methyl pseudouridine separated by cell line are depicted in FIG. 14. The top performing tails with pseudouridine or Nl-methyl pseudouridine are detailed in Table 15. The top performing tails with pseudouridine or Nl-methyl pseudouridine separated by cell line are detailed in Table 16- 18Table 15: Top performing tails with pseudouridine or Nl-methyl pseudouridine aggregate across cell linesTable 16: Top performing tails with pseudouridine or Nl-methyl pseudouridine in RPE1Table 17: Top performing tails with pseudouridine or Nl-methyl pseudouridine in Huh7Table 18: Top performing tails with pseudouridine or Nl-methyl pseudouridine in BNL.CL.2

[0193] Performance of template RNA tails without intervening stability elements between the 3' module and adenine-rich region are depicted in FIG. 16. Performance of template RNA tails without intervening stability elements between the 3' module and adenine-rich region separated by cell lines are depicted in FIG. 18. The top performing tails without intervening stability elements are detailed in Table 19. The top performing tails without intervening stability elements separated by cell line are detailed in Table 20-22.Table 19: Top performing tails without intervening stability elements aggregate across cell linesTable 20: Top performing tails without intervening stability elements in RPE1Table 21: Top performing tails without intervening stability elements in Huh7Table 22: Top performing tails without intervening stability elements in BNL.CL.2

[0194] Performance of template RNA tails with pseudouridine or N1 -methyl pseudouridine and without intervening stability elements between the 3' module and adenine-rich region are depicted in FIG. 17. Performance of template RNA tails with pseudouridine or N1 -methyl pseudouridine and without intervening stability elements between the 3' module and adenine-rich region separated by cell lines are depicted in FIG. 19. The top performing tails with pseudouridine or N1 -methyl pseudouridine and without intervening stability elements are detailed in Table 23. The top performing tails with pseudouridine or Nl-methyl pseudouridine and without intervening stability elements separated by cell line are detailed in Table 24-26.Table 23: Top performing tails with pseudouridine or Nl-methyl pseudouridine and without intervening stability elements aggregated across cell linesTable 24: Top performing tails with pseudouridine or Nl-methyl pseudouridine and without intervening stability elements in RPE1Table 25: Top performing tails with pseudouridine or Nl-methyl pseudouridine and without intervening stability elements in Huh7Table 26: Top performing tails with pseudouridine or Nl-methyl pseudouridine and without intervening stability elements in BNL.CL.2

[0195] Performance of template RNA tails with Nl-methyl pseudouridine and without intervening stability elements between the 3' module and adenine-rich region are depicted in FIG. 20. Performance of template RNA tails with Nl-methyl pseudouridine and without intervening stability elements between the 3' module and adenine-rich region separated by cell lines are depicted in FIG. 21. The top performing tails with Nl-methyl pseudouridine and without intervening stability elements are detailed in Table 27. The top performing tails with Nl- methyl pseudouridine and without intervening stability elements separated by cell line are detailed in Table 28-30.Table 27: Top performing tails with Nl-methyl pseudouridine and without intervening stability elements aggregate across cell linesTable 28: Top performing tails with Nl-methyl pseudouridine and without intervening stability elements in RPE1Table 29: Top performing tails with Nl-methyl pseudouridine and without intervening stability elements in Huh7Table 30: Top performing tails with Nl-methyl pseudouridine and without intervening stability elements in BNL.CL.2

[0196] Performance of template RNA tails with pseudouridine and without intervening stability elements between the 3' module and adenine-rich region are depicted in FIG. 22. Performance of template RNA tails with pseudouridine and without intervening stability elements between the 3' module and adenine-rich region separated by cell lines are depicted in FIG. 23. The top performing tails with pseudouridine and without intervening stability elements are detailed in Table 31. The top performing tails with pseudouridine and without intervening stability elements separated by cell line are detailed in Tables 32-34.Table 31: Top performing tails with pseudouridine and without intervening stability elements aggregated across cell linesTable 32: Top performing tails with pseudouridine and without intervening stability elements inRPE1Table 33: Top performing tails with pseudouridine and without intervening stability elements inHuh7Table 34: Top performing tails with pseudouridine and without intervening stability elements inBNL.CL.2

[0197] Performance of template RNA tails with uridine and without intervening stability elements between the 3' module and adenine-rich region are depicted in FIG. 24. Performance of template RNA tails with uridine and without intervening stability elements between the 3' module and adenine-rich region separated by cell lines are depicted in FIG. 25.

[0198] Table 35 reports data of performance of tails tested.Table 35: Performance of tailsExample 4: Calculation of fraction of non-adenine nucleotides

[0199] Calculations of fraction of non-adenine nucleotides (e.g., calculated fraction of non- adenine nts in Table 43) were calculated based on the number of non-adenine nucleobases in the adenine-rich region without the tail end over the length of the adenine-rich region without the tail end. The tail end is indicated as the terminal three nucleotides of the adenine rich region (FIG. 1).

[0200] For example, the tail region comprising having a sequence corresponding to SEQ ID NO: 37 comprises an adenine-rich region having a sequence corresponding to SEQ ID NO: 174. The calculated fraction of non-adenine nucleotides for the tail region is 0.31 and is calculated by the number of non-adenine nucleobases in the adenine-rich region without the tail end (9) over the length of the adenine-rich region without the tail end (29).Sequence Listing

Claims

CLAIMSWHAT IS CLAIMED IS:

1. A nucleic acid construct comprising (a) a 5’ region, (b) a template sequence, and (c) a 3’ region comprising a tail region, wherein: the tail region comprises (i) a four-nucleotide complementarity region that is complementary to a region of an rRNA gene, and (ii) an adenine-rich region having at least 19 nucleotides; the adenine-rich region is located 3’ of the four-nucleotide complementarity region in the nucleic acid construct; at least 60% of nucleobases in the adenine-rich region are adenines; and at least 10% of nucleobases in the adenine-rich region are pyrimidine nucleobases.

2. The nucleic acid construct of claim 1, wherein the adenine-rich region comprises a sequence having at least 80%, at least 90%, or at least 95% sequence identity to any one of SEQ ID NOs: 174, 173, 143, 171, 130, 135, 165, 170, 139, 137, 177, 129, 178, or 167.

3. The nucleic acid construct of claim 1 or 2, wherein the adenine-rich region comprises any one of SEQ ID NOs: 174, 173, 143, 171, 130, 135, 165, 170, 139, 137, 177, 129, 178, or 167.

4. The nucleic acid construct of claim 1, wherein at least 20% of nucleobases in the adenine-rich region are pyrimidine nucleobases.

5. The nucleic acid construct of claim 4, wherein the adenine-rich region comprises a sequence having at least 80%, at least 90%, or at least 95% sequence identity to any one of SEQ ID NOs: 174, 173, 143, 171, 135, 170, 139, 137, 177, or 178.

6. The nucleic acid construct of claim 4 or 5, wherein the adenine-rich region comprises any one of SEQ ID NOs: 174, 173, 143, 171, 135, 170, 139, 137, 177, or 178.

7. The nucleic acid construct of claim 1 or 4, wherein at least 30% of nucleobases in the adenine-rich region are pyrimidine nucleobases.

8. The nucleic acid construct of claim 7, wherein the adenine-rich region comprises a sequence having at least 80%, at least 90%, or at least 95% sequence identity to any one of SEQ ID NOs: 174, 173, 143, 139, 177, or 178.

9. The nucleic acid construct of claim 7 or 8, wherein the adenine-rich region comprises any one of SEQ ID NOs: 174, 173, 143, 139, 177, or 178.

10. The nucleic acid construct of any one of claims 1, 4, or 7, wherein the pyrimidine nucleobases comprise uracil.

11. The nucleic acid construct of claim 10, wherein all of the pyrimidine nucleobases of the adenine-rich region are uracils.

12. The nucleic acid construct of any one of claims 1, 4, or 7, wherein the pyrimidine nucleobases comprise cytosine.

13. The nucleic acid construct of claim 12, wherein all of the pyrimidine nucleobases of the adenine-rich region are cytosines.

14. The nucleic acid construct of any one of claims 1-13, wherein the adenine-rich region comprises, at its 3’ end, a tail end sequence, wherein the tail end sequence is selected from the group consisting of: AAA, AAX, and AXX; wherein X is C, G, or U.

15. The nucleic acid construct of any one of claims 1-14, wherein there is no additional sequence that is 3’ to the adenine-rich region.

16. The nucleic acid construct of any one of claims 1-13, wherein the nucleic acid construct further comprises a tail end sequence that is located 3’ of the adenine rich region, wherein the tail end sequence is selected from the group consisting of: AAA, AAX, and AXX; wherein X is C, G, or U.

17. The nucleic acid construct of claim 16, wherein there is no additional sequence that is 3’ to the tail end sequence.

18. The nucleic acid construct of any one of claims 1-17, wherein there is no intervening sequence between the four-nucleotide complementarity region and the adenine-rich region.

19. The nucleic acid construct of any one of claims 1-18, wherein the nucleic acid construct comprises a sequence having at least 80%, at least 90%, or at least 95% sequence identity to any one of SEQ ID NOs: 37, 58, 48, 35, 73, 68, 43, 51, 32, 81, 16, 69, 44, 42, or 56.

20. The nucleic acid construct of any one of claims 1-19, wherein the nucleic acid construct comprises any one of SEQ ID NOs: 37, 58, 48, 35, 73, 68, 43, 51, 32, 81, 16, 69, 44, 42, or 56.

21. The nucleic acid construct of any one of claims 1-20, wherein the nucleic acid construct comprises RNA.

22. The nucleic acid construct of any one of claims 1-21, wherein the nucleic acid construct comprises a modified uridine.

23. The nucleic acid construct of any one of claims 1-22, wherein the modified uridine is a pseudouridine or a N1 -methylpseudouridine.

24. The nucleic acid construct of any one of claims 1-23, wherein the modified uridine is a pseudouridine.

25. The nucleic acid construct of any one of claims 1-23, wherein the modified uridine is a N1 -methylpseudouridine.

26. The nucleic acid construct of any one of claims 1-25, wherein the nucleic acid construct is a single-stranded nucleic acid construct.

27. The nucleic acid construct of any one of claims 1-26, wherein the adenine-rich region comprises at least 22 nucleotides.

28. The nucleic acid construct of any one of claims 1-27, wherein the adenine-rich region comprises at least 29 nucleotides.

29. The nucleic acid construct of any one of claims 1-28, wherein the complementarity region is complementary to a region in a 28S rRNA gene.

30. The nucleic acid construct of any one of claims 1-29, wherein the complementarity region has a sequence of UAGC or TAGC.

31. The nucleic acid construct of any one of claims 1-30, wherein the nucleic acid construct has increased stability as compared to an otherwise identical nucleic acid construct having a tail region consisting of SEQ ID NO: 1.

32. The nucleic acid construct of any one of claims 1-31, wherein the nucleic acid construct has increased exonuclease resistance as compared to an otherwise identical nucleic acid construct having a tail region consisting of SEQ ID NO: 1.

33. The nucleic acid construct of any one of claims 1-32, wherein the nucleic acid construct is more efficiently inserted into a genome by target primed reverse transcription as compared to an otherwise identical nucleic acid construct having a tail region consisting of SEQ ID NO: 1.

34. The nucleic acid construct of any one of claims 1-33, wherein the template sequence encodes a gene.

35. The nucleic acid construct of claim 34, wherein the gene is a therapeutic gene or a diagnostic gene.

36. The nucleic acid construct of any one of claims 1-35, wherein the 5’ region comprises a 5 ’-module, a promoter, and a 5’ UTR.

37. The nucleic acid construct of claim 36, wherein the 5’ module comprises a 5’-flanking self-cleaving ribozyme motif.

38. The nucleic acid construct of claim 37, wherein the 5’ module comprises a 5’ UTR derived from a same R2 retroelement as the 5 ’-flanking self-cleaving ribozyme motif.

39. The nucleic acid construct of any one of claims 1-38, wherein the 3’ region further comprises a 3’ UTR and a 3’ module.

40. A system for modifying a target nucleic acid in a cell of a subject, the system comprising:(a) the nucleic acid construct of any one of claims 1-39; and(b) an R2 retroelement polypeptide or a second nucleic acid encoding the R2 retroelement polypeptide; wherein the R2 retroelement polypeptide inserts the template sequence or a reverse complement thereof into the target nucleic acid.

41. A cell comprising the nucleic acid construct of any one of claims 1-39 or the system of claim 40.

42. The cell of claim 41, wherein the cell is a liver cell.

43. The cell of claim 41 or 42, wherein the cell is a cancer cell.

44. The cell of claim 41, wherein the cell is a retinal cell.

45. The cell of any one of claims 41-44, wherein the cell comprises an rRNA gene.

46. The cell of any one of claims 41-45, wherein the nucleic acid construct is not endogenous to the cell.

47. The cell of any one of claims 41-46, wherein the cell is a mammalian cell.

48. The cell of any one of claims 41-47, wherein the cell is a human cell.

49. A method of modifying a target nucleic acid in a cell, the method comprising providing the cell with:(a) the nucleic acid construct of any one of any one of claims 1-39; and(b) an R2 retroelement polypeptide or a second nucleic acid encoding the R2 retroelement polypeptide;wherein the R2 retroelement polypeptide inserts the template sequence of the nucleic acid construct, or a reverse complement thereof, into the target nucleic acid.

50. The method of claim 49, wherein the target nucleic acid comprises an rRNA gene.

51. The method of claim 49 or 50, wherein the cell is derived from a human subject.

52. A nucleic acid construct comprising (a) a 5’ region, (b) a template sequence, and (c) a 3’ region comprising a tail region, wherein the tail region comprises a four-nucleotide complementarity region and an adenine-rich region of at least 22 nucleotides, wherein at least 70% of the nucleobases in the adenine-rich region are adenine.

53. The nucleic acid construct of claim 52, wherein the tail region further comprises a stability element.

54. The nucleic acid construct of claim 53, wherein the stability element is 5’ to the adenine-rich region.

55. The nucleic acid construct of claim 53, wherein the stability element is 3’ to the adenine-rich region.

56. The nucleic acid construct of any one of claims 53-55, wherein the four- nucleotide complementarity region is 5’ to the adenine-rich region and / or the stability element.

57. The nucleic acid construct of any one of claims 52-56, wherein the complementarity region is complementary to a region in a 28S rRNA gene.

58. The nucleic acid construct of any one of claims 52-57, wherein the complementarity region has a sequence of UAGC or TAGC.

59. The nucleic acid construct of any one of claims 52-58, wherein the adenine-rich region is between 22 and 41 nucleotides in length.

60. The nucleic acid construct of any one of claims 52-59, wherein the adenine-rich region is 22, 32, or 41 nucleotides in length.

61. The nucleic acid construct of any one of claims 52-60, wherein the adenine-rich region comprises one or more non-adenine nucleobases.

62. The nucleic acid construct of claim 61, wherein the adenine-rich region has no more than 10%, 15%, or 20% or 25% non-adenine nucleobases.

63. The nucleic acid construct of claim 61 or 62, wherein the one or more non- adenine nucleobases comprise any one of cytosine, guanine, uridine, or a modified uridine.

64. The nucleic acid construct of claim 63, wherein the modified uridine is pseudouridine or N1 -methylpseudouridine.

65. The nucleic acid construct of any one of claims 52-64, wherein the adenine-rich region comprises any one of SEQ ID NOs: 118-178.

66. The nucleic acid construct of any one of claims 52-65, wherein the terminal three nucleotides of the adenine-rich region each comprise adenine.

67. The nucleic acid construct of any one of claims 52-65, wherein the terminal nucleotide of the adenine-rich region does not comprise adenine.

68. The nucleic acid construct of claim 67, wherein the two most terminal nucleotides of the adenine-rich region do not comprise adenine.

69. The nucleic acid construct of any one of claims 53-68, wherein the stability element is selected from K3, K4, eK5, MALATl mm, MALATl hs, MALATl hs-core, PAN KSHV, dENE os, HSL, U7, K9, K16, KEMCV, or KBDV1.

70. The nucleic acid construct of any one of claims 53-69, wherein the stability element comprises any one of SEQ ID NOs: 90-117.

71. The nucleic acid construct of claim 52, wherein the tail region comprises any one of SEQ ID NOs: 2-89.

72. The nucleic acid construct of any one of claims 52-71, wherein the nucleic acid construct has increased stability as compared to an otherwise identical nucleic acid construct having a tail region consisting of SEQ ID NO: 1.

73. The nucleic acid construct of any one of claims 52-72, wherein the nucleic acid construct has increased exonuclease resistance as compared to an otherwise identical nucleic acid construct having a tail region consisting of SEQ ID NO: 1.

74. The nucleic acid construct of any one of claims 52-73, wherein the nucleic acid construct is more efficiently inserted into a genome by target primed reverse transcription as compared to an otherwise identical nucleic acid construct having a tail region consisting of SEQ ID NO: 1.

75. The nucleic acid construct of any one of claims 52-74, wherein the template sequence encodes a gene.

76. The nucleic acid construct of claim 75, wherein the gene is a therapeutic gene or a diagnostic gene.

77. The nucleic acid construct of any one of claims 52-76, wherein the 5’ region comprises a 5 ’-module, a promoter, and a 5’ UTR.

78. The nucleic acid construct of claim 77, wherein the 5’ module comprises a 5’- flanking self-cleaving ribozyme motif.

79. The nucleic acid construct of claim 77 or 78, wherein the 5’ module comprises a 5’ UTR derived from a same R2 retroelement as the 5 ’-flanking self-cleaving ribozyme motif.

80. The nucleic acid construct of any one of claims 52-79, wherein the 3’ region further comprises a 3’ UTR and a 3’ module.

81. A system for modifying a target nucleic acid in a cell of a subject, the system comprising:(a) the nucleic acid construct of any one of claims 52-80; and(b) an R2 retroelement polypeptide or a second nucleic acid encoding the R2 retroelement polypeptide; wherein the R2 retroelement polypeptide inserts the template sequence or a reverse complement thereof into the target nucleic acid.

82. A method of modifying a target nucleic acid in a cell of a subject, the method comprising providing the cell with:(a) the nucleic acid construct of any one of any one of claims 52-80; and(b) an R2 retroelement polypeptide or a second nucleic acid encoding the R2 retroelement polypeptide; wherein the R2 retroelement polypeptide inserts the template sequence of the nucleic acid construct, or a reverse complement thereof, into the target nucleic acid.

83. A nucleic acid construct comprising (a) a 5’ region, (b) a template sequence, and (c) a 3’ region comprising a tail region, wherein the tail region comprises (i) a four-nucleotide complementarity region that is complementary to a region of an rRNA gene, and (ii) an adenine- rich region having at least 19 nucleotides, wherein the adenine-rich region is located 3’ of the four-nucleotide complementarity region in the nucleic acid construct, wherein at least 60% of the nucleobases in the adenine-rich region are adenines, and wherein at least one nucleobase in the adenine-rich region is a non-adenine nucleobase.

84. A nucleic acid construct comprising (a) a 5’ region, (b) a template sequence, and(c) a 3’ region comprising a tail region, wherein the tail region comprises (i) a four-nucleotide complementarity region that is complementary to a region of an rRNA gene, and (ii) an adenine- rich region having at least 22 nucleotides, wherein the adenine-rich region is located 3’ of the four-nucleotide complementarity region in the nucleic acid construct, wherein at least 50% of the nucleobases in the adenine-rich region are adenines, and wherein at least one nucleobase in the adenine-rich region is a non-adenine nucleobase.

85. The nucleic acid construct of claim 83 or 84, wherein the non-adenine nucleobase is a pyrimidine nucleobase.

86. The nucleic acid construct of any one of claims 83-85, wherein the non-adenine nucleobase is a uracil.

87. The nucleic acid construct of any one of claims 83-85, wherein the non-adenine nucleobase is a cytosine.

88. The nucleic acid construct of any one of claims 83-87, wherein the nucleic acid construct comprises a sequence having at least 80%, at least 90%, or at least 95% sequence identity to any one of SEQ ID NOs: 118-178.

89. The nucleic acid construct of any one of claims 83-88, wherein the nucleic acid construct comprises any one of SEQ ID NOs: 118-178.

90. The nucleic acid construct of any one of claims 83-89, wherein the nucleic acid construct comprises a sequence having at least 80%, at least 90%, or at least 95% sequence identity to any one of SEQ ID NOs: 174, 173, 143, 171, 118, 120-122, 125, 126, 129, 130, 135, 137, 139, 165, 167, 170, 177, and 178.

91. The nucleic acid construct of any one of claims 83-90, wherein the nucleic acid construct comprises any one of SEQ ID NOs: 174, 173, 143, 171, 118, 120-122, 125, 126, 129, 130, 135, 137, 139, 165, 167, 170, 177, and 178.

92. The nucleic acid construct of any one of claims 83-91, wherein the nucleic acid construct comprises a sequence having at least 80%, at least 90%, or at least 95% sequence identity to any one of SEQ ID NOs: 2-89.

93. The nucleic acid construct of any one of claims 83-92, wherein the nucleic acid construct comprises any one of SEQ ID NOs: 2-89.

94. The nucleic acid construct of any one of claims 83 or 84, wherein at least 10% of the nucleobases in the adenine-rich region are non-adenine nucleobases.

95. The nucleic acid construct of claim 94, wherein the at least 10% of the nucleobases in the adenine-rich region are pyrimidine nucleobases.

96. The nucleic acid construct of claim 95, wherein the pyrimidine nucleobases comprise uracil.

97. The nucleic acid construct of claim 95 or 96, wherein all of the pyrimidine nucleobases of the adenine-rich region are uracils.

98. The nucleic acid construct of claim 95, wherein the pyrimidine nucleobases comprise cytosine.

99. The nucleic acid construct of claim 95 or 96, wherein all of the pyrimidine nucleobases of the adenine-rich region are cytosines.

100. The nucleic acid construct of claim 94, wherein the nucleic acid construct comprises a sequence having at least 80%, at least 90%, or at least 95% sequence identity to any one of SEQ ID NOs: 128-178.

101. The nucleic acid construct of claim 94 or 100, wherein the nucleic acid construct comprises any one of SEQ ID NOs: 128-178.

102. The nucleic acid construct of any one of claims 83, 84, or 94, wherein at least 20% of the nucleobases in the adenine-rich region are non-adenine nucleobases.

103. The nucleic acid construct of claim 102, wherein the at least 20% of the nucleobases in the adenine-rich region are pyrimidine nucleobases.

104. The nucleic acid construct of claim 103, wherein the pyrimidine nucleobases comprise uracil.

105. The nucleic acid construct of claim 103 or 104, wherein all of the pyrimidine nucleobases of the adenine-rich region are uracils.

106. The nucleic acid construct of claim 103, wherein the pyrimidine nucleobases comprise cytosine.

107. The nucleic acid construct of claim 103 or 104, wherein all of the pyrimidine nucleobases of the adenine-rich region are cytosines.

108. The nucleic acid construct of claim 102, wherein the nucleic acid construct comprises a sequence having at least 80%, at least 90%, or at least 95% sequence identity to any one of SEQ ID NOs: 133-143, 151-162, or 168-178.

109. The nucleic acid construct of claim 102 or 108, wherein the nucleic acid construct comprises any one of SEQ ID NOs: 133-143, 151-162, or 168-178.

110. The nucleic acid construct of any one of claims 83, 84, 94, or 102, wherein at least 30% of the nucleobases in the adenine-rich region are non-adenine nucleobases.

111. The nucleic acid construct of claim 110, wherein the at least 30% of the nucleobases in the adenine-rich region are pyrimidine nucleobases.

112. The nucleic acid construct of claim 111, wherein the pyrimidine nucleobases comprise uracil.

113. The nucleic acid construct of claim 111 or 112, wherein all of the pyrimidine nucleobases of the adenine-rich region are uracils.

114. The nucleic acid construct of claim 111, wherein the pyrimidine nucleobases comprise cytosine.

115. The nucleic acid construct of claim 111 or 112, wherein all of the pyrimidine nucleobases of the adenine-rich region are cytosines.

116. The nucleic acid construct of claim 110, wherein the nucleic acid construct comprises a sequence having at least 80%, at least 90%, or at least 95% sequence identity to any one of SEQ ID NOs: 139-143, 157-162, or 173-178.

117. The nucleic acid construct of claim 110 or 116, wherein the nucleic acid construct comprises any one of SEQ ID NOs: 139-143, 157-162, or 173-178.

118. The nucleic acid construct of any one of claims 83-117, wherein the adenine-rich region comprises, at its 3’ end, a tail end sequence, wherein the tail end sequence is selected from the group consisting of: AAA, AAX, and AXX, wherein X is C, G, or U.

119. The nucleic acid construct of any one of claims 83-118, wherein there is no additional sequence that is 3’ to the adenine-rich region.

120. The nucleic acid construct of any one of claims 83-117, wherein the nucleic acid construct further comprises a tail end sequence that is located 3’ of the adenine rich region, wherein the tail end sequence is selected from the group consisting of: AAA, AAX, and AXX, wherein X is C, G, or U.

121. The nucleic acid construct of claim 120, wherein there is no additional sequence that is 3’ to the tail end sequence.

122. The nucleic acid construct of any one of claims 83-121, wherein there is no intervening sequence between the four-nucleotide complementarity region and the adenine-rich region.

123. The nucleic acid construct of any one of claims 83-122, wherein the nucleic acid construct comprises RNA.

124. The nucleic acid construct of any one of claims 83-123, wherein the nucleic acid construct comprises a modified uridine.

125. The nucleic acid construct of claim 124, wherein the modified uridine is a pseudouridine or a N1 -methylpseudouridine.

126. The nucleic acid construct of claim 124 or 125, wherein the modified uridine is a pseudouridine.

127. The nucleic acid construct of claim 124 or 125, wherein the modified uridine is a N 1 -methylpseudouridine.

128. The nucleic acid construct of any one of claims 83-127, wherein the tail region further comprises a stability element.

129. The nucleic acid construct of any one of claims 83-128, wherein the complementarity region is complementary to a region in a 28S rRNA gene.

130. The nucleic acid construct of any one of claims 83-129, wherein the complementarity region has a sequence of UAGC or TAGC.

131. The nucleic acid construct of any one of claims 83-130, wherein the nucleic acid construct has increased stability as compared to an otherwise identical nucleic acid construct having a tail region consisting of SEQ ID NO: 1.

132. The nucleic acid construct of any one of claims 83-131, wherein the nucleic acid construct has increased exonuclease resistance as compared to an otherwise identical nucleic acid construct having a tail region consisting of SEQ ID NO: 1.

133. The nucleic acid construct of any one of claims 83-132, wherein the nucleic acid construct is more efficiently inserted into a genome by target primed reverse transcription as compared to an otherwise identical nucleic acid construct having a tail region consisting of SEQ ID NO: 1.

134. The nucleic acid construct of any one of claims 83-133, wherein the template sequence encodes a gene.

135. The nucleic acid construct of claim 134, wherein the gene is a therapeutic gene or a diagnostic gene.

136. The nucleic acid construct of any one of claims 83-135, wherein the 5’ region comprises a 5 ’-module, a promoter, and a 5’ UTR.

137. The nucleic acid construct of claim 136, wherein the 5’ module comprises a 5’- flanking self-cleaving ribozyme motif.

138. The nucleic acid construct of claim 137, wherein the 5’ module comprises a 5’ UTR derived from a same R2 retroelement as the 5 ’-flanking self-cleaving ribozyme motif.

139. The nucleic acid construct of any one of claims 83-138, wherein the 3’ region further comprises a 3’ UTR and a 3’ module.

140. A system for modifying a target nucleic acid in a cell of a subject, the system comprising:(a) the nucleic acid construct of any one of claims 83-139; and(b) an R2 retroelement polypeptide or a second nucleic acid encoding the R2 retroelement polypeptide; wherein the R2 retroelement polypeptide inserts the template sequence or a reverse complement thereof into the target nucleic acid.

141. A cell comprising the nucleic acid construct of any one of claims 83-139 or the system of claim 140.

142. The cell of claim 141, wherein the cell is a liver cell.

143. The cell of claim 141 or 142, wherein the cell is a cancer cell.

144. The cell of claim 141, wherein the cell is a retinal cell.

145. The cell of any one of claims 141-144, wherein the cell comprises an rRNA gene.

146. The cell of any one of claims 141-145, wherein the nucleic acid construct is not endogenous to the cell.

147. The cell of any one of claims 141-146, wherein the cell is a mammalian cell.

148. The cell of any one of claims 141-147, wherein the cell is a human cell.

149. A method of modifying a target nucleic acid in a cell, the method comprising providing the cell with:(a) the nucleic acid construct of any one of claims 83-139; and(b) an R2 retroelement polypeptide or a second nucleic acid encoding the R2 retroelement polypeptide; wherein the R2 retroelement polypeptide inserts the template sequence of the nucleic acid construct, or a reverse complement thereof, into the target nucleic acid.

150. The method of claim 149, wherein the target nucleic acid comprises an rRNA gene.

151. The method of claim 149 or 150, wherein the cell is derived from a human subject.

Citation Information

Patent Citations

  • Site-Specific Gene Modifications

    US20230340523A1

  • Methods and compositions for prime editing nucleotide sequences

    US20230383289A1

  • Genome insertions in cells

    US20250230471A1

  • Multicomponent systems for site-specific genome modifications

    WO2023215727A2